In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the rapid iteration of LLMs, objective, quantitative, and comprehensive evaluation of their capabilities has become a critical link in advancing technological development. Currently, the mainstream static benchmark dataset-based evaluation methods face challenges such as the diversity of task types, inconsistent evaluation criteria, and fragmentation of data and processing workflows, making it difficult to efficiently conduct cross-domain and large-scale model evaluation. To address the aforementioned issues, this paper proposes and open-sources OpenCompass, a one-stop, scalable, and high-concurrency-supported general-purpose LLM evaluation platform. Adhering to the design philosophy of modularization and component decoupling, the platform boasts three core advantages: high compatibility, flexibility, and high concurrency. The core architecture of OpenCompass comprises five key components: the Configuration System, Task Partitioning Module, Execution and Scheduling Module, Task Execution Unit, and Result Visualization Module. Its workflow provides rule-based, LLM-as-a-Judge, and cascaded evaluators to adapt to the requirements of different task scenarios. Supporting mainstream benchmark datasets across multiple domains, including knowledge, reasoning, computation, science, language, code, etc., the platform offers a unified and efficient LLM evaluation tool for both academia and industry, facilitating the accurate identification of strengths and weaknesses of LLMs as well as their subsequent optimization.
This study aimed to compare the diagnostic value of [68Ga]Ga-DOTA-FGFR1 and [18F]FDG PET/CT in the evaluation of lung cancer patients. A prospective study was conducted between March 2023 and July 2023. Patients with high clinical suspicion of lung cancer were recruited. Each participant underwent PET/CT scanning using [68Ga]Ga-DOTA-FGFR1 and [18F]FDG within 6 days. Histopathology and clinical follow-up results serve as reference criteria for final diagnosis. We used a paired samples t-test or a Wilcoxon signed-rank test to compare the uptake of [68Ga]Ga-DOTA-FGFR1 and [18F]FDG. The diagnostic performance between the two tracers was compared using the McNemar χ² test. A total of 101 participants were included (mean age 63.267 ± 9.344 [range 39–86 years]). In benign lung lesions, [68Ga]Ga-DOTA-FGFR1 had lower TBR and SUVmax than [18F]FDG (2.924 vs. 5.705, P < 0.001;1.395 vs. 4.014, P < 0.001). The TBR of [68Ga]Ga-DOTA-FGFR1 in benign lymph nodes was also lower than [18F]FDG (0.880 vs. 1.25, P < 0.001). [68Ga]Ga-DOTA-FGFR1 had a higher diagnostic specificity for primary tumors than [18F]FDG (52
Code large language models (Code LLMs) have demonstrated remarkable performance in code generation. Nonetheless, most existing works focus on boosting code LLMs from the perspective of programming capabilities, while their natural language capabilities receive less attention. To fill this gap, we thus propose a novel framework, comprising two modules: AttentionExtractor, which is responsible for extracting key phrases from the user's natural language requirements, and AttentionCoder, which leverages these extracted phrases to generate target code to solve the requirement. This framework pioneers an innovative idea by seamlessly integrating code LLMs with traditional natural language processing tools. To validate the effectiveness of the framework, we craft a new code generation benchmark, called MultiNL-H, covering five natural languages. Extensive experimental results demonstrate the effectiveness of our proposed framework.
In vertical federated learning (VFL), commercial entities collaboratively train a model while preserving data privacy. However, a malicious participant's poisoning attack may degrade the performance of this collaborative model. The main challenge in achieving the poisoning attack is the absence of access to the server-side top model, leaving the malicious participant without a clear target model. To address this challenge, we introduce an innovative end-to-end poisoning framework P-GAN. Specifically, the malicious participant initially employs semi-supervised learning to train a surrogate target model. Subsequently, this participant employs a GAN-based method to produce adversarial perturbations to degrade the surrogate target model's performance. Finally, the generator is obtained and tailored for VFL poisoning. Besides, we develop an anomaly detection algorithm based on a deep auto-encoder (DAE), offering a robust defense mechanism to VFL scenarios. Through extensive experiments, we evaluate the efficacy of P-GAN and DAE, and further analyze the factors that influence their performance.
The task of code generation aims to generate code solutions based on given programming problems. Recently, code large language models (code LLMs) have shed new light on this task, owing to their formidable code generation capabilities. While these models are powerful, they seldom focus on further improving the accuracy of library-oriented API invocation. Nonetheless, programmers frequently invoke APIs in routine coding tasks. In this paper, we aim to enhance the proficiency of existing code LLMs regarding API invocation by mimicking analogical learning, which is a critical learning strategy for humans to learn through differences among multiple instances. Motivated by this, we propose a simple yet effective approach, namely DiffCoder, which excels in API invocation by effectively training on the differences (diffs) between analogical code exercises. To assess the API invocation capabilities of code LLMs, we conduct experiments on seven existing benchmarks that focus on mono-library API invocation. Additionally, we construct a new benchmark, namely PanNumEval, to evaluate the performance of multi-library API invocation. Extensive experiments on eight benchmarks demonstrate the impressive performance of DiffCoder. Furthermore, we develop a VSCode plugin for DiffCoder, and the results from twelve invited participants further verify the practicality of DiffCoder.
The impressive performance of large language models (LLMs) on code-related tasks has shown the potential of fully automated software development. In light of this, we introduce a new software engineering task, namely Natural Language to code Repository (NL2Repo). This task aims to generate an entire code repository from its natural language requirements. To address this task, we propose a simple yet effective framework CodeS, which decomposes NL2Repo into multiple sub-tasks by a multi-layer sketch. Specifically, CodeS includes three modules: RepoSketcher, FileSketcher, and SketchFiller. RepoSketcher first generates a repository's directory structure for given requirements; FileSketcher then generates a file sketch for each file in the generated structure; SketchFiller finally fills in the details for each function in the generated file sketch. To rigorously assess CodeS on the NL2Repo task, we carry out evaluations through both automated benchmarking and manual feedback analysis. For benchmark-based evaluation, we craft a repository-oriented benchmark, SketchEval, and design an evaluation metric, SketchBLEU. For feedback-based evaluation, we develop a VSCode plugin for CodeS and engage 30 participants in conducting empirical studies. Extensive experiments prove the effectiveness and practicality of CodeS on the NL2Repo task.
Vertical federated learning (VFL) enables multiple parties to collaboratively train a model while preserving privacy. However, recent studies have raised concerns about the susceptibility of VFL models, including those using logistic regression and neural networks, to feature inference attacks. Meanwhile, the non-differentiable characteristics of decision tree ensembles make conducting such attacks impractical. To address this challenge, we introduce a feature inference attack framework FIA-TE tailored for decision tree ensembles, including gradient boosted decision trees (GBDT) and random forest. Specifically, we distill the knowledge from trees into neural networks by leaf embedding and structure distillation to create a targeted model for the inference attack. We then employ a generative model based on the deconvolutional network for capturing correlation features and reconstructing the target features. Through extensive experiments on table and image data, we evaluate the effectiveness of our framework and provide an analysis of potential influencing factors.
PURPOSE:There is no specific literature on the best implantation position of the Femoral Neck System (FNS) for treating Pauwels type III femoral neck fracture in young adults.METHODS:Use finite-element analysis to compare the mechanical properties of implantation positions: FNS in the central position, FNS in the low position, and FNS in the low position combined with cannulated screw (CS). The CT data of the femur were imported into the mimics20.0 to obtain the three-dimensional model of the femur; imported into geomagic2017 and SolidWorks 2017 for optimizations; models of FNS and CS are built on the basis of the device manuals. Grouping is as follows: FNS group, FNS-LOW group, and FNS-CS group. Assemble and import them into abaques6.14 for load application. The displacement distribution and von Mises Stress value of them were compared.RESULTS:On femoral stability and stress distribution, the FNS-CS group performs best, followed by the FNS-LOW group, and finally FNS group. The FNS-LOW group has an improvement over the FNS group but not by much.CONCLUSION:In operations, when the implantation position of the central guide wire is not at the center of the femoral neck but slightly lower, it is recommended not to adjust the wire repeatedly in pursuit of the center position; for femoral neck fractures that are extremely unstable at the fracture end or require revision, the insertion strategy of FNS in the low position combined with CS can be adopted to obtain better fixation effects.
ObjectiveTo evaluate the clinical effects of the posterior unilateral approach with 270° spinal canal decompression and three-column reconstruction using double titanium mesh cage (TMC) for thoracic and lumbar burst fractures.Materials and methodsFrom May 2013 to May 2018, 27 patients with single-level thoracic and lumbar burst fractures were enrolled. Every patient was followed for at least 18 months. Demographic data, neurologic status, back pain, canal compromise, anterior body compression, operative time, estimated blood loss and surgical-related complications were evaluated. Radiographs were reviewed to assess deformity correction, anterior body height correction, bony fusion and TMC subsidence.ResultsThe average preoperative percentages of canal compromise and anterior body height compression were 58.4% and 50.5%, respectively. All surgeries were successfully completed in one phase, the operative time was 151.5 ± 25.5 min (range: 115–220 min), the estimated blood loss was 590.7 ± 169.9 ml (range: 400–1,000 ml). Neurological function recovery was significantly improved except for 3 grade A patients. The preoperative visual analog scale (VAS) scores for back pain were significantly decreased compared with the values at the last follow-up (P = 0.000). The correct deformity angle was 12.4 ± 4.7° (range: 3.9–23.3°), and the anterior body height recovery was 96.7%. The TMC subsidence at the last follow-up was 1.3 ± 0.7 mm (range: 0.3–3.1 mm). Bony fusion was achieved in all patients.ConclusionThe posterior unilateral approach with 270° spinal canal decompression and three-column reconstruction using double TMC is a clinically feasible, safe and alternative treatment for thoracic and lumbar burst fractures.
As one of the GEN-IV nuclear reactor systems proposed by GEN-IV International Forum, the sodium-cooled fast reactor features the energy generation with fast spectrum and a closed fuel cycle for fuel breeding and actinide management. In order to validate system codes for reactor system analyses and to improve code performance, benchmark activities of experimental transients of fast reactors were performed by worldwide institutions. The Shut-down Heat Removal Tests in Experimental Breeder Reactor II were proposed in the context of the International Atomic Energy Agency. In this work, loss-of-flow tests performed in Experimental Breeder Reactor II are analyzed using system thermal-hydraulic code CATHARE. Evolutions of important reactor parameters such as core power, coolant flow rate, temperatures in core and in the intermediate heat exchanger are predicted and compared against experimental data. The sodium pool is further modeled by using a 3D model in CATHARE-3 and its verification is also conducted in the benchmark. Qualitative agreement is obtained for the reactor parameters predicted by CATHARE. The inherent passive safety characteristics of Experimental Breeder Reactor II in unprotected loss-of-flow transient is demonstrated, and the stratification effects in sodium pool are able to be identified with the 3D model. The results obtained in this work are expectable to provide some valuable insights for future sodium-cooled fast reactor transient accident analyses by coupling system thermal-hydraulic codes and CFD tools.
Objective Explore the difference of oncology outcome of laparotomy and laparoscopy in the new FIGO2018 stage of early cervical squamous cell carcinoma without any high risk pathological factors. Methods The 5-years OS and DFS of cervical squamous cell carcinoma undergoing laparotomy and laparoscopy from 2004 to 2018 were compared by the total study population and propensity score from China. Result There was no difference in 5-year OS between laparotomy (2,478 cases) and laparoscopy (1,504 cases), but the 5-year DFS of laparotomy was higher (92.2 %vs. 90.4%, P=0.022). Cox analysis showed that laparoscopy was not an independent risk factor for the death of cervical squamous cell carcinoma (OS: P=0.598), but it was an independent risk factor for the recurrence/death (HR = 1.468,95% CI 1.131 ~ 1.906, P=0.004). There was no difference in 5-year OS between laparotomy (2,391 cases) and laparoscopy (1,495 cases) after 1:2 PSM, but the 5-year DFS of laparotomy was higher (92.7% vs. 90.8%, P = 0.006), Cox analysis showed that laparoscopy was not an independent risk factor for the death of cervical squamous cell carcinoma (OS: P=0.521), but it was an independent risk factor for the recurrence/death (HR=1.512, 95%CI 1.151~1.971, P=0.002). Conclusion There is no difference in 5-year OS between these groups for early cervical squamous cell carcinoma in new stage of FIGO2018 without any high-risk pathological factors, the 5-year DFS of laparotomy is higher than that of laparoscopy group, and laparoscopy is an independent risk factor for recurrence/death, so laparoscopy has a higher risk of recurrence.
Subsets with "good" properties in finite fields are widely applied in coding and cryptography. In this paper we introduce pseudorandom measures for subsets in finite fields and prove lower bounds for the pseudorandom measure. The pseudorandom properties of support of some Boolean functions are studied and the properties of cyclotomic classes in finite fields have also been discussed.
Vertical federated learning has a great potential of driving a great variety of business cooperation among enterprises in many fields. In machine learning, decision tree ensembles such as gradient boosting decision trees (GBDT) and random forest are widely applied powerful models with high interpretability and modeling efficiency. However, state-of-art framework for decision tree ensembles in vertical federated learning frameworks adapt anonymous features to avoid possible data breaches, makes the interpretability of the model compromised. To address this issue in the inference process, in this paper, we firstly make a problem analysis about the necessity of disclosure meanings of feature to Guest Party in vertical federated learning. We protect data privacy and allow the disclosure of feature meaning by concealing decision paths and adapt a communication-efficient secure computation method for inference outputs. The advantages of Fed-EINI will be demonstrated through both theoretical analysis and extensive numerical results. We improve the interpretability of the model by disclosing the meaning of features while ensuring efficiency and accuracy.
Background: Intrawound treatments have been reported to have favorable efficacy for preventing surgical site infection (SSI); however, the best strategy remains unknown. Objective: The aim of this systematic review and network meta-analysis was to evaluate the efficacy of intrawound treatments to prevent SSI after spine surgery. Study Design: A systematic review and network meta-analysis. Methods: We searched the Cochrane Library, EMbase, PubMed, Chinese Science and Technology Periodical Database (VIP), China National Knowledge Infrastructure (CNKI), and Wanfang Data from the date of inception to March 2, 2020. The randomized controlled trials (RCTs) and cohort studies were identified and extracted by 2 reviewers independently. We performed a traditional pairwise meta-analysis to evaluate overall efficacy of intrawound treatments. Meanwhile, a network metaanalysis was performed to compare and rank the treatment efficacy using frequentist approach. Results: Thirty-three publications (6 RCTs and 27 retrospective cohort studies) were included, involving 22,763 patients. For pairwise meta-analysis, the combined results showed that the intrawound treatment had a significantly lower SSI rate than the control group (CG) (odds ratio [OR] = 0.41; 95% confidence interval [CI], 0.31-0.55). For network meta-analysis, the treatment of vancomycin (VA) (OR = 0.53; 95% CI, 0.39-0.71), povidone-iodine (PI) (OR = 0.10; 95% CI, 0.04 - 0.23), and vancomycin + povidone-iodine (VA+PI) (OR = 0.25; 95% CI, 0.11-0.58) were found to be significantly more efficacious than CG on reduction of SSI rate. PI ranked first on reducing SSI, followed by PI+HP, VA+PI, gentamicin (GM), VA, and hydrogen peroxide (HP); CG ranked last. Limitations: Firstly, only 6 RCTs are included in this systematic review. Retrospective cohort studies tend to exaggerate the real results, although most of them are high-quality according to the Newcastle-Ottawa Quality Assessment Scale (NOQAS). More high-quality RCTs need to be included to obtain convincing conclusions. Secondly, the population of this study involves both adult and pediatric cohorts, patients with tumor, congenital disease, or degenerative disease. There is no subgroup analysis for ages and type of diseases, which might have influence on the overall pooled analysis. Thirdly, we define the application of saline solution and no intrawound treatment as the control group, which might ignore their heterogeneity. Fourthly, follow-up periods are variable and the sample size of HP is small. Finally, additional research is needed to compare the complications of different treatments and the benefits of various dosages. Conclusion: We found that VA and PI show promising results on reducing SSI. PI is recommended as the most efficacious intrawound treatment to prevent SSI after spine surgery.
"Inflammaging" refers to the chronic, low-grade inflammation that characterizes aging. Aging, like obesity, is associated with visceral adiposity and insulin resistance. Adipose tissue macrophages (ATMs) have played a major role in obesity-associated inflammation and insulin resistance. Macrophages are elevated in adipose tissue in aging. However, the changes and also possibly functions of ATMs in aging and aging-related diseases are unclear. In this review, we will summarize recent advances in research on the role of adipose tissue macrophages with aging-associated insulin resistance and discuss their potential therapeutic targets for preventing and treating aging and aging-related diseases.
During the long-term evolution of the host environment,Mycobacterium tuberculosis can completely or partially adapt to survive in the host cells by avoiding or modifying the host response to infection.Many factors contribute to the successful invasion of this pathogen into host cells.Over the past few decades,a great deal of researches have led to deeper understanding of the complex pathogenesis of Mycobacterium tuberculosis,especially its unique Ⅶ secretion system and the cell membrane with a variety of complex lipid have become the hot topic of the aspects.This review summarizes the recent research results of Mycobacterium tuberculosis associated virulence factors,hoping to find new drug targets from these virulence factors.
Using the firm-level data over 1989-2012 from 53 countries, we find religiosity in a country is positively associated with trade credit use by local firms. Specifically, after controlling for firm- and country-level factors as well as industry and year effects, we show that trade credit use is higher in more religious countries. Moreover, both creditor rights and social trust in a country enhance the positive association between religiosity and trade credit use, while the quality of national-level disclosure mitigates the aforementioned positive association. These results are robust to alternative measures of religiosity, alternative sampling requirements and potential endogeneity concerns.
This study aims at evaluating the effects of RTS (rotation softened trauma fixation system) compared with PCPSF (percutaneous conventional pedicle screw fixation) on type A thoracolumbar fractures. In this retrospective cohort study, 116 patients with type A thoracolumbar fractures from March 2014 to June 2018 were enrolled. PCPSF was performed in 60 patients, meanwhile the other 56 patients accepted RTS. VAS scores, Cobb angle, anterior vertebral height (AVH) and perioperative data were compared between the two groups. Both groups were consistent with baseline on demographic and clinical characteristics. No significant difference was observed in VAS score between-group before and after operation. One year after surgery, the VAS score of RTS group was lower than that of PCPSF group (0.7 ± 0.3 vs. 1.5 ± 0.4). The postoperative AVH (%) in PCPSF was 82.3% (95%CI, 81.7–84.6), and 91.78% (95% CI, 91.1–92.4) in RTS. The postoperative improvement rate of AVH (%) in RTS was higher than that in PCPSF (30.6 ± 5.0 [95% CI, 29.2–32.0] vs. 22.0 ± 7.3 [95% CI, 20.2–24.2]). The postoperative Cobb angle (°) in PCPSF was 2.6 ± 3.4 (95%CI,11.7–13.5), and 7.5 ± 2.0 (95%CI,7.0–8.0) in RTS. The postoperative correction of Cobb angle (°) in RTS was higher than that in PCPSF (16.1 ± 3.8 95%CI,15.1–17.1] vs. 11.6 ± 5.2 95%CI,10.3–13.1]). Compared with PCPSF, RTS has advantages in restoring the anterior vertebral height and reducing local kyphosis.
Background A healthy diet in a college student life is critical to ensure their normal growth, study and development. Aim In order to accurately assess the dietary pattern of college students and guide it, our study aims to evaluate the validity of instant photography as an alternative dietary assessment method in college students. Methods Nine participants were enrolled and given a presentation on dietary assessment methods, including weighing, 24-hour recall, and instant photography. The participants took pictures of their foods from three angles before and after eating for constant seven days, foods weighing was completed by others. Then, the participants recalled the foods’ weights after 24 hours. Two trained observers estimated food weight from the digital images (n = 285) gathered at the end of the study with the aid of Chinese food atlas reference. Results Instant photograph showed significant correlation with weighing method on food weights of grains, tubers, vegetables, fruits, meat and eggs (all P ? 0.01 ). 24-hour method had similar correlation with weighing method on food weights except fruits. Compared with 24-hour recall, instant photograph displayed underestimation on weights of grains, tubers, vegetables, and meat. However, instant photograph had more accurate estimations on weights of fruits and egg. Furthermore, compared with nutrients data from weighing method, both instant photography and 24-hour recall methods showed promising estimations on the amounts of energy, protein, fat, carbohydrate, vitamin A, vitamin C, vitamin E, calcium, iron and zinc (all P < 0.001 ). Compared with 24-hour recall, instant photograph displayed underestimation on the amounts of energy, protein, fat, carbohydrate, vitamin A, vitamin C, vitamin E, zinc. However, instant photograph had a more accurate estimation on calcium. Conclusion Instant photography is an easy and feasible dietary assessment method for college students, with valid estimations.