BACKGROUND The growing disparity between the rising demand for liver transplantation (LT) and the limited availability of donor organs has prompted a greater reliance on older liver grafts. Traditionally, utilizing livers from elderly donors has been associated with outcomes inferior to those achieved with grafts from younger donors. By accounting for additional risk factors, we hypothesize that the utilization of older liver grafts has a relatively minor impact on both patient survival and graft viability. AIM To evaluate the impact of donor age on LT outcomes using multivariate analysis and comparing young and elderly donor groups. METHODS In the period from April 2013 to December 2018, 656 adult liver transplants were performed at the University Hospital Merkur. Several multivariate Cox proportional hazards models were developed to independently assess the significance of donor age. Donor age was treated as a continuous variable. The approach involved univariate and multivariate analysis, including variable selection and assessment of interactions and transformations. Additionally, to exemplify the similarity of using young and old donor liver grafts, the group of 87 recipients of elderly donor liver grafts (≥ 75 years) was compared to a group of 124 recipients of young liver grafts (≤ 45 years) from the dataset. Survival rates of the two groups were estimated using the Kaplan-Meier method and the log-rank test was used to test the differences between groups. RESULTS Using multivariate Cox analysis, we found no statistical significance in the role of donor age within the constructed models. Even when retained during the entire model development, the donor age's impact on survival remained insignificant and transformations and interactions yielded no substantial effects on survival. Consistent insignificance and low coefficient values suggest that donor age does not impact patient survival in our dataset. Notably, there was no statistical evidence that the five developed models did not adhere to the proportional hazards assumption. When comparing donor age groups, transplantation using elderly grafts showed similar early graft function, similar graft (P = 0.92), and patient survival rates (P = 0.86), and no significant difference in the incidence of postoperative complications. CONCLUSION Our center's experience indicates that donor age does not play a significant role in patient survival, with elderly livers performing comparably to younger grafts when accounting for other risk factors.
Background: Hepatocellular carcinoma (HCC) is one of the leading indications for liver transplantation (LT) however, selection criteria remain controversial. We aimed to identify survival factors and predictors for tumour recurrence using machine learning (ML) methods. We also compared ML models to the Cox regression model. Methods: Thirty pretransplant donor and recipient general and tumour specific parameters were analysed from 170 patients who underwent orthotopic liver transplantation for HCC between March 2013 and December 2019 at the University Hospital Merkur, Zagreb. Survival rates were calculated using the Kaplan-Meier method and multivariate analysis was performed using the Cox proportional hazards regression model. Data was also processed through Coxnet (a regularized Cox regression model), Random Survival Forest (RSF), Survival Support Vector Machine (SVM) and Survival Gradient Boosting models, which included pre-processing, variable selection, imputation of missing data, training and cross-validation of the models. The cross-validated concordance index (CI) was used as an evaluation metric and to determine the best performing model. Results: Kaplan-Meier curves for 5-year survival time showed survival probability of 80% for recipient survival and 82% for graft survival. The 5-year HCC recurrence was observed in 19% of patients. The best predictive accuracy was observed in the RSF model with CI of 0.72, followed by the Survival SVM model (CI 0.70). Overall ML models outperform the Cox regression model with respect to their limitations. Random Forest analysis provided several relevant outcome predictors: alpha fetoprotein (AFP), donor C-reactive protein (CRP), recipient age and neutrophil to lymphocyte ratio (NLR). Cox multivariate analysis showed similarities with RSF models in identifying detrimental variables. Some variables such as donor age and number of transarterial chemoembolization treatments (TACE) were pointed out, but these were not influential in our RSF model. Conclusions: Using ML methods in addition to classical statistical analysis, it is possible to develop sufficient prognostic models, which, compared to established risk scores, could help us quantify survival probability and make changes in organ utilization.
Transmission losses through the building envelope account for a large proportion of building energy balance. One of the most important parameters for determining transmission losses is thermal transmittance. Although thermal transmittance does not take into account dynamic parameters, it is traditionally the most commonly used estimation of transmission losses due to its simplicity and efficiency. It is challenging to estimate the thermal transmittance of an existing building element because thermal properties are commonly unknown or not all the layers that make up the element can be found due to technical-drawing information loss. In such cases, experimental methods are essential, the most common of which is the heat-flux method (HFM). One of the main drawbacks of the HFM is the long measurement duration. This research presents the application of deep learning on HFM results by applying long-short term memory units on temperature difference and measured heat flux. This deep-learning regression problem predicts heat flux after the applied model is properly trained on temperature-difference input, which is backpropagated by measured heat flux. The paper shows the performance of the developed procedure on real-size walls under the simulated environmental conditions, while the possibility of practical application is shown in pilot in-situ measurements.
District heating systems are an important part of the future smart energy systems and are seen in the European Union as a vehicle for reaching energy efficiency targets. Integrating different energy systems requires high prediction accuracy for all energy sub-systems. Within this paper a data analysis was made with the goal of identifying a high accuracy prediction model and ranking the most influential parameters on heat consumption of final consumers in district heating systems. The data set consisted of the actual billing data comprising of 260 buildings and it was additionally supplemented by the behavioural data obtained from interviews and questionnaires conducted on the demonstration building in Zagreb, Croatia. The authors choose regression trees, random forest and regression support vector machines as algorithms for testing prediction accuracy and evaluating the variable importance ranking on the data set. The best performing algorithm was random forest, resulting with high prediction accuracy and the root mean squared error of prediction of specific annual heat consumption below 1 kWh/m 2 . Furthermore, all analysed machine learning algorithms ranked importance variables for both technical and behavioural parameters, giving the indication what parameters should be influenced in order to reach specific targets, such as energy savings. (C)2020 Elsevier Ltd. All rights reserved.
In this paper, we present the preliminary results on the analysis of deep learning terms used for natural language processing (NLP) tasks. We propose a statistical analysis of papers published from 2012 to 2015 in the main ACL conferences. Our aim is investigating which DL term, and consequently which DL method, is mostly used for each specific NLP task, since its introduction in the field. In order to do this, our first contribution is the development of two terminological lists, referring respectively to DL methods for text analysis and NLP tasks. The list of deep learning terms contains 41 terms and acronyms, as well as a NLP term list contains 145 terms and acronyms. From our corpus, the frequencies of various terms have been extracted with respect to the ACL conference and the publication year. After the preliminary data analysis, we decided to restrict the extraction process to abstract texts. We applied multivariate techniques called correspondence analysis in order to visualize and evaluate the joint behavior of our variables.
Automatic Term Extraction (ATE) extracts terminology from domain-specific corpora. ATE is used in many NLP tasks, including Computer Assisted Translation, where it is typically applied to individual documents rather than the entire corpus. While corpus-level ATE has been extensively evaluated, it is not obvious how the results transfer to document-level ATE. To fill this gap, we evaluate 16 state-of-the-art ATE methods on full-length documents from three different domains, on both corpus and document levels. Unlike existing studies, our evaluation is more realistic as we take into account all gold terms. We show that no single method is best in corpus-level ATE, but C-Value and KeyConceptRelatendess surpass others in document-level ATE.
We present a preliminary study on predicting news values from headline text and emotions. We perform a multivariate analysis on a dataset manually annotated with news values and emotions, discovering interesting correlations among them. We then train two competitive machine learning models – an SVM and a CNN – to predict news values from headline text and emotions as features. We find that, while both models yield a satisfactory performance, some news values are more difficult to detect than others, while some profit more from including emotion information.
In this paper, we describe our preliminary study on annotating event mention as a part of our research on high-precision news event extraction models. To this end, we propose a two-layer annotation scheme, designed to separately capture the functional and conceptual aspects of event mentions. We hypothesize that the precision of models can be improved by modeling and extracting separately the different aspects of news events, and then combining the extracted information by leveraging the complementarities of the models. In addition, we carry out a preliminary annotation using the proposed scheme and analyze the annotation quality in terms of inter-annotator agreement.
Information retrieval tasks (e.g., recommendation, query expansion) could benefit from identification, formalizing, and relevance ranking of conceptual links between documents. Using concept graphs (i.e., knowledge bases) to characterize relatedness between documents have opened up opportunities for formalizing semantic search and browsing over text collections -- by representing documents as subgraphs in the concept graph, one can establish conceptual and structural similarity of document pairs. In some use scenarios, it is appropriate to learn the metrics from a collective input by the users. In others, the user needs are highly idiosyncratic and require inference from a few examples with a strong user bias. In this work, we investigate to what extent can an algorithm exploiting knowledge from an external knowledge base (KB) capture the conceptual links between texts that humans identify as relevant. We first characterize the concept paths between any two concepts linked from two different documents and define a metric for concept relatedness by scoring these paths. Second, we aggregate the pairwise concept relatedness to define a composite metrics for capturing most salient concepts linking the two input documents. We report on the experiments that involve the use of paraphrased explanations of the document relatedness by the end-users. The results of the preliminary study, conducted with the goal of informing the design of the graph operators to enable content search and recommendations, indicate that with a KB-based algorithm we are able to match human-identified conceptual commonalities between texts significantly better than by comparing simple bag-of-word representations of documents.
Software defect prediction research relies on data that must be collected from otherwise separate repositories. To achieve greater generalization of the results, standardized protocols for data collection and validation are necessary. This paper presents an exhaustive survey of techniques and approaches used in the data collection process. It identifies some of the issues that must be addressed to minimize dataset bias and also provides a number of measures that can help researchers to compare their data collection approaches and evaluate their data quality. Moreover, we present a data collection procedure that uses a bug-code linking technique based on regular expression. The detailed comparison and root cause analysis of inconsistencies with a number of popular data collection approaches and their publicly available datasets, reveals that our procedure achieves the most favorable results. Finally, we implement our data collection procedure in a data collection tool we name the Bug-Code (BuCo) Analyzer.
Recent research has explored the use of Knowledge Bases (KBs) to represent documents as subgraphs of a KB concept graph and define metrics to characterize semantic relatedness of documents in terms of properties of the document concept graphs. However, none of the studies so far have examined to what degree such metrics capture a user-perceived relatedness of documents. Considering the users' explanations of how pairs of documents are related, the aim is to identify concepts in a KB graph that express the same notion of document relatedness. Our algorithm generates paths through the KB graph that originate from the terms in two documents. KB concepts where these paths intersect capture the semantic relatedness of the two starting terms and therefore the two documents. We consider how such intersecting concepts relate to the concepts in the users' explanations. The higher the users' concepts appear in the ranked list of intersecting concepts, the better the method in capturing the users' notion of document relatedness. Our experiments show that our approach outperforms a simpler graph method that uses properties of the concept nodes alone.
Knowledge is an essential element of modern business and increasing attention is given to its acquisition, distribution and exploitation in everyday business activities.Therefore, KONČAR launched the development of a knowledge management system for its own demands and initiated a collaboration with the academic community for scientific research purposes and potential broader social significance of the project.With regard to the multidisciplinary nature of knowledge management, an agreement was reached with the University of Zagreb, the Faculty of Humanities and Social Sciences and the Faculty of Electrical Engineering and Computing.The knowledge management system will enable an effective management of all segments of intellectual capital of an organization, resulting in increase in productivity and higher market competitiveness, as well as an increased capability for generating new values for all parties to the agreement.
When tweeting on a topic, Twitter users often post messages that convey the same or similar meaning. We describe TweetingJay, a system for detecting paraphrases and semantic similarity of tweets, with which we participated in Task 1 of SemEval 2015. TweetingJay uses a supervised model that combines semantic overlap and word alignment features, previously shown to be effective for detecting semantic textual similarity. TweetingJay reaches 65.9% F1-score and ranked fourth among the 18 participating systems. We additionally provide an analysis of the dataset and point to some peculiarities of the evaluation setup.
Software Defect Prediction (SDP) deals with localization of potentially faulty areas of the source code. Classification models are the main tool for performing the prediction and the search for a model of utmost performance is an ongoing activity. This paper explores the performance of Rotation Forest classification al gorithm in the SDP problem domain. Rotation Forest is a novel algorithm that exhibited excellent performance in several studies. However, it was not systematically used in the SDP. Furthermore, it is very important to perform the case studies in various contexts. This study uses 5 subsequent releases of Eclipse JDT as the objects of the analysis. The performance evaluation is based on comparison with two other, known classification models that exhibited very good performance so far. The results of our case study concur with other studies that recognize the Rotation forest to be the state of the art classification algorithm
Software Defect Prediction (SDP) empirical studies are highly biased with the quality of data and widely suffer from limited generalizations. The main reasons are the lack of data and its systematic data collection procedures. Our research aims at producing the first systematically defined data collection procedure for SDP datasets that are obtained by linking separate development repositories. This paper is the first step to achieving that objective, performing an exploratory study. We review the existing literature on approaches and tools used in the collection of SDP datasets, derive a detailed collection procedure and test it in this exploratory study. We quantify the bias that may be caused by the issues we identified and we review 35 tools for software product metrics collection. The most critical issues are many-to-many relation between bug-file links, duplicated bug-file links and the issue of untraceable bugs. Our research provides more detailed, experience based data collection procedure, crucial for further development of SDP body of knowledge. Furthermore, our findings enabled us to develop the automatic data collection tool.
Annie Morin合作论文数IFSIC( Institut de Formation en Informatique et Communication)3