6607 Background: RWD derived from Electronic Health Records (EHR) has detailed clinical information about patient journeys that can assist in clinical research, trial design, safety assessments etc. However, much of the vital information is locked away in unstructured clinical texts and needs to be converted to structured format to be useful for downstream applications. We demonstrate how this can be achieved at scale with a high degree of accuracy through NLP. Methods: NLP models were developed to extract data for 11 clinical variables from unstructured notes of ~98k lung cancer patients and merged with the structured data into a common data model (Table). These models were a combination of domain knowledge, rule-based models, machine learning models, and deep learning models. The increase in fill rate per variable over structured data only was used to quantify the improvement by NLP. The accuracy of the models was assessed against a manually curated dataset comprising of 752 patients. Results: The NLP models significantly improved the fill rate of key clinical variables and were able to extract the information from clinical notes with high accuracy (Table). For some variables such as NSCLC/SCLC status, surgery, tumor grade and histology, all or most of the data was extracted via NLP. Metastatic status via NLP included distant metastasis, locally advanced disease and no metastasis whereas in the structured data, only data for distant metastasis was present. In the case of Performance Status (PS), even though a significant number of patients had at least 1 PS recorded in the structured data, NLP significantly increased longitudinal capture, thus increasing the density of this variable per patient. Conclusions: NLP models can be developed and used to enrich structured RWD data by extracting information from unstructured documents thus significantly improving the utility of this data for downstream applications. Given the high accuracy of these models and the scale at which they can be run, this can be a good alternative to human curation or can augment human curation enabling the creation of very large-scale datasets for clinical research. [Table: see text]
Background Immunotherapy is one of the most prominent therapies for NSCLC patients. While there is a lot of promise, adverse events (AEs) due to immunotherapies are a concern. Entering the era of COVID-19, the interaction of COVID-19 vaccination status with immunotherapy is not fully understood.1-2 As most newly diagnosed NSCLC patients will be vaccinated, understanding this interaction is important for managing their treatment. This study aims at determining whether COVID-19 vaccination status has any significant effect on AEs and outcomes of aNSCLC patients treated with immunotherapies in 1st line. Methods This retrospective study leverages ConcertAI's NSCLC Patient360TM dataset, a deeply curated real-world oncology dataset with patients from across the United States. aNSCLC patients who started 1st line treatment containing an immunotherapy at least 30 days after their last COVID-19 vaccine were included in the vaccine-primed cohort (N= 138). 1st line treatment in these patients started between January 2021 – April 2022. Similarly, a cohort of vaccine-naïve patients was created by including all patients in the dataset who received their 1st line immunotherapy treatment between January 2019 – April 2020 (N=1537) to ensure none of them received COVID-19 vaccine prior to immunotherapy treatment. Descriptive analysis on these cohorts showed no significant differences in terms of age, race, gender and treatment patterns. AEs for each patient during the course of 1st line immunotherapy treatment were identified. These AEs were categorised into 5 levels (table 1). To normalise the effect of length of treatment, AE/time on immunotherapy was calculated. Progression-Free Survival (PFS) and Overall Survival (OS) from start of L1 was also compared between the two cohorts. Results 56% vaccine-naïve and 54% vaccine-primed patients had an AE while on immunotherapy. The distribution of severity of AEs between the two cohorts was also quite similar (table 2). Although the AE/time was higher in the vaccine-naïve cohort (p-value=0.03) (figure 1), this effect was mostly driven by 41 (2.6%) outlier patients who had many AEs in a very short span of time after starting immunotherapy. We believe such outliers were not seen in the vaccine-primed cohort primarily due to its smaller sample size. OS and PFS were similar between the two cohorts (figures 2 and 3). Conclusions COVID-19 vaccination status does not affect frequency or severity of immunotherapy related AEs or have a significant impact on patients' outcomes. As more data becomes available on the vaccine-primed cohort the impact on rarer patient sub-populations can be evaluated Acknowledgements The author would like to thank the following people from the Data Science Solutions team for help in generating accurate OS and PFS plots Andrew Noble, Judith Mueller, Rahul Das, Jericho Cain, Somasekhar Suryadevara – Data Science Solutions Concert AI USA and Rishi Jajoo formerly at ConcertAI References Mei Q, Hu G, Yang Y, et al Impact of COVID-19 vaccination on the use of PD-1 inhibitor in treating patients with cancer: a real-world study. Journal for ImmunoTherapy of Cancer 2022;10:e004157. doi: 10.1136/jitc-2021-004157. Brest P, Mograbi B, Hofman P, et al. COVID-19 vaccination and cancer immunotherapy: should they stick together?. Br J Cancer 2022;126, 1–3. https://doi.org/10.1038/s41416-021-01618-0
1064 Background: CDK4/6 inhibitors plus endocrine therapy are approved for treatment of HR+/HER2- metastatic breast cancer (MBC) and have shown to provide a significant progression free survival benefit over endocrine therapy alone. But not all patients benefit from this treatment and some develop resistance over time. The molecular mechanisms governing this resistance are poorly understood. We have developed a real world dataset that includes data elements from structured EMR tables as well as deeply curated unstructured data from BC patients (ConcertAI Genome360 BC Dataset) who have been treated with CDK4/6 inhibitors and have undergone DNA sequencing to identify somatic mutations. We have leveraged this linked clinical-genomics dataset to identify genetic drivers of resistance and response to CDK4/6 inhibitors. Methods: This retrospective study uses the Genome360 BC Dataset (N = 1249). The patient’s eligibility to be included in this study (N = 456) was HR+/HER2- MBC patients with age > 18 years treated with at least one of the CDK4/6 inhibitors and have response data based on RECIST criteria (responders = 231, non-responders = 225). For each patient in both cohorts, all pathogenic gene mutations and copy number changes were identified and enrichment analysis was performed. Biomarkers with Z value > 1.96 (p value < 0.05) were considered for further analysis. Pathway analysis was performed using these biomarkers and the CDK4/6 pathway to identify pathways and genes that can potentially be targeted to overcome resistance based on the mutational landscape of the patients receiving therapy. Results: We identified 7 potential segments (similar groups of genes) which predicted response or resistance to CDK4/6 inhibitors. Here we present data on 3 such segments which are closely related. Loss of function mutations in RB1 were enriched in the non-responder population (Z value = 2.33; p value = 0.026; N = 31). This is consistent with previously reported findings. In addition, amplifications and gain of function mutations in MYC and associated genes were also significantly enriched in the non-responder population (Z value = 2.71; P value = 0.01; N = 44). Interestingly, loss of function mutations in TSC1/2 genes which are downstream of MYC were predictors of good response to CDK4/6 inhibitors (Z value = 2.19; P value = 0.036; N = 30), strengthening the role of the parallel MYC signaling pathway in resistance to CDK4/6 inhibitors. Conclusions: Using our Genome360 BC Dataset, we have identified genetic markers affecting response to CDK4/6 inhibitors. In addition to the known role of RB1 in resistance to CDK4/6 inhibitors, the MYC signaling pathway emerged as a strong candidate. Based on these results, patients with mutations in these pathways may benefit from addition of mTOR or PKL1 inhibitors to CDK4/6 inhibitors to overcome resistance and prolong their effect.
Background ICIs are a promising class of drugs that have improved the treatment of a broad spectrum of cancers. Biomarkers such as high Tumor Mutation Burden (TMB), high PD-L1 expression, or high Microsatellite Instability (MSI) are predictive of improved response to ICIs.1-3 However, not all patients with these biomarkers respond well to ICIs and some develop resistance over time. The underlying molecular mechanisms of resistance are poorly known. We leveraged our deeply curated real world clinico-genomics database (ConcertAI Genome360TM) of NSCLC patients treated with ICIs to identify the drivers of resistance and response. Methods This retrospective study used the ConcertAI Genome360TM NSCLC dataset. The eligible patients had advanced NSCLC treated with ICIs (N=2532). A subset of these patients also had high TMB/PD-L1/MSI (TPM) status (N=986). The following analysis was performed for both the TPM unselected and high TPM populations. The patients were subdivided into responder and non-responder cohorts based on their response to ICIs. For each patient in both cohorts, genes with pathogenic mutations, fusions and copy number changes were identified and enrichment analysis was performed between cohorts. Biomarkers with p value less than 0.05 were further considered for pathway analysis along with the immune signaling network to identify pathways and genes responsible for response and resistance to ICIs in-spite of TPM high status. Results We identified segments which predicted response or resistance to ICIs (table 1). TERT promoter mutations and loss of function (LOF) mutations in STAG2 promote high TMB and PD-L1 expression respectively.4,5 We also see them enriched in our overall responder cohort, however, the effect of STAG2 LOF mutations towards response to ICIs is seen even in the TPM high population indicating the effect is due to more than just upregulation of PD-L1. LOF mutations in ATM/ATR genes were also enriched in the overall responder cohort in-line with previous observations in murine cells.6,7 Gene amplifications in CDK12 and CEBPA genes were predictive of resistance to ICIs and interestingly, amplifications in CDK12 and RET were also highly predictive of resistance to ICIs in the TPM high cohort which was heavily biased towards response to therapy. Pathway analysis on the combined results identified that genes in the cGAS-STING pathway are playing a vital role in determining response. Conclusions Our analysis highlights some of the underlying mechanisms for response or resistance to ICIs which can provide clues for designing new combination trials for patients whose tumor progresses on ICIs. References Xu Y, Wan B, Chen X, Zhan P, Zhao Y, Zhang T, Liu H, Afzal MZ, Dermime S, Hochwald SN, Hofman P, Borghaei H, Lin D, Lv T, Song Y; written on behalf of AME Lung Cancer Collaborative Group. The association of PD-L1 expression with the efficacy of anti-PD-1/PD-L1 immunotherapy and survival of non-small cell lung cancer patients: a meta-analysis of randomized controlled trials. Transl Lung Cancer Res. 2019 Aug;8(4):413–428. Petrelli F, Ghidini M, Ghidini A, Tomasello G. Outcomes Following Immune Checkpoint Inhibitor Treatment of Patients With Microsatellite Instability-High Cancers: A Systematic Review and Meta-analysis. JAMA Oncol. 2020 Jul 1;6(7):1068–1071. Palmeri M, Mehnert J, Silk AW, Jabbour SK, Ganesan S, Popli P, Riedlinger G, Stephenson R, de Meritens AB, Leiser A, Mayer T, Chan N, Spencer K, Girda E, Malhotra J, Chan T, Subbiah V, Groisberg R. Real-world application of tumor mutational burden-high (TMB-high) and microsatellite instability (MSI) confirms their utility as immunotherapy biomarkers. ESMO Open. 2022 Feb;7(1):100336. Li H, Li J, Zhang C, Zhang C, Wang H. TERT mutations correlate with higher TMB value and unique tumor microenvironment and may be a potential biomarker for anti-CTLA4 treatment. Cancer Med. 2020 Oct;9(19):7151–7160. Nie Z, Gao W, Zhang Y, Hou Y, Liu J, Li Z, Xue W, Ye X, Jin A. STAG2 loss-of-function mutation induces PD-L1 expression in U2OS cells. Ann Transl Med. 2019 Apr;7(7):127. Hu M, Zhou M, Bao X, Pan D, Jiao M, Liu X, Li F, Li CY. ATM inhibition enhances cancer immunotherapy by promoting mtDNA leakage and cGAS/STING activation. J Clin Invest. 2021 Feb 1;131(3):e139333. Sheng H, Huang Y, Xiao Y, Zhu Z, Shen M, Zhou P, Guo Z, Wang J, Wang H, Dai W, Zhang W, Sun J, Cao C. ATR inhibitor AZD6738 enhances the antitumor activity of radiotherapy and immune checkpoint inhibitors by potentiating the tumor immune microenvironment in hepatocellular carcinoma. J Immunother Cancer. 2020 May;8(1):e000340.
e21172 Background: Analysis of Real World Data (RWD) from Electronic Health Records (EHR) for applications such as Health Economics and Outcomes Research (HEOR) or regulatory submissions requires identification of the lines of therapy (LoT) patients have received. LoTs are typically not captured in EHR and must be manually abstracted. As the use of RWD increases, there is a growing need to create algorithms that can work on RWD to extract LoT information in an automated manner with high accuracy. We present here the results of such an algorithm created on NSCLC RWD. Methods: 10950 advanced NSCLC patients from the ConcertAI Oncology RWD database who had received anti-neoplastic treatment after advanced diagnosis were used to build and validate this algorithm. These data were further enriched by expert nurse curators to fill in missing oral drug information and identify progression events. We developed a progression-based LoT (pLoT) model that identified LoT changes in sync with tumor progressions. If patients received multiple regimens before progression they were captured as nested regimens within the LoT. The algorithm uses complex rules to define combination of drugs as regimens (combination rule), identify resumption of regimens (gap rule) or dropping of drugs from regimens as new lines and to handle noisiness in RWD etc. Results: The LoT model accurately captures line changes triggered by progression events as well as any nested regimen changes due to adverse events etc. Patient level validation of LoT was carried out by clinical experts using an in-house tool and found to be consistent with literature & individual drug data. Cohort level analysis of top 3 combinations of therapies used in 1st & 2nd line treatment between 2015-2020 (8200 patients) are shown in Table. Sensitivity analysis on the combination rule showed that this parameter can be changed between 28-33 days without significantly impacting the LoT output (<1% impact). We use a 30 day combination rule as the default. Similarly, the gap rule parameter is quite robust and does not show significant variation between 45 – 90 days (<2% impact). We use 63 days. Conclusions: We have developed a robust algorithm to derive pLoT on RWD at scale assuming availability of curated progression data which can be used to support use cases such as HEOR, clinical development and regulatory submissions. pLoT is better suited for outcomes analysis compared to regimen based LoT since it distinguishes changes in treatment due to progression events from changes due to toxicity, drug availability, etc., and allows analysis on a more homogeneous patient population relating to their past clinical experience. [Table: see text]
e21540 Background: Metastatic status is a crucial variable in most oncology studies but is not available in claims data. The objective of this study is to develop a machine learning model for Imputation of metastatic status from claims data with ground. Truth is derived from highly curated electronic medical record data. Methods: We used a set of 11389 melanoma patients from the ConcertAI real world database of intersecting claims and EMR data that includes data from CancerLinQ Discovery. Using features from claims and our gold standard labels from EMR we built an ML model using (XGBoost) extreme gradient boosting, an algorithm that iteratively combines a set of decision trees into a single model. We used 60% of the data for training, 20% for hyper-parameter tuning, and 20% for holdout testing. The model was built using 55 features. Results: The table below summarizes results. Metrics are on the final hold out set which was unseen by the model and entirely composed of highly curated EMR data. Conclusions: We are able to build a high precision model for the imputation of metastatic melanoma status using claims data. This could enable significantly better use of claims data stemming from the ability to find a metastatic cohort with very few false positives. Providing more precise cohort identification for comparative effectiveness studies. We found features such as secondary neoplasm diagnosis, anti-neoplastic meds, and radiation ranking highly in our analysis of model feature importances. Using techniques to analyze non-linear feature interactions in our AI model we found an interaction relationship between long term anti-neoplastic therapy, reported pain and metastatic status which we plan to further study. This work is preliminary and we are working to further improve model performance.[Table: see text]
e14056 Background: Determination of the metastatic status of a patient is important for outcomes research and candidacy for clinical trials. Structured data in EMR may not always capture the metastatic status, and it is useful to extract it automatically from physician notes. Contextual understanding of the notes is important to resolve issues such as a) local vs distal metastasis b) statements involving family history of metastasis or physician instructing the patient to look for certain signs of metastasis c) text indicating suspicion of metastasis or absence of metastasis d) indirect utterances, e.g. cancer has spread to the bone. e) corrections to previous findings. Methods: We used a set of 20138 breast cancer patients from Concerto HealthAI real world oncology dataset that includes data from CancerLinQ Discovery to build & validate the set of NLP algorithms. 5300 sentences from 1500 patients were annotated & algorithms manually validated by data abstractors for 500 patients. The algorithms developed were the following: 1) Classification of a sentence into 3 classes: Distal/Local metastasis, Suspicious & Other 2) Classification of a sentence into 2 classes: Distal or Local 3) Classification of a patient into 2 classes: Distal metastasis or not distal metastasis 4) Multi label classification for detecting sites of metastasis. Sentence level algorithms were built using Deep Learning and patient level aggregation of sentence level prediction was done using ML approaches including temporal features. Pretrained ULMFiT model was fine-tuned with Concerto HealthAI’s corpus for sentence classification tasks. Results: At a sentence level, we obtained an accuracy of 0.85 for the distal/local vs suspicious vs irrelevant model and 0.97 for the distal vs not distal metastasis model. Our patient level metrics are shown in the table. The classes used for sites of metastasis are Brain, Bone, Lung, Liver, Distant Lymph nodes & Unknown sites. Subset accuracy (mean fraction of labels which match ) of 0.93 was obtained on the hold out test set at patient level. Conclusions: Metastatic status & site of metastasis can be reliably extracted automatically from clinical notes using deep learning techniques. This information will be valuable for clinical trial matching, outcomes research and other applications. [Table: see text]
e13078 Background: Models that can dynamically predict risk of metastatic breast cancer (MBC) recurrence based on cumulative historical clinical data could help guide patient care & surveillance decisions. The objectives of this study were to predict risk of MBC recurrence dynamically from any point after 1 year of initial diagnosis in a BC patient’s journey. We show representative results for predicting 4 year risk post 1 year of date of diagnosis. There are established models to predict risk of distant recurrence at the time of diagnosis but we have not found much work on dynamic risk scores. Methods: We used a set of 3807 patients from the Concerto HealthAI database of oncology EMR data that includes clinical data from CancerLinQ Discovery to build this model that were further enriched by expert nurse curators. The average age at diagnosis was 58 & the average follow up period for this cohort of patients was 6.6 years. The cohort included patients of all breast cancer subtypes. 628 patients had metastatic recurrence within 4 years post 1 year of date of diagnosis. We used 60% of the data for training, 20% for hyper-parameter tuning & 20% for testing & tried out various machine learning (ML) algorithms including Linear Regression, Lasso, Random Forest, Extremely Random Forests, & XGBoost. Extremely Random Forest built using 330 features had the best performance. Results: The performance of various ML algorithms for predicting metastatic BC recurrence within 4 years post 1 year from date of diagnosis is provided in the table below with sensitivity held constant at 0.7. Key variables influencing the results in each model are also indicated. The Extremely Random Forest model for predicting risk of metastatic recurrence within 1 year from 1 year post diagnosis yielded an AUC of 0.814 & a balanced accuracy of 0.719. Conclusions: An AI model to predict risk of metastatic recurrence in breast cancer patients built using a real world dataset yielded promising results. Furthermore, analysis of input variables provided insights not only into the key features driving metastatic recurrence risk such as previous surgery, tumor subtype, stage & age at diagnosis etc. Such a model could be a useful for assessing patient risk & treatment options at various points in a breast cancer patients journey as well as stratify patients for different levels of surveillance. [Table: see text]
e19318 Background: ECOG PS is a prognostic indicator of outcomes, and scores of 0-1 (good ECOG PS) are often required for clinical trial enrollment. Patients treated in non-trial settings often lack ECOG PS scores limiting the ability of Real World Data from these patients to be used in external control arms (ECAs) or to provide optimal specificity for clinical effectiveness research. Machine Learning can be used to impute ECOG PS scores from other clinical data at various points during treatment. Methods: We developed a series of models using logistic regression (LR) or XGBoost (XGB) that impute ECOG PS at initial diagnosis, metastatic diagnosis and final evaluation using a curated Non-Small Cell Lung Cancer cohort of 31,425 patients with at least one ECOG PS score. Results: AUC-ROC values of up to 0.81 could be obtained for imputing a patient’s final ECOG PS, with lower AUC values when imputing ECOG PS at initial and metastatic diagnosis using large numbers (i.e. thousands) of features. We developed more interpretable models with 110 or 40 features with reduced but still satisfactory AUC, with accuracy of predicting good ECOG PS scores of around 80%. Key features were obtained from lab tests, physical exams, comorbidities, medications, age and metastatic status. The table below shows the results of several of these models. Where the models misclassify ECOG PS, the error was rarely greater than 1 grade. Conclusions: ECOG PS is subjective, suggesting that ML based cohort assignment will be sufficiently accurate to support their use in research. Further work will be required to assess if the ML predicted cohorts have different outcomes. [Table: see text]
e21596 Background: There are ongoing efforts to understand and predict exceptional response to existing cancer therapies, but few clinical characteristics of these patients are known. We trained a machine learning model using the Concerto HealthAI database of oncology EMR data that includes clinical data from CancerLinQ Discovery to predict slow progression, a proxy for exceptional response, in aNSCLC in the second line setting. Methods: We trained an XGBoost model to predict patients with a progression free survival (PFS) greater than 180 days from the start of second line therapy (index date). This cutoff approximately determines the top 20% of PFS values in our database (median PFS = 86 days). Patients were included from the study if they (1) were pathologically confirmed aNSCLC without other primary cancer diagnoses and (2) started their second-line therapy between 2013 and 2017. Patients were labeled as slow progressors if they (1) had no evidence of progression or death within 180 days of index and (2) were evaluated for progression for at least 180 days post-index. The model considered data up to 120 days prior to index date. Risk factors in the model included demographics, vitals, common labs, common medical conditions, ECOG performance status, stage, histology, prior cancer treatment patterns, prior progression/response assessments, and medication history. Feature importance was evaluated using SHapley Additive exPlanations (SHAP). Results: 2205 patients met selection criteria of the study. Of these, 420 were labeled as slow progressors. 1776 patients were used for model training and 429 were set aside for model validation. The final model was able to predict slow progression with an AUCROC of 0.75 (F-score 0.48, precision 0.39, recall 0.6). The performance compares favorably to that of a logistic regression model (0.66 AUCROC). Top features that indicated slow progression included a low number of prior progression events or regimens, absence of metastatic disease, lower stage/t-stage/ECOG, absence of COPD, previous treatment with an EGFR inhibitor, normal Alk-Phos/WBC (versus elevated), absence of tachycardia, and a normal BMI (versus low). Conclusions: Machine learning and real world-data provided promising results in predicting slow progression in aNSCLC and may be useful in discovering novel drivers of favorable response.
6556 Background: Survival prediction models for lung cancer patients could help guide their care and therapy decisions. The objectives of this study were to predict probability of survival beyond 90, 180 and 360 days from any point in a lung cancer patient’s journey. Methods: We developed a Gradient Boosting model (XGBoost) using data from 55k lung cancer patients in the ASCO CancerLinQ database that used 3958 unique variables including Dx and Rx codes, biomarkers, surgeries and lab tests from ≤1 year prior to the prediction point, which was chosen at random for each patient. We used 40% data for training, 25% for hyper-parameter tuning, 20% for testing and 15% for holdout validation. Death date available in the Electronic Health Record was cross checked by linkage to death registries. Results: The model was validated on the holdout set of 8,468 patients. The Area Under the Curve (AUC) for the model was 0.79. The precision and recall for predicting survival beyond the three time points were between 0.7-0.8 and 0.8-0.9 respectively (see table). This compares favourably to other lung cancer survival models created using different machine learning techniques (Jochems 2017, Dekker 2009). A Cox-PH model created using the top 20 variables also had a significantly lower performance (see table). Analysis of input variables yielded distinctive patterns for patient subgroups and time points. Tumor status, medications, lab values and functional status were found to be significant in patient sub cohorts. Conclusions: An AI model to predict survival of lung cancer patients built using a large real world dataset yielded high accuracy. This general model can further be used to predict survival of sub cohorts stratified by variables such as stage or various treatment effects. Such a model could be useful for assessing patient risk and treatment options, evaluating cost and quality of care or determining clinical trial eligibility. [Table: see text]
The ability to automatically learn task specific feature representations has led to a huge success of deep learning methods. When large training data is scarce, such as in medical imaging problems, transfer learning has been very effective. In this paper, we systematically investigate the process of transferring a Convolutional Neural Network, trained on ImageNet images to perform image classification, to kidney detection problem in ultrasound images. We study how the detection performance depends on the extent of transfer. We show that a transferred and tuned CNN can outperform a state-of-the-art feature engineered pipeline and a hybridization of these two techniques achieves 20% higher performance. We also investigate how the evolution of intermediate response images from our network. Finally, we compare these responses to state-of-the-art image processing filters in order to gain greater insight into how transfer learning is able to effectively manage widely varying imaging regimes.
Histograms are widely used in medical imaging, network intrusion detection, packet analysis and other stream-based high throughput applications. However, while porting such software stacks to the GPU, the computation of the histogram is a typical bottleneck primarily due to the large impact on kernel speed by atomic operations. In this work, we propose a stream-based model implemented in CUDA, using a new adaptive kernel that can be optimized based on latency hidden CPU compute. We also explore the tradeoffs of using the new kernel vis-\`a-vis the stock NVIDIA SDK kernel, and discuss an intelligent kernel switching method for the stream based on a degeneracy criterion that is adaptively computed from the input stream.
Many high performance-computing algorithms are bandwidth limited, hence the need for optimal data rearrangement kernels as well as their easy integration into the rest of the application. In this work, we have built a CUDA library of fast kernels for a set of data rearrangement operations. In particular, we have built generic kernels for rearranging m dimensional data into n dimensions, including Permute, Reorder, Interlace/De-interlace, etc. We have also built kernels for generic Stencil computations on a two-dimensional data using templates and functors that allow application developers to rapidly build customized high performance kernels. All the kernels built achieve or surpass best-known performance in terms of bandwidth utilization.
We consider the problem of analyzing influences in financial networks by studying correlations in stock price movements of companies in the S&P 500 index and measures of influence that can be attributed to each company. We demonstrate that under a novel and natural measure of influence involving cross-correlations of stock market returns and market capitalization, the resulting network of financial influences is Scale Free. This is further corroborated by the existence of an intuitive set of highly influential hub nodes in the network. Finally, it is also shown that companies that have been deleted from the S&P 500 index had low values of influence.
We consider the issue of model selection for some prediction problems in consumer finance. In particular, we look at performance metrics in the context of classification problems. Example areas considered include response modeling, profitability modeling and default prediction in the framework of a customer relationship management (CRM) system. We propose some guidelines for choosing the appropriate performance measure for the predictive model based on the decision framework it is part of.