Background: Gastroesophageal reflux disease (GERD) symptoms substantially impair patients' quality of life. The use of patient-reported outcome (PRO) instruments for symptom measurement has been advocated by regulatory authorities. However, current tools for GERD symptom evaluation are limited by recall bias. To improve the real-time characterization of GERD symptoms, we developed an electronic diary (e-diary) for daily symptom monitoring. Objective: This study aimed to develop and optimize a PRO-based e-diary for GERD symptom evaluation and to examine the effect of symptom frequency on adherence. Methods: The GERD e-diary evaluated 8 daytime (acid regurgitation, cough, heartburn, sour taste in the mouth, hiccups, hoarseness, dysphagia, and chest pain) and 2 nighttime symptoms (acid regurgitation and cough) for 8 consecutive weeks. Adherence, defined as the daily completion rate of e-diary, was analyzed and optimized across three stages: (1) no reminder, (2) sending reminder SMS text messaging upon the detection of missing data (no reminders during the first 3-5 days after enrollment), and (3) immediate installation of reminder system at enrollment. Weekly symptom frequency was calculated as the sum of symptomatic days per week. Multiple regression analyses were performed to examine the effects of system optimization and symptom frequency on adherence after controlling for confounders. Results: A total of 138 patients with GERD (70 men, 68 women; mean [SD] age 52.9 [12.3] years) were recruited. At the first stage, the adherence was 47.2%, 40%, and 57.6% for nighttime, daytime, and overall symptoms. System optimization significantly improved the adherence of nighttime symptoms by 12.5% (95% CI 3.7-21.3) and 10.9% (95% CI 2.6-19.2), daytime symptoms by 21.7% (95% CI 14.2-29.2) and 20.8% (95% CI 13.7-27.9), and overall symptoms by 16.5% (95% CI 9.8-23.2) and 18.5% (95% CI 12.2-4.8) in the second and third stages, respectively. Symptom frequency was positively associated with adherence, increasing by 0.7% (95% CI 0.6-0.8) for overall symptoms and 0.9% (95% CI 0.7-1) for both daytime and nighttime symptoms per additional symptom frequency. Adherence gradually decreased along the study period. (first vs eighth week: nighttime 80.1% vs 61.5%, beta=-18.6, 95% CI -26.9 to -10.3; daytime 85.1% vs 66.8%, beta=-18.3, 95% CI -25.6 to -11; overall 95.1% vs 78%, beta=-17.2, 95% CI -23.5 to -10.9). Conclusions: The adherence of the GERD e-diary can be optimized by using SMS text messaging reminders. Higher symptom frequency was associated with increased adherence, although engagement declined over time. This innovative PRO-based e-diary with prolonged recording provides a real-time, prospective tool that overcomes the recall and ecological biases inherent in traditional short-term retrospective GERD symptom assessments. This advancement empowers patients through improved self-awareness and provides physicians with precise, long-term data, facilitating tailored therapeutic interventions and supporting personalized GERD management.
The integration of medical open databases with artificial intelligence (AI) technologies marks a transformative era in biomedical research and health care innovation. Over the past 25 years, initiatives like PhysioNet have revolutionized data access, fostering unprecedented levels of collaboration and accelerating medical discoveries. This rise of medical open databases presents challenges, particularly in harmonizing research enablement with patient confidentiality. In response, privacy laws such as the Health Insurance Portability and Accountability Act have been established, and privacy-enhancing technologies have been adopted to maintain this delicate balance. Privacy-enhancing technologies, including differential privacy, secure multiparty computation, and notably, federated learning (FL), have become instrumental in safeguarding personal health information. FL, in particular, represents a significant advancement by enabling the development and training of AI models on decentralized data. In Taiwan, significant strides have been made in aligning with these global data-sharing and privacy standards. We have actively promoted the sharing of medical data through the development of dynamic consent systems. These systems enable individuals to control and adjust their data-sharing preferences, ensuring transparency and continuity of consent in the ever-evolving landscape of digital health. Despite the challenges associated with privacy protections, the benefits, including improved diagnostics and treatment, are substantial. The availability of open databases has notably accelerated AI research, leading to significant advancements in medical diagnostics and treatments. As the landscape of health care research continues to evolve with open science and FL, the role of medical open databases remains crucial in shaping the future of medicine, promising enhanced patient outcomes and fostering a global research community committed to ethical integrity and privacy.
Han Chinese people comprise nearly 20% of the global population but remain under-represented in genetic studies1,2, so there is an urgent need for large-scale cohorts to advance precision medicine. Here we present the Taiwan Precision Medicine Initiative (TPMI), established by Academia Sinica in collaboration with 16 major medical centres around Taiwan, which has recruited 565,390 participants who consent to provide DNA samples for genetic profiling and grant access to their electronic medical records (EMRs) for research. EMR access is both retrospective and prospective, allowing longitudinal studies. Genetic profiling is done with population-optimized arrays of single-nucleotide polymorphisms for people of Han Chinese ancestry, which enable genome-wide association3,4, phenome-wide association5,6 and polygenic risk score7,8 studies to be performed to evaluate common disease risk and pharmacogenetic response. Participants also agreed to be re-contacted for future research and receive personalized genetic risk profiles with health management recommendations. The TPMI has established the TPMI Data Access Platform, a central database and analysis platform that both safeguards the security of the data and facilitates academic research. As a large cohort of individuals with non-European ancestry that merges genetic profiles with EMR data and enables longitudinal follow-up, TPMI provides a unique resource that could be used to validate genetic risk prediction models, perform clinical trials of risk-based health management and inform health policies. Ultimately, the TPMI cohort will contribute to global genetic research and serve as a model for population-based precision medicine.
Climate change has resulted in frequent extreme disasters and scarce resources, leading to a massive population into cities for favorable survival conditions, and also increasing urban air pollution burdens. It is urgent to assess population health risks related with urban air pollution, which usually relies on census method and meteorological measurement data. However, health impacts may be underestimated, because of challenges to represent the dynamic population mobility and perform unified analysis of different pollution hazards. The contribution of this work is to combine census data with Location Based Service to identify the spatiotemporal mobility pattern of urban population, and then population-weighted exposure (PWE) and health impacts of various air pollution (PM2.5, O3, and NO2) are synergistically evaluated. Taking Nanjing as the study area, it was found that the pollution peak areas correlated with population mobility in the study region, shifting from urban suburbs to the center during the daytime, with the maximum concentration exceeding 165 mu g/m3. O3 caused a relatively high PWE level and had a greater health impact than PM2.5 and NO2, adding the mortality by up to 5 % especially on weekdays. The annual health impact of O3 was approximately twice that of PM2.5 and NO2. Human-centered regulation strategies of urban air pollution were proposed in terms of personnel behaviors, government control, and urban design towards mitigation of air pollution risk and sustainable urban development.
AbstractDNA sequencing of patients with rare disorders has been highly successful in identifying “causal variants” for numerous conditions. However, there are many reports of healthy individuals who harbor these deleterious variants, leading to the concept of incomplete penetrance and doubt about the utility of genetic testing in clinical practice and population screening. As the deleterious variants are rare, the penetrance of these variants in the population is largely unknown. We analyzed the genetic and clinical data from 486,956 participants of the Taiwan Precision Medicine Initiative (TPMI) to determine the risk difference between those with and without deleterious variants. In all, we analyzed 292 disease-relevant variants and their clinical outcomes to assess their association. We found that only 15 variants show a risk difference exceeding 5% between those with or without the variants. In essence, 87.3% of deleterious variants exhibit minimal risk differences, suggesting a limited impact on the individual and population levels. Our analysis revealed increasing trends with age in six cardiovascular and degenerative diseases and bell-shaped trends in two cancers. Additionally, we identified three clinical outcomes exhibiting a dose-response relationship with the number of deleterious variants. Our findings show that large-scale testing of deleterious variants found in the literature is not warranted, except for those exhibiting large disease risk differences.
The Taiwan Precision Medicine Initiative (TPMI), a project initiated by the Academia Sinica in collaboration with 16 major medical centers around Taiwan, has recruited 565,390 participants who consented to provide DNA samples for genetic profiling and grant access to their electronic medical records (EMR) for studies to develop precision medicine. Access to the EMR is both retrospective and prospective, allowing researchers to conduct prospective studies over time. Genetic profiling is done with population-optimized SNP arrays for the Han Chinese populations that enable genetic analyses such as genome-wide association, phenome-wide association, and polygenic risk score studies to evaluate common disease risk and pharmacogenetic response. Furthermore, the TPMI participants agree to be contacted for future research opportunities related to their genetic risks and receive personalized genetic risk profiles with health management recommendations. TPMI has established the TPMI Data Access Platform (TDAP), a central database and analysis platform that both safeguards the security of the data and facilitates academic research. The TPMI is the largest non-European cohort that merges genetic profiles with EMR in the world. With a cohort that can be followed over time, it can be utilized to validate genetic risk prediction models, conduct clinical trials to show the efficacy of risk-based health management, and optimize health policies based on genetic risks. In this report, we describe the TPMI study design, the population and genetic characteristics of the TPMI cohort, and the power it provides to conduct crucial studies in developing precision medicine on a population and personal level. As Han Chinese represent almost 20% of the world's population, the results of TPMI studies will benefit >1.4 billion people around the world and serve as a model for developing population-based precision medicine. ### Competing Interest Statement The authors have declared no competing interest.
Introduction Surveys are common research tools, and questionnaires revisions are a common occurrence in longitudinal studies. Revisions can, at times, introduce systematic shifts in measures of interest. We formulate that questionnaire revision are a stochastic process with transition matrices. Thus, revision shifts can be reduced by first estimating these transition matrices, which can be utilized in estimation of interested measures. Materials and method An ideal survey response model is defined by mapping between the true value of a participant’s response to an interval in the grouped data type scale. A population completed surveys multiple times, as modeled with multiple stochastic process. This included stochastic processes related to true values and intervals. While multiple factors contribute to changes in survey responses, here, we explored the method that can mitigate the effects of questionnaire revision. We proposed the Version Alignment Method (VAM), a data preprocessing tool, which can separate the transitions according to revisions from all transitions via solving an optimization problem and using the revision-related transitions to remove the revision effect. To verify VAM, we used simulation data to study the estimation error and a real life MJ dataset containing large amounts of long-term questionnaire responses with several questionnaire revisions to study its feasibility. Result We compared the difference of the annual average between consecutive years. Without adjustment, the difference is 0.593 when the revision occurred, while VAM brought it down to 0.115, where difference between years without revision was in the 0.005, 0.125 range. Furthermore, our method rendered the responses to the same set of intervals, thus comparing the relative frequency of items before and after revisions became possible. The average estimation error in L infinity was 0.0044 which occupied the 95% CI which was constructed by bootstrap analysis. Conclusion Questionnaire revisions can induce different response bias and information loss, thus causing inconsistencies in the estimated measures. Conventional methods can only partly remedy this issue. Our proposal, VAM, can estimate the aggregate difference of all revision-related systematic errors and can reduce the differences, thus reducing inconsistencies in the final estimations of longitudinal studies.
The high quality of simulation systems depends on precise input parameters. The underlying population model is crucial for an agent-based simulation system that studies the spread of infectious diseases. To build the population model, the system requires a household structure (HSD) of the simulated area, including a household list with recorded age information for each member. Previous research has shown that changes in household structure significantly impact disease-spreading patterns. However, with the increasing frequency of severe infectious diseases like SARS (the year 2002), H1N1 (the year 2009), and COVID-19 (the year 2019), using outdated HSD data is inappropriate. This paper proposes a Monte-Carlo-based approach to approximate the HSD for a given year using aggregated information from a range of years. The validation of our algorithm shows good matches with different diseases. As a result, we obtained an HSD for 2020 to study the spread of new COVID-19 variants and future outbreaks.
: In the kernel of an agent-based disease-spreading simulation system, the key factor is the commuting flows of students and workers during weekdays, which gives the movement of people between their residents and offices/schools. During commuting, people who lived in different areas mixed, which increases the spatial spreading of the virus temporally. It is difficult to extract the exact flow from data such as the census. However, small-scale survey examples and aggregated information, such as the size of schools and dormitories and transportation utilization, are known. Using the above, together with information on transportation routes and public transits, in this paper, we give a method based on the well-known flow conservation principle to construct a commuting flow in Taiwan. Validations are given to show such constructed data to fairly describe the real flow by observing our simulation system’s behaviors against what happened in previous pandemics.
Purpose: Retinopathy screening via digital imaging is promising for early detection and timely treatment, and tracking retinopathic abnormality over time can help to reveal the risk of disease progression. We developed an innovative physician-oriented artificial intelligence-facilitating diagnosis aid system for retinal diseases for screening multiple retinopathies and monitoring the regions of potential abnormality over time. Approach: Our dataset contains 4908 fundus images from 304 eyes with image-level annotations, including diabetic retinopathy, age-related macular degeneration, cellophane maculopathy, pathological myopia, and healthy control (HC). The screening model utilized a VGG-based feature extractor and multiple-binary convolutional neural network-based classifiers. Images in time series were aligned via affine transforms estimated through speeded-up robust features. Heatmaps of retinopathy were generated from the feature extractor using gradient-weighted class activation mapping++, and individual candidate retinopathy sites were identified from the heatmaps using clustering algorithm. Nested cross-validation with a train-to-test split of 80% to 20% was used to evaluate the performance of the screening model. Results: Our screening model achieved 99% accuracy, 93% sensitivity, and 97% specificity in discriminating between patients with retinopathy and HCs. For discriminating between types of retinopathy, our model achieved an averaged performance of 80% accuracy, 78% sensitivity, 94% specificity, 79% F1-score, and Cohen's kappa coefficient of 0.70. Moreover, visualization results were also shown to provide reasonable candidate sites of retinopathy. Conclusions: Our results demonstrated the capability of the proposed model for extracting diagnostic information of the abnormality and lesion locations, which allows clinicians to focus on patient-centered treatment and untangles the pathological plausibility hidden in deep learning models.
The kernel of an agent based simulation system for spreading of infectious disease needs a so called household structure (HSD) of the area being simulated which contains a list of households with the age of each member in the household being recorded. Such a household structure is available in a Census that is usually released every 10 years. Previous researches have shown the changing of the household structure has a great impact on disease spreading patterns. It is observed that the changing of the household structure e.g., the average citizen ages and household size, is at a faster speed. However, serious infectious diseases, such as SARS (year 2002), H1N1 (year 2009) and COVID-19 (year 2019), occur with a higher frequency now than previous eras. For example, it would be bad to use HSD2010 built using Census 2010 to simulate COVID-19. In view of this situation, we need a better way to obtain a good household structure in between the Census years in order for an agent-based simulation system to be effective. Note that though a detailed Census is not available every year, aggregated information such as the number of households with a particular size, and the number of people of a particular age are usually available almost monthly. Given HSDx, the household structure for year x, and the aggregated information from year y where y > x, we propose a Monte-Carlo based approach "patching" HSDx to get an approximated HSDy. To validate our algorithm, we pick x and y - x + 10 which both Censuses are available and find out the root-mean-square error (RMSE) between Census's HSDy and generated HSDy is fairly small for x = 1990 and 2000. The spreading patterns obtained by our simulation system have good matches. We hence obtain HSD2020 to be used in your system for studying the spreading of COVID-19.
In this paper, we report some initial results obtained from the agent-based simulation system SimTW about the changing of spreading dynamics, e.g. speed, magnitude and affected people of different ages, when the target society is aging. A disease model of influenza is built and then is invoked with two different social structures, e.g., population and household distribution, and working and schooling patterns based on Census 2000 and Census 2010 of Taiwan. In the 10 years time, the average population age in the country increases from 33.0 to 37.6 while the average household size decreases from 3.19 to 2.94. From the simulation results, we find that in the more aging year-2010 society, the pandemic, if occurred, is smaller, in terms of the total number of infected persons and slower in terms of the date of the peak number of daily new cases, but is more serious both in terms of the numbers of needed hospital beds and death cases. Using this finding, we hope to motivate further discussions on adapting public health policies to this inevitable global trend of aging.
Deep Neural Networks (DNNs) are very popular in many machine learning domains. To achieve higher accuracy, DNNs have become deeper and larger. However, the improvement in accuracy comes with the price of the longer inference time and energy consumption. The marginal cost to increase a unit of accuracy has become higher as the accuracy itself is rising.The Branchynet, known as early exits, is an architecture to address increasing marginal cost for improving accuracy. The Branchynet adds extra side classifiers to a DNN model. The inference on a significant portion of the samples can exit from the network earlier via these side branches if they already have high confidence in the results.The Branchynet requires manually tuning the learning hyperparameters, e.g., the locations of branches and the confidence threshold for early exiting. The effectiveness of this manual tuning dramatically impacts the efficiency of the tuned networks. To the best of our knowledge, there are no efficient algorithms to find the best branch location, which is a trade-off between the accuracy and inference time on the Branchynet.We propose an algorithm to find the optimal branch locations for the Branchynet. We formulate the problem of finding the optimal branch location for the branchynet as an optimization problem, and prove that the branch placement problem is an NPcomplete problem. We then derive dynamic programming that runs in pseudo-polynomial time and solves the branch placement problem optimally.We also implement our algorithm and solve the branch placement problems on four types of VGG networks. The experiment results indicate that our dynamic programming can find the optimal branch locations for generating the maximum number of correct classifications within a given time budget. We also run the four VGG models on a GeForce RTX-3090 GPU with the branch combination found by the dynamic programming. The experiment results show that our dynamic programming accurately predicts the number of correct classifications and the execution time on the GPU.
Dementia is related to the cellular accumulation of β-amyloid plaques, tau aggregates, or α-synuclein aggregates, or to neurotransmitter deficiencies in the dopaminergic and cholinergic pathways. Cellular and neurochemical changes are both involved in dementia pathology. However, the role of dopaminergic and cholinergic networks in metabolic connectivity at different stages of dementia remains unclear. The altered network organisation of the human brain characteristic of many neuropsychiatric and neurodegenerative disorders can be detected using persistent homology network (PHN) analysis and algebraic topology. We used 18F-fluorodeoxyglucose positron emission tomography (18F-FDG PET) imaging data to construct dopaminergic and cholinergic metabolism networks, and used PHN analysis to track the evolution of these networks in patients with different stages of dementia. The sums of the network distances revealed significant differences between the network connectivity evident in the Alzheimer’s disease and mild cognitive impairment cohorts. A larger distance between brain regions can indicate poorer efficiency in the integration of information. PHN analysis revealed the structural properties of and changes in the dopaminergic and cholinergic metabolism networks in patients with different stages of dementia at a range of thresholds. This method was thus able to identify dysregulation of dopaminergic and cholinergic networks in the pathology of dementia.
Simulation systems are human artifacts to capture the abstraction and simplification of the real world. Study the output of simulation systems can help us understand the real world better. Deep learning system needs large volume and high quality data, therefore, a perfect match with simulation systems. We use the data from an agent based simulation system for disease transmission, to train the deep neural network to perform several prediction tasks. The model reaches 80 percent accuracy to predict the infectious level of virus, the prediction of the peak date is off by at most 8 days 90 percent of the time, and the prediction of the peak value is off at most 20 percent 90 percent of the time at the end of the 7th week. We use some preprocessing tricks and relative error leveling to resolve the magnitude problem. Among all these encouraging results, we did encounter some difficulty when predicting the index date given information at the middle of an epidemic. We note that if some interesting concepts are difficult to predict in a simulated world, it sheds some lights on the difficulty for real world scenarios. To learn the effects of mitigation strategies is an interesting and sensible next step.
The deep learning approach has been successfully applied to various disciplines. When using optimization algorithms, there is a need to evaluate the performance of solutions found so far. The simulation system usually serve as the evaluator. However, to speedup the process, an approximation function, called surrogate, can replace the time consuming simulator. We propose to use deep learning to construct the surrogate function in epidemiology. The simulator is an agent-based stochastic model for influenza and the optimization problem is to find vaccination strategy to minimize the number of infected cases or economical impact. The optimizer is a genetic algorithm and the fitness function is the simulation program. An attempt to use the surrogate function with table lookup and interpolation was reported before. The results show that the surrogate constructed by deep learning approach outperforms the interpolation based one for both total case and economical impact. The average of the absolute value of relative error is less than 0.27%, which is quite close to the intrinsic limitation of the stochastic variation of the simulation software 0.2%, and the rank coefficients are all above 0.999. The vaccination strategy recommended is still to vaccine the school age children first which is consistent with the previous studies for minimizing total infected cases. As to minimize economical impact, the priority goes to the middle schoolers then to young working adults The results are encouraging and it should be a worthy effort to use machine learning approach to explore the vast parameter space of simulation models in epidemiology.
In 2011, the Ministry of Health and Welfare of Taiwan established the National Electronic Medical Record Exchange Center (EEC) to permit the sharing of medical resources among hospitals. This system can presently exchange electronic medical records (EMRs) among hospitals, in the form of medical imaging reports, laboratory test reports, discharge summaries, outpatient records, and outpatient medication records. Hospitals can send or retrieve EMRs over the virtual private network by connecting to the EEC through a gateway. International standards should be adopted in the EEC to allow users with those standards to take advantage of this exchange service. In this study, a cloud-based EMR-exchange prototyping system was implemented on the basis of the Integrating the Healthcare Enterprise's Cross-Enterprise Document Sharing integration profile and the existing EMR exchange system. RESTful services were used to implement the proposed prototyping system on the Microsoft Azure cloud-computing platform. Four scenarios were created in Microsoft Azure to determine the feasibility and effectiveness of the proposed system. The experimental results demonstrated that the proposed system successfully completed EMR exchange under the four scenarios created in Microsoft Azure. Additional experiments were conducted to compare the efficiency of the EMR-exchanging mechanisms of the proposed system with those of the existing EEC system. The experimental results suggest that the proposed RESTful service approach is superior to the Simple Object Access Protocol method currently implemented in the EEC system, according to the irrespective response times under the four experimental scenarios.
Accurate, complete, and timely disease surveillance data are vital for disease control. We report a national scale effort to automatically extract information from electronic medical records as well as electronic laboratory systems. The extracted information is then transferred to the centers of disease control after a proper confirmation process. The coverage rates of the automated reporting systems are over 50%. Not only is the workload of surveillance greatly reduced, but also reporting is completed in near real-time. From our experiences, a system sustainable strategy, well-defined working plan, and multifaceted team coordination work effectively. Knowledge management reduces the cost to maintain the system. Training courses with hands-on practice and reference documents are useful for LOINC adoption.
Churn Jung Liau合作论文数Institute of Information Science27
Der-Tsai Lee (李德財)合作论文数Institute of Information Science & Research Center for IT Innovation, Academia Sinica3