We study the discrete point-based fragment of Ensemble Logic () over the natural numbers, a logic combining displacement φ_u, bounded metric modalities _t and _t with additive bounds, Boolean connectives, and first-order quantification over . Motivated by the need for a unified symbolic layer for biomedical knowledge with temporal, spatial, genomic, and multimodal metric content, we develop the foundational discrete theory of the formalism. We give syntax and semantics, and prove a forward embedding of () over a finite proposition set 𝒫 into first-order monadic Presburger arithmetic (,<,+;𝒫). This embedding yields the analytical upper bounds, while a reduction from nondeterministic two-counter machines with recurring control states proves that satisfiability is Σ^1_1-complete and validity is dually Π^1_1-complete. Expressively, () strictly extends the star-free ω-languages and is incomparable with the ω-regular languages: it defines the non-ω-regular counting language {a^mb^mc^md^m| m≥ 1}^ω, whereas a delimited parity language remains outside the logic by classical Presburger-arithmetic lower bounds. On the proof-theoretic side, we present a sound Hilbert system and establish completeness relative to monadic Presburger validity as oracle, noting that completeness relative to plain Presburger arithmetic is impossible.
The BRAIN Initiative Cell Atlas Network (BICAN) is generating large-scale multimodal datasets to profile cell types in the human, non-human primate, and mouse brain. The diversity of single-cell and spatial transcriptomic and epigenomic assays, combined with varied experimental contexts, multiple data-generating laboratories and distributed infrastructure, poses substantial challenges for data integration and reuse in BICAN. To address this, we implemented a standards framework that enables layered integration of these data into knowledge-ready products for interoperable brain cell atlases. This framework organizes data based on three progressively structured layers. First, we introduced an assay-agnostic modeling layer that unifies the representation of single-cell and spatial omics data using a common set of biological entities and processes assessed by diverse experimental techniques. Second, we implemented harmonized metadata standards that capture key experimental features linked to biospecimen provenance across heterogeneous tissue sources, species, and preparations, supporting integration and validation while minimizing burden on data contributors. Third, we present an extensible representation for data-driven cell type taxonomies that integrates molecular data with annotations, ontology mappings, and evidence. Together, these contributions represent an end-to-end framework that transforms heterogeneous datasets into structured, interoperable resources that support broad community reuse via mapping algorithms, annotation systems, and visualization platforms. This approach links biospecimen provenance with cell-level outputs and embeds these in a standardized taxonomy format, enabling downstream applications such as cross-dataset integration, reference mapping, and knowledge-driven analysis. More broadly, our work demonstrates a generalizable strategy for enabling an efficient data-to-knowledge pipeline in a large-scale consortium setting.
Seizure frequency is a key outcome measure in epilepsy management and a critical factor in assessing Sudden Unexpected Death in Epilepsy (SUDEP) risk. However, this information is often embedded in unstructured narratives within Epilepsy Monitoring Unit (EMU) reports, limiting its utility for large-scale data-driven research. In this study, we present a framework for extracting structured seizure frequency information from EMU reports using large language models (LLMs). We curated and annotated 1,242 seizure frequency instances from relevant sections of 800 reports from six institutions, capturing attributes containing temporal information, quantitative details, and normalized seizure semiology ontology terms. We fine-tuned and evaluated five LLMs (Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, Qwen2.5-14B-Instruct-1M, MedGemma-27B-Instruct, and GPT-4.1) on this dataset. Performance was evaluated at the frequency and attribute levels, with attribute-level results reported both per attribute and as an overall micro-averaged score across attributes. Among these, GPT-4.1 achieved the highest frequency-level F1-score of 79.09% and overall attribute-level F1-score of 95.85%, followed by MedGemma that achieved F1-scores of 74.12% and 94.24% at the frequency-level and overall attribute-level, respectively. All fine-tuned models demonstrated strong per-attribute performance, with GPT-4.1 consistently outperforming the others. Performance was highest for event, category, and duration attributes, and comparatively lower for temporal and quantitative attributes. These findings highlight the potential of fine-tuned generative LLMs to extract structured frequency information from relevant clinical narrative sections in EMU reports, enabling scalable seizure frequency-based analyses such as SUDEP risk assessment.
Large sleep-study repositories contain rich time-stamped physiological annotations, but cohort discovery is still commonly implemented as ad hoc scripts or scalar-index filters. We present a logic-based temporal cohort discovery engine that brings formal semantics, model checking, specialized indexing, and empirical evaluation into a unified biomedical informatics framework. We adopt Rational Ensemble Logic (QEL) as a dense-time formal foundation for sleep-data querying and represent each annotated polysomnogram as a Biomedical Event Structure Temporal Model (BEST), a finite mapping from event labels to non-overlapping rational interval ensembles. Cohort discovery is formulated as model checking of QEL formulas over BEST databases. We organize common sleep-research requirements into three reusable temporal query patterns: single-event retrieval, dual-event temporal pattern matching, and event data extraction. The prototype cohort discovery engine was implemented in Python with in-memory and MongoDB-backed execution modes and evaluated on synthetic interval datasets containing up to 90 million intervals and on real-world National Sleep Research Resource annotations from the Cleveland Children's Sleep and Health Study (CCSHS) containing 515 subjects, 202,587 intervals, 23 event labels. 2DFC constructs indexes in linear space and linear build time, reducing build time at 90 million intervals from 11,549 s with RTFC and 23,902 seconds with 2DRT to 3,655 seconds. On CCSHS, cohort-selection queries executed at sub-second latency at native scale and under 45 seconds at 1,000 times scale. This work is a part of the Symbolic Biomedicine program championed by the corresponding author.
BACKGROUND AND OBJECTIVES:Severe hypoxemia after generalized convulsive seizures (GCSs) can trigger neural injury and is a potential biomarker for sudden unexpected death in epilepsy (SUDEP). Some degree of variability in interbreath interval is normal, but increased variability may suggest dysfunctional breathing control and may be associated with severe postictal hypoxemia. We evaluated the relationship between interictal breathing variability and severity and duration of hypoxemia after GCS. METHODS:We prospectively collected video-EEG, respiratory flow and effort, pulse oximetry (SpO2), and ECG from people with epilepsy (PWE). Measures of interictal interbreath interval variability (coefficient of variation, root mean square of successive differences [RMSSD], and long-term [SD-2] variability from Poincaré plots) from interictal asleep and awake periods and other relevant variables were evaluated as covariates for primary outcomes: (1) hypoxemia duration (length of time SpO2 <90%) and (2) severity of hypoxemia (SpO2 nadir), and secondary outcome: occurrence of combined prolonged and pronounced hypoxemia. Univariable and multivariable models were created for primary outcomes, but only univariable analyses were performed for the secondary outcome. RESULTS:Of 2,506 participants enrolled, 257 (141 [∼54%] female; mean age = 37.9 years) had ≥1 GCS, but only 152 GCS in 123 had evaluable respiratory data. Multivariable model for hypoxemia duration showed that SpO2 nadir (mean ratio [MR] = 0.88, 95% CI 0.81-0.96, p = 0.002) and SD-2 of the awake interbreath interval (MR = 1.06, 95% CI 1.01-1.13, p = 0.04) were significantly associated. RMSSD of the non-REM interbreath interval (mean difference = -5.01, 95% CI -8.10 to -1.93, p = 0.002) was the only variable significantly associated with hypoxemia severity after controlling for duration of postictal generalized EEG suppression, SD-2 of the awake interbreath interval, and body mass index. Univariable analyses for combined prolonged and pronounced hypoxemia showed SD-2 of the awake interbreath interval, temporal lobe epilepsy, ictal central apnea, and a shorter tonic phase duration were significantly associated. DISCUSSION:Measures of interictal respiratory variability are associated with severe and prolonged hypoxemia after GCS. Increased interictal respiratory variability suggests baseline respiratory dysregulation in some PWE and may be a surrogate for SUDEP risk.
This book analyzes auditing, curation, and enhancement of biomedical ontologies, revolving around a novel formal method called "formal concept analysis"
The American Academy of Sleep Medicine (AASM) Manual is the clinical standard for polysomnography (PSG) scoring, but its narrative rules can admit multiple reasonable interpretations, contributing to inter-scorer variability and implementation differences across studies and software systems. We present a formal framework for translating sleep-scoring rules into Rational Ensemble Logic (QEL), a dense-time (i.e., a continuous, rational-valued timeline rather than discrete steps) formalism that combines first-order quantification with metric temporal operators. Using an extraction-and-compilation procedure, we identified 18 unique atomic propositions and derived 12 final specifications corresponding to clinically scoreable AASM events. Back-translation of QEL specifications into clinician-facing language retained high semantic fidelity to the original scoring narratives (embedding cosine similarity: 79.3, 95
The goal of the NIH BRAIN Initiative Cell Atlas Network’s (BICAN) is to create a comprehensive cell census and atlas of the human brain to accelerate progress in neuroscience and translational medicine. This goal requires not only large-scale genomic and imaging data generation but also accurate spatial alignment of brain specimens to a common coordinate framework. To address this need, we present BRAVE, a web-based Neuroanatomy-anchored Brain Region Annotation, Visualization, and Exploration engine for BICAN. BRAVE enables interactive annotation of brain specimens using the Allen Brain Atlas, links data to the Whole Brain Spatial Ontology, and provides 2D and 3D visualization of annotated regions. Its visualization capabilities allow researchers to explore complex spatial relationships, examine the anatomical distribution of generated datasets, and access associated data across the entire brain. By combining ontology-driven annotation with interactive visualization, BRAVE supports reproducible, scalable mapping of BICAN data, facilitates integrative analyses and cross-study comparisons, and advances the standardization and accessibility of large-scale neuroscience datasets.
The reliance on unstructured free text for documenting clinical trial protocols creates a significant barrier to automated reasoning, cohort discovery, and trial simulation. The lack of formal structure obscures critical temporal phenotypes, such as dynamic eligibility criteria and event timing constraints. Although Temporal Ensemble Logic (TEL) offers an expressive framework for modeling these elements, manual encoding remains a prohibitive bottleneck. We introduce the CT-TEL workflow: a scalable pipeline leveraging Large Language Models (LLMs) to translate narrative clinical protocols into TEL formulas. We applied CT-TEL to generate logical models for 23 real-world trials from ClinicalTrials.gov. We evaluated translation fidelity via a back-translation approach, using LLMs to convert TEL formulas back into natural language and measuring semantic similarity against source texts. The resulting semantic retention suggests that LLMs may offer a pathway for mapping informal protocols to computable logic, providing preliminary evidence toward scalable clinical trial emulation within the emerging "Symbolic Biomedicine" paradigm championed by the corresponding author.
Rationale:Conventional measures of obstructive sleep apnea severity, particularly the apnea-hypopnea index, do not adequately capture event-level neurophysiologic responses to respiratory events. Whether post-apnea/hypopnea arousal dynamics provide prognostic information beyond established metrics remains unknown. Objectives:To determine whether post-apnea/hypopnea arousal dynamics are associated with all-cause and cardiovascular mortality. Methods:We conducted a retrospective analysis of in-home polysomnography data from 8,053 adults across four community-based cohorts. Peak time (PT; latency to maximal arousal probability), peak height (PH; maximal arousal probability), and area under the curve (AUC; cumulative arousal probability) were derived from peri-stimulus time histograms aligned to event termination. Associations with mortality were examined using multivariable Cox models and random-effects meta-analysis. Measurements and Main Results:PT, but not PH or AUC, was associated with mortality. In pooled analyses, each 1-second delay in PT was associated with higher all-cause mortality in males (hazard ratio [HR], 1.04; 95% confidence interval [CI], 1.02-1.06) and females (HR, 1.03; 95% CI, 1.00-1.06). For cardiovascular mortality, each 1-second delay in PT was associated with higher risk in males (HR, 1.05; 95% CI, 1.02-1.08) but not females (HR, 1.04; 95% CI, 0.99-1.10). Associations were driven primarily by non-rapid eye movement sleep and remained materially unchanged after additional adjustment for apnea-hypopnea index, arousal index, and hypoxic burden. Conclusions:Delayed arousal timing after apnea/hypopnea termination was associated with increased mortality risk independent of conventional measures of obstructive sleep apnea severity. Event-level arousal timing may provide prognostic information beyond count-based and hypoxemia-based metrics.
Mapping epilepsy variants in ClinVar to standard disease classifications remains challenging due to incomplete phenotype information. We integrated the Mondo Disease Ontology (Mondo) with the International League Against Epilepsy (ILAE) Classification to systematically identify, organize, and analyze age-dependent epilepsy-associated variants from the worlds largest public variant repository. Starting from over three million ClinVar entries, we filtered monogenic disorders tagged with 421 Mondo terms representing neonatal/infantile, childhood, or variable-onset epilepsy syndromes. Gene-disease validity was further refined using multiple expert-curated lists, resulting in 71,942 unique variant submissions in 171 genes. Approximately 70% of pathogenic and uncertain variants clustered in 50 genes, highlighting the concentration of clinically relevant findings in a relatively small subset of known epilepsy genes. SCN1A, GRIN2A, and DEPDC5 stood out as top contributors to neonatal/infantile, childhood, and variable-onset syndromes, respectively. Notably, two-thirds of genes were specific to a single age group, but the remaining one-third spanned multiple syndromes, reflecting both phenotypic and genetic heterogeneity. We also observed substantial overlap with neurodevelopmental and autism-related genes, especially among early-onset epilepsies, emphasizing shared molecular pathways. Our Mondo-driven approach demonstrates the feasibility of large-scale, ontology-based curation of ClinVar for domain-specific analyses, offering greater clinical precision than traditional text-based queries. This resource can guide epilepsy-focused gene panel design, inform new classification updates, and pave the way for age-targeted therapeutics. By bridging the gap between extensive genomic repositories and detailed clinical ontologies, we provide a robust, standardized framework to accelerate progress in precision epilepsy genetics. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study did not receive any funding. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: All the used sources containing human data are publicly available. We specifically used ClinVar data (https://www.ncbi.nlm.nih.gov/clinvar/), SFARI Gene database (https://gene.sfari.org/), and SySNDD database (https://sysndd.dbmr.unibe.ch/). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data produced in the present study are available in the manuscript text and supplemental files. Any additional data is available upon reasonable request to the authors.
COVID-SPHERE is a self-service web application designed to advance clinical research informatics by facilitating secondary use of Electronic Health Records (EHR) for COVID-19 research. The system employs a flexible EHR concept framework that defines hierarchical concepts and ontologies, enabling clinical researchers to build complex temporal queries through an intuitive, single-click interface without requiring database expertise. Our method dynamically generates MongoDB queries in real-time and offers interactive, faceted visualizations to analyze longitudinal patient activities and integrated health records, supporting both individual patient analysis and population-level research. Hosted on a server managing over 5 TB of data encompassing 30 billion health records spanning 15 years from more than 8.8 million patients, this work demonstrated its generalizability by supporting multiple published research studies investigating various COVID related research topics on epidemiology, treatment outcomes, and long-term sequelae since November 2020. By simplifying the cohort discovery process, COVID-SPHERE reduces the informatics barriers between researchers and EHR data, enhancing the efficiency of clinical and translational research while promoting data-driven insights for COVID-19 surveillance and intervention. Its architecture is applicable to other large-scale clinical research data warehouses, offering a model for future public health informatics systems.
One of the desirable properties of the resulting graph structure is that the subsumption relationship (is-a hierarchy) should form a lattice [81]. There are in general two types of lattice-based approaches to ontology quality assurance. One involves the direct application of Formal Concept Analysis (FCA [82]), mostly for auditing semantic completeness or missing concepts [76]. The second involves the extraction of lattice-violating fragments (discussed in Sects. 4.1 and 4.2), or non-lattice fragments (note that non-lattice fragments and non-lattice subgraphs are used interchangeably in this book), which represent violations of the FCA principle that systematic engineering approaches for constructing concept hierarchies always result in order structures that are lattices in the sense of lattice theory [82]. This non-lattice approach for ontology quality assurance involves the extraction of graph substructures (i.e., sub-orders) that violate the lattice property, which states that any two concept nodes have at most one minimal shared (common) ancestor and at most one maximal shared descendant.
Apnea-hypopnea index (AHI) considers only the rate of sleep apnea events, providing an incomplete view of sleep apnea and its adverse health outcomes. In contrast, this study considers temporal interactions among events from different organ systems. We analyzed novel temporal relationships between sleep apnea (from the respiratory system) and arousal (from the brain), thereby generalizing the concept of respiratory effort-related arousal (RERA). We studied 3,502 participants in the Sleep Heart Health Study (SHHS) Visit 1. The temporal relationships between apnea and arousal were characterized using peri-event time histograms (PETHs) of arousals time-aligned to the end of sleep apnea. Binomial tests were utilized to identify the significant time regions compared with the null distribution independent of sleep apnea. Cox regression models adjusted for common covariates were used to test the associations between the characteristics of the PETHs vs. CVD mortality and all-cause mortality. We also compared the hazard ratio (HR) to that of AHI and hypoxic burden. For participants with arousals significantly enriched at certain time regions of apnea, the peak time of the PETH curves (i.e., time from the end of apnea to the following most arousal-intense time point) showed associations with both all-cause and CVD mortality. In females, the adjusted HRs per one-second increase in peak time were 1.04 [95% CI: 1.01–1.08, p=0.013] for all-cause mortality and 1.06 [1.01–1.12, p=0.019] for CVD mortality, which were more significant than those of AHI for all-cause mortality [HR=1.00 (1.00-1.01), p=0.373] and CVD mortality [HR=1.00 (0.99-1.01), p=0.991], and those of hypoxic burden (log-transformed and standardized) for all-cause mortality [HR=1.05 (0.94-1.18), p=0.344], and CVD mortality [HR = 0.97 (0.80-1.19), p=0.797]. Similar patterns were observed in males, with stronger associations than those for conventional indices. These preliminary results suggest that, at the event level, temporal relationships between sleep apneas and arousals are important predictors for all-cause and CVD mortality. Future studies will extend the study-population to larger-scale sleep cohorts and incorporate additional sleep microstructure events, forming a network of multiple organ systems dynamically interacting with each other during sleep. NIH (R01NS126690,R01NS116287,R01NS102190,R01NS102574, R01NS107291, RF1AG064312, RF1NS120947, R01AG073410, R01HL161253),AASM (RF1AG064312), NIA (R21AG085495,R01AG083836), NSF (2014431)
In this chapter, we introduce lexical approaches to systematically detect potential subtype (or is-a relation) inconsistencies that can be generally applied to biomedical terminologies. These approaches utilize lexical information in concept names to uncover and suggest fixes to ontology defects.
BACKGROUND:Sudden unexpected death in epilepsy (SUDEP) is the leading cause of epilepsy-related mortality. Generalised-particularly nocturnal-convulsive seizures, longstanding epilepsy, and solitary living have been identified retrospectively as risk factors. No definitive electroclinical biomarkers have been prospectively ascertained. This study aimed to identify SUDEP risk markers using multimodality data with long-term follow-up. METHODS:This prospective, multicentre, observational cohort study, conducted at nine centres (eight in the USA and one in the UK), recruited children and adults with epilepsy who were undergoing prolonged video-electroencephalographic (EEG) monitoring. Inclusion criteria were diagnosis of epilepsy by an epilepsy specialist, with or without drug resistance; age older than 2 months; admission to the epilepsy monitoring unit of a participating centre, with video-EEG monitoring; and completion of at least one 6-month follow-up. Demographic, electroclinical, and cardiorespiratory data were collected at baseline. Participants were followed up long term through routine clinic visits, review of electronic health records, and telephone interviews to collect information about seizure frequency, medication status, and mortality. The primary endpoint was time to SUDEP. Cox proportional hazards models were used to assess significant risk factors. FINDINGS:Between Sept 17, 2011, and Dec, 30, 2021, 2632 children and adults with epilepsy were enrolled in this study; 164 were lost to follow-up. 38 (1·54%) of 2468 participants died from SUDEP (12 definite, 18 probable, and eight possible SUDEP cases) and two had near-SUDEP events. Incident SUDEP mortality rate was 4·76 (95% CI 3·37-6·53) cases per 1000 person-years, from a cohort of 7982 person-years. Living alone (hazard ratio 7·62, 95% CI 3·94-14·71), three or more generalised convulsive seizures in the previous year (3·1, 1·64-5·87]), longer ictal central apnoea (1·11, 1·05-1·18), and longer postictal central apnoea (1·32, 1·14-1·54]) were significant predictors of increased SUDEP risk. In a subanalysis excluding possible and near-SUDEP cases, longer ictal central apnoea was not significant. INTERPRETATION:This study shows an association between premortem peri-ictal apnoea and increased SUDEP risk. Cardiorespiratory monitoring during seizures might benefit assessments of epilepsy mortality risk. Together with solitary living and convulsive seizure frequency, peri-ictal apnoea (>14 s for postictal central apnoea and >17 s for ictal central apnoea) could inform the development of a validatable SUDEP risk index. FUNDING:US National Institutes of Health.
Ensuring the completeness of IS-A relations in SNOMED CT is crucial for maintaining its accuracy in clinical applications. In this study, we propose a hybrid approach leveraging non-lattice subgraphs and pre-trained language models (PLMs) to identify missing IS-A relations in SNOMED CT. We fine-tuned four BERT-based models: BERT, DistillBERT, DeBERTa, and BioClinicalBERT, and four generative large language models (LLMs): BioMistral, Llama3, Gemma2, and Phi-4. Missing IS-A relations were identified through consensus predictions by all eight models. De-BERTa achieved the best performance (precision: 0.96, recall: 0.97, F1-score: 0.965) for IS-A relation prediction. Our approach identified 678 potential missing IS-A relations in SNOMED CT (March 2023 US Edition), of which 100 randomly selected cases were manually reviewed by a domain expert, confirming 93 as valid (93% precision). These results demonstrate the effectiveness of fine-tuned PLMs in detecting missing IS-A relations within non-lattice subgraphs, offering a promising avenue for improving SNOMED CT's quality.
This chapter presents a variety of approaches that leverage non-lattice substructures (i.e. non-lattice fragments or non-lattice subgraphs) for ontological analyses.
Sudden Unexpected Death in Epilepsy (SUDEP) is a major cause of death for epilepsy patients having uncontrolled seizures. Understanding the complex neural circuits within the central nervous system is crucial for understanding the mechanisms underlying cardiorespiratory regulation, particularly in the context of SUDEP. This study explores the potential of GPT-4o, a cutting-edge language model, to automate the extraction of neural projections from scientific literature. We developed prompts to extract neuroscientific structures, extract projections, and perform synonym harmonization. Applying the approach to four neuroscientific articles, the method extracted 205 projections. A random sample of 100 projections identified was handed over to a domain expert for review where 95 were found to be correct. Therefore, GPT-4o was determined to be accurate in parsing complex scientific texts in extracting neural projections. Future work will involve extracting additional entities like techniques and species information for the projections identified.