
INTRODUCTION:Medical services routinely transmit patient data using PDF, even as FHIR emerges as the standard for structured healthcare interoperability. This mismatch reflects a broader fragmentation in digital documentation, where pragmatic workflows often outpace technical ideals. METHODS:We propose embedding FHIR bundles into PDF to enable structured data reuse without disrupting established processes. These hybrid documents can be processed via FHIR Binary endpoints, allowing downstream systems to extract, validate, and map the embedded data to interoperable resources. RESULTS:A proof-of-concept using German emergency medical services records demonstrates that vital parameters and timestamps can be transmitted as PDF while preserving machine-readable structure. CONCLUSION:This approach respects existing transport mechanisms and accommodates heterogeneous IT landscapes. By bridging legacy formats with modern standards, the method offers a scalable migration path toward interoperability-ready to deploy wherever PDF are already in use. Thus, our approach provides a migration pathway for integrating routine data into interoperable research infrastructures, enabling structured reuse without altering clinical workflows.
INTRODUCTION:In Germany, there is growing interest in linking clinical data from routine care available in data integration centres (DIC) with external data, such as medical registries. However, a suitable technical and organisational infrastructure is required for the secondary use and linkage of registry data. This paper presents the concept of a framework for the integration of routine data and registry data using the structures of a DIC. METHODS:The development of the framework followed a step-by-step approach: (1) literature research, (2) derivation of the theoretical foundations of the framework, and (3) design and modelling of the framework. RESULTS:Challenges for data linkage and matching solutions from initiatives such as the Medical Informatics Initiative (MII) or Network of University Medicine (NUM) were identified to create the theoretical basis of the framework. The initial design of the framework on increasingly detailed hierarchical levels includes functional components, processes and support units of a DIC to fulfil specific use cases from registry setup to data linkage. DISCUSSION:The framework was developed to serve as a blueprint for the setup of a DIC as a bidirectional link between registries and routine care. At this stage it represents a theoretical concept as the result of formative research which remains to be systematically evaluated, further specified and tested in pilot projects. Despite the challenges, both registries and DIC could benefit from a future practical implementation of the framework.
INTRODUCTION:Entering the content of manually filled documents in the database of an electronic clinical Trial Master File (eTMF) is a tedious and time-consuming task. METHODS:We report experiments for automatic transcription with an optical document recognition pipeline and a web-based (global) and a local multimodal large language model (LLM). RESULTS:Different approaches are best suited for different column types of the table-based documents. CONCLUSIONS:Drastic time savings compared to manual transcription are possible.
INTRODUCTION:Removal of identifying information from data used in clinical studies protects patient privacy and maintains confidentiality. It ensures compliance with laws like data protection regulations, which safeguard personal data. Medical imaging data may contain various sensitive personal data. An image is composed of pixel data and structured metadata. The de-facto standard for medical imaging, the Digital Imaging and Communication in Medicine (DICOM) format, defines more than 5,000 unique DICOM metadata items, including so-called private tags that can contain arbitrary information. Therefore, de-identification of DICOM objects may become a notable challenge, even without consideration of identifying information in pixel data such as burnt-in identifiers or visible series numbers of implants. To decide which tool best to use for defined requirements, this paper aims to define a workflow for comparative evaluations of DICOM de-identification tools. METHODS:We assessed requirements and performance indicators for DICOM de-identification tools. We employ open data and open-source tools to assess the different requirements in a comparable way, and apply the workflow on the widely used java-based de-identification tool Clinical Trials Processor (CTP) and the current state of a recently developed python-based anonymisation tool, both provided by the Radiological Society of North America (RSNA). RESULTS:The proposed workflow proved helpful in systematically comparing de-identification tools and highlighting key differences in functionality and use case. The comparison of CTP and RSNA Anonymizer (RDA) de-identification tools reveals that both tools effectively remove identifying information from datasets, but with distinct differences in their approaches and compliance with the DICOM standard. CONCLUSION:The evaluation of two DICOM de-identification tools using open data and methods revealed decisive differences in functionality and practical suitability, highlighting the value of a structured comparison. De-identification tools need to be selected based on specific use cases, especially given ongoing challenges with metadata and image-based identifiers.
INTRODUCTION:The lawful processing of health data in medical research necessitates robust mechanisms for managing patient consent and objections, aligning with national and european regulations. While the initial version of the HL7 standard Consent Management" primarily focused on opt-in scenarios, evolving legal landscapes and practical implementation challenges highlight the need for comprehensive solutions encompassing both opt-in and opt-out approaches, including withdrawals and objections. This paper details the systematic revision of the latest HL7 FHIR-based "Consent Management 2.0" standard to address these limitations. METHODS:Our methodology involved a critical assessment of the 2021 standard against three years of practical experience and emerging regulatory requirements. RESULTS:Key improvements include enhanced support for diverse document types (consent, withdrawal, refusal, objection), refined technical specifications for automated conversion of questionnaire responses into machine-readable Consent Resources, and the introduction of a novel "ResultType" category. This new category enables use-case-specific aggregation of consent information, simplifying downstream processing and reducing interpretation ambiguities. Additionally, uniform FHIR search parameters were defined, and comprehensive examples were integrated into the implementation guide. The revised standard successfully underwent the HL7 ballot process in April 2025, with early practical implementations already demonstrating its utility. CONCLUSION:This extended standard significantly enhances the interoperability and legal robustness of consent management in complex research infrastructures, fostering improved patient autonomy and trust in digital health data reuse.
INTRODUCTION:Assessing the ever-growing number of publications in evidence-based medicine by means of their risk of biases is as essential as it is challenging. This is especially true for the field of complementary and alternative medicine (CAM), a field that remains underrepresented in systematic review collections such as those by the Cochrane Review Groups. METHODS:In this work, we present CAMIH, a semantic wiki platform that offers clinicians a collaborative space to find, summarize, and discuss CAM evidence. CAMIH is built on semantic web technologies and structures information using semantic triplets. By structuring like this, CAMIH goes beyond simple data collection. Our goal is to enable a deeper understanding and organization of evidence, thereby acting as a CAM-specific supplement to existing evidence-synthesis frameworks inspired by the Cochrane methodology. RESULTS:We anticipate the implemented platform to make evidence synthesis and risk of bias assessment more efficient, but also reduce the time required to derive treatment strategies. Given its foundation in semantic web technologies, it serves both as a practical tool for clinicians and as a methodological blueprint for other research domains seeking to systematically organize gathered evidence. DISCUSSION:Given the advantages of the platform, it requires, in its current state, manual efforts to be kept up to date. However, our goal is too semi-automize this process to sustainably keep CAMIH relevant. CONCLUSION:This work provides an addition to the evidence database-landscape for the CAM field. We hope it will enable clinicians to create, discuss, and synthesize evidence while also providing a blueprint for other research areas that want to organize evidence.
INTRODUCTION:Hearing loss, affecting over 19% of the global population, is a major disability worldwide, with its prevalence expected to increase due to demographic changes. Cochlear implants (CIs) provide a crucial treatment for severe to profound sensorineural hearing loss when conventional hearing aids fail. Although technological and surgical advancements have expanded CI indications, hearing preservation (HP) after implantation remains unpredictable and varies significantly among patients. Recent studies indicate that machine learning (ML) methods could offer improved prediction. Therefore, this study aimed to evaluate the feasibility of predicting HP in potential CI users. METHODS:Clinical data from 225 CI patients (mean age: 59.9 years) implanted at Hannover Medical School (MHH) between 2009 and 2024 were retrospectively analyzed. ML models were developed and compared with baseline models such as linear regression and a mean predictor. RESULTS:Among all models, the Random Forest (RF) achieved the best predictive performance. Electrode insertion angle and age at implantation were identified as the most influential features for predicting HP, contributing 61.0% and 24.3% respectively. Despite the results of the RF model, limitations such as prediction error and a small dataset were acknowledged. CONCLUSION:The study highlights the potential of ML methods for predicting HP in CI users but underscores the need for the integration of more surgical and objective data.
INTRODUCTION:The transition from proprietary, paper-based care transition records (CTRs) to standardized digital formats like HL7 FHIR remains a significant challenge for healthcare institutions. Variability in document layouts, coupled with slow adoption of new interoperability standards, complicates efforts to digitize patient records while preserving data integrity and privacy. METHODS:This study presents a machine learning-based pipeline for automated information extraction from scanned CTRs. Synthetic training data was generated using a custom CTR generator. A Detectron2-based object detection model, integrated with LayoutParser for document structure analysis and Tesseract OCR for text recognition, was trained on this synthetic dataset. Checkbox detection was performed via an image-processing pipeline based on pixel density analysis. Extracted information was mapped to the FHIR-based PIO-ULB (Pflegeinformationsobjekt - Überleitungsbogen) format using a custom serialization tool. RESULTS:A synthetic dataset of 10,000 CTR samples was used for training and evaluation. The model achieved high values for accuracy, precision, recall, and F1-score metrics for synthetic data (97%, 98%, 95%, 97%) and showed robust performance for real-world data (85%, 86%, 83%, 85%). Lower performance on real-world data was attributed to layout variability and scanning artifacts absent from the synthetic training set. DISCUSSION:The results demonstrate the feasibility of using machine learning for automated extraction and standardization of CTRs, particularly when relying on synthetic data to overcome data privacy constraints during development. While accuracy declines with real-world document variability, the approach provides a possible interim solution for facilitating interoperability in healthcare documentation. Future work will focus on extending data generation to cover complex document layouts and integrating advanced OCR and handwriting recognition methods to further improve extraction performance.
INTRODUCTION:The heterogeneity of metadata continues to be a key challenge in the healthcare sector. The Data Dictionary Minimal Information Model (DDMIM) aims to meet the need for interoperability between different standards and data dictionaries to facilitate the exchange of metadata. OBJECTIVE:This paper presents the conception, and the development of a metadata search portal based on the DDMIM specification, designed to improve the discoverability and accessibility of health datasets and enhance interoperability. METHODS:We conducted a literature review of existing metadata repositories to select potentially relevant ones for further work. A mapping was created to transform metadata from different MDRs into the DDMIM format. In parallel, the requirements for a prototype search portal are being evaluated, which integrates metadata from various public repositories. RESULTS:The results show that a DDMIM-based search portal can effectively integrate heterogeneous metadata sources and improve the finding of health datasets. DISCUSSION:Such a portal supports the integration of heterogeneous metadata sources and ensures compliance with FAIR principles to optimize the use of health data for research and clinical applications. It is therefore of great importance to address the existing challenges in the field of medical data integration and utilization.
INTRODUCTION:The integration of Retrieval-Augmented Generation (RAG) into domain-specific systems enables context-aware and traceable information retrieval. This study explores chunking and embedding strategies for a RAG-based question-answering system tailored to administrative documents at University Hospital Halle, focusing on model selection, parameter tuning, and retrieval performance. The insights gained from this study should serve as the foundation for the future development of a Retrieval-Augmented Generation (RAG) based chatbot system that aims to facilitate access to document pool contents for hospital staff. METHODS:A corpus of 1,219 documents was preprocessed and chunked using varied parameters, including soft/hard character limits and overlaps. Eight embedding models were evaluated using Similarity Score and Maximum Marginal Relevance (MMR) retrievers. Top models Jinaai-v3 and Aari1995 were further analyzed across eight parameter configurations and ensemble retrievers using weight (w) and context (c) parameters. RESULTS:Aari1995 reached the highest Top10 score (92.3%) with stable performance across chunk sizes and retriever configurations. Jinaai-v3 showed slightly stronger Top5 (84.6%) and Top3 (76.9%) scores but with greater sensitivity to parameter variations. Ensemble retrievers improved retrieval quality for both models, particularly when tuned via w-values. The c-parameter showed negligible influence. Runtime evaluation revealed that Jinaai-v3 generated vector stores more than four times faster than Aari1995. Overall, the similarity score retriever consistently outperformed MMR, both standalone and in ensemble configurations. CONCLUSION:Chunking and embedding choices significantly affect retrieval in domain-specific RAG systems. While both Jinaai-v3 and Aari1995 were effective, they differed in stability, accuracy, and efficiency. Findings support deploying a locally executable RAG system for administrative use, guiding future optimization of chunking and parameter robustness.
INTRODUCTION:In 2024, the GeMTeX project launched the largest ever de-identification campaign for German-language clinical reports, and, as a pilot study, published GraSCCoPHI, the first de-identified German-language gold standard corpus of synthetic discharge summaries. METHODS:GeMTeX's de-identification workflow is described here - including annotation tool management and, pre-annotation experience, such as assembling and training annotation groups and the evolution of guidelines. RESULTS:We present the project's progress in the first year with respect to de-identification efforts and the challenges we faced during the rollout at six hospital sites in four German states. The refinement of the annotation guidelines became an ongoing process, often with unforeseen hurdles to overcome as we moved from testing to production. From our current internal interim corpus (9,000 documents with about 20 million tokens), we are publishing the first quantitative insights, such as the average amount of identifiable information per document, a list of confounding factors we did not anticipate at the beginning of the project, and three key lessons learned. CONCLUSION:We note that the unforeseen hurdles behave like the Pareto principle and fall into the set of less than 20% of the annotations.
INTRODUCTION:As part of the German Medical Informatics Initiative (MII) and Network University Medicine (NUM), a central research terminology service (TS) is provided by the Service Unit Terminology Services (SU-TermServ). This HL7 FHIR-based service depends on the timely and comprehensive availability of FHIR terminology resources to provide the necessary interactions for the distributed MII/NUM infrastructure. While German legislation has recently instituted a national terminology service for medical classifications and terminologies, the scope of the MII and NUM extends beyond routine patient care, encompassing the need for supplementary or specialized services and terminologies that are not commonly utilized elsewhere. METHODS:The SU-TermServ's processes are based on established FHIR principles and the recently-proposed Canonical Resources Management Infrastructure Implementation Guide, which are outlined in this paper. RESULTS:The strategy and processes implemented within the project can deliver the needed resources both to the central FHIR terminology service, but also to the local data integration centers, in a transparent and consistent fashion. The service currently provides approximately 7000 resources to users via the standardized FHIR API. CONCLUSION:The professionalized distribution and maintenance of these terminological resources and the provision of a powerful TS implementation aids both the development of the Core Data Set and the data integration centers, and ultimately biomedical researchers requesting access to this rich data.
Introduction: Detecting negations in clinical text is crucial for accurate documentation and decision-making. Methods: This study assesses open-source Large Language Models (LLMs) for detecting negations in German clinical discharge letters, comparing them to the rule-based approach (GeNeg) and human annotations. Results: While Llama 3.3 and Deepseek-R1 (70B) showed slight accuracy improvements, their high computational costs limit practicality compared to GeNeg. Llama 3.3 achieved the highest accuracy (.9670) and F1-score (.9620), outperforming all other models and slightly exceeding GeNeg in accuracy and F1-score. However, it required significantly more computational time (5.9 sec/sent) when compared to GeNeg’s processing time (.005 sec/sent). Conclusion: The study results suggest hybrid approaches combining rule-based efficiency paired with LLMs’ linguistic capabilities. In addition, future work should therefore optimize prompts and integrate LLMs with traditional methods to balance accuracy and efficiency.
INTRODUCTION:Manual ICD-10 coding of German clinical texts is time-consuming and error-prone. This project aims to develop a semi-automated pipeline for efficient coding of unstructured medical documentation. STATE OF THE ART:Existing approaches often rely on fine-tuned language models that require large datasets and perform poorly on rare codes, particularly in low-resource languages such as German. CONCEPT:The proposed system integrates Named Entity Recognition, semantic and lexical retrieval, abbreviation resolution, and context-aware normalization within a Retrieval-Augmented Generation (RAG) framework using a compact generative model. IMPLEMENTATION:The pipeline utilizes Sentence-BERT embeddings, FAISS indexing, and the Mistral-Small-Instruct model. ICD codes are assigned through a combination of semantic similarity and generative refinement among the top retrieval candidates. LESSONS LEARNED:Major sources of error were found in semantic retrieval and diagnosis normalization. Future improvements should focus on domain-specific German embeddings, more robust abbreviation handling, and enhanced context-aware prompting to increase accuracy and usability in clinical environments.
INTRODUCTION:Mitotic figure (MF) density has been established as a key biomarker for certain tumors. Recently, the differentiation between atypical MFs (AMF) and normal MFs (NMFs) has gained increased interest in research, as AMFs density could be an independent biomarker. This results in the challenge of finding an automated, deterministic way to differentiate between AMFs and NMFs. METHODS:In this study, the AUCMEDI deep learning framework is applied to the recently published AMi-Br dataset to get a first bearing on the complexity of the task at hand. The dataset includes eight mitotic subclasses derived from breast cancer samples, four for NMFs and four for AMF. Using a patient-level cross- validation strategy and a ConvNeXt-based ensemble, we trained and evaluated an eight-class subtype classification model. RESULTS:Our results show high specificity across all classes (≥ 90%), but sensitivity varies significantly between mitotic subclasses (0-82%), reflecting the dataset's inherent challenges. The mean AUC of 85.90% outperforms the binary classification baseline (69.8%). CONCLUSION:The results highlight the promise of progress in subclass-level mitotic analysis while pointing to areas for further model refinement.
INTRODUCTION:The growing number of connected medical devices in hospitals poses serious operational technology (OT) security challenges. Effective countermeasures require a structured analysis of the communication interfaces and security configurations of individual devices. STATE OF THE ART:Although Manufacturer Disclosure Statements for Medical Device Security (MDS2, Version 2019) offer relevant information, they are rarely integrated into cybersecurity workflows. Existing studies are limited in scope and lack scalable methodologies for systematic evaluation. CONCEPT:This study analyzed 209 MDS2 documents and 161 security white papers to extract structured information on ports, protocols, and protective measures. Over 52,000 question-answer pairs were converted into a machine-readable format using customized parsing and validation routines. The aim was to establish whether this dataset could inform risk assessments and future applications involving Large Language Models (LLMs). IMPLEMENTATION:The analysis revealed 367 distinct ports, including common protocols such as HTTPS (443), DICOM (104), and RDP (3389), as well as vendor-specific proprietary ports. Approximately 40% of the devices used over 20 ports, indicating a broad attack surface. OCR errors and inconsistent formatting required manual corrections. A consolidated dataset was developed to support clustering, comparison across vendors and versions, and preparation for downstream LLM use, particularly via structured SBOM and configuration data. LESSONS LEARNED:Although no model training was conducted, the structured dataset can support AI-based OT security workflows. The findings highlight the critical need for up-to-date, machine-readable manufacturer data in standardized formats and schemas. Such information could greatly enhance the automation, comparability, and scalability of hospital cybersecurity measures.
INTRODUCTION:EyeMatics (Eye disease "treated" with interoperable medical informatics) is a clinical use case in Germany's Medical Informatics Initiative (MII). The objective of EyeMatics is to improve the understanding of the treatment effects of intravitreal injections, the most frequent procedure to treat retinal diseases. To achieve this, as part of the efforts, medical data of multiple hospital sites must be integrated into a local FHIR repositories using a common set of FHIR profiles. METHODS:A panel of medical and technical experts was formed to create a project-specific core data set. The core data set is compiled from different preliminary work like previously existing data models, the data donation of a synthetic patient encounter scenario, screenshots of UI forms used at the local sites and relevant MII profiles. RESULTS:Complex profiling and custom code systems were required to represent central ophthalmic concepts such as visual acuity using FHIR. A common pattern for representing eye laterality in FHIR Observation resources was established. The resulting profiles are available in a public GitHub repository. CONCLUSION:While our FHIR profiles will be sufficient for EyeMatics, our experience shows that the clinical complexity of basic ophthalmological observations requires some deviations from standard FHIR modelling patterns and would be limited by gaps in terminology if deployed elsewhere. Our experience indicates that project-specific FHIR profile development is possible within a short time frame but may harbor risks of contributing to a growing number of fragmented implementations unless closely coordinated or based on common data models. To counteract this, future ophthalmic MII extension modules should attempt to integrate both profiles described here and international standard data models for ophthalmology.
INTRODUCTION:In the context of precision oncology, patients often have complex conditions that require treatment based on specific and up-to-date knowledge of guidelines and research. This entails considerable effort when preparing such cases for molecular tumor boards (MTBs). Large language models (LLMs) could help to lower this burden if they could provide such information quickly and precisely on demand. Since out-of-the-box LLMs are not specialized for clinical contexts, this work aims to investigate their usefulness for answering questions arising during MTB preparation. As such questions can contain sensitive data, we evaluated medium-scale models suitable for running on-premise using consumer grade hardware. METHODS:Three recent LLMs to be tested were selected based on established benchmarks and unique characteristics like reasoning capability. Exemplary questions related to MTBs were collected from domain experts. Six of those were selected for the LLMs to generate responses to. Response quality and correctness was evaluated by experts using a questionnaire. RESULTS:Out of 60 contacted domain experts, 5 fully completed the survey, with another 5 completing it partially. The evaluation revealed a modest overall performance. Our findings identified significant issues, where a large percentage of answers contained outdated or incomplete information, as well as factual errors. Additionally, a high discordance between evaluators regarding correctness and varying rater confidence has been observed. CONCLUSION:Our results seem to be indicating that medium-scale LLMs are currently insufficiently reliable for use in precision oncology. Common issues include outdated information and confident presentation of misinformation, which indicates a gap between benchmark- and real-world performance. Future research should focus on mitigating limitations with advanced techniques such as Retrieval-Augmented-Generation (RAG), web search capability or advanced prompting, while prioritizing patient safety.
INTRODUCTION:The medical care of patients with rare diseases is a cross-border concern across the EU. This is also reflected in the usage statistics of the SE-ATLAS, where most access occurs via browser languages set to German, English, French, or Polish. The SE-ATLAS website provides information on healthcare services and patient organisations for rare diseases in Germany. As SE-ATLAS currently offers its content almost exclusively in German, non-German-speaking users may encounter language barriers. Against this background, this paper explores whether common machine translation systems can translate medical texts into other languages at a reasonable level of quality. METHODS:For this purpose, the translation systems DeepL, ChatGPT, and Google Translate were analysed. Translation quality was assessed using the standardised metrics BLEU, METEOR, and COMET. In contrast to subjective human assessments, these automated metrics allow for objective and reproducible evaluation. The analysis focused on machine-generated translations of German-language texts from the OPUS corpus into English, French, and Polish, each compared against existing reference translations. RESULTS:BLEU scores were generally lower than those of the other metrics, whereas METEOR and COMET indicated moderate to high translation quality. Translations into English were consistently rated higher than those into French and Polish. CONCLUSION:As the three analysed translation systems showed hardly any statistically significant differences in translation quality and all delivered acceptable results, further criteria should be taken into account when choosing an appropriate system. These include factors such as data protection, cost-efficiency, and ease of integration.
INTRODUCTION:The COVID-19 pandemic exposed both direct and collateral health impacts especially on vulnerable populations, underscoring the need for more targeted and equitable crisis response strategies. Health-related dashboards could support better information sharing, research, and care delivery, but current dashboards often fail to address the needs of vulnerable groups. This study aimed to assess expert perspectives on key aspects of a new crisis response health dashboard to protect vulnerable populations intended to be used by medical professionals and affected persons. METHODS:A prospective, participatory workshop was conducted with a multidisciplinary group of researchers from the COLLPAN consortium (n = 20). The workshop employed the 6-3-5 method developed by Bernd Rohrbach. Data were collected through semi-structured textual responses, transcribed, and analyzed using a thematic analysis with MAXQDA (version 24.5.0). RESULTS:The envisioned dashboard targets a wide range of users-including patients, healthcare professionals, researchers, policymakers-with particular attention to those with limited digital literacy. Core functionalities include data visualization, management, analysis, networking, and administrative support, enhanced by multilingual, app-based, and artificial intelligence assisted features. The proposed content encompasses resource availability, epidemiological indicators, disease burden with regional and international comparisons, and the inclusion of individual risk profiling. Data sources include health, administrative, socioeconomic, and demographic datasets. The limitations identified relate to technical, regulatory, user-centered, definitional, and resource-based challenges. DISCUSSION AND CONCLUSION:The study highlights the importance of inclusive, user-centered design in the development of health-related dashboards, particularly to address the needs of vulnerable populations. By involving diverse stakeholders at an early stage and strengthening the technical foundations, digital solutions have the potential to reduce health inequalities rather than reinforcing them.