Type 2 diabetes, one of the most prevalent chronic diseases worldwide, represents a critical challenge for healthcare systems due to its progressive course and the need for continuous management to prevent long-term complications. Although previous research has demonstrated the possibility of remission under specific conditions, a more comprehensive understanding of the underlying determinants remains essential. This study presents a scoping review on data repository technologies and proposes a Data Lakehouse-based architecture for the analysis of factors associated with type 2 diabetes remission. The proposed framework enables the ingestion, integration, storage, and advanced analysis of large-scale, heterogeneous, and multi-format data sets, facilitating the identification of patterns, correlations, and key variables linked to remission. To assess feasibility, a functional prototype was implemented, incorporating a large language model to semantically classify scientific articles, with outputs integrated into a Lakehouse infrastructure using Apache Kafka, Spark, and Iceberg. Furthermore, a relational data model is designed to enable longitudinal monitoring of clinical variables such as dietary habits, pharmacological treatments, physical activity, and medical interventions. By leveraging structured, interoperable, and traceable data, the proposed approach aims to enhance clinical decision-making, support personalized therapeutic strategies, and ultimately contribute to more effective disease management and improved patient outcomes.
This article introduces the Evidence Propagation Calculus (EPiC), an operational framework for first-order reasoning built on a simple but productive observation: familiar inference patterns such as Modus Ponens and Modus Tollens behave like the movement of evidential markers across a structured graph. Positive evidence at an antecedent propagates forward to the consequent; negative evidence at a consequent propagates backward. When both markers coexist at a node, the system is locally inconsistent but not operationally broken. To make this observation precise, EPiC grounds reasoning in a four-valued evidential domain V={N,T,F,B}, where N denotes absence of evidence, T positive evidence, F negative evidence, and B their coexistence. Each logical connective is assigned a local evidential table, and inference is treated uniformly as the progressive restriction of admissible configurations under an evidential order: inadmissible values are eliminated, minimal surviving values are selected as the next effective evidential states, and the resulting restrictions propagate across shared variables. Compound formulas are decomposed into families of local unary and binary constraints through auxiliary variables, making the propagation process explicit and structurally uniform. Within this setting, Modus Ponens, Modus Tollens, and polarity-switching negation are not postulated as primitive rules. They emerge as derived consequences of the same local table calculus. The framework distinguishes different operational routes of justification. In some cases, positive support reaches the target formula directly through successive local restrictions. In others, propagation first stabilizes the relevant components and the target occurrence is then fixed by the corresponding connective table. Consistency is not a second basic notion of justification but a distinguished property of certain justified outcomes. The article establishes local and global soundness, conservativity over the classical fragment, and a conditional adequacy result. It further develops a translation between decomposed formulas and informational graphs, with a reverse reconstruction theorem for well-formed graphs. The result is a unified operational account of first-order reasoning situated between model-theoretic and proof-theoretic approaches, in which semantics, propagation, and graphical structure are mutually supporting rather than independently layered.
The use of open-ended survey questions for data collection has increased significantly across various areas, as has the application of machine learning (ML) and natural language processing (NLP) techniques to analyze respondents’ opinions. In this study, we conducted a scoping review of 79 studies that analyze open-ended answers given in surveys. We structured our review around six main criteria: application of supervised learning, unsupervised learning, Supervised Descriptive Rule Discovery (SDRD), open-ended questions, NLP, and opinion comparison. This approach allowed us to identify the most used tasks, algorithms, and technologies in ML and NLP, revealing areas of opportunity and the main future challenges. We based our review on the methodological framework of Arksey and O’Malley and adapted PRISMA for reporting systematic reviews. Our findings suggest that most studies addressing surveys with open-ended questions were published in 2020 and 2022, predominantly focusing on research and health domains.
One of the main problems faced by database administrators for optimizing analytic workloads is fragmentation. Therefore, in recent decades, several fragmentation methods for analytical platforms have been proposed because this technique is able to improve the performance of OLAP (Online Analytical Processing) queries. In this study, we conducted an exploratory review of horizontal fragmentation methods for analytical repositories such as data warehouses, data lakes, and data lakehouses. This study presents a scoping review conducted using Arksey and O’Malley’s methodological framework and reported according to the PRISMA guidelines, covering 58 primary studies on horizontal fragmentation published from 2015 to 2025. Our analysis focuses on five aspects: (1) determining the main techniques used in horizontal fragmentation works for analytical repositories, (2) the classification of these studies, (3) the performance metrics considered when evaluating the horizontal fragmentation scheme, (4) the information type indexed by the repositories, and (5) the technologies most used by the approaches. Our findings suggest that horizontal fragmentation is a good opportunity to improve the performance of analytical workloads in most cases. The results of this scoping review will provide important guidelines for future research on horizontal fragmentation methods. In addition, the results will provide clues about the use of OLAP technologies for professionals and academics considering future directions.
Nowadays, the management of digital medical images faces increasing challenges due to the volume, diversity, and need for interoperability between systems. The DICOM standard has become the main format for storing and transmitting medical images, enabling the integration of data from studies such as MRI, CT, and USG. However, its complex structure and the increasing volume of data generated by interconnected devices, including IoT sensors, demand new strategies for efficient storage and retrieval. Clinical databases must support large volumes of heterogeneous data while ensuring fast access, availability, and secure information exchange. This review explores the integration of the DICOM standard with medical database systems, emphasizing the role of sensors as a primary source in clinical data management. The findings aim to support the development of more effective strategies for data retrieval and exchange, such as database fragmentation, to reduce query response times and improve information systems used by healthcare professionals and patients. Additionally, various data sets or benchmarks used in the analyzed studies are described. As a result, two approaches are identified as particularly noteworthy among the reviewed works, serving as a reference for future applications and technological developments in healthcare.
El aprendizaje matemático es crucial, y las habilidades cognitivas son clave para impulsarlo. En México, las deficiencias en matemáticas desde educación básica representan un reto educativo. Este artículo presenta una arquitectura escalable y modular de un videojuego serio de realidad virtual, basada en MVC, para fomentar el aprendizaje matemático y el desarrollo cognitivo personalizado. Aplicando una metodología híbrida (Design Thinking, ADDIE y el Proceso de Desarrollo de Videojuegos), la propuesta es una solución educativa que usa algoritmos de personalización, además del estilo de vida y desempeño cognitivo del usuario, para adaptar un plan de estrategias que potencie sus habilidades cognitivas y aprendizaje matemático.
The reduction of autopsies presents a challenge for medicine, impacting diagnostic accuracy and hospital decision-making. Uncertainty persists regarding the factors behind this decline and the tools best suited to analyze medical opinions. This study conducts a scoping review of contrast set mining (CSM) and contrast pattern mining (CPM) based on Arksey and O'Malley's framework and PRISMA guidelines, aiming to assess their use in studies that compare medical views on declining autopsy rates. Of the 1,292 identified articles, 39 were reviewed. The review reveals usage patterns of CSM and CPM across domains while identifying a gap in the literature concerning their application in opinion comparison. These findings underscore the need for innovative approaches to examine perceptions of autopsy reduction and open a new research avenue focused on applying CSM to the analysis of medical surveys. This result provides a deeper understanding of the factors influencing autopsy assessment and improves hospital care processes.
Subgroup discovery (SD) is a data mining technique that allows us to obtain the properties of each element given a particular population; these properties are of interest for a specific study, finding the most important or significant subgroups of the population. Also, the larger the population, the more successful the analysis and the creation of the subgroups, since, on this basis, the possibility of finding more unusual characteristics among the elements of the population is greater. The principal purpose of SD is not to obtain a predictive function, but to achieve a result that users can comprehend and interpret easily, and at the same time provide a more complete and suggestive description of the data. In this paper, we present an application of this technique to the medical field to analyze the opinions of physicians on the decreasing rates of autopsies in Mexican hospitals, utilizing five SD algorithms. The results obtained are the rules that allow for the comparison of medical opinions in three hospitals.
Age-related macular degeneration (AMD) is one of the leading causes of vision loss in elderly adults around the world and is among the main visual impairments in Mexico. The difficulty of diagnosing AMD in its early stages motivates the use of advanced deep-learning methods that offer significant potential to improve diagnostic accuracy in retinal image analysis. In recent years, Transformer architectures for computer vision, such as Vision Transformer (ViT), Swin Transformer and BERT Pre-training of Image Transformers (BEiT) have provided a novel perspective for image analysis. This study presents a comparative analysis of these architectures, applied to AMD detection, focusing on each model's capability to classify the early stages of the disease. Although the small size of medical image datasets represented a challenge, our results suggest that ViT-based architectures and their derivatives achieve significant performance in AMD detection. BEiT is particularly notable for its consistently superior performance.
Age-related macular degeneration (AMD) is a leading cause of vision loss among older adults. This study evaluates and compares the performance of five vision transformer (ViT) models (Classic ViT, Swin Transformer, BEiT, Swin Transformer V2, and SwiftFormer) in detecting dry AMD using fundus images. We used an initial dataset of 305 images, divided into training, validation, and test sets. The data set was increased by employing data augmentation techniques to enhance the models' generalization capabilities in different phases of the disease: No ADM, Mild, Moderate, and Advanced. Metrics of accuracy, precision, recall, F1-score, ROC curves, and computational efficiency were evaluated. Classic ViT and BEiT models achieved the best overall performance, excelling in accuracy (83.69% and 82.60%) and F1-score (85.47% and 86.36%), while SwiftFormer stood out in computational efficiency with a shorter inference time (14.19 ms/img), lower memory consumption (3.86 GB), and higher energy efficiency (131.46 W). This study provides evidence-based guidance for selecting ViT models for early detection of AMD, streamlining clinical diagnosis, and improving patient outcomes.
A vision system comprises several steps, with each step exerting a significant impact on the final outcome. One of these crucial steps is segmentation, which isolates the region of interest within an object and removes the background. Segmentation is vital because it enhances the quality of the isolated region, improves the accuracy of extracted features, and reduces the noise introduced by poor-quality features. Several segmentation techniques are available in the literature, each requiring one or more adjustable parameters. Selecting the most appropriate technique for segmenting a particular image dataset can be challenging. Several factors can affect the segmentation quality, including preprocessing, the chosen segmentation method, parameter fine-tuning approaches, and the characteristics of the dataset. Moreover, variations in lighting and intensity can further influence segmentation quality. At times, an expert may need to manually choose the segmentation technique and fine-tune its associated parameters. This paper presents the development of an automated algorithm for the selection of segmentation techniques and their associated parameters. The developed techniques are implemented and compared using diverse datasets, and the resulting experimental outcomes are thoroughly discussed and analyzed. The algorithm aims to streamline and simplify the process of selecting appropriate segmentation techniques, determining the required parameters, and selecting suitable pre-processing techniques.
El metaverso es una red de mundos virtuales tridimensionales, persistentes y concurrentes centrados en la conexión y la interacción social que permiten realizar actividades que no se harían en el mundo físico. En los últimos años, se ha incrementado significativamente el número de trabajos de investigación enfocados en el uso del metaverso en la educación. En ellos se resumen cómo el metaverso se puede utilizar para una amplia gama de objetivos educativos como el desarrollo de competencias digitales, la creación de experiencias de aprendizaje más inmersivas, el fomento de la creatividad y la colaboración entre estudiantes. Sin embargo, pocas veces se implementa en la educación primaria. Además, en el contexto específico de México, el reporte del uso del metaverso en la educación es limitado y poco frecuente. En este artículo se explora el metaverso integrado en Roblox como una herramienta alterna viable para mejorar el proceso de la enseñanza de matemáticas en el primer grado de educación primaria en México. El enfoque de enseñanza se centra en el conteo de números y sumas de dos dígitos, proporcionando una experiencia interactiva y enriquecedora para los estudiantes. Se espera que este trabajo de investigación proporcione una comprensión clara del potencial del metaverso en la educación, y a su vez fomente más investigaciones sobre la educación basada en el metaverso en un futuro cercano.
El conocimiento de las letras es el primer paso para lograr la capacidad para leer y escribir adecuadamente, por lo que es de suma importancia en el proceso enseñanza-aprendizaje desarrollar actividades que capten la atención y el interés del niño por aprender. Actualmente el uso de las tecnologías de la información (TICs) ofrece la posibilidad de ampliar las estrategias didácticas que se pueden utilizar para el mejoramiento del quehacer educativo. El uso de la Realidad Aumentada es una de las tecnologías de mayor aplicación en el área educativa, por lo que en este trabajo se presenta el desarrollo de una herramienta que combina el uso de Realidad Aumentada con un control de interfaz a través del movimiento de las manos para introducir a los niños en el conocimiento de las letras. Esta aplicación despliega modelos tridimensionales de las letras y objetos que las utilizan y permite a los usuarios interactuar con ellos a través de movimientos de las manos. Su aplicación con niños de preescolar demostró ser una herramienta entretenida que ayudó de forma positivo al aprendizaje de las letras
An autopsy is a widely recognized procedure to guarantee ongoing enhancements in medicine. It finds extensive application in legal, scientific, medical, and research domains. However, declining autopsy rates in hospitals constitute a worldwide concern. For example, the Regional Hospital of Rio Blanco in Veracruz, Mexico, has substantially reduced the number of autopsies at hospitals in recent years. Since there are no documented historical records of a decrease in the frequency of autopsy cases, it is crucial to establish a methodological framework to substantiate any actual trends in the data. Emerging pattern mining (EPM) allows for finding differences between classes or data sets because it builds a descriptive data model concerning some given remarkable property. Data set description has become a significant application area in various contexts in recent years. In this research study, various EPM (emerging pattern mining) algorithms were used to extract emergent patterns from a data set collected based on medical experts’ perspectives on reducing hospital autopsies. Notably, the top-performing EPM algorithms were iEPMiner, LCMine, SJEP-C, Top-k minimal SJEPs, and Tree-based JEP-C. Among these, iEPMiner and LCMine demonstrated faster performance and produced superior emergent patterns when considering metrics such as Confidence, Weighted Relative Accuracy Criteria (WRACC), False Positive Rate (FPR), and True Positive Rate (TPR).
Four out of ten Mexican adults have high cholesterol, according to the National Institute of Cardiology. Cholesterol is essential for the production of substances in our body, such as hormones and vitamin D metabolism; it is essential for the absorption of calcium and bile acids. However, excess cholesterol causes hardening and narrowing in the walls of the arteries and can form a clot that causes a heart attack or stroke. Taking into account this problem, in this chapter, we present PREDIROL: Predicting Cholesterol Saturation Levels Using Big Data, Logistic Regression, and Dissipative Particle Dynamics Simulation, which presents an approach with Big Data and mesoscopic simulation techniques with a method of Particle Dynamics DPD (Dissipative Particle Dynamics). Parallel computing using CUDA was implemented to build the DPD model that would represent the cholesterol and blood molecules. However, considering the quantity of cholesterol and blood molecules generated in 3D, which required high computing power, we opted for the 3Dmol.js library based on WebGL for rendering 3D graphics within any web browser. PREDIROL seeks to raise awareness about the care of cholesterol concentration levels since having high levels is detrimental to health, but having low concentration levels, the body does not produce cells in the body. This is a tool for preventive medicine and to improve the lifestyle of users before they develop more serious ailments and even heart attacks or strokes.
Multimedia databases store high-volume data, which causes problems in efficient information retrieval, and increases execution costs and response times of the queries. To solve this problem, data fragmentation techniques exist to improve query performance, increase information availability, and efficiently execute more operations accessing less irrelevant data. This article presents a comprehensive review of 34 methods related to hybrid fragmentation and subsequently proposes the design of a hybrid fragmentation method that adapts the scheme according to workload changes to maintain efficient retrieval of multimedia data. The proposed technologies are Java as a programming language, Java Server Faces (JSF) as a framework, MySQL and MongoDB database management systems, and NetBeans as an Integrated Development Environment (IDE), following the UWE methodology (Unified Modeling Language-based Web Engineering).
Con los avances tecnológicos, distintos proyectos desarrollaron soluciones satisfactorias a problemáticas de aprendizaje principalmente relacionadas con el área de las matemáticas y la medicina utilizando realidad aumentada e interfaces humano-máquina, siendo un área de oportunidad la aplicación de estas tecnologías en el aprendizaje de la lectoescritura. Dado que según la última prueba PISA, aplicada en 2018, México se sitúa en el nivel 2 de comprensión lectora, por debajo del promedio de los países miembros de la Organización para la Cooperación y el Desarrollo Económico (OCDE), en respuesta a esta problemática, se desarrolló una herramienta que combina la realidad aumentada e interfaces humano-máquina aplicado a la lectoescritura. Esta aplicación despliega modelos tridimensionales renderizados a través de un dispositivo con cámara, permitiendo a los usuarios interactuar con estos modelos mediante el reconocimiento de movimientos de las manos. La aplicación fue probada exitosamente en estudiantes de jardín de niños divididos en dos grupos: rezagados y adelantados en el aprendizaje de las letras. Los resultados demostraron que esta herramienta es entretenida y efectiva, proporcionando un recurso significativo para profesores y maestros en la enseñanza de la lectoescritura.
Cloud-based platforms have gained popularity over the years because they can be used for multiple purposes, from synchronizing contact information to storing and managing user fitness data. These platforms are still in constant development and, so far, most of the data they store is entered manually by users. However, more and better wearable devices are being developed that can synchronize with these platforms to feed the information automatically. Another aspect that highlights the link between wearable devices and cloud-based health platforms is the improvement in which the symptomatology and/or physical status information of users can be stored and syn-chronized in real-time, 24 h a day, in health platforms, which in turn enables the possibility of synchronizing these platforms with specialized medical software to promptly detect important variations in user symptoms. This is opening opportunities to use these platforms as support for monitoring disease symptoms and, in general, for monitoring the health of users. In this work, the characteristics and possibilities of use of four popular platforms currently available in the market are explored, which are Apple Health, Google Fit, Samsung Health, and Fitbit.
Rafael Valencia-Garcia合作论文数Universidad de Murcia2