
В данной работе рассматривается проектирование, разработка и испытания прототипа прикладной системы с применением технологий искусственного интеллекта для автоматизированного анализа документации атомной отрасли, включая проверку соответствия регламентирующим документам и нормативным требованиям. Система использует технологии больших языковых моделей (LLM) и поисково-дополненной генерации (RAG). Рассмотрены варианты применения данного стека технологий для использования специализированных доменов знаний (массива нормативной документации и технических инструкций атомной отрасли) для качественной оперативной экспертизы рабочих документов с применением искусственного интеллекта. Повышение качества и скорости квалифицированных экспертиз для сложных высокотехнологичных отраслей может зачастую иметь критическую важность для реализации серьезных проектов. В связи с наличием значимой чувствительной информации технического и экономического характера в отрасли, важным является организация работ без внешнего доступе к анализируемой проектной документации. В работе рассмотрены способы организации работ с применением как облачных, так и внутриконтурных LLM. Разработан прототип инструмента оркестрации разнородных моделей с доступом в закрытый контур компании через VPN-каналы. Выполнено сравнение решения задачи интеллектуальным ассистентом эксперта в облачном и внутриконтурном режиме. Описан разработанный прототип системы, функционирующей без подключения к внешним сетям, обеспечивающей высокий уровень конфиденциальности и импортозамещения. Прототип обеспечивает разделение ролей предметного пользователя (эксперта) и администратора, предоставляя возможность защищённого взаимодействия с системой. Разработанный интеллектуальный ассистент реализует продвинутый контекстный поиск и аналитику/экспертизу проектной документации на основании информации из специализированных доменов знаний ограниченного доступа на базе применения технологии RAG. Реализация решения выполнена на базе цифрового стенда-полигона Высшей инженерной школы МИФИ и рассчитана на применение для образовательных задач, учебно-практических кейсах индустриальных партнеров и для выполнения работ в интересах заказчиков из производственной сферы, в первую очередь – предприятий Росатома (включая, но не ограничиваясь). This paper discusses the design, development and testing of a prototype of an application system using artificial intelligence technologies for automated analysis of nuclear industry documentation, including verification of compliance with regulatory documents and regulatory requirements. The system uses large language model (LLM) and search-augmented generation (RAG) technologies. Options for using this technology stack for the use of specialized knowledge domains (an array of regulatory documents and technical instructions of the nuclear industry) for high-quality operational examination of working documents using artificial intelligence are considered. Improving the quality and speed of qualified expertise for complex high-tech industries can often be critical for the implementation of serious projects. Due to the availability of significant sensitive information of a technical and economic nature in the industry, it is important to organize work without external access to the analyzed design documentation. The paper discusses methods of organizing work using both cloud and in-loop LLMs. A prototype of a tool for orchestrating heterogeneous models with access to the company's closed loop via VPN channels has been developed. A comparison of the solution of the problem by the intelligent assistant of the expert in the cloud and intra-contour modes is carried out. The developed prototype of a system operating without connection to external networks, providing a high level of confidentiality and import substitution, is described. The prototype provides for the separation of the roles of the subject user (expert) and the administrator, providing the possibility of secure interaction with the system. The developed intelligent assistant implements advanced contextual search and analytics/examination of project documentation based on information from specialized knowledge domains of limited access based on the use of RAG technology. The solution was implemented on the basis of a digital test site of the Higher School of Engineering MEPhI and is designed for use for educational tasks, educational and practical cases of industrial partners and for performing work in the interests of customers from the production sector, primarily Rosatom enterprises (including, but not limited to).
Данный номер журнала International Journal of Open Information Technologies является очередной шагом уже ставшего традиционным и систематическим взаимодействия между журналом INJOIT и Высшей инжиниринговой школой НИЯУ МИФИ. В настоящий номер журнала вошли работы, выполненных в 2024-2025 учебном году преподавателями, сотрудниками, аспирантами и магистрантами ВИШ НИЯУ МИФИ по тематикам системной инженерии, цифрового инжиниринга и инженерии данных. This issue of the International Journal of Open Information Technologies is the next step in the already traditional and systematic interaction between the INJOIT journal and the Higher Engineering School of the National Research Nuclear University MEPHI. This issue of the journal includes works performed in the 2024-2025 academic year by teachers, staff, graduate students and undergraduates of the HES MEPHI on the topics of systems engineering, digital engineering and data engineering.
В статье представлена методология построения цифровых моделей для анализа, проектирования и оптимизации производственно-технологических цепочек (ПТЦ) в отраслях с непрерывной или квазинепрерывной переработкой сырья. Основу подхода составляет формализация понятий «технологический поток», «передел» и «воздействие», что позволяет создавать цифровые двойники на всех этапах жизненного цикла технологии — от лабораторной верификации до опытно-промышленной и промышленной реализации. Рассмотрены онтологическая структура моделирования, параметризация потоков, классификация моделей по уровням сложности (структурные, балансовые, динамические, вариационные, сценарные), а также подходы к построению цифровых моделей как отдельных переделов, так и комплексных ПТЦ. Подчёркивается значение интеграции с автоматизированными системами управления, применения no-code/low-code платформ и использования цифровых экспериментов при выборе и оптимизации производственных схем. Концепция направлена на снижение рисков внедрения инноваций в условиях нереферентных технологий и сформирована на основе работ авторов в широком спектре промышленных отраслей, включая высокоинтенсивную аквакультуру, целлюлозно-бумажное производство и переработку жидких радиоактивных отходов в гражданской атомной энергетике. The article presents a methodology for building digital models for the analysis, design and optimization of production and technological chains (PTC) in industries with continuous or quasi-continuous processing of raw materials. The approach is based on the formalization of the concepts of "technological flow", "redistribution" and "impact", which makes it possible to create digital twins at all stages of the technology life cycle, from laboratory verification to pilot and industrial implementation. The ontological structure of modeling, the parameterization of flows, the classification of models by levels of complexity (structural, balance, dynamic, variational, scenario), as well as approaches to the construction of digital models of both individual processing stages and complex PTCs are considered. The importance of integration with automated control systems, the use of no-code/low-code platforms and the use of digital experiments in the selection and optimization of production schemes is emphasized. The concept is aimed at reducing the risks of introducing innovations in the context of non-reference technologies and is formed on the basis of the authors' work in a wide range of industrial sectors, including high-intensity aquaculture, pulp and paper production and processing of liquid radioactive waste in civil nuclear energy.
Превышение плановых сроков реализации объектов строительства остается одной из наиболее острых проблем, снижающих эффективность строительной отрасли России. Ключевым препятствием использования исторических данных и методов машинного обучения является крайняя разрозненность и денормализованность исходных данных, в частности, значительная вариативность и нестандартность наименований строительных работ. Авторами предложена методология автоматической унификации вариативности в наименованиях строительных работ с применением машинного обучения. Данная унификация рассмотрена как критический шаг для построения систем прогнозирования сроков в рамках государственной стратегии сокращения строительных циклов на 30% к 2030 г. Предложена комплексная методология: предобработка данных, векторизация текста и автоматическая унификация. Исследовано применение двух подходов: кластеризации (K-Means, DBSCAN) для выявления групп схожих работ и классификации (включая Логистическую регрессию, SVM, Random Forest, XGBoost, LSTM) для отнесения описаний к унифицированным классам. Проведен сравнительный анализ эффективности алгоритмов по метрикам точность/полнота/F1. Доказана высокая эффективность методов машинного обучения (особенно XGBoost, Random Forest, LSTM) для унификации, что позволяет формировать структурированный массив исторических данных. Предложено комплексное решение для формирования рекомендаций для оптимизации календарных планов на основе анализа схожих исторических объектов методами машинного обучения. Сформирован структурированный массив исторических данных, включающий информацию о наименованиях работ и их длительностях, который после будет использован для внедрения решений, позволяющих получать обоснованные прогнозы сроков, корректировать плановые длительности строительных работ и минимизировать риски срыва, повышая эффективность управления проектами. Exceeding the planned deadlines for the implementation of construction projects remains one of the most pressing problems reducing the efficiency of the Russian construction industry. The key obstacle to using historical data and machine learning methods is the extreme fragmentation and denormalization of the source data, in particular, significant variability and non-standardization of the names of construction works. The authors propose a methodology for the automatic unification of variability in the names of construction works using machine learning. This unification is considered as a critical step in building systems for forecasting deadlines within the framework of the state strategy to reduce construction cycles by 30% by 2030. A comprehensive methodology is proposed: data preprocessing, text vectorization and automatic unification. The use of two approaches is studied: clustering (K-Means, DBSCAN) to identify groups of similar works and classification (including Logistic regression, SVM, Random Forest, XGBoost, LSTM) to assign descriptions to unified classes. A comparative analysis of the efficiency of algorithms by the metrics accuracy/recall/F1 was conducted. High efficiency of machine learning methods (especially XGBoost, Random Forest, LSTM) for unification was proven, which allows to form a structured array of historical data. A comprehensive solution for generating recommendations for optimizing calendar plans based on the analysis of similar historical objects using ML was proposed. A structured array of historical data was formed, including information on the names of works and their durations, which will then be used to implement solutions that allow to obtain reasonable forecasts of deadlines, adjust the planned duration of construction work and minimize the risks of failure, increasing the efficiency of project management.
The safety of nuclear power plants (NPPs) is achieved through the implementation of a multi-layered protection system and safety functions, tailored to the configuration of the nuclear facility. One of the critical systems in ensuring safety is the reactor protection system, which ensures the automatic transition of the reactor to a controlled safe state during emergency situations and in the event of deviations from normal operating conditions. An effective assessment of the functional safety of reactor protection systems during the design, implementation, and operation phases serves as a vital tool for confirming compliance with safety requirements and ensuring reliable protection of NPPs. In contemporary contexts, the evaluation of functional safety effectiveness is conducted based on a risk-oriented approach, which is implemented in accordance with the regulatory requirements set forth by the IAEA and national regulatory bodies. This article presents a systematic study that investigates the risk-oriented approach to assessing the functional safety of reactor protection systems in NPPs. The research is grounded in a systematic mapping of findings from a sample of 72 articles, selected according to predefined criteria. Consequently, the current state of the field has been defined, and key methodologies, methods, and tools have been identified. Additionally, significant trends have been highlighted, including the integration of modern modeling techniques with traditional approaches. Furthermore, areas for growth and promising directions for future research have been identified, as noted in the reviewed articles.
В современных условиях, когда регионы регулярно сталкиваются с экономическими, эпидемиологическими, геополитическими и климатическими шоками, повышение адаптивности региональных социально-экономических систем приобретает ключевое значение. Особенно это актуально для сектора МСП, который, несмотря на свою уязвимость, играет ключевую роль в устойчивом развитии регионов. Повышение шокоустойчивости регионов требует использования математического моделирования, позволяющего не только оперативно оценивать риски, но и прогнозировать последствия кризисов, а также поддерживать принятие обоснованных управленческих решений. Цель настоящей работы – адаптировать модель Леонтьева для прогнозирования последствий шоковых воздействий на МСП региона. Ее практическая значимость выражается в возможности применения полученной новой прогнозной модели в функции системы поддержки принятия решений органами власти при планировании и реализации политик и конкретных мер поддержки сектора МСП в условиях неопределенности. Авторы обосновывают выбор модели Леонтьева и оригинального подхода, предусматривающего использование в качестве «отраслей» секторов крупного бизнеса и МСП. На реальных данных о динамике социально-экономических показателей Липецкой области проведены расчеты количественных значений элементов Леонтьевских матриц и краткосрочного урона, вызванного последствиями шоков 2019 и 2020 гг., предложена интерпретация полученных данных. In modern conditions, when regions regularly face economic, epidemiological, geopolitical and climatic shocks, increasing the adaptability of regional socio-economic systems is of key importance. This is especially true for the SME sector, which, despite its vulnerability, plays a key role in the sustainable development of regions. Increasing the resilience of regions requires the use of modern digital solutions that can not only promptly assess risks, but also predict the consequences of crises, and support the adoption of informed management decisions. The purpose of this paper is to adapt the Leontief’s model to forecast the consequences of shock impacts on regional SMEs. Its practical significance is expressed in the possibility of using the obtained new forecast model as a decision support system for authorities when planning and implementing policies and specific measures to support the SME sector under conditions of uncertainty. The authors substantiate the choice of the Leontief’s model and the original approach, which provides for the use of large business and SME sectors as “industries”. Based on real data on the dynamics of socio-economic indicators of Lipetsk region, quantitative values of the elements of Leontief’s matrices of inter-industry balance and short-term damage from the shocks of 2019 and 2020 were calculated, and an interpretation of the obtained data was proposed.
В статье рассматривается программная библиотека, предназначенная для автоматического шифрования и дешифрования конфиденциальных данных на уровне приложения с использованием отечественных алгоритмов ГОСТ. Актуальность решения обоснована повышенными требованиями информационной безопасности к защите персональных данных и соответствием нормативным актам Российской Федерации в области криптографической защиты информации. Проанализированы существующие подходы к шифрованию полей баз данных – от механизмов СУБД до средств объектно-реляционного отображения – и показаны их ограничения в контексте российских стандартов. Предлагаемая библиотека представляет собой обёртку над Spring Data JPA (с возможностью работы на уровне Hibernate), обеспечивающую прозрачное шифрование полей, помеченных специальной аннотацией, при сохранении в базу данных и автоматическое расшифрование при чтении. Описана архитектура решения: использование аннотаций, перехват операций сериализации/десериализации посредством встроенных механизмов JPA (конвертеры атрибутов и слушатели событий), интеграция в цикл работы Spring Data, а также управление криптографическими ключами. Особое внимание уделяется алгоритмам шифрования ГОСТ 28147-89 и ГОСТ Р 34.12-2015 («Магма» и «Кузнечик»), их режимам работы и применимости для задачи шифрования данных в СУБД. Обсуждается возможность использования сертифицированных криптографических провайдеров (например, CryptoPro CSP) для реализации алгоритмов в соответствии с требованиями регуляторов. Приведено сравнение с существующими аналогами, показаны отличия и научная новизна предлагаемого подхода. This paper introduces a software library that enables transparent, application-level encryption and decryption of sensitive data fields while remaining fully compliant with Russian cryptographic regulations. The solution targets growing information-security requirements for personal-data protection and aligns with domestic legislation governing cryptographic safeguards. We survey current techniques for encrypting database fields—from DBMS-native mechanisms to object–relational mapping (ORM) extensions—and identify their limitations in the context of Russian GOST standards. The proposed library acts as a lightweight wrapper around Spring Data JPA (with optional direct Hibernate integration). Data members marked with a custom annotation are automatically encrypted on persistence and decrypted on retrieval, without altering business logic. The architecture leverages JPA attribute converters and event listeners to intercept serialization and deserialization, integrates seamlessly into the Spring Data lifecycle, and incorporates flexible key-management facilities.Particular attention is paid to GOST 28147-89 and GOST R 34.12-2015 (“Magma” and “Kuznechik”) block-cipher algorithms, their operating modes, and suitability for database-encryption workloads. We also discuss the use of certified cryptographic providers such as CryptoPro CSP to satisfy regulator requirements. A comparative analysis with existing solutions highlights the advantages and scientific novelty of the proposed approach, demonstrating its ability to deliver fine-grained, standards-compliant security with minimal impact on developer productivity or system performance.
Целью настоящей работы является исследование возможностей моделей глубокого обучения в диагностике заболеваний надпочечников по изображениям компьютерного томографа (КТ). Для поддержки принятия врачебных решений и автоматизированной классификации найденных новообразований по КТ брюшной полости разработан двухэтапный нейросетевой подход, сочетающий сегментацию и классификацию патологий. Он позволяет визуализировать найденные образования и работать с несколькими образованиями, находящихся на одном КТ-снимке. В исследовании использован набор данных, предоставленный НМИЦ эндокринологии им. академика И.И. Дедова, на момент написания статьи состоящий из 228 КТ-снимков. Для оптимизации времени обработки снимки были конвертированы в видеоформат MP4 (54 кадра в видео), что сократило объем данных без существенной потери диагностической ценности. Изображения прошли предварительную обработку, для борьбы с дисбалансом классов была использована аугментация. Каждый КТ-снимок имеет по три метки, каждая из которых соответствует наличию новообразования соответствующего вида, а именно: злокачественный, доброкачественный или неопределенный фенотип. Для инстанса сегментации новообразований применена модель YOLOv11-seg с предобучением на датасете COCO. Для классификации использована модель 3DResNet-50, обученная на выделенных областях. Предлагаемый комбинированный двухступенчатый подход реализуется в программном комплексе «Ассистент врача-эндокринолога». This study investigates the potential of deep learning models for the diagnosis of adrenal gland diseases using computed tomography (CT) images. To support clinical decision-making and automate the classification of identified neoplasms in abdominal CT scans, a two-stage neural network approach, combining segmentation and classification, was developed. This approach allows for the visualization of detected lesions and accommodates multiple lesions within a single CT image. The study utilized a dataset provided by the Academician I.I. Dedov National Medical Research Center of Endocrinology, comprising 228 CT scans at the time of writing. To optimize processing time, images were converted to MP4 video format (54 frames per video), reducing data volume without significantly compromising diagnostic value. Images underwent preprocessing, and data augmentation was employed to address class imbalance. Each CT scan was annotated with three labels, corresponding to the presence of a neoplasm with a malignant, benign, or indeterminate phenotype. For lesion instance segmentation, a YOLOv11-seg model pre-trained on the COCO dataset was implemented. A 3DResNet-50 model, trained on the segmented regions, was used for classification. The proposed combined two-stage approach is implemented in a software suite designated “Assistant Endocrinologist”.
В работе представлен кейс миграции высоконагруженной системы с Redis 6.2.x на Apache Cassandra 4.1.x в конфигурации двух дата‑центров (RF=3+3, `LOCAL_QUORUM`). Приведена воспроизводимая методика нагрузочного тестирования (YCSB‑A, Zipf, 100 млн ключей, 1 КБ запись, прогрев/замер) и сравнение задержек p95/p99 и пропускной способности с учётом конфигураций операционной системы, файловой структуры, дискового пространства и конфигурации виртуальной JAVA-машины. Показаны результаты испытаний на отказ при потере узла и потере дата-центра. Представлены регламенты эксплуатации, включая восстановление и резервное копирование. Обсуждены варианты компромиссной оптимизации выбора стратегии уплотнения данных с анализом альтернативных подходов, таких как стратегия равномерного уплотнения LCS и стратегия для временных рядов TWCS). Рассмотрено влияние фоновых задач уплотнения и синхронизация данных между узлами в высоконагруженной распределённой системе хранения данных на «хвосты задержек» - временные скачки длительности задержки, ведущие к деградации производительности системы при высоких нагрузках. The paper presents a case study of migrating a high-load system from Redis 6.2.x to Apache Cassandra 4.1.x in a configuration of two data centers (RF=3+3, 'LOCAL_QUORUM'). A reproducible load testing methodology (YCSBA, Zipf, 100 million keys, 1 KB write, warm-up/metering) and a comparison of p95/p99 latency and bandwidth are given, taking into account the configurations of the operating system, file structure, disk space and the configuration of the JAVA virtual machine. The results of failure tests for node loss and data center loss are shown. Operating procedures, including recovery and backup, are presented. Options for compromise optimization of the choice of data compaction strategy are discussed with the analysis of alternative approaches, such as the LCS uniform compaction strategy and the TWCS time series strategy. The influence of background compaction tasks and data synchronization between nodes in a high-load distributed data storage system on "latency tails" - temporary jumps in latency leading to system performance degradation under high loads is considered.
Статья описывает концепцию, архитектуру и результаты опытной эксплуатации «Цифрового полигона» ВИШ НИЯУ МИФИ — цифровой инфраструктуры для объединения образования, НИР и прикладных проектов с индустриальными партнёрами. Показано, как разрыв ожиданий между академической и промышленной средами (в терминах уровней технологической готовности, TRL) трансформируется в требования к сервисам, политике доступа и наблюдаемости. Предложена целевая архитектура, основанная на виртуализации (Proxmox VE), инфраструктуре как коде (Ansible) и едином контуре наблюдаемости (Prometheus/Grafana, распределённый трейсинг), а также регламенты эксплуатации для учебно-исследовательских и пилотных производственных сценариев. В рамках кейсов продемонстрированы: локальный инференс больших языковых моделей на рабочих станциях с несколькими GPU и стенд микросервисного приложения (≈15 сервисов) с трассировкой и бизнес-метриками. По итогам внедрения зафиксированы улучшения эксплуатационных показателей: рост доступности до ~99,5%, снижение числа инцидентов на ~78%, сокращение времени развёртывания типовых служб до 1–2 часов; для микросервисного стенда — уменьшение времени обнаружения и устранения отказов (TTD/TTR) на 67% и 58% соответственно. Научная и практическая новизна работы — в интеграции TRL-подхода к взаимодействию с индустрией с воспроизводимым инженерным шаблоном кампусной платформы (виртуализация + IaC + наблюдаемость) и показом её применимости для on-prem задач ИИ. Намечены шаги дальнейшего развития: кластеризация, унификация IAM/SSO и политика резервного копирования. The article describes the concept, architecture and results of the trial operation of the "Digital Polygon" of the Higher School of National Research Nuclear University MEPhI - a digital infrastructure for combining education, research and applied projects with industrial partners. It is shown how the gap in expectations between academic and industrial environments (in terms of levels of technological readiness, TRL) is transformed into requirements for services, access policy, and observability. A target architecture based on virtualization (Proxmox VE), infrastructure as code (Ansible) and a single observability loop (Prometheus/Grafana, distributed tracing) is proposed, as well as operating regulations for educational research and pilot production scenarios. The cases demonstrated: local inference of large language models on workstations with multiple GPUs and a microservice application bench (≈15 services) with tracing and business metrics. As a result of the implementation, improvements in operational indicators were recorded: an increase in availability up to ~99.5%, a decrease in the number of incidents by ~78%, a reduction in the deployment time of typical services to 1-2 hours; for a microservice bench – a reduction in the time of detection and elimination of failures (TTD/TTR) by 67% and 58%, respectively. The scientific and practical novelty of the work lies in the integration of the TRL approach to interaction with the industry with a reproducible engineering template of the campus platform (virtualization + IaC + observability) and the demonstration of its applicability for on-prem AI tasks. Further development steps are outlined: clustering, IAM/SSO unification and backup policy.
В данной статье рассматривается концепция использования онтологической модели и цифрового двойника для управления жизненным циклом сложного инженерного объекта (СИО) на этапе проектирования и строительства. Основное внимание уделено проблематике увеличения сроков строительства из-за отсутствия единого сквозного процесса управления. Рассмотрены этапы жизненного цикла СИО, включая проектирование, строительно-монтажные работы (СМР) и пуско-наладочные работы (ПНР). Описаны ключевые участники процесса, структура планирования и преобразование входных данных в выходные результаты. Результатами научно-исследовательской работы должна стать концепция онтологической модели, техническое задание на её разработку, на основе функциональных и нефункциональных требования к системе, а также готовая онтологическая модель с настроенными взаимосвязями. This article considers the concept of using an ontological model and a digital double to manage the life cycle of a complex engineering object (SIS) at the design and construction stage. Emphasis is placed on the problem of increasing construction lead times due to the lack of a single end-to-end management process. The life cycle stages of SIS, including design, construction and installation work (CTP) and commissioning (PDP), have been considered. Key players in the process, planning structure and transformation of input data into output results are described. The results of research work should be the concept of ontological model, technical task for its development on the basis of functional and non-functional requirements to the system, as well as a ready ontological model with customized relationships.
Цель исследования - оценить в пилотном сравнительном исследовании влияние модульной адаптации (сокращения) текста клинических рекомендаций, используемых как непосредственный источник информации, на качество (релевантность и безопасность) ответов большой языковой модели при решении реальных клинических задач, касающихся лечения артериальной гипертензии у взрослых. В исследование включены 45 клинических кейсов, сформулированных практикующими врачами, не имевшими опыта взаимодействия с LLM. Кейсы были рандомизированы в три группы: (1) адаптированный текст клинической рекомендации (исследуемая группа), (2) неадаптированный полный текст (контрольная 2), (3) отсутствие документа (контрольная 1). Ответы LLM (GPT-4) генерировались врачом-оператором с использованием кастомного промта (инструкции к языковой модели) и оценивались независимыми экспертами по шкале из трёх критериев: клиническая адекватность, безопасность, соответствие положениям клинических рекомендаций. Исследуемая группа показала наивысшие значения по всем критериям (средние баллы: КА — 8,8; БР — 9,5; СКР — 9,2). Контрольная группа 2 продемонстрировала высокую безопасность (9,1), но меньшую нормативную точность (СКР — 7,3). Контрольная группа 1 существенно отставала по всем показателям (КА — 6,3; БР — 7,5; СКР — 5,1). Установлено, что модульная адаптация (сокращение) клинических рекомендаций, направленная на удаление нерелевантных к задачам разделов, обеспечивает существенное повышение релевантности и нормативной точности медицинских ответов LLM. Подход может быть использован при проектировании систем поддержки принятия врачебных решений на базе LLM, в том числе в рамках цифровой трансформации клинической практики. The objective of this study was to evaluate, in a pilot setting, the effect of modular adaptation of clinical guideline texts on the quality of medical responses generated by a large language model (LLM) when solving real-world clinical tasks related to the management of arterial hypertension in adults. A total of 45 clinical scenarios formulated by practicing physicians were randomized into three groups: (1) adapted modular version of the clinical guideline (intervention group), (2) full unadapted guideline text (control group 2), and (3) no document support (control group 1). Responses were generated by GPT-4 using a custom prompt and assessed independently by experts using three separate criteria: clinical adequacy, safety, and compliance with the official recommendations, rated on a 0–10 scale. The adapted guideline group demonstrated the highest mean scores across all criteria (clinical adequacy – 8.8; safety – 9.5; compliance – 9.2), with statistically significant differences confirmed by the Mann–Whitney U test. The results show that excluding structurally irrelevant sections of the guideline (e.g. definitions, epidemiology, methodological notes) significantly improves the relevance and regulatory accuracy of LLM-generated recommendations. This approach may be applicable in the design of AI-based clinical decision support systems and the broader context of digital transformation in healthcare.
Статья посвящена внедрению системы контроля персонала на строительных площадках посредством использования умных часов. Рассматриваются преимущества и особенности данного решения, такие как возможность мониторинга местоположения сотрудников, отслеживания их активности и состояния здоровья. Отмечается повышение уровня безопасности труда благодаря своевременному реагированию на чрезвычайные ситуации и предупреждению несчастных случаев. Рассматриваются аспекты организации рабочего процесса, оптимизации взаимодействия между сотрудниками и руководством, а также снижение рисков травматизма и улучшение дисциплины работников. Приводятся практические рекомендации по выбору устройств и интеграции их в существующие корпоративные системы управления персоналом. Целью данного исследования является разработка и оценка эффективности внедрения системы контроля персонала на строительных площадках с использованием умных часов для повышения уровня безопасности труда и улучшения организационных процессов. Методами исследования выбраны: экспериментальное внедрение - проведение пилотного проекта на крупном строительном объекте, оснащение сотрудников умными часами с функциями GPS-трекинга, пульсометра и акселерометра; интервьюирование участников внедрения - получение обратной связи от сотрудников и руководства относительно удобства использования устройства, влияния на рабочий процесс и психологическое состояние; статистический анализ полученных результатов - сравнение ключевых показателей эффективности до и после внедрения технологии. The article is devoted to the implementation of a personnel control system on construction sites through the use of smart watches. The advantages and features of this solution are considered, such as the ability to monitor the location of employees, track their activity and health status. There is an increase in the level of occupational safety due to timely response to emergencies and accident prevention. Aspects of organizing the workflow, optimizing interaction between employees and management, as well as reducing injury risks and improving employee discipline are considered. Practical recommendations on the choice of devices and their integration into existing corporate personnel management systems are given. The purpose of this study is to develop and evaluate the effectiveness of implementing a personnel control system on construction sites using smart watches to improve occupational safety and improve organizational processes. The following research methods are selected: experimental implementation - conducting a pilot project at a large construction site, equipping employees with smart watches with GPS tracking, heart rate monitor and accelerometer functions; interviewing implementation participants - receiving feedback from employees and management regarding the convenience of using the device, its impact on the workflow and psychological state; statistical analysis of the results obtained is a comparison of key performance indicators before and after technology implementation.
The aim of this article is to explore the capabilities of modern neural networks in analyzing cytological whole slide images in svs format in accordance with the Bethesda categorization system. The article presents the results of the proposed neural network models for solving tasks of automatic segmentation and classification of both various types of individual cells and their clusters. For the diagnosis of thyroid cancer using computer vision, the following cell types are identified: Hurthle cells, cells with pseudoinclusions, C-cells, and clusters of cells forming papillary structures, shapeless structures with ordered and unordered cell arrangements. The proposed models have demonstrated their effectiveness in solving tasks related to the intelligent analysis of both whole-slide cytological images and tiled image segmentation. The results obtained for the segmentation of individual cells are as follows: mean Dice coefficient (DC) = 90.9% for pseudoinclusions, DC = 86.2% for Hurthle cells, DC = 90.2% for C-cells. For cell cluster segmentation, the mean Intersection over Union (IoU) is 84%, and DC is 91%. The classification accuracy of cell clusters into three classes is 77.9%.
Целью работы является комплексный анализ механизмов раннего коллапса языковых моделей, работающих с медицинскими текстами, при рекурсивном их обучении на примере архитектур Mistral-7B и LLaMA-3. Проведено экспериментальное исследование динамики изменения перплексии, метрик BLEU и ROUGE, а также распределения вероятностей токенов в процессе многопоколенческого синтетического обучения. Выявлены два типа коллапса моделей: ранний (характеризующийся быстрой деградацией вероятностных распределений) и поздний (с постепенным снижением разнообразия генерации). Установлено, что модель Mistral демонстрирует большую устойчивость к коллапсу данных по сравнению с LLaMA, что обусловлено особенностями ее архитектуры с механизмом скользящего внимания (sliding window attention). Работа предлагает новый методологический подход к количественной оценке деградации языковых моделей и формулирует практические рекомендации по предотвращению потери модельного разнообразия при рекурсивном обучении. Исследование проводилось на текстовых цитологических данных, используемых при диагностике заболеваний щитовидной железы. The aim of the work is a comprehensive analysis of the mechanisms of early collapse of language models working with medical texts during their recursive training using the example of the Mistral-7B and LLaMA-3 architectures. An experimental study of the dynamics of perplexity change, BLEU and ROUGE metrics, as well as the probability distribution of tokens in the process of multi-generation synthetic training was conducted. Two types of model collapse are identified: early (characterized by rapid degradation of probability distributions) and late (with a gradual decrease in the diversity of generation). It is established that the Mistral model demonstrates greater resistance to data collapse compared to LLaMA, which is due to the features of its architecture with a sliding window attention mechanism. The paper proposes a new methodological approach to quantifying the degradation of language models and formulates practical recommendations for preventing the loss of model diversity during recursive learning. The study was conducted on text cytological data used in the diagnosis of thyroid diseases.
В настоящей статье рассмотрена цифровая трансформация предприятия и связанное с ней принятие стратегических решений, направленных на получение стабильных результатов в долгосрочной перспективе, а также типовые практические ошибки, совершаемые в ходе подобной трансформации. Освещенные вопросы цифровой трансформации применимы в практике частных или государственных предприятий и в равной степени затрагивают как производственные, так и непроизводственные сферы деятельности. Статья раскрывает суть сформировавшихся подходов в области цифровой трансформации и их различия при рассмотрении в краткосрочной и долгосрочной перспективах развития компании. Наряду с этим, предлагаются простые и универсальные пути подготовки цифровой трансформации на предприятии. Отмечается важность работы с бизнес-процессами и архитектурой будущих систем для наиболее полного понимания предполагаемых границ цифровой трансформации и определения ожидаемых эффектов не только в виде финансовых и временных показателей, но и лучшего понимания того, в каком виде эта цифровая трансформация должна быть реализована. Отдельное внимание в данной статье уделяется невозможности цифровой трансформации без предварительного обследования и осмысления текущего состояния предметной области. Так, отмечается большая вероятность получения отрицательного результата проведенных мероприятий по улучшениям при немедленной цифровизации какой-либо предметной области без её переосмысления. На выходе, как правило, можно получить те же неупорядоченные процессы, только в цифровом виде и гораздо более дорогие в обслуживании, поскольку после цифровизации процесс будет использовать цифровой инструментарий (например, программу-редактор, парсер и т.д.), либо полноценную информационную систему, нуждающуюся в постоянной технической поддержке и обновлениях. В заключении можно отметить выводы, касающиеся, как процессной и архитектурной составляющей цифровой трансформации, так и потребность в оценке эффективности ранее внедренных цифровых решений, реальную эффективность которых в ряде случаев можно поставить под сомнение. Таким образом, автором подсвечиваются потребности исследования не только текущего, но и ретроспективного состояния той или иной предметной области компании, не показывающих анонсированные при внедрении цифровых решений эффекты. This article examines the digital transformation of an enterprise and the associated strategic decision-making aimed at achieving stable results in the long term, as well as typical practical mistakes made during such a transformation. The highlighted issues of digital transformation are applicable in the practice of private or state-owned enterprises and equally affect both production and non-production areas of activity. The article reveals the essence of the established approaches in the field of digital transformation and their differences when considered in the short and long-term development prospects of the company. Along with this, we offer simple and universal ways to prepare for digital transformation in an enterprise. It is noted that it is important to work with business processes and architecture of future systems in order to fully understand the expected boundaries of digital transformation and determine the expected effects not only in the form of financial and time indicators, but also to better understand the form in which this digital transformation should be implemented. Special attention in this article is paid to the impossibility of digital transformation without a preliminary examination and current state of the subject area. Thus, there is a high probability of obtaining a negative result of the improvement measures carried out with the immediate digitalization of any subject area without rethinking it. At the output, as a rule, you can get the same disordered processes, only in digital form and much more expensive to maintain, since after digitalization the process will use digital tools (for example, an editor program, a parser, etc.), or a full-fledged information system in need of constant technical support and updates. In conclusion, we can note the conclusions concerning both the process and architectural components of digital transformation, as well as the need to assess the effectiveness of previously implemented digital solutions, the real effectiveness of which in some cases can be questioned. Thus, the author highlights the research needs of not only the current, but also the retrospective state of a particular company’s subject area, which does not show the effects announced during the implementation of digital solutions.
Рассмотрен подход к организации репозитория метаданных как одного из элементов системы оценки качества данных. Для формирования репозитория предлагается использовать объектно-ориентированную модель данных, обогащенную проверками качества данных. Дано описание основных элементов репозитория такого типа, а также описан подход по привязке проверок качества данных к объектам модели данных. Предложен способ хранения рассмотренного репозитория метаданных, разработаны специальные алгоритмы для его обработки, включая алгоритм «распаковки» атрибутов класса и алгоритм определения актуальных проверок данных с учетом возможных переопределений. На основе описанных теоретических положений был реализован прототип репозитория метаданных. Указанный прототип был использован для организации проверок на примере задачи оценки качества данных операторов, осуществляющих обработку персональных данных. В сравнении с реализацией репозитория метаданных на основе физической модели данных, применение описанного в настоящем исследовании подхода отличается сокращением объема описанных атрибутов и проверок качества данных на 23% и 27% соответственно при одновременном сохранении количества реально запускаемых («реализованных») проверок. Исследуемый в статье подход может быть полезен в практических задачах анализа качества данных как потенциальный способ снижения трудозатрат на управление проверками качества. The paper considers the approach to organizing a metadata repository as one of the elements of the data quality assessment system. An object-oriented data model enriched with data quality checks is proposed for repository formation. A description of the key repository elements is given, along with an approach to linking data quality checks to objects in the data model. The article provides a method for storing the discussed metadata repository and special algorithms for its processing, including the "unpacking" algorithm for class attributes and the algorithm for determining the relevant data checks considering possible overrides. Based on the described theoretical propositions, a prototype of the metadata repository was implemented. The prototype was used to organize checks for assessing the data quality of personal data operator’s registry. In comparison with the implementation of a metadata repository based on a physical data model, the application of the approach described in this research results in a reduction of attribute and data quality check description by 23% and 27%, respectively, while maintaining the same quantity of executed checks. The investigated approach can be useful in practical tasks related to data quality analysis as a potential way to reduce the workload of data quality check management.
В работе рассматривается одна из основных задач при разработке интерпретатора для инструментария BlockSet – проектирование синтеза SQL запросов на основе метамодели, а так программного обеспечения для его применения. Авторами рассмотрен как инструментарий в целом, так и роль самого алгоритма в его контексте. Подробно рассмотрены периферийные инструменты для управления алгоритмом в разках языка BML. Продемонстрированы этапы его формирования и технические тонкости реализации. Помимо прямого алгоритма так же рассмотрен обратный алгоритм генерации запроса для поиска факторов частных событий, который необходим при реализации ресурса на основе событийно-ориентированного подхода. Для наглядности разобран пример обработки события появления “сообщения”, в результате которого необходимо определять id “пользователей” -- отправителя и получателя, которых нужно уведомить о появлении указанного “сообщения”. Подведены выводы, демонстрирующие, результаты проделанной работы. The paper examines one of the main tasks in developing an interpreter for the BlockSet toolkit - designing the synthesis of SQL queries based on the metamodel, as well as software for its use. The authors examined both the toolkit as a whole and the role of the algorithm itself in its context. Peripheral tools for managing the algorithm in the BML language are discussed in detail. The stages of its formation and technical details of implementation are demonstrated. In addition to the direct algorithm, a reverse algorithm for generating a query to search for factors of private events, which is necessary when implementing a resource based on an event-oriented approach, is also considered. For clarity, an example of processing the “message” event has been analyzed, as a result of which it is necessary to determine the id of “users” - the sender and the recipient, who need to be notified about the appearance of the specified “message”. Conclusions are drawn demonstrating the results of the work done.
This article discusses the problem of blocking nodes in road networks and percolation thresholds for the transport infrastructure of a modern metropolis. For metropolitan networks, the values of percolation thresholds are calculated and displayed, considering the different density of connections between network nodes. Further, it is shown that the dependence of the values of the percolation thresholds on the network’s density can be described by functional dependencies with a high degree of correlation. The obtained results can be used to assess the reliability of transport infrastructure and to check the increase in the capacity of selected sections of the road network. Further, the article discusses obtaining a description of road infrastructure from open sources (obtaining data from OpenStreetMap using the SUMO - Simulation of Urban MObility package), after which it is possible to build a graph of the road network and determine its percolation properties. For the constructed nodes of the graph described a model of the stochastic dynamics of blocking a single road lane in the transport network. The threshold value L of number of cars that can be placed in the lane (based on the length of the road lane) is used as constraints and the incoming and outgoing flows of cars are also determined as income parameters of a model. The constructed model allows us to obtain the predicted blocking time of a road network line, for a given probability of such blocking, where the probability of blocking a single road network line is taken from the percolation properties of this network discussed earlier. The resulting blocking times of road network nodes make it possible to build an algorithm for controlling traffic light regulation.
The widespread use of heterogeneous computing platforms, as well as the incorporation of computationally expensive implementations of intelligent data analysis algorithms into modern software systems leads to the demand in moving software fragments to most suitable hardware accelerators that are available on a heterogeneous computing platform. In this research, we propose an approach to the generation of recommendations for improving the performance of software systems by finding candidate algorithm implementations for hardware acceleration, and by suggesting the most suitable hardware accelerator among the specialized processors that are available on a given heterogeneous computing platform. The proposed approach is based on a code-to-code search technique, which extracts code fragments from an abstract syntax tree (AST), converts them into vectors containing program features, and compares the vectors with the query program vector. The obtained results confirm that the use of automatically recommended hardware accelerators for the code fragments identified using the proposed approach indeed allows to increase the performance of software systems solving machine learning tasks.