目的 探讨影响人工智能心电分类算法抗扰性的因素.方法 使用公开心电数据库和开源的人工智能心电分类算法模型,在算法训练、测试环节引入实验室测量的噪声,同时改变训练集的样本量,观察算法测试结果的变化趋势.结果 当测试集单独添加噪声时,算法分类的总体准确性下降超过3%;当训练集、测试集分别添加相同类型噪声时,算法分类的总体准确性的降幅不超过0.5%;对于不同类型的心拍,训练集样本量对算法分类准确率的影响趋势各不相同.结论 人工智能心电分类算法在训练阶段可引入必要的噪声,以加强算法本身的抗扰性,但同时应关注各分类之间的差异.训练集的扩增并非必然提升算法的抗扰性.
Objective. To explore a centralized approach to build test sets and assess the performance of an artificial intelligence medical device (AIMD) which is intended for computer-aided diagnosis of diabetic retinopathy (DR). Method. A framework was proposed to conduct data collection, data curation, and annotation. Deidentified colour fundus photographs were collected from 11 partner hospitals with raw labels. Photographs with sensitive information or authenticity issues were excluded during vetting. A team of annotators was recruited through qualification examinations and trained. The annotation process included three steps: initial annotation, review, and arbitration. The annotated data then composed a standardized test set, which was further imported to algorithms under test (AUT) from different developers. The algorithm outputs were compared with the final annotation results (reference standard). Result. The test set consists of 6327 digital colour fundus photographs. The final labels include 5 stages of DR and non-DR, as well as other ocular diseases and photographs with unacceptable quality. The Fleiss Kappa was 0.75 among the annotators. The Cohen’s kappa between raw labels and final labels is 0.5. Using this test set, five AUTs were tested and compared quantitatively. The metrics include accuracy, sensitivity, and specificity. The AUTs showed inhomogeneous capabilities to classify different types of fundus photographs. Conclusions. This article demonstrated a workflow to build standardized test sets and conduct algorithm testing of the AIMD for computer-aided diagnosis of diabetic retinopathy. It may provide a reference to develop technical standards that promote product verification and quality control, improving the comparability of products.
目的 研究AI医疗器械质量评价需求和标准体系设计.方法 围绕国内外AI医疗器械监管政策与标准化动态,开展文献调研;结合产品测试经验,分析AI医疗器械标准发展的方向和定位.结果 AI医疗器械的标准体系需要面向产品全生命周期监管与质量管理的需要,协调推进基础标准、方法标准的制修订,在时机成熟时进一步发展产品标准、管理标准.结论 本文研究的策略有助于推动AI医疗器械标准化工作,完善质量评价体系.
目的 旨在研究人工智能(Artificial Intelligence,AI)医疗器械质量管理的发展趋势,促进标准规范研究.方法 根据AI医疗器械相关监管、法规的最新动态,梳理质量管理面临的特殊问题;结合产品检测实践情况,分析AI医疗器械质量管理的需求与发展方向.结果 在AI医疗器械的质量管理过程中,可溯源性、数据集管理、质量控制等方面都需进一步增强.结论 本文研究的思路和方法有助于建立适合AI医疗器械的专用标准规范.
为了在系列政策的支持下,我国医用机器人产业得到了飞速的发展,在临床诊疗中得到了较为广泛的应用,在手术治疗、术后康复领域显现了巨大的优势和先进性.但医用机器人是一个融合了多学科、多技术的复杂医疗器械系统,具有较大的潜在风险,需要制定专用的标准体系来对该类产品进行约束和规范,从而引导产业向高端化发展.本文阐述了医用机器人相关标准现状,梳理了国内外医用机器人的标准化需求,从通用定义、可靠性、可用性、系统集成、部件级性能及测试方法、产品级性能及安全六方面进行分析,提出专用标准体系的设计思路,以期后续为我国医用机器人领域标准化工作的科学有序开展提供参考.
Introduction Computer aided detection and diagnosis (CADe and CADx) products are an emerging branch of medical device industry. However, limited technical standard has been developed for product verification and validation. It will be helpful to investigate the current practice of preclinical and clinical evaluation of approved products and provide insights for future standardization. Areas covered Document review was conducted on 56 products approved by the United States Food and Drug Administration, including Summary of Safety and Effectiveness Data, 510(k) decision andde novodecision summaries. Key parameters describing product characteristics, preclinical studies and clinical studies were collected. Evaluation strategies for CADe/CADx products were analyzed and assessed. Expert opinion Preclinical studies were widely adopted in the verification of CADe/CADx products. Standalone performance testing was a common procedure, but the selection of testing dataset and performance metrics showed significant variability and flexibility among manufacturers. Clinical studies were reported by all class III products and some class II products, and Multi-Reader Multi-Case design was commonly used. However, statistical analysis and presentation/interpretation of results was oftentimes incomplete. To resolve above issues, systematic development of standards of CADe/CADx is encouraged, which can be implemented at different aspects through the product lifecycle.
目的 探讨手术机器人关键性能指标的评价方法,并建立测试方案.方法 以图像引导式手术机器人为例,分析系统定位准确性及系统延时的各影响因素,通过标准模体的设计,对比模体端与机械臂端空间两向量的长度及夹角进行系统定位精度评价;通过非接触式实验装置的设计,监测从患者移动到机械臂移动的时间差进行系统延时的评价.结果 实现系统精度与系统延时的测量.结论 试验方案的设计应充分模拟临床应用状态,关注部件级性能的同时应关注系统性能的评价.
目的 针对手术机器人的可用性评价方法进行研究,制定相应的评估方法,以期对医用机器人可用性评价提供思路.方法 手术机器人可用性评价采用无干扰观察法.通过观察研究用户操作使用手术机器人以及操作任务中的流畅度等信息进行数据收集和记录测试,从而实现对机器人的可用性进行评价.结果 结合手术机器人的特点,本文提出了手术机器人可用性测试的基本流程和基本人员要求,针对手术机器人可用性评价的核心的功能,可用性测试涉及的操作规范,可用性测试应用场景要求和任务的起点和终点等方面进行了讨论.结论 总体来说,针对手术机器人进行可用性评价,有利于科学有效地评价手术机器人在真实手术过程中可能遇到的风险,有助于形成合理的手术机器人质量评价规范.
目的:梳理人工智能医疗器械的伦理现状与面临的问题,探讨医疗器械领域可遵循的人工智能伦理准则和发展原则.方法:查阅国内外已发布的人工智能伦理准则,国内外监管、医学伦理相关规范要求文件.结果 与结论:人工智能医疗器械在发展与应用过程中仍面临着社会影响、个人数据保护、人工智能算法和医学伦理等问题.人工智能医疗器械在沿着推动医学人工智能伦理标准化、保持数据完整性方向发展的同时,还要保障数据隐私.
随着信息技术和互联网行业的发展,全球进入大数据时代,数据的开发、挖掘和分析应用越来越广泛,对数据的质量要求也越来越高.目前,国内外的专家学者对医疗领域人工智能产品都进行了很多研发,人工智能产品的研发需要依托海量的医学临床数据.为了保证这类产品的质量,必须从源头进行必要的筛选和清洗,以保障数据质量,支持后续的产品研发与验证过程.本文对DICOM格式的数据清洗问题进行分析,开发了对原始数据进行清洗和审核的流程,在实践中进行了测试,证明能够有效地发现数据缺陷,为今后开展医学人工智能专用数据集的质控工作起到借鉴作用.
从工程、主观、经济三方面探讨相应的评价内容、指标和方法,旨在全面、有效地评价国内外磁共振设备的优劣,使得不同品牌、不同类型的磁共振设备之间得以向比较,为我国磁共振设备的研发、质控、售后提供参考.