Coronary artery disease (CAD), one of the leading causes of mortality worldwide, necessitates effective risk assessment strategies, with coronary artery calcium (CAC) scoring via computed tomography (CT) being a key method for prevention. Traditional methods, primarily based on UNET architectures implemented on pre-built models, face challenges like the scarcity of annotated CT scans containing CAC and imbalanced datasets, leading to reduced performance in segmentation and scoring tasks. In this study, we address these limitations by introducing DINO-LG, a novel label-guided extension of DINO (self-distillation with no labels) that incorporates targeted augmentation on annotated calcified regions during self-supervised pre-training. Our three-stage pipeline integrates Vision Transformer (ViT-Base/8) feature extraction via DINO-LG trained on 914 CT scans comprising 700 gated and 214 non-gated acquisitions, linear classification to identify calcified slices, and U-NET segmentation for CAC quantification and Agatston scoring. DINO-LG achieved 89% sensitivity and 90% specificity for detecting CAC-containing CT slices, compared to standard DINO's 79% sensitivity and 77% specificity, reducing false-negative and false-positive rates by 49% and 57% respectively. The integrated system achieves 90% accuracy in CAC risk classification on 45 test patients, outperforming standalone U-NET segmentation (76% accuracy) while processing only the relevant subset of CT slices. This targeted approach enhances CAC scoring accuracy by feeding the UNET model with relevant slices, improving diagnostic precision while lowering healthcare costs by minimizing unnecessary tests and treatments.
Reading the Herculaneum papyri is challenging because both the scrolls and the ink, which is carbon-based, are carbonized. In X-ray radiography and tomography, ink detection typically relies on density- or composition-driven contrast, but carbon ink on carbonized papyrus provides little attenuation contrast. Building on earlier X-ray phase contrast tomographic and microscopic observations suggesting that relief contributes to letter visibility, we show that surface morphology of written regions contains enough signal to distinguish ink from papyrus. To this end, we train a deep-learning model on three-dimensional optical profilometry of mechanically opened Herculaneum papyri to separate inked and uninked areas. Although our measurements do not show general bulk letter relief or universal roughness cues, we show that high-resolution topography alone contains a usable signal for ink detection on the collected dataset. We further quantify how lateral sampling governs learnability and how a native-resolution model behaves on coarsened inputs. Leave-one-papyrus-out transfer is heterogeneous. Diminishing segmentation performance with decreasing lateral resolution provides insight into the characteristic spatial scales that must be resolved on our dataset to exploit the morphological signal. These findings define a proof-of-concept and conditionally inform resolution needs for morphology-based reading of closed scrolls, which will also depend on the imaging modality.
Multispectral Imaging (MSI) was used to recover the Archimedes Palimpsest. However, the work did not report to have any ink bleed-through. We develop an MSI pipeline that helps in visualizing the content and the ink bleed-through. As for data, we use damaged pages from the sacramental journal of the Holy Trinity Greek Orthodox Community (HTGOC) Church, situated in New Orleans. During Hurricane Katrina, this church faced destruction leading to damaged sacramental journals, like ink wash-away, ink bleed-through, torn patches of pages, and mold development. In this work, we try to discover any ink signals that show up in a wide spectrum of light, from ultraviolet (UV) to infrared (IR). We also try to study the ink signal response to the different combinations of wavelengths and perform an ablation study that gives us a list of combinations of different spectral bands that generate maximum information about the data. The manuscript reports 12 such combinations after the ablation study, and our data use all the wavelengths to produce ink signals, which are in better contrast in comparison to the images captured under visible light.
Through the annals of time, writing has slowly scrawled its way from the painted surfaces of stone walls to the grooves of inscriptions to the strokes of quill, pen, and ink. While we still inscribe stone (tombstones, monuments) and we continue to write on skin (tattoos abound), our quotidian method of writing on paper is increasingly abandoned in favor of the quick-to-generate digital text. And even though the stone-inscribed text of epigraphy offers demonstrably better permanence than that of writing on skin and paper—even better than that of the memory system of the modern computer (Bollacker in Am Sci 98:106, 2010)—this field of study has also made the digital leap. Today’s scholarly analyses of epigraphic content increasingly rely on high-tech approaches involving data science and computer models. This essay discusses how advances in a number of exciting technologies are enabling the digital analysis of epigraphic texts and accelerating the ability of scholars to preserve, renew, and reinvigorate the study of the inscriptions that remain from throughout history.
We present a complete software pipeline for revealing the hidden texts of the Herculaneum papyri using X-ray CT images. This enhanced virtual unwrapping pipeline combines machine learning with a novel geometric framework linking 3D and 2D images. We also present EduceLab-Scrolls, a comprehensive open dataset representing two decades of research effort on this problem. EduceLab-Scrolls contains a set of volumetric X-ray CT images of both small fragments and intact, rolled scrolls. The dataset also contains 2D image labels that are used in the supervised training of an ink detection model. Labeling is enabled by aligning spectral photography of scroll fragments with X-ray CT images of the same fragments, thus creating a machine-learnable mapping between image spaces and modalities. This alignment permits supervised learning for the detection of "invisible" carbon ink in X-ray CT, a task that is "impossible" even for human expert labelers. To our knowledge, this is the first aligned dataset of its kind and is the largest dataset ever released in the heritage domain. Our method is capable of revealing accurate lines of text on scroll fragments with known ground truth. Revealed text is verified using visual confirmation, quantitative image metrics, and scholarly review. EduceLab-Scrolls has also enabled the discovery, for the first time, of hidden texts from the Herculaneum papyri, which we present here. We anticipate that the EduceLab-Scrolls dataset will generate more textual discovery as research continues.
This article describes the first effort to read inside a damaged codex using X-Ray micro-CT imaging, which has an additional complication beyond most unrolled scrolls, for which the process has been successful: there is writing on both sides. The project is a collaboration between a humanist, a team of computer scientists and engineers, as well as librarians and conservators, to undertake the x-ray micro-CT imaging of codex M.910, a fifth- or sixth-century parchment codex of Acts of the Apostles which is too damaged to open in its current state. The first round of image processing was conducted in December 2017 at the Morgan Library and Museum, and a second round in November 2019; work on restoring the text using machine learning is ongoing, and has already resulted in the identification of some words and phrases. We first describe codex M.910, including the basics of its codicology, and its potential significance for early Christian book culture, as well as the history of the biblical text. We then provide an overview of the manuscript imaging process, at a level of technical detail intended for a general audience, with the hope of providing a reference for future work in this expanding field of research. The key initial step was the preparation of the manuscript’s mount, which had to take into account the necessities of both conservation safety and micro-CT imaging. We also break down the imaging process itself, which was carried out by Skyscan micro-CT scanner, donated for use in this project by Micro Photonics. Finally, we give a brief discussion of the ongoing preliminary data analysis.
Abstract:Ancient documents pose many challenges for the scholars who painstakingly study and elucidate them. Natural deterioration occurs over time, erasing words and sentences that were once apparent. Water and fire damage can render text completely unreadable. Wrinkles and folds obscure content essential to meaning. Thankfully, old and new imaging methods can today be combined to rescue “lost” text and make it once again accessible to scholars.Using computer vision techniques like registration, historical images that often represent the most faithful record of the original content of a document can now be combined with those produced by newer technologies, like spectral imaging and 3D modeling. The result is a diachronic digital compilation that enables new scholarly discoveries. Using the collection of fragments from an opened ancient scroll from Herculaneum, PHerc.118, the work outlined in this paper prototypes a process that capitalizes on the best of old and new images to create a single, definitive digital model for scholarly study.
In this essay, I summarize a few ideas inspired by my involvement in the “Coronavation” working group, which spanned 2020’s COVID-19 crisis. Health-care practitioners, computer scientists, and engineers alike, we strive to meet the challenges associated with practice under threat of pandemic with the same ideals driving the rapid, positive developments in health care today: innovation, collaboration and technology convergence, and acquisition of valuable data that leads to better approaches and new ideas. The ideas sketched here, forged by the need for practical pandemic responses, are rooted in those ideals.
Virtual unwrapping is a software pipeline for the noninvasive recovery of texts inside damaged manuscripts via the analysis of three dimensional tomographic data, typically X-ray micro-CT. Recent advancements to the virtual unwrapping pipeline include the use of trained models to perform the “texturing” phase, where the content written upon a surface is extracted from the 3D volume and projected onto a surface mesh representing that page. Trained models are critical for their ability to discern subtle changes that indicate the presence or absence of writing at a given point on the surface. The unique datasets and computational pipeline required to train and make use of these models make it a challenge to develop succinct, reliable, and reproducible research infrastructure. This paper presents our response to that challenge and outlines our framework designed to support the ongoing development of machine learning models to advance the capability of virtual unwrapping. Our approach is designed on the principles of visualization, automation, data access, metadata, and consistent benchmarks.
Today’s digital libraries consist of much more than simple 2D images of manuscript pages or paintings. Advanced imaging techniques – 3D modeling, spectral photography, and volumetric x-ray, for example – can be applied to all types of cultural objects and can be combined to create complex digital representations comprising many disparate parts. In addition, emergent technologies like virtual unwrapping and artificial intelligence (AI) make it possible to create “born digital” versions of unseen features, such as text and brush strokes, that are “hidden” by damage and therefore lack verifiable analog counterparts. Thus, the need for transparent metadata that describes and depicts the set of algorithmic steps and file combinations used to create such complicated digital representations is crucial. At EduceLab, we create various types of complex digital objects, from virtually unwrapped manuscripts that rely on machine learning tools to create born-digital versions of unseen text, to 3D models that consist of 2D photos, multi- and hyperspectral images, drawings, and 3D meshes. In exploring ways to document the digital provenance chain for these complicated digital representations and then support the dissemination of the metadata in a clear, concise, and organized way, we settled on the use of the Metadata Encoding Transmission Standard (METS). This paper outlines our design to exploit the flexibility and comprehensiveness of METS, particularly its behaviorSec, to meet emerging digital provenance metadata needs.
Non-invasive volumetric imaging can now capture the internal structure and detailed evidence of ink-based writing from within the confines of damaged and deteriorated manuscripts that cannot be physically opened. As demonstrated recently on the En-Gedi scroll, our "virtual unwrapping" software pipeline enables the recovery of substantial ink-based text from damaged artifacts at a quality high enough for serious critical textual analysis. However, the quality of the resulting images is defined by the subjective evaluation of scholars, and a choice of specific algorithms and parameters must be available at each stage in the pipeline in order to maximize the output quality.
The noninvasive digital restoration of ancient texts written in carbon black ink and hidden inside artifacts has proven elusive, even with advanced imaging techniques like x-ray-based micro-computed tomography (micro-CT). This paper identifies a crucial mistaken assumption: that micro-CT data fails to capture any information representing the presence of carbon ink. Instead, we show new experiments indicating a subtle but detectable signature from carbon ink in micro-CT. We demonstrate a new computational approach that captures, enhances, and makes visible the characteristic signature created by carbon ink in micro-CT. This previously “unseen” evidence of carbon inks, which can now successfully be made visible, is a discovery that can lead directly to the noninvasive digital recovery of the lost texts of Herculaneum.
Scanning macro‐X‐ray fluorescence (XRF) spectroscopy on works of art provides researchers with rich data sets containing information about material composition and technique of material use in a compelling visual format in the form of element‐specific distribution maps. The accuracy of these maps, however, is influenced by the topography of the object, which ideally is two dimensional, relatively flat and able to be placed parallel to the data collection x-ray optics. In reality, few works of art are truly flat. Small nuances in the visualized elemental intensity may be introduced into element distribution maps by the presence of topography, whether the curve of a centuries-old panel painting, the natural warping of works on paper or parchment, or, in the most extreme cases, in actual three dimensional objects. The inability to confidently ascribe a change in signal intensity to actual elemental composition versus topographically-induced variance, therefore, presents a challenge, particularly when attempting to identify markers of artists’ techniques, compare several objects, or overlay/register images from scanning XRF with those from other imaging modalities. To address this challenge, this paper introduces a new methodology for post-processing scanning XRF data sets to correct for elemental intensity variations as a function of topography. The method augments the acquired XRF data based on a three-dimensional reconstruction of an object and a set of elemental intensity/distance response functions. These response functions act as a calibrated guide for modifying the intensity map based on depth variation. The geometry-based parameters of local surface shape (curvature), distance of the XRF detector from the surface, region of intersection of the incident fluorescence beam with the surface, and the orientation of the incident beam with respect to the surface normal, are each accounted for in the calibration phase as a large set of pre-acquired examples. This provides a mechanism for capturing and understanding the anticipated variations in the macro-XRF data, interpolating the examples in order to smoothly estimate variations, and applying those variations as corrections to macro-XRF data collected on non-planar surfaces.The acquisition and representation of the macro-XRF variation as a function of the geometry is explained, with an emphasis on understanding the parameters that induce the most severe errors in the XRF estimates. The representational framework for collecting, storing, and summarizing calibration data over a large number of scans is discussed, followed by several proof of concept examples, including data from one of the masterpieces of the J. Paul Getty Museum collection: Mummy portrait of a woman (JPGM #81.AP.42), also known as Isidora. This 1st century Romano-Egyptian funeral portrait on wood was originally included in mummy wrappings, and is therefore curved to match the natural curves of the embalmed subject. An XRF scan of Isidora was recently undertaken as part of a long-standing project – Ancient Panel Paintings: Examination, Analysis, and Research (APPEAR) – that seeks to increase our knowledge on the materials and manufacture of paintings of this type. The natural curvature of this panel painting, together with the rich texture typical of the encaustic technique, makes Isidora the perfect candidate to test the proposed methodology.
This paper presents a software framework for the registration and visualization of layered image sets. To demonstrate the utility of these tools, we apply them to the St. Chad Gospels manuscript, relying on images of each page of the document as it appeared over time. An automated pipeline is used to perform non-rigid registration on each series of images. To visualize the differences between copies of the same page, a registered image viewer is constructed that enables direct comparisons of registered images. The registration pipeline and viewer for the resulting aligned images are generalized for use with other data sets.
Computer imaging techniques are commonly used to preserve and share readable manuscripts, but capturing writing locked away in ancient, deteriorated documents poses an entirely different challenge. This software pipeline-referred to as "virtual unwrapping"-allows textual artifacts to be read completely and noninvasively. The systematic digital analysis of the extremely fragile En-Gedi scroll (the oldest Pentateuchal scroll in Hebrew outside of the Dead Sea Scrolls) reveals the writing hidden on its untouchable, disintegrating sheets. Our approach for recovering substantial ink-based text from a damaged object results in readable columns at such high quality that serious critical textual analysis can occur. Hence, this work creates a new pathway for subsequent textual discoveries buried within the confines of damaged materials.
The use of secondary task performance to assess mental workload in a primary task is appealing because the method clearly reflects a central goal of workload assessment – to determine what other functions an operator can undertake while satisfactorily performing the ongoing (primary) technical challenges of a job. For example, does a surgeon performing a suturing task have the cognitive reserves to maintain situation awareness, deal with unanticipated events, or coordinate the efforts of other team members? Unfortunately, secondary task measures have a reputation for being intrusive, artificial, and difficult to use. In the current article, we describe procedures to minimize these concerns, specifically when using an interval production secondary task. Although our suggestions for implementing interval production are based on experience in surgical training environments, the method is grounded in workload assessment research from a variety of other contexts over the past two decades. The methodology appears to be highly adaptable.
The Google Cultural Institute Platform1 is a large-scale system for ingesting, archiving, organizing, and interacting with digital assets of cultural material. This paper explains the components through which the platform contextualizes individual assets in order to enable storytelling. Contextualization is an inverse problem: given assets that are instances of cultural material, infer their precise context and use that as a way to support the storytelling process. The approach is based on three components: extraction, knowledge, and scale. Extraction is the inference of context from two sources of information: explicitly provided metadata, and automatically extracted features. Knowledge is the use of a large reference fact database for further contextualizing an asset based on its descriptors. And scale, achieved through global self-serve, enables massively expanded coverage of the knowledge database and crowdsource potential for metadata refinement. Together these components sustain a storytelling framework and a compelling user experience that has the potential to become the largest repository of cultural information and coherent narrative in history.
As more designers allow users to customize the look and feel of interfaces, users will be required to recognize the implications of their choices on their future performance, comfort, and enjoyment. Understanding the limits of people’s predictive capabilities may be an important component in identifying why people choose one product over another based on ease of use or why people have difficulties identifying tasks that can be performed together. The purpose of this study is to further explore users’ biases for utilizing the amount of white space in the stimulus as a predictor of task difficulty, to validate discrepancies between predicted task difficulty and performance outcomes found in previous research and the human factors literature, and to identify task-specific strategies that are used to anticipate task difficulty. The study uses the NASA Task Load Index (NASA-TLX) in prospective difficulty judgments for these three types of tasks: 1) a stimulus-response compatibility task, 2) a target acquisition task and 3) a perceptual search task. In general, participants predicted lower task demand for designs with more intervening white space. For the visual search task, these estimates of demand were consistent with participants’ actual performance reaction times. However, for the stove design and Fitts’ tasks participants rated tasks that were likely to result in more errors as less challenging suggesting that the type of task is an important factor in participants’ abilities to predict relative task difficulty.
J. Griffioen合作论文数University of Kentucky;Department of Computer Science 5
Rajendra Yavatkar合作论文数System-on-Chip (SoC) Architecture for the Intel Architecture Group at Intel Corporation2