We present a bicontinuous, minimal surface (the helicoid) as a scaffold on which to define the topology and geometry of yarns in a weft-knitted fabric. Modeling with helicoids offers a geometric approach to simulating a physical manufacturing process, which should generate geometric models suitable for downstream analyses. The centerline of a yarn in a knitted fabric is specified as a geodesic path, with constrained boundary conditions, running along a helicoid at a fixed distance. The shape of the yarn’s centerline is produced via an optimization process over a polyline. The distances between the vertices of the polyline are shortened and a repulsive potential keeps the vertices at a specified distance from the helicoid. These actions and constraints are formulated into a single “cost” function, which is then minimized. The yarn geometry is generated as a tube around the centerline. The optimized configuration, defined for a half loop, is duplicated, reflected, and shifted to produce the centerlines for the multiple stitches that make up a fabric. The approach provides a promising framework for estimating the mechanical behavior/properties of weftknitted fabrics. Fabric-level deformation energy may be estimated by scaling the helicoid scaffold, computing new yarn paths, determining the amount of ensuing yarn stretch, and computing the total amount of yarn stretching energy. Computational results are calibrated and verified with measurements taken from actual yarns and fabrics.
Scientific laboratory notebooks, particularly those in analog, handwritten form, represent a significant yet underutilized data source for computational studies. This paper reports on our research to further develop a pipeline for transforming analog lab notebooks to AI-Ready digital archives. The research is conducted within the framework for Computational Archival Science (CAS), extending CAS principles, drawing from archival practice and computational thinking. We provide background context on laboratory notebook history and current day use, explore CAS as a framework for study, followed by our research goals and methods. Automated extraction results for table records found in the notebooks have an error rate under 5% on a per cell basis. The framework, methods, and our findings seek to advance pipelines for making analog records, both historical and current, accessible and curated for computational research. The findings presented underscore both the accelerating pace of extraction technologies and the importance of more structured, consistent analog documentation practices to support computational transformation and AI-readiness. The conclusion summarizes results and identifies next steps.
Abstract Image‐based machine learning tools are an ascendant ‘big data’ research avenue. Citizen science platforms, like iNaturalist, and museum‐led initiatives provide researchers with an abundance of data and knowledge to extract. These include extraction of metadata, species identification, and phenomic data. Ecological and evolutionary biologists are increasingly using complex, multi‐step processes on data. These processes often include machine learning techniques, often built by others, that are difficult to reuse by other members in a collaboration. We present a conceptual workflow model for machine learning applications using image data to extract biological knowledge in the emerging field of imageomics. We derive an implementation of this conceptual workflow for a specific imageomics application that adheres to FAIR principles as a formal workflow definition that allows fully automated and reproducible execution, and consists of reusable workflow components. We outline technologies and best practices for creating an automated, reusable and modular workflow, and we show how they promote the reuse of machine learning models and their adaptation for new research questions. This conceptual workflow can be adapted: it can be semi‐automated, contain different components than those presented here, or have parallel components for comparative studies. We encourage researchers—both computer scientists and biologists—to build upon this conceptual workflow that combines machine learning tools on image data to answer novel scientific questions in their respective fields.
Shape analysis tasks, including mesh classification, segmentation, and retrieval demonstrate symmetries in Euclidean space and should be invariant to geometric transformations such as rotation and translation. However, existing methods in mesh analysis often rely on extensive data augmentation and more complex analysis models to handle 3D rotations. Despite these efforts, rotation invariance is not guaranteed, which can significantly reduce accuracy when test samples undergo arbitrary rotations, because the analysis method struggles to generalize to the unknown orientations of the test samples. To address these challenges, our work presents a novel approach that employs graph neural networks (GNNs) to analyze mesh-structured data. Our proposed GNN layer, aggregation function, and local pooling layer are equivariant to the rotation, reflection and translation of 3D shapes, making them suitable building blocks for our proposed rotation-invariant network for the classification of mesh models. Therefore, our proposed approach does not need rotation augmentation, and we can maintain accuracy even when test samples undergo arbitrary rotations. Extensive experiments on various datasets demonstrate that our methods achieve state-of-the-art performance.
Collections of analog lab notebooks are an invaluable source of data about research conditions, steps, and outcomes, and in aggregate have the potential to provide new insights into the successes, failures and pedagogy of research laboratories. Unfortunately, these artifacts are increasingly at risk of being lost from the historical scientific record, given limited archiving and an absence of computational and AI readiness. This paper reports on research addressing this challenge by testing mechanisms for transforming digital scans of analog lab notebooks into AI-ready data resources. The research being pursued is framed by the field of computational archival science (CAS) and the aim to utilize analog, research lab notebook data for scientific study. The paper presents background context on archival lab notebooks and CAS, discusses MOF (metal organic frameworks) and COF (covalent organic frameworks) synthesis – the scientific domain of the lab notebooks under study, and details our research methods. We demonstrate a promising approach that automatically segments pages into discrete entry types, extracts the contents of those entries, refines the output and assesses the automated results. These efforts represent a first step towards developing a framework for both improving the usability of archival lab notebooks, and enabling their contents to be used in subsequent scientific inquiry.
Wisdom-of-Crowds-Bots (WoC-Bots) are simple, modular agents working together in a multi-agent environment to collectively make binary predictions. The agents represent a knowledge-diverse crowd, with each agent trained on a subset of available information. A honey-bee-derived swarm aggregation mechanism is used to elicit a collective prediction with an associated confidence value from the agents. Due to their multi-agent design, WoC-Bots can be distributed across multiple hardware nodes, include new features without re-training existing agents, and the aggregation mechanism can be used to incorporate predictions from other sources, thus improving overall predictive accuracy of the system. In addition to these advantages, we demonstrate that WoC-Bots are competitive with other top classification methods on three datasets and apply our system to a real-world sports betting problem, producing a consistent return on investment from 1 January 2021 through 15 November 2022 on most major sports.
Computational archival science (CAS) provides new pathways for research. Biologists, for example, can perform scientific studies by applying AI/ML to digital biological specimen collections and explore questions that were not possible in the analog world. One such approach is the application of computational methods for specimen outlining to assist with specimen identification, morphometry, and other scientific questions. The challenge is to determine how to computationally generate and represent a specimen’s outline. The research presented in this paper addresses this challenge, through the deployment of elliptical Fourier descriptors (EFDs). The paper describes the image processing pipeline for extracting fish outlines, a key morphological feature, and representing the outlines using EFDs. In addition, our research presents the application of machine learning classification on the EFDs. The resulting dataset is well suited for a variety of machine learning-based downstream analyses, including classification by genus and species. Overall, the classification tests produced a 96.3% accuracy, demonstrating the distinguishing nature of the EFDs, and by proxy, the fish outlines as a whole. Broadly, these results indicate the effectiveness of archival specimen usage in machine learning applications, and demonstrate specimen outlining via Fourier descriptors as a computational archival science approach.
Support taking through bracing or leaning while performing manual tasks is known to enhance the capability of the operator.However, simulation of this natural and biomechanically signicant behaviour in a DHM environment is either not possible or calls for signicant expertise and planning on the part of the simulation engineer.While manual simulation is time-consuming and error-prone, an algorithmic procedure is expected to enhance eciency and versatility in the simulation of diverse work environments and what-if scenarios.This paper presents a computational method for determining the location of and reaction at a support point on a given surface that is most advantageous for performing a task.The method also evaluates dierent possible support combinations and the associated optimal postures for performing a given task.The method is illustrated through one-handed reach and supported sitting tasks.Given the task and the environment, the simulation is performed without the need for any user intervention.
In this paper, we describe algorithms that perform loop order analysis of weft-knitted textiles, which build upon the foundational TopoKnit topological data structure and associated query functions. During knitting, loops of yarn may be overlayed on top of each other and then stitched together with another piece of yarn. Loop order analysis aims to determine the front-to-back ordering of these overlapping loops, given a stitch pattern that defines the knitted fabric. Loop order information is crucial for the simulation of electrical current, water, force, and heat flow within functional fabrics. The new algorithms are based on the assumption that stitch instructions are executed row-by-row and for each row the instructions can be executed in any temporal order. To make our algorithms knitting-machine-independent, loop order analysis utilizes precedence rules that capture the order that stitch commands are executed when a row of yarn loops are being knitted by a two-bed flat weft knitting machine. Basing the algorithms on precedence rules allows them to be modified to adapt to the analysis of fabrics manufactured on a variety of knitting machines that may execute stitch commands in different temporal orders. Additionally, we have developed visualization methods for displaying the loop order information within the context of a TopoKnit yarn topology graph.
Researchers seeking to apply computational methods are increasingly turning to scientific digital archives containing images of specimens. Unfortunately, metadata errors can inhibit the discovery and use of scientific archival images. One such case is the NSF-sponsored Biology Guided Neural Network (BGNN) project, where an abundance of metadata errors has significantly delayed development of a proposed, new class of neural networks. This paper reports on research addressing this challenge. We present a prototype workflow for specimen scientific name metadata verification that is grounded in Computational Archival Science (CAS), report on a taxonomy of specimen name metadata error types with preliminary solutions. Our 3-phased workflow includes tag extraction, text processing, and interactive assessment. A baseline test with the prototype workflow identified at least 15 scientific name metadata errors out of 857 manually reviewed, potentially erroneous specimen images, corresponding to a ∼0.2% error rate for the full image dataset. The prototype workflow minimizes the amount of time domain experts need to spend reviewing archive metadata for correctness and AI-readiness before these archival images can be utilized in downstream analysis.
Metadata is a key data source for researchers seeking to apply machine learning (ML) to the vast collections of digitized biological specimens that can be found online. Unfortunately, the associated metadata is often sparse and, at times, erroneous. This paper extends previous research conducted with the Illinois Natural History Survey (INHS) collection (7244 specimen images) that uses computational approaches to analyze image quality, and then automatically generates 22 metadata properties representing the image quality and morphological features of the specimens. In the research reported here, we demonstrate the extension of our initial work to University of the Wisconsin Zoological Museum (UWZM) collection (4155 specimen images). Further, we enhance our computational methods in four ways: (1) augmenting the training set, (2) applying contrast enhancement, (3) upscaling small objects, and (4) refining our processing logic. Together these new methods improved our overall error rates from 4.6 to 1.1%. These enhancements also allowed us to compute an additional set of 17 image-based metadata properties. The new metadata properties provide supplemental features and information that may also be used to analyze and classify the fish specimens. Examples of these new features include convex area, eccentricity, perimeter, skew, etc. The newly refined process further outperforms humans in terms of time and labor cost, as well as accuracy, providing a novel solution for leveraging digitized specimens with ML. This research demonstrates the ability of computational methods to enhance the digital library services associated with the tens of thousands of digitized specimens stored in open-access repositories world-wide by generating accurate and valuable metadata for those repositories.
Machine knitted textiles are complex multi-scale material structures increasingly important in many industries, including consumer products, architecture, composites, medical, and military. Computational modeling, simulation, and design of industrial fabrics require efficient representations of the spatial, material, and physical properties of such structures. We propose a process-oriented representation, TopoKnit, that defines a foundational data structure for representing the topology of weft-knitted textiles at the yarn scale. Process space serves as an intermediary between the machine and fabric spaces, and supports a concise, computationally efficient evaluation approach based on on-demand, near constant-time queries. In this paper, we define the properties of the process space, and design a data structure to represent it and algorithms to evaluate it. We demonstrate the effectiveness of the representation scheme by providing results of evaluations of the data structure in support of common topological operations in the fabric space.
The requirement in Scotland for online/remote teaching during the first two university semesters of the 2020-21 academic year, meant that staff needed to redesign significant parts of the practical curriculum. While interactive online lab teaching resources are available (as an example see reference 1), they can be difficult to adapt to existing learning outcomes. Additionally, online resources often cannot be conjugated sequentially in a way that allows a student (and their mistakes) to follow them through the experiment from 'start to finish'. The authors designed and built 'digital mimics' to allow students to perform the normal experiments in silico using Excel spreadsheets for both UV analysis and HPLC experiments. Methods The digital mimics consisted of three components: solution preparation, spectra/trace generation and analysis. Solution preparation allowed students to do several serial dilutions with fixed glassware sizes. The resultant concentration calculations included volumetric errors and the results were hidden from students. UV spectra and HPLC traces were generated using Gaussian and exponentially modified Gaussian (2) equations respectively. The models contained simulated noise generated using an accumulative MOD expression of a prime number: this avoided the volatile RAND/RANDBETWEEN Excel functions. Students extracted their results by selecting absorbances, or performing peak integration. Several cells and worksheets were hidden from students, thereby ensuring that they were only presented with the information they would normally see in a real lab. Results and conclusion The use of Excel in this manner is a straightforward way of designing digital mimics that can align with existing teaching material (including assessment and marking schemes). Furthermore, Excel is relatively accessible for most University students, and works on both Mac and PCs. The use of digital mimics allows students to make mistakes and rectify them before assignment submission (or the real lab). Spreadsheets were shared online under a CC BY-SA licence (3). This approach was also adapted to mimic tablet manufacture using published mathematical models.
Whole slide images are examined by pathologists and scored according to the Gleason grading system. It is a time-consuming task and may involve assessing variability between different pathologists. In this work, a deep learning system is presented that generates classification maps for whole slide images. This system produces patch-level results first and then predicts a classification map for each prostate cancer slide. The classification maps contain regional cancer severity for each biopsy and are compared with provided mask images. Both provided mask images and predicted mask images are then reviewed by an experienced pathologist to evaluate classification performance. Most state-of-the-art deep learning methods cannot explain how they output classification results. With this work’s classification maps, pathologists can see the regional classification results that explain the algorithm’s classification.
Metadata are key descriptors of research data, particularly for researchers seeking to apply machine learning (ML) to the vast collections of digitized specimens. Unfortunately, the available metadata is often sparse and, at times, erroneous. Additionally, it is prohibitively expensive to address these limitations through traditional, manual means. This paper reports on research that applies machine-driven approaches to analyzing digitized fish images and extracting various important features from them. The digitized fish specimens are being analyzed as part of the Biology Guided Neural Networks (BGNN) initiative, which is developing a novel class of artificial neural networks using phylogenies and anatomy ontologies. Automatically generated metadata is crucial for identifying the high-quality images needed for the neural network’s predictive analytics. Methods that combine ML and image informatics techniques allow us to rapidly enrich the existing metadata associated with the 7,244 images from the Illinois Natural History Survey (INHS) used in our study. Results show we can accurately generate many key metadata properties relevant to the BGNN project, as well as general image quality metrics (e.g. brightness and contrast). Results also show that we can accurately generate bounding boxes and segmentation masks for fish, which are needed for subsequent machine learning analyses. The automatic process outperforms humans in terms of time and accuracy, and provides a novel solution for leveraging digitized specimens in ML. This research demonstrates the ability of computational methods to enhance the digital library services associated with the tens of thousands of digitized specimens stored in open-access repositories world-wide.
Knitting is a manufacturing technique that manipulates yarns to create textiles. This method of producing textiles has been employed by humans for several millennia and is increasingly im portant to many industries. Despite their long-time existence and significant capabilities, com putational modeling, simulation, and design tools have been underutilized for textiles in general, limiting the ability of knitted textiles to be widely deployed and to reach their full industrial po tential. These computational tools require a robust representation and efficient evaluation of the spatial, material and physical properties of textile structures. An example of an efficient model ing method for knitted fabrics is TopoKnit, a process-oriented representation for capturing the topology of weft-knitted textiles. In this paper, we extend TopoKnit and present new algorithms that may be used to determine additional topological structures and assess the manufacturability and structural stability of knitted textiles modeled by this foundational data structure. We com- pare our results to outputs from a commercial software system to confirm the effectiveness and validity of our algorithms
We present a flexible, multi-agent approach to predictive classification problems which uses simple, modular agents that interact and share information socially in an arena with a variable number of participants. Opinion aggregation is accomplished using a honey-bee-derived optimization algorithm that improves accuracy and reduces variance compared with existing weighted and unweighted voter mechanisms. Confidence metrics may be derived from the agent interactions. We apply our system to a data set of 483 de-identified breast cancer patients to predict node-positive or node-negative disease with over 78.5% accuracy in general. When eliminating low-confidence predictions, which leaves 79.5% of patients, classification accuracy improves to 84.5%.
Helicoids have been utilized as a scaffold on which to define the topology and geometry of yarns in a weft-knitted fabric. The centerline of a yarn in the fabric is specified as a geodesic path, with constrained boundary conditions, running along a helicoid at a fixed distance. The properties and constraints of the yarn are formulated into a single "energy" function, which is then minimized to produce the desired resulting models. We present improvements to this approach that address the deficiencies of the original work and extend its capabilities to more complex stitches, such as transfer, tuck and miss. A single bicontinuous surface is described, which replaces discrete helicoids and produces higher quality, continuous yarn models. A new computational method is employed that significantly speeds up the optimization computations. Including offset surfaces with the scaffold, as well as removing sections of the scaffold, allow for the modeling of complex stitches. The improved approach produces superior geometric results, consisting of complex knitting stitches, at a fraction of the computational cost of the previous method.
Knitted fabrics are widely used in clothing because of their distinctive ability to be shaped and formed, which is fundamentally different from the behavior of woven cloth. Since stitches produce complex interactions between yarns, the macroscopic behavior of knitted fabrics depends more on their loop structure and stitch patterns than on the physical properties of the yarn. In order to explore the unique mechanical properties of knitted textiles we have developed a yarn-level model for weft-knitted fabrics that can be used in Finite Element Analysis (FEA) simulations. Producing geometric models of yarns in a knitted material is framed as an optimization problem. In this computing context, a single "cost" function is defined that captures the various required features of the final geometric model. The function is specified in such a way that finding the variable values that results in a minimum function evaluation produces the desired geometric result. The centerlines of the fabric's yarns are defined as Catmull-Rom splines, and the cost function is minimized by adjusting the locations of the spline's control points. The optimization is based on physical parameters such as yarn interpenetration, length of the yarn and bending energy. The optimized models are written to a file which can be directly read by an FEA software. The results show that our approach can create yarn-level models of weft-knitted fabrics consisting of an arbitrary pattern of knit and purl stitches, with a range of sizes, that are suitable for FEA simulations. (C) 2020 Elsevier B.V. All rights reserved.
Mihran Tuceryan合作论文数Department of Computer and Information Science,Indiana University Purdue University Indianapolis7