This article addresses several critical aspects of using optically pumped magnetometers (OPMs), focusing on both metrological issues and the enhancement of signal quality. We present a quantitative methodology for OPM measurements standardization and quality evaluation, which is crucial in biomedical applications like magnetoencephalography (MEG). Additionally, we introduce a novel, cost-effective portable active magnetic shielding system-digital adaptive suppression system (DASS)-that represents a significant advancement over traditional analog active shielding. Through comprehensive experimental research, we evaluate first- and third-generation commercial OPMs from QuSpin Inc. in two distinct magnetic environments. Our results demonstrate that the DASS ensures optimal and reliable OPM performance, even in noisy urban settings, surpassing the effectiveness of conventional analog shielding. These findings highlight the need for advanced magnetic shielding solutions to enhance the accuracy and reproducibility of OPM measurements.
The pore-scale numerical modeling of CO 2 injection into natural rock saturated with oil–water mixture was performed using the density functional hydrodynamics approach. The detailed 3D digital model of the sandstone core sample contained over 7 billion cells, which allowed us to perform analysis of oil displacement efficiency at different scales. Utilization of large-size detailed numerical models make it possible to characterize, both qualitatively and quantitatively, the processes at pore scale to the level of detail not achievable on smaller models. The obtained results indicate large-scale effects even on relatively heterogeneous core indicating possible need for multiscale hierarchical models even in heterogeneous cases. This fact imposes the demand for scalability performance on both the software and hardware used in such simulations, as well as the need for adequate modeling upscaling methods.
In the realm of predictive toxicology for small molecules, the applicability domain of QSAR models is often limited by the coverage of the chemical space in the training set. Consequently, classical models fail to provide reliable predictions for wide classes of molecules. However, the emergence of innovative data collection methods such as intensive hackathons have promise to quickly expand the available chemical space for model construction. Combined with algorithmic refinement methods, these tools can address the challenges of toxicity prediction, enhancing both the robustness and applicability of the corresponding models. This study aimed to investigate the roles of gradient boosting and strategic data aggregation in enhancing the predictivity ability of models for the toxicity of small organic molecules. We focused on evaluating the impact of incorporating fragment features and expanding the chemical space, facilitated by a comprehensive dataset procured in an open hackathon. We used gradient boosting techniques, accounting for critical features such as the structural fragments or functional groups often associated with manifestations of toxicity.
Solubility is crucial in organic chemistry and holds significant value in the field of medicinal chemistry. Employing computational and QSPR modeling for solubility estimation is favorable as it reduces experimental costs. However, high-quality experimental data is essential for training these QSPR models. In our study, we compiled a dataset consisting of 54,273 experimental solubility values within a temperature range of 243.15 to 403.15 K in various organic solvents and water. This dataset can be used as a reference for individual values or training solubility QSPR models. We conducted a statistical analysis and identified prevalent patterns in the data. Furthermore, we developed an interactive, parametric t-SNE-based tool to explore the chemical space of solutes. Utilizing this tool, we characterized common scaffolds in the dataset and demonstrated that the chemical space of solutes is extensive and diverse.
Leaf area and biomass are important morphological parameters for in situ plant monitoring since a leaf is vital for perceiving and capturing the environmental light as well as represents the overall plant development. The traditional approach for leaf area and biomass measurements is destructive requiring manual labor and may cause damages for the plants. In this work, we report on the AI-based approach for assessing and predicting the leaf area and plant biomass. The proposed approach is able to estimate and predict the overall plants biomass at the early stage of growth in a non-destructive way. For this reason we equip an industrial greenhouse for cucumbers growing with the commercial off-the-shelf environmental sensors and video cameras. The data from sensors are used to monitor the environmental conditions in the greenhouse while the top-down images are used for training Fully Convolutional Neural Networks (FCNN). The FCNN performs the segmentation task for leaf area calculation resulting in 82% accuracy. Application of trained FCNNs to the sequences of camera images allowed the reconstruction of per-plant leaf area and their growth-dynamics. Then we established the dependency between the average leaf area and biomass using the direct measurements of the biomass. This in turn allowed for reconstruction and prediction of the dynamics of biomass growth in the greenhouse using the image data with 10% average relative error for the 12 days prediction horizon. The actual deployment showed the high potential of the proposed data-driven approaches for plant growth dynamics assessment and prediction. Moreover, it closes the gap towards constructing fully closed autonomous greenhouses for harvests and plants biological safety.
Multi-task learning in deep neural networks has become a topic of growing importance in many research fields, including drug discovery. However, applying multi-task learning poses new challenges in improving prediction performance. This study investigated the potential of training data enrichment to enhance multi-task model prediction quality in drug discovery. The study evaluated four scenarios with varying degrees of information capacity of the training data and applied two types of test data to evaluate prediction performance. We used three datasets: ViralChEMBL, which consisted of binary activities of compounds against viral species, was applied for the classification task; pQSAR(159) and pQSAR(4267), which consisted of bio-activities of compounds and assays from the research of the profile-QSAR method, were applied for regression tasks. We built multi-task models based on the feed-forward DNNs using the PyTorch framework. Our findings showed that training data enrichment could be an effective means of enhancing prediction performance in multi-task learning, but the degree of improvement depends on the quality of the training data. The more unique compounds and targets the training data included, the more new compound-target interactions are required for prediction improvement. Also, we found out that even using multi-task learning, one could not predict the interactions of compounds that are highly dissimilar from those used for model training. The study provides some recommendations for effectively employing multi-task learning in drug discovery to improve prediction accuracy and facilitate the discovery of novel drug candidates.
The rise of deep learning in various scientific and technology areas promotes the development of AI-based tools for information retrieval. Optical recognition of organic structures is a key part of the automated extraction of chemical information. However, this is a challenging task because there is a large variety of representation styles. In this research, we present a Transformer-based artificial neural network to convert images of organic structures to molecular structures. To train the model, we created a comprehensive data generator that stochastically simulates various drawing styles, functional groups, functional group placeholders (R-groups), and visual contamination. We demonstrate that the Transformer-based architecture can gather chemical insights from our generator with almost absolute confidence. That means that, with Transformer, one can fully concentrate on data simulation to build a good recognition model. A web demo of our optical recognition engine is available online at Syntelly platform.
For interpretation of electroencephalography (EEG) and magnetoencephalography (MEG) data, multiple solutions of the respective forward problems are needed. In this paper, we assess performance of the mixed-hybrid finite element method (MHFEM) applied to EEG and MEG modeling. The method provides an approximate potential and induced currents and results in a system with a positive semi-definite matrix. The system thus can be solved with a variety of standard methods (e.g. the preconditioned conjugate gradient method). The induced currents satisfy discrete charge conservation law making the method conservative. We studied its performance on unstructured tetrahedral grids for a layered spherical head model as well as a realistic head model. We also compared its accuracy versus the conventional nodal finite element method ( P1 FEM). To avoid modeling singular sources, we completed our computations with a subtraction approach; the derived expression for the MEG response different from earlier published and involves integration of finite quantities only. We conclude that although the MHFEM is more computationally demanding than the P1 FEM, its use is justified for EEG and MEG modeling on low-resolution head models where P1 FEM loses accuracy.
NMDA (N-methyl-d-aspartate) receptor antagonists are promising tools for the treatment of a wide variety of central nervous system impairments including major depressive disorder. We present here the activity optimization process of a biphenyl-based NMDA negative allosteric modulator (NAM) guided by free energy calculations, which led to a 100 times activity improvement (IC50 = 50 nM) compared to a hit compound identified in virtual screening. Preliminary calculation results suggest a low affinity for the human ether-a-go-go-related gene ion channel (hERG), a high affinity for which was earlier one of the main obstacles for the development of first-generation NMDA-receptor negative allosteric modulators. The docking study and the molecular dynamics calculations suggest a completely different binding mode (ifenprodil-like) compared to another biaryl-based NMDA NAM EVT-101.
Dissociation induced by the accumulation of internal energy via collisions of ions with neutral molecules is one of the most important fragmentation techniques in mass spectrometry (MS), and the identification of small singly charged molecules is based mainly on the consideration of the fragmentation spectrum. Many research studies have been dedicated to the creation of databases of experimentally measured tandem mass spectrometry (MS/MS) spectra (such as MzCloud, Metlin, etc.) and developing software for predicting MS/MS fragments in silico from the molecular structure (such as MetFrag, CFM-ID, CSI:FingerID, etc.). However, the fragmentation mechanisms and pathways are still not fully understood. One of the limiting obstacles is that protomers (positive ions protonated at different sites) produce different fragmentation spectra, and these spectra overlap in the case of the presence of different protomers. Here, we are proposing to use a combination of two powerful approaches: computing fragmentation trees that carry information of all consecutive fragmentations and consideration of the MS/MS data of isotopically labeled compounds. We have created PyFragMS-a web tool consisting of a database of annotated MS/MS spectra of isotopically labeled molecules (after H/D and/or 16O/18O exchange) and a collection of instruments for computing fragmentation trees for an arbitrary molecule. Using PyFragMS, we investigated how the site of protonation influences the fragmentation pathway for small molecules. Also, PyFragMS offers capabilities for performing database search when MS/MS data of the isotopically labeled compounds are taken into account.
Neural Architecture Search (NAS) is a promising and rapidly evolving research area. Training a large number of neural networks requires an exceptional amount of computational power, which makes NAS unreachable for those researchers who have limited or no access to high-performance clusters and supercomputers. A few benchmarks with precomputed neural architectures performances have been recently introduced to overcome this problem and ensure reproducible experiments. However, these benchmarks are only for the computer vision domain and, thus, are built from the image datasets and convolution-derived architectures. In this work, we step outside the computer vision domain by leveraging the language modeling task, which is the core of natural language processing (NLP). Our main contribution is as follows: we have provided search space of recurrent neural networks on the text datasets and trained 14k architectures within it; we have conducted both intrinsic and extrinsic evaluation of the trained models using datasets for semantic relatedness and language understanding evaluation; finally, we have tested several NAS algorithms to demonstrate how the precomputed results can be utilized. We consider that the benchmark will provide more reliable empirical findings in the community and stimulate progress in developing new NAS methods well suited for recurrent architectures.
We performed large-scale numerical simulations using a composite model to investigate the infection spread in a supermarket during a pandemic. The model is composed of the social force, purchasing strategy and infection transmission models. Specifically, we quantified the infection risk for customers while in a supermarket that depended on the number of customers, the purchase strategies and the physical layout of the supermarket. The ratio of new infections compared to sales efficiency (earned profit for customer purchases) was computed as a factor of customer density and social distance. Our results indicate that the social distance between customers is the primary factor influencing infection rate. Supermarket layout and purchasing strategy do not impact social distance and hence the spread of infection. Moreover, we found only a weak dependence of sales efficiency and customer density. We believe that our study will help to establish scientifically-based safety rules that will reduce the social price of supermarket business.
Nowadays, automatic video analysis using state-of-the-art machine learning techniques is highly relevant and finds lots of industrial applications. This approach gives the possibility to solve a set of the problems such as finding the main part of the video, tracking in time by the duration of an action that is labor-intensive and almost impossible to be solved manually or by old-fashioned methods. One of the most promising approaches is the Weakly-supervised method. The Weakly-supervised technique can be used for temporal action localization, which is a challenging task because it uses only video-level labels during the training. The main benefit of this method is that it is not required to create large-scale datasets with temporal annotations, which are quite time-consuming. In this work, we propose to improve the current state-of-the-art in temporal action localization by introducing the new feature extraction procedure. We achieved better results on the RGB stream by making the new feature extraction using more powerful neural network models on the benchmark THUMOS”14 dataset [1]. We also provide a deep analysis of the importance of the feature extraction procedure in the whole workflow for temporal action detection. The results of our investigations can be applied directly for solving a wide range of industrial problems in a robust and accurate manner.
We developed a Transformer-based artificial neural approach to translate between SMILES and IUPAC chemical notations: Struct2IUPAC and IUPAC2Struct. The overall performance level of our model is comparable to the rule-based solutions. We proved that the accuracy and speed of computations as well as the robustness of the model allow to use it in production. Our showcase demonstrates that a neural-based solution can facilitate rapid development keeping the required level of accuracy. We believe that our findings will inspire other developers to reduce development costs by replacing complex rule-based solutions with neural-based ones.
This is the supplementary data for the manuscript: Image2SMILES: Transformer-based Molecular Optical Recognition Engine It contains pairs of image-string, generated from 1M SMILES strings. These strings were randomly chosen from PubChem database.It was prepared using the code, published at https://github.com/syntelly/img2smiles_generator/ To unpack do:tar xvf subset_1M.tar.xz && tar xvf subset_1M_dump.tar.gz && rm subset_1M_dump.tar.gz You'll get the following data: subset_1M.smi - list of 1M source SMILES subset_1M_dump - directory with images subset_1M_result.csv - list of pairs FGSMILES - pathcode, first 3 chars of pathcode are corresponding subdirs in subset_1M_dump subset_1M_fails.csv - list of failed molecules from subset_1M.smi subset_1M_grpcounter.lst - list of counted groups, used in this generation You can generate your own data using https://github.com/syntelly/img2smiles_generator/
The novel high-precision measurement method was proposed. The considered method possesses laser-interferometric precision, does not require to move any external mirrors over full measuring basis, and can be applied to measure distances of 10 − 2−105 m.
We propose a novel approach optimizing passive photonic integrated component topology computations. It is based on Green’s Function Integral Equation method, utilizes weighted optimization methods and fast Toeplitz-like matvecs for GMRES, and exploits GPGPU accelerators.
Humans prefer visual representations for the analysis of large databases. In this work, we suggest a method for the visualization of the chemical reaction space. Our technique uses the t-SNE approach that is parameterized using a deep neural network (parametric t-SNE). We demonstrated that the parametric t-SNE combined with reaction difference fingerprints could provide a tool for the projection of chemical reactions on a low-dimensional manifold for easy exploration of reaction space. We showed that the global reaction landscape projected on a 2D plane corresponds well with the already known reaction types. The application of a pretrained parametric t-SNE model to new reactions allows chemists to study these reactions in a global reaction space. We validated the feasibility of this approach for two commercial drugs, darunavir and montelukast. We believe that our method can help to explore reaction space and will inspire chemists to find new reactions and synthetic ways.
BACKGROUND:Humans are exposed to tens of thousands of chemical substances that need to be assessed for their potential toxicity. Acute systemic toxicity testing serves as the basis for regulatory hazard classification, labeling, and risk management. However, it is cost- and time-prohibitive to evaluate all new and existing chemicals using traditional rodent acute toxicity tests. In silico models built using existing data facilitate rapid acute toxicity predictions without using animals. OBJECTIVES:The U.S. Interagency Coordinating Committee on the Validation of Alternative Methods (ICCVAM) Acute Toxicity Workgroup organized an international collaboration to develop in silico models for predicting acute oral toxicity based on five different end points: Lethal Dose 50 (LD50 value, U.S. Environmental Protection Agency hazard (four) categories, Globally Harmonized System for Classification and Labeling hazard (five) categories, very toxic chemicals [LD50 (LD50≤50mg/kg)], and nontoxic chemicals (LD50>2,000mg/kg). METHODS:An acute oral toxicity data inventory for 11,992 chemicals was compiled, split into training and evaluation sets, and made available to 35 participating international research groups that submitted a total of 139 predictive models. Predictions that fell within the applicability domains of the submitted models were evaluated using external validation sets. These were then combined into consensus models to leverage strengths of individual approaches. RESULTS:The resulting consensus predictions, which leverage the collective strengths of each individual model, form the Collaborative Acute Toxicity Modeling Suite (CATMoS). CATMoS demonstrated high performance in terms of accuracy and robustness when compared with in vivo results. DISCUSSION:CATMoS is being evaluated by regulatory agencies for its utility and applicability as a potential replacement for in vivo rat acute oral toxicity studies. CATMoS predictions for more than 800,000 chemicals have been made available via the National Toxicology Program's Integrated Chemical Environment tools and data sets (ice.ntp.niehs.nih.gov). The models are also implemented in a free, standalone, open-source tool, OPERA, which allows predictions of new and untested chemicals to be made. https://doi.org/10.1289/EHP8495.
In the current article, we present the first solid-state sensor feasible for magnetoencephalography (MEG) that works at room temperature. The sensor is a fluxgate magnetometer based on yttrium-iron garnet films (YIGM). In this feasibility study, we prove the concept of usage of the YIGM in terms of MEG by registering a simple brain induced field-the human alpha rhythm. All the experiments and results are validated with usage of another kind of high-sensitive magnetometers-optically pumped magnetometer, which currently appears to be well-established in terms of MEG.