Layer-by-layer (LbL) assembly of polyelectrolytes is a universal method for controlling composite coatings. When combined with atomic force microscopy (AFM), it is useful strategy to analyze morphology. Optimisation of the LbL technique is bottlenecked by manual sample preparation, and there is also a lack of standardised analytical pipelines to convert characterisation data into machine-readable descriptors. We address both bottlenecks by combining a collaborative robotic dip-coating platform with an automated feature-extraction backend for AFM topography. The robotic arm handles the substrate (Si wafer), immerses it in polyelectrolyte solution, and prepares samples for AFM study without operator intervention. Resulting AFM scans are transferred into a reproducible pipeline that applies a logged preprocessing chain and computes twelve groups of physically interpretable descriptors per sample: ISO 25,178 areal roughness, height-distribution statistics, radial power spectrum with Hurst and fractal exponents, two-dimensional autocorrelation, local-patch minima/maxima distributions, persistent homology of patch point clouds, gliding-box lacunarity, and a recipe-derived block of layer composition, sequence k-grams and RDKit monomer descriptors. Every descriptor is stored with the originating method version and parameters, supporting idempotent recomputation and side-by-side methodological evolution. The pipeline was applied to polyethyleneimine (PEI)/polystyrene sulfonate (PSS) assemblies on Si substrates, and an unsupervised principal component projection of the scalar descriptors already separates samples by the number of deposited layers, demonstrating that the resulting feature vectors carry coating-state information at a level useful for downstream machine learning. This work establishes a traceable, data‑driven bridge from synthesis instructions to high‑dimensional surface descriptors, demonstrating for the first time an end‑to‑end pipeline where robotic LbL assembly, versioned AFM processing, and hybrid topological‑chemical feature extraction converge into a single machine‑learning‑ready dataset.
This study investigates the ability to characterize the functional-group composition of coking feedstocks using infrared spectroscopy combined with molecular modeling and data-driven analysis. FTIR spectra of industrial feedstocks were analyzed using descriptor-based metrics capturing aromatic, saturated, and oxygen-containing contributions. To evaluate spectral-composition relationships under controlled conditions, molecular dynamics (MD) simulations were used to generate synthetic IR spectra of representative hydrocarbon mixtures with known compositions. Linear chemometric models were found to adequately recover bulk compositional fractions, while neural network model improved prediction of CH3/CH2 ratio. Multivariate curve resolution of experimental spectra identified three reproducible latent components that correlate consistently with predefined spectral indices. The proposed workflow enables semi-quantitative, descriptor-based comparison of feedstock compositions prior to downstream processing.
The present study focuses on the problem of taste recognition of coffee by the means of electrochemical imprints. Cyclic voltammetry facilities and machine learning (ML) techniques made it possible to create a combined method of discrimination of taste profile operating with such sensory categories as sweetness, bitterness, acidity as well as overall quality. The electrochemical responses of four different electrodes in nearly 200 different samples of coffee contributed to the essential databases used for training ML-models via supervised algorithms. Best performance of quality recognition was achieved with the help of LogisticRegression using a gold electrode as the sensor (F1=0.89), while acidity and sweetness were recognized in the most efficient way by boosting algorithm at the Ni and Cu sensor electrodes (F1=0.87 and F1=0.72), respectively, and XGBClassifier was the most effective algorithm to estimate bitterness at the gold electrode (F1=0.63). Gas chromatography-mass spectrometry (GC-MS) identified key volatiles, while Density functional theory (DFT)/docking simulations confirmed electrode adsorption of caffeine, acids, and furanmethanol, supporting electrochemical fingerprints as taste proxies.
This work proposes several machine learning models that predict B3LYP-D4/def-TZVP outputs from HF-3c outputs for supramolecular structures. The data set consists of 1031 entries of dimer, trimer, and tetramer cyclic structures, containing both molecules with heteroatoms in the ring and without. Six quantum chemistry descriptors and features are calculated by using both computational methods: Gibbs energy, electronic energy, entropy, enthalpy, dipole moment, and band gap. Statistical analysis shows a good correlation between energy properties and bad correlation only for the dipole moment. Machine learning models are separated into three groups: linear, tree-based, and neural networks. The best models for the prediction of density functional theory features are LASSO for linear, XGBoost for tree-based, and single-layer perceptron for neural networks with energy-related features having the best prediction values and dipole moment having the worst.
Macrocyclic compounds enable diverse applications utilizing their cavity for guest encapsulation. Current approaches to cavity volume calculation in tunnel-like cavities suffer from weak predictability due to mouth opening ambiguity. Here tessellation and divide-and-conquer approaches to the cavity volume prediction in calixarene macrocycles featuring convex-hull based determination of boundaries of tunnel-like cavities and a user-friendly CaviDAC tool for academic researchers and specialists in chemo(bio)informatics are presented. The cavity volumes of the basket-, barrel-, and sandglass-shaped calixarenes calculated by triangular tessellation and quickhull algorithms show the divergence of less than 3.0%. CaviDAC software outperforms common cavity volume calculation software in terms of accuracy in the cavitands with ill-defined cavity opening, which is promising for the computational screening of novel materials with molecular-level porosity.
In this report, we present electrochemical immunosensors for the detection of S. aureus bacteria on the basis of SPCE/PEI/& Acy;BSA/PSS layer-by-layer assembly as a recognition element. QCM measurements and AFM imaging ensure effective adhesion of S. aureus antibody to PEI surface and its strong interactions with analyte through the PSS polyelectrolyte layer. Impedimetric detection of S. aureus gives the LOD of 1000 CFU/mL and the linear range from 104 to 107 CFU/mL and features facile assembly of recognition element and easy sampling. Voltammetric detection of the formation of the sandwich immunocomplex with secondary antibody in the outermost layer (AB-AG-AB-HRP) not only decreases the detection limit to 230 CFU/mL and expands the linear range of detection to 103-108 CFU/mL, but also could detect S. aureus bacteria with a portable open-source custom potentiostat in voltammetric mode, which is promising for non-invasive point-of-care monitoring of pathogens and addresses issues of antibody-based sensors, such as high cost and difficult chemical modification.
The increasing complexity in designing nanostructured materials for electronics, biomedicine, and energy applications requires advanced computational methods to enhance research efficiency and minimize experimental costs. This study proposes an innovative agent-based retrieval-augmented generation (RAG) system integrated with large language models (LLMs) to automate the extraction and analysis of scientific information from extensive literature databases, specifically targeting nanostructured materials developed via two-photon polymerization (2PP). In addition to extracting and analyzing scientific data, our approach emphasizes understanding how these nanostructured materials interact with cells, which is crucial for controlling their application in biomedicine. The developed platform demonstrates robust semantic accuracy (cosine similarity: 0.82) and high overall task precision (0.81), significantly reducing the likelihood of misinformation by incorporating dynamic query refinement mechanisms. The intuitive, user-friendly interface facilitates quick access to relevant scientific data, thereby improving researchers' productivity and enabling more accurate experimental planning. Although the system exhibits certain limitations regarding domain-specific terminology coverage, further fine-tuning and specialized training are anticipated to enhance its performance and reliability for advanced scientific applications.
This study presents a machine learning (ML)/Artificial Intelligence (AI) approach to classify types of sparkling wines (champagnes) and their respective containers using image data of bubble patterns. Sparkling wines are oversaturated with dissolved CO2, which results in extensive bubbling when the wine bottle is uncorked. The nucleation and properties of bubbles depend on the chemical composition of the wine, the properties of the glass, and the concentration of CO2. For carbonated liquids supersaturated with CO2, the interaction of natural and cavitation bubbles is a non-trivial matter. We study ultrasonic cavitation bubbles in two types of sparkling wines and two types of glasses with the computer vision (CV) analysis of video images and clustering using an artificial neural network (NN) approach. By integrating a segmentation NN to filter out irrelevant frames and applying the Contrastive Language-Image Pre-Training (CLIP) NN for feature embedding, followed by TabNet for classification, we demonstrate a novel application of ML/AI for distinguishing champagne characteristics. The results show that the bubbles are significantly different to be classified by the ML techniques for different types of wine and glasses. Consequently, our study demonstrates that CV/AI/ML analysis of ultrasound cavitation bubbles can be used to analyze carbonated liquids.
This study introduces a novel heuristic phenomenological model for analyzing the evolution of contact areas on rough surface. Contrasting with traditional methods, it employs a cut-off threshold approach to track numerical and topological metrics across different deformation stages. The model quantifies contact area distributions, nested sub-regions, and self-affine parameters, revealing universal trends across scales spanning nanometers to kilometers. Metrics for synthetically generated isotropic surfaces with Hurst exponents H = 2.5 and 3.5 correlate closely with those from AFM and SEM experimental datasets, respectively. In addition, the model has been tested on NASA's SRTM datasets. Cross-correlation demonstrate significant similarities in numerical and topological metrics across diverse measurement techniques, surface types, and scales, highlighting the method's robustness and calibration-free scale invariance. This approach bridges gaps in multiscale tribological analysis, offering deeper insights into frictional transitions and surface interactions. Beyond tribology and materials science, this general approach enables fundamental characterization of surface morphology as such, making it applicable to diverse fields including geomorphology, biomimetics, and nanotechnology.
Hit identification is a central challenge in early drug discovery, traditionally requiring substantial experimental resources. Recent advances in artificial intelligence, particularly large language models (LLMs), have enabled virtual screening methods that reduce costs and improve efficiency. However, the growing complexity of these tools has limited their accessibility to wet-lab researchers. Multi-agent systems offer a promising solution by combining the interpretability of LLMs with the precision of specialized models and tools. In this work, we present MADD, a multi-agent system that builds and executes customized hit identification pipelines from natural language queries. MADD employs four coordinated agents to handle key subtasks in de novo compound generation and screening. We evaluate MADD across seven drug discovery cases and demonstrate its superior performance compared to existing LLM-based solutions. Using MADD, we pioneer application of AI-first drug design to five biological targets and release the identified hit molecules. Finally, we introduce a new benchmark of query-molecule pairs and docking scores for over three million compounds to contribute to the agentic future of drug design.
An overview of modern technologies in the field of laboratory automation, such as robotics and the use of computer vision, is presented. The main methods and equipment required for the automation process are discussed. A brief description of the implementation of collaborative robots and computer vision systems in automated laboratory work-stations is given, and a specific example of a centrifugation procedure is considered.
This study demonstrated a machine learning approach to predict the photocatalytic properties of graphitic carbon nitride (g-C3N4) 3 N 4 ) depending on its synthesis parameters to enhance photocatalytic hydrogen production. In connection with the task, a database was experimentally formed to prepare g-C3N4 3 N 4 samples by heat treatment of nitrogen-containing precursors in air at a temperature of 450-600 degrees C with varying time and heating rates of the synthesis. Physicochemical analyses characterized the materials, including X-ray diffraction and low- temperature nitrogen adsorption. Several machine learning algorithms were used to process the obtained data, which showed a high-efficiency R2 2 above 0.9. Hyperparameters were optimized for each model using different preprocessing methods. In addition, the importance of features for selecting the most effective sample was assessed. For convenience, a web application was created with the ability to expand the database to work with machine learning models.
The use of collaborative robots (Cobots) for materials development in chemical laboratories is currently of high priority. Herein, the Cobot is used for autonomous continued analysis and synthesis of graphene oxide–polyethyleneimine‐based membrane to unify a method and prospects for big data collection are shown. Membranes have already demonstrated a selective affinity to potassium cations and promised to adjust permeability for other cations by changing pH. The Cobot allows a variation of membrane properties by its composition modification. The present strategy combines a novel perspective of material production by Cobots and the application of machine learning. Moreover, the current approach can be adapted for different modern chemical laboratories for various scientific research, and the proper workflow is provided.
We present an electrochemical platform designed to reduce time of Escherichia coli bacteria detection from 24-48-hours to 30 minutes. The presented approach is based on a system which includes gallium-indium (eGaIn) alloy to provide conductivity and a hydrogel system to preserve bacteria and their metabolic species during the analysis. The work is dedicated to accurate and fast detection of Escherichia coli bacteria in different environments with the supply of machine learning methods. Electrochemical data obtained during the analysis is processed via multilayer perceptron model to identify i.e. predict bacterial concentration in the samples. The performed approach provides the effectiveness of bacteria identification in the range of 102 to 109 colony forming units per ml with the average accuracy of 97%. The proposed bioelectrochemical system combined with machine learning model is prospective for food analysis, agriculture, biomedicine.
The nanoscale topographic features of surfaces, such as vertical, lateral, and multiscale structures, are essential for understanding and identifying correlations with properties in various applications. These features are critical in applications ranging from biomedical devices to electronic components, as they play a significant role in determining the functionality of these systems. Despite the capabilities of traditional surface analysis methods, which rely on standard vertical profile and area measurements, these techniques often fail to capture the subtle and specific features that differentiate between surfaces with different morphologies. To address this challenge, this paper proposes applying of a topological data analysis of atomic force microscopy data to create a unique topological signature for each surface. This approach can help us better understand complex relationships between topography and functional characteristics in various applications and enable further advanced surface comparison.
In the realm of predictive toxicology for small molecules, the applicability domain of QSAR models is often limited by the coverage of the chemical space in the training set. Consequently, classical models fail to provide reliable predictions for wide classes of molecules. However, the emergence of innovative data collection methods such as intensive hackathons have promise to quickly expand the available chemical space for model construction. Combined with algorithmic refinement methods, these tools can address the challenges of toxicity prediction, enhancing both the robustness and applicability of the corresponding models. This study aimed to investigate the roles of gradient boosting and strategic data aggregation in enhancing the predictivity ability of models for the toxicity of small organic molecules. We focused on evaluating the impact of incorporating fragment features and expanding the chemical space, facilitated by a comprehensive dataset procured in an open hackathon. We used gradient boosting techniques, accounting for critical features such as the structural fragments or functional groups often associated with manifestations of toxicity.
Machine-vision analysis of a frame with a gas bubble in the resonance mode ( n = 8).
The surface roughness of layer-by-layer (LbL) polyelectrolytes is studied by atomic force microscopy (AFM) and analyzed with novel methods including topological data analysis (TDA) and machine learning (ML) to correlate multiscale roughness with the number of bilayers and to recognize the types of polyelectrolytes (PEs). LbL PEs composed of one to four bilayers of (1) polyethylenimine (PEI)/poly(sodium 4-styrenesulfonate) (PSS), (2) PEI/poly(acrylic acid) (PAA), and (3) PEI/MXene rigid flakes are deposited on a smooth silicon wafer. With a growing number of bilayers, the roughness changes from a smooth surface to an equilibrium rough profile. The AFM study of the surface morphology demonstrates that surface roughness is multiscale, with smaller features imposed on larger ones. Roughness data is filtered from measurement resolution artifacts, and several methods are applied: correlation length, statistics of the distribution of extremes in trimmed images, and TDA barcodes and persistence diagrams of simplexes in 8D data space. An ML algorithm is used to determine the number of bilayers in a PE. Roughness analysis indicates a gradual transition from a smooth to a rough surface with saturation at three to four bilayers and the existence of multiscale roughness invariance.