Ligand-based drug discovery (LBDD) relies on making use of known binders to a protein target to find structurally diverse molecules similarly likely to bind. This process typically involves a brute force search of the known binder (query) against a molecular library using some metric of molecular similarity. One popular approach overlays the pharmacophore-shape profile of the known binder to 3D conformations enumerated for each of the library molecules, computes overlaps, and picks a set of diverse library molecules with high overlaps. While this virtual screening workflow has had considerable success in hit diversification, scaffold hopping, and patent busting, it scales poorly with library sizes and restricts candidate generation to existing library compounds. Leveraging recent advances in voxel-based generative modelling, we propose a pharmacophore-based generative model and workflows that address the scaling and fecundity issues of conventional pharmacophore-based virtual screening. We introduce VoxCap, a voxel captioning method for generating SMILES strings from voxelised molecular representations. We propose two workflows as practical use cases as well as benchmarks for pharmacophore-based generation: de-novo design, in which we aim to generate new molecules with high pharmacophore-shape similarities to query molecules, and fast search, which aims to combine generative design with a cheap 2D substructure similarity search for efficient hit identification. Our results show that VoxCap significantly outperforms previous methods in generating diverse de-novo hits. When combined with our fast search workflow, VoxCap reduces computational time by orders of magnitude while returning hits for all query molecules, enabling the search of large libraries that are intractable to search by brute force.
Machine learning models for 3D molecular property prediction typically rely on atom-based representations, which may overlook subtle physical information. Electron density maps – the direct output of X-ray crystallography and cryo-electron microscopy – offer a continuous, physically grounded alternative. We compare three voxel-based input types for 3D convolutional neural networks (CNNs): atom types, raw electron density, and density gradient magnitude, across two molecular tasks – protein-ligand binding affinity prediction (PDBbind) and quantum property prediction (QM9). We focus on voxel-based CNNs because electron density is inherently volumetric, and voxel grids provide the most natural representation for both experimental and computed densities. On PDBbind, all representations perform similarly with full data, but in low-data regimes, density-based inputs outperform atom types, while a shape-based baseline performs comparably – suggesting that spatial occupancy dominates this task. On QM9, where labels are derived from Density Functional Theory (DFT) but input densities from a lower-level method (XTB), density-based inputs still outperform atom-based ones at scale, reflecting the rich structural and electronic information encoded in density. Overall, these results highlight the task- and regime-dependent strengths of density-derived inputs, improving data efficiency in affinity prediction and accuracy in quantum property modeling.
Generative modeling is increasingly important for data-driven computational design. Conventional approaches pair a generative model with a discriminative model to select or guide samples toward optimized designs. Yet discriminative models often struggle in data-scarce settings, common in scientific applications, and are unreliable in the tails of the distribution where optimal designs typically lie. We introduce generative property enhancer (GPE), an approach that implicitly guides generation by matching samples with lower property values to higher-value ones. Formulated as conditional density estimation, our framework defines a target distribution with improved properties, compelling the generative model to produce enhanced, diverse designs without auxiliary predictors. GPE is simple, scalable, end-to-end, modality-agnostic, and integrates seamlessly with diverse generative model architectures and losses. We demonstrate competitive empirical results on standard _in silico_ offline (non-sequential) protein fitness optimization benchmarks. Finally, we propose iterative training on a combination of limited real data and self-generated synthetic data, enabling extrapolation beyond the original property ranges.
Generative models for structure-based drug design are often limited to a specific modality, restricting their broader applicability. To address this challenge, we introduce FuncBind, a framework based on computer vision to generate target-conditioned, all-atom molecules across atomic systems. FuncBind uses neural fields to represent molecules as continuous atomic densities and employs score-based generative models with modern architectures adapted from the computer vision literature. This modality-agnostic representation allows a single unified model to be trained on diverse atomic systems, from small to large molecules, and handle variable atom/residue counts, including non-canonical amino acids. FuncBind achieves competitive in silico performance in generating small molecules, macrocyclic peptides, and antibody complementarity-determining region loops, conditioned on target structures. FuncBind also generated in vitro novel antibody binders via de novo redesign of the complementarity-determining region H3 loop of two chosen co-crystal structures. As a final contribution, we introduce a new dataset and benchmark for structure-conditioned macrocyclic peptide generation.
This paper emphasizes the need to broaden organizational perspectives through Open X, which promotes sharing and collaboration over selfishness and competition, instead of that industrial intellectual protection through patents can divert resources essential for the growth of organizations. Faced with new realities, organizations need different management approaches with the potential to transform the reindustrialization resulting from deindustrialization into a Neo-industrialization 2.0. It does not mean tearing down or creating new boundaries but an open culture where organizational efforts have social relevance. In the face of economic interests, Open X can make organizational outcomes more plentiful and robust.
We presents VoxBind, a new score-based generative model for 3D molecules conditioned on protein structures. Our approach represents molecules as 3D atomic density grids and leverages a 3D voxel-denoising network for learning and generation. We extend the neural empirical Bayes formalism (Saremi & Hyvärinen, 2019) to the conditional setting and generate structure-conditioned molecules with a two-step procedure: (i) sample noisy molecules from the Gaussian-smoothed conditional distribution with underdamped Langevin MCMC using the learned score function and (ii) estimate clean molecules from the noisy samples with single-step denoising. Compared to the current state of the art, our model is simpler to train, significantly faster to sample from, and achieves better results on extensive in silico benchmarks—the generated molecules are more diverse, exhibit fewer steric clashes, and bind with higher affinity to protein pockets.
We present NEBULA, the first latent 3D generative model for scalable generation of large molecular libraries around a seed compound of interest. Such libraries are crucial for scientific discovery, but it remains challenging to generate large numbers of high quality samples efficiently. 3D-voxel-based methods have recently shown great promise for generating high quality samples de novo from random noise (Pinheiro et al., 2023). However, sampling in 3D-voxel space is computationally expensive and use in library generation is prohibitively slow. Here, we instead perform neural empirical Bayes sampling (Saremi Hyvarinen, 2019) in the learned latent space of a vector-quantized variational autoencoder. NEBULA generates large molecular libraries nearly an order of magnitude faster than existing methods without sacrificing sample quality. Moreover, NEBULA generalizes better to unseen drug-like molecules, as demonstrated on two public datasets and multiple recently released drugs. We expect the approach herein to be highly enabling for machine learning-based drug discovery. The code is available at https://github.com/prescient-design/nebula
We introduce a new functional representation for 3D molecules based on their continuous atomic density fields. Using this representation, we propose a new model based on neural empirical Bayes for unconditional 3D molecule generation in the continuous space using neural fields. Our model, FuncMol, encodes molecular fields into latent codes using a conditional neural field, samples noisy codes from a Gaussian-smoothed distribution with Langevin MCMC, denoises these samples in a single step and finally decodes them into molecular fields. FuncMol performs all-atom generation of 3D molecules without assumptions on the molecular structure and scales well with the size of molecules, unlike most existing approaches. Our method achieves competitive results on drug-like molecules and easily scales to macro-cyclic peptides, with at least one order of magnitude faster sampling. The code is available at https://github.com/prescient-design/funcmol.
One of the most frequent, most expensive and potentially more impactful tasks in crop management is surveying and scouting the fields for problems in crop development. Any biotic / abiotic stress undetected becomes a bigger problem to solve later, with a potentially cascading effect on yield and/or quality and, subsequently, crop value. For annual crops (such as corn, soy, etc.) this can be solved in a cost-effective way with Sentinel data. For permanent crops planted in rows (such as vineyards), the interference from the inter-row makes it much more challenging. Under a contract for the European Space Agency (ESA), Spin.Works has been developing an early anomaly detection system based on fusion of Sentinel-2 and UAV imagery, targeting an update rate of 5 days. The early anomaly detection is applied to vineyards, particularly, for nutrient and water stresses. The early anomaly detection system is integrated into Spin.Works’ MAPP.it platform and its development is being carried out in close cooperation with the internal R&D group of Sogrape Vinhos, Portugal's largest winemaker and a long-standing MAPP.it user.
We propose a new score-based approach to generate 3D molecules represented as atomic densities on regular grids. First, we train a denoising neural network that learns to map from a smooth distribution of noisy molecules to the distribution of real molecules. Then, we follow the _neural empirical Bayes_ framework [Saremi and Hyvarinen, 2019] and generate molecules in two steps: (i) sample noisy density grids from a smooth distribution via underdamped Langevin Markov chain Monte Carlo, and (ii) recover the "clean" molecule by denoising the noisy grid with a single step. Our method, _VoxMol_, generates molecules in a fundamentally different way than the current state of the art (ie, diffusion models applied to atom point clouds). It differs in terms of the data representation, the noise model, the network architecture and the generative modeling algorithm. Our experiments show that VoxMol captures the distribution of drug-like molecules better than state of the art, while being faster to generate samples.
Spin.Works has been developing its MAPP.it platform and implementing features in close cooperation with the internal R&D group of Sogrape Vinhos, Portugal's largest winemaker and a long-standing MAPP.it user. Borne of such cooperation were a number of tools that are currently available or in late-stage development in MAPP.it: Information register and filtering capabilities for all plots in a property; combining high spatial resolution data from drone with high temporal resolution data from satellite; availability of past years' data enabling inquiry into historical comparisons and trends; simple statistical analysis such as plant distribution perpercentile, dynamic cut-off points for zoning tools or smoothing; identification, counting, and georeferencing of gaps in the vineyards (dead or otherwise lost plants); plot variability measurement; high degree of exportability and interoperability, such as ability to download both raster and vector data or export maps/analysis as pdf files; mobile app enabling in-field data consultation and analysis, as well as georeferenced notes and photos. Using MAPP.it, Sogrape has streamlined its viticulture management, supporting more efficient daily planning from vineyard managers, evaluating the effect of management decisions on annual and monthly time-frames, explaining the underpinning reasons for observed vineyard block variability and scheduling harvests according to plant vigour and maturity levels (combination of MAPP.it and maturity control data). MAPP.it and Sogrape will continue to cooperate in the eco-development of the plat form to improve the features and functionality of the MAPP.it service taking advantage of developments in satellite data availability and computer support edgeomatic analysis, hopefully leading to easy, quick, and accurate methods for estimating water stress risk, carbon balance potentials, and ecosystem management with nature and biodiversity conservation indicators.
Agility is characterized particularly by rapid response (high velocity of response) and proactivity. These features are very important considering Industry 4.0 skills to be acquired by engineering students. Social Network based Education is a methodology for students' effective learning and adoption of Industry 4.0 concepts and skills, and, in the context of this paper, of the agility concept and skill. This paper contributes to the definition of agility measures of student groups to evaluate students group agility in realization of different tasks, i.e. different student assignments, and consequently, of the effectiveness of learning agility concept. For the case study, two groups of students, from different school years, were monitored and their agility was measured based on the defined agility measures, providing objective agility measures of each group. The measures provide an evaluation of the effectiveness of learning agility concept, as well as which group of students adopted more effectively the concept of agility. The measures proposed could be further considered as the agility measures to be applied in companies.
Accurately modeling and predicting RNA biology has been a long-standing challenge, bearing significant clinical ramifications for variant interpretation and the formulation of tailored therapeutics. We describe a foundation model for RNA biology, “BigRNA”, which was trained on thousands of genome-matched datasets to predict tissue-specific RNA expression, splicing, microRNA sites, and RNA binding protein specificity from DNA sequence. Unlike approaches that are restricted to missense variants, BigRNA can identify pathogenic non-coding variant effects across diverse mechanisms, including polyadenylation, exon skipping and intron retention. BigRNA accurately predicted the effects of steric blocking oligonucleotides (SBOs) on increasing the expression of 4 out of 4 genes, and on splicing for 18 out of 18 exons across 14 genes, including those involved in Wilson disease and spinal muscular atrophy. We anticipate that BigRNA and foundation models like it will have widespread applications in the field of personalized RNA therapeutics.
The design process is unrepeatable, dynamic, fluid, and dependent on the context mood. It is part of an ever-changing social system. The purpose of design is to create, to bring out solutions to regulate the increasing complexity of social, economic, environmental, and cultural systems. Blending science and art, the designer chooses from several hypotheses, deconstructs reality, and rebuilds it.
Machine learning systems are typically trained and tested on the same distribution of data. However, in the real world, models and agents must adapt to data distributions that change over time. Previous work in computer vision has proposed using image corruptions to model this change. In contrast, we propose studying models under a setting more similar to what an agent might encounter in the real world. In this setting, models must adapt online without labels to a test distribution that changes in semantics. We define two types of semantic distribution shift, one or both of which can occur: \emph{static shift}, where the test set contains labels unseen at train time, and \emph{continual shift}, where the distribution of labels changes throughout the test phase. Using a dataset that contains both class and attribute labels for image instances, we generate shifts by changing the joint distribution of class and attribute labels. We compare to previously proposed methods for distribution adaptation that optimize a fixed self-supervised criterion at test time or a meta-learning criterion at train time. Surprisingly, these provide little improvement in this more difficult setting, with some even underperforming a static model that does not change parameters at test time. In this setting, we introduce two models that ``learn to adapt''---via recurrence and learned Hebbian update rules. These models outperform both previous work and static models under both \emph{static} and \emph{continual} semantic shifts, suggesting that ``learning to adapt'' is a useful capability for models and agents in a changing world.
One of the key challenges in future space explorers is the ability to carry out complex mission profiles while avoiding constant ground support until arrival at the mission target. A key point is precise self-knowledge of location and attitude. Over the last several years there have been many demonstrations of how to use visual cues to enable safe and precise execution of key mission phases, including in large-scale missions (most recently on NASA's Mars Perseverance). Nevertheless, this transition is sure to occur at a faster pace on small missions due to their comparatively low cost. We have investigated how to forego entirely ground-based navigation throughout a mission - between launch separation and target arrival. We propose to use primarily just three small optical instruments (two star trackers and one high-resolution camera), along with a high-performance processing unit, while considering complementary sensors such as IMUs and ranging instruments for critical events. We describe two different mission profiles, a lunar landing and an asteroid mission. We have calculated suitable trajectories to reach our targets, and describe appropriate image processing techniques to reach the required positioning performance. We also describe the covariance analyses that guide both trajectory correction timeline and the observation schedule. We have built prototype hardware instrument to test our progress towards achieving this goal, and have tested it under conditions representative of a real mission. Finally, we are currently qualifying cameras for In-Orbit Demonstrations in early 2023 to inform our next steps.
AEROS is a 3U CubeSat pathfinder toward a future ocean-observing constellation, targeting the Portuguese Atlantic region. AEROS features a miniaturized, high-resolution Hyperspectral Imager (HSI), a 5MP RGB camera, and a Software Defined Radio (SDR). The sensor generated data will be processed and aggregated for end-users in a new web-based Data Analysis Center (DAC). The HSI has 150 spectrally contiguous bands covering visible to near-infrared with 10 nm bandwidth. The HSI collects ocean color data to support studies of oceanographic characteristics known to influence the spatio-temporal distribution and movement behavior of marine organisms. Usage of an SDR expands AEROS's operational and communication range and allows for remote reconfiguration. The SDR receives, demodulates, and retransmits short duration messages, from sources including tagged marine organisms, autonomous vehicles, subsurface floats, and buoys. The future DAC will collect, store, process, and analyze acquired data, taking advantage of its ability to disseminate data across the stakeholders and the scientific network. Correlation of animal-borne Argos platform locations and oceanographic data will advance fisheries management, ecosystem-based management, monitoring of marine protected areas, and bio-oceanographic research in the face of a rapidly changing environment. For example, correlation of oceanographic data collected by the HSI, geolocated with supplementary images from the RGB camera and fish locations, will provide researchers with near real-time estimates of essential oceanographic variables within areas selected by species of interest.
The paper addresses the question: what is engineering? We intuitively know engineering applications such as manufacturing, production, industry, management, business. The answer is not consensual because it is not easy. Furthermore, the ontological question brings us to a second question. What distinguishes engineering from other areas? It is the creative ability that distinguishes engineering. And this artificial faculty only exists in Design. Epistemology in science promotes the existence of herds, increasingly specialized groups of knowledge production. Nevertheless, engineers assume themselves as makers, and in the growing diversity promoted by specialization, they will certainly give different answers when asked about their work. We aggregate all of them as sign-makers. Therefore, engineering is Design and only Design. We reject other views. The argument presented on the phenomenological level considers them false. This paper demonstrates that it is mandatory to create a distinctive sign, which places engineering as relevant in organizations. Without the sign described in semiotics, engineering, which could pretend to be everything, becomes trivial.
Learning Factory could be considered as an instrument for effective learning and training of advanced manufacturing concepts, through true connection between universities and companies. A supporting infrastructure, i.e. an implementation architecture, should be designed in such way to strengthen this objectives. This paper presents a contribution to the Learning Factory architecture implementation, considering different implementation infrastructures: physical stationary infrastructure, physical mobile infrastructure, internet-based infrastructure and blended infrastructure. A Learning Factory implementation framework is presented considering three dimensions: education paradigm, implementation infrastructure, and the learning object. Additionally, two types of the Learning Factory architectures, the physical stationary and internet-based implementation infrastructures, designed and implemented in an ongoing course on industrial engineering are presented as well.
The paper presents the ICARUS Pedagogical Framework, or Reference model to address development of innovative pedagogical approaches to overcome the actual effectiveness problems in development, acquisition and application of the required knowledge and skills for the concepts related to Industry 4.0. The aim of the Pedagogical Framework is minimum twofold: 1) to define and guide and applications of innovative pedagogical methods to explore and address the needs of HEI educators and learners, and 2) to provide a model of the pedagogical methodology design space for future development and adaptations. In the send part of the paper a contribution to the formalization of the model using set theory approach is presented as well.