Visual attention mechanisms play a crucial role in human perception and aesthetic evaluation. Recent advances in Vision Transformers (ViTs) have demonstrated remarkable capabilities in computer vision tasks, yet their alignment with human visual attention patterns remains underexplored, particularly in aesthetic contexts. This study investigates the correlation between human visual attention and ViT attention mechanisms when evaluating handcrafted objects. We conducted an eye-tracking experiment with 30 participants (9 female, 21 male, mean age 24.6 years) who viewed 20 artisanal objects comprising basketry bags and ginger jars. Using a Pupil Labs eye-tracker, we recorded gaze patterns and generated heatmaps representing human visual attention. Simultaneously, we analyzed the same objects using a pre-trained ViT model with DINO (Self-DIstillation with NO Labels), extracting attention maps from each of the 12 attention heads. We compared human and ViT attention distributions using four complementary metrics-Kullback-Leibler divergence, Structural Similarity Index (SSIM), Pearson's Correlation Coefficient (CC), and Similarity (SIM)-across varying Gaussian parameters ([Formula: see text]), yielding 1,152,000 distance evaluations. Additionally, we performed Areas of Interest (AOI) analysis to quantify ViT attention concentration within object regions. Statistical analysis revealed optimal correlation at [Formula: see text], with attention head #12 showing the strongest alignment with human visual patterns across all metrics. Significant differences were found between attention heads, with heads #7 and #9 demonstrating the greatest divergence from human attention ([Formula: see text]), Tukey HSD test). AOI analysis confirmed that all ViT heads concentrated attention significantly more within object regions than background areas ([Formula: see text]), with heads #12, #1, and #3 achieving lift values of +30 to +40 percentage points. Results indicate that while ViTs exhibit more global attention patterns compared to human focal attention, certain attention heads can approximate human visual behavior, particularly for specific object features like buckles in basketry items. These findings suggest potential applications of ViT attention mechanisms in product design and aesthetic evaluation, while highlighting fundamental differences in attention strategies between human perception and current AI models.
Although neuroscience has made considerable progress in recent decades by proposing robust models that explain the mechanisms of attention and perception in humans, emulating this capability using computational techniques remains complex. It was not until the development of models such as Visual Transformers (ViT) that it became possible to partially replicate this uniquely human trait. The main objective of this study was to explore the extent to which attention models, such as ViT, can reproduce the manner in which people distribute their visual attention when exposed to various stimuli, particularly in the context of handcrafted objects. Human fixations (i.e., attention) were recorded using an eye tracker, while the ViT model processed the same images to generate attention maps to evaluate the degree of similarity between the two patterns. For this purpose, heatmaps were constructed, and quantitative metrics were applied to assess their similarity. The results revealed areas of convergence and significant differences, highlighting the current limitations of computational models in capturing the more subtle aspects of human perception. This comparison not only helps us better understand the capabilities of ViT but also provides a foundation for reflecting on future improvements in automated attention models and their potential applications in contexts where visual interpretation is crucial.
Femoroacetabular impingement syndrome (FAIS) is a condition that implies increased intra-articular forces due to abnormal morphology, leading to pain, reduction in the range of motion, and even early development of hip osteoarthritis [1]. Diagnosing FAIS is challenging, and there is a lack of tools that implement the state-of-the-art techniques based on the process of 3D modeling, considering that these are time-consuming, require a large amount of data, and user input [2]-[4].A 3D Slicer extension was developed, implementing a Machine learning pipeline that allows for generating a 3D reconstruction of the proximal femur in 16.2 +/- 1.9 seconds. This reconstruction has an error close to other state-of-the-art techniques with a 95% Hausdorff distance of 4.8 +/- 2.1 [mm]. This pipeline starts with an automatic Deep learning based segmentation using a 3D Unet architecture that was trained with a set of 3T Flash DIXON MRI images (27 images with 208 slices, 256x256 px, 0.7 mm thickness) of patients diagnosed with FAIS, which performed with 0.91 and 0.82 for the DICE index. This segmentation was later improved using morphological operations and the DBSCAN algorithm. This work allows us to obtain a 3D model that is faithful to the patient's proximal femur, reducing segmentation time without losing the accuracy of manual segmentation. Finally, the extension it's available for downloading and use at this link: https://github.com/VenjaminRodriguezR/CAMalyzer
Las técnicas de balanceo de mecanismos han sido utilizadas a lo largo de los años para extender la vida útil de las máquinas reduciendo las vibraciones, el desgaste y la fatiga de sus componentes. Sin embargo, al tiempo que estas técnicas han crecido en efectividad lo han hecho también en complejidad matemática. Hoy en día la mayoría de los métodos de balanceo utilizan coordenadas cartesianas, las cuales generan ecuaciones complejas con funciones trigonométricas difíciles de simplificar. En este artículo se presenta el uso de coordenadas naturales para la obtención de los parámetros de balanceo de un mecanismo manivela-biela-corredera simplificado, evitando de esta forma el uso de funciones trigonométricas. La optimización del balanceo del mecanismo se lleva a cabo utilizando un algoritmo de optimización estocástico basado en poblaciones, permitiendo así la reducción del Momento de Sacudimiento (ShM) en un 97,76% y la reducción de la Fuerza de Sacudimiento (ShF) en un 94,58%.
This research is framed within the study of automatic recognition of emotions in artworks, proposing a methodology to improve performance in detecting emotions when a network is trained with an image type different from the entry type, which is known as the cross-depiction problem. To achieve this, we used the QuickShift algorithm, which simplifies images’ resources, and applied it to the Open Affective Standardized Image (OASIS) dataset as well as the WikiArt Emotion dataset. Both datasets are also unified under a binary emotional system. Subsequently, a model was trained based on a convolutional neural network using OASIS as a learning base, in order to then be applied on the WikiArt Emotion dataset. The results show an improvement in the general prediction performance when applying QuickShift (73% overall). However, we can observe that artistic style influences the results, with minimalist art being incompatible with the methodology proposed.
Images are capable of conveying emotions, but emotional experience is highly subjective. Advances in artificial intelligence have enabled the generation of images based on emotional descriptions. However, the level of agreement between the generative images and human emotional responses has not yet been evaluated. In order to address this, 20 artistic landscapes were generated using StyleGAN2-ADA. Four variants evoking positive emotions (contentment and amusement) and negative emotions (fear and sadness) were created for each image, resulting in 80 pictures. An online questionnaire was designed using this material, in which 61 observers classified the generated images. Statistical analyses were performed on the collected data to determine the level of agreement among participants between the observers’ responses and the generated emotions by AI. A generally good level of agreement was found, with better results for negative emotions. However, the study confirms the subjectivity inherent in emotional evaluation.
Visual representation as a means of communication uses elements to build a narrative. We propose using computer analysis to perform a quantitative analysis of the elements used in the visual creations that have been produced in reference to the epidemic, using 927 images compiled from The Covid Art Museum's Instagram account. This process has been carried out with techniques based on deep learning to detect objects contained in each study image. The research reveals the elements that are repeated in images to create narratives and the relations of association that are established in the sample. The predominant discourses in the sample do not show concern for the effects of illness. On the contrary, the impact and effects of confinement, through the prominent presence of elements such as human figures, windows, and buildings, are the most expressed experiences in the creations analyzed.
Color is a complex communicative element. At the level of artistic creation, this component influences both formal aspects and symbolic weight, directly affecting the construction of the message, and its associated emotion. During the COVID-19 pandemic, people generated countless images transmitting the subjective experiences of this event, and the social network Instagram was used to share this visual material. Using the repository of images created in the Instagram account CAM (The COVID Art Museum), we propose a methodology to understand the use of color and its emotional relationship in this context. The proposed methodology consists of creating a model that learns to recognize emotions via a convolutional neural network using the ArtEmis database. This model will subsequently be applied to recognize emotions in the CAM dataset, also extracting color attributes and their harmonies. Once both processes are completed, we combine the results, generating an expanded discussion on the usage of color and emotion. The results indicate that warm colors and analog compositions prevail in the sample. The relationship between emotions and composition shows a trend in positive emotions, reinforced by the results of the emotional relationship analysis of color attributes (hue, saturation, and lighting).
Point matching in multiple images is an open problem in computer vision because of the numerous geometric transformations and photometric conditions that a pixel or point might exhibit in the set of images. Over the last two decades, different techniques have been proposed to address this problem. The most relevant are those that explore the analysis of invariant features. Nonetheless, their main limitation is that invariant analysis all alone cannot reduce false alarms. This paper introduces an efficient point-matching method for two and three views, based on the combined use of two techniques: (1) the correspondence analysis extracted from the similarity of invariant features and (2) the integration of multiple partial solutions obtained from 2D and 3D geometry. The main strength and novelty of this method is the determination of the point-to-point geometric correspondence through the intersection of multiple geometrical hypotheses weighted by the maximum likelihood estimation sample consensus (MLESAC) algorithm. The proposal not only extends the methods based on invariant descriptors but also generalizes the correspondence problem to a perspective projection model in multiple views. The developed method has been evaluated on three types of image sequences: outdoor, indoor, and industrial. Our developed strategy discards most of the wrong matches and achieves remarkable F-scores of 97%, 87%, and 97% for the outdoor, indoor, and industrial sequences, respectively.
Until a safe and effective vaccine to fight the SARS-CoV-2 virus is developed and available for the global population, preventive measures, such as wearable tracking and monitoring systems supported by Internet of Things (IoT) infrastructures, are valuable tools for containing the pandemic. In this review paper we analyze innovative wearable systems for limiting the virus spread, early detection of the first symptoms of the coronavirus disease COVID-19 infection, and remote monitoring of the health conditions of infected patients during the quarantine. The attention is focused on systems allowing quick user screening through ready-to-use hardware and software components. Such sensor-based systems monitor the principal vital signs, detect symptoms related to COVID-19 early, and alert patients and medical staff. Novel wearable devices for complying with social distancing rules and limiting interpersonal contagion (such as smart masks) are investigated and analyzed. In addition, an overview of implantable devices for monitoring the effects of COVID-19 on the cardiovascular system is presented. Then we report an overview of tracing strategies and technologies for containing the COVID-19 pandemic based on IoT technologies, wearable devices, and cloud computing. In detail, we demonstrate the potential of radio frequency based signal technology, including Bluetooth Low Energy (BLE), Wi-Fi, and radio frequency identification (RFID), often combined with Apps and cloud technology. Finally, critical analysis and comparisons of the different discussed solutions are presented, highlighting their potential and providing new insights for developing innovative tools for facing future pandemics.
This paper presents a novel wearable system devoted to assist the mobility of blind and visually impaired people in urban environments with the simple use of a smartphone and tactile feedback. The system exploits the positioning data provided by the smartphone’s GPS sensor to locate in real-time the user in the environment and to determine the directions to a destination. The resulting navigational directions are encoded as vibrations and conveyed to the user via an on-shoe tactile display. To validate the pertinence of the proposed system, two experiments were conducted with test users. The first one involved a group of 20 voluntary normally sighted subjects that were requested to recognize the navigational instructions displayed by the tactile-foot device. The results show high recognition rates for the task. The second experiment consisted of guiding two blind voluntary subjects along public urban spaces to target destinations. Results show that the task was successfully accomplished and suggest that the system enhances independent safe navigation of visually impaired and blind people. Moreover, results show the potentials of smartphones and tactile-foot devices in assistive technology. Keywords: assistive technology, GPS localization, mobility of blind people, tactile-foot stimulation, vibrotactile display, wearable system.
Magnetotactic bacteria (MTB) are endowed with an exquisite orientation mechanism allowing them to swim along the geomagnetic field lines [1-3]. This mechanism consists of a chain of bio-synthesized magnetic nano-crystals [1,4,5]. Although the physics behind the minimum size of this biological compass is well understood [6], it is yet unclear what sets its maximum size [6-8]. Here, by using a macroscopic experiment inspired in these microorganisms, we show that larger magnetic moments will drive the collective behavior of MTB into a phase where bacteria are unable to swim freely, in detriment of their evolutive fitness [3,9]. Simple macroscopic experiments, numerical simulations, and analytic estimates are used to explain the upper limit for the size of the chain of nano-crystals in MTB. Our macroscopic bio-inspired experiment, and physical model provide new opportunities to explore and understand the phases of magnetic active matter at all scales.
The detection of cracks is an important monitoring task in civil engineering infrastructure devoted to ensuring durability, structural safety, and integrity. It has been traditionally performed by visual inspection, and the measurement of crack width has been manually obtained with a crack-width comparator gauge (CWCG). Unfortunately, this technique is time-consuming, suffers from subjective judgement, and is error-prone due to the difficulty of ensuring a correct spatial measurement as the CWCG may not be correctly positioned in accordance with the crack orientation. Although algorithms for automatic crack detection have been developed, most of them have specifically focused on solving the segmentation problem through Deep Learning techniques failing to address the underlying problem: crack width evaluation, which is critical for the assessment of civil structures. This paper proposes a novel automated method for surface cracking width measurement based on digital image processing techniques. Our proposal consists of three stages: anisotropic smoothing, segmentation, and stabilized central points by k-means adjustment and allows the characterization of both crack width and curvature-related orientation. The method is validated by assessing the surface cracking of fiber-reinforced earthen construction materials. The preliminary results show that the proposal is robust, efficient, and highly accurate at estimating crack width in digital images. The method effectively discards false cracks and detects real ones as small as 0.15 mm width regardless of the lighting conditions.
This paper reports on the progress of a wearable assistive technology (AT) device designed to enhance the independent, safe, and efficient mobility of blind and visually impaired pedestrians in outdoor environments. Such device exploits the smartphone's positioning and computing capabilities to locate and guide users along urban settings. The necessary navigation instructions to reach a destination are encoded as vibrating patterns which are conveyed to the user via a foot-placed tactile interface. To determine the performance of the proposed AT device, two user experiments were conducted. The first one requested a group of 20 voluntary normally sighted subjects to recognize the feedback provided by the tactile-foot interface. The results showed recognition rates over 93%. The second experiment involved two blind voluntary subjects which were assisted to find target destinations along public urban pathways. Results show that the subjects successfully accomplished the task and suggest that blind and visually impaired pedestrians might find the AT device and its concept approach useful, friendly, fast to master, and easy to use.
In classical mechanics, solutions can be classified according to their stability. Each of them is part of the possible trajectories of the system. However, the signatures of unstable solutions are hard to observe in an experiment, and most of the times if the experimental realization is adiabatic, they are considered just a nuisance. Here we use a small number of XY magnetic dipoles subject to an external magnetic field for studying the origin of their collective magnetic response. Using bifurcation theory we have found all the possible solutions being stable or unstable, and explored how those solutions are naturally connected by points where the symmetries of the system are lost or restored. Unstable solutions that reveal the symmetries of the system are found to be the culprit that shape hysteresis loops in this system. The complexity of the solutions for the nonlinear dynamics is analyzed using the concept of boundary basin entropy, finding that the damping time scale is critical for the emergence of fractal structures in the basins of attraction. Furthermore, we numerically found domain wall solutions that are the smallest possible realizations of transverse walls and vortex walls in magnetism. We experimentally confirmed their existence and stability showing that our system is a suitable platform to study domain wall dynamics at the macroscale.
Horizontal displacements of a multiple-anchor pile wall in a 28.5 m deep excavation using the top–down construction method have been monitored using optical fiber (Brillouin optical time-domain reflectometry (BOTDR)), strain gauges, inclinometers, and a topographic survey. This work presents a comparison between these different techniques to measure horizontal displacements in the pile at several stages of the soil excavation process. It was observed that displacements can be separated into two components: Rigid body motion and pile flexural deformation. Measurements using optical fiber and inclinometers are considered the most adequate and easy to install. A numerical model allows us to evaluate the influence of earth pressure on the estimated horizontal displacements. It is shown that using soil pressure on the wall given by p = 0.65Kaγh, on a simplified modeled wall, provides a close deduction of horizontal displacements compared to observed values on the field.
As an alternative to conventional batteries and other energy scavenging techniques, this paper introduces the idea of using micro-turbines to extract energy from wind forces at the microscale level and to supply power to battery-less microsystems. Fundamental research efforts on the design, fabrication, and test of micro-turbines with blade lengths of just 160 µm are presented in this paper along with analytical models and preliminary experimental results. The proof-of-concept prototypes presented herein were fabricated using a standard polysilicon surface micro-machining silicon technology (PolyMUMPs) and could effectively transform the kinetic energy of the available wind into a torque that might drive an electric generator or directly power supply a micro-mechanical system. Since conventional batteries do not scale-down well to the microscale, wind micro-turbines have the potential for becoming a practical alternative power source for microsystems, as well as for extending the operating range of devices running on batteries.
Medical knowledge is accumulated in scientific research papers along time. In order to exploit this knowledge by automated systems, there is a growing interest in developing text mining methodologies to extract, structure, and analyze in the shortest time possible the knowledge encoded in the large volume of medical literature. In this paper, we use the Latent Dirichlet Allocation approach to analyze the correlation between funding efforts and actually published research results in order to provide the policy makers with a systematic and rigorous tool to assess the efficiency of funding programs in the medical area. We have tested our methodology in the Revista Médica de Chile, years 2012-2015. 50 relevant semantic topics were identified within 643 medical scientific research papers. Relationships between the identified semantic topics were uncovered using visualization methods. We have also been able to analyze the funding patterns of scientific research underlying these publications. We found that only 29% of the publications declare funding sources, and we identified five topic clusters that concentrate 86% of the declared funds. Our methodology allows analyzing and interpreting the current state of medical research at a national level. The funding source analysis may be useful at the policy making level in order to assess the impact of actual funding policies, and to design new policies.
The CO2 and water vapor exchange between leaf and atmosphere are relevant for plant physiology. This process is done through the stomata. These structures are fundamental in the study of plants since their properties are linked to the evolutionary process of the plant, as well as its environmental and phytohormonal conditions. Stomatal detection is a complex task due to the noise and morphology of the microscopic images. Although in recent years segmentation algorithms have been developed that automate this process, they all use techniques that explore chromatic characteristics. This research explores a unique feature in plants, which corresponds to the stomatal spatial distribution within the leaf structure. Unlike segmentation techniques based on deep learning tools, we emphasize the search for an optimal threshold level, so that a high percentage of stomata can be detected, independent of the size and shape of the stomata. This last feature has not been reported in the literature, except for those results of geometric structure formation in the salt formation and other biological formations.
Commercial polypropylene fibers are incorporated as reinforcement of cement-based materials to improve their mechanical and damage performances related to properties such as tensile and flexural strength, toughness, spalling and impact resistance, delay formation of cracks and reducing crack widths. Yet, the production of these polypropylene fibers generates economic costs and environmental impacts and, therefore, the use of alternative and more sustainable fibers has become more popular in the research materials community. This paper addresses the characterization of recycled polypropylene fibers (RPFs) obtained from discarded domestic plastic sweeps, whose morphological, physical and mechanical properties are provided in order to assess their implementation as fiber-reinforcement in cement-based mortars. An experimental program addressing the incorporation of RPFs on the mechanical-damage performance of mortars, including a sensitivity analysis on the volumes and lengths of fiber, is developed. Using analysis of variance, this paper shows that RPFs statistically enhance flexural toughness and impact strength for high dosages and long fiber lengths. On the contrary, the latter properties are not statistically modified by the incorporation of low dosages and short lengths of RPFs, but still in these cases the incorporation of RPFs in mortars have the positive environmental impact of waste encapsulation. In the case of average compressive and flexural strength of mortars, these properties are not statistically modified when adding RPFs.