Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription remains insufficiently understood. In this work, we benchmark traditional OCR systems and VLM-based approaches on the Berrutti dataset, a challenging collection of Uruguayan dictatorship-era documents derived from microfilm scans. While VLMs consistently outperform traditional methods in terms of Character Error Rate (CER) and Word Error Rate (WER), we show that these improvements hide a more complex picture. Through a detailed qualitative analysis, we uncover systematic failure modes that are invisible to standard metrics, including orthographic normalization, spurious content generation, and semantic substitutions that preserve fluency while altering meaning. Errors affecting named entities are particularly critical, as they can introduce substantial semantic distortions with minimal impact on CER and WER. These findings reveal a critical gap between quantitative OCR performance and transcription fidelity in real-world archival settings, and highlight the need for evaluation frameworks that go beyond character-level accuracy to capture the semantic reliability of generated transcriptions.
We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. This provides a tool that accelerates philological analysis. Using character prototypes derived from clustering and aligning automatically extracted character images, the method defines a typeface distance between any two books. To produce actionable outputs, we develop an a contrario statistical framework to interpret the significance of the computed typeface distances. We apply the method to the philological study of 17 th -century Spanish printed theatre chapbooks in a quantity that exceeds the capabilities of systematic visual inspection by human experts. Our method enables the automatic comparison of Roman and Italic types extracted from different books. After validation by human experts, our method has led to new printer attributions being discovered, and former printer attributions being revised. This success strongly suggests that our method has the potential to enable digital bibliography on a larger scale than was previously possible.
Tree ring marking remains a key step in dendrometry and dendrochronology, but it is often performed manually, making the process time-consuming, subjective, and difficult to scale to large image datasets. We present the Tree Ring Analyzer Suite (TRAS), an open-source graphical software for automatic delineation, manual correction, and measurement of tree rings in wood cross-sectional images. TRAS integrates three complementary detection algorithms: the classical image-processing method CS-TRD and two deep-learning approaches, DeepCS-TRD and INBD. The interface allows users to refine automatic detections, remove false positives, and manually add missing rings. It also computes dendrochronological metrics such as earlywood and latewood areas, ring perimeter, equivalent ring width, and custom path-based ring-width measurements. TRAS was evaluated on 18 expertly annotated Pinus taeda L. cross-section images. DeepCS-TRD achieved the best automatic detection performance, with an F-score of 81.0 TRAS provides a flexible and reproducible solution for tree-ring analysis on Windows, macOS, and Linux. Code is available at the https://hmarichal93.github.io/tras.
In this article, we analyze and propose a Python implementation of the method "Pith Estimation on Rough Log End images using Local Fourier Spectrum Analysis", by Rudolf Schraml and Andreas Uhl. The algorithm is tested over two datasets.
Abstract Tree ring marking is a key step in dendrometry, essential for studies in dendrochronology, forest dynamics, and growth analysis. However, it is still often done manually, which makes the process time-consuming, subjective, and hard to scale to large image datasets. Here, we present the Tree Ring Analyzer Suite (TRAS), an open-source software tool with a graphical interface that enables automatic delineation and manual correction of tree rings in wood cross-sectional images. The software incorporates three complementary algorithms for automatically detecting tree rings: a classic image processing method (CS-TRD) and two deep-learning approaches (DeepCS-TRD and INBD). TRAS provides an intuitive graphical user interface that allows users to refine automatic ring detection results, remove erroneous detections, and manually add undetected rings. The software also calculates various dendrochronological metrics, including earlywood and latewood areas, ring perimeter, and equivalent ring width, with options for custom measurements along specified paths. A dataset of 18 expertly annotated tree cross-section images of Pinus taeda L. was used for validation and accuracy assessment, with additional manual measurements serving as ground truth. For 1D measurements (ring width), a comparative analysis with a traditional tool (CooRecorder) was performed. TRAS achieved accurate automatic detection of tree rings P. taeda samples, with DeepCS-TRD obtaining the highest F-score (81.0%) and precision (86.4%). Automatic detection reduces the annotation effort to $\sim $20% of ring boundaries requiring manual correction, representing a substantial time saving over fully manual delineation. TRAS measurements showed excellent agreement with CooRecorder ($r> 0.99$), and common detection errors, such as jump propagation or false positives near knots, were easily corrected using the postprocessing interface. TRAS offers a robust and flexible solution for researchers in dendrometry, enhancing the efficiency and accuracy of tree ring analysis. The software runs on Windows, macOS, and Linux.
Shrub-ring analysis is increasingly used to assess climate-growth relationships in Arctic and alpine ecosystems. However, manual ring measurement remains labor-intensive and time-consuming, limiting the scale of ecological inference. To address this challenge, here we evaluated the Iterative Next Boundary Detection (INBD) deep learning method for automated ring detection using a new dataset of 50 manually annotated Salix glauca cross-section images from western Greenland. The model achieved intermediate performance, successfully detecting rings in morphologically clear samples but showing limited accuracy in more complex cases. We further evaluated three image resizing strategies and found that normalizing images to a fixed largest dimension of 1504 pixels improved segmentation accuracy and reduced training time compared to fixed downsampling approaches. We compared ring traces from automated and manual delineations, calculating basal area increment (BAI) from both approaches, along with six additional metrics derived from the manual ring traces. Growth patterns and ring counts from automated delineations were generally consistent with manually traced rings. Correlation analyses showed positive relationships between summer temperature and growth, with BAI (both automatic and manual) showing non-significant trends. In contrast, most one-dimensional metrics exhibited significant positive correlations, highlighting the potential influence of measurement approach on inferred climate sensitivity. Linear mixed-effects models further revealed consistent, significant positive relationships between shrub growth and mean summer temperature across all metrics, with the model based on automatically derived BAI explaining the largest proportion of variance. Our findings highlight both the potential and current limitations of the INBD method for automated shrub ring analysis in Salix glauca. Despite existing accuracy issues, the method can currently produce ecologically meaningful ring delineations and growth patterns. With species-specific training, refinement, and further testing, automated ring detection can accelerate data extraction from shrub rings and expand dendrochronological research in cold-climate regions.
Shrub-ring analysis is increasingly used to assess climate-growth relationships in Arctic and alpine ecosystems. However, manual ring measurement remains labor-intensive and time-consuming, limiting the scale of ecological inference. To address this challenge, here we evaluated the Iterative Next Boundary Detection (INBD) deep learning method for automated ring detection using a new dataset of 50 manually annotated Salix glauca cross-section images from western Greenland. The model achieved intermediate performance, successfully detecting rings in morphologically clear samples but showing limited accuracy in more complex cases. We further evaluated three image resizing strategies and found that normalizing images to a fixed largest dimension of 1504 pixels improved segmentation accuracy and reduced training time compared to fixed downsampling approaches. We compared ring traces from automated and manual delineations, calculating basal area increment (BAI) from both approaches, along with six additional metrics derived from the manual ring traces. Growth patterns and ring counts from automated delineations were generally consistent with manually traced rings. Correlation analyses showed positive relationships between summer temperature and growth, with BAI (both automatic and manual) showing non-significant trends. In contrast, most one-dimensional metrics exhibited significant positive correlations, highlighting the potential influence of measurement approach on inferred climate sensitivity. Linear mixed-effects models further revealed consistent, significant positive relationships between shrub growth and mean summer temperature across all metrics, with the model based on automatically derived BAI explaining the largest proportion of variance. Our findings highlight both the potential and current limitations of the INBD method for automated shrub ring analysis in Salix glauca. Despite existing accuracy issues, the method can currently produce ecologically meaningful ring delineations and growth patterns. With species-specific training, refinement, and further testing, automated ring detection can accelerate data extraction from shrub rings and expand dendrochronological research in cold-climate regions.
Here, we propose Deep CS-TRD, a new automatic algorithm for detecting tree rings in whole cross-sections. It substitutes the edge detection step of CS-TRD by a deep-learning-based approach (U-Net), which allows the application of the method to different image domains: microscopy, scanner or smartphone acquired, and species (Pinus taeda, Gleditsia triachantos and Salix glauca). Additionally, we introduce two publicly available datasets of annotated images to the community. The proposed method outperforms state-of-the-art approaches in macro images (Pinus taeda and Gleditsia triacanthos) while showing slightly lower performance in microscopy images of Salix glauca. To our knowledge, this is the first paper that studies automatic tree ring detection for such different species and acquisition conditions. The dataset and source code are available in https://hmarichal93.github.io/deepcstrd/ .
This work describes a Tree Ring Detection method for complete Cross-Sections of trees (CS-TRD). The method is based on the detection, processing, and connection of edges corresponding to the tree's growth rings. The method depends on the parameters for the Canny Devernay edge detector ($\sigma$ and two thresholds), a resize factor, the number of rays, and the pith location. The first five parameters are fixed by default. The pith location can be marked manually or using an automatic pith detection algorithm. Besides the pith localization, the CS-TRD method is fully automated and achieves an F-Score of 89\% in the UruDendro dataset (of Pinus Taeda) with a mean execution time of 17 seconds and of 97\% in the Kennel dataset (of Abies Alba) with an average execution time 11 seconds.
The automatic detection of tree-ring boundaries and other anatomical features using image analysis has progressed substantially over the past decade with advances in machine learning and imagery technology, as well as increasing demands from the dendrochronology community. This paper presents a publicly available dataset of 64 annotated images of transverse sections of commercially grown Pinus taeda L. trees from northern Uruguay, presenting 17 to 24 annual rings. The collection contains several challenging features for automatic ring detection, including illumination and surface preparation variation, fungal infection (blue stains), knot formation, missing bark or interruptions in outer rings, and radial cracking. This dataset can be used to develop and test automatic tree ring detection algorithms. The dataset presented here was used to develop the Cross-Section Tree-Ring Detection (CS-TRD) method, an open-source automated ring-detection algorithm for cross-sectioned images. Dataset access at https://doi.org/10.5281/zenodo.15110647 . Access to the metadata describing the data set: https://metadata-afs.nancy.inra.fr/geonetwork/srv/fre/catalog.search#/metadata/5fdbd411-9ae1-4ce6-8ef0-cdfa2fbd7a6a .
Automatic sign language translation has gained particular interest in the computer vision and computational linguistics communities in recent years. Given each sign language country’s particularities, machine translation requires local data to develop new techniques and adapt existing ones. This work presents iLSU-T, an open dataset of interpreted Uruguayan Sign Language RGB videos with audio and text transcriptions. This type of multimodal and curated data is paramount for developing novel approaches to understand or generate tools for sign language processing. iLSU-T comprises more than 185 hours of interpreted sign language videos from public TV broadcasting. It covers diverse topics and includes the participation of 18 professional interpreters of sign language. A series of experiments using three state-of-the-art translation algorithms is presented. The aim is to establish a baseline for this dataset and evaluate its usefulness and the proposed pipeline for data processing. The experiments highlight the need for more localized datasets for sign language translation and understanding, which are critical for developing novel tools to improve accessibility and inclusion of all individuals. Our data and code can be accessed at https://github.com/ariel-e-stassi/iLSU-T.
Current OCR systems are based on deep learning models trained on large amounts of data. Although they have shown some ability to generalize to unseen data, especially in detection tasks, they can struggle with recognizing low-quality data. This is particularly evident for printed documents, where intra-domain data variability is typically low, but inter-domain data variability is high. In that context, current OCR methods do not fully exploit each document's redundancy. We propose an unsupervised method by leveraging the redundancy of character shapes within a document to correct imperfect outputs of a given OCR system and suggest better clustering. To this aim, we introduce an extended Gaussian Mixture Model (GMM) by alternating an Expectation-Maximization (EM) algorithm with an intra-cluster realignment process and normality statistical testing. We demonstrate improvements in documents with various levels of degradation, including recovered Uruguayan military archives and 17th to mid-20th century European newspapers.
Tree-ring growth represents the annual wood increment for a tree, and quantifying it allows researchers to assess which silvicultural practices are best suited for each species. Manual measurement of this growth is time-consuming and often imprecise, as it is typically performed along 4 to 8 radial directions on a cross-sectional disc. In recent years, automated algorithms and datasets have emerged to enhance accuracy and automate the delineation of annual rings in cross-sectional images. To address the scarcity of wood cross-section data, we introduce the UruDendro4 dataset, a collection of 102 image samples of Pinus taeda L., each manually annotated with annual growth rings. Unlike existing public datasets, UruDendro4 includes samples extracted at multiple heights along the stem, allowing for the volumetric modeling of annual growth using manually delineated rings. This dataset (images and annotations) allows the development of volumetric models for annual wood estimation based on cross-sectional imagery. Additionally, we provide a performance baseline for automatic ring detection on this dataset using state-of-the-art methods. The highest performance was achieved by the DeepCS-TRD method, with a mean Average Precision of 0.838, a mean Average Recall of 0.782, and an Adapted Rand Error score of 0.084. A series of ablation experiments were conducted to empirically validate the final parameter configuration. Furthermore, we empirically demonstrate that training a learning model including this dataset improves the model's generalization in the tree-ring detection task.
The automatic detection of tree-ring boundaries and other anatomical features using image analysis has progressed substantially over the past decade with advances in machine learning and imagery technology, as well as increasing demands from the dendrochronology community. This paper presents a publicly available database of 64 scanned images of transverse sections of commercially grown Pinus taeda trees from northern Uruguay, ranging from 17 to 24 years old. The collection contains several challenging features for automatic ring detection, including illumination and surface preparation variation, fungal infection (blue stains), knot formation, missing cortex or interruptions in outer rings, and radial cracking. This dataset can be used to develop and test automatic tree ring detection algorithms. This paper presents to the dendrochronology community one such method, Cross-Section Tree-Ring Detection (CS-TRD), which identifies and marks complete annual rings in cross-sections for tree species presenting a clear definition between early and latewood. We compare the CS-TRD performance against the ground truth manual delineation of all rings over the UruDendro dataset. The CS-TRD software identified rings with an average F-score of 89% and RMSE error of 5.27px for the entire database in less than 20 seconds per image. Finally, we propose a robust measure of the ring growth using the \emph{equivalent radius} of a circle having the same area enclosed by the detected tree ring. Overall, this study contributes to the dendrochronologist's toolbox of fast and low-cost methods to automatically detect rings in conifer species, particularly for measuring diameter growth rates and stem transverse area using entire cross-sections.
This work analyzes the BigColor method, a fully automatic colorization approach that aims to meet the challenge of providing realistic and vivid colorization for complex and diverse images in real -world scenarios. The method is a BigGAN-inspired encoder -generator network, using a spatial feature map, enabling single forward -pass colorization, supporting arbitrary input resolutions, and producing multimodal colorization results. We provide a short analysis of the method's results and highlight some limitations alongside its achievements.
A fully automated technique for wood pith detection (APD), relying on the concentric shape of the structure of wood ring slices, is introduced. The method estimates the ring’s local orientations using the 2D structure tensor and finds the pith position, optimizing a cost function designed for this problem. We also present a variant (APD-PCL) using the parallel coordinate space that enhances the method’s effectiveness when there are no clear tree ring patterns. Furthermore, refining Kurdthongmee’s work, a YoloV8 net is trained for pith detection, producing a deep learning-based approach (APD-DL). All methods were tested on seven datasets, including images captured under diverse conditions (controlled laboratory settings, sawmill, and forest) and featuring various tree species (Pinus taeda, Douglas fir, Abies alba, and Gleditsia triacanthos). All proposed approaches outperform existing state-of-the-art methods and can be used in CPU-based real-time applications. Additionally, we provide a novel dataset comprising images of gymnosperm and angiosperm species. Dataset and source code are available at http://github.com/hmarichal93/apd .
This paper briefly describes and analyzes iColoriT, a hybrid colorization method based on a Vision Transformer that propagates user hints to relevant regions of a grayscale image while using color priors learned from a large image dataset. This approach gives users more control over color inference and shows a quick way to achieve results.
This work presents the INBD network proposed by Gillert et al. in CVPR-2023 and studies its application for delineating tree rings in RGB images of Pinus taeda cross sections captured by a smartphone (UruDendro dataset), which are images with different characteristics from the ones used to train the method. The INBD network operates in two stages: first, it segments the background, pith, and ring boundaries. In the second stage, the image is transformed into polar coordinates, and ring boundaries are iteratively segmented from the pith to the bark. Both stages are based on the U-Net architecture. The method achieves an F-Score of 77.5, a mAR of 0.540, and an ARAND of 0.205 on the evaluation set. The code for the experiments is available at https://github.com/hmarichal93/mlbrief_inbd.
An automated method to analyze personal record cards generated by Organismo Coordinador de Operaciones Antisubversivas (OCOA) during the civic-military dictatorship in Uruguay between 1973 and 1985 is presented. These personal records are part of Archivo Berrutti, a collection of digitized documents, partially processed by the Uruguayan government. The main goal of this study is to extract the maximum amount of information from the personal record cards to ease the analysis by specialized teams. To achieve the goal, a methodology which combines image processing and pattern recognition techniques has been developed. This methodology takes advantage of the known geometric structure of the cards to straighten them, identify and extract relevant pieces, classify them, extract relevant fields, and transcribe crucial information such as names and identification numbers.