Coral reefs are declining worldwide due to climate change and local stressors. To inform effective conservation or restoration, monitoring at the highest possible spatial and temporal resolution is necessary. Conventional coral reef surveying methods are limited in scalability due to their reliance on expert labor time, motivating the use of computer vision tools to automate the identification and abundance estimation of live corals from images. However, the design and evaluation of such tools has been impeded by the lack of large high quality datasets. We release the Coralscapes dataset, the first general-purpose dense semantic segmentation dataset for coral reefs, covering 2075 images, 39 benthic classes, and 174k segmentation masks annotated by experts. Coralscapes has a similar scope and the same structure as the widely used Cityscapes dataset for urban scene segmentation, allowing benchmarking of semantic segmentation models in a new challenging domain which requires expert knowledge to annotate. We benchmark a wide range of semantic segmentation models, and find that transfer learning from Coralscapes to existing smaller datasets consistently leads to state-of-the-art performance. Coralscapes will catalyze research on efficient, scalable, and standardized coral reef surveying methods based on computer vision, and holds the potential to streamline the development of underwater ecological robotics.
In light of the critical threat to coral reefs worldwide due to human activity, innovative monitoring strategies are needed that are efficient, standardized, scalable, and economical. This paper presents the results of the first large-scale transnational coral reef surveying endeavor in the Red Sea using DeepReefMap, which provides automatic analysis of video transects by employing neural networks for 3D semantic mapping. DeepReefMap is trained using imagery from low-cost underwater cameras, allowing surveys to be conducted and analyzed in just a few minutes. This initiative was carried out in Djibouti, Jordan, and Israel, with over 184 hours of collected video footage for training the neural network for 3D reconstruction. We created a semantic segmentation dataset of video frames with over 200,000 annotated polygons from 39 benthic classes, down to the resolution of prominent visually identifiable genera found in the Red Sea. We analyzed 365 video transects from 45 sites using the deep-learning based mapping system, demonstrating the method’s robustness across environmental conditions and input video quality. We show that the surveys are consistent in characterizing the benthic composition, therefore showcasing the potential of DeepReefMap for monitoring. This research pioneers deep learning for practical 3D underwater mapping and semantic segmentation, paving the way for affordable, widespread deployment in reef conservation and ecology with tangible impact.
Coral reefs are crucial for biodiversity and provide vital resources for humankind. But despite such a central role, they are confronted to increasing threats linked to climate change, pollution, and local stressors. To ensure effective conservation, efficient and scalable monitoring is key: this necessitates automated identification of benthic classes and their states on a large scale through semantic segmentation. However, segmentation of underwater videos is challenging, because of visual similarities between benthic classes, underwater distortions and limited available datasets, making it harder to create accurate and robust models. In this paper, we present a method for training a semantic segmentation model on a small dataset of video frames of coral scenes, by fine-tuning a large transformer model. Our approach uses transfer learning on the Segment Anything Model (SAM), incorporating specific training and prediction strategies. We benchmark our model against a CNN for semantic segmentation as a baseline. Our results demonstrate a substantial improvement in model performance, particularly for benthic classes that often appear as small objects and rarer classes, highlighting the potential of our approach in advancing coral reef mapping and monitoring.
Deployable systems for CubeSat applications have significantly matured in recent years as ever more ambitious missions drive innovations in design. Collecting high frequency signals (5-30 MHz) is particularly challenging due to the corresponding increase in size needed for a receiving antenna to operate in that frequency range. This paper describes the mechanical design, analysis, test, build, and deployment of a 6 meter x 6 meter crossed dipole CubeSat antenna. The antenna system was paired with a galactic noise limited receiver on a flight mission to better characterize high frequency signals from both Earth and extragalactic sources, and serve as a precursor to future radio missions throughout the solar system. The antenna is composed of four 3 meter long tape spring elements that stow within a 100 mm x 100 mm x 22 mm volume and weigh less than 450 grams. All four elements are simultaneously released by the energization of a single hot knife mechanism and use strain energy to fully deploy. This is a significant departure from most large aperture satellite antennas which typically require a motor to control deployment. The antenna was designed, tested, and built on a compressed schedule of less than a year and with a budget of less than half of what would be considered typical at NASA's Jet Propulsion Laboratory (JPL). The system was successfully deployed on orbit, but science return was limited due to a persistent attitude perturbation likely caused by a solar-flutter phenomenon acting on the antenna elements.
This paper provides a closed-form formulation to estimate the lowest natural frequency of a panel supported by tape spring hinges, a strain energy-based hinge that is commonly used in deployable space structures. The formulation is derived for the bending-shearing deformation modes, idealizing the hinge as a Timoshenko beam. The formulation provides a complete description of each geometric parameter's role for the natural frequency. We further extend the analysis by defining beta(x) and beta(y) non dimensional length scales, that describe the transition between deformation modes. The analytic formulation has been validated using a series of finite element analyses, where excellent agreement is observed for longer hinge lengths. As the hinge length becomes shorter and the separation between the tape springs (.) increases, the accuracy of the analytic prediction reduces. Current work includes an experimental campaign, where preliminary results have shown the importance of properly mounting the tape spring with the panel. Several mounting brackets have been fabricated in-house that can preserve the geometric properties of the tape spring at the connecting point.
Coral reefs are among the most diverse ecosystems on our planet, and are depended on by hundreds of millions of people. Unfortunately, most coral reefs are existentially threatened by global climate change and local anthropogenic pressures. To better understand the dynamics underlying deterioration of reefs, monitoring at high spatial and temporal resolution is key. However, conventional monitoring methods for quantifying coral cover and species abundance are limited in scale due to the extensive manual labor required. Although computer vision tools have been employed to aid in this process, in particular SfM photogrammetry for 3D mapping and deep neural networks for image segmentation, analysis of the data products creates a bottleneck, effectively limiting their scalability. This paper presents a new paradigm for mapping underwater environments from ego-motion video, unifying 3D mapping systems that use machine learning to adapt to challenging conditions under water, combined with a modern approach for semantic segmentation of images. The method is exemplified on coral reefs in the northern Gulf of Aqaba, Red Sea, demonstrating high-precision 3D semantic mapping at unprecedented scale with significantly reduced required labor costs: a 100 m video transect acquired within 5 minutes of diving with a cheap consumer-grade camera can be fully automatically analyzed within 5 minutes. Our approach significantly scales up coral reef monitoring by taking a leap towards fully automatic analysis of video transects. The method democratizes coral reef transects by reducing the labor, equipment, logistics, and computing cost. This can help to inform conservation policies more efficiently. The underlying computational method of learning-based Structure-from-Motion has broad implications for fast low-cost mapping of underwater environments other than coral reefs.
Underwater scenes are challenging for computer vision methods due to color degradation caused by the water column and detrimental lighting effects such as caustic caused by sunlight refracting on a wavy surface. These challenges impede widespread use of computer vision tools that could aid in ecological surveying of underwater environments or in industrial applications. Existing algorithms for alleviating caustics and descattering the image to recover colors are often impractical to implement due to the need for ground-truth training data, the necessity for successful alignment of an image within a 3D scene, or other assumptions that are infeasible in practice. In this paper, we propose a solution to tackle those problems in underwater computer vision: our method is based on two neural networks: CausticsNet, for single-image caustics removal, and BackscatterNet, for backscatter removal. Both neural networks are trained using an objective formulated with the aid of self-supervised monocular SLAM on a collection of underwater videos. Thus, our method does not requires any ground-truth color images or caustics labels, and corrects images in real-time. We experimentally demonstrate the fidelity of our caustics removal method, performing similarly to state-of-the-art supervised methods, and show that the color restoration and caustics removal lead to better downstream performance in Structure-from-Motion image keypoint matching than a wide range of methods.
Countless signal processing applications include the reconstruction of signals from few indirect linear measurements. The design of effective measurement operators is typically constrained by the underlying hardware and physics, posing a challenging and often even discrete optimization task. While the potential of gradient-based learning via the unrolling of iterative recovery algorithms has been demonstrated, it has remained unclear how to leverage this technique when the set of admissible measurement operators is structured and discrete. We tackle this problem by combining unrolled optimization with Gumbel reparametrizations, which enable the computation of low-variance gradient estimates of categorical random variables. Our approach is formalized by GLODISMO (Gradient-based Learning of DIscrete Structured Measurement Operators). This novel method is easy-to-implement, computationally efficient, and extendable due to its compatibility with automatic differentiation. We empirically demonstrate the performance and flexibility of GLODISMO in several prototypical signal recovery applications, verifying that the learned measurement matrices outperform conventional designs based on randomization as well as discrete optimization baselines.
Countless signal processing applications include the reconstruction of an unknown signal from very few indirect linear measurements. Because the measurement operator is commonly constrained by the hardware or the physics of the observation process, finding measurement matrices that enable accurate signal recovery poses a challenging discrete optimization task. Meanwhile, recent advances in the field of machine learning have highlighted the effectiveness of gradient-based optimization methods applied to large computational graphs such as those arising naturally when unrolling iterative algorithms for signal recovery. However, it has remained unclear how to leverage this technique when the set of admissible measurement matrices is both discrete and sparse. In this paper, we tackle this problem and propose an efficient and flexible method for learning structured sparse measurement matrices. Our approach uses unrolled optimization in conjunction with Gumbel reparametrizations. We empirically demonstrate the effectiveness of our method in two prototypical compressed sensing situations.
Submitted by Tibor Kremic and Mike Amato, co-chairs, Venus Surface Platform Study Team Leads Martha Gilmore, Walter Kiefer, Natasha Johnson, Jonathan Sauder, Gary Hunter, and Thomas Thompson
It is well-established that many iterative sparse reconstruction algorithms can be unrolled to yield a learnable neural network for improved empirical performance. A prime example is learned ISTA (LISTA) where weights, step sizes and thresholds are learned from training data. Recently, Analytic LISTA (ALISTA) has been introduced, combining the strong empirical performance of a fully learned approach like LISTA, while retaining theoretical guarantees of classical compressed sensing algorithms and significantly reducing the number of parameters to learn. However, these parameters are trained to work in expectation, often leading to suboptimal reconstruction of individual targets. In this work we therefore introduce Neurally Augmented ALISTA, in which an LSTM network is used to compute step sizes and thresholds individually for each target vector during reconstruction. This adaptive approach is theoretically motivated by revisiting the recovery guarantees of ALISTA. We show that our approach further improves empirical performance in sparse reconstruction, in particular outperforming existing algorithms by an increasing margin as the compression ratio becomes more challenging.
Language models trained with Maximum Likelihood Estimation (MLE) have been considered as a mainstream solution in Natural Language Generation (NLG) for years. Recently, various approaches with Generative Adversarial Nets (GANs) have also been proposed. While offering exciting new prospects, GANs in NLG by far are nevertheless reportedly suffering from training instability and mode collapse, and therefore outperformed by conventional MLE models. In this work, we propose techniques for improving GANs in NLG, namely Best Student Forcing (BSF), a novel yet simple adversarial training mechanism in which generated sequences of high quality are selected as temporary ground-truth to further train the generator. We also use an ensemble of discriminators to increase training stability and sample diversity. Evaluation shows that the combination of BSF and multiple discriminators consistently performs better than previous GAN approaches over various metrics, and outperforms a baseline MLE in terms of Frechet Distance, a recently proposed metric capturing both sample quality and diversity.
Point clouds provide a flexible and natural representation usable in countless applications such as robotics or self-driving cars. Recently, deep neural networks operating on raw point cloud data have shown promising results on supervised learning tasks such as object classification and semantic segmentation. While massive point cloud datasets can be captured using modern scanning technology, manually labelling such large 3D point clouds for supervised learning tasks is a cumbersome process. This necessitates effective unsupervised learning methods that can produce representations such that downstream tasks require significantly fewer annotated samples. We propose a novel method for unsupervised learning on raw point cloud data in which a neural network is trained to predict the spatial relationship between two point cloud segments. While solving this task, representations that capture semantic properties of the point cloud are learned. Our method outperforms previous unsupervised learning approaches in downstream object classification and segmentation tasks and performs on par with fully supervised methods.
Point clouds provide a flexible and natural representation usable in countless applications such as robotics or self-driving cars. Recently, deep neural networks operating on raw point cloud data have shown promising results on supervised learning tasks such as object classification and semantic segmentation. While massive point cloud datasets can be captured using modern scanning technology, manually labelling such large 3D point clouds for supervised learning tasks is a cumbersome process. This necessitates methods that can learn from unlabelled data to significantly reduce the number of annotated samples needed in supervised learning. We propose a self-supervised learning task for deep learning on raw point cloud data in which a neural network is trained to reconstruct point clouds whose parts have been randomly rearranged. While solving this task, representations that capture semantic properties of the point cloud are learned. Our method is agnostic of network architecture and outperforms current unsupervised learning approaches in downstream object classification tasks. We show experimentally, that pre-training with our method before supervised training improves the performance of state-of-the-art models and significantly improves sample efficiency.