Floods are one of the most common natural disasters, and the rate of flood related disasters has more than doubled since the year 2000. Accurate and timely warnings are critical for mitigating flood risks, especially for major events that can potentially impact thousands. For that, we develop a global early warning system for major flood events, using discharge predictions from the Global Hydrological Model developed by Google. We use agglomerative clustering to cluster discharge predictions at Hydro-ATLAS sub-basins that exceed certain return period thresholds. These extended space and time clusters indicate potential flood events. A supervised model is then employed to extract clusters associated with a higher risk of major flood events. We use the Dartmouth Flood Observatory (DFO) and GDACS datasets of historical flood events as ground truth for training the classification model. We use various data sources to extract cluster features, such as estimate of affected population, prior event density at the cluster’s location, total drainage area associated with the cluster etc. Our initial experiments show promising results, demonstrating the potential of a machine learning-based early warning system to accurately predict major flood events on a global scale, using global hydrological discharge predictions. While the results are encouraging, we believe that further refinement and validation of the model, coupled with efforts to improve the quality of ground truth data and further cluster features, could increase the accuracy of the system.
BACKGROUND CONTEXT: The annular epiphysis (AE) is a peripheral ring of cortical bone that forms a secondary ossification center in the superior and inferior surfaces of vertebral bodies (VBs). The AE is the last ossification site in the skeleton, typically forming at about the 25th year of life. The AE functions jointly with vertebral endplates to anchor the intervertebral discs to the VBs. PURPOSE: To establish accurate data on the sizes of the AE of the cervical spine (C3-C7); to compare the ratios between areas and the ratios of the AE to VBs; to compare the ratios between the superior and inferior VB surface areas; and to compare AE lengths between the posterior and anterior midsagittal areas. STUDY DESIGN: Measurement of 424 cervical spines (C3-C7) obtained from the skeletal col-lection of the Natural History Museum, Cleveland, Ohio (USA). METHODS: The sample was characterized by sex, age, and ethnic origin. The following measure-ments were recorded for each vertebra: (1) the surface area of the VBs and the AE, (2) the midsagit-tal anterior and posterior length of the AE, (3) the ratios between the AE and VB surface areas, and (4) the ratios between the superior and inferior disc surface areas. RESULTS: The study revealed that the AE and VBs in men were larger than in women. With age, the AE and VBs became larger; the ratio between the AE and VB surface was approximately 0.5 through-out the middle to lower cervical spine. The ratio of superior to inferior VBs was approximately 0.8. We found no differences between African Americans versus European Americans or between the ante-rior versus the posterior midsagittal length of the AE of the superior and inferior VBs. CONCLUSIONS: The ratios between the superior and inferior VBs are >0.8, and the ratio is the same for the entire middle to lower spine. Thus, the ratio between the superior and inferior VBs to the AE is > 0.5. Men had larger AEs and VBs than women did, with both VBs and AEs becoming larger with age. Knowing these relationships are important so that orthopedic surgeons can best correct these issues in young patients (<25 years old) during spine surgery. The data reported here provide, for the first time, all the relevant sizes of the AE and VB. In future studies, AEs and VBs of living patients can be measured with computed tomography. CLINICAL SIGNIFICANCE: The ER location and function are clinically significant showing any changes during life that might lead to clinical issues related to intervertebral discs such as intervertebral disc asymmetry, disc herniation, nerve pressure, cervical osteophytes and neck pain. & COPY; 2023 Elsevier Inc. All rights reserved.
Abstract The increasing intensity and frequency of floods is one of the many consequences of our changing climate. In this work, we explore ML techniques that improve the flood detection module of an operational early flood warning system. Our method exploits an unlabeled dataset of paired multi-spectral and synthetic aperture radar (SAR) imagery to reduce the labeling requirements of a purely supervised learning method. Prior works have used unlabeled data by creating weak labels out of them. However, from our experiments, we noticed that such a model still ends up learning the label mistakes in those weak labels. Motivated by knowledge distillation and semi-supervised learning, we explore the use of a teacher to train a student with the help of a small hand-labeled dataset and a large unlabeled dataset. Unlike the conventional self-distillation setup, we propose a cross-modal distillation framework that transfers supervision from a teacher trained on richer modality (multi-spectral images) to a student model trained on SAR imagery. The trained models are then tested on the Sen1Floods11 dataset. Our model outperforms the Sen1Floods11 baseline model trained on the weak-labeled SAR imagery by an absolute margin of $ 6.53\% $ intersection over union (IoU) on the test split.
Google's operational flood forecasting system was developed to provide accurate real-time flood warnings to agencies and the public with a focus on riverine floods in large, gauged rivers. It became operational in 2018 and has since expanded geographically. This forecasting system consists of four subsystems: data validation, stage forecasting, inundation modeling, and alert distribution. Machine learning is used for two of the subsystems. Stage forecasting is modeled with the long short-term memory (LSTM) networks and the linear models. Flood inundation is computed with the thresholding and the manifold models, where the former computes inundation extent and the latter computes both inundation extent and depth. The manifold model, presented here for the first time, provides a machine-learning alternative to hydraulic modeling of flood inundation. When evaluated on historical data, all models achieve sufficiently high-performance metrics for operational use. The LSTM showed higher skills than the linear model, while the thresholding and manifold models achieved similar performance metrics for modeling inundation extent. During the 2021 monsoon season, the flood warning system was operational in India and Bangladesh, covering flood-prone regions around rivers with a total area close to 470 000 km2, home to more than 350 000 000 people. More than 100 000 000 flood alerts were sent to affected populations, to relevant authorities, and to emergency organizations. Current and future work on the system includes extending coverage to additional flood-prone locations and improving modeling capabilities and accuracy.
Diagnosis of a specific learning disability such as dysgraphia impacts children's academic progress and well-being. Dysgraphia is diagnosed by clinicians based on children's written product and educational staff's impressions. This process is time consuming and subjective. Consequently, many children with mild dysgraphia remain undiagnosed, especially those from lower socioeconomic backgrounds. In this work, a method for automatic identification and characterization of dysgraphia in third-grade children is described. The method is based on analyzing the child's writing dynamics by sampling the pressure the pen exerts on the paper as well as the pen's position and orientation by using a standard digital writing pad. Ninety-nine samples were collected from writers with dysgraphia and proficient writers. A wide range of features covering dynamic properties of the writing and typographic (i.e., visual) properties were extracted for each participant. Machine learning methodologies were used to infer a statistical model, which is capable of discriminating dysgraphic products from proficient products with approximately 90% accuracy. The model was analyzed to conclude which handwriting features are most discriminative. Since the model provides 90% sensitivity for a specificity of 90%, it is the first step toward future use as an effective standard indicator for dysgraphia detection.
Let G = (V, E), vertical bar V vertical bar = n, be a simple connected graph. An edge-colored graph G is rainbow edge-connected if any two vertices are connected by a path whose edges are colored by distinct colors. The rainbow connection number of a connected graph G, denoted by rc(G), is the smallest number of colors that are needed in order to make G rainbow edge connected. In this paper we obtain tight bounds for rc(G). We use our results to generalize previous results for graphs with delta(G) >= 3.
We discuss numerical modeling attacks on several proposed strong physical unclonable functions (PUFs). Given a set of challenge-response pairs (CRPs) of a Strong PUF, the goal of our attacks is to construct a computer algorithm which behaves indistinguishably from the original PUF on almost all CRPs. If successful, this algorithm can subsequently impersonate the Strong PUF, and can be cloned and distributed arbitrarily. It breaks the security of any applications that rest on the Strong PUF's unpredictability and physical unclonability. Our method is less relevant for other PUF types such as Weak PUFs. The Strong PUFs that we could attack successfully include standard Arbiter PUFs of essentially arbitrary sizes, and XOR Arbiter PUFs, Lightweight Secure PUFs, and Feed-Forward Arbiter PUFs up to certain sizes and complexities. We also investigate the hardness of certain Ring Oscillator PUF architectures in typical Strong PUF applications. Our attacks are based upon various machine learning techniques, including a specially tailored variant of logistic regression and evolution strategies. Our results are mostly obtained on CRPs from numerical simulations that use established digital models of the respective PUFs. For a subset of the considered PUFs—namely standard Arbiter PUFs and XOR Arbiter PUFs—we also lead proofs of concept on silicon data from both FPGAs and ASICs. Over four million silicon CRPs are used in this process. The performance on silicon CRPs is very close to simulated CRPs, confirming a conjecture from earlier versions of this work. Our findings lead to new design requirements for secure electrical Strong PUFs, and will be useful to PUF designers and attackers alike.
All askers who post questions in Community-based Question Answering (CQA) sites such as Yahoo! Answers, Quora or Baidu's Zhidao, expect to receive an answer, and are frustrated when their questions remain unanswered. We propose to provide a type of "heads up" to askers by predicting how many answers, if at all, they will get. Giving a preemptive warning to the asker at posting time should reduce the frustration effect and hopefully allow askers to rephrase their questions if needed. To the best of our knowledge, this is the first attempt to predict the actual number of answers, in addition to predicting whether the question will be answered or not. To this effect, we introduce a new prediction model, specifically tailored to hierarchically structured CQA sites.We conducted extensive experiments on a large corpus comprising 1 year of answering activity on Yahoo! Answers, as opposed to a single day in previous studies. These experiments show that the F 1 we achieved is 24% better than in previous work, mostly due the structure built into the novel model.
In Web search, users may remain unsatisfied for several reasons: the search engine may not be effective enough or the query might not reflect their intent. Years of research focused on providing the best user experience for the data available to the search engine. However, little has been done to address the cases in which relevant content for the specific user need has not been posted on the Web yet. One obvious solution is to directly ask other users to generate the missing content using Community Question Answering services such as Yahoo! Answers or Baidu Zhidao. However, formulating a full-fledged question after having issued a query requires some effort. Some previous work proposed to automatically generate natural language questions from a given query, but not for scenarios in which a searcher is presented with a list of questions to choose from. We propose here to generate synthetic questions that can actually be clicked by the searcher so as to be directly posted as questions on a Community Question Answering service. This imposes new constraints, as questions will be actually shown to searchers, who will not appreciate an awkward style or redundancy. To this end, we introduce a learning-based approach that improves not only the relevance of the suggested questions to the original query, but also their grammatical correctness. In addition, since queries are often underspecified and ambiguous, we put a special emphasis on increasing the diversity of suggestions via a novel diversification mechanism. We conducted several experiments to evaluate our approach by comparing it to prior work. The experiments show that our algorithm improves question quality by 14% over prior work and that adding diversification reduced redundancy by 55%.
Community-based Question Answering sites, such as Yahoo! Answers or Baidu Zhidao, allow users to get answers to complex, detailed and personal questions from other users. However, since answering a question depends on the ability and willingness of users to address the asker's needs, a significant fraction of the questions remain unanswered. We measured that in Yahoo! Answers, this fraction represents 15% of all incoming English questions. At the same time, we discovered that around 25% of questions in certain categories are recurrent, at least at the question-title level, over a period of one year. We attempt to reduce the rate of unanswered questions in Yahoo! Answers by reusing the large repository of past resolved questions, openly available on the site. More specifically, we estimate the probability whether certain new questions can be satisfactorily answered by a best answer from the past, using a statistical model specifically trained for this task. We leverage concepts and methods from query-performance prediction and natural language processing in order to extract a wide range of features for our model. The key challenge here is to achieve a level of quality similar to the one provided by the best human answerers. We evaluated our algorithm on offline data extracted from Yahoo! Answers, but more interestingly, also on online data by using three "live" answering robots that automatically provide past answers to new questions when a certain degree of confidence is reached. We report the success rate of these robots in three active Yahoo! Answers categories in terms of both accuracy, coverage and askers' satisfaction. This work presents a first attempt, to the best of our knowledge, of automatic question answering to questions of social nature, by reusing past answers of high quality.
Modern consumers are inundated with choices. A variety of products are offered to consumers, who have unprecedented opportunities to select products that meet their needs. The opportunity for selection also presents a time-consuming need to select. This has led to the development of recommender systems that direct consumers to products expected to satisfy them. One area in which such systems are particularly useful is that of media products, such as movies, books, television, and music. We study the details of media recommendation by focusing on a large scale music recommender system. To this end, we introduce a music rating data set that is likely to be the largest of its kind, in terms of both number of users, items, and total number raw ratings. The data were collected by Yahoo! Music over a decade. We formulate a detailed recommendation model, specifically designed to account for the data set properties, its temporal dynamics, and the provided taxonomy of items. The paper demonstrates a design process that we believe to be useful at many other recommendation setups. The process is based on gradual modeling of additive components of the model, each trying to reflect a unique characteristic of the data.
Computational classification of gene expression profiles into distinct disease phenotypes has been highly successful to date. Still, robustness, accuracy, and biological interpretation of the results have been limited, and it was suggested that use of protein interaction information jointly with the expression profiles can improve the results. Here, we study three aspects of this problem. First, we show that interactions are indeed relevant by showing that co-expressed genes tend to be closer in the network of interactions. Second, we show that the improved performance of one extant method utilizing expression and interactions is not really due to the biological information in the network, while in another method this is not the case. Finally, we develop a new kernel method—called NICK—that integrates network and expression data for SVM classification, and demonstrate that overall it achieves better results than extant methods while running two orders of magnitude faster.
Computational classi cation of gene expression pro les into distinct disease phenotypes has been highly successful to date. Still, robustness, accuracy and biological interpretation of the results have been limited, and it was suggested that use of protein interaction information jointly with the expression pro les can improve the results. Here, we study three aspects of this problem. First, we show that interactions are indeed relevant by showing that co-expressed genes tend to be closer in the network of interactions. Second, we show that the improved performance of one extant method utilizing expression and interactions is not really due to the biological information in the network, while in another method this is not the case. Finally, we develop a new kernel method called NICK that integrates network and expression data for SVM classi cation, and demonstrate that overall it achieves better results than extant methods while running two orders of magnitude faster.
Daniel L. Silver合作论文数Jodrey School of Computer Science3
Misha Tsodyks合作论文数Department of Neurobiology
Weizmann Institute of Science3