Tables are ubiquitously used to share relational data in various media and formats. Particularly, an enormous number of HTML tables are contained in web pages. They are a valuable data source for applications of web mining, question-answering, and knowledge base construction. However, not all HTML tables are genuine (i.e. containing relational data). Most of them are utilised as means of layout and navigation. In turn, the genuine tables can have different features of layouts, formatting, and content. One of the prevalent problems in web table extraction is to determine main functional and layout types and properties of tables wide-spread on the Web. Currently, a wide range of taxonomies of table types and properties is available to researchers and practitioners. All of these taxonomies could be utilised for table type classification in order to choose further type-specific treatment of tabular data. The existing taxonomies provide similar table types but use the confusing terminology. This paper is an attempt to overview the existing taxonomies of table types by matching their terminology and comparing them qualitatively.
Baikal Natural Territory (BNT) is the territory that adjacent to Lake Baikal, which is a unique natural object and, in accordance with the UNESCO Convention, a "World Natural Heritage". Baikal is in the central part of the Baikal Rift Zone (BRZ) – the most active seismic zone located in the middle of Russia. The development of the BRZ leads to the emergence of dangerous geological processes that can lead to a violation of the balance in the Lake Baikal ecological system and the surrounding area. In addition, these processes and phenomena pose a real threat to the smooth functioning of mainline communications, hydroelectric power plants and strategically important industries in the region, which, according to the classification of the Ministry of Emergency Situations of Russia, belongs to the first category of danger. To ensure systematic monitoring and forecasting of the environmental situation of the BNT, systematic observations are organized, as well as obtaining and analysing information about the activity of hazardous geological processes in digital form. The digital transformation of monitoring of hazardous geological processes, resulting from the digitalization of processes and the development of appropriate infrastructure, provides the possibility of using new models and methods, more flexible approaches to the analysis of ongoing processes and the prediction of possible extreme events. In this paper, a digital platform is proposed that provides support for the digital transformation of the monitoring of hazardous geological processes using the example of BNT. The platform under consideration may be used for ecological monitoring of BNT area.
Большой объем нередактируемых документов публикуется и распространяется в формате PDF. Часто они являются “неразмеченными”, т. е. не сопровождаются аннотацией о собственной структуре, в них нет метаданных о месторасположении заголовков, параграфов, абзацев, таблиц, списков, рисунков, колонтитулов и пр. Анализ компоновки документов состоит в распознавании перечисленных элементов структуры. Базовой частью этого процесса является сегментация текста внутри страниц на блоки, которые затем можно классифицировать как заголовки, абзацы, ячейки таблиц и пр. Известные алгоритмы сегментации страниц в основном предназначены для работы либо с растровыми изображениями документов, либо с печатно-ориентированным ASCII-текстом. По сравнению с этими форматами данных PDF предоставляет дополнительную информацию (порядок рендеринга, шрифтовые метрики, линейки и пр.), которая может улучшить качество анализа компоновки документов. В работе излагается опыт адаптации некоторых существующих алгоритмов сегментации текста внутри страниц изображений документов и ASCII-текста, для того чтобы сделать их применимыми напрямую к формату PDF - неразмеченным случаям. Currently, a large amount of non-editable documents are published and distributed in PDF (Portable Document Format). Often, they are “untagged”, i. e. there are no annotation about their structure, including headings, paragraphs, tables, lists, figures, footers, etc. The document layout analysis consists in recognizing the listed elements of the structure. A basic part of this process is the segmentation of page text into blocks that can be classified as headings, paragraphs, table cells, etc. The well-known page segmentation algorithms are mainly designed to deal with either bitmap images of document pages or print-oriented ASCII text. Compared to these data formats, PDF provides additional information (rendering order, font metrics, ruling lines, etc.) that can improve document layout analysis. The paper describes our experience on the adaptation of some existing algorithms for segmenting page text in document images and ASCII text to make them applicable directly for PDF format - untagged cases.
The latest forecasts indicate wildfire activity in many parts of the world. Wildfire smoke contains hazardous air pollutants such as carbon monoxide, nitrogen dioxide, ozone, particulate matter et cetera. However, prediction of this impact and on time medical care are difficult due to the lack of digital decision-making systems. The aim of this study is to assess population health risks associated with the sub-daily exposure to wildfire smoke produced by massive foci of combustion near the populated areas and at a significant distance from them. We consider reflex reactions as a response to a short-term exposure. The maximum value of the 95th percentile from the series of observations at the monitoring point was used to assess the hazard. For the mathematical description of the “concentration-effect” relationship, the model of individual thresholds is applicable. This model describes a dependence as a straight line under the condition that the concentration is expressed in the form of a normal-probabilistic scale. The frequency of additional cases is determined by studying the number of requests for medical assistance (including calls for ambulance) with complaints of respiratory disorders, lacrimation, etc. on the territories affected by wildfires smokes. The indicator is calculated per 1000 population. The probability of negative biological effects in response to the impact of wildfire smoke is associated mainly with the content of CO and TPM in the conditions of the Baikal region. The frequency of additional requests for medical care ranged from 0.137 to 0.933 per 1000 exposed population during the fire period in settlements where risk levels are >0.01. We developed a digital environment that allows us to get information about harmful substances in the outdoor air from different sources and in different formats and data schemes. The digital environment supports implementation of models for assessing hazards to human body organs.
Monitoring and analysing data on environmental pollution during forest fires and their health impacts in different geographical areas will help to improve the quality of the risks of determining adverse effects on health and identifying vulnerable groups of the population. The algorithm for assessing the potential and realised health risk in a dangerous period includes several consecutive stages. For the algorithmic implementation of methods for assessing the impact of air quality on public health, a service-oriented geoportal system is being created. Specialised original services have been developed. They allow users to perform all the basic operations with the file system on the server through the user's browser only. For web-services effective applying, Jupiter Notebook is used. In addition to standard libraries, it is possible to use WPS services that increase data processing capabilities.
The Web stores a large volume of web-tables with semi-structured data. The Semantic Web community considers them as a valuable source for the knowledge graph population. Interrelated named entities can be extracted from web-tables and mapped to a knowledge graph. It generally requires reconstructing the semantics missing in web-tables to interpret them according to their meaning. This paper discusses prospects of an end-to-end solution for the knowledge graph population by entities extracted from web-tables of predefined types. The discussion covers theoretical foundations both for transforming data from web-tables to entity sets (table analysis) and for mapping entities, attributes, and relations to a knowledge graph (semantic table annotation). Unlike general-purpose text mining and web-scraping tools, we aim at developing a solution that takes into account the relational nature of the information represented in web-tables. In contrast to the table-specific proposals, our approach implies both the table analysis and the semantic table annotation.
Spreadsheet tables are one of the most commonly used formats to organise and store sets of statistical, financial, accounting and other types of data. This form of data representation is widely used in science, education, engineering, and business. The key feature of spreadsheet tables that they are generally created by people in order to be further used by other people rather than by automated programs. During spreadsheet creation, commonly, no consideration is given to the possibility of further automated data processing. This leads to a large variety of possible spreadsheet table structures and further complicates automated extraction of table content and table understanding. One of the key factors that influence on the quality of table understanding by machines is the correctness of the header structure, for example, position and relation between cells. In this paper, we present a case study of a tabular data extraction approach and estimate its performance on a variety of datasets. The rule-driven software platform TabbyXL was used for tabular data extraction and canonicalisation. The experiment was conducted on real-world tables of SAUS200 (The 2010 Statistical Abstract of the United States) corpora. For the evaluation, we used spreadsheet tables as they are presented in SAUS; the same tables, but with an automatically corrected header structure; and tables where the structure of the header was corrected by experts. The case study results demonstrate the importance of header structure correctness for automated table processing and understanding. The ground-truth preparation procedures, example of rules describing relationships between table elements, and results of the evaluation are presented in the paper.
The freely available tabular data represented in various digital formats, such as print-oriented documents, spreadsheets, and web pages, are a valuable source to populate knowledge graphs. However, difficulties that inevitably arise with the extraction and integration of the tabular data often hinder their intensive use in practice. TabbyDOC project aims at elaborating a theoretical basis and developing open software for data extraction from arbitrary tables. Previously, it was devoted to the following issues: (i) table extraction tables from print-oriented documents, (ii) data transformation from spreadsheet tables to relational and linked data. This paper summarizes the project’s results that are intended for the following tasks: (i) automation of fine-tuning artificial neural networks for table detection in document images, (ii) a synthesis of programs for spreadsheet data transformation driven by user-defined rules of table analysis and interpretation, and (iii) generating RDF-triples from entities extracted from relational tables.
The territories of the Baikal Region and Mongolia belong to the areas with evaluated seismic activity. In turn, these territories are at heightened risk of potentially damaging events for human socio-economic activity. Consequently, seismic activity recording and forecasting would allow us to minimise possible damages. These issues require collecting and processing large volumes of heterogeneous data. In order to effectively process such datasets, it would be necessary to use state-of-the-art information technologies and expandable analytical systems. Such systems should have a set of tools to enable collection, generation, transformation, visualisation and analysis of data. However, while implementing products of these types, developers, normally tend to use low-level tools for programming (various general-purpose programming languages and standard DBMS capabilities). On the other hand, developers tend to create highly specialised systems that are closely related to a specific automation object and focus on certain data structures. The paper considers an approach to the development of an informationanalytical system as an infrastructure element for assessing the seismic hazard of large lithospheric blocks of the Baikal region and Mongolia.
Nowadays methods and software for extracting tables from document images and portable documents (PDF) continue to be actively developed. One of the promising approaches to this task is the usage of fine-tuned object detection models. However, this approach involves many manipulations with data preparation and training process configuration. This paper proposes an automated workflow for fine-tuning deep neural network models for the table detection in document images. It enables us to automate two sub-tasks: (i) preparing a training dataset in the PascalVOC format with image transformation and augmentation; (ii) training a table detection model by using the well-known Faster R-CNN architecture. Implementation of the workflow design simplifies the use of the approach proposed by decreasing the number of required manipulations.
A spreadsheet is one of the most commonly used forms of representation for datasets of similar type. Spreadsheets provide considerable flexibility for data structure organisation. As a result of this flexibility, tables with very complex data structures could be created. In turn, such complexity makes automatic table processing and data extraction a challenging task. Therefore, table preproccessing step is often required in the data extraction pipeline. This paper proposes a heuristic algorithm for the correction of a table header in a spreadsheet. The aim of the proposed algorithm is to transform a machine-readable structure of the table header into its visual representation. The algorithm achieves this aim by iterating through table header cells and merging some of them according to proposed heuristics. The transformed structure, in turn, allows to improve quality of spreadsheet understanding and data extraction further in the pipeline. The proposed algorithm was implemented in the TabbyXL toolset.
Аннотация.Исследование сейсмического потенциала территорий является одной из важных задач, оказывающих значимое влияние на их социальноэкономическое развитие.Такие исследования чрезвычайно актуальны для территорий сейсмоактивного Монголо-Байкальского региона.Оценка сейсмического потенциала, в том числе построение карт энергии сейсмотектонического деформирования литосферы, требует обработки большого объема пространственных данных в значительные временные интервалы.Обработка такого массива данных является трудозатратной.Также в случае использования данных, полученных из различных источников, требуется их предобработка, очистка для осуществления возможности совместного использования и интеграции.В рамках данной работы предложен подход к автоматизации исследования трудоемкой задачи сейсмического районирования и построения карт энергий
Authoring documents is based on analysis of existing documents, collecting and processing their data, which is time consuming creative work. New documents may use parts of existing ones, e.g., header data, names of persons and titles of organizations, tabular data and their description, footers, etc. Semantic description of frequently used classes of documents allows simplifying the authoring process. The paper presents an approach to creation of open document catalogs that involves technologies of Linked Open Data, document templates, and logical inference for deducing parts of documents from data of the catalogs. The approach is being realized as open software toolset and services on the base of other open source systems. The software under development is used to construct information system components in various application areas, e.g., for synthesis educational processes regulations in departments of Irkutsk State University and organizations of joint accounting for source data collection.
The paper is devoted to the problem of an end-to-end table transformation from untagged portable documents (PDF) to linked data. It covers the issues of the table extraction from documents, the reconstruction of logical table structure, the conceptualization of their natural-language content, and the linking of extracted data with external vocabularies. We consider some perspective approaches for the deeplearning-based table detection, heuristic-based table structure recognition, rule-based table analysis, and knowledge-based table interpretation. They can be used as a basis to develop a consistent solution for this problem. Our application experience confirms that such solutions are demanded for populating databases and generating ontologies with tabular data being extracted from weakly and semi-structured documents.
Tables in electronic documents (spreadsheets) contain large volumes of useful information about different domains. Efficient extraction of data from document tables plays a crucial role in its further usage including analysis and integration. The visual or logical structure of table elements might differ from its physical structure. Such differences cause difficulties for automated table processing and understanding. Automated correction from physical form to visual allows to simplify tables processing operations. In this paper, we propose a heuristic approach for transformation of tables’ header cells. The main goal of the proposed approach is to provide an algorithm and software tool for recovering a physical structure of a spreadsheet header. The proposed approach is illustrated by application to the Statistical Abstract of the United States (SAUS) dataset.
This paper presents an approach to rule-based spreadsheet data extraction and transformation. We determine a table object model and domain-specific language of table analysis and interpretation rules. In contrast to the existing data transformation languages, we draw up this process as consecutive steps: role analysis, structural analysis, and interpretation. To the best of our knowledge, there are no languages for expressing rules for transforming tabular data into the relational form in terms of the table understanding. We also consider a tool for transforming spreadsheet data from arbitrary to relational tables. The performance evaluation has been done automatically for both (role and structural) stages of table analysis with the prepared ground-truth data. It shows high F-score from 95.82% to 99.04% for different recovered items in the existing dataset of 200 arbitrary tables of the same genre (government statistics).
The problem of software modeling having various models as sources and its transformation based on logical inference is considered. Source models are converted into graphs of RDF and then processed with knowledge based system organized in a network of objects. Various notation can be used to represent models, such as UML, SysML, CMMN, BPMN2.0, as well as RDF graphs and analyzed source code, having implemented a corresponding converter. The objects are represented in the LogTalk programming language. Objects query graphs and other objects implementing a scenario of a software system synthesis within Model Driven Architecture paradigm. Usage of such kind of transformation approach allows us to develop software system carcasses on the level of abstract models, involve various sources of model data in the transformation, define and structuring conversion knowledge as objects. An example of a dataflow environment synthesis encapsulating Mothur library for new generation sequencing is presented.