Classifying public tenders is a useful task for both companies that are invited to participate and for inspecting fraudulent activities. To facilitate the task for both participants and public administrations, the European Union presented a common taxonomy (Common Procurement Vocabulary, CPV) which is mandatory for tenders of certain importance; however, the contracts in which a CPV label is mandatory are the minority compared to all the Public Administrations activities. Classifying over a real-world taxonomy introduces some difficulties that can not be ignored. First of all, some fine-grained classes have an insufficient (if any) number of observations in the training set, while other classes are far more frequent (even thousands of times) than the average. To overcome those difficulties, we present a zero-shot approach, based on a pre-trained language model that relies only on label description and respects the label taxonomy. To train our proposed model, we used industrial data, which comes from contrattipubblici.org, a service by: Spazio Dati.s.r.l that collects public contracts stipulated in Italy in the last 25 years. Results show that the proposed model achieves better performance in classifying low-frequent classes compared to three different baselines, and is also able to predict never-seen classes.
The CoBiS is a network formed by 65 libraries. The project is a pilot for Piedmont that is aiming to provide the Committee with an infrastructure for LOD publishing, thus creating a triplification pipeline designed to be easy to automate and replicate. This is being realized with open source technologies, such as the RML mapping language or the JARQL tool that uses Linked Data to describe the conversion of XML, JSON or tabular data into RDF. The first challenge consisted in making possible the dialog of heterogeneous data sources, coming from four different library software (Clavis, Erasmo, SBNWeb and BIBLIOWin 5.0web) and different types of data (bibliographic, multimedia, and archival). The information contained in the catalogs is progressively interlinked with external data sources, such as Wikidata, VIAF, LoC and BNF authority files, Wikipedia and the Dizionario Biografico degli Italiani. Partners of the CoBiS LOD Project are: National Institute for Astrophysics (INAF), Turin Academy of Sciences, Olivetti Historical Archives Association, Alpine Club National Library, Deputazione Subalpina di Storia Patria, National Institute for Metrological Research (INRIM). The technical realization of the project is entrusted to Synapta, and it is partially sponsored by Piedmont Region.
The Italian anti-corruption Act (law n. 190/2012) requires all public administrations to spread procurement information as open data. Each body is obliged to yearly release standardized XML files, on its public website, containing data that describes all issued public contracts. Though this information is currently available on a machine-readable format, the data is fragmented and published in different files on different websites, without a unified and human-readable view of the information. The ContrattiPubblici.org project aims at developing a semantic knowledge graph based on linked data principles in order to overcome the fragmentation of existent datasets, to allow easy analysis, and to enable the reuse of information. The objectives are to increase public awareness about public spending, to improve transparency on the public procurement chain, and to help companies to retrieve useful knowledge for their business activities.
Public Procurement (PP) information, made available as Open Government Data (OGD), leads to tangible benefits to identify government spending for goods and services. Nevertheless, making data freely available is a necessary, but not sufficient condition for improving transparency. Fragmentation of OGD due to diverse processes adopted by different administrations and inconsistency within data affect opportunities to obtain valuable information. In this article, we propose a solution based on linked data to integrate existing datasets and to enhance information coherence. We present an application of such principles through a semantic layer built on Italian PP information available as OGD. As result, we overcame the fragmentation of datasources and increased the consistency of information, enabling new opportunities for analyzing data to fight corruption and for raising competition between companies in the market.
The CoBiS is a network formed by 65 libraries. The project is a pilot for Piedmont aiming to provide the libraries with an infrastructure for LOD publishing, creating a triplification pipeline designed to be easy to automate and replicate. This was realized with open source technologies, such as the TARQL and JARQL tools that use SPARQL queries to describe the conversion of tables (CSV) or trees (JSON) into graphs (RDF data). The first challenge consisted in making possible the dialog of heterogeneous data sources, coming from four different library applications and different types of data. As a second step, the information contained in the catalogs was interlinked with external data sources.
Le Pubbliche Amministrazioni (PA) accumulano dati. Lo fanno per poter funzionare e per dimostrare di aver ben funzionato. La rivoluzione digitale rende trascurabile il costo di mettere a disposizione tali dati per il riutilizzo ed aumenta il costo opportunità di limitarne l'uso alla finalità per cui sono stati raccolti. La normativa incentiva tale riutilizzo, ed esistono numerosi standard tecnici e buone pratiche utili a renderlo anche praticamente fattibile e sostenibile. In sintesi, la pubblicazione di dati aperti è oggi una buona pratica, ma anche un dovere per le PA. Il presente articolo amplia ed organizza i concetti di cui sopra, nell'ottica di fornire gli elementi per comprendere lo stato dell'arte dei dati aperti (intesi come approccio all'amministrazione pubblica, campo di ricerca e movimento). Dal punto di vista normativo, ci si sofferma sul contesto europeo ed italiano. Dal punto di vista tecnico, si offrono alcuni approfondimenti relativi al formalismo “linked data”. Il lavoro prosegue esemplificando più in dettaglio il caso dei dati pubblici relativi al patrimonio immobiliare delle PA italiane. Tale esempio è significativo in quanto al confine tra il mondo dei dati aperti per il riutilizzo tradizionalmente intesi e quello della trasparenza amministrativa; inoltre, la presenza di riferimenti territoriali nei dati (come minimo, il numero civico) fornisce una chiave per l'incrocio con altri dataset. Un’analisi delle pratiche di apertura di tali dati offre anche ottimi spunti per illustrare i limiti di una pubblicazione priva del necessario coordinamento e delle indispensabili linee guida tecniche.
Public Administrations (PA) collect data. They do that to function and for accountability purposes. The digital revolution implies that the cost of making these data available for reuse is negligible, while it increases the opportunity cost of limiting their use to the purpose for which they were originally collected. The law encourages such reuse, and there is a growing number of technical standards and good practices making that easier and sustainable. In short, nowadays, the publication of open data is a good practice, but also a duty for PAs. The paper at hand discusses the aforementioned concepts, with the purpose of providing the elements needed to understand the state of the art of open data (as an approach to government, a field of research and a movement). From the legal point of view, the focus in on the European and Italian jurisdictions. From the technical point of view, the “linked data” formalism is discussed in some details. The last part of the paper analyses the case of open data concerning government real estate in Italy. Such example is relevant since it is at the border between the open data and transparency domains; moreover, the presence of geographical references (at minimum, the address) provides a key to cross these data with other datasets. An analysis of the current publication practices in this domain is also functional to showing the limitations of a publication duty which is not accompanied by the necessary degree of coordination and by detailed technical guidelines.
The diffusion of Open Government Data (OGD) in recent years kept a very fast pace. However, evidence from practitioners shows that disclosing data without proper quality control may jeopardize dataset reuse and negatively affect civic participation. Current approaches to the problem in literature lack a comprehensive theoretical framework. Moreover, most of the evaluations concentrate on open data platforms, rather than on datasets. In this work, we address these two limitations and set up a framework of indicators to measure the quality of Open Government Data on a series of data quality dimensions at most granular level of measurement. We validated the evaluation framework by applying it to compare two cases of Italian OGD datasets: an internationally recognized good example of OGD, with centralized disclosure and extensive data quality controls, and samples of OGD from decentralized data disclosure (municipality level), with no possibility of extensive quality controls as in the former case, hence with supposed lower quality.Starting from measurements based on the quality framework, we were able to verify the difference in quality: the measures showed a few common acquired good practices and weaknesses, and a set of discriminating factors that pertain to the type of datasets and the overall approach. On the basis of this evaluation, we also provided technical and policy guidelines to overcome the weaknesses observed in the decentralized release policy, addressing specific quality aspects. (C) 2016 Elsevier Inc. All rights reserved.
The Web's evolution into a Semantic Web and the continuous increase in the amount of data published as linked data open up new opportunities for annotation and categorization systems to reuse these data as semantic knowledge bases. Accordingly, information extraction systems use linked data to exploit semantic knowledge bases, which can be interconnected and structured to increase the precision and recall of annotation and categorization mechanisms. TellMeFirst classifies and enriches textual documents written in English and Italian. Although various works present solutions for text annotation and classification, this article describes and studies the use case of a telecommunications operator that has adopted TellMeFirst to add value to two services available to its users: FriendTV and Society.
P2P technologies enable dissemination of content in an efficient way, especially if compared to the traditional techiques of content transmission over the Internet (by means of server-client network protocols, such as FTP, HTTP, etc). However, such efficiency must face the limits imposed by the law. In particular, file-sharing of protected subject matter is prohibited in all those cases where a prior authorization by right holders is absent. Such authorization is almost always missing, due to the extremely high transactive costs connected with its negotiation. This situation represents a huge market failure and restricts the freedom to access knowledge as granted by art. 27 sec. 1 of the Universal Declaration of Human Rights, with consequences that have a heavy negative impact on the cultural and economic development of our society. There has been few efforts in seeking mechanisms intended to ease the meeting between supply and demand of digital content. On the contrary, much effort has, in recent years, been put into limiting such a phenomena by leverage of the dissuasive power of criminal laws, and to involve access providers (ISP) in surveillance activities. This approach is clearly in contrast with fundamental and constitutional rights, and does not represent a solution to the market failure above mentioned. International and EU Community legislation allows for exceptions to the exclusive rights of reproduction and of making available to the public on-demand, provided that right holders are remunerated. This paper seeks to address the issue of file-sharing of copyrighted works, by analysing a variety of legal mechanisms and their compliance with international and European law. Among the possible solutions that can be taken (general taxation, special purpose tax, mandatory or voluntary licenses), we suggest that a system of collective extended licenses, if properly tuned, may solve the problems connected with the current situation of P2P, benefit the affected players economically, and increase the general welfare of society through a more efficient, fair, and open dissemination of culture and knowledge.
To explore the socio-technical aspects of the Internet requires infrastructures to properly foster interdisciplinary work and the development of appropriate research methods. To this end we present a platform called EINS Evidence Base (EINS-EB) which is developed as part of the EINS project. The EINS-EB also aims to empower researchers, academics, organisations and society to engage with Internet Science research independent of background. Currently, it provides for the collection and discovery of data resources, of analytic and simulation tools, and, in the future, of the methodologies behind those tools an of relevant scholarly activity. We explore issues of data representation, dataset description, dataset catalogues and method catalogues for Internet Science. The evidence base adopts semantic technologies to provide an interoperable catalogue of online resources related to Internet science. We also present activities on making the evidence base interoperable with related e-Science activities by communities engaging in relevant interdisciplinary collaboration.
There are now almost 700 Open Access policies around the world, two thirds of them in universities and research institutes. There is considerable variation across these policies in terms of the conditions they lay down for authors and of their effectiveness. This briefing paper lays out the main issues that affect the effectiveness of a policy in providing high levels of Open Access research material.
The Sui Generis Database Rights (SGDR) protection grants an exclusive right on databases when a substantial investment is required to collect and arrange the database contents. Since this specific protection makes any re-use of such contents impossible without an explicit permission, therefore directly impacting on the exploitation of Open Data, managing SGDR (where exsisting) - e.g. by adopting a license - is crucial for any public body who wants to make its data available for re-use. The paper examines the new features introduced in the 4.0 version of the Creative Commons Public Licenses, with particular attention to the treatment of SGDR to describe the suitability of the 4.0 version in the specific field of Open Data licensing and re-use. The evaluation has been conducted in light of the current EU legal framework on database rights, also considering the issue of interoperability with other existing database licenses
Context The diffusion of Linked Data and Open Data in recent years kept a very fast pace. However evidence from practitioners shows that disclosing data without proper quality control may jeopardize datasets reuse in terms of apps, linking, and other transformations. Objective Our goals are to understand practical problems experienced by open data users in using and integrating them and build a set of concrete metrics to assess the quality of disclosed data and better support the transition towards linked open data. Method We focus on Open Government Data (OGD), collecting problems experienced by developers and mapping them to a data quality model available in literature. Then we derived a set of metrics and applied them to evaluate a few samples of Italian OGD. Result We present empirical evidence concerning the common quality problems experienced by open data users when using and integrating datasets. The measurements effort showed a few acquired good practices and common weaknesses, and a set of discriminant factors among datasets. Conclusion The study represents the first empirical attempt to evaluate the quality of open datasets at an operational level. Our long-term goal is to support the transition towards Linked Open Government Data (LOGD) with a quality improvement process in the wake of the current practices in Software Quality.
M. Torchiano合作论文数Dept. of Control and Computer Engineering6