Knowledge graphs have successfully been adopted by academia, governement and industry to represent large scale knowledge bases. Open and collaborative knowledge graphs such as Wikidata capture knowledge from different domains and harmonize them under a common format, making it easier for researchers to access the data while also supporting Open Science.Wikidata keeps getting bigger and better, which subsumes integration use cases. Having a large amount of data such as the one presented in a scopeless Wikidata offers some advantages, e.g., unique access point and common format, but also poses some challenges, e.g., performance.Regular wikidata users are not unfamiliar with running into frequent timeouts of submitted queries. Due to its popularity, limits have been imposed to allow for fair access to many.However this suppreses many interesting and complex queries that require more computational power and resources. Replicating Wikidata on one's own infrastructure can be a solution which also offers a snapshot of the contents of wikidata at some given point in time. There is no need to replicate Wikidata in full, it is possible to work with subsets targeting, for instance, a particular domain. Creating those subsets has emerged as an alternative to reduce the amount and spectrum of data offered by Wikidata. Less data makes more complex queries possible while still keeping the compatibility with the whole Wikidata as the model is kept. In this paper we report the tasks done as part of a Wikidata subsetting project during the Virtual BioHackathon Europe 2020 and SWAT4(HC)LS 2021, which had already started at NBDC/DBCLS BioHackathon 2019 in Japan, SWAT4(HC)LS hackathon 2019, and Virtual COVID-19 BioHackathon 2019. We describe some of approaches we identified to create subsets and some susbsets from the Life Sciences domain as well as other use cases we also discussed.
There are thousands of data repositories on the Web, providing access to millions of datasets. National and regional governments, scientific publishers and consortia, commercial data providers, and others publish data for fields ranging from social science to life science to high-energy physics to climate science and more. Access to this data is critical to facilitating reproducibility of research results, enabling scientists to build on others' work, and providing data journalists easier access to information and its provenance. In this paper, we discuss Google Dataset Search, a dataset-discovery tool that provides search capabilities over potentially all datasets published on the Web. The approach relies on an open ecosystem, where dataset owners and providers publish semantically enhanced metadata on their own sites. We then aggregate, normalize, and reconcile this metadata, providing a search engine that lets users find datasets in the "long tail" of the Web. In this paper, we discuss both social and technical challenges in building this type of tool, and the lessons that we learned from this experience.
Recent breakthroughs in the machine learning community have sometimes been characterized as a kind of steamroller. Having first crushed earlier results in image understanding, deep learning researchers have also been making impressive progress with both audio and video understanding and generation, and have had high profile and evocative results with natural language processing. While it is not surprising if this trend can seem threatening (in the examples given, it threatens to overturn previous methods) to a number of traditional subfields of AI (and their associated approaches), Semantic Web researchers and advocates would be ill-advised to consider deep learning a rival or threatening effort. In fact, we argue that there is much to learn from the success of these methods.
Separation between content and presentation has always been one of the important design aspects of the Web. Historically, however, even though most Web sites were driven off structured databases, they published their content purely in HTML. Services such as Web search, price comparison, reservation engines, etc. that operated on this content had access only to HTML. Applications requiring access to the structured data underlying these Web pages had to build custom extractors to convert plain HTML into structured data. These efforts were often laborious and the scrapers were fragile and error-prone, breaking every time a site changed its layout.
RDF Schema er et vokabularium til datamodellering af RDF-data. RDF Schema er en ekstension af det grundlæggende RDF-vokabularium.
Leo Obrst a,∗, Michael Gruninger b, Ken Baclawski c, Mike Bennett d, Dan Brickley e, Gary Berg-Cross f, Pascal Hitzler g, Krzysztof Janowicz h, Christine Kapp i, Oliver Kutz j, Christoph Lange k, Anatoly Levenchuk l, Francesca Quattri m, Alan Rector n, Todd Schneider o, Simon Spero p, Anne Thessen q, Marcela Vegetti r, Amanda Vizedom s, Andrea Westerinen t, Matthew West u and Peter Yim v a The MITRE Corporation, McLean, VA, USA b University of Toronto, Toronto, Canada c Northeastern University, Boston, MA, USA d Hypercube Ltd., London, UK e Google, London, UK f Knowledge Strategies, Washington, DC, USA g Wright State University, Dayton, OH, USA h University of California, Santa Barbara, Santa Barbara, CA, USA i JustIntegration, Inc., Kissimmee, FL, USA j Otto von Guericke University Magdeburg, Magdeburg, Germany k University of Bonn, Bonn, Germany; Fraunhofer IAIS, Sankt Augustin, Germany l TechInvestLab.ru, Moscow, Russia m The Hong Kong Polytechnic University, Hong Kong n University of Manchester, Manchester, UK o PDS, Inc., Arvada, CO, USA p University of North Carolina, Chapel Hill, NC, USA q Arizona State University, Phoenix, AZ, USA r INGAR (CONICET/UTN), Santa Fe, Argentina s Criticollab, LLC, Durham, NC, USA t Nine Points Solutions, LLC, Potomac, MD, USA u Information Junction, Fareham, UK v CIM Engineering, Inc., San Mateo, CA, USA
The Open Annotation Core Data Model specifies an interoperable framework for creating associations between related resources, annotations, using a methodology that conforms to the Architecture of the World Wide Web. Open Annotations can easily be shared between platforms, with sufficient richness of expression to satisfy complex requirements while remaining simple enough to also allow for the most common use cases, such as attaching a piece of text to a single web resource.An Annotation is considered to be a set of connected resources, typically including a body and target, where the body is somehow about the target. The full model supports additional functionality, enabling semantic annotations, embedding content, selecting segments of resources, choosing the appropriate representation of a resource and providing styling hints for consuming clients.n
The Open Annotation Core Data Model specifies an interoperable framework for creating associations between related resources, annotations, using a methodology that conforms to the Architecture of the World Wide Web. Open Annotations can easily be shared between platforms, with sufficient richness of expression to satisfy complex requirements while remaining simple enough to also allow for the most common use cases, such as attaching a piece of text to a single web resource.An Annotation is considered to be a set of connected resources, typically including a body and target, where the body is somehow about the target. The full model supports additional functionality, enabling semantic annotations, embedding content, selecting segments of resources, choosing the appropriate representation of a resource and providing styling hints for consuming clients.
AGRIS ofrece una gran recopilacion de referencias bibliograficas, tales como articulos de investigacion, estudios y tesis. Cada una incluye metadatos como conferencias, investigadores, editores, instituciones y palabras clave de diferentes tesauros como AGROVOC. Con el aumento de busquedas de texto completo y la disponibilidad en linea de un mayor volumen de material de investigacion, la funcion de los metadatos bibliograficos parece redundante. Cuando se consideran como una forma de modelacion que enfatiza las relaciones, conexiones y enlaces, los metadatos bibliograficos crecen en valor en la medida que la Web crece en conectividad. Pueden proporcionar a los investigadores un mapa de la comunidad global de investigacion, vinculando los resultados formales (articulos, datos) con la literatura gris (preimpresos, proyectos) mas amplia y con plataformas de comunicacion (blogs, foros) que ayudan a los investigadores a posicionar los resultados formales en un contexto mas amplio. Este trabajo describe la evolucion de la funcion de la base de datos bibliografica de AGRIS para convertirse en un nucleo de literatura pertinente a la investigacion agricola. Esta gran base de datos de 3 millones de recursos agricolas, recopilada por mas de 150 instituciones en los ultimos 35 anos, se esta convirtiendo en el punto de partida para tener acceso a los diversos conocimientos en ciencia y tecnologia agricola disponibles a nivel mundial en la Web.
AGRIS fournit une grande collection de references bibliographiques, comme les articles de recherche, etudes et theses, qui comprennent des metadonnees comme les conferences, chercheurs, editeurs, institutions, et mots-cles de differents thesaurus tel AGROVOC. Avec la hausse de la recherche sur plein texte et de la disponibilite en ligne de plus de materiel de recherche, le role des metadonnees bibliographiques apparait redondant. Quand elles sont considerees comme une forme de modelage qui souligne des relations, les connexions et les liens, la valeur des metadonnees bibliographiques augmente avec l’augmentation de la connectivite du Web. Elles peuvent fournir aux chercheurs une carte de la communaute de recherche globale, reliant des productions formelles (articles et donnees) avec une litterature grise plus large (pre-imprimes, brouillons), avec des plateformes de communication (blogs, forums) qui aident les chercheurs a mettre des conclusions formelles dans un contexte plus large. Cet article decrit l’evolution du role de la base de donnees bibliographiques AGRIS, qui devient un moyeu de la litterature sur la recherche agricole. La grande base de donnees de 3 millions de ressources agricoles, alimentee par plus de 150 institutions durant ces 35 dernieres annees, devient le point de depart pour acceder a la diverse connaissance des sciences et technologies agricoles disponibles mondialement sur le Web.
The Future Television workshop at EuroITV 2010 will explore how emerging Semantic Web and Social Web technologies can be integrated into the (increasingly Web-based) television experience to create new services and content offers around TV programming. It will bring together visionary minds from the TV, Social Web and Semantic Web communities to discuss and plan for a Future Television which will be both semantic and social.
Depuis des annees AGRIS offre une enorme collection de references bibliographiques, notamment des documents de recherche, des etudes et des theses, toutes accompagnees de metadonnees, telles que conferences, chercheurs, editeurs, institutions, et de mots-cles provenant de differents thesaurus comme AGROVOC. Avec l’accroissement de la recherche en texte integral et la disponibilite en ligne de documents de recherche de plus en plus nombreux, le role des metadonnees bibliographiques peut paraitre superflu. En revanche si on les considere comme une forme de modelisation qui met en evidence les relations, les connexions et les liens, leur valeur augmente en meme temps que la connectabilite sur le Web, et elles peuvent offrir une carte de la communaute mondiale des chercheurs, etablissant le lien entre les produits conventionnels (documents, donnees) et une litterature grise plus abondante (publications preliminaires, projets de textes) et des plateformes de communication (blogues, forums), qui aident les chercheurs a presenter des resultats officiels dans un contexte plus large. Le present document cherche a decrire le role en pleine evolution de la base de donnees bibliographiques AGRIS qui devient un centre de documentation scientifique agricole. Le reservoir gigantesque de 3 millions de sources d’informations agricoles, rassemblees par plus de 150 institutions depuis 35 ans, devient le point d’entree pour acceder a la diversite des connaissances dans le domaine des sciences et technologies agricoles qui sont disponibles a l’echelle mondiale sur le Web.
Shannon Bradshaw合作论文数Department of Computer Science;Intelligent Information Laboratory3
Jacco Van Ossenbruggen合作论文数Centrum voor Wiskunde en Informatica ( CWI )3
Mike Dean合作论文数Raytheon BBN Technologies2