
Making stand-off markup and semantic web technologies converge
The amount of information that scholars need to associate to documents, in particular literary documents, is often staggering. Often what is needed are multiple, independent markup annotations that overlap each other on the same content, which requires ways to overcome the natural hierarchical characteristic of XML. Analogously, new research trends such as Linked Data, Open Data, Microformats and Semantic Publishing impose on the scholars the need to annotate their texts in such a way as to increase their availability, their understandability and their linkability by other parties connected over the Internet. In my talk I will describe a novel approach to address these issues in an integrated way, using as example a technology developed at the University of Bologna, called EARMARK. On the one hand, we will discuss how OWL ontologies, requiring only standard Semantic Web tools (including SPARQL, SWRL and simple SW engines) can be used to characterize and define markup languages, from the grammar to syntactical requirements for hierarchical containment and sequences of document parts; on the other hand, we will discuss how to overcome the basic syntactic limitations of XML as a markup metalanguage and propose how something based on RDF and graphs may be used to create markup languages that by construction are free of such limitations, allowing out of the box multiple hierarchies, overlaps and self-overlaps, and non-contiguity of fragments. On yet another hand, we will show the usefulness of unifying in a single vocabulary the characteristics of markup languages and embedded annotations (including microformats, schema.org and RDFa), so as to provide a simple and common approach to both the structural and the semantic annotation of a text resource. On the fourth and final hand of this strange Kali-like monster, we will show how, with the proper conceptual modeling (including a not-so-superficial visit of FRBR), we can annotate text resources with a very large number of facts, from the cultural context of the text to the scientific context of the publication, to the semantics of the structure, to the organization of the argumentation, to the actual semantics of the text content. Although we will make frequent examples from our own EARMARK proposal, the talk will be more or less independent of actual technical choices, and will strive to provide pointers to advanced conceptual reasonings about text resources, rather than ready-made recipes or technologies. A basic understanding of markup languages and semantic web is necessary. The talk will be (even dynamically) flexible to the competencies found in the audience with regard to details and introductions to overlapping, Linked data, embedded annotation formats, and FRBR.
2022
2021
2020
2019
- Home
- Schedule
- Workshops
- Lectures (public)
- Projects (public)
- Poster Session (public)
- Panel (public)
- Teasers (public)
- Cultural programme
- Experts
- Lecturers
- Scientific Committee
- Important dates (new)
- Application
- Scholarships (updated)
- Participation fees
- Refund policy
- T-Shirts
- Child care
- Birthday thoughts












