
Compilation, Annotation and Analysis of Written Text Corpora
In this course we will provide an introduction to existing methods for the creation and analysis of linguistic corpora. This will be done to a large extend by hands-on demonstrations. We primarily focus on historical texts, though more recent literature will also be taken into account. Participants' own material can be discussed as well.
Firstly, we demonstrate and discuss the steps towards the creation of corpora, focusing on text selection, treatment of metadata, transcription and representation of the material, methods of (linguistic and text structural) annotation, and finally the presentation and provision of corpora. In this context, standards and best practices for data recognition and annotation will be introduced and applied by example, e. g. the guidelines of the Text Encoding Initiative (TEI) and specific formats for linguistic text annotation.
Secondly, we will introduce methods of corpus analysis, esp. under consideration of what the CLARIN infrastructure offers in this field. In this context, we will demonstrate the transformation of scientific research questions into complex corpus queries based on linguistic and structural phenomena. Participants will get to know and learn how to use specific tools for corpus analysis, e. g. how to perform complex linguistic searches on the Deutsches Textarchiv corpus considering spelling variance or how to combine services of CLARIN’s WebLicht.
2022
2021
2020
2019
- Home
- Schedule
- Workshops
- Lectures (public)
- Projects (public)
- Poster Session (public)
- Panel (public)
- Teasers (public)
- Cultural programme
- Experts
- Lecturers
- Scientific Committee
- Important dates (new)
- Application
- Scholarships (updated)
- Participation fees
- Refund policy
- T-Shirts
- Child care
- Birthday thoughts
2018
2017
2016
- Home
- Schedule
- Workshops
- XML-TEI encoding, structuring and rendering
- Compilation, Annotation and Analysis of Written Text Corpora
- Comparing Corpora
- Digital Editions and Editorial Theory
- Searching Linguistic Patterns in Large Text Corpora for Digital Humanities Research
- Lexicometric text analysis using CLARIN-D Webservices and R
- Stylometry
- Spoken Language and Multimodal Corpora
- Digital Lexica, Terminological Databases and Encyclopaedias
- Exploring art and technology within contemporary network culture
- From Text to Map: Modeling Historical Humanties Data in Mapping Environments
- Project Management
- Data management for the humanities
- Digital Research Infrastructures in the Humanities: How to Use, Build and Maintain Them
- Lectures (public)
- Projects & Posters (public)
- Panel
- Teasers (public)
- Slams
- Experts
- Lecturers
- Scientific Committee
- Important dates
- Application
- Scholarships
- Fees
- Refund policy
- Flyer
- Child Care












