Text Mining with Canonical Text Services - Using a Text Reference System for Citation Analysis, Text Alignment and more
A Canonical Text Service provides text passages that are specified by URN like references. It is specified in a way that allows to create CTS URNs for any possible text passage in a document.
The data can be requested using GET requests that are provided in an URL. Each request must contain one parameter request which specifies the CTS function to use. Function specific parameters - like the URN - are added as additional GET parameters.
For example, the following CTS request returns the text content of chapter 3 of the book Genesis of the English King James Bible: http://cts.informatik.uni-leipzig.de/pbc/cts/?request=GetPassage&urn=urn:cts:pbc:bible.parallel.eng.kingjames:1.3
Further information about CTS can be researched here.
The workshop aims to introduce the CTS protocol to new users and provides the tools to set up individual instances of CTS based on prepared data sets. At the end of the first two days, each participant can expect to have a running instance of CTS available online.
Once the CTS instances are up and running, participants will learn, how text data can be shared with other researchers and cloned between different instances of the system. The various tools and methods will be introduced, including two text alignment tools, a comprehensive CTS text mining framework and a workflow for citation analysis.
Programming skills are not required. Graphical managment tools for the work with CTS instances are available. The work with the text mining framework and the citation analysis requires a basic understanding of command line terminals (UNIX). Participants will work on pre prepared virtual machines. It is expected that participants are familiar with TEI/XML markup for digital documents. Teaching TEI/XML is not part of the workshop.
Participants may bring their own data sets into the workshop. For compliance, these documents should be encoded as UTF-8 and use a generic "TEI/XML div-type notation" similiar to this example. Other TEI/XML formats will propably also work. Non TEI/XML is currently not supported. Every participant must make sure that online publication of the texts does not violate license agreements.
Participants will get access to the programs and the freely available data sets that are part of Leipzigs CTS infrastructure, including documents from the Parallel Bible Corpus, the Deutsche Textarchiv, the TED Talk Transcripts and many more and are invited to use them after the workshop.
The workshop will last 1 week.
Introduction course to TEI/XML in week 1: From Print and Manuscript to Electronic Version: Text Digitization and Annotation). We aim at synchronizing the courses to enable students to re-use the results of week 1 if they want. The course is also open for students that did not attend the TEI/XML course.
2022
2021
2020
2019
2018
2017
- Important dates
- Schedule
- Workshops
- XML-TEI document encoding, structuring, rendering and transformation
- Hands on Humanities Data Workshop - Creation, Discovery and Analysis
- Introduction to programming for the Web
- From Print and Manuscript to Electronic Version: Text Digitization and Annotation
- Text processing for linguists and literary scholars with R
- Spoken Language and Multimodal Corpora
- Stylometry
- The Iconic Turn. Image Driven Digital Art History
- Humanities Data and Mapping Environments
- Working with SQL and graph databases
- Canonical Text Services
- Data Management and legal and ethical issues
- Teasers / Specials
- Lectures (public)
- Projects (public)
- Panel (public)
- Cultural Programme
- Experts
- Lecturers
- ConfTool
- Fees
- Refund Policy
- T-Shirt
- Child care
- Flyer
- Scientific Committee
- Scholarships
- Application