
Lexicometric text analysis using CLARIN-D Webservices and R
In data analysis the programming language R recently gained a lot of attention also in the humanities and social sciences. Although it genuinely was developed for analysis of structured numerical data, it has been proven useful for text mining and lexicometric analysis as well. A crucial point in lexicometrics is the linguistic preprocessing of natural language texts for further analyses. For elaborated preprocessing the CLARIN-D infrastructure provides a number of useful webservices, e.g. sentence segmentation, tokenization, lemmatization, POS-tagging and named entity recognition. The course introduces into the use of CLARIN-D webservices within R to consciously prepare and perform lexicometric analysis such as key term extraction and cooccurrence analysis.
The workshop consists of theoretical and practical lessons. After a short introduction some important concepts of Text Mining will be presented. The course will focus on dealing with Text Mining in R and WebLicht webservices. It is aimed at users who are new to R covering information on the installation and basic concepts like data import/ export, data types, operators etc. There will also be a short introduction into WebLicht covering its capabilities and usage. This knowledge is then used to work with WebLicht webservices within R.
The practical parts will be based on examples so that participants can apply knowledge from the theoretical parts. We will cover where to get and how to import different text resources into R and how to examine data. The participants will learn how to preprocess text data which includes techniques such as sentence segmentation and tokenization. Such preprocessing steps will then allow to apply further techniques like Part-of-Speech-Tagging or Named-Entity-Recognition. WebLicht webservices will be used within R to support these tasks. Based on the resulting annotations, R will be used for frequency-related analyses, e.g. cooccurrences. Additionally, participants will learn how to visualize the results from within R without using external tools.
2022
2021
2020
2019
- Home
- Schedule
- Workshops
- Lectures (public)
- Projects (public)
- Poster Session (public)
- Panel (public)
- Teasers (public)
- Cultural programme
- Experts
- Lecturers
- Scientific Committee
- Important dates (new)
- Application
- Scholarships (updated)
- Participation fees
- Refund policy
- T-Shirts
- Child care
- Birthday thoughts
2018
2017
2016
- Home
- Schedule
- Workshops
- XML-TEI encoding, structuring and rendering
- Compilation, Annotation and Analysis of Written Text Corpora
- Comparing Corpora
- Digital Editions and Editorial Theory
- Searching Linguistic Patterns in Large Text Corpora for Digital Humanities Research
- Lexicometric text analysis using CLARIN-D Webservices and R
- Stylometry
- Spoken Language and Multimodal Corpora
- Digital Lexica, Terminological Databases and Encyclopaedias
- Exploring art and technology within contemporary network culture
- From Text to Map: Modeling Historical Humanties Data in Mapping Environments
- Project Management
- Data management for the humanities
- Digital Research Infrastructures in the Humanities: How to Use, Build and Maintain Them
- Lectures (public)
- Projects & Posters (public)
- Panel
- Teasers (public)
- Slams
- Experts
- Lecturers
- Scientific Committee
- Important dates
- Application
- Scholarships
- Fees
- Refund policy
- Flyer
- Child Care












