
Digital Archives: Reading and Manipulating Large-Scale Catalogues, Curating and Creating Small-Scale Archives
The purpose of this two-weeks workshop is to develop practical and critical skills toward representation of knowledge in digital archives and to build a small-scaled digital archive.
In the first week, we learn OpenRefine, performing ‘distant reading’ of catalogues. Using OpenRefine, students will gain critical insights on how to read, inspect the content and structure, clean and enrich digital catalogues in order to make them machine actionable and human readable. We will learn different ways of organizing the archive and enriching the data using authority files and linked open data, such as data from library of congress, VIAF and more.
In the second week, building on the practical experience and critical insights acquired in the first week, students will design and implement a digital archive for a collection of documents and will design its metadata. We will work with Omeka-s and Tropy, get into deeper details, experimenting with various forms of data representation.
The content of the workshop is both theoretical (concepts of archives, authority files, ontologies) as well as practical (hands-on work with specific tools).
No prior knowledge is needed or assumed.
First week - Reading and working with data / collections in OpenRefine
Digital data in various formats stand today in the heart of the humanist’s work: sometimes the files are very large, sometimes messy, sometimes structured and sometimes not. How can you tell what there is in the data? How is it organized? Can it be improved and enriched ? What can be understood from the data about the world, the research domain, about its creators?
OpenRefine was first developed by Google and then given to the community. It is an open tool for data processing and a skill which can be most valuable to anyone working with data. Purposed to do ‘data wrangling’, i.e., data cleaning, organizing, unifying and more, it allows one to get familiar with the content of her data: research results in spreadsheet tables, library or museum catalogues in MARC or XML, family trees in GEDCOM format, JSON files, tweets and more text files.
Learning to work in OpenRefine acquaintances with notions and tools that are useful in many aspects in DH. The workshop, therefore, goes beyond the mere familiarity with a tool (which is powerful by itself and deserves attention) - it gives the students a variety of skills, including utilizing API’s, harnessing the power of linked open data (LOD), scraping web pages, understanding clustering algorithms, and developing some programming notions:
- Different file types (CSV, TSV, Spreadsheets, JSON, XML TEI)
- Working effectively with regular expressions
- Writing expressions with GREL (the programming language)
- Working with API (geonames)
- Working with LOD (wikidata, Kima)
At the end of this course, the participants will be familiar with OpenRefine and with some of its advanced possibilities and experiment with various workflows such as scraping data from the web and organizing it and creating a map from a text.
Class 1: Introduction, loading a file and faceting Class 2: Regular expressions Class 3: Clustering Class 4: Enriching - Fetching data using REST API (working with GeoNames) HandsOn Session (working on your data, practicing administrative tasks such as changing working directory or memory size) Class 5: Reconciliation: enriching with Wikidata Class 6: Working with Jsons and XMLs Class 7: Web Scraping Class 8: From text to map, using OpenRefine Class 9: Summary
*Exact plan is prone to changes
Second week - Building a Digital Archive
Many researchers, libraries or archives own relatively small collections of documents that need digitization and structuring in order to be investigated, represented, searched. In many cases, documents as part of a larger narrative that needs to be told and presented.
But how to do that? Is there a necessity to employ a software engineer to implement it? And if so, how to communicate the subtle needs of the project, the nuances, the best practice that you acquire as a digital humanist?
In this workshop, we’ll move from the large-scale catalogues, and deal with the details. Structuring a digital archive is a series of choices, and digitization forces us to re-think of traditional methods of curating, presenting and representing archives.
We’ll begin with an overview of your collections and review examples of archives. We’ll discuss possible metadata schemes and the various considerations in the choice process - thinking of an archive as a body of knowledge that exists on its own. Then we’ll use and understand ‘best practice’ tools for the implementation.
Class 1: Theory of archives - An introduction Class 2: Digital archives - Examples, your collections Class 3: Working with primary resources, “Scanning party?” Class 4: Metadata - Methods of description, issues, dilemmas Class 5: Omeka-S - Introduction, creating an account Class 6: Tropy - Basic features, interface with Omeka Hands On Session - Working on your collection Class 7: Linking and integration with external resources Class 8: Publishing - Design Omeka pages Class 9: Summary
2022
- Home
- Important dates
- Application
- Workshops
- XML-TEI document encoding, structuring, rendering and transformation
- Hands on Humanities Data Workshop - Creation, Discovery and Analysis
- Stylometry
- Text processing for linguists and literary scholars with R
- Distant Reading in R. From Text Analysis to Mapping
- Visual Artificial Intelligence for the Digital Humanities
- Procedural Creativity and Digital Humanities Scholarship
- Digital Archives: Reading and Manipulating Large-Scale Catalogues, Curating and Creating Small-Scale Archives
- Making an edition of a text in many versions
- Project Planning and Management in the Digital Humanities
- Experts
- ConfTool
- Scholarships etc.
- Participation fees
- Moodle
- Scientific Committee
2021
2020
2019
- Home
- Schedule
- Workshops
- Lectures (public)
- Projects (public)
- Poster Session (public)
- Panel (public)
- Teasers (public)
- Cultural programme
- Experts
- Lecturers
- Scientific Committee
- Important dates (new)
- Application
- Scholarships (updated)
- Participation fees
- Refund policy
- T-Shirts
- Child care
- Birthday thoughts












