Uploaded July 2019 | Updated September 2026, 2 weeks ago
Dr Tanja Säily, tenure-track assistant professor in English Language at the University of Helsinki, and Dr Eetu Mäkelä, tenure-track assistant professor in Human Sciences–Computing Interaction at the University of Helsinki, presented their research on neologism use in the Corpus of Early English Correspondence and how the OED, as well as large historical text collections, are semi-automatically consulted within the project.
Dr Säily and Dr Mäkelä provided an overview of the project and the possibilities of using the OED in digital humanities research in general.
This session covered:
· An overview of the neologism research project, with examples of new vocabulary and its social embedding in the Corpus of Early English Correspondence
· Semi-automatic workflows for moving between the OED and historical text collections
· Building a pipeline from the corpus, through automated spelling normalization, lemmatization, filtering, comparison, and analysis to results
· The barriers faced with regard to historical spelling variation, noise from optical character recognition and material bias
· A discussion of the solutions applied, including statistical and neural machine translation, look-ups against comparison corpora and user interfaces for manual checking of results
· What the researchers wished they knew before they started
· What the future looks like for this and other projects using the OED
· Q&A session
Who is this talk for?
· Researchers working in similar projects – using the OED or not
· Anyone interested in historical sociolinguistics, historical lexicology, or historical lexicography
· People interested in the complexities of applying computational methods to historical data
Dr Tanja Säily, tenure-track assistant professor in English Language at the University of Helsinki, and Dr Eetu Mäkelä, tenure-track assistant professor in Human Sciences–Computing Interaction at the University of Helsinki, presented their research on neologism use in the Corpus of Early English Correspondence and how the OED, as well as large historical text collections, are semi-automatically consulted within the project.
Dr Säily and Dr Mäkelä provided an overview of the project and the possibilities of using the OED in digital humanities research in general.
This session covered:
· An overview of the neologism research project, with examples of new vocabulary and its social embedding in the Corpus of Early English Correspondence
· Semi-automatic workflows for moving between the OED and historical text collections
· Building a pipeline from the corpus, through automated spelling normalization, lemmatization, filtering, comparison, and analysis to results
· The barriers faced with regard to historical spelling variation, noise from optical character recognition and material bias
· A discussion of the solutions applied, including statistical and neural machine translation, look-ups against comparison corpora and user interfaces for manual checking of results
· What the researchers wished they knew before they started
· What the future looks like for this and other projects using the OED
· Q&A session
Who is this talk for?
· Researchers working in similar projects – using the OED or not
· Anyone interested in historical sociolinguistics, historical lexicology, or historical lexicography
· People interested in the complexities of applying computational methods to historical data










