About Curatr
Project Overview
Curatr is a bespoke online platform designed to improve the accessibility of the British Library Nineteenth Century Digitised Books Collection. The collection is provided by British Library Labs and the platform was developed by the ERC-funded VICTEUR project at UCD English, Drama and Film, in collaboration with researchers at the Insight Research Ireland Centre for Data Analytics as part of Insight's Cultural Analytics Research Initiative. Curatr hosts digitised plain-text versions of the English-language portion of the collection, referred to here as BL19, corresponding to 35,884 unique out-of-copyright titles, both fiction and non-fiction, from 1700 to 1899. Accounting for multi-volume works, this totals 46,403 volumes of text.
The platform combines search tools and visual semantic network maps to bring overlooked texts to light and reveal connections between related terms and ideas across the collection. By integrating machine learning techniques into literary scholarship, Curatr opens BL19 to a range of scholarly approaches through multiple entry points, making it accessible to researchers without specialist training in digital humanities. The platform also incorporates the first digitised version of the topical classification index used by the British Museum Library from the early nineteenth century until 1973, and subsequently by the British Library until the end of 1985. Since legibility varies in earlier texts, the platform is also designed to help researchers locate and evaluate relevant titles before consulting them in person at the Library.
Platform Features
Curatr provides a range of tools for searching, browsing, and analysing the BL19 collection:
- Collection Search: a searchable index of 46,403 volumes and 12,322,488 text segments, filterable by author, title, year, classification, publication location, and document type, with sorting by relevance, date, or title.
- Classification Index: a browsable version of the hierarchical topical index used by the British Museum Library from the early nineteenth century until 1973, and subsequently by the British Library until the end of 1985, from broad categories such as "Fiction" and "Geography" down to more fine-grained sub-topics.
- Catalogue: a sortable and searchable table of all books in the collection, browsable by title, author, and year of publication.
- Authors: browse the collection by author, with links to all associated volumes.
- Ngram Viewer: plot the frequency of one or more words across the collection over time, with the option to click through to the corresponding search results for any given year. Results can be exported as a CSV file.
- Semantic Networks: visualise conceptual relationships in the collection by constructing interactive semantic networks from seed words, with associated words identified using word embedding models.
- Concordance: identify every occurrence of a particular word or phrase within the collection, presented alongside its immediate linguistic context.
- Word Lexicons: create and manage curated lists of keywords related to a given research topic. Lexicons can be expanded automatically using a word embedding model to suggest semantically similar terms, and used to drive searches and sub-corpus exports.
- Sub-Corpora: define and export smaller, topic-specific sub-corpora of the collection, filtered thematically, chronologically, and by classification, for close reading and offline analysis.
- Bookmarks: save volumes and individual text segments of interest to a personal reading list.
Resources & Contact
A series of short instructional videos from VICTEUR project researchers on the use and functionality of Curatr is available on the VICTEUR project website.
For queries on the use and development of Curatr, please contact us via email. Code and data is made openly available on Github.
Citing Curatr
If you use Curatr in an academic publication, we would appreciate citations to the following paper:
Leavy, S., Meaney, G., Wade, K. and Greene, D. (2019) Curatr: A Platform for Semantic Analysis and Curation of Historical Literary Texts, in Proceedings of the 13th International Conference on Metadata and Semantics Research (MTSR 2019) (pp. 354-366). Springer International Publishing. [PDF] [BibTeX] [RIS]
If you wish to cite a specific version of a text identified on the Curatr platform, we recommend the following citation format:
Dickens, C. (1892) The Old Curiosity Shop. Available at: curatr.ucd.ie (Accessed: 16 March 2026).
Acknowledgements
This work is part of the VICTEUR project, which has received funding from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (grant agreement No 884951), and is being undertaken by members of the UCD School of English, Drama and Film, in collaboration with researchers from the Insight Research Ireland Centre for Data Analytics at the UCD School of Computer Science. The original British Library Nineteenth Century Digitised Books collection was provided by British Library Labs. Curatr by UCD Centre for Cultural Analytics is licensed under a Creative Commons BY-NC-ND 4.0 Licence. Background image created by Steven Cadman.



