Text Service Infrastructure
The plain text inventory format is
URN [TAB] Title [TAB] Year [TAB] Author [TAB] Copyright restricted [TAB] lang [NL]
| Public Data Instances |
| Namespace | State | Access Endpoint (Inventory) | Inventory Structure (Treemap) | Lang | Copyright Restricted (*) | Content | Data Provenance | Downloadable Resources |
| ancJewLit | Stable | 📚 | 🖌 | heb,arc | Free | Classical Jewish sources | AncJewLit GitHub | DB+SRC VRT Word2Vec Model Current Database |
| dhd | Stable | 📚 | 🖌 | deu | Free | Abstracts of DHD conference published by Verband Digital Humanities im deutschsprachigen Raum e.V. | DHd-Verband GitHub | DB+SRC VRT Word2Vec Model Current Database |
| dsb | WIP, URNs not persistent |  | 🖌 | dsb(,deu,hsb) | Mixed | Lower Sorbian text corpus | Serbski Institute Operator / Admin | |
| dta | WIP | 📚 | 🖌 | deu | Free | Deutsches Textarchiv Kernkorpus + Erweiterung | Projektseite (BBAW) | |
| edh | Stable | 📚 | 🖌 | Multi | Free | The Epigraphic Database Heidelberg contains the texts of Latin and bilingual (i.e. Latin-Greek) inscriptions of the Roman Empire | Epigraphic Database Heidelberg | DB+SRC VRT Word2Vec Model Current Database |
| folgershakespeare | Stable | 📚 | 🖌 | eng | Free | All Shakespeare's works | Folger Shakespeare Library | DB+SRC VRT Word2Vec Model Current Database |
| gps4 | Stable | 📚 | 🖌 | deu | Free | German Political Speeches corpus compiled by Adrien Barbaresi (6'685 documents) | German Political Speeches Corpus and Visualization | DB+SRC VRT Word2Vec Model Current Database |
| gwtc | Stable | 📚 | 🖌 | eng | Fully | The Game Walkthrough Corpus reduced to 8'616 Neoseeker walkthroughs. | Game Walkthrough Corpus | |
| humboldtdigital | Stable | 📚 | 🖌 | deu | Free | Tagebücher, Briefe, Dokumente, Forschungsbeiträge, Chronologieeinträge der Edition Humboldt Digital (Version 11.0.1) | TELOTA (BBAW) GitHub | DB+SRC VRT Word2Vec Model Current Database |
| jeanpaulbriefe | Stable | 📚 | 🖌 | deu | Free | Daten der digitalen Briefedition "Jean Paul – Sämtliche Briefe digital" (https://www.jeanpaul-edition.de) | TELOTA (BBAW) GitHub | DB+SRC VRT Word2Vec Model Current Database |
| kant | Stable | 📚 | 🖌 | deu,lat,fra | Free | Kant’s gesammelte Schriften. Neuedition der Abteilung I (https://kant-digital.bbaw.de/) | TELOTA (BBAW) GitHub | DB+SRC VRT Word2Vec Model Current Database |
| lebenswelten | Stable | 📚 | 🖌 | deu | Free | Lebenswelten, Erfahrungsräume, politische Horizonte (Familie Lehndorff 18.-20. Jhdt) and Neuzeitlich-bäuerlicher Lebenswelten (ostpreußische Gutsarchive) | TELOTA (BBAW) GitHub | DB+SRC VRT Word2Vec Model Current Database |
| openarabicpe | Stable | 📚 | 🖌 | ara | Free | Open Arabic Periodical Editions (Muqtabas, Manar, Ustadh, Haqaiq, Lughat, Zuhur) | OpenArabicPE GitHub | DB+SRC VRT Word2Vec Model Current Database |
| pbc | Stable | 📚 | 🖌 | Multi | Free | 20 copyright-free parallel bible translations | Parallel Bible Corpus | DB+SRC VRT Word2Vec Model Current Database |
| pcp | Stable | 📚 | 🖌 | fra | Free | Chrétien de Troyes's Le Chevalier de la Charrette (Lancelot, ca. 1180) | The Princeton Charrette Project | DB+SRC VRT Word2Vec Model Current Database |
| textgrid | Stable | 📚 | 🖌 | deu | Free | Textgrid | The Digital Library in Textgrid | DB+SRC VRT Word2Vec Model Current Database |
| tgap | Stable | 📚 | 🖌 | eng,cat | Free | Thomas Gray Archiv Poems | Thomas Gray Archive | DB+SRC VRT Word2Vec Model Current Database |
| voth | Stable | 📚 | 🖌 | Multi | Free | David Boder: Voices of the Holocaust | David Boder: Voices of the Holocaust | DB+SRC VRT Word2Vec Model Current Database |
(*) The current list of requests that are available for copyright restricted texts is documented here.
|
| Online Tools |
Namespace Resolver - Resolve URN namespaces to API endpoints
oTAPIlot - Explore and query registered API endpoints
Openleaf - Read texts in ebook style
|
| Resources and Source Code |
Source Code Repositories (Git hosted via Bitbucket.org)
|
| Suggested Citation |
Jochen Tiepmar. 2025. Text Service Infrastructure. URL https://urncts.eu, requested on
|
Programming Interfaces & Export Files
Usage Examples are available via the left column
Selected Collections in E-Book Style
Pillar 2: Text Mining
| Text Mining Instances |
Default instances use the generic Text Miner. Custom instances may provide corpus specific features. See the source repository for installation and configuration. Subcorpus setups require a file named "urnlist.txt" as provided in the table.
|
| Source Code Resources |
Source Code and Installation (Git hosted via Bitbucket.org)
|
| Suggested Citation |
Jochen Tiepmar. Text Service Infrastructure. URL https://urncts.eu, requested on
|
Cooperation and Service Contracts
Need a custom text infrastructure? I provide CTS/DTS, text mining, corpus tools, APIs and data hosting. Contact me for a service contract.
Adding a TEI/XML corpus to the existing infrastructure is free of charge.
State of Infrastructure
Chronicle
Availability and Software-Versions (versionsoftware.php):
🖌 📚
Database Versions (versiondata.php):
🖌 📚
Central Enquiry
FAQ
What is this?
tl;dr: This is a decentralized text data service based on Canonical Text Service, Distributed Text Service and additional text APIs.
Is the implementation feature complete?
CTS: Stable. DTS: Work in Progress. See the Roadmap for planned features.
How do data requests work for texts that are copyright restricted?
Copyright-restricted data can be accessed through temporary, dataset-specific access tokens for authorized users.
Can I host my data on my own server ?
Yes. Data sets are available with database and source via Zenodo and can be deployed on your own server.
Are CTS URNs persistently citable?
CTS URNs are designed for persistent citation. Work-in-progress corpora are explicitly marked. Mutable data are versioned via the namespace.
How reliable is this service?
The software and copyright-free data sets are openly available and can be independently deployed. A future cloning mechanism will further reduce the dependence on individual servers.
Is it finally "A Library of a Billion Words"?
Academic Mentions
⇒
Jacob Langeloh (2024) : Unleash the Apparatus? Towards a shared representation of knowledge about connections between primary sources.
⇒
Thibault Clérice, Hugh Cayless, Jonathan Robie, Ian W. Scott (2026) : Distributed Text Services 1.0: An API Standard for Publishing and Extending TEI Document Collections.
Legal & Data Protection
Non-commercial academic research and data service.
Impressum and Data Protection