Text Service Infrastructure




Pillar 1: Decentralized Text APIs
Canonical Text Service (CTS) or Distributed Text Service (DTS)

The plain text inventory format is
URN [TAB] Title [TAB] Year [TAB] Author [TAB] Copyright restricted [TAB] lang [NL]

Public Data Instances
NamespaceStateAccess Endpoint (Inventory)Inventory Structure (Treemap)LangCopyright Restricted (*)ContentData ProvenanceDownloadable Resources
ancJewLitStable📚🖌heb,arcFreeClassical Jewish sourcesAncJewLit GitHubDB+SRC
VRT
Word2Vec Model
Current Database
dhdStable📚🖌deuFreeAbstracts of DHD conference published by Verband Digital Humanities im deutschsprachigen Raum e.V. DHd-Verband GitHubDB+SRC
VRT
Word2Vec Model
Current Database
dsbWIP, URNs not persistent🖌dsb(,deu,hsb)MixedLower Sorbian text corpus Serbski Institute
Operator / Admin
dtaWIP📚🖌deuFreeDeutsches Textarchiv Kernkorpus + Erweiterung Projektseite (BBAW)
edhStable📚🖌MultiFreeThe Epigraphic Database Heidelberg contains the texts of Latin and bilingual (i.e. Latin-Greek) inscriptions of the Roman EmpireEpigraphic Database HeidelbergDB+SRC
VRT
Word2Vec Model
Current Database
folgershakespeareStable📚 🖌engFreeAll Shakespeare's worksFolger Shakespeare LibraryDB+SRC
VRT
Word2Vec Model
Current Database
gps4Stable📚🖌deuFreeGerman Political Speeches corpus compiled by Adrien Barbaresi (6'685 documents)German Political Speeches Corpus and Visualization DB+SRC
VRT
Word2Vec Model
Current Database
gwtcStable📚🖌engFullyThe Game Walkthrough Corpus reduced to 8'616 Neoseeker walkthroughs. Game Walkthrough Corpus
humboldtdigitalStable📚🖌deuFreeTagebücher, Briefe, Dokumente, Forschungsbeiträge, Chronologieeinträge der Edition Humboldt Digital (Version 11.0.1)TELOTA (BBAW) GitHubDB+SRC
VRT
Word2Vec Model
Current Database
jeanpaulbriefeStable📚🖌deuFreeDaten der digitalen Briefedition "Jean Paul – Sämtliche Briefe digital" (https://www.jeanpaul-edition.de)TELOTA (BBAW) GitHubDB+SRC
VRT
Word2Vec Model
Current Database
kantStable📚🖌deu,lat,fraFreeKant’s gesammelte Schriften. Neuedition der Abteilung I (https://kant-digital.bbaw.de/)TELOTA (BBAW) GitHubDB+SRC
VRT
Word2Vec Model
Current Database
lebensweltenStable📚🖌deuFreeLebenswelten, Erfahrungsräume, politische Horizonte (Familie Lehndorff 18.-20. Jhdt) and Neuzeitlich-bäuerlicher Lebenswelten (ostpreußische Gutsarchive)TELOTA (BBAW) GitHubDB+SRC
VRT
Word2Vec Model
Current Database
openarabicpeStable📚🖌araFreeOpen Arabic Periodical Editions (Muqtabas, Manar, Ustadh, Haqaiq, Lughat, Zuhur)OpenArabicPE GitHubDB+SRC
VRT
Word2Vec Model
Current Database
pbcStable📚🖌MultiFree20 copyright-free parallel bible translations Parallel Bible CorpusDB+SRC
VRT
Word2Vec Model
Current Database
pcpStable📚🖌fraFreeChrétien de Troyes's Le Chevalier de la Charrette (Lancelot, ca. 1180)The Princeton Charrette ProjectDB+SRC
VRT
Word2Vec Model
Current Database
textgridStable📚🖌deuFreeTextgridThe Digital Library in TextgridDB+SRC
VRT
Word2Vec Model
Current Database
tgapStable📚🖌eng,catFreeThomas Gray Archiv Poems Thomas Gray ArchiveDB+SRC
VRT
Word2Vec Model
Current Database
vothStable📚🖌MultiFreeDavid Boder: Voices of the Holocaust David Boder: Voices of the HolocaustDB+SRC
VRT
Word2Vec Model
Current Database
(*) The current list of requests that are available for copyright restricted texts is documented here.
Online Tools Namespace Resolver - Resolve URN namespaces to API endpoints
oTAPIlot - Explore and query registered API endpoints
Openleaf - Read texts in ebook style
Resources and Source Code Source Code Repositories (Git hosted via Bitbucket.org)
Suggested Citation Jochen Tiepmar. 2025. Text Service Infrastructure. URL https://urncts.eu, requested on

Programming Interfaces & Export Files

Usage Examples are available via the left column

Python Programming Library
Javascript Programming Library
Verticalized Text
Corpus Workbench
Sketch-Engine

(WIP)
File Generation Script
Example File
Word Embeddings
Word2Vec
Word Vectors
Language Models

(WIP)
File Generation Script
Example File
Canonical Text Service
CTS Specifications
Distributed Text Service
(WIP)
DTS Specifications
Text Search
(WIP)
Exakte Suche

Selected Collections in E-Book Style

Autor Hans Christian Andersen | Wilhelm Busch | Johann Wolfgang von Goethe | Grimm's Fairy Tales | Franz Kafka | Karl May | Friedrich Schiller | Shakespeare





Pillar 2: Text Mining

Text Mining Instances
Text Mining CorporaSource Document ListTime SeriesType
ancJewLitFullNodefault
dhdFullYesdefault
dsb
Virtual Hardware @SI
Copyright-FreeYescustom
dsb
Dedicated Hardware @urncts.eu
Data Source @SI
Copyright-FreeYescustom
dtaFullYesdefault
folgershakespeareFullNodefault
edhFullNodefault
gps4FullYesdefault
jeanpaulbriefeFullYesdefault
humboldtdigital
Chronology Letter
Full
Chronology Letter
Yesdefault
gwtcFullNodefault
kantFullNodefault
lebensweltenFullYesdefault
openarabicpeFullNodefault
pbc
French English German
Full
.fra. .eng. .deu.
Yesdefault
pcpFullNodefault
textgridFullNodefault
tgapFullNodefault
voth
English German
Full
.eng. .deu.
Nodefault
Default instances use the generic Text Miner. Custom instances may provide corpus specific features. See the source repository for installation and configuration. Subcorpus setups require a file named "urnlist.txt" as provided in the table.
Source Code Resources Source Code and Installation (Git hosted via Bitbucket.org)
Suggested Citation Jochen Tiepmar. Text Service Infrastructure. URL https://urncts.eu, requested on




Cooperation and Service Contracts

Need a custom text infrastructure? I provide CTS/DTS, text mining, corpus tools, APIs and data hosting. Contact me for a service contract.
Adding a TEI/XML corpus to the existing infrastructure is free of charge.




State of Infrastructure

Chronicle

Availability and Software-Versions (versionsoftware.php): 🖌 📚
Database Versions (versiondata.php): 🖌 📚

Central Enquiry

AuthorsLanguagesYearsCopyright
Restricted
List
Interactive List
Chart
List
Interactive List
Chart
List
Interactive List
Chart
List
Interactive List
Chart




FAQ

What is this?

tl;dr: This is a decentralized text data service based on Canonical Text Service, Distributed Text Service and additional text APIs.

Is the implementation feature complete?

CTS: Stable. DTS: Work in Progress. See the Roadmap for planned features.

How do data requests work for texts that are copyright restricted?

Copyright-restricted data can be accessed through temporary, dataset-specific access tokens for authorized users.

Can I host my data on my own server ?

Yes. Data sets are available with database and source via Zenodo and can be deployed on your own server.

Are CTS URNs persistently citable?

CTS URNs are designed for persistent citation. Work-in-progress corpora are explicitly marked. Mutable data are versioned via the namespace.

How reliable is this service?

The software and copyright-free data sets are openly available and can be independently deployed. A future cloning mechanism will further reduce the dependence on individual servers.

Is it finally "A Library of a Billion Words"?




Academic Mentions

Jacob Langeloh (2024) : Unleash the Apparatus? Towards a shared representation of knowledge about connections between primary sources.

Thibault Clérice, Hugh Cayless, Jonathan Robie, Ian W. Scott (2026) : Distributed Text Services 1.0: An API Standard for Publishing and Extending TEI Document Collections.



Legal & Data Protection

Non-commercial academic research and data service.
Impressum and Data Protection