De verbo vitæ

About Vox Vulgata

The Clementine Vulgate, read aloud in Ecclesiastical Latin, with the Douay-Rheims translation alongside and every word followed as it is spoken.

73books
1,334chapters
35,809verses
89.86hours of audio

Mission

Vox Vulgata, literally “the common voice,” is a play on words. It is the common voice because it is God’s revelation, which has been preached to the ends of the world. It is also the common voice because it is the translation of the Vulgata Clementina, the version of the Latin Bible approved by the Council of Trent.

Vox Vulgata is owned and promoted by the Lingua Sacra Institute, an apostolate aimed at promoting the use and appreciation of Latin throughout the Church and the world.

The texts

The Latin is the Clementine Vulgate of 1592, the edition promulgated after the Council of Trent and long standard in the Roman Church. The text comes from the Clementine Vulgate Project, through the BibleGet digital edition. Its spelling, ligatures, punctuation, and verse order are preserved for display.

The English is the Douay-Rheims Bible in Bishop Challoner’s 1749–1752 revision, taken from the BibleCorps public-domain transcription. Latin supplies the structure of the reader; English is joined verse by verse wherever the two editions have a direct counterpart. Eleven verses remain Latin-only rather than being silently dropped or forced into a false match.

How the audio was made

The readers are synthetic. Br. Jerome and Br. Benedict are generated voices, not recordings of human readers, and the names are ours rather than anyone’s. Each chapter was generated with Meta’s facebook/mms-tts-lat model after the displayed Latin was normalized into an Ecclesiastical Latin pronunciation layer. The original text remains unchanged on the page, and every reader speaks the same words with the same timings.

  1. Synthesis. Chapters were rendered in passage-sized chunks, with punctuation translated into pause cues the model could follow.
  2. Rhythm. Those cues were rebuilt as exact digital silence, and unintended hesitations before difficult words were gently shortened.
  3. One timbre. ChatterboxVC retargeted the chunks to a single reference voice so a long chapter would not wander between different speaker characters. Each reader has its own frozen reference, which is the only thing that differs between them.
  4. Delivery. The finished chapters were encoded as Opus at 24 kHz, then decoded and checked again for clipping before release.

Why the words stay in step

The highlighting is not guessed by listening back with a forced aligner. Word durations come from the speech model’s own duration predictor—the same timing plan used to create the sound. Whenever pauses, compression, or voice conversion changed the length of a passage, the timing map was carried through the same transformation. More than 612,000 word cues were verified against the Latin displayed in the reader.

Provenance and use

Both source texts are in the public domain. The generated audio was made with Meta’s facebook/mms-tts-lat model, licensed under CC BY-NC 4.0. Vox Vulgata adapted the output through Ecclesiastical Latin pronunciation processing, pause reconstruction, compression, voice conversion, timing transformation, and Opus encoding. No endorsement by Meta is implied.

Vox Vulgata is part of the Lingua Sacra Institute. Donations are entirely optional and help cover the website’s hosting costs while supporting the Institute’s broader apostolate.

Contact

Please direct all bug reports, recommendations, questions, and other feedback to [email protected].

Analytics and privacy

Vox Vulgata uses Cloudflare Web Analytics and a first-party Lingua Sacra event collector to understand aggregate use. Controlled events may include the page viewed, book and chapter identifiers, genuine active listening time, completion, player controls, reference-search success, offline downloads, installation, donation interest, and categorized application errors.

Vox Vulgata does not send raw reference-search text, names, email addresses, user IDs, or intentionally collected personal information in these events. Visitor profiles, analytics identifiers, advertising analytics, and session replay are not used. Events are not queued for later replay while the device is offline.

Raw custom events are retained for up to 3 months. Optional aggregate analytics is on by default and can be turned off at any time. The preference is shared across Lingua Sacra applications in the visitor’s browser.

Cloudflare still processes ordinary requests needed to deliver and protect the websites even when optional analytics is off. Privacy questions may be sent to [email protected].