We’re learning how to measure the quality of AI translations of the Bahá’í Writings, and how to use those measurements to improve them. Our standard is Shoghi Effendi. His translations, aligned word by word with their Arabic and Persian originals, record thousands of his choices across many kinds of writing.
This post explains how CTAI’s translation pipeline works, how we test it, and what each stage adds. It’s a work in progress.
The pipeline
Translation happens in five stages:
- Style classification. Which of six modes is this paragraph written in?
- Relevant examples. How did Shoghi Effendi translate similar paragraphs, phrases and words?
- Style and book glossary. How did he render key terms in this kind of writing, and how has this book rendered them so far?
- Translation committee. Claude Opus writes the translation.
- Editorial pass. A model from a different family checks it against the original and corrects it.
Stages 1 to 3 make up the Jafar Report. Jafar, our AI research assistant, assembles it from our concordance of Shoghi Effendi’s translations, which are aligned word by word with their originals.
1. Style classification
Shoghi Effendi didn’t translate everything in one voice. A prayer reads differently from a commentary or a proclamation, so the first stage is to recognise what kind of writing a paragraph is.
We use six modes: the five the Báb Himself named (verses, prayers, commentaries, learned discourse and Persian writings) and the Qur’anic. Each mode has finer styles within it, such as praise, supplication, exegesis and narrative.
A small, fast model labels each paragraph and says how confident it is. (This is System-1, our layer for quick decisions, which runs on Jev.) A whole book takes the style most of its paragraphs share, so it keeps one voice from start to finish.
2. Relevant examples
Jafar then gathers examples of how Shoghi Effendi handled similar material:
- Paragraphs. A few of his paragraphs in the same style, to set the tone.
- Phrases. His renderings of each phrase in the paragraph.
- Words. His renderings of the words left over. Where a word has several meanings, System-1 picks the one this passage needs.
Examples are ranked by meaning first, then by how close their style is to this paragraph’s. For Qur’anic passages, Jafar also draws on every verse of the Qur’an that Shoghi Effendi translated: 283 passages in all.
3. Style and book glossary
For each style, a glossary records how Shoghi Effendi rendered key terms in that kind of writing. Each book adds its own layer of names, recurring phrases and settled choices, so a term reads the same way from the first page to the last.
4. Translation committee
Claude Opus writes the translation. It gets the original, the Jafar Report, the neighbouring paragraphs and the document’s context: its title, author, recipient, and the figures and Qur’anic allusions it involves. Its instructions follow Shoghi Effendi’s principles, such as choosing elevated words for sacred things: “the Crimson Ark”, not “the red boat”.
5. Editorial pass
A GPT model reads the original, the draft and the same report, and returns exact corrections. Code applies a correction only if the text it replaces appears exactly once in the draft. A model from another family catches mistakes the writer’s family misses.
The writer and editor also attach notes where a reading is uncertain, a name is veiled or a phrase alludes to the Qur’an. Automatic checks and an accuracy score go with every paragraph.
How we test
We test only on Arabic and Persian passages that have never been translated, so no model can recall an English version. Accuracy and closeness to Shoghi Effendi are scored separately, because a translation can be faithful without sounding like him, or sound like him and get the meaning wrong.
| Score | The question | How it’s measured |
|---|---|---|
| Accuracy | Is the meaning right? | Judges from two model families read the original and count errors: omissions, additions, reversed negations, misread references, wrong word senses |
| Precedent | Does it use his renderings? | How often it follows his renderings of the same words and phrases |
| Grammar | Are its sentences shaped like his? | Compared with his English in the same style |
| Closeness judge | Would he have rendered it this way? | A judge reads it beside his renderings and two of his paragraphs in the same style |
Closeness is always measured within the passage’s own style, since he rendered a word one way in a prayer and another in an argument.
Before relying on the closeness score, we checked that it can recognise Shoghi Effendi. We took passages he did translate, and scored his translation and another translation of the same passage side by side. A good score should pick his every time. Ours usually does:
It isn’t perfect, so we use it as a guide, not a verdict. To find out what one stage adds, we run the pipeline with and without it, keeping everything else the same.
What each stage adds
The glossary made terms more consistent without costing any accuracy.
The editor, together with fuller context for the writer, tested on 24 untranslated passages:
| Writer alone | Writer and editor | |
|---|---|---|
| Meaning (1–10) | 8.50 | 8.94 |
| Passages with no major error | 58% | 75% |
| Closeness judge | 0.84 | 0.88 |
Accuracy improved most: three in four passages now come through with no major error, up from three in five. Closeness held, and improved on the judge. The editorial pass raises the cost from about $0.10 to $0.17 per 1,000 characters of source text.
Not everything helped. Explaining Shoghi Effendi’s balance of literal and interpretive rendering to the writer made no measurable difference; showing it his actual renderings did. We’ve kept the explanation, since it does no harm and future models may use it. The same went for asking the writer to reason as a committee of three experts in Shoghi Effendi’s manner (a scholar of Shakespeare and the King James Bible, an expert in Persian poetry, and an expert in Islamic philosophy and literature): on 24 untranslated passages, accuracy and closeness came out the same.
How far there is to go
- Shoghi Effendi interpreted about 43% of his renderings; we interpret about 31%. That’s roughly 72% of the way to his balance.
- His precedents are thinnest for the Báb’s own styles, where most of the untranslated material is.
- Every number here comes from model judges and small samples.
Next we’ll add blind ratings from readers who know the source languages, to check the model judges, and build a larger fixed set of test passages.
Everything the pipeline produces is labelled pre-provisional: a researched draft for a scholar to start from, not a finished translation.
Further reading: The Writer Sets the Style, the Editor Catches the Errors · Phrases, Not Words