Skip to main content

Measuring Translation Against Shoghi Effendi

Chad Jones /

We’re learning how to measure the quality of AI translations of the Bahá’í Writings, and how to use those measurements to improve them. Our standard is Shoghi Effendi. His translations, aligned word by word with their Arabic and Persian originals, record thousands of his choices across many kinds of writing.

This post explains how CTAI’s translation pipeline works, how we test it, and what each stage adds. It’s a work in progress.

The pipeline

Translation happens in five stages:

  1. Style classification. Which of six modes is this paragraph written in?
  2. Relevant examples. How did Shoghi Effendi translate similar paragraphs, phrases and words?
  3. Style and book glossary. How did he render key terms in this kind of writing, and how has this book rendered them so far?
  4. Translation committee. Claude Opus writes the translation.
  5. Editorial pass. A model from a different family checks it against the original and corrects it.

Stages 1 to 3 make up the Jafar Report. Jafar, our AI research assistant, assembles it from our concordance of Shoghi Effendi’s translations, which are aligned word by word with their originals.

1. Style classification

Shoghi Effendi didn’t translate everything in one voice. A prayer reads differently from a commentary or a proclamation, so the first stage is to recognise what kind of writing a paragraph is.

We use six modes: the five the Báb Himself named (verses, prayers, commentaries, learned discourse and Persian writings) and the Qur’anic. Each mode has finer styles within it, such as praise, supplication, exegesis and narrative.

A small, fast model labels each paragraph and says how confident it is. (This is System-1, our layer for quick decisions, which runs on Jev.) A whole book takes the style most of its paragraphs share, so it keeps one voice from start to finish.

2. Relevant examples

Jafar then gathers examples of how Shoghi Effendi handled similar material:

  • Paragraphs. A few of his paragraphs in the same style, to set the tone.
  • Phrases. His renderings of each phrase in the paragraph.
  • Words. His renderings of the words left over. Where a word has several meanings, System-1 picks the one this passage needs.

Examples are ranked by meaning first, then by how close their style is to this paragraph’s. For Qur’anic passages, Jafar also draws on every verse of the Qur’an that Shoghi Effendi translated: 283 passages in all.

3. Style and book glossary

For each style, a glossary records how Shoghi Effendi rendered key terms in that kind of writing. Each book adds its own layer of names, recurring phrases and settled choices, so a term reads the same way from the first page to the last.

4. Translation committee

Claude Opus writes the translation. It gets the original, the Jafar Report, the neighbouring paragraphs and the document’s context: its title, author, recipient, and the figures and Qur’anic allusions it involves. Its instructions follow Shoghi Effendi’s principles, such as choosing elevated words for sacred things: “the Crimson Ark”, not “the red boat”.

5. Editorial pass

A GPT model reads the original, the draft and the same report, and returns exact corrections. Code applies a correction only if the text it replaces appears exactly once in the draft. A model from another family catches mistakes the writer’s family misses.

The writer and editor also attach notes where a reading is uncertain, a name is veiled or a phrase alludes to the Qur’an. Automatic checks and an accuracy score go with every paragraph.

How we test

We test only on Arabic and Persian passages that have never been translated, so no model can recall an English version. Accuracy and closeness to Shoghi Effendi are scored separately, because a translation can be faithful without sounding like him, or sound like him and get the meaning wrong.

ScoreThe questionHow it’s measured
AccuracyIs the meaning right?Judges from two model families read the original and count errors: omissions, additions, reversed negations, misread references, wrong word senses
PrecedentDoes it use his renderings?How often it follows his renderings of the same words and phrases
GrammarAre its sentences shaped like his?Compared with his English in the same style
Closeness judgeWould he have rendered it this way?A judge reads it beside his renderings and two of his paragraphs in the same style

Closeness is always measured within the passage’s own style, since he rendered a word one way in a prayer and another in an argument.

Before relying on the closeness score, we checked that it can recognise Shoghi Effendi. We took passages he did translate, and scored his translation and another translation of the same passage side by side. A good score should pick his every time. Ours usually does:

Recognising Shoghi Effendi: how often his translation scored higher
His vs Claude Opus (37 passages)
78%
His vs DeepSeek (37 passages)
84%
His vs the early "red boat" translations (8 passages)
100%

It isn’t perfect, so we use it as a guide, not a verdict. To find out what one stage adds, we run the pipeline with and without it, keeping everything else the same.

What each stage adds

Style classification: how often System-1 was right, on 30 passages we labelled by hand
Mode
93%
Finer style
70%
The Jafar Report: each part switched off and on
WithoutWith
The report: following his phrase renderings
0.55
0.77
The report: closeness judge
0.83
0.89
Style-ranked examples: fit both meaning and style
37%
47%
Glossary: recurring terms rendered the same way
50%
58%

The glossary made terms more consistent without costing any accuracy.

The editor, together with fuller context for the writer, tested on 24 untranslated passages:

Writer aloneWriter and editor
Meaning (1–10)8.508.94
Passages with no major error58%75%
Closeness judge0.840.88

Accuracy improved most: three in four passages now come through with no major error, up from three in five. Closeness held, and improved on the judge. The editorial pass raises the cost from about $0.10 to $0.17 per 1,000 characters of source text.

Not everything helped. Explaining Shoghi Effendi’s balance of literal and interpretive rendering to the writer made no measurable difference; showing it his actual renderings did. We’ve kept the explanation, since it does no harm and future models may use it. The same went for asking the writer to reason as a committee of three experts in Shoghi Effendi’s manner (a scholar of Shakespeare and the King James Bible, an expert in Persian poetry, and an expert in Islamic philosophy and literature): on 24 untranslated passages, accuracy and closeness came out the same.

How far there is to go

Interpretive (rather than literal) renderings, with Shoghi Effendi as the goal
Shoghi Effendi (the goal)
43%
CTAI
31% · 72% of his
  • Shoghi Effendi interpreted about 43% of his renderings; we interpret about 31%. That’s roughly 72% of the way to his balance.
  • His precedents are thinnest for the Báb’s own styles, where most of the untranslated material is.
  • Every number here comes from model judges and small samples.

Next we’ll add blind ratings from readers who know the source languages, to check the model judges, and build a larger fixed set of test passages.

Everything the pipeline produces is labelled pre-provisional: a researched draft for a scholar to start from, not a finished translation.

Further reading: The Writer Sets the Style, the Editor Catches the Errors · Phrases, Not Words