Skip to main content

The Writer Sets the Style, the Editor Catches the Errors

Chad Jones /

CTAI stands for Committee Translation AI. The idea behind the name is that no single model should translate the Bahá’í Writings alone. A committee does: a concordance that remembers Shoghi Effendi’s choices, a fast decision system, a writer, an editor, and in the end a human reader.

My working assumption was simple. The best translation should come from giving the committee everything at once: good metadata about the work, the surrounding paragraphs, a full Jafar report with examples and a whole model paragraph, a strong writer, and a second strong model editing. It seemed obvious. So we tested it.

How we tested the committee: everything, then minus one

The design is an ablation. Start with every ingredient switched on. Then take out one ingredient at a time and see what gets worse. If removing something changes nothing, it wasn’t pulling its weight. If removing it makes things worse, now we know how much it was worth, measured in the presence of everything else.

The full setup, which I’ll call “everything”, looked like this:

Committee memberRole in “everything”
Document fieldsWork, author, recipient and date where the Phelps inventory has them, genre, named figures, Qur’ánic allusions
Neighbouring paragraphsThe source paragraphs just before and after, so the writer sees where a thought comes from and where it goes
The Jafar reportShoghi Effendi’s renderings of the phrases and words in the passage, relevant examples, and one or two whole paragraphs he translated, given to both writer and editor
The writerClaude Opus 5.5 at medium effort
The editorGPT-6.1-sol from a different model family, returning exact-text corrections that code applies

Then came the variants: everything minus the document fields, minus the neighbours, minus the Jafar report, minus the examples (keeping only the compact report), minus the editor, and minus the editor’s evidence. We also swapped the roles (GPT writes, Opus edits), tried a lighter writer with everything else in place, and ran today’s translation API and an earlier lighter setup for comparison.

The test passages: untranslated Arabic and Persian

All 24 test passages are genuinely untranslated, taken from SifterSearch’s queue of originals that have no English anywhere, so no model could have seen a published answer.

  • 12 Arabic and 12 Persian
  • Prayers, epistles, exposition and commentary, and Qur’ánic-style writing (much of it the Báb’s)
  • 15 of the 24 have no punctuation at all
  • Three are pairs of consecutive paragraphs, to test continuity

Every setup translated all 24.

Scoring accuracy and style separately

A translation can be faithful and flat, or beautiful and wrong. Those are different failures, and a reader may care about one more than the other. So we scored them separately, in separate judging passes with separate instructions, so neither judgement could colour the other.

Accuracy, judged by reading the source:

  1. Overall meaning, 1 to 10
  2. Omissions
  3. Additions
  4. Reversals and negation errors
  5. Misread referents and relations
  6. Wrong sense of a word or phrase
  7. Unresolved ambiguity

Style, judged against two genuine Shoghi Effendi paragraphs of the same genre:

  1. Cadence
  2. Register
  3. Terminology as Shoghi Effendi uses it
  4. Rhetorical structure

Two judges from different model families scored every translation, and we averaged them. Before trusting them, we checked both on passages with known planted errors and on clean text.

What each ingredient contributes

Differences are measured against “everything”. For accuracy and style, higher is better; for errors, lower is better.

SetupMeaning (/10)Errors per passageMajor errors per passageStyle (/10)Est. cost per 1,000 English words
Everything8.901.100.197.97$0.41
− document fields−0.121.150.23−0.08$0.40
− neighbouring paragraphs−0.081.290.21−0.05$0.33
− Jafar report+0.231.120.08−0.10$0.22
− examples (compact report only)−0.121.210.21−0.16$0.26
− editor (writer only)−0.291.420.31+0.01$0.26
− evidence for the editor+0.170.790.15−0.06$0.31
Roles swapped (GPT writes, Opus edits)−0.291.250.31−0.67$0.39
Lighter writer + everything else−0.501.460.40−0.57$0.31
Today’s translation API−0.441.900.42−0.38$0.11
Earlier lighter setup−0.902.150.50−0.68$0.04

Bold differences are the ones whose 95% range excludes zero. With 24 passages, anything smaller than about 0.3 is noise. Costs are estimates from the measured token usage of each setup at the providers’ published prices (September 2026), with prompt caching on.

The writer sets the style

The clearest result in the study: style belongs to the writer. Every setup where Opus wrote the first draft scored between 7.8 and 8.0 on style. Every setup where GPT wrote the draft scored 7.3 to 7.4, even after Opus edited it. An editor fixes sentences; it doesn’t change the voice they were written in.

That also explains why the lighter writer lost on both counts. Starting from a weaker draft costs accuracy that the editor only partly recovers, and style that it doesn’t recover at all.

A second model family buys accuracy

The editor buys accuracy, not style. Taking the editor away raised the error count by about a third, and major errors by a similar share. It cost nothing in style, because the editor makes targeted corrections rather than rewriting.

The order matters. Opus writing and GPT editing beat the reverse on both accuracy and style. A different family for the editor is part of the point: a model is poor at catching its own habits, and a second family reads the passage with different blind spots.

The Jafar report didn’t help on these texts

This was the surprise, and the most useful lesson. Removing the Jafar report didn’t hurt. Accuracy came out slightly higher without it, with the fewest major errors of any setup. Style was the same. The editor also did slightly better without seeing it.

One signal pointed the other way. Our offline voice classifier, which measures closeness to Shoghi Effendi’s English vocabulary, rated the English lower without the report. The report does push his vocabulary into the translation; the judges just didn’t reward it on these passages.

The likely reason is the texts themselves. Much of the queue is the Báb’s Qur’ánic-style Arabic, the Persian Bayán and similar works. Shoghi Effendi translated very little of this kind of writing, so many of the report’s phrase matches and examples come from passages that don’t really fit. Evidence that almost fits can mislead a writer as easily as help one.

That is a design lesson for Jafar, not a reason to drop it. The report needs to know when its evidence is close to the passage at hand and when it isn’t, and to stay quiet when it has nothing relevant to say.

Metadata and neighbours: small, steady gains

Document fields and neighbouring paragraphs both helped a little on accuracy and style. Neither effect is large enough to separate from noise with 24 passages, but both point the same way, and neither adds much to the prompt. We’re keeping them.

What the committee will do differently

The recommended committee for untranslated works:

  1. A strong writer drafts the translation, with the document fields and the neighbouring paragraphs.
  2. An editor from a different model family reads the source and the draft and returns exact corrections, which code checks and applies.
  3. The Jafar report steps back for works far from Shoghi Effendi’s own translations, until it can judge its own relevance.
  4. A human reader gets the final say on the most important works.

On these 24 passages this setup scored 9.12 on meaning, with fewer major errors than any other setup, and 7.87 on style. That’s within noise of the best style score. Today’s translation API scored 8.46 and 7.59.

These are estimates of what translating would cost with each setup, from the usage we measured, not what this study cost to run.

SetupEst. cost per 1,000 English wordsA 10,000-word workSifterSearch’s untranslated queue (about 10.5 million characters)
Recommended committee (strong writer + fields + neighbours, editor from another family)$0.22about $2.20about $740, or about $370 with batch pricing
Everything, including the Jafar report$0.41about $4.10about $1,380, or about $690 with batch pricing
Today’s translation API$0.11about $1.10about $400, or about $200 with batch pricing

Batch pricing halves the cost for work that can wait a few hours, which suits a library-scale queue. Dropping the Jafar report on these texts nearly halves the cost of “everything” while scoring as well or better, so the recommended committee is both the most accurate setup and one of the lighter ones.

Limits of this study

  • The judges are models. They were checked against known errors before use, but they aren’t people. One judge also flagged too many clean passages as having major errors, so its major-error counts run high.
  • Own-family bias. Each judge showed a hint of preferring output from its own model family. Averaging the two judges offsets this, but doesn’t remove it.
  • Small sample. 24 passages is enough to see the large effects (the writer, the editor, the role order) but not the small ones.
  • No human ratings yet. Human ratings of a subset are in progress, with meaning and style rated separately.

Next steps for the translation committee

Three things come out of this for the next round.

  1. An agreed measurement plan up front. We should set 5 to 10 named measures with the people who will use the results before the experiment runs, not after. This study came close, with about ten, but they weren’t agreed in advance.
  2. Human ratings. Model judges tell us where to look. A source-competent human reader tells us what’s true.
  3. Better Jafar evidence for texts outside Shoghi Effendi’s corpus. The report was built around the Writings he translated. The untranslated queue is mostly other works, so the report has to learn when its evidence is relevant and when it isn’t.

Further reading on CTAI’s translation research