BLOMEGA

The same Gemini interpreter scored 40.3% or 75.2% on hard conversations, depending on the brief

Lab note · 19 September 2026 · BLOMEGA

Abstract technical illustration of two separate clusters of light connected by a translucent relay bridge whose three layered channels thin out from bottom to top, cyan and orange on a dark ground

On 1,090 hard conversational scenarios, Gemini 3.1 Pro acting as an interpreter passed 40.3% of communicative-goal checks with a plain "translate this" instruction and 75.2% with a brief describing the setting, the speakers and the cultural norms of the language pair. Same model, same scenarios, same judge: a 34.9-point swing. The benchmark, from KAIST and published 17 September 2026, also finds every one of 10 interpreter setups losing ground from meaning to social fit, and standard MT metrics blind to most of the loss.

KAIST released a benchmark for interpreter agents, not sentences, on 17 September 2026

17 September 2026. Faiz Ghifari Haznitrama and Alice Oh of the KAIST School of Computing posted Evaluating Communicative Success in Machine-Translated Conversation, arXiv:2609.19885. It targets the product category shipping fastest in localisation right now: an agent that sits between two people who share no language and relays a live conversation.

The setup: 2,812 OpenSubtitles dialogue segments chosen for difficulty across six language pairs among Arabic, Bengali, Indonesian and Korean, each used in both directions for 5,624 scenarios and 12 translation directions, with no English in any pair. Each scenario gets a checklist of yes/no criteria in three layers (semantic, pragmatic, cultural-social), written by Gemini 3.1 Pro with the conversation context in view. A simulated native-language recipient replies to each translation, and a judge scores translation and reply against the checklist. Ten interpreter setups were run, giving 56,240 scored outputs.

Ten interpreters, and all ten lose points from meaning to social fit

Table 1. Single-turn communicative-goal pass rate (%) over 5,624 directed OpenSubtitles scenarios, 12 directions among Arabic, Bengali, Indonesian and Korean. "Strict" zeroes any output that GlotLID says is in the wrong language. L1 semantic, L2 pragmatic, L3 cultural-social. LLM interpreters received the cultural-context brief; the three MT systems received raw source text only. Source: Haznitrama and Oh, arXiv:2609.19885, Figure 3.
InterpreterTypePass rateStrictL1 meaningL2 intent, toneL3 social fitL1 minus L3Source
Gemini 3.1 Profrontier LLM91.28998938711Fig. 3
DeepSeek V4 Profrontier LLM82.98193867617Fig. 3
Qwen3.5 Flashefficient LLM78.57789817217Fig. 3
Gemini 3.1 Flash Liteefficient LLM78.27791797219Fig. 3
GPT-5.4 Miniefficient LLM77.27691806922Fig. 3
DeepSeek V4 Flashefficient LLM75.87491796724Fig. 3
Google TranslateMT system54.35379564336Fig. 3
NLLB-200 3.3BMT system52.15170534426Fig. 3
Tiny Ayasmall LLM464462464022Fig. 3
SeamlessM4T v2MT system24.22443231726Fig. 3
Every interpreter loses ground from meaning to social fit Pass rate (%) by checklist layer. Google Translate keeps 79% of meaning criteria and 43% of social ones. 0 25 50 75 100 L1 semantic L2 pragmatic L3 cultural-social Gemini 3.1 Pro 98 > 93 > 87 DeepSeek V4 Pro 93 > 86 > 76 Google Translate 79 > 56 > 43 SeamlessM4T v2 43 > 23 > 17 grey: 6 other setups
Pass rate by checklist layer for all ten setups. The ordering L1 > L2 > L3 holds for every system. Source: arXiv:2609.19885, Figure 3.

The gap between layers is the paper's central result, and its size tracks capability. Gemini 3.1 Pro loses 11 points from L1 to L3. The four efficient LLMs lose 17 to 24. Google Translate keeps 79% of meaning criteria and 43% of social-fit criteria, a 36-point fall, the steepest in the table. Two systems that finish close overall can reach that total with very different layer profiles: Gemini 3.1 Flash Lite and Qwen3.5 Flash are 0.3 points apart overall and 2 points apart on L1.

Language direction matters in both roles (p < 0.001 for source and for target). Averaged across the ten systems, Bengali is 10 points harder than a system's own mean as a source language and 6 points easier as a target; Arabic is the reverse, 10 points harder as a target. The authors read Bengali as a comprehension problem and Arabic as an adaptation problem.

Conversation does not rescue weak systems. In six-turn scripted conversations the three LLM configurations score 92% to 94% per turn and NLLB and Tiny Aya 45% to 47%, and every system scores lower on conversation-level criteria (cross-turn consistency, goal completion) than on its own turns: 40.7% against 45.0% for NLLB. Giving Gemini 3.1 Flash Lite the translated history improved all six pairs, significantly in three, by 1.3 to 2.4 points.

The checklist catches register and honorific failures that CometKiwi and MQM judges score as perfect

What the benchmark checks that a fidelity metric does not User A Bengali only Interpreter agent + translation brief User B Korean only utterance translation reply LLM judge (Gemini 3.1 Pro) context + source + translation + reply answers each criterion yes / no L1 semantic propositions, entities, quantities, grammar Gemini Pro 98 / Google 79 L2 pragmatic speech act, intent, tone, naturalness Gemini Pro 93 / Google 56 L3 cultural-social register, honorifics, face, relational stance Gemini Pro 87 / Google 43 One case from the paper (Table 11), Bengali to Korean: A senior official's rhetorical jab at a junior officer is rendered with upward-polite Korean endings, as a real question. CometKiwi 0.89 (top quartile). GEMBA-MQM: 0 errors. MQM-APE: 0 errors. Checklist: 4 of 26 criteria met. The meaning is right. The power relationship, and the sarcasm, are gone. None of the three metrics has a slot for that. Checklists average 16.2 criteria per scenario (range 7 to 34).
The three actors, the three checklist layers, and one case where every fidelity metric and the checklist disagree. Source: arXiv:2609.19885, Figure 2, Section 3 and Table 11.

The paper quantifies the disagreement on Gemini 3.1 Pro's 5,624 outputs (Table 10). Of the outputs each metric places in its top quartile, 10.3% (RATE) to 12.2% (MQM-APE) score 0.6 or lower on the checklist; CometKiwi's figure is 10.7%. Of the outputs each metric places in its bottom quartile, 62.0% (CometKiwi) to 69.6% (MQM-APE) still pass 90% or more of their criteria. Within a strong system, the metrics are ranking by something other than whether the conversation worked.

Pooling all ten systems hides this. Across about 56,000 outputs CometKiwi correlates with the checklist at Spearman ρ = 0.321, GEMBA-MQM 0.436, MQM-APE 0.301 and RATE 0.529. The authors show every one of those pooled values exceeds every within-system value for the two strongest interpreters, with bootstrap intervals no wider than ±0.03: the pooled correlation is the metric correctly ranking SeamlessM4T below Gemini, not the metric seeing tone.

On hard scenarios, the brief alone moves Gemini 3.1 Pro by 34.9 points Pass rate (%), four briefs, same scenarios, same judge. 20 40 60 80 Gemini 3.1 Pro 40.3 / 44.9 / 64.1 / 75.2 DeepSeek V4 Pro 39.4 / 43.4 / 53.4 / 61.2 GPT-5.4 Mini 39.8 / 45.1 / 51.0 / 53.7 DeepSeek V4 Flash 35.7 / 39.2 / 46.6 / 50.5 Gemini 3.1 Flash Lite 41.0 / 46.2 / 55.1 / 45.9 Qwen3.5 Flash 32.4 / 34.3 / 44.8 / 41.9 Tiny Aya 25.9 / 28.7 / 28.1 / 30.2 Pooled, 7 setups 36.4 / 40.3 / 49.0 / 51.3 direct + context spec-aware cultural context
Pass rate under four briefs on the same 1,087 to 1,092 hard scenarios per model. The gain is largest for the strongest model; for Gemini 3.1 Flash Lite and Qwen3.5 Flash the specification-aware brief beats the cultural-context one, which the authors partly attribute to the hard subset having been curated from Flash Lite's own failures. Source: arXiv:2609.19885, Table 1 and Table 9.

Pooled over seven setups, scenario context adds about 4 points (36.4 to 40.3), a structured specification-aware brief about 9 more (49.0), and the pair-specific cultural context about 2 more (51.3). The briefs also get longer as they add content, and the authors note the gains are not length-matched against a same-length control.

If you ship an interpreter feature, the brief is part of the product and fidelity scores are not acceptance tests

Write the brief like a localisation kit. The strongest lever in the paper is not the model; it is telling the model who is speaking to whom, in what setting, and what the target culture expects of that relationship. For Gemini 3.1 Pro that was worth 34.9 points on hard cases. For Tiny Aya it was worth 4.3. A small model cannot use a brief a large model can.

Test L3 separately, in every target language. A single pass rate hides where the loss sits. Korean speech levels, Bengali honorific pronouns, Indonesian face-preserving formality and Arabic register with gender agreement are the four mechanisms the paper's failure cases span. Each needs criteria written by someone who knows that language's rules.

Do not accept an interpreter on CometKiwi or an MQM judge alone. Among a strong system's outputs, roughly one in ten of those the metrics rank highest fail the conversation, and roughly two in three of those they rank lowest are fine. The first is deployment risk. The second is a false rejection of a translation that correctly adapted to context.

Dedicated MT systems were tested with a handicap, so do not over-read their rows. Google Translate, NLLB-200 and SeamlessM4T v2 received raw source text only, while LLM interpreters received the full brief. L2 and L3 criteria test tone and honorifics that need exactly that context. The authors flag this confound themselves and did not run a matched condition.

A judgement, marked as one: Gemini 3.1 Pro wrote the checklists, served as the primary judge, and is the top-scoring interpreter. The authors tested for this. Human-judge agreement was 75.4% (κ = 0.403) for Gemini against 69.2% (κ = 0.328) for GPT-5.4 on the same items, and DeepSeek V4 Pro was rejected as a judge after its accuracy fell from 96.6% to 38.6% under polarity inversion. That is good practice and still leaves the top row graded by its own model family. Read Gemini's 91.2 with that in mind; the layer ordering, which holds for every system and every judge, is the sturdier finding.

Check it yourself

# The paper: Figure 3 (page 5), Tables 1-2 (page 7), Tables 9-11 (pages 26-31)
curl -L -o comm.pdf https://arxiv.org/pdf/2609.19885
python3 -c "import pypdf; r=pypdf.PdfReader('comm.pdf'); print(r.pages[25].extract_text()[:2500])"

# Source dialogue: OpenSubtitles (ODC-By 1.0), e.g. the Bengali-Korean slice
#   https://huggingface.co/datasets/Helsinki-NLP/open_subtitles

# The scoring rule, so you can run the same metric on your own interpreter logs:
python3 - <<'EOF'
# p_s = share of a scenario's criteria met; benchmark score = unweighted mean of p_s
scenarios = [
  {"L1": [1,1,1,1], "L2": [1,1,0], "L3": [0,0,1]},   # meaning kept, register lost
  {"L1": [1,1,1],   "L2": [1,1],   "L3": [1,1]},
]
p = [sum(sum(v) for v in s.values()) / sum(len(v) for v in s.values()) for s in scenarios]
print(round(100*sum(p)/len(p), 1))                 # 85.0
for layer in ("L1","L2","L3"):
    print(layer, round(100*sum(sum(s[layer])/len(s[layer]) for s in scenarios)/len(scenarios), 1))
# L1 100.0 / L2 83.3 / L3 66.7
EOF

The paper says its scoring code, derived scenarios and system outputs are released, but gives no repository URL in the arXiv version, and we could not locate one on 19 September 2026. The numbers above can be checked against the PDF. To reproduce the method on your own product, the judge and checklist-generation prompts are printed in Appendix K.

What would prove this wrong

The central claim is that interpreter quality degrades from semantic to cultural-social success for every system, and that the translation brief is a first-order variable. The layer claim is wrong if a replication on at least 1,000 scenarios in at least three non-English pairs, using a judge from a different model family than the top interpreter, finds any of the top three systems scoring higher on cultural-social criteria than on semantic ones. The brief claim is weakened if the same replication, with a non-Gemini judge, finds Gemini 3.1 Pro's gain from direct instruction to cultural-context brief below 15 points, against the 34.9 reported. We will look for either by 30 June 2027.

Sources

  1. Faiz Ghifari Haznitrama and Alice Oh (KAIST), Evaluating Communicative Success in Machine-Translated Conversation, arXiv:2609.19885v1, 17 September 2026. Figure 3 (pass rates by system, language and layer), Figure 4 (multi-turn), Table 1 and Table 9 (brief ablation), Table 2 (judge tests), Section 7.3 (human agreement), Table 10 (metric quartile disagreement and pooled correlations), Table 11 (qualitative cases), Appendix A (context asymmetry).
  2. Pierre Lison and Jörg Tiedemann, OpenSubtitles2016, LREC 2016, aclanthology.org/L16-1147. Dataset mirror: Helsinki-NLP/open_subtitles.
  3. Ricardo Rei et al., CometKiwi, WMT 2022; Tom Kocmi and Christian Federmann, GEMBA-MQM, WMT 2023. The two families of fidelity metric the benchmark is compared against.
  4. Yuto Kayano and Saku Sugawara, Specification-aware machine translation and evaluation for purpose alignment, WMT 2025. The brief the paper's cultural-context brief extends.

Related BLOMEGA guides: When human post-editing wins, by discourse dependency · Live dubbing latency at IWSLT 2026 · Multilingual LLM judges and translationese bias

FAQ

How do you evaluate an AI interpreter or live translation agent?

Score each translated turn against a scenario-specific checklist of yes/no criteria at three layers: semantic (meaning, entities, quantities), pragmatic (speech act, intent, tone) and cultural-social (register, honorifics, face). KAIST's benchmark of 17 September 2026 does this over 5,624 OpenSubtitles scenarios in Arabic, Bengali, Indonesian and Korean, with a judge that also sees the recipient's reply.

Is Google Translate good enough for live conversation?

On this benchmark it passed 54.3% of communicative-goal criteria against 91.2% for Gemini 3.1 Pro, and fell from 79% on meaning criteria to 43% on cultural-social criteria. It received raw source text only, while LLM interpreters received a context brief, so part of that gap is the missing context.

How much does the translation prompt matter for an LLM interpreter?

A lot for strong models. On about 1,090 hard scenarios Gemini 3.1 Pro scored 40.3% with a direct instruction and 75.2% with a brief describing the setting, speakers and the cultural norms of the language pair. Pooled over seven models the gain was 36.4% to 51.3%. Tiny Aya gained only 4.3 points.

Do CometKiwi and MQM judges catch cultural errors in translation?

Mostly not. Among Gemini 3.1 Pro's outputs, 10.3% to 12.2% of those each metric ranked in its top quartile failed the checklist at 0.6 or below, and 62.0% to 69.6% of those ranked in the bottom quartile passed at 0.9 or above. In one Bengali to Korean case CometKiwi gave 0.89 and both MQM judges found zero errors while the checklist passed 4 of 26 criteria.