The first-placed system in IWSLT 2026's 2-to-4 second latency class ran at 52.7 seconds on YouTube audio

IWSLT 2026 sorted simultaneous speech translation systems into a low-latency class (0 to 2 seconds) and a high-latency class (2 to 4 seconds) using their development-set numbers. On the English-to-Chinese test sets, the system that placed first in the high-latency class measured 28.3 seconds of LongYAAL on ACL talks, 36.3 on Bloomberg business news and 52.7 on YouTube-sourced YODAS audio. The regime label held on the development set and broke everywhere else. If you are buying live dubbing this quarter, the number in the contract should be measured on your audio, not on a conference recording.
Two things landed in the same month, and only one of them published latencies
11 September 2026, Amsterdam. At IBC 2026, CAMB.AI announced a Streaming SDK and Streaming Dashboard for live multilingual dubbing and subtitles, integrated into NVIDIA's Holoscan for Media reference architecture, built on its MARS and BOLI models, supporting 150 or more languages, with Ligue 1+, NASCAR and Eurovision Sport named as users and a claim that a broadcaster can start multilingual output in as little as three lines of code (Sports Video Group, TV Tech). In the same week NVIDIA described NDI using its LipSync NIM and Active Speaker Detection NIM microservices for real-time translation and lip-synced dubbing inside existing broadcast workflows. Neither announcement carries an end-to-end latency figure.
The published latency number on CAMB.AI's own site is a different measurement: its low-latency streaming post puts MARS8-Flash at roughly 100 milliseconds time-to-first-byte. That is how fast the speech synthesiser starts talking once it has been handed text. It says nothing about how long the translation policy waited before handing over that text.
July 2026, San Diego. The 23rd IWSLT evaluation campaign did measure that wait, across 10 shared tasks and more than 30 teams, and published every number. The findings paper is Speech Translation and Metrics in 2026, 87 pages, 60 authors, open access. Its Simultaneous track ran on raw unsegmented audio in four directions: English into German, Chinese and Italian, and Czech into English. Quality is XCOMET-XL, latency is LongYAAL, both computed by OmniSTEval on re-segmented output. Ten systems from eight teams submitted.
The latency class was assigned on the development set and did not survive the test sets
Participants submitted development-set logs, the organisers computed LongYAAL, and that number decided which class a system competed in. Table 1 below takes the first-placed English-to-Chinese system in the high-latency class and follows it across every evaluation set, then shows what the systems it beat were doing at the same time.
| Evaluation set | Pl. | System | COMET | BLEU | chrF | LongYAAL (s) | Source |
|---|---|---|---|---|---|---|---|
| MCIF dev | 1 | NEMO | 0.84 | 47.48 | 41.01 | 15.4 | Table 36 |
| ACL test | 1 | NEMO | 0.83 | 47.64 | 40.30 | 28.3 | Table 36 |
| Bloomberg test | 1 | NEMO | 0.75 | 30.46 | 27.05 | 36.3 | Table 36 |
| YODAS test | 1 | NEMO | 0.68 | 22.50 | 21.06 | 52.7 | Table 36 |
| ACL test | 2 | MLLP-VRAIN UPV | 0.82 | 50.03 | 42.75 | 4.8 [5.2] | Table 36 |
| ACL test | 3 | CPII-HK | 0.78 | 45.04 | 40.46 | 3.0 | Table 36 |
| ACL test | 5 | CUHKSZ | 0.74 | 42.57 | 35.82 | 2.3 | Table 36 |
| YODAS test | 2 | Baseline | 0.65 | 22.57 | 21.25 | 1.0 [1.2] | Table 36 |
| ACL test (En-De) | 1 | NEMO | 0.93 | 45.01 | 71.44 | 4.7 | Table 35 |
| YODAS test (En-De) | 1 | NEMO | 0.79 | 29.57 | 56.73 | 6.0 | Table 35 |
Read the last two rows first. The same team, the same campaign, English into German: 4.7 seconds on ACL talks, 6.0 on YODAS. Slow for a 2-to-4 second class, but recognisable. The blowup is specific to English into Chinese.
Look at the YODAS row for the Baseline: COMET 0.65 at 1.0 second, against the first-placed system's 0.68 at 52.7 seconds. Three hundredths of COMET for fifty-one seconds. On a live stream that is not a tradeoff, it is a different product.
The findings paper acknowledges the pattern in one sentence: "Due to the aforementioned latency measurement issues, several systems exhibit higher latencies on the test sets than their development set metrics initially suggested at the time of submission, explaining the unusually large latency values reported in the results tables." The word "aforementioned" is the paper's only occurrence of it in that section, and the Simultaneous track write-up does not describe the issues it refers back to. So the published numbers stand without an explanation attached to them.
A 100 millisecond time-to-first-byte measures the last box in the chain
A live dubbing pipeline has at least five places where seconds accumulate, and the vendor metrics in circulation cover one of them. The diagram below puts the IWSLT 2026 measurements next to the marketing measurement, on the same timeline.
Two of the campaign's smaller results explain why the orange span is hard to shrink honestly.
The smallest model paid the largest compute penalty. CUNI-POCKET was the only end-to-end submission, a single 1B-parameter Canary model aimed at edge deployment. It showed the biggest gap between computation-aware and non-computation-aware latency, more than 0.4 seconds, against a campaign average of about 0.3. The findings paper flags this as running against the expectation that smaller models incur lower overhead. Model size is not a proxy for responsiveness.
Extra context helps quality and costs nothing in latency, if you integrate it properly. The 2026 edition added a track where systems could read the source paper PDF for each ACL talk. MLLP-VRAIN UPV gained an average of 2.75 COMET across three language directions and both latency regimes. NEMO gained roughly nothing over its context-free run. CUHKSZ went slightly backwards on English to Chinese. Same input, three outcomes, which puts the value in the integration rather than the context.
What a localization buyer should put in the contract
Ask for the metric, not the adjective. "Real-time" is not a number. LongYAAL, StreamLAAL and LongDAL are numbers, they are defined in public, and OmniSTEval computes them. Ask a vendor which one they report, on which audio, and whether it is computation-aware. The IWSLT 2026 tables print the computation-aware value in brackets next to the unaware one precisely because the two differ.
Require the measurement on your own content. The 1.0 second YODAS penalty is the whole argument. A system tuned on conference talks met its declared class on conference talks. The same system on YouTube-style audio was slower and worse: English to Chinese COMET fell from 0.83 to 0.68 and BLEU from 47.64 to 22.50. Sports commentary, reality formats and user-generated video are further from ACL talks than YODAS is.
Treat quality metrics as disagreeing witnesses. On English to Chinese, MLLP-VRAIN UPV beat the higher-ranked NEMO submission by an average of 1.6 chrF and 1.7 BLEU while scoring 0.067 lower on COMET. If your acceptance test uses one metric and your vendor optimised another, you will argue about a real disagreement rather than a rounding error.
Budget the latency you can hide. A pre-recorded stream can absorb a 30 second delay behind a broadcast delay buffer. A live sports call cannot: the crowd noise arrives 30 seconds before the commentary that explains it. That difference, not model quality, is what decides whether the current state of simultaneous translation is usable for a given format. This is a judgement, not a measurement.
Nine of ten systems were cascades. The findings paper states that with the exception of a single end-to-end submission, all participating teams adopted cascaded approaches: separate ASR, then MT, then emission policy. Vendors selling a single end-to-end model as inherently faster are selling against the way the field's own best systems are actually built in 2026.
Check it yourself
Every number above is in one open-access PDF and one open-source toolkit.
# 1. the findings paper (87 pages, open access, no paywall)
curl -sL -o iwslt2026.pdf https://aclanthology.org/2026.iwslt-1.39.pdf
python3 -c "
import pypdf
t='\n'.join(p.extract_text() or '' for p in pypdf.PdfReader('iwslt2026.pdf').pages)
open('iwslt2026.txt','w').write(t)
i=t.find('Table 36')
print(t[i-6000:i+200])" # Table 36 = English to Chinese
# Table 35 = English to German, Table 34 = Czech to English, Table 37 = English to Italian
# 2. the latency metrics, as code
git clone https://github.com/pe-trik/OmniSTEval
# LongYAAL is the primary metric; StreamLAAL, LongLAAL and LongDAL are also reported
# 3. the audio that produced the 1.0 second penalty
# https://huggingface.co/datasets/espnet/yodas (partition en003 is held out from training)
# 4. reproduce the regime-versus-reality gap in the table above
python3 - <<'PY'
# NEMO, English to Chinese, high-latency regime, LongYAAL seconds (Table 36)
sets = {"MCIF dev":15.4, "ACL test":28.3, "Bloomberg test":36.3, "YODAS test":52.7}
comet = {"MCIF dev":0.84, "ACL test":0.83, "Bloomberg test":0.75, "YODAS test":0.68}
for k,v in sets.items():
print(f"{k:16s} {v:6.1f} s COMET {comet[k]:.2f} over declared 4 s ceiling by {v-4:.1f} s")
PY
# MCIF dev 15.4 s COMET 0.84 over declared 4 s ceiling by 11.4 s
# ACL test 28.3 s COMET 0.83 over declared 4 s ceiling by 24.3 s
# Bloomberg test 36.3 s COMET 0.75 over declared 4 s ceiling by 32.3 s
# YODAS test 52.7 s COMET 0.68 over declared 4 s ceiling by 48.7 s
To test a vendor, hand over 30 minutes of your own audio, ask for the output log with per-token emission timestamps in SimulStream or SimulEval JSONL format, and run OmniSTEval yourself. A vendor who cannot produce emission timestamps cannot produce a latency number either.
What would prove this wrong
The claim under test is that the quality-latency frontier for unsegmented live speech translation sits well above one second, and that out-of-domain audio moves it further. It is wrong if, by 31 July 2027, a system in the IWSLT 2027 Simultaneous track reports a computation-aware LongYAAL at or under 1.0 second on the YODAS test set while matching the 2026 best COMET on that set, which was 0.68 for English to Chinese and 0.79 for English to German. The closest 2026 result was the organisers' baseline at 1.0 second and 0.65 COMET, which trades 0.03 COMET for the speed and is therefore already close on one axis and not the other.
A second prediction, marked as judgement: no live dubbing vendor will publish a LongYAAL or StreamLAAL figure on customer audio before 31 July 2027. Time-to-first-byte is the metric being marketed because it is the flattering one. If a vendor publishes an end-to-end emission-timestamp latency on non-ACL audio before that date, this reading was too cynical.
FAQ
How much latency does live AI dubbing actually add?
On the IWSLT 2026 test sets, English to German ran between 2.0 and 7.7 seconds of non-computation-aware LongYAAL for the systems that behaved, and the first-placed English to Chinese system measured 28.3 seconds on ACL talks, 36.3 on Bloomberg news and 52.7 on YODAS. Those are text latencies. Synthesis, packaging and delivery sit on top.
Does a 100 millisecond time-to-first-byte mean sub-second live dubbing?
No. Time-to-first-byte measures how quickly the speech synthesiser starts emitting audio once it has text. It excludes the wait the read/write policy imposes before committing that text, which is what LongYAAL measures. In 2026 that wait ran from 1.5 seconds to over 50.
Why is YouTube audio harder than conference talks?
The findings paper reports that YODAS consistently induced the highest latencies across all language pairs, about 1.0 second higher than the other sets, and attributes it to domain mismatch and distinct acoustic characteristics. Quality fell alongside: best English to Chinese COMET dropped from 0.83 on ACL talks to 0.68 on YODAS.
Are cascaded or end-to-end systems better for simultaneous speech translation?
Cascaded, in practice. All 2026 submissions except one were cascades. The single end-to-end system, a 1B-parameter Canary model, showed the largest computation-aware penalty, over 0.4 seconds against a campaign average of about 0.3.
Sources
- IWSLT 2026 organisers (60 authors), Speech Translation and Metrics in 2026: Findings of the IWSLT Campaign, Proceedings of the 23rd International Conference on Spoken Language Translation, San Diego, July 2026. Track V (Simultaneous): latency regimes, metrics, YODAS latency spike, computation-aware gap, extra-context results; Tables 34 to 37 (per-direction results). PDF.
- Zeyu Yang and Satoshi Nakamura, CUHKSZ Simultaneous Speech Translation System for IWSLT 2026, IWSLT 2026. Qwen3-Omni-30B-A3B backbone with a learned wait token.
- Javier Iranzo-Sanchez et al., MLLP-VRAIN UPV System for the IWSLT 2026 Simultaneous Speech Translation Task, IWSLT 2026. Parakeet plus Qwen 3.5 cascade with a relaxed longest-common-prefix policy.
- Sports Video Group, IBC 2026: CAMB.AI Unveils the World's First Multilingual Broadcasting Agent Powered by NVIDIA, 11 September 2026, and TV Tech, CAMB.AI Unveils Multilingual Broadcasting AI Agent at IBC2026, 11 September 2026. Streaming SDK and Dashboard, MARS and BOLI, 150+ languages, Ligue 1+, NASCAR, Eurovision Sport.
- NVIDIA, NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC, September 2026. LipSync NIM and Active Speaker Detection NIM in NDI workflows.
- CAMB.AI, Lowest-Latency TTS for Live Sports Streaming. MARS8-Flash at roughly 100 milliseconds time-to-first-byte.
- OmniSTEval (latency and quality toolkit used by the campaign) and espnet/yodas (the YouTube-sourced evaluation audio).
Related BLOMEGA guides: Prime Video's visual dubbing · The cost of a dubbed minute in 2026 · EU AI Act Article 50 and AI dubbing disclosure.