Data Analytics report
Do LLM Agents Work Equally Well Across Languages?
A visual evidence review of 30 languages, 16 projects, and 15 charts; current frontier models separated from earlier flagships and test resources.
Executive Summary
No. Published evaluations show that the same LLM agent can perform materially differently across languages. The practical conclusion is to evaluate model × language × task × localization setting. English performance alone does not establish reliability in another language. The evidence supports unequal performance; it does not support a permanent ranking of languages.
- Arabic: repeated gaps in research and desktop tasks. In audited, tool-assisted GAIA tasks, Arabic trails English by 19.4–30.3 percentage points across three earlier flagships. In macOSWorld, both specialized computer-use agents score lower in Arabic when instructions and the operating-system interface switch together. These results concern generic Arabic labels; they do not establish Egyptian Arabic performance. GAIA-v2-LILT, macOSWorld v4.
- Thai: the risk becomes clearer when the whole business workflow is localized. Kimi K2.5 scores 56.1% in English, 57.3% with Thai dialogue only, and 32.7% with the full Thai retail workflow localized. Policies, tools, and database content matter beyond conversational fluency. Earlier web-shopping results also flag Thai. SEATauBench, X-WebAgentBench.
- Japanese and Hindi require narrower conclusions. They are among Kimi K3's lower document-parsing observations, but Japanese desktop results change direction by agent, and Hindi GAIA scores rise substantially after benchmark auditing. Weakness depends on the task and evaluation design. Kimi K3 results.
For the newest frontier models, the 30-language evidence is still incomplete. Most controlled multilingual execution results come from earlier flagships. Current-model failures are useful deployment signals, but a failure observed in Russian, for example, does not prove that Russian caused it without an English control. This review covers 16 benchmark/resource projects and preserves counterexamples, missing evidence, and openness limits. Evidence cutoff: September 5, 2026; no new model evaluations were run.
Auditing the benchmark narrows language gaps without eliminating them
Arabic remains below each model's English reference after the GAIA audit. The first chart shows English minus audited-language pass@1: higher bars mean a larger remaining gap. German, Hindi, Korean, and Brazilian Portuguese have smaller gaps, with different magnitudes by model.
GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | English minus audited (pp) | Model |
|---|---|---|
| Arabic (generic)† | 19.4 | GPT-5.4 |
| German | 3.1 | GPT-5.4 |
| Hindi | 6.7 | GPT-5.4 |
| Korean | 4.3 | GPT-5.4 |
| Portuguese (Brazil) | 8.5 | GPT-5.4 |
| Arabic (generic)† | 21.8 | Gemini 3.1 Pro |
| German | 7.2 | Gemini 3.1 Pro |
| Hindi | 10.3 | Gemini 3.1 Pro |
| Korean | 9.1 | Gemini 3.1 Pro |
| Portuguese (Brazil) | 8.4 | Gemini 3.1 Pro |
| Arabic (generic)† | 30.3 | Claude Opus 4.6 |
| German | 12.7 | Claude Opus 4.6 |
| Hindi | 17 | Claude Opus 4.6 |
| Korean | 20.6 | Claude Opus 4.6 |
| Portuguese (Brazil) | 16.4 | Claude Opus 4.6 |
Every audited language improves for every evaluated model. The second chart shows audited minus minimally translated scores. Audits modify answer alignment, cultural context, and difficulty together: neither the entire improvement nor the entire residual is a pure language effect. Table 2.
GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Audit improvement (pp) | Model |
|---|---|---|
| Arabic (generic)† | 15.2 | GPT-5.4 |
| German | 16.3 | GPT-5.4 |
| Hindi | 25.4 | GPT-5.4 |
| Korean | 29.1 | GPT-5.4 |
| Portuguese (Brazil) | 10.9 | GPT-5.4 |
| Arabic (generic)† | 17.5 | Gemini 3.1 Pro |
| German | 17 | Gemini 3.1 Pro |
| Hindi | 25.4 | Gemini 3.1 Pro |
| Korean | 30.2 | Gemini 3.1 Pro |
| Portuguese (Brazil) | 17 | Gemini 3.1 Pro |
| Arabic (generic)† | 17 | Claude Opus 4.6 |
| German | 17 | Claude Opus 4.6 |
| Hindi | 32.7 | Claude Opus 4.6 |
| Korean | 25.5 | Claude Opus 4.6 |
| Portuguese (Brazil) | 13.9 | Claude Opus 4.6 |
Both charts use the same experiment, not independent replications. MAPS-GAIA and its audited version also share task ancestry. These are GPT-5.4, Gemini 3.1 Pro, and Opus 4.6 results, with 165 tasks per language in an actual tool-assisted workflow. They do not measure today's models directly.
Conversational fluency does not establish local business competence
Thai retail performance falls when localization extends beyond dialogue. The chart compares an English baseline, localized dialogue, and fully localized business content for the same earlier flagship. Vietnamese full-localization retail performance slightly exceeds English, an important counterexample to a universal decline.
SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Task success rate | Setting |
|---|---|---|
| Vietnamese | 56.1% | All English |
| Vietnamese | 68.7% | Dialogue localized |
| Vietnamese | 56.7% | Fully localized business |
| Thai | 56.1% | All English |
| Thai | 57.3% | Dialogue localized |
| Thai | 32.7% | Fully localized business |
| Indonesian | 56.1% | All English |
| Indonesian | 64% | Dialogue localized |
| Indonesian | 43.3% | Fully localized business |
| Filipino† | 56.1% | All English |
| Filipino† | 67.5% | Dialogue localized |
| Filipino† | 44.4% | Fully localized business |
Thai fully localized airline and telecom results also fall below their English references; those observations remain in the chart data. The user simulator changes language too and can make errors, so the measured decline belongs to the overall system. It cannot all be assigned to the evaluated agent. Results use three trials per task; domains are not pooled. SEATauBench Tables 9, 10, and 13.
Desktop results repeat the Arabic signal and complicate the Japanese story
Both specialized computer-use agents score lower in Arabic than in English. Japanese and Russian move in opposite directions across agents: above English for OpenAI CUA, below English for Claude CUA. Compare each agent with its own English bar.
macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Task success rate | Agent |
|---|---|---|
| English reference | 44.4% | Claude CUA |
| Arabic (generic)† | 31.6% | Claude CUA |
| Japanese | 36.8% | Claude CUA |
| Russian | 40.9% | Claude CUA |
| English reference | 33.3% | OpenAI CUA |
| Arabic (generic)† | 28.1% | OpenAI CUA |
| Japanese | 35.1% | OpenAI CUA |
| Russian | 39.2% | OpenAI CUA |
These 2025 model versions run 171 comparable tasks. Instructions and the OS interface change together, combining reading, layout, and planning effects. The authors document Arabic interface localization failures, including positioning errors. This motivates targeted interface testing, without establishing that right-to-left text alone causes the gap. macOSWorld v4 Table 3 and cases.
Translating the workflow into English can make performance worse
Thai is the lowest original-language shopping score among the 11 target languages shown. Translating to English helps French but hurts several others, particularly Arabic and Urdu. The y-values are WebShop task scores, not binary purchase-success percentages.
X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Task score | Strategy |
|---|---|---|
| French | 42.7 | Original-language BaseAgent |
| French | 48.33 | Google Translate to English |
| Spanish | 37.31 | Original-language BaseAgent |
| Spanish | 25.97 | Google Translate to English |
| German | 34.56 | Original-language BaseAgent |
| German | 37.41 | Google Translate to English |
| Russian | 36.41 | Original-language BaseAgent |
| Russian | 38 | Google Translate to English |
| Turkish | 43.18 | Original-language BaseAgent |
| Turkish | 33.99 | Google Translate to English |
| Arabic (generic)† | 41.34 | Original-language BaseAgent |
| Arabic (generic)† | 2.18 | Google Translate to English |
| Vietnamese | 37.93 | Original-language BaseAgent |
| Vietnamese | 35.19 | Google Translate to English |
| Thai | 18.51 | Original-language BaseAgent |
| Thai | 12.69 | Google Translate to English |
| Hindi | 36.75 | Original-language BaseAgent |
| Hindi | 33.47 | Google Translate to English |
| Swahili | 42.1 | Original-language BaseAgent |
| Swahili | 36.49 | Google Translate to English |
| Urdu | 34.15 | Original-language BaseAgent |
| Urdu | 2.3 | Google Translate to English |
This earlier GPT-4o experiment adds a different task family to the Thai concern. It cannot be averaged with SEATau into a language risk score. Swahili's original-language result is comparatively strong, contradicting a simple rule that lower-resource languages always perform worst. The figure compares two published strategies, not every strategy in the paper. X-WebAgentBench Table 2.
Corroboration should include disagreement and missing controls
More sources should constrain the conclusion as well as support it. A paper, its GitHub repository, and its leaderboard form one evidence chain. An audited derivative is not a fresh independent task sample. The following table is a reading guide, not a ranking.
Language signals alongside counterevidence
Researcher synthesis, not risk score or statistical ranking. Shared task ancestry, historical models, components, and non-frontier baselines remain distinguished. Recommendations are unrun tests. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language group | Supporting evidence | Supported interpretation | Counterevidence or limit | Action |
|---|---|---|---|---|
| 01 Thai | SEA retail and other domains; X-Web shopping; K3 documents | Weakness across tasks; current-frontier evidence is component-only | Thai is the highest one-shot language for Llama 3.1 405B in MASSIVE; models/prompts change order | Retest fully localized business and shopping; avoid all-model claims |
| 02 Arabic (generic)† | Audited GAIA: three flagships below English; macOS: two CUAs lower; K3 component | Aligned historical signals across two execution families | GAIA improves after audit; GUI also changes RTL layout; dialects untested | Retest search, RTL UI, and numeric arguments; do not extend to Egyptian Arabic |
| 03 Japanese | macOS CUA; native code tasks; K3 component | Component weakness; execution depends on model | OpenAI CUA exceeds English; Claude CUA falls below English | Separate document reading, click targeting, and display-width tests |
| 04 Hindi | Audited GAIA; native code; MASSIVE; K3 component | Several task families; low scores sensitive to task quality | GAIA audit gains 25.4–32.7 pp; English instructions do not remove all native-task difficulty in Hindi/German | Audit task alignment, entities, and regional rules; distinguish reading from execution |
| 05 Russian | Current RuBench execution and GorillaHard plans; historical macOS | Concrete current failures, more direct than a component | RuBench/MERA lack English pairs; historical OpenAI CUA Russian exceeds English | Regress specific failures; do not rank Russian worst overall |
| 06 Vietnamese / Indonesian / Filipino† | SEA full business; MASSIVE; Vietnamese X-Web | Several tasks; language ordering varies by domain | Kimi Vietnamese retail slightly exceeds English; Filipino simulator also errs | Accept by business domain and verify Filipino–Tagalog variety alignment |
| 07 German / Korean / Portuguese | Audited GAIA; MASSIVE; K3; German/Korean LILT | More evidence does not establish deployment readiness | Audited German gap is 3.1 pp for GPT-5.4, 12.7 pp for Opus 4.6 | Test the target model and actual workflow; no universal safe-language list |
| 08 Amharic / Kannada | Historical MASSIVE zero-shot AST; Amharic synthetic tasks | Risk signal mainly from one historical task family | Lowest language differs by model; no current-frontier execution replication | Prioritize function arguments; evidence cannot rank current models |
| 09 Swahili / Urdu | X-Web; MASSIVE; non-frontier MAST retrieval | Task- and model-dependent signals; no uniform direction | Swahili original-language GPT-4o shopping is not low; speech and retrieval scores are not interchangeable | Measure retrieval, calls, and final state separately with English controls |
| 10 Other coverage / gaps | Bengali/Tamil/Telugu tool or retrieval resources; Pidgin speech component | Consult all 30 rows; missing evidence is not low ability | Egyptian Arabic/Western Punjabi exact matches missing; Hausa lacks current-frontier baseline | Use native tasks: pcm≠en, arz≠ar, pnb≠pa, jv≠id |
Amharic and Kannada warrant additional testing, with a narrower basis: in MASSIVE-Agents' 10k, zero-shot AST setting, Amharic is Nova Premier's lowest language and Kannada is Claude 3.5 v2 Sonnet's lowest. These are historical static function-call results. A later synthetic dataset does not independently replicate those performance findings. MASSIVE-Agents Table 2.
The broader evidence map distinguishes results from test resources
Nine named benchmark columns cover 27 of the 30 target languages in some form. A value of 2 means a verified per-language result in this broader historical/resource layer; 1 means task or aggregate coverage. Neither is an ability score. A dagger marks related Arabic or Filipino labels rather than an exact target-variety match.
Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker |
|---|---|---|---|---|---|---|---|---|---|
| Hindi | 2 | 2 | 2 | 0 | 0 | 1 | 1 | 2 | 1 |
| Spanish | 2 | 0 | 2 | 0 | 0 | 1 | 0 | 2 | 0 |
| Modern Standard Arabic† | 2 | 2 | 2 | 2 | 0 | 1 | 0 | 2 | 0 |
| French | 2 | 0 | 2 | 0 | 0 | 0 | 0 | 2 | 0 |
| Bengali | 2 | 0 | 0 | 0 | 0 | 0 | 1 | 2 | 0 |
| Portuguese | 2 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Indonesian | 2 | 0 | 0 | 0 | 2 | 0 | 0 | 0 | 0 |
| Urdu | 2 | 0 | 2 | 0 | 0 | 0 | 0 | 2 | 0 |
| Russian | 2 | 0 | 2 | 2 | 0 | 0 | 0 | 2 | 0 |
| German | 2 | 2 | 2 | 0 | 0 | 1 | 0 | 2 | 0 |
| Japanese | 2 | 0 | 0 | 2 | 0 | 1 | 0 | 0 | 0 |
| Nigerian Pidgin | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Egyptian Arabic | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Marathi | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 |
| Vietnamese | 2 | 0 | 2 | 0 | 2 | 0 | 0 | 0 | 0 |
The first half combines static calls, execution tasks, and reusable resources. Dense coverage means more ways to investigate a language, not evidence that today's frontier models have passed them.
Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker | Coverage marker |
|---|---|---|---|---|---|---|---|---|---|
| Telugu | 2 | 0 | 0 | 0 | 0 | 0 | 1 | 2 | 0 |
| Swahili | 2 | 0 | 2 | 0 | 0 | 0 | 0 | 2 | 1 |
| Hausa | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| Turkish | 2 | 0 | 2 | 0 | 0 | 1 | 0 | 0 | 0 |
| Western Punjabi | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Tagalog† | 2 | 0 | 0 | 0 | 2 | 0 | 0 | 0 | 0 |
| Tamil | 2 | 0 | 0 | 0 | 0 | 0 | 1 | 2 | 0 |
| Iranian Persian | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Korean | 2 | 2 | 0 | 0 | 0 | 1 | 0 | 0 | 0 |
| Amharic | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| Thai | 2 | 0 | 2 | 0 | 2 | 0 | 0 | 2 | 0 |
| Javanese | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Italian | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Gujarati | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2 | 0 |
| Kannada | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 2 | 0 |
Javanese has historical MASSIVE function-call evidence; Marathi has voice-tool resources; Gujarati has MAST retrieval coverage. Hausa's synthetic tasks do not establish current-model performance. Nigerian Pidgin has a separate code-switching speech component resource, outside this agent-results grid. Egyptian Arabic and Western Punjabi still lack an included exact-match result.
What has actually failed on current models?
The following evidence retains the review's current-model observations. It can identify execution failures and component weaknesses. Without a matched language control, it cannot identify how much of a failure is caused by the language.
Valid formatting can hide substantial task errors
Sol passes formatting on 99.9% of Russian tool-plan samples but passes the full sample on 69.1%. Each bar splits all samples into three mutually exclusive parts. The middle segment shows errors that a JSON-format check would miss.
Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed. SQLite VALUES reproduces reviewed observations; not a live source query.
| Model | Share | Share | Share |
|---|---|---|---|
| GPT-5.6 Sol | 69.1% | 30.8% | 0.1% |
| Claude Opus 5 | 67.7% | 27.2% | 5.1% |
| Grok 4.6 | 73.7% | 26.3% | 0% |
GorillaHard grades static outputs covering calls, abstention, and clarification; it does not execute tools. The documented dataset has 1,169 items, while submission-specific evaluated counts and repetitions are not separately disclosed. Model settings are not fully aligned, so this chart does not select a model winner. Sol, Opus 5, Grok 4.6.
Correct tool selection can still produce incorrect arguments
About 8.1%–16.3% of call-required samples select matching tools without matching all arguments. The middle segment points to concrete acceptance checks: entities, amounts, dates, and argument transfer between steps.
Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference. SQLite VALUES reproduces reviewed observations; not a live source query.
| Model | Share | Share | Share |
|---|---|---|---|
| GPT-5.6 Sol | 70.2% | 16.3% | 13.5% |
| Claude Opus 5 | 67.7% | 11.3% | 21% |
| Grok 4.6 | 74.3% | 8.1% | 17.6% |
This denominator contains only call-required items, unlike the previous figure. The nested metrics share a denominator within this chart, enabling subtraction; the two figures cannot be joined into a funnel. No English comparison establishes a Russian penalty. Scoring implementation.
Knowing when to stop needs its own acceptance test
On abstention-required items, roughly 33.9%–41.7% lack the correct abstention. False abstention on call-required items is much rarer. The chart separates these conditional error rates; the denominators appear in the category labels.
1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions. SQLite VALUES reproduces reviewed observations; not a live source query.
| Required action | Conditional error rate | Model |
|---|---|---|
| Abstention required (127) | 41.7% | GPT-5.6 Sol |
| Call required (1,032) | 0.9% | GPT-5.6 Sol |
| Abstention required (127) | 35.4% | Claude Opus 5 |
| Call required (1,032) | 0.1% | Claude Opus 5 |
| Abstention required (127) | 33.9% | Grok 4.6 |
| Call required (1,032) | 0.4% | Grok 4.6 |
These rates cannot be added, and a static output error is not an observed unauthorized real-world action. Abstention conditions follow task rules and are not all safety-policy refusals. GorillaHard definition.
Real code execution reveals reproducible task failures
RuBench asks agents to modify real repositories from Russian issue instructions, then runs regression tests. The chart separates audit-retained passes, raw passes removed for answer exposure, and original failures.
Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control. SQLite VALUES reproduces reviewed observations; not a live source query.
| Configuration | Share | Share | Share |
|---|---|---|---|
| GPT-5.6 Sol (23 tasks × 3) | 71% | 11.6% | 17.4% |
| Claude Opus 5 (20 tasks × 3) | 76.7% | 1.7% | 21.7% |
Opus 5 fully passes 0/3 runs on each of two HTTP fixes: method/header case rules and an empty-filename parsing exception. These are concrete task failures, without evidence that Russian caused them. Sol and Opus use different task sets and harnesses. Audit deductions re-score existing network-enabled runs; they are not offline reruns. Pinned run record.
Four document languages deserve targeted end-to-end checks
Thai, Japanese, generic Arabic, and Hindi sit at the lower end of Kimi K3's 13 included language observations. The median line locates observations within this sample; it is not a deployment threshold.
Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Parsing quality score | Observed position |
|---|---|---|
| Thai | 72.1 | Lower four observations |
| Japanese | 74.9 | Lower four observations |
| Arabic (unspecified variety) | 77.4 | Lower four observations |
| Hindi | 77.5 | Lower four observations |
| French | 80 | Other nine observations |
| Spanish | 80.2 | Other nine observations |
| Russian | 82.4 | Other nine observations |
| Vietnamese | 84.8 | Other nine observations |
| Indonesian | 86.9 | Other nine observations |
| Portuguese | 88.9 | Other nine observations |
| German | 89.1 | Other nine observations |
| Korean | 89.9 | Other nine observations |
| Italian | 92.7 | Other nine observations |
These scores combine text, layout, formula, and table parsing. They are not agent success rates. Language documents are not paired for identical content, and Arabic varieties are not separated. Validate the complete sequence: read a field, call the API, and inspect the resulting state. Pinned author results.
The script groups overlap substantially, and Korean scores 89.9. A blanket non-Latin-script weakness label would conceal this counterexample.
Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges. SQLite VALUES reproduces reviewed observations; not a live source query.
| Script group | Parsing quality score | Parsing quality score | Parsing quality score | Parsing quality score | Parsing quality score |
|---|---|---|---|---|---|
| Latin script (7) | 80 | 82.5 | 86.9 | 89 | 92.7 |
| Other scripts (6) | 72.1 | 75.53 | 77.45 | 81.18 | 89.9 |
Boxes show the middle half of language observations; whiskers show minima and maxima. This summarizes 13 observations, not confidence intervals or the causal effect of a writing system. Parsing metric definitions.
Evidence for the current frontier is still sparse
These two maps include only this review's September current-model cohort. The broader historical results and resources above are excluded. A value of 1 marks included evidence, 0.5 marks the related Arabic label, and 0 means no included result was found.
Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Evidence marker | Evidence marker | Evidence marker | Evidence marker |
|---|---|---|---|---|
| 01 Hindi | 1 | 0 | 0 | 0 |
| 02 Spanish | 1 | 0 | 0 | 0 |
| 03 Modern Standard Arabic | 0.5 | 0 | 0 | 0 |
| 04 French | 1 | 0 | 0 | 0 |
| 05 Bengali | 0 | 0 | 0 | 0 |
| 06 Portuguese | 1 | 0 | 0 | 0 |
| 07 Indonesian | 1 | 0 | 0 | 0 |
| 08 Urdu | 0 | 0 | 0 | 0 |
| 09 Russian | 1 | 1 | 1 | 0 |
| 10 German | 1 | 0 | 0 | 0 |
| 11 Japanese | 1 | 0 | 0 | 0 |
| 12 Nigerian Pidgin | 0 | 0 | 0 | 0 |
| 13 Egyptian Arabic | 0 | 0 | 0 | 0 |
| 14 Marathi | 0 | 0 | 0 | 0 |
| 15 Vietnamese | 1 | 0 | 0 | 0 |
In the first half, Russian has code execution and static tool-plan results. French, Spanish, and other languages have document-component results, which do not establish browser or multi-turn business reliability.
Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages. SQLite VALUES reproduces reviewed observations; not a live source query.
| Language | Evidence marker | Evidence marker | Evidence marker | Evidence marker |
|---|---|---|---|---|
| 16 Telugu | 0 | 0 | 0 | 0 |
| 17 Swahili | 0 | 0 | 0 | 0 |
| 18 Hausa | 0 | 0 | 0 | 0 |
| 19 Turkish | 0 | 0 | 0 | 0 |
| 20 Western Punjabi | 0 | 0 | 0 | 0 |
| 21 Tagalog | 0 | 0 | 0 | 0 |
| 22 Tamil | 0 | 0 | 0 | 0 |
| 23 Iranian Persian | 0 | 0 | 0 | 0 |
| 24 Korean | 1 | 0 | 0 | 0 |
| 25 Amharic | 0 | 0 | 0 | 0 |
| 26 Thai | 1 | 0 | 0 | 0 |
| 27 Javanese | 0 | 0 | 0 | 0 |
| 28 Italian | 1 | 0 | 0 | 0 |
| 29 Gujarati | 0 | 0 | 0 | 0 |
| 30 Kannada | 0 | 0 | 0 | 0 |
In the second half, Korean, Thai, and Italian have document-component results. Historical Javanese function-call evidence does not supply a current execution guarantee. Gurmukhi Punjabi cannot substitute for Western Punjabi. No current-model matched-language experiment was included in the paired-comparison column.
Turn suspected mechanisms into verifiable tasks
Urdu and Iranian Persian offer concrete test conditions involving numerals, dates, and mixed-script identifiers. These are proposed fixtures, not measured current-agent failure rates. Unicode number conventions.
Proposed acceptance tests
Researcher-proposed fixtures based on Unicode conventions and interface contracts. No measured model outcomes or failure incidence. Preserve distinctions pcm≠en, arz≠ar, pnb≠pa, jv≠id. SQLite VALUES reproduces reviewed observations; not a live source query.
| Test target | Input condition | Expected final result |
|---|---|---|
| Amounts | Urdu / Persian | ۱٬۲۵۰٫۵۰; explicit currency | Pass 1250.5 according to the API contract; preserve currency |
| Dates | explicit locale / calendar | 05/09/2026; with/without locale control | Convert correctly when defined; clarify per rules when ambiguous |
| Dialect / neighboring language | four gaps | Native pcm, arz, pnb, jv requests | Preserve native constraints; no substitution by neighboring-language scores |
| Entity IDs | all languages | INV-01250/AB; mixed scripts | Preserve string, leading zeros, separators; act on the correct entity |
The most informative next experiment holds model, tools, budget, and business objective fixed, changes the language, and checks the final state. Add locally authored tasks alongside those controls: matched tasks help estimate language gaps; native tasks test practical usefulness.
Scope, metrics, and sources
This review checks published papers, author repositories, and dataset cards without running models. Its 16 core benchmark/resource entries are not an exhaustive census of public research. The fixed 30-language scope follows total-speaker coverage after excluding English and Chinese varieties, using a pinned public transcription. Population determines inclusion, not training resources or ability.
The September 5 frontier scope includes Astra, Sol, Fable 5.1/5, Opus 5, Muse Spark 1.3, Grok 4.6, Kimi K3, GLM 5.3, and near-boundary Gemini 3.8 Flash/Qwen3.8. Per-language usable results exist for only a subset. Earlier flagships are labeled separately throughout. Cohort reference.
Three current-model evidence layers
Review synthesis of the audited RuBench, GorillaHard, and MDP configurations. Do not equate static plans, document components, executed tasks, or language-causal comparisons. SQLite VALUES reproduces reviewed observations; not a live source query.
| Evidence | Measurement | Key limitation |
|---|---|---|
| GorillaHard | Static output; documented N=1,169 total, 1,032 calls, 127 abstentions | Documented counts, not separately disclosed run N. Incomplete settings; Sol source commit unresolved; no English pairs. |
| MDPBench | Kimi K3 document parsing quality; 13 target observations | Unpaired content/layout; unspecified Arabic variety; composite score is not failure probability. |
| RuBench | Russian instructions → repository edits → regression tests | Sol 23×3, Opus 20×3; different sets/harnesses. Earlier Fable fallback omitted. |
Model-only aggregate scores from MCP Atlas and a tiny German evaluation without original traces do not fill per-language gaps here. “Not found” refers to this search and its inclusion rules. The lookup below keeps all 30 target languages and identifies useful evidence and next checks.
All 30 languages: evidence and next checks
Researcher mapping retains all target languages, named evidence and inference limits. No invented scores; related labels do not replace exact language varieties. SQLite VALUES reproduces reviewed observations; not a live source query.
| No. | Language | Code | Available evidence | Interpretation and next check |
|---|---|---|---|---|
| 1Source: All 30 languages: evidence and next checks | Hindi | hi | MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks | Audited GAIA recovers substantially; K3 reading is low. Separate benchmark quality from local-rule execution. |
| 2Source: All 30 languages: evidence and next checks | Spanish | es | MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval | GAIA parent set, shopping, function calls, and retrieval coverage. No complete current-frontier comparison; do not presume European-language reliability. |
| 3Source: All 30 languages: evidence and next checks | Modern Standard Arabic | ar | MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval | GAIA and macOS give aligned historical signals; K3 component score is lower. Generic Arabic does not precisely establish MSA or Egyptian dialect performance. |
| 4Source: All 30 languages: evidence and next checks | French | fr | MASSIVE calls; X-Web shopping; MAST retrieval | Shopping, function calls, retrieval baselines, and a current document component. Retest full business execution on the target model. |
| 5Source: All 30 languages: evidence and next checks | Bengali | bn | MASSIVE calls; Voice tools; MAST retrieval | Historical calls, voice-tool tasks, and retrieval resources. Current-frontier execution remains unestablished; test entities and arguments. |
| 6Source: All 30 languages: evidence and next checks | Portuguese | pt | MASSIVE calls; GAIA audited | Audited GAIA uses Brazilian Portuguese; MASSIVE uses Portugal locale. Regional business rules are not interchangeable. |
| 7Source: All 30 languages: evidence and next checks | Indonesian | id | MASSIVE calls; SEA multi-turn | SEA covers full multi-turn workflows with domain-dependent effects. Indonesian does not substitute for Javanese. |
| 8Source: All 30 languages: evidence and next checks | Urdu | ur | MASSIVE calls; X-Web shopping; MAST retrieval | Historical/non-frontier calls, shopping, and retrieval. Numerals and dates are proposed mechanisms, not measured current-model failure rates. |
| 9Source: All 30 languages: evidence and next checks | Russian | ru | MASSIVE calls; X-Web shopping; macOS use; MAST retrieval | Current code execution and static calls contain failures. Without matched English tasks, Russian causation is unestablished. |
| 10Source: All 30 languages: evidence and next checks | German | de | MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval | Audited GAIA and native code coverage are relatively rich. Gaps differ by model; component strength is not deployment acceptance. |
| 11Source: All 30 languages: evidence and next checks | Japanese | ja | MASSIVE calls; macOS use; LILT native code | Current parsing is low; historical CUA gaps change direction by model. Test visual targeting, text, and local software rules separately. |
| 12Source: All 30 languages: evidence and next checks | Nigerian Pidgin | pcm | SwitchBoard speech component | Code-switching speech resources found; current-frontier tool execution unverified. Do not substitute English. |
| 13Source: All 30 languages: evidence and next checks | Egyptian Arabic | arz | No matching execution result included | Generic Arabic studies cannot fill this dialect. No included exact-dialect agent result. |
| 14Source: All 30 languages: evidence and next checks | Marathi | mr | Voice tools | VoiceAgentBench provides speech-tool tasks; do not extend its English-only multi-turn results to Marathi. |
| 15Source: All 30 languages: evidence and next checks | Vietnamese | vi | MASSIVE calls; X-Web shopping; SEA multi-turn | SEA, shopping, function calls, and documents provide multiple sources. Kimi fully localized retail slightly exceeds English: preserve the counterexample. |
30 results · Showing first 15
Public evidence has several different levels of openness
A public paper, an open task set, evaluation code, and raw runs are different assets. Terminal-Bench-LILT exposes samples but requires contacting the authors for the full set; the community LILTBench collection is separate. MASSIVE-Agents' converted dataset is publicly available, superseding an earlier “not released” characterization.
16 projects: measurements and openness
Project inventory, not independent experiment count or exhaustive literature census. Primary papers, author repositories, and data cards checked. Unverified does not mean unavailable. No full repository-by-repository reproduction or legal license review. MAST baseline only as of cutoff. SQLite VALUES reproduces reviewed observations; not a live source query.
| No. | Project | Language coverage | Models and evidence layer | Verified openness | Interpretation limits |
|---|---|---|---|---|---|
| 1Source: 16 projects: measurements and openness | RuBench | Russian | Sol, Opus 5 / current direct | Tasks and run results public; full grading tests withheld | No matched English; models use different task sets |
| 2Source: 16 projects: measurements and openness | MERA GorillaHard | Russian | Sol, Opus 5, Grok 4.6 / current direct | Framework, metrics, scores public; incomplete test truth/configuration | Tools not executed; submission settings not fully aligned |
| 3Source: 16 projects: measurements and openness | MDPBench / Kimi K3 | 13 target observations, including related Arabic label | Kimi K3 / current component | Official per-language scores public; no rerun here | Not agent success; unpaired documents |
| 4Source: 16 projects: measurements and openness | GAIA-v2-LILT | ar de hi ko pt-BR + en | GPT-5.4, Gemini 3.1 Pro, Opus 4.6 / earlier flagships | Evaluation code, HF tasks, per-language table public | Audit changes function, culture, difficulty together; not pure language causality |
| 5Source: 16 projects: measurements and openness | MAPS | Final paper: 12 languages; older HF card: 11 | Historical models / context | Paper and HF data public; pin subset and version | MATH is not execution; GAIA shares ancestry with preceding row |
| 6Source: 16 projects: measurements and openness | SEATauBench | vi th id Filipino zh + en | Kimi K2.5, GPT-5-mini, Qwen3 / historical | Code, domain data, analysis framework public; MIT repository | Simulator can also fail; no isolated agent-only language effect |
| 7Source: 16 projects: measurements and openness | macOSWorld | ar ja ru zh + en | Claude 3.7 CUA, 2025 OpenAI CUA / earlier flagships | Tasks, environment, scoring code public; per-language paper results | Instructions and UI switch together; v4 revises some scores |
| 8Source: 16 projects: measurements and openness | X-WebAgentBench | 14 non-English settings; 11 target languages | GPT-4o and others / historical | Code and data entry points public; paper results | Task Score is not success rate; derived from WebShop |
| 9Source: 16 projects: measurements and openness | MASSIVE-Agents | 52 languages; 24 targets including related labels | Nova Premier, Claude 3.5 and others / earlier flagships | Converted HF data CC BY 4.0; per-language paper tables | Not multi-turn execution; filtering changes samples/function coverage by language |
| 10Source: 16 projects: measurements and openness | Terminal-Bench-LILT | ar cs de es hi ja ko sr tr zh | GPT-5.5, Opus 4.8 and others / earlier flagships | Paper, aggregate scores, samples public; full tasks require contacting authors | Blog 300 versus leaderboard 324 tasks: denominators not mixed; full set not public |
| 11Source: 16 projects: measurements and openness | LILT multilingual tau | de ko + en | GPT-5.4, Opus 4.8, Gemini 3.1 Pro / earlier flagships | Public board and run notes; dynamic per-language scores not extracted here | Different English/target-language simulators confound scores |
| 12Source: 16 projects: measurements and openness | LILTBench community | Authors report 31 languages; not exhaustively checked task by task | Opus 4.6 / challenging test resources | Tasks, verifiers, leaderboard public; Apache-2.0 repository | Adversarial hard-task selection cannot estimate population failure rates |
| 13Source: 16 projects: measurements and openness | VoiceAgentBench | en hi bn mr ta te ml | SpeechLM / ASR+LLM / non-frontier resources | Evaluation code and HF data entry public; custom community license | Multi-turn subset is English only; predicted calls are not final environment success |
| 14Source: 16 projects: measurements and openness | TelcoAgent-Bench | ar + en | 3B–8B models / non-frontier resources | Blueprints, tasks, prediction/score JSON public; license unverified | Intent/resolution similarity is not strict success; cannot extrapolate to flagships |
| 15Source: 16 projects: measurements and openness | MultiAgent-X | 12 languages; target hi sw ha am | No verified current-frontier baseline / resources | Samples, scripts, structure public; full license unverified | Synthetic; structural validity is not native quality; pa is not pnb |
16 results · Showing first 15
VoiceAgentBench and TelcoAgent add task coverage, without current-frontier results. MAST currently provides a non-frontier baseline; its scheduled final competition results are still in the future at the review cutoff. These entries expand test resources, not the flagship-performance charts.
What to test next
Start with languages that have repeated task-specific signals, then fill commercially important blind spots. Test Thai dialogue-only, tool, and full-business localization separately. Separate standard Arabic, dialects, and right-to-left interfaces. Split Japanese reading, visual targeting, and execution. Audit Hindi answer keys and regional rules before interpreting low scores. Prioritize other languages by local usage and error cost, not by how little research exists.
Fix the exact model, tool versions, task budget, and acceptance criteria. Retain every tool call and final state; for speech, retain the original audio and a correct-transcript control. The open question is which current-model failures persist after separating hearing, reading, tool selection, argument formation, and execution. None of these proposed follow-up tests has been run for this review.
Sources
- All samples: format and content outcomes
Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Model", "Full sample passed", "Valid format, sample failed", "Invalid format", "Submission date", "Evaluated N", "Configuration limits", "Source", "Documented dataset N", "Scope") AS ( VALUES ('GPT-5.6 Sol', 0.691, 0.308, 0.001, '2026-08-18', 'Not separately disclosed', 'openrouter / 9ebf388', 'https://mera.a-ai.ru/en/text/submits/2.0/8', 1169, 'All call, abstention, and clarification items'), ('Claude Opus 5', 0.677, 0.272, 0.051, '2026-08-18', 'Not separately disclosed', 'openrouter / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/1', 1169, 'All call, abstention, and clarification items'), ('Grok 4.6', 0.737, 0.263, 0, '2026-08-18', 'Not separately disclosed', 'local-chat-completions / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/11', 1169, 'All call, abstention, and clarification items') ) SELECT * FROM reviewed_published_rows; - Call-required items: tools and arguments
Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Model", "Tools and arguments matched", "Tools matched, arguments incomplete", "Tools not fully matched", "Submission date", "Evaluated N", "Configuration limits", "Source", "Documented dataset N", "Scope") AS ( VALUES ('GPT-5.6 Sol', 0.702, 0.163, 0.135, '2026-08-18', 'Not separately disclosed', 'openrouter / 9ebf388', 'https://mera.a-ai.ru/en/text/submits/2.0/8', 1032, 'Call-required items only'), ('Claude Opus 5', 0.677, 0.113, 0.21, '2026-08-18', 'Not separately disclosed', 'openrouter / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/1', 1032, 'Call-required items only'), ('Grok 4.6', 0.743, 0.081, 0.176, '2026-08-18', 'Not separately disclosed', 'local-chat-completions / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/11', 1032, 'Call-required items only') ) SELECT * FROM reviewed_published_rows; - Abstention errors under two conditions
1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Required action", "Model", "Conditional error rate", "Documented denominator", "Scoring interpretation", "Evaluated N", "Source") AS ( VALUES ('Abstention required (127)', 'GPT-5.6 Sol', 0.417, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/8'), ('Call required (1,032)', 'GPT-5.6 Sol', 0.009, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/8'), ('Abstention required (127)', 'Claude Opus 5', 0.354, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/1'), ('Call required (1,032)', 'Claude Opus 5', 0.001, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/1'), ('Abstention required (127)', 'Grok 4.6', 0.339, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/11'), ('Call required (1,032)', 'Grok 4.6', 0.004, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/11') ) SELECT * FROM reviewed_published_rows; - Russian code tasks: audited run outcomes
Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Configuration", "Audit-retained passes", "Raw passes removed by audit", "Original failures", "Retained pass count", "Removed pass count", "Original failure count", "Total runs", "Distinct tasks", "Harness", "Setting", "Source") AS ( VALUES ('GPT-5.6 Sol (23 tasks × 3)', 0.7101449275362319, 0.11594202898550725, 0.17391304347826086, 49, 8, 12, 69, 23, 'Codex CLI', 'xhigh', 'https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md'), ('Claude Opus 5 (20 tasks × 3)', 0.7666666666666667, 0.016666666666666666, 0.21666666666666667, 46, 1, 13, 60, 20, 'Claude Code', 'xhigh', 'https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md') ) SELECT * FROM reviewed_published_rows; - Document parsing across 13 language observations
Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "Code", "Parsing quality score", "Observed position", "Script group", "Order", "Median across 13 observations", "Source") AS ( VALUES ('Thai', 'th', 72.1, 'Lower four observations', 'Other scripts', 1, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Japanese', 'ja', 74.9, 'Lower four observations', 'Other scripts', 2, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Arabic (unspecified variety)', 'ar', 77.4, 'Lower four observations', 'Other scripts', 3, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Hindi', 'hi', 77.5, 'Lower four observations', 'Other scripts', 4, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('French', 'fr', 80.0, 'Other nine observations', 'Latin script', 5, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Spanish', 'es', 80.2, 'Other nine observations', 'Latin script', 6, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Russian', 'ru', 82.4, 'Other nine observations', 'Other scripts', 7, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Vietnamese', 'vi', 84.8, 'Other nine observations', 'Latin script', 8, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Indonesian', 'id', 86.9, 'Other nine observations', 'Latin script', 9, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Portuguese', 'pt', 88.9, 'Other nine observations', 'Latin script', 10, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('German', 'de', 89.1, 'Other nine observations', 'Latin script', 11, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Korean', 'ko', 89.9, 'Other nine observations', 'Other scripts', 12, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Italian', 'it', 92.7, 'Other nine observations', 'Latin script', 13, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e') ) SELECT * FROM reviewed_published_rows; - Parsing score distributions by script group
Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Script group", "Minimum", "First quartile", "Median", "Third quartile", "Maximum", "Observation count", "Mean", "Language", "Source") AS ( VALUES ('Latin script (7)', 80.0, 82.5, 86.9, 89.0, 92.7, 7, 86.08571428571429, 'French; Spanish; Vietnamese; Indonesian; Portuguese; German; Italian', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), ('Other scripts (6)', 72.1, 75.525, 77.45, 81.17500000000001, 89.9, 6, 79.03333333333333, 'Thai; Japanese; Arabic (unspecified variety); Hindi; Russian; Korean', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e') ) SELECT * FROM reviewed_published_rows; - Current-model evidence map · 1–15
Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "Document component", "Static tool plan", "End-to-end execution", "Matched language control", "Marker definition") AS ( VALUES ('01 Hindi', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('02 Spanish', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('03 Modern Standard Arabic', 0.5, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('04 French', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('05 Bengali', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('06 Portuguese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('07 Indonesian', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('08 Urdu', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('09 Russian', 1, 1, 1, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('10 German', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('11 Japanese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('12 Nigerian Pidgin', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('13 Egyptian Arabic', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('14 Marathi', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('15 Vietnamese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability') ) SELECT * FROM reviewed_published_rows; - Current-model evidence map · 16–30
Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "Document component", "Static tool plan", "End-to-end execution", "Matched language control", "Marker definition") AS ( VALUES ('16 Telugu', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('17 Swahili', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('18 Hausa', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('19 Turkish', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('20 Western Punjabi', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('21 Tagalog', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('22 Tamil', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('23 Iranian Persian', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('24 Korean', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('25 Amharic', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('26 Thai', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('27 Javanese', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('28 Italian', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('29 Gujarati', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'), ('30 Kannada', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability') ) SELECT * FROM reviewed_published_rows; - Proposed acceptance tests
Researcher-proposed fixtures based on Unicode conventions and interface contracts. No measured model outcomes or failure incidence. Preserve distinctions pcm≠en, arz≠ar, pnb≠pa, jv≠id. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Test target", "Input condition", "Expected final result") AS ( VALUES ('Amounts | Urdu / Persian', '۱٬۲۵۰٫۵۰; explicit currency', 'Pass 1250.5 according to the API contract; preserve currency'), ('Entity IDs | all languages', 'INV-01250/AB; mixed scripts', 'Preserve string, leading zeros, separators; act on the correct entity'), ('Dates | explicit locale / calendar', '05/09/2026; with/without locale control', 'Convert correctly when defined; clarify per rules when ambiguous'), ('Dialect / neighboring language | four gaps', 'Native pcm, arz, pnb, jv requests', 'Preserve native constraints; no substitution by neighboring-language scores') ) SELECT * FROM reviewed_published_rows; - Three current-model evidence layers
Review synthesis of the audited RuBench, GorillaHard, and MDP configurations. Do not equate static plans, document components, executed tasks, or language-causal comparisons. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Evidence", "Measurement", "Key limitation") AS ( VALUES ('GorillaHard', 'Static output; documented N=1,169 total, 1,032 calls, 127 abstentions', 'Documented counts, not separately disclosed run N. Incomplete settings; Sol source commit unresolved; no English pairs.'), ('MDPBench', 'Kimi K3 document parsing quality; 13 target observations', 'Unpaired content/layout; unspecified Arabic variety; composite score is not failure probability.'), ('RuBench', 'Russian instructions → repository edits → regression tests', 'Sol 23×3, Opus 20×3; different sets/harnesses. Earlier Fable fallback omitted.') ) SELECT * FROM reviewed_published_rows; - RuBench v1: construction, protocol and Fable fallback
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- RuBench task metadata
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- AIOH4 original Russian task statement
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- GorillaHard dataset description and Russian sample
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- GorillaHard task configuration
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- MERA team's Text 2.0 launch explanation
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- MERA GPT 5.6 Sol submission
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- MERA Claude Opus 5 submission
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- MERA Grok 4.6 submission
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- MDPBench parsing metrics
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- Pinned public language population transcription
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- Frontier cohort reference
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- Score improvement after task auditing
GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "Code", "Model", "Before audit", "After audit", "English reference", "Audit improvement (pp)", "English minus audited (pp)", "Tasks per language", "Model cohort", "Execution harness", "Source") AS ( VALUES ('Arabic (generic)†', 'ar', 'GPT-5.4', 32.1, 47.3, 66.7, 15.2, 19.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('German', 'de', 'GPT-5.4', 47.3, 63.6, 66.7, 16.3, 3.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Hindi', 'hi', 'GPT-5.4', 34.6, 60.0, 66.7, 25.4, 6.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Korean', 'ko', 'GPT-5.4', 33.3, 62.4, 66.7, 29.1, 4.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Portuguese (Brazil)', 'pt', 'GPT-5.4', 47.3, 58.2, 66.7, 10.9, 8.5, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Arabic (generic)†', 'ar', 'Gemini 3.1 Pro', 34.6, 52.1, 73.9, 17.5, 21.8, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('German', 'de', 'Gemini 3.1 Pro', 49.7, 66.7, 73.9, 17.0, 7.2, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Hindi', 'hi', 'Gemini 3.1 Pro', 38.2, 63.6, 73.9, 25.4, 10.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Korean', 'ko', 'Gemini 3.1 Pro', 34.6, 64.8, 73.9, 30.2, 9.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Portuguese (Brazil)', 'pt', 'Gemini 3.1 Pro', 48.5, 65.5, 73.9, 17.0, 8.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Arabic (generic)†', 'ar', 'Claude Opus 4.6', 32.1, 49.1, 79.4, 17.0, 30.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('German', 'de', 'Claude Opus 4.6', 49.7, 66.7, 79.4, 17.0, 12.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Hindi', 'hi', 'Claude Opus 4.6', 29.7, 62.4, 79.4, 32.7, 17.0, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Korean', 'ko', 'Claude Opus 4.6', 33.3, 58.8, 79.4, 25.5, 20.6, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Portuguese (Brazil)', 'pt', 'Claude Opus 4.6', 49.1, 63.0, 79.4, 13.9, 16.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1') ) SELECT * FROM reviewed_published_rows; - Remaining gap from each model’s English reference
GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "Code", "Model", "Before audit", "After audit", "English reference", "Audit improvement (pp)", "English minus audited (pp)", "Tasks per language", "Model cohort", "Execution harness", "Source") AS ( VALUES ('Arabic (generic)†', 'ar', 'GPT-5.4', 32.1, 47.3, 66.7, 15.2, 19.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('German', 'de', 'GPT-5.4', 47.3, 63.6, 66.7, 16.3, 3.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Hindi', 'hi', 'GPT-5.4', 34.6, 60.0, 66.7, 25.4, 6.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Korean', 'ko', 'GPT-5.4', 33.3, 62.4, 66.7, 29.1, 4.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Portuguese (Brazil)', 'pt', 'GPT-5.4', 47.3, 58.2, 66.7, 10.9, 8.5, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Arabic (generic)†', 'ar', 'Gemini 3.1 Pro', 34.6, 52.1, 73.9, 17.5, 21.8, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('German', 'de', 'Gemini 3.1 Pro', 49.7, 66.7, 73.9, 17.0, 7.2, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Hindi', 'hi', 'Gemini 3.1 Pro', 38.2, 63.6, 73.9, 25.4, 10.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Korean', 'ko', 'Gemini 3.1 Pro', 34.6, 64.8, 73.9, 30.2, 9.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Portuguese (Brazil)', 'pt', 'Gemini 3.1 Pro', 48.5, 65.5, 73.9, 17.0, 8.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Arabic (generic)†', 'ar', 'Claude Opus 4.6', 32.1, 49.1, 79.4, 17.0, 30.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('German', 'de', 'Claude Opus 4.6', 49.7, 66.7, 79.4, 17.0, 12.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Hindi', 'hi', 'Claude Opus 4.6', 29.7, 62.4, 79.4, 32.7, 17.0, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Korean', 'ko', 'Claude Opus 4.6', 33.3, 58.8, 79.4, 25.5, 20.6, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'), ('Portuguese (Brazil)', 'pt', 'Claude Opus 4.6', 49.1, 63.0, 79.4, 13.9, 16.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1') ) SELECT * FROM reviewed_published_rows; - Retail success by extent of localization
SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "Code", "Setting", "Task success rate", "Model", "Displayed domain", "Trials per task", "Airline English", "Airline fully localized", "Telecom English", "Telecom fully localized", "Sample definition", "Source") AS ( VALUES ('Vietnamese', 'vi', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Vietnamese', 'vi', 'Dialogue localized', 0.687, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Vietnamese', 'vi', 'Fully localized business', 0.567, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Thai', 'th', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Thai', 'th', 'Dialogue localized', 0.573, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Thai', 'th', 'Fully localized business', 0.327, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Indonesian', 'id', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Indonesian', 'id', 'Dialogue localized', 0.64, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Indonesian', 'id', 'Fully localized business', 0.433, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Filipino†', 'tl', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Filipino†', 'tl', 'Dialogue localized', 0.675, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'), ('Filipino†', 'tl', 'Fully localized business', 0.444, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1') ) SELECT * FROM reviewed_published_rows; - Desktop task success by agent and language
macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "Code", "Agent", "Task success rate", "Exact model version", "Task count", "Difference from English (pp)", "UI and instruction languages", "Source") AS ( VALUES ('English reference', 'en', 'Claude CUA', 0.444, 'claude-3-7-sonnet-20250219', 171, 0.0, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'), ('Arabic (generic)†', 'ar', 'Claude CUA', 0.316, 'claude-3-7-sonnet-20250219', 171, -12.8, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'), ('Japanese', 'ja', 'Claude CUA', 0.368, 'claude-3-7-sonnet-20250219', 171, -7.6, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'), ('Russian', 'ru', 'Claude CUA', 0.409, 'claude-3-7-sonnet-20250219', 171, -3.5, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'), ('English reference', 'en', 'OpenAI CUA', 0.33299999999999996, 'computer-use-preview-2025-03-11', 171, 0.0, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'), ('Arabic (generic)†', 'ar', 'OpenAI CUA', 0.281, 'computer-use-preview-2025-03-11', 171, -5.2, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'), ('Japanese', 'ja', 'OpenAI CUA', 0.35100000000000003, 'computer-use-preview-2025-03-11', 171, 1.8, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'), ('Russian', 'ru', 'OpenAI CUA', 0.392, 'computer-use-preview-2025-03-11', 171, 5.9, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4') ) SELECT * FROM reviewed_published_rows; - Web shopping scores: original language versus English translation
X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "Code", "Strategy", "Task score", "Model", "Instructions per language", "Metric", "Source") AS ( VALUES ('French', 'fr', 'Original-language BaseAgent', 42.7, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('French', 'fr', 'Google Translate to English', 48.33, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Spanish', 'es', 'Original-language BaseAgent', 37.31, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Spanish', 'es', 'Google Translate to English', 25.97, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('German', 'de', 'Original-language BaseAgent', 34.56, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('German', 'de', 'Google Translate to English', 37.41, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Russian', 'ru', 'Original-language BaseAgent', 36.41, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Russian', 'ru', 'Google Translate to English', 38.0, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Turkish', 'tr', 'Original-language BaseAgent', 43.18, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Turkish', 'tr', 'Google Translate to English', 33.99, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Arabic (generic)†', 'ar', 'Original-language BaseAgent', 41.34, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Arabic (generic)†', 'ar', 'Google Translate to English', 2.18, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Vietnamese', 'vi', 'Original-language BaseAgent', 37.93, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Vietnamese', 'vi', 'Google Translate to English', 35.19, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Thai', 'th', 'Original-language BaseAgent', 18.51, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Thai', 'th', 'Google Translate to English', 12.69, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Hindi', 'hi', 'Original-language BaseAgent', 36.75, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Hindi', 'hi', 'Google Translate to English', 33.47, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Swahili', 'sw', 'Original-language BaseAgent', 42.1, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Swahili', 'sw', 'Google Translate to English', 36.49, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Urdu', 'ur', 'Original-language BaseAgent', 34.15, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'), ('Urdu', 'ur', 'Google Translate to English', 2.3, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1') ) SELECT * FROM reviewed_published_rows; - 16 projects: measurements and openness
Project inventory, not independent experiment count or exhaustive literature census. Primary papers, author repositories, and data cards checked. Unverified does not mean unavailable. No full repository-by-repository reproduction or legal license review. MAST baseline only as of cutoff. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("No.", "Project", "Language coverage", "Task layer", "Models and evidence layer", "Verified openness", "Interpretation limits", "Task ancestry", "Source") AS ( VALUES (1, 'RuBench', 'Russian', 'Code execution', 'Sol, Opus 5 / current direct', 'Tasks and run results public; full grading tests withheld', 'No matched English; models use different task sets', 'Independent native code tasks', 'https://github.com/eugeneshilow/rubench'), (2, 'MERA GorillaHard', 'Russian', 'Static tool plans', 'Sol, Opus 5, Grok 4.6 / current direct', 'Framework, metrics, scores public; incomplete test truth/configuration', 'Tools not executed; submission settings not fully aligned', 'MERA / GorillaHard', 'https://mera.a-ai.ru/en/text'), (3, 'MDPBench / Kimi K3', '13 target observations, including related Arabic label', 'Document component', 'Kimi K3 / current component', 'Official per-language scores public; no rerun here', 'Not agent success; unpaired documents', 'Document parsing component', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'), (4, 'GAIA-v2-LILT', 'ar de hi ko pt-BR + en', 'Retrieval / tool execution', 'GPT-5.4, Gemini 3.1 Pro, Opus 4.6 / earlier flagships', 'Evaluation code, HF tasks, per-language table public', 'Audit changes function, culture, difficulty together; not pure language causality', 'Audited MAPS-GAIA branch', 'https://arxiv.org/html/2604.24929v1'), (5, 'MAPS', 'Final paper: 12 languages; older HF card: 11', 'GAIA/SWE/safety/math mixture', 'Historical models / context', 'Paper and HF data public; pin subset and version', 'MATH is not execution; GAIA shares ancestry with preceding row', 'Parent GAIA, SWE, MATH, ASB sets', 'https://aclanthology.org/2026.findings-eacl.42/'), (6, 'SEATauBench', 'vi th id Filipino zh + en', 'Multi-turn tools and final state', 'Kimi K2.5, GPT-5-mini, Qwen3 / historical', 'Code, domain data, analysis framework public; MIT repository', 'Simulator can also fail; no isolated agent-only language effect', 'Derived from tau2; shares framework with LILT tau', 'https://github.com/SEACrowd/SEATauBench'), (7, 'macOSWorld', 'ar ja ru zh + en', 'Interactive desktop', 'Claude 3.7 CUA, 2025 OpenAI CUA / earlier flagships', 'Tasks, environment, scoring code public; per-language paper results', 'Instructions and UI switch together; v4 revises some scores', 'Native macOS tasks', 'https://arxiv.org/html/2506.04135v4'), (8, 'X-WebAgentBench', '14 non-English settings; 11 target languages', 'Interactive web shopping', 'GPT-4o and others / historical', 'Code and data entry points public; paper results', 'Task Score is not success rate; derived from WebShop', 'WebShop derivative', 'https://github.com/WPENGxs/X-WebAgentBench'), (9, 'MASSIVE-Agents', '52 languages; 24 targets including related labels', 'Static function / argument matching', 'Nova Premier, Claude 3.5 and others / earlier flagships', 'Converted HF data CC BY 4.0; per-language paper tables', 'Not multi-turn execution; filtering changes samples/function coverage by language', 'MASSIVE derivative; BFCL scoring', 'https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id'), (10, 'Terminal-Bench-LILT', 'ar cs de es hi ja ko sr tr zh', 'Native multilingual code execution', 'GPT-5.5, Opus 4.8 and others / earlier flagships', 'Paper, aggregate scores, samples public; full tasks require contacting authors', 'Blog 300 versus leaderboard 324 tasks: denominators not mixed; full set not public', 'Native Terminal-Bench format; separate from community LILTBench', 'https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark'), (11, 'LILT multilingual tau', 'de ko + en', 'Multi-turn tools and final state', 'GPT-5.4, Opus 4.8, Gemini 3.1 Pro / earlier flagships', 'Public board and run notes; dynamic per-language scores not extracted here', 'Different English/target-language simulators confound scores', 'tau2 derivative; not independent framework replication', 'https://benchmarks.lilt.com/'), (12, 'LILTBench community', 'Authors report 31 languages; not exhaustively checked task by task', 'Native code / English pairs', 'Opus 4.6 / challenging test resources', 'Tasks, verifiers, leaderboard public; Apache-2.0 repository', 'Adversarial hard-task selection cannot estimate population failure rates', 'Community native tasks; distinct from commercial full set', 'https://github.com/lilt/liltbench-tasks-public'), (13, 'VoiceAgentBench', 'en hi bn mr ta te ml', 'Speech-to-tool-call scoring', 'SpeechLM / ASR+LLM / non-frontier resources', 'Evaluation code and HF data entry public; custom community license', 'Multi-turn subset is English only; predicted calls are not final environment success', 'Synthetic speech-tool tasks', 'https://github.com/ola-krutrim/VoiceAgentBench'), (14, 'TelcoAgent-Bench', 'ar + en', 'Diagnostic tool sequences / summaries', '3B–8B models / non-frontier resources', 'Blueprints, tasks, prediction/score JSON public; license unverified', 'Intent/resolution similarity is not strict success; cannot extrapolate to flagships', 'Telecom blueprint tasks', 'https://github.com/BrahiM-Mefgouda/TelcoAgent'), (15, 'MultiAgent-X', '12 languages; target hi sw ha am', 'Function-call data', 'No verified current-frontier baseline / resources', 'Samples, scripts, structure public; full license unverified', 'Synthetic; structural validity is not native quality; pa is not pnb', 'New synthetic tasks; quoted MASSIVE scores are not replication', 'https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark'), (16, 'MAST FIRE 2026', '21-language track union; 14 targets', 'Multi-turn retrieval / answers', 'Public Tongyi-30B baseline / non-frontier', 'Task/corpus entry points and per-language baseline public', 'September 30 final results are future at cutoff; baseline only; English documents and answers', 'BrowseComp-Plus derivative; three Indic languages overlap tracks', 'https://mast-benchmark.github.io/') ) SELECT * FROM reviewed_published_rows; - Broader evidence map · 1–15
Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "MASSIVE calls", "GAIA audited", "X-Web shopping", "macOS use", "SEA multi-turn", "LILT native code", "Voice tools", "MAST retrieval", "MultiAgent-X tasks", "Code", "Included projects", "Mapping limitation", "Meaning") AS ( VALUES ('Hindi', 2, 2, 2, 0, 0, 1, 1, 2, 1, 'hi', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Spanish', 2, 0, 2, 0, 0, 1, 0, 2, 0, 'es', 'MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Modern Standard Arabic†', 2, 2, 2, 2, 0, 1, 0, 2, 0, 'ar', 'MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval', 'Generic Arabic/Filipino are related labels, not exact MSA/Tagalog matches', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('French', 2, 0, 2, 0, 0, 0, 0, 2, 0, 'fr', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Bengali', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'bn', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Portuguese', 2, 2, 0, 0, 0, 0, 0, 0, 0, 'pt', 'MASSIVE calls; GAIA audited', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Indonesian', 2, 0, 0, 0, 2, 0, 0, 0, 0, 'id', 'MASSIVE calls; SEA multi-turn', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Urdu', 2, 0, 2, 0, 0, 0, 0, 2, 0, 'ur', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Russian', 2, 0, 2, 2, 0, 0, 0, 2, 0, 'ru', 'MASSIVE calls; X-Web shopping; macOS use; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('German', 2, 2, 2, 0, 0, 1, 0, 2, 0, 'de', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Japanese', 2, 0, 0, 2, 0, 1, 0, 0, 0, 'ja', 'MASSIVE calls; macOS use; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Nigerian Pidgin', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'pcm', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Egyptian Arabic', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'arz', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Marathi', 0, 0, 0, 0, 0, 0, 1, 0, 0, 'mr', 'Voice tools', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Vietnamese', 2, 0, 2, 0, 2, 0, 0, 0, 0, 'vi', 'MASSIVE calls; X-Web shopping; SEA multi-turn', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability') ) SELECT * FROM reviewed_published_rows; - Broader evidence map · 16–30
Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language", "MASSIVE calls", "GAIA audited", "X-Web shopping", "macOS use", "SEA multi-turn", "LILT native code", "Voice tools", "MAST retrieval", "MultiAgent-X tasks", "Code", "Included projects", "Mapping limitation", "Meaning") AS ( VALUES ('Telugu', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'te', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Swahili', 2, 0, 2, 0, 0, 0, 0, 2, 1, 'sw', 'MASSIVE calls; X-Web shopping; MAST retrieval; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Hausa', 0, 0, 0, 0, 0, 0, 0, 0, 1, 'ha', 'MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Turkish', 2, 0, 2, 0, 0, 1, 0, 0, 0, 'tr', 'MASSIVE calls; X-Web shopping; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Western Punjabi', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'pnb', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Tagalog†', 2, 0, 0, 0, 2, 0, 0, 0, 0, 'tl', 'MASSIVE calls; SEA multi-turn', 'Generic Arabic/Filipino are related labels, not exact MSA/Tagalog matches', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Tamil', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'ta', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Iranian Persian', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'fa', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Korean', 2, 2, 0, 0, 0, 1, 0, 0, 0, 'ko', 'MASSIVE calls; GAIA audited; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Amharic', 2, 0, 0, 0, 0, 0, 0, 0, 1, 'am', 'MASSIVE calls; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Thai', 2, 0, 2, 0, 2, 0, 0, 2, 0, 'th', 'MASSIVE calls; X-Web shopping; SEA multi-turn; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Javanese', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'jv', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Italian', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'it', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Gujarati', 0, 0, 0, 0, 0, 0, 0, 2, 0, 'gu', 'MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'), ('Kannada', 2, 0, 0, 0, 0, 0, 0, 2, 0, 'kn', 'MASSIVE calls; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability') ) SELECT * FROM reviewed_published_rows; - Language signals alongside counterevidence
Researcher synthesis, not risk score or statistical ranking. Shared task ancestry, historical models, components, and non-frontier baselines remain distinguished. Recommendations are unrun tests. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("Language group", "Supporting evidence", "Supported interpretation", "Counterevidence or limit", "Action") AS ( VALUES ('01 Thai', 'SEA retail and other domains; X-Web shopping; K3 documents', 'Weakness across tasks; current-frontier evidence is component-only', 'Thai is the highest one-shot language for Llama 3.1 405B in MASSIVE; models/prompts change order', 'Retest fully localized business and shopping; avoid all-model claims'), ('02 Arabic (generic)†', 'Audited GAIA: three flagships below English; macOS: two CUAs lower; K3 component', 'Aligned historical signals across two execution families', 'GAIA improves after audit; GUI also changes RTL layout; dialects untested', 'Retest search, RTL UI, and numeric arguments; do not extend to Egyptian Arabic'), ('03 Japanese', 'macOS CUA; native code tasks; K3 component', 'Component weakness; execution depends on model', 'OpenAI CUA exceeds English; Claude CUA falls below English', 'Separate document reading, click targeting, and display-width tests'), ('04 Hindi', 'Audited GAIA; native code; MASSIVE; K3 component', 'Several task families; low scores sensitive to task quality', 'GAIA audit gains 25.4–32.7 pp; English instructions do not remove all native-task difficulty in Hindi/German', 'Audit task alignment, entities, and regional rules; distinguish reading from execution'), ('05 Russian', 'Current RuBench execution and GorillaHard plans; historical macOS', 'Concrete current failures, more direct than a component', 'RuBench/MERA lack English pairs; historical OpenAI CUA Russian exceeds English', 'Regress specific failures; do not rank Russian worst overall'), ('06 Vietnamese / Indonesian / Filipino†', 'SEA full business; MASSIVE; Vietnamese X-Web', 'Several tasks; language ordering varies by domain', 'Kimi Vietnamese retail slightly exceeds English; Filipino simulator also errs', 'Accept by business domain and verify Filipino–Tagalog variety alignment'), ('07 German / Korean / Portuguese', 'Audited GAIA; MASSIVE; K3; German/Korean LILT', 'More evidence does not establish deployment readiness', 'Audited German gap is 3.1 pp for GPT-5.4, 12.7 pp for Opus 4.6', 'Test the target model and actual workflow; no universal safe-language list'), ('08 Amharic / Kannada', 'Historical MASSIVE zero-shot AST; Amharic synthetic tasks', 'Risk signal mainly from one historical task family', 'Lowest language differs by model; no current-frontier execution replication', 'Prioritize function arguments; evidence cannot rank current models'), ('09 Swahili / Urdu', 'X-Web; MASSIVE; non-frontier MAST retrieval', 'Task- and model-dependent signals; no uniform direction', 'Swahili original-language GPT-4o shopping is not low; speech and retrieval scores are not interchangeable', 'Measure retrieval, calls, and final state separately with English controls'), ('10 Other coverage / gaps', 'Bengali/Tamil/Telugu tool or retrieval resources; Pidgin speech component', 'Consult all 30 rows; missing evidence is not low ability', 'Egyptian Arabic/Western Punjabi exact matches missing; Hausa lacks current-frontier baseline', 'Use native tasks: pcm≠en, arz≠ar, pnb≠pa, jv≠id') ) SELECT * FROM reviewed_published_rows; - All 30 languages: evidence and next checks
Researcher mapping retains all target languages, named evidence and inference limits. No invented scores; related labels do not replace exact language varieties. SQLite VALUES reproduces reviewed observations; not a live source query.
SQL query
WITH reviewed_published_rows ("No.", "Language", "Code", "Available evidence", "Interpretation and next check") AS ( VALUES (1, 'Hindi', 'hi', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks', 'Audited GAIA recovers substantially; K3 reading is low. Separate benchmark quality from local-rule execution.'), (2, 'Spanish', 'es', 'MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval', 'GAIA parent set, shopping, function calls, and retrieval coverage. No complete current-frontier comparison; do not presume European-language reliability.'), (3, 'Modern Standard Arabic', 'ar', 'MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval', 'GAIA and macOS give aligned historical signals; K3 component score is lower. Generic Arabic does not precisely establish MSA or Egyptian dialect performance.'), (4, 'French', 'fr', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Shopping, function calls, retrieval baselines, and a current document component. Retest full business execution on the target model.'), (5, 'Bengali', 'bn', 'MASSIVE calls; Voice tools; MAST retrieval', 'Historical calls, voice-tool tasks, and retrieval resources. Current-frontier execution remains unestablished; test entities and arguments.'), (6, 'Portuguese', 'pt', 'MASSIVE calls; GAIA audited', 'Audited GAIA uses Brazilian Portuguese; MASSIVE uses Portugal locale. Regional business rules are not interchangeable.'), (7, 'Indonesian', 'id', 'MASSIVE calls; SEA multi-turn', 'SEA covers full multi-turn workflows with domain-dependent effects. Indonesian does not substitute for Javanese.'), (8, 'Urdu', 'ur', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Historical/non-frontier calls, shopping, and retrieval. Numerals and dates are proposed mechanisms, not measured current-model failure rates.'), (9, 'Russian', 'ru', 'MASSIVE calls; X-Web shopping; macOS use; MAST retrieval', 'Current code execution and static calls contain failures. Without matched English tasks, Russian causation is unestablished.'), (10, 'German', 'de', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval', 'Audited GAIA and native code coverage are relatively rich. Gaps differ by model; component strength is not deployment acceptance.'), (11, 'Japanese', 'ja', 'MASSIVE calls; macOS use; LILT native code', 'Current parsing is low; historical CUA gaps change direction by model. Test visual targeting, text, and local software rules separately.'), (12, 'Nigerian Pidgin', 'pcm', 'SwitchBoard speech component', 'Code-switching speech resources found; current-frontier tool execution unverified. Do not substitute English.'), (13, 'Egyptian Arabic', 'arz', 'No matching execution result included', 'Generic Arabic studies cannot fill this dialect. No included exact-dialect agent result.'), (14, 'Marathi', 'mr', 'Voice tools', 'VoiceAgentBench provides speech-tool tasks; do not extend its English-only multi-turn results to Marathi.'), (15, 'Vietnamese', 'vi', 'MASSIVE calls; X-Web shopping; SEA multi-turn', 'SEA, shopping, function calls, and documents provide multiple sources. Kimi fully localized retail slightly exceeds English: preserve the counterexample.'), (16, 'Telugu', 'te', 'MASSIVE calls; Voice tools; MAST retrieval', 'Calls, voice tools, and retrieval resources are reusable. Static calls do not establish long-horizon execution.'), (17, 'Swahili', 'sw', 'MASSIVE calls; X-Web shopping; MAST retrieval; MultiAgent-X tasks', 'Historical task coverage, but GPT-4o original-language shopping is not weak. MAST small-model scores cannot rank frontier risk.'), (18, 'Hausa', 'ha', 'MultiAgent-X tasks', 'MultiAgent-X supplies synthetic calls. Exploratory safety logs were excluded from reliable risk assessment; no current-frontier baseline.'), (19, 'Turkish', 'tr', 'MASSIVE calls; X-Web shopping; LILT native code', 'Calls, shopping, and native code tasks exist. Add current-model end-to-end and local-rule tests.'), (20, 'Western Punjabi', 'pnb', 'No matching execution result included', 'MultiAgent-X and MAST pa/Gurmukhi cannot substitute for Western Punjabi; exact matching results remain missing.'), (21, 'Tagalog', 'tl', 'MASSIVE calls; SEA multi-turn', 'MASSIVE tl-PH and SEA Filipino are related labels. Confirm the user variety, then test multi-turn business execution.'), (22, 'Tamil', 'ta', 'MASSIVE calls; Voice tools; MAST retrieval', 'Calls, voice-tool tasks, and MAST retrieval are reusable. Current-frontier complete execution remains unestablished.'), (23, 'Iranian Persian', 'fa', 'MASSIVE calls', 'MASSIVE fa-IR provides historical calls. Current models still need Persian numerals, calendars, and identifier-preservation tests.'), (24, 'Korean', 'ko', 'MASSIVE calls; GAIA audited; LILT native code', 'Audited GAIA improves and K3 parsing is relatively high. LILT multi-turn results have simulator confounds; broad reliability is unestablished.'), (25, 'Amharic', 'am', 'MASSIVE calls; MultiAgent-X tasks', 'MASSIVE historical weakness is specific; MultiAgent-X supplies resources, not independent performance replication. Retest current flagships.'), (26, 'Thai', 'th', 'MASSIVE calls; X-Web shopping; SEA multi-turn; MAST retrieval', 'Historical SEA and shopping weaknesses plus low K3 parsing motivate retesting. They do not imply every model/task is weak.'), (27, 'Javanese', 'jv', 'MASSIVE calls', 'MASSIVE jv-ID provides historical calls. The missing-evidence claim concerns current execution; Indonesian is not a substitute.'), (28, 'Italian', 'it', 'MASSIVE calls', 'MASSIVE, MAPS parent coverage, and current documents exist. Highest parsing does not imply the most reliable complete agent.'), (29, 'Gujarati', 'gu', 'MAST retrieval', 'MAST supplies retrieval tasks and a non-frontier baseline. Current-frontier tools/business execution evidence remains insufficient.'), (30, 'Kannada', 'kn', 'MASSIVE calls; MAST retrieval', 'Claude 3.5 is historically low on MASSIVE; MAST adds retrieval resources. No current-frontier multi-task replication.') ) SELECT * FROM reviewed_published_rows; - sea_code
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- gaia_code
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- maps
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- maps_data
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- mac_code
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- xweb_code
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- massive_data
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- voice
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- voice_code
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- telco
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- telco_code
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- terminal
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- lilt
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- liltbench
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- multi
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- mast
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- judge
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- switch
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- GAIA-v2-LILT Table 2
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- SEATauBench Tables 9/10/13
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- macOSWorld v4 Table 3 and cases
Primary source supporting adjacent evidence; reviewed September 5, 2026.
- X-WebAgentBench Table 2
Primary source supporting adjacent evidence; reviewed September 5, 2026.