{
  "surface": "report",
  "manifest": {
    "version": 1,
    "surface": "report",
    "title": "Do LLM Agents Work Equally Well Across Languages?",
    "generatedAt": "2026-09-05",
    "description": "A visual evidence review of 30 languages, 16 projects, and 15 charts; current frontier models separated from earlier flagships and test resources.",
    "blocks": [
      {
        "id": "text-0",
        "type": "markdown",
        "body": "# Do LLM Agents Work Equally Well Across Languages?"
      },
      {
        "id": "text-1",
        "type": "markdown",
        "body": "## Executive Summary\n\n**No. Published evaluations show that the same LLM agent can perform materially differently across languages.** The practical conclusion is to evaluate **model × language × task × localization setting**. English performance alone does not establish reliability in another language. The evidence supports unequal performance; it does not support a permanent ranking of languages.\n\n- **Arabic: repeated gaps in research and desktop tasks.** In audited, tool-assisted GAIA tasks, Arabic trails English by **19.4–30.3 percentage points** across three earlier flagships. In macOSWorld, both specialized computer-use agents score lower in Arabic when instructions and the operating-system interface switch together. These results concern generic Arabic labels; they do not establish Egyptian Arabic performance. [GAIA-v2-LILT](https://arxiv.org/html/2604.24929v1#S6), [macOSWorld v4](https://arxiv.org/html/2506.04135v4#S5).\n- **Thai: the risk becomes clearer when the whole business workflow is localized.** Kimi K2.5 scores **56.1% in English, 57.3% with Thai dialogue only, and 32.7% with the full Thai retail workflow localized**. Policies, tools, and database content matter beyond conversational fluency. Earlier web-shopping results also flag Thai. [SEATauBench](https://arxiv.org/html/2606.28715v1#A6), [X-WebAgentBench](https://arxiv.org/html/2505.15372v1#S3).\n- **Japanese and Hindi require narrower conclusions.** They are among Kimi K3's lower document-parsing observations, but Japanese desktop results change direction by agent, and Hindi GAIA scores rise substantially after benchmark auditing. Weakness depends on the task and evaluation design. [Kimi K3 results](https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e).\n\n**For the newest frontier models, the 30-language evidence is still incomplete.** Most controlled multilingual execution results come from earlier flagships. Current-model failures are useful deployment signals, but a failure observed in Russian, for example, does not prove that Russian caused it without an English control. This review covers 16 benchmark/resource projects and preserves counterexamples, missing evidence, and openness limits. Evidence cutoff: **September 5, 2026**; no new model evaluations were run."
      },
      {
        "id": "text-2",
        "type": "markdown",
        "body": "## Auditing the benchmark narrows language gaps without eliminating them\n\n**Arabic remains below each model's English reference after the GAIA audit.** The first chart shows English minus audited-language pass@1: higher bars mean a larger remaining gap. German, Hindi, Korean, and Brazilian Portuguese have smaller gaps, with different magnitudes by model."
      },
      {
        "id": "chart-gaia_residual-3",
        "type": "chart",
        "chartId": "gaia_residual",
        "layout": "full"
      },
      {
        "id": "text-4",
        "type": "markdown",
        "body": "**Every audited language improves for every evaluated model.** The second chart shows audited minus minimally translated scores. Audits modify answer alignment, cultural context, and difficulty together: neither the entire improvement nor the entire residual is a pure language effect. [Table 2](https://arxiv.org/html/2604.24929v1#S6)."
      },
      {
        "id": "chart-gaia_lift-5",
        "type": "chart",
        "chartId": "gaia_lift",
        "layout": "full"
      },
      {
        "id": "text-6",
        "type": "markdown",
        "body": "Both charts use the same experiment, not independent replications. MAPS-GAIA and its audited version also share task ancestry. These are GPT-5.4, Gemini 3.1 Pro, and Opus 4.6 results, with 165 tasks per language in an actual tool-assisted workflow. They do not measure today's models directly."
      },
      {
        "id": "text-7",
        "type": "markdown",
        "body": "## Conversational fluency does not establish local business competence\n\n**Thai retail performance falls when localization extends beyond dialogue.** The chart compares an English baseline, localized dialogue, and fully localized business content for the same earlier flagship. Vietnamese full-localization retail performance slightly exceeds English, an important counterexample to a universal decline."
      },
      {
        "id": "chart-sea_localization-8",
        "type": "chart",
        "chartId": "sea_localization",
        "layout": "full"
      },
      {
        "id": "text-9",
        "type": "markdown",
        "body": "Thai fully localized airline and telecom results also fall below their English references; those observations remain in the chart data. The user simulator changes language too and can make errors, so the measured decline belongs to the overall system. It cannot all be assigned to the evaluated agent. Results use three trials per task; domains are not pooled. [SEATauBench Tables 9, 10, and 13](https://arxiv.org/html/2606.28715v1#A6)."
      },
      {
        "id": "text-10",
        "type": "markdown",
        "body": "## Desktop results repeat the Arabic signal and complicate the Japanese story\n\n**Both specialized computer-use agents score lower in Arabic than in English.** Japanese and Russian move in opposite directions across agents: above English for OpenAI CUA, below English for Claude CUA. Compare each agent with its own English bar."
      },
      {
        "id": "chart-macos_language-11",
        "type": "chart",
        "chartId": "macos_language",
        "layout": "full"
      },
      {
        "id": "text-12",
        "type": "markdown",
        "body": "These 2025 model versions run 171 comparable tasks. Instructions and the OS interface change together, combining reading, layout, and planning effects. The authors document Arabic interface localization failures, including positioning errors. This motivates targeted interface testing, without establishing that right-to-left text alone causes the gap. [macOSWorld v4 Table 3 and cases](https://arxiv.org/html/2506.04135v4#S5)."
      },
      {
        "id": "text-13",
        "type": "markdown",
        "body": "## Translating the workflow into English can make performance worse\n\n**Thai is the lowest original-language shopping score among the 11 target languages shown.** Translating to English helps French but hurts several others, particularly Arabic and Urdu. The y-values are WebShop task scores, not binary purchase-success percentages."
      },
      {
        "id": "chart-xweb_translation-14",
        "type": "chart",
        "chartId": "xweb_translation",
        "layout": "full"
      },
      {
        "id": "text-15",
        "type": "markdown",
        "body": "This earlier GPT-4o experiment adds a different task family to the Thai concern. It cannot be averaged with SEATau into a language risk score. Swahili's original-language result is comparatively strong, contradicting a simple rule that lower-resource languages always perform worst. The figure compares two published strategies, not every strategy in the paper. [X-WebAgentBench Table 2](https://arxiv.org/html/2505.15372v1#S3)."
      },
      {
        "id": "text-16",
        "type": "markdown",
        "body": "## Corroboration should include disagreement and missing controls\n\n**More sources should constrain the conclusion as well as support it.** A paper, its GitHub repository, and its leaderboard form one evidence chain. An audited derivative is not a fresh independent task sample. The following table is a reading guide, not a ranking."
      },
      {
        "id": "table-triangulation-17",
        "type": "table",
        "tableId": "triangulation",
        "layout": "full"
      },
      {
        "id": "text-18",
        "type": "markdown",
        "body": "Amharic and Kannada warrant additional testing, with a narrower basis: in MASSIVE-Agents' 10k, zero-shot AST setting, Amharic is Nova Premier's lowest language and Kannada is Claude 3.5 v2 Sonnet's lowest. These are historical static function-call results. A later synthetic dataset does not independently replicate those performance findings. [MASSIVE-Agents Table 2](https://aclanthology.org/2025.findings-emnlp.1099.pdf)."
      },
      {
        "id": "text-19",
        "type": "markdown",
        "body": "## The broader evidence map distinguishes results from test resources\n\n**Nine named benchmark columns cover 27 of the 30 target languages in some form.** A value of 2 means a verified per-language result in this broader historical/resource layer; 1 means task or aggregate coverage. Neither is an ability score. A dagger marks related Arabic or Filipino labels rather than an exact target-variety match."
      },
      {
        "id": "chart-broad_atlas_1-20",
        "type": "chart",
        "chartId": "broad_atlas_1",
        "layout": "full"
      },
      {
        "id": "text-21",
        "type": "markdown",
        "body": "The first half combines static calls, execution tasks, and reusable resources. Dense coverage means more ways to investigate a language, not evidence that today's frontier models have passed them."
      },
      {
        "id": "chart-broad_atlas_2-22",
        "type": "chart",
        "chartId": "broad_atlas_2",
        "layout": "full"
      },
      {
        "id": "text-23",
        "type": "markdown",
        "body": "Javanese has historical MASSIVE function-call evidence; Marathi has voice-tool resources; Gujarati has MAST retrieval coverage. Hausa's synthetic tasks do not establish current-model performance. Nigerian Pidgin has a separate [code-switching speech component resource](https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch), outside this agent-results grid. Egyptian Arabic and Western Punjabi still lack an included exact-match result."
      },
      {
        "id": "text-24",
        "type": "markdown",
        "body": "## What has actually failed on current models?\n\nThe following evidence retains the review's current-model observations. It can identify execution failures and component weaknesses. Without a matched language control, it cannot identify how much of a failure is caused by the language."
      },
      {
        "id": "text-25",
        "type": "markdown",
        "body": "## Valid formatting can hide substantial task errors\n\n**Sol passes formatting on 99.9% of Russian tool-plan samples but passes the full sample on 69.1%.** Each bar splits all samples into three mutually exclusive parts. The middle segment shows errors that a JSON-format check would miss."
      },
      {
        "id": "chart-format_partition-26",
        "type": "chart",
        "chartId": "format_partition",
        "layout": "full"
      },
      {
        "id": "text-27",
        "type": "markdown",
        "body": "GorillaHard grades static outputs covering calls, abstention, and clarification; it does not execute tools. The documented dataset has 1,169 items, while submission-specific evaluated counts and repetitions are not separately disclosed. Model settings are not fully aligned, so this chart does not select a model winner. [Sol](https://mera.a-ai.ru/en/text/submits/2.0/8), [Opus 5](https://mera.a-ai.ru/en/text/submits/2.0/1), [Grok 4.6](https://mera.a-ai.ru/en/text/submits/2.0/11)."
      },
      {
        "id": "text-28",
        "type": "markdown",
        "body": "## Correct tool selection can still produce incorrect arguments\n\n**About 8.1%–16.3% of call-required samples select matching tools without matching all arguments.** The middle segment points to concrete acceptance checks: entities, amounts, dates, and argument transfer between steps."
      },
      {
        "id": "chart-argument_partition-29",
        "type": "chart",
        "chartId": "argument_partition",
        "layout": "full"
      },
      {
        "id": "text-30",
        "type": "markdown",
        "body": "This denominator contains only call-required items, unlike the previous figure. The nested metrics share a denominator within this chart, enabling subtraction; the two figures cannot be joined into a funnel. No English comparison establishes a Russian penalty. [Scoring implementation](https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/utils.py)."
      },
      {
        "id": "text-31",
        "type": "markdown",
        "body": "## Knowing when to stop needs its own acceptance test\n\n**On abstention-required items, roughly 33.9%–41.7% lack the correct abstention.** False abstention on call-required items is much rarer. The chart separates these conditional error rates; the denominators appear in the category labels."
      },
      {
        "id": "chart-decision_errors-32",
        "type": "chart",
        "chartId": "decision_errors",
        "layout": "full"
      },
      {
        "id": "text-33",
        "type": "markdown",
        "body": "These rates cannot be added, and a static output error is not an observed unauthorized real-world action. Abstention conditions follow task rules and are not all safety-policy refusals. [GorillaHard definition](https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md)."
      },
      {
        "id": "text-34",
        "type": "markdown",
        "body": "## Real code execution reveals reproducible task failures\n\nRuBench asks agents to modify real repositories from Russian issue instructions, then runs regression tests. The chart separates **audit-retained passes, raw passes removed for answer exposure, and original failures**."
      },
      {
        "id": "chart-execution_partition-35",
        "type": "chart",
        "chartId": "execution_partition",
        "layout": "full"
      },
      {
        "id": "text-36",
        "type": "markdown",
        "body": "Opus 5 fully passes **0/3 runs on each of two HTTP fixes**: method/header case rules and an empty-filename parsing exception. These are concrete task failures, without evidence that Russian caused them. Sol and Opus use different task sets and harnesses. Audit deductions re-score existing network-enabled runs; they are not offline reruns. [Pinned run record](https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md)."
      },
      {
        "id": "text-37",
        "type": "markdown",
        "body": "## Four document languages deserve targeted end-to-end checks\n\n**Thai, Japanese, generic Arabic, and Hindi sit at the lower end of Kimi K3's 13 included language observations.** The median line locates observations within this sample; it is not a deployment threshold."
      },
      {
        "id": "chart-document_profile-38",
        "type": "chart",
        "chartId": "document_profile",
        "layout": "full"
      },
      {
        "id": "text-39",
        "type": "markdown",
        "body": "These scores combine text, layout, formula, and table parsing. **They are not agent success rates.** Language documents are not paired for identical content, and Arabic varieties are not separated. Validate the complete sequence: read a field, call the API, and inspect the resulting state. [Pinned author results](https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e).\n\nThe script groups overlap substantially, and Korean scores **89.9**. A blanket non-Latin-script weakness label would conceal this counterexample."
      },
      {
        "id": "chart-script_distribution-40",
        "type": "chart",
        "chartId": "script_distribution",
        "layout": "full"
      },
      {
        "id": "text-41",
        "type": "markdown",
        "body": "Boxes show the middle half of language observations; whiskers show minima and maxima. This summarizes 13 observations, not confidence intervals or the causal effect of a writing system. [Parsing metric definitions](https://arxiv.org/html/2603.28130v1#S3.SS4)."
      },
      {
        "id": "text-42",
        "type": "markdown",
        "body": "## Evidence for the current frontier is still sparse\n\n**These two maps include only this review's September current-model cohort.** The broader historical results and resources above are excluded. A value of 1 marks included evidence, 0.5 marks the related Arabic label, and 0 means no included result was found."
      },
      {
        "id": "chart-atlas_1-43",
        "type": "chart",
        "chartId": "atlas_1",
        "layout": "full"
      },
      {
        "id": "text-44",
        "type": "markdown",
        "body": "In the first half, Russian has code execution and static tool-plan results. French, Spanish, and other languages have document-component results, which do not establish browser or multi-turn business reliability."
      },
      {
        "id": "chart-atlas_2-45",
        "type": "chart",
        "chartId": "atlas_2",
        "layout": "full"
      },
      {
        "id": "text-46",
        "type": "markdown",
        "body": "In the second half, Korean, Thai, and Italian have document-component results. Historical Javanese function-call evidence does not supply a current execution guarantee. Gurmukhi Punjabi cannot substitute for Western Punjabi. No current-model matched-language experiment was included in the paired-comparison column."
      },
      {
        "id": "text-47",
        "type": "markdown",
        "body": "## Turn suspected mechanisms into verifiable tasks\n\nUrdu and Iranian Persian offer concrete test conditions involving numerals, dates, and mixed-script identifiers. These are proposed fixtures, not measured current-agent failure rates. [Unicode number conventions](https://www.unicode.org/reports/tr35/tr35-numbers.html)."
      },
      {
        "id": "table-action_map-48",
        "type": "table",
        "tableId": "action_map",
        "layout": "full"
      },
      {
        "id": "text-49",
        "type": "markdown",
        "body": "**The most informative next experiment holds model, tools, budget, and business objective fixed, changes the language, and checks the final state.** Add locally authored tasks alongside those controls: matched tasks help estimate language gaps; native tasks test practical usefulness."
      },
      {
        "id": "text-50",
        "type": "markdown",
        "body": "## Scope, metrics, and sources\n\nThis review checks published papers, author repositories, and dataset cards without running models. Its 16 core benchmark/resource entries are not an exhaustive census of public research. The fixed 30-language scope follows total-speaker coverage after excluding English and Chinese varieties, using a [pinned public transcription](https://en.wikipedia.org/w/index.php?title=List_of_languages_by_total_number_of_speakers&oldid=1370552712). Population determines inclusion, not training resources or ability.\n\nThe September 5 frontier scope includes Astra, Sol, Fable 5.1/5, Opus 5, Muse Spark 1.3, Grok 4.6, Kimi K3, GLM 5.3, and near-boundary Gemini 3.8 Flash/Qwen3.8. Per-language usable results exist for only a subset. Earlier flagships are labeled separately throughout. [Cohort reference](https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2)."
      },
      {
        "id": "table-methods-51",
        "type": "table",
        "tableId": "methods",
        "layout": "full"
      },
      {
        "id": "text-52",
        "type": "markdown",
        "body": "Model-only aggregate scores from MCP Atlas and a tiny German evaluation without original traces do not fill per-language gaps here. “Not found” refers to this search and its inclusion rules. The lookup below keeps all 30 target languages and identifies useful evidence and next checks."
      },
      {
        "id": "table-language_lookup-53",
        "type": "table",
        "tableId": "language_lookup",
        "layout": "full"
      },
      {
        "id": "text-54",
        "type": "markdown",
        "body": "## Public evidence has several different levels of openness\n\n**A public paper, an open task set, evaluation code, and raw runs are different assets.** Terminal-Bench-LILT exposes samples but requires contacting the authors for the full set; the community LILTBench collection is separate. MASSIVE-Agents' converted dataset is publicly available, superseding an earlier “not released” characterization."
      },
      {
        "id": "table-evidence_inventory-55",
        "type": "table",
        "tableId": "evidence_inventory",
        "layout": "full"
      },
      {
        "id": "text-56",
        "type": "markdown",
        "body": "VoiceAgentBench and TelcoAgent add task coverage, without current-frontier results. MAST currently provides a non-frontier baseline; its scheduled final competition results are still in the future at the review cutoff. These entries expand test resources, not the flagship-performance charts."
      },
      {
        "id": "text-57",
        "type": "markdown",
        "body": "## What to test next\n\n**Start with languages that have repeated task-specific signals, then fill commercially important blind spots.** Test Thai dialogue-only, tool, and full-business localization separately. Separate standard Arabic, dialects, and right-to-left interfaces. Split Japanese reading, visual targeting, and execution. Audit Hindi answer keys and regional rules before interpreting low scores. Prioritize other languages by local usage and error cost, not by how little research exists.\n\nFix the exact model, tool versions, task budget, and acceptance criteria. Retain every tool call and final state; for speech, retain the original audio and a correct-transcript control. The open question is which current-model failures persist after separating hearing, reading, tool selection, argument formation, and execution. None of these proposed follow-up tests has been run for this review."
      }
    ],
    "charts": [
      {
        "id": "format_partition",
        "title": "All samples: format and content outcomes",
        "subtitle": "Each bar = 100%; documented 1,169 items including calls, abstention, clarification",
        "description": "Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed.",
        "showDescription": true,
        "type": "horizontalStackedBar100",
        "dataset": "format_partition",
        "sourceId": "format_partition",
        "intent": "composition",
        "layout": "full",
        "question": "All samples: format and content outcomes",
        "rationale": "Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed.",
        "encodings": {
          "x": {
            "field": "Model",
            "label": "Model"
          },
          "y": {
            "fields": [
              "Full sample passed",
              "Valid format, sample failed",
              "Invalid format"
            ],
            "format": "percent",
            "label": "Share"
          }
        },
        "valueFormat": "percent",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        }
      },
      {
        "id": "argument_partition",
        "title": "Call-required items: tools and arguments",
        "subtitle": "Each bar = 100%; documented 1,032 call-required items",
        "description": "Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference.",
        "showDescription": true,
        "type": "horizontalStackedBar100",
        "dataset": "argument_partition",
        "sourceId": "argument_partition",
        "intent": "composition",
        "layout": "full",
        "question": "Call-required items: tools and arguments",
        "rationale": "Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference.",
        "encodings": {
          "x": {
            "field": "Model",
            "label": "Model"
          },
          "y": {
            "fields": [
              "Tools and arguments matched",
              "Tools matched, arguments incomplete",
              "Tools not fully matched"
            ],
            "format": "percent",
            "label": "Share"
          }
        },
        "valueFormat": "percent",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        }
      },
      {
        "id": "decision_errors",
        "title": "Abstention errors under two conditions",
        "subtitle": "Left: required abstention missing; right: false abstention on call items",
        "description": "1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions.",
        "showDescription": true,
        "type": "bar",
        "dataset": "decision_errors",
        "sourceId": "decision_errors",
        "intent": "comparison",
        "layout": "full",
        "question": "Abstention errors under two conditions",
        "rationale": "1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions.",
        "encodings": {
          "x": {
            "field": "Required action",
            "label": "Required action"
          },
          "y": {
            "field": "Conditional error rate",
            "format": "percent",
            "label": "Conditional error rate"
          },
          "color": {
            "field": "Model",
            "label": "Model"
          }
        },
        "valueFormat": "percent",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "options": {
          "orientation": "vertical",
          "grouping": "grouped"
        }
      },
      {
        "id": "execution_partition",
        "title": "Russian code tasks: audited run outcomes",
        "subtitle": "Sol: 49 + 8 + 12 = 69 runs; Opus 5: 46 + 1 + 13 = 60",
        "description": "Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control.",
        "showDescription": true,
        "type": "horizontalStackedBar100",
        "dataset": "execution_partition",
        "sourceId": "execution_partition",
        "intent": "composition",
        "layout": "full",
        "question": "Russian code tasks: audited run outcomes",
        "rationale": "Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control.",
        "encodings": {
          "x": {
            "field": "Configuration",
            "label": "Configuration"
          },
          "y": {
            "fields": [
              "Audit-retained passes",
              "Raw passes removed by audit",
              "Original failures"
            ],
            "format": "percent",
            "label": "Share"
          }
        },
        "valueFormat": "percent",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        }
      },
      {
        "id": "document_profile",
        "title": "Document parsing across 13 language observations",
        "subtitle": "Kimi K3; quality score 0–100; median 82.4; observed range 20.6 points",
        "description": "Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold.",
        "showDescription": true,
        "type": "bar",
        "dataset": "document_profile",
        "sourceId": "document_profile",
        "intent": "comparison",
        "layout": "full",
        "question": "Document parsing across 13 language observations",
        "rationale": "Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "field": "Parsing quality score",
            "format": "number",
            "label": "Parsing quality score"
          },
          "color": {
            "field": "Observed position",
            "label": "Observed position"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "options": {
          "orientation": "horizontal",
          "grouping": "grouped"
        },
        "referenceLines": [
          {
            "axis": "y",
            "value": 82.4,
            "label": "Median 82.4",
            "color": "neutral",
            "lineStyle": "dashed"
          }
        ],
        "maxRows": 13
      },
      {
        "id": "script_distribution",
        "title": "Parsing score distributions by script group",
        "subtitle": "Latin script: 7 observations; other scripts: 6; Korean scores 89.9",
        "description": "Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges.",
        "showDescription": true,
        "type": "boxPlot",
        "dataset": "script_distribution",
        "sourceId": "script_distribution",
        "intent": "distribution",
        "layout": "full",
        "question": "Parsing score distributions by script group",
        "rationale": "Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges.",
        "encodings": {
          "x": {
            "field": "Script group",
            "label": "Script group"
          },
          "y": {
            "fields": [
              "Minimum",
              "First quartile",
              "Median",
              "Third quartile",
              "Maximum"
            ],
            "format": "number",
            "label": "Parsing quality score"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "none"
        },
        "labels": {
          "values": "none"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        }
      },
      {
        "id": "atlas_1",
        "title": "Current-model evidence map · 1–15",
        "subtitle": "1 = included match; 0.5 = related label; 0 = no included result",
        "description": "Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related.",
        "showDescription": true,
        "type": "heatmap",
        "dataset": "atlas_1",
        "sourceId": "atlas_1",
        "intent": "relationship",
        "layout": "full",
        "question": "Current-model evidence map · 1–15",
        "rationale": "Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "fields": [
              "Document component",
              "Static tool plan",
              "End-to-end execution",
              "Matched language control"
            ],
            "format": "number",
            "label": "Evidence marker"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "none"
        },
        "labels": {
          "values": "none"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "sequential",
          "name": "blue"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "maxRows": 15
      },
      {
        "id": "atlas_2",
        "title": "Current-model evidence map · 16–30",
        "subtitle": "1 = included match; 0.5 = related label; 0 = no included result",
        "description": "Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages.",
        "showDescription": true,
        "type": "heatmap",
        "dataset": "atlas_2",
        "sourceId": "atlas_2",
        "intent": "relationship",
        "layout": "full",
        "question": "Current-model evidence map · 16–30",
        "rationale": "Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "fields": [
              "Document component",
              "Static tool plan",
              "End-to-end execution",
              "Matched language control"
            ],
            "format": "number",
            "label": "Evidence marker"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "none"
        },
        "labels": {
          "values": "none"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "sequential",
          "name": "blue"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "maxRows": 15
      },
      {
        "id": "gaia_lift",
        "title": "Score improvement after task auditing",
        "subtitle": "April 2026 flagships; 165 tasks per language; three models",
        "description": "GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim.",
        "showDescription": true,
        "type": "bar",
        "dataset": "gaia_lift",
        "sourceId": "gaia_lift",
        "intent": "comparison",
        "layout": "full",
        "question": "Score improvement after task auditing",
        "rationale": "GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "field": "Audit improvement (pp)",
            "label": "Audit improvement (pp)",
            "format": "number"
          },
          "color": {
            "field": "Model",
            "label": "Model"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "options": {
          "orientation": "vertical",
          "grouping": "grouped"
        },
        "maxRows": 40
      },
      {
        "id": "gaia_residual",
        "title": "Remaining gap from each model’s English reference",
        "subtitle": "April 2026 flagships; 165 tasks per language; three models",
        "description": "GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals.",
        "showDescription": true,
        "type": "bar",
        "dataset": "gaia_residual",
        "sourceId": "gaia_residual",
        "intent": "comparison",
        "layout": "full",
        "question": "Remaining gap from each model’s English reference",
        "rationale": "GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "field": "English minus audited (pp)",
            "label": "English minus audited (pp)",
            "format": "number"
          },
          "color": {
            "field": "Model",
            "label": "Model"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "options": {
          "orientation": "vertical",
          "grouping": "grouped"
        },
        "maxRows": 40
      },
      {
        "id": "sea_localization",
        "title": "Retail success by extent of localization",
        "subtitle": "Kimi K2.5; earlier flagship; three trials per task; repeated English reference",
        "description": "SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English.",
        "showDescription": true,
        "type": "bar",
        "dataset": "sea_localization",
        "sourceId": "sea_localization",
        "intent": "comparison",
        "layout": "full",
        "question": "Retail success by extent of localization",
        "rationale": "SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "field": "Task success rate",
            "label": "Task success rate",
            "format": "percent"
          },
          "color": {
            "field": "Setting",
            "label": "Setting"
          }
        },
        "valueFormat": "percent",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "options": {
          "orientation": "vertical",
          "grouping": "grouped"
        },
        "maxRows": 40
      },
      {
        "id": "macos_language",
        "title": "Desktop task success by agent and language",
        "subtitle": "2025 specialized CUAs; 171 tasks per language; instructions and OS switch together",
        "description": "macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs.",
        "showDescription": true,
        "type": "bar",
        "dataset": "macos_language",
        "sourceId": "macos_language",
        "intent": "comparison",
        "layout": "full",
        "question": "Desktop task success by agent and language",
        "rationale": "macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "field": "Task success rate",
            "label": "Task success rate",
            "format": "percent"
          },
          "color": {
            "field": "Agent",
            "label": "Agent"
          }
        },
        "valueFormat": "percent",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "options": {
          "orientation": "vertical",
          "grouping": "grouped"
        },
        "maxRows": 40
      },
      {
        "id": "xweb_translation",
        "title": "Web shopping scores: original language versus English translation",
        "subtitle": "GPT-4o; 2025 study; 200 instructions per language; 11 target languages",
        "description": "X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics.",
        "showDescription": true,
        "type": "bar",
        "dataset": "xweb_translation",
        "sourceId": "xweb_translation",
        "intent": "comparison",
        "layout": "full",
        "question": "Web shopping scores: original language versus English translation",
        "rationale": "X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "field": "Task score",
            "label": "Task score",
            "format": "number"
          },
          "color": {
            "field": "Strategy",
            "label": "Strategy"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "all"
        },
        "labels": {
          "values": "all"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "categorical",
          "name": "blue-orange"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "options": {
          "orientation": "horizontal",
          "grouping": "grouped"
        },
        "maxRows": 40
      },
      {
        "id": "broad_atlas_1",
        "title": "Broader evidence map · 1–15",
        "subtitle": "2 = per-language result; 1 = task/aggregate coverage; 0 = not included",
        "description": "Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger.",
        "showDescription": true,
        "type": "heatmap",
        "dataset": "broad_atlas_1",
        "sourceId": "broad_atlas_1",
        "intent": "relationship",
        "layout": "full",
        "question": "Broader evidence map · 1–15",
        "rationale": "Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "fields": [
              "MASSIVE calls",
              "GAIA audited",
              "X-Web shopping",
              "macOS use",
              "SEA multi-turn",
              "LILT native code",
              "Voice tools",
              "MAST retrieval",
              "MultiAgent-X tasks"
            ],
            "label": "Coverage marker",
            "format": "number"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "none"
        },
        "labels": {
          "values": "none"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "sequential",
          "name": "blue"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "maxRows": 40
      },
      {
        "id": "broad_atlas_2",
        "title": "Broader evidence map · 16–30",
        "subtitle": "2 = per-language result; 1 = task/aggregate coverage; 0 = not included",
        "description": "Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists.",
        "showDescription": true,
        "type": "heatmap",
        "dataset": "broad_atlas_2",
        "sourceId": "broad_atlas_2",
        "intent": "relationship",
        "layout": "full",
        "question": "Broader evidence map · 16–30",
        "rationale": "Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists.",
        "encodings": {
          "x": {
            "field": "Language",
            "label": "Language"
          },
          "y": {
            "fields": [
              "MASSIVE calls",
              "GAIA audited",
              "X-Web shopping",
              "macOS use",
              "SEA multi-turn",
              "LILT native code",
              "Voice tools",
              "MAST retrieval",
              "MultiAgent-X tasks"
            ],
            "label": "Coverage marker",
            "format": "number"
          }
        },
        "valueFormat": "number",
        "valueLabels": {
          "mode": "none"
        },
        "labels": {
          "values": "none"
        },
        "legend": {
          "position": "bottom",
          "sort": "spec"
        },
        "palette": {
          "kind": "sequential",
          "name": "blue"
        },
        "settings": {
          "sort": "none",
          "categoryLabelPolicy": "wrap"
        },
        "surface": {
          "viewMode": "visualization",
          "interactiveLegend": true,
          "showControls": true
        },
        "maxRows": 40
      }
    ],
    "tables": [
      {
        "id": "action_map",
        "title": "Proposed acceptance tests",
        "dataset": "action_map",
        "sourceId": "action_map",
        "columns": [
          {
            "field": "Test target",
            "label": "Test target"
          },
          {
            "field": "Input condition",
            "label": "Input condition"
          },
          {
            "field": "Expected final result",
            "label": "Expected final result"
          }
        ],
        "description": "From input conditions to final state; these tests have not been run",
        "showDescription": true,
        "defaultSort": {
          "field": "Test target",
          "direction": "asc"
        },
        "density": "comfortable"
      },
      {
        "id": "methods",
        "title": "Three current-model evidence layers",
        "dataset": "methods",
        "sourceId": "methods",
        "columns": [
          {
            "field": "Evidence",
            "label": "Evidence"
          },
          {
            "field": "Measurement",
            "label": "Measurement"
          },
          {
            "field": "Key limitation",
            "label": "Key limitation"
          }
        ],
        "description": "Metric and denominator boundaries",
        "showDescription": true,
        "defaultSort": {
          "field": "Evidence",
          "direction": "asc"
        },
        "density": "comfortable"
      },
      {
        "id": "evidence_inventory",
        "title": "16 projects: measurements and openness",
        "dataset": "evidence_inventory",
        "sourceId": "evidence_inventory",
        "columns": [
          {
            "field": "No.",
            "label": "No."
          },
          {
            "field": "Project",
            "label": "Project"
          },
          {
            "field": "Language coverage",
            "label": "Language coverage"
          },
          {
            "field": "Models and evidence layer",
            "label": "Models and evidence layer"
          },
          {
            "field": "Verified openness",
            "label": "Verified openness"
          },
          {
            "field": "Interpretation limits",
            "label": "Interpretation limits"
          }
        ],
        "description": "Public paper, tasks, code, and runs are separate assets",
        "showDescription": true,
        "defaultSort": {
          "field": "No.",
          "direction": "asc"
        },
        "density": "comfortable"
      },
      {
        "id": "triangulation",
        "title": "Language signals alongside counterevidence",
        "dataset": "triangulation",
        "sourceId": "triangulation",
        "columns": [
          {
            "field": "Language group",
            "label": "Language group"
          },
          {
            "field": "Supporting evidence",
            "label": "Supporting evidence"
          },
          {
            "field": "Supported interpretation",
            "label": "Supported interpretation"
          },
          {
            "field": "Counterevidence or limit",
            "label": "Counterevidence or limit"
          },
          {
            "field": "Action",
            "label": "Action"
          }
        ],
        "description": "Reading order only; every conclusion includes an alternative explanation",
        "showDescription": true,
        "defaultSort": {
          "field": "Language group",
          "direction": "asc"
        },
        "density": "comfortable"
      },
      {
        "id": "language_lookup",
        "title": "All 30 languages: evidence and next checks",
        "dataset": "language_lookup",
        "sourceId": "language_lookup",
        "columns": [
          {
            "field": "No.",
            "label": "No."
          },
          {
            "field": "Language",
            "label": "Language"
          },
          {
            "field": "Code",
            "label": "Code"
          },
          {
            "field": "Available evidence",
            "label": "Available evidence"
          },
          {
            "field": "Interpretation and next check",
            "label": "Interpretation and next check"
          }
        ],
        "description": "Fixed coverage, including exact-variety gaps",
        "showDescription": true,
        "defaultSort": {
          "field": "No.",
          "direction": "asc"
        },
        "density": "comfortable"
      }
    ],
    "sources": [
      {
        "id": "format_partition",
        "label": "All samples: format and content outcomes",
        "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
        "notes": "Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed.",
        "query": {
          "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
          "description": "Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Model\", \"Full sample passed\", \"Valid format, sample failed\", \"Invalid format\", \"Submission date\", \"Evaluated N\", \"Configuration limits\", \"Source\", \"Documented dataset N\", \"Scope\") AS (\n  VALUES\n    ('GPT-5.6 Sol', 0.691, 0.308, 0.001, '2026-08-18', 'Not separately disclosed', 'openrouter / 9ebf388', 'https://mera.a-ai.ru/en/text/submits/2.0/8', 1169, 'All call, abstention, and clarification items'),\n    ('Claude Opus 5', 0.677, 0.272, 0.051, '2026-08-18', 'Not separately disclosed', 'openrouter / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/1', 1169, 'All call, abstention, and clarification items'),\n    ('Grok 4.6', 0.737, 0.263, 0, '2026-08-18', 'Not separately disclosed', 'local-chat-completions / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/11', 1169, 'All call, abstention, and clarification items')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "All samples: format and content outcomes"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed."
          ]
        }
      },
      {
        "id": "argument_partition",
        "label": "Call-required items: tools and arguments",
        "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/utils.py",
        "notes": "Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference.",
        "query": {
          "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/utils.py",
          "description": "Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Model\", \"Tools and arguments matched\", \"Tools matched, arguments incomplete\", \"Tools not fully matched\", \"Submission date\", \"Evaluated N\", \"Configuration limits\", \"Source\", \"Documented dataset N\", \"Scope\") AS (\n  VALUES\n    ('GPT-5.6 Sol', 0.702, 0.163, 0.135, '2026-08-18', 'Not separately disclosed', 'openrouter / 9ebf388', 'https://mera.a-ai.ru/en/text/submits/2.0/8', 1032, 'Call-required items only'),\n    ('Claude Opus 5', 0.677, 0.113, 0.21, '2026-08-18', 'Not separately disclosed', 'openrouter / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/1', 1032, 'Call-required items only'),\n    ('Grok 4.6', 0.743, 0.081, 0.176, '2026-08-18', 'Not separately disclosed', 'local-chat-completions / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/11', 1032, 'Call-required items only')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Call-required items: tools and arguments"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference."
          ]
        }
      },
      {
        "id": "decision_errors",
        "label": "Abstention errors under two conditions",
        "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
        "notes": "1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions.",
        "query": {
          "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
          "description": "1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Required action\", \"Model\", \"Conditional error rate\", \"Documented denominator\", \"Scoring interpretation\", \"Evaluated N\", \"Source\") AS (\n  VALUES\n    ('Abstention required (127)', 'GPT-5.6 Sol', 0.417, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/8'),\n    ('Call required (1,032)', 'GPT-5.6 Sol', 0.009, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/8'),\n    ('Abstention required (127)', 'Claude Opus 5', 0.354, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/1'),\n    ('Call required (1,032)', 'Claude Opus 5', 0.001, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/1'),\n    ('Abstention required (127)', 'Grok 4.6', 0.339, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/11'),\n    ('Call required (1,032)', 'Grok 4.6', 0.004, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/11')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Abstention errors under two conditions"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions."
          ]
        }
      },
      {
        "id": "execution_partition",
        "label": "Russian code tasks: audited run outcomes",
        "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "notes": "Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control.",
        "query": {
          "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
          "description": "Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Configuration\", \"Audit-retained passes\", \"Raw passes removed by audit\", \"Original failures\", \"Retained pass count\", \"Removed pass count\", \"Original failure count\", \"Total runs\", \"Distinct tasks\", \"Harness\", \"Setting\", \"Source\") AS (\n  VALUES\n    ('GPT-5.6 Sol (23 tasks × 3)', 0.7101449275362319, 0.11594202898550725, 0.17391304347826086, 49, 8, 12, 69, 23, 'Codex CLI', 'xhigh', 'https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md'),\n    ('Claude Opus 5 (20 tasks × 3)', 0.7666666666666667, 0.016666666666666666, 0.21666666666666667, 46, 1, 13, 60, 20, 'Claude Code', 'xhigh', 'https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Russian code tasks: audited run outcomes"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control."
          ]
        }
      },
      {
        "id": "document_profile",
        "label": "Document parsing across 13 language observations",
        "href": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e",
        "notes": "Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold.",
        "query": {
          "url": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e",
          "description": "Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Parsing quality score\", \"Observed position\", \"Script group\", \"Order\", \"Median across 13 observations\", \"Source\") AS (\n  VALUES\n    ('Thai', 'th', 72.1, 'Lower four observations', 'Other scripts', 1, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Japanese', 'ja', 74.9, 'Lower four observations', 'Other scripts', 2, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Arabic (unspecified variety)', 'ar', 77.4, 'Lower four observations', 'Other scripts', 3, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Hindi', 'hi', 77.5, 'Lower four observations', 'Other scripts', 4, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('French', 'fr', 80.0, 'Other nine observations', 'Latin script', 5, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Spanish', 'es', 80.2, 'Other nine observations', 'Latin script', 6, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Russian', 'ru', 82.4, 'Other nine observations', 'Other scripts', 7, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Vietnamese', 'vi', 84.8, 'Other nine observations', 'Latin script', 8, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Indonesian', 'id', 86.9, 'Other nine observations', 'Latin script', 9, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Portuguese', 'pt', 88.9, 'Other nine observations', 'Latin script', 10, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('German', 'de', 89.1, 'Other nine observations', 'Latin script', 11, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Korean', 'ko', 89.9, 'Other nine observations', 'Other scripts', 12, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Italian', 'it', 92.7, 'Other nine observations', 'Latin script', 13, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Document parsing across 13 language observations"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold."
          ]
        }
      },
      {
        "id": "script_distribution",
        "label": "Parsing score distributions by script group",
        "href": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e",
        "notes": "Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges.",
        "query": {
          "url": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e",
          "description": "Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Script group\", \"Minimum\", \"First quartile\", \"Median\", \"Third quartile\", \"Maximum\", \"Observation count\", \"Mean\", \"Language\", \"Source\") AS (\n  VALUES\n    ('Latin script (7)', 80.0, 82.5, 86.9, 89.0, 92.7, 7, 86.08571428571429, 'French; Spanish; Vietnamese; Indonesian; Portuguese; German; Italian', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Other scripts (6)', 72.1, 75.525, 77.45, 81.17500000000001, 89.9, 6, 79.03333333333333, 'Thai; Japanese; Arabic (unspecified variety); Hindi; Russian; Korean', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Parsing score distributions by script group"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges."
          ]
        }
      },
      {
        "id": "atlas_1",
        "label": "Current-model evidence map · 1–15",
        "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "notes": "Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related.",
        "query": {
          "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
          "description": "Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"Document component\", \"Static tool plan\", \"End-to-end execution\", \"Matched language control\", \"Marker definition\") AS (\n  VALUES\n    ('01 Hindi', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('02 Spanish', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('03 Modern Standard Arabic', 0.5, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('04 French', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('05 Bengali', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('06 Portuguese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('07 Indonesian', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('08 Urdu', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('09 Russian', 1, 1, 1, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('10 German', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('11 Japanese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('12 Nigerian Pidgin', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('13 Egyptian Arabic', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('14 Marathi', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('15 Vietnamese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Current-model evidence map · 1–15"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related."
          ]
        }
      },
      {
        "id": "atlas_2",
        "label": "Current-model evidence map · 16–30",
        "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "notes": "Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages.",
        "query": {
          "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
          "description": "Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"Document component\", \"Static tool plan\", \"End-to-end execution\", \"Matched language control\", \"Marker definition\") AS (\n  VALUES\n    ('16 Telugu', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('17 Swahili', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('18 Hausa', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('19 Turkish', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('20 Western Punjabi', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('21 Tagalog', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('22 Tamil', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('23 Iranian Persian', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('24 Korean', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('25 Amharic', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('26 Thai', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('27 Javanese', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('28 Italian', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('29 Gujarati', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('30 Kannada', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Current-model evidence map · 16–30"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages."
          ]
        }
      },
      {
        "id": "action_map",
        "label": "Proposed acceptance tests",
        "href": "https://www.unicode.org/reports/tr35/tr35-numbers.html",
        "notes": "Researcher-proposed fixtures based on Unicode conventions and interface contracts. No measured model outcomes or failure incidence. Preserve distinctions pcm≠en, arz≠ar, pnb≠pa, jv≠id.",
        "query": {
          "url": "https://www.unicode.org/reports/tr35/tr35-numbers.html",
          "description": "Researcher-proposed fixtures based on Unicode conventions and interface contracts. No measured model outcomes or failure incidence. Preserve distinctions pcm≠en, arz≠ar, pnb≠pa, jv≠id. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Test target\", \"Input condition\", \"Expected final result\") AS (\n  VALUES\n    ('Amounts | Urdu / Persian', '۱٬۲۵۰٫۵۰; explicit currency', 'Pass 1250.5 according to the API contract; preserve currency'),\n    ('Entity IDs | all languages', 'INV-01250/AB; mixed scripts', 'Preserve string, leading zeros, separators; act on the correct entity'),\n    ('Dates | explicit locale / calendar', '05/09/2026; with/without locale control', 'Convert correctly when defined; clarify per rules when ambiguous'),\n    ('Dialect / neighboring language | four gaps', 'Native pcm, arz, pnb, jv requests', 'Preserve native constraints; no substitution by neighboring-language scores')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Proposed acceptance tests"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Researcher-proposed fixtures based on Unicode conventions and interface contracts. No measured model outcomes or failure incidence. Preserve distinctions pcm≠en, arz≠ar, pnb≠pa, jv≠id."
          ]
        }
      },
      {
        "id": "methods",
        "label": "Three current-model evidence layers",
        "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
        "notes": "Review synthesis of the audited RuBench, GorillaHard, and MDP configurations. Do not equate static plans, document components, executed tasks, or language-causal comparisons.",
        "query": {
          "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
          "description": "Review synthesis of the audited RuBench, GorillaHard, and MDP configurations. Do not equate static plans, document components, executed tasks, or language-causal comparisons. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Evidence\", \"Measurement\", \"Key limitation\") AS (\n  VALUES\n    ('GorillaHard', 'Static output; documented N=1,169 total, 1,032 calls, 127 abstentions', 'Documented counts, not separately disclosed run N. Incomplete settings; Sol source commit unresolved; no English pairs.'),\n    ('MDPBench', 'Kimi K3 document parsing quality; 13 target observations', 'Unpaired content/layout; unspecified Arabic variety; composite score is not failure probability.'),\n    ('RuBench', 'Russian instructions → repository edits → regression tests', 'Sol 23×3, Opus 20×3; different sets/harnesses. Earlier Fable fallback omitted.')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Three current-model evidence layers"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Review synthesis of the audited RuBench, GorillaHard, and MDP configurations. Do not equate static plans, document components, executed tasks, or language-causal comparisons."
          ]
        }
      },
      {
        "id": "evidence-rubench_paper",
        "label": "RuBench v1: construction, protocol and Fable fallback",
        "href": "https://arxiv.org/html/2607.06411v1",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://arxiv.org/html/2607.06411v1",
          "tables_used": [
            "RuBench v1: construction, protocol and Fable fallback"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "evidence-rubench_tasks",
        "label": "RuBench task metadata",
        "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/tasks.json",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/tasks.json",
          "tables_used": [
            "RuBench task metadata"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "evidence-rubench_aioh4",
        "label": "AIOH4 original Russian task statement",
        "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/tasks/AIOH4-case-sensitivity/task.md",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/tasks/AIOH4-case-sensitivity/task.md",
          "tables_used": [
            "AIOH4 original Russian task statement"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "evidence-mera_dataset",
        "label": "GorillaHard dataset description and Russian sample",
        "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/datasets/GorillaHard/README.md",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/datasets/GorillaHard/README.md",
          "tables_used": [
            "GorillaHard dataset description and Russian sample"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "evidence-mera_config",
        "label": "GorillaHard task configuration",
        "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/gorillahard.yaml",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/gorillahard.yaml",
          "tables_used": [
            "GorillaHard task configuration"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "evidence-mera_author",
        "label": "MERA team's Text 2.0 launch explanation",
        "href": "https://habr.com/ru/companies/sberbank/articles/1075592/",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://habr.com/ru/companies/sberbank/articles/1075592/",
          "tables_used": [
            "MERA team's Text 2.0 launch explanation"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "evidence-mera_sol",
        "label": "MERA GPT 5.6 Sol submission",
        "href": "https://mera.a-ai.ru/en/text/submits/2.0/8",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://mera.a-ai.ru/en/text/submits/2.0/8",
          "tables_used": [
            "MERA GPT 5.6 Sol submission"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "evidence-mera_opus",
        "label": "MERA Claude Opus 5 submission",
        "href": "https://mera.a-ai.ru/en/text/submits/2.0/1",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://mera.a-ai.ru/en/text/submits/2.0/1",
          "tables_used": [
            "MERA Claude Opus 5 submission"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "evidence-mera_grok",
        "label": "MERA Grok 4.6 submission",
        "href": "https://mera.a-ai.ru/en/text/submits/2.0/11",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://mera.a-ai.ru/en/text/submits/2.0/11",
          "tables_used": [
            "MERA Grok 4.6 submission"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "linked-20",
        "label": "MDPBench parsing metrics",
        "href": "https://arxiv.org/html/2603.28130v1#S3.SS4",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://arxiv.org/html/2603.28130v1#S3.SS4",
          "tables_used": [
            "MDPBench parsing metrics"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "linked-21",
        "label": "Pinned public language population transcription",
        "href": "https://en.wikipedia.org/w/index.php?title=List_of_languages_by_total_number_of_speakers&oldid=1370552712",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://en.wikipedia.org/w/index.php?title=List_of_languages_by_total_number_of_speakers&oldid=1370552712",
          "tables_used": [
            "Pinned public language population transcription"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "linked-22",
        "label": "Frontier cohort reference",
        "href": "https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2",
        "query": {
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "url": "https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2",
          "tables_used": [
            "Frontier cohort reference"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "gaia_lift",
        "label": "Score improvement after task auditing",
        "href": "https://arxiv.org/html/2604.24929v1",
        "notes": "GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim. Material inputs: https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS",
        "query": {
          "url": "https://arxiv.org/html/2604.24929v1",
          "description": "GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Model\", \"Before audit\", \"After audit\", \"English reference\", \"Audit improvement (pp)\", \"English minus audited (pp)\", \"Tasks per language\", \"Model cohort\", \"Execution harness\", \"Source\") AS (\n  VALUES\n    ('Arabic (generic)†', 'ar', 'GPT-5.4', 32.1, 47.3, 66.7, 15.2, 19.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'GPT-5.4', 47.3, 63.6, 66.7, 16.3, 3.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'GPT-5.4', 34.6, 60.0, 66.7, 25.4, 6.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'GPT-5.4', 33.3, 62.4, 66.7, 29.1, 4.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'GPT-5.4', 47.3, 58.2, 66.7, 10.9, 8.5, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Arabic (generic)†', 'ar', 'Gemini 3.1 Pro', 34.6, 52.1, 73.9, 17.5, 21.8, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'Gemini 3.1 Pro', 49.7, 66.7, 73.9, 17.0, 7.2, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'Gemini 3.1 Pro', 38.2, 63.6, 73.9, 25.4, 10.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'Gemini 3.1 Pro', 34.6, 64.8, 73.9, 30.2, 9.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'Gemini 3.1 Pro', 48.5, 65.5, 73.9, 17.0, 8.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Arabic (generic)†', 'ar', 'Claude Opus 4.6', 32.1, 49.1, 79.4, 17.0, 30.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'Claude Opus 4.6', 49.7, 66.7, 79.4, 17.0, 12.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'Claude Opus 4.6', 29.7, 62.4, 79.4, 32.7, 17.0, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'Claude Opus 4.6', 33.3, 58.8, 79.4, 25.5, 20.6, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'Claude Opus 4.6', 49.1, 63.0, 79.4, 13.9, 16.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Score improvement after task auditing"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim."
          ]
        }
      },
      {
        "id": "gaia_residual",
        "label": "Remaining gap from each model’s English reference",
        "href": "https://arxiv.org/html/2604.24929v1",
        "notes": "GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals. Material inputs: https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS",
        "query": {
          "url": "https://arxiv.org/html/2604.24929v1",
          "description": "GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Model\", \"Before audit\", \"After audit\", \"English reference\", \"Audit improvement (pp)\", \"English minus audited (pp)\", \"Tasks per language\", \"Model cohort\", \"Execution harness\", \"Source\") AS (\n  VALUES\n    ('Arabic (generic)†', 'ar', 'GPT-5.4', 32.1, 47.3, 66.7, 15.2, 19.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'GPT-5.4', 47.3, 63.6, 66.7, 16.3, 3.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'GPT-5.4', 34.6, 60.0, 66.7, 25.4, 6.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'GPT-5.4', 33.3, 62.4, 66.7, 29.1, 4.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'GPT-5.4', 47.3, 58.2, 66.7, 10.9, 8.5, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Arabic (generic)†', 'ar', 'Gemini 3.1 Pro', 34.6, 52.1, 73.9, 17.5, 21.8, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'Gemini 3.1 Pro', 49.7, 66.7, 73.9, 17.0, 7.2, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'Gemini 3.1 Pro', 38.2, 63.6, 73.9, 25.4, 10.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'Gemini 3.1 Pro', 34.6, 64.8, 73.9, 30.2, 9.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'Gemini 3.1 Pro', 48.5, 65.5, 73.9, 17.0, 8.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Arabic (generic)†', 'ar', 'Claude Opus 4.6', 32.1, 49.1, 79.4, 17.0, 30.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'Claude Opus 4.6', 49.7, 66.7, 79.4, 17.0, 12.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'Claude Opus 4.6', 29.7, 62.4, 79.4, 32.7, 17.0, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'Claude Opus 4.6', 33.3, 58.8, 79.4, 25.5, 20.6, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'Claude Opus 4.6', 49.1, 63.0, 79.4, 13.9, 16.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Remaining gap from each model’s English reference"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals."
          ]
        }
      },
      {
        "id": "sea_localization",
        "label": "Retail success by extent of localization",
        "href": "https://arxiv.org/html/2606.28715v1",
        "notes": "SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English. Material inputs: https://arxiv.org/html/2606.28715v1 ; https://github.com/SEACrowd/SEATauBench",
        "query": {
          "url": "https://arxiv.org/html/2606.28715v1",
          "description": "SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Setting\", \"Task success rate\", \"Model\", \"Displayed domain\", \"Trials per task\", \"Airline English\", \"Airline fully localized\", \"Telecom English\", \"Telecom fully localized\", \"Sample definition\", \"Source\") AS (\n  VALUES\n    ('Vietnamese', 'vi', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Vietnamese', 'vi', 'Dialogue localized', 0.687, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Vietnamese', 'vi', 'Fully localized business', 0.567, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Thai', 'th', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Thai', 'th', 'Dialogue localized', 0.573, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Thai', 'th', 'Fully localized business', 0.327, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Indonesian', 'id', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Indonesian', 'id', 'Dialogue localized', 0.64, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Indonesian', 'id', 'Fully localized business', 0.433, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Filipino†', 'tl', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Filipino†', 'tl', 'Dialogue localized', 0.675, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Filipino†', 'tl', 'Fully localized business', 0.444, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Retail success by extent of localization"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English."
          ]
        }
      },
      {
        "id": "macos_language",
        "label": "Desktop task success by agent and language",
        "href": "https://arxiv.org/html/2506.04135v4",
        "notes": "macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs. Material inputs: https://arxiv.org/html/2506.04135v4 ; https://github.com/showlab/macosworld",
        "query": {
          "url": "https://arxiv.org/html/2506.04135v4",
          "description": "macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Agent\", \"Task success rate\", \"Exact model version\", \"Task count\", \"Difference from English (pp)\", \"UI and instruction languages\", \"Source\") AS (\n  VALUES\n    ('English reference', 'en', 'Claude CUA', 0.444, 'claude-3-7-sonnet-20250219', 171, 0.0, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Arabic (generic)†', 'ar', 'Claude CUA', 0.316, 'claude-3-7-sonnet-20250219', 171, -12.8, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Japanese', 'ja', 'Claude CUA', 0.368, 'claude-3-7-sonnet-20250219', 171, -7.6, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Russian', 'ru', 'Claude CUA', 0.409, 'claude-3-7-sonnet-20250219', 171, -3.5, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('English reference', 'en', 'OpenAI CUA', 0.33299999999999996, 'computer-use-preview-2025-03-11', 171, 0.0, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Arabic (generic)†', 'ar', 'OpenAI CUA', 0.281, 'computer-use-preview-2025-03-11', 171, -5.2, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Japanese', 'ja', 'OpenAI CUA', 0.35100000000000003, 'computer-use-preview-2025-03-11', 171, 1.8, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Russian', 'ru', 'OpenAI CUA', 0.392, 'computer-use-preview-2025-03-11', 171, 5.9, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Desktop task success by agent and language"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs."
          ]
        }
      },
      {
        "id": "xweb_translation",
        "label": "Web shopping scores: original language versus English translation",
        "href": "https://arxiv.org/html/2505.15372v1",
        "notes": "X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics. Material inputs: https://arxiv.org/html/2505.15372v1 ; https://github.com/WPENGxs/X-WebAgentBench",
        "query": {
          "url": "https://arxiv.org/html/2505.15372v1",
          "description": "X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Strategy\", \"Task score\", \"Model\", \"Instructions per language\", \"Metric\", \"Source\") AS (\n  VALUES\n    ('French', 'fr', 'Original-language BaseAgent', 42.7, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('French', 'fr', 'Google Translate to English', 48.33, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Spanish', 'es', 'Original-language BaseAgent', 37.31, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Spanish', 'es', 'Google Translate to English', 25.97, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('German', 'de', 'Original-language BaseAgent', 34.56, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('German', 'de', 'Google Translate to English', 37.41, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Russian', 'ru', 'Original-language BaseAgent', 36.41, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Russian', 'ru', 'Google Translate to English', 38.0, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Turkish', 'tr', 'Original-language BaseAgent', 43.18, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Turkish', 'tr', 'Google Translate to English', 33.99, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Arabic (generic)†', 'ar', 'Original-language BaseAgent', 41.34, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Arabic (generic)†', 'ar', 'Google Translate to English', 2.18, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Vietnamese', 'vi', 'Original-language BaseAgent', 37.93, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Vietnamese', 'vi', 'Google Translate to English', 35.19, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Thai', 'th', 'Original-language BaseAgent', 18.51, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Thai', 'th', 'Google Translate to English', 12.69, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Hindi', 'hi', 'Original-language BaseAgent', 36.75, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Hindi', 'hi', 'Google Translate to English', 33.47, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Swahili', 'sw', 'Original-language BaseAgent', 42.1, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Swahili', 'sw', 'Google Translate to English', 36.49, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Urdu', 'ur', 'Original-language BaseAgent', 34.15, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Urdu', 'ur', 'Google Translate to English', 2.3, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Web shopping scores: original language versus English translation"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics."
          ]
        }
      },
      {
        "id": "evidence_inventory",
        "label": "16 projects: measurements and openness",
        "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "notes": "Project inventory, not independent experiment count or exhaustive literature census. Primary papers, author repositories, and data cards checked. Unverified does not mean unavailable. No full repository-by-repository reproduction or legal license review. MAST baseline only as of cutoff. Material inputs: https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md ; https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md ; https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e ; https://arxiv.org/html/2606.28715v1 ; https://github.com/SEACrowd/SEATauBench ; https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://aclanthology.org/2026.findings-eacl.42/ ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS ; https://arxiv.org/html/2506.04135v4 ; https://github.com/showlab/macosworld ; https://arxiv.org/html/2505.15372v1 ; https://github.com/WPENGxs/X-WebAgentBench ; https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2510.07978v1 ; https://github.com/ola-krutrim/VoiceAgentBench ; https://arxiv.org/html/2604.06209v1 ; https://github.com/BrahiM-Mefgouda/TelcoAgent ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://benchmarks.lilt.com/ ; https://github.com/lilt/liltbench-tasks-public ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark ; https://mast-benchmark.github.io/ ; https://arxiv.org/html/2604.04532v1 ; https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
        "query": {
          "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
          "description": "Project inventory, not independent experiment count or exhaustive literature census. Primary papers, author repositories, and data cards checked. Unverified does not mean unavailable. No full repository-by-repository reproduction or legal license review. MAST baseline only as of cutoff. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"No.\", \"Project\", \"Language coverage\", \"Task layer\", \"Models and evidence layer\", \"Verified openness\", \"Interpretation limits\", \"Task ancestry\", \"Source\") AS (\n  VALUES\n    (1, 'RuBench', 'Russian', 'Code execution', 'Sol, Opus 5 / current direct', 'Tasks and run results public; full grading tests withheld', 'No matched English; models use different task sets', 'Independent native code tasks', 'https://github.com/eugeneshilow/rubench'),\n    (2, 'MERA GorillaHard', 'Russian', 'Static tool plans', 'Sol, Opus 5, Grok 4.6 / current direct', 'Framework, metrics, scores public; incomplete test truth/configuration', 'Tools not executed; submission settings not fully aligned', 'MERA / GorillaHard', 'https://mera.a-ai.ru/en/text'),\n    (3, 'MDPBench / Kimi K3', '13 target observations, including related Arabic label', 'Document component', 'Kimi K3 / current component', 'Official per-language scores public; no rerun here', 'Not agent success; unpaired documents', 'Document parsing component', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    (4, 'GAIA-v2-LILT', 'ar de hi ko pt-BR + en', 'Retrieval / tool execution', 'GPT-5.4, Gemini 3.1 Pro, Opus 4.6 / earlier flagships', 'Evaluation code, HF tasks, per-language table public', 'Audit changes function, culture, difficulty together; not pure language causality', 'Audited MAPS-GAIA branch', 'https://arxiv.org/html/2604.24929v1'),\n    (5, 'MAPS', 'Final paper: 12 languages; older HF card: 11', 'GAIA/SWE/safety/math mixture', 'Historical models / context', 'Paper and HF data public; pin subset and version', 'MATH is not execution; GAIA shares ancestry with preceding row', 'Parent GAIA, SWE, MATH, ASB sets', 'https://aclanthology.org/2026.findings-eacl.42/'),\n    (6, 'SEATauBench', 'vi th id Filipino zh + en', 'Multi-turn tools and final state', 'Kimi K2.5, GPT-5-mini, Qwen3 / historical', 'Code, domain data, analysis framework public; MIT repository', 'Simulator can also fail; no isolated agent-only language effect', 'Derived from tau2; shares framework with LILT tau', 'https://github.com/SEACrowd/SEATauBench'),\n    (7, 'macOSWorld', 'ar ja ru zh + en', 'Interactive desktop', 'Claude 3.7 CUA, 2025 OpenAI CUA / earlier flagships', 'Tasks, environment, scoring code public; per-language paper results', 'Instructions and UI switch together; v4 revises some scores', 'Native macOS tasks', 'https://arxiv.org/html/2506.04135v4'),\n    (8, 'X-WebAgentBench', '14 non-English settings; 11 target languages', 'Interactive web shopping', 'GPT-4o and others / historical', 'Code and data entry points public; paper results', 'Task Score is not success rate; derived from WebShop', 'WebShop derivative', 'https://github.com/WPENGxs/X-WebAgentBench'),\n    (9, 'MASSIVE-Agents', '52 languages; 24 targets including related labels', 'Static function / argument matching', 'Nova Premier, Claude 3.5 and others / earlier flagships', 'Converted HF data CC BY 4.0; per-language paper tables', 'Not multi-turn execution; filtering changes samples/function coverage by language', 'MASSIVE derivative; BFCL scoring', 'https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id'),\n    (10, 'Terminal-Bench-LILT', 'ar cs de es hi ja ko sr tr zh', 'Native multilingual code execution', 'GPT-5.5, Opus 4.8 and others / earlier flagships', 'Paper, aggregate scores, samples public; full tasks require contacting authors', 'Blog 300 versus leaderboard 324 tasks: denominators not mixed; full set not public', 'Native Terminal-Bench format; separate from community LILTBench', 'https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark'),\n    (11, 'LILT multilingual tau', 'de ko + en', 'Multi-turn tools and final state', 'GPT-5.4, Opus 4.8, Gemini 3.1 Pro / earlier flagships', 'Public board and run notes; dynamic per-language scores not extracted here', 'Different English/target-language simulators confound scores', 'tau2 derivative; not independent framework replication', 'https://benchmarks.lilt.com/'),\n    (12, 'LILTBench community', 'Authors report 31 languages; not exhaustively checked task by task', 'Native code / English pairs', 'Opus 4.6 / challenging test resources', 'Tasks, verifiers, leaderboard public; Apache-2.0 repository', 'Adversarial hard-task selection cannot estimate population failure rates', 'Community native tasks; distinct from commercial full set', 'https://github.com/lilt/liltbench-tasks-public'),\n    (13, 'VoiceAgentBench', 'en hi bn mr ta te ml', 'Speech-to-tool-call scoring', 'SpeechLM / ASR+LLM / non-frontier resources', 'Evaluation code and HF data entry public; custom community license', 'Multi-turn subset is English only; predicted calls are not final environment success', 'Synthetic speech-tool tasks', 'https://github.com/ola-krutrim/VoiceAgentBench'),\n    (14, 'TelcoAgent-Bench', 'ar + en', 'Diagnostic tool sequences / summaries', '3B–8B models / non-frontier resources', 'Blueprints, tasks, prediction/score JSON public; license unverified', 'Intent/resolution similarity is not strict success; cannot extrapolate to flagships', 'Telecom blueprint tasks', 'https://github.com/BrahiM-Mefgouda/TelcoAgent'),\n    (15, 'MultiAgent-X', '12 languages; target hi sw ha am', 'Function-call data', 'No verified current-frontier baseline / resources', 'Samples, scripts, structure public; full license unverified', 'Synthetic; structural validity is not native quality; pa is not pnb', 'New synthetic tasks; quoted MASSIVE scores are not replication', 'https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark'),\n    (16, 'MAST FIRE 2026', '21-language track union; 14 targets', 'Multi-turn retrieval / answers', 'Public Tongyi-30B baseline / non-frontier', 'Task/corpus entry points and per-language baseline public', 'September 30 final results are future at cutoff; baseline only; English documents and answers', 'BrowseComp-Plus derivative; three Indic languages overlap tracks', 'https://mast-benchmark.github.io/')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "16 projects: measurements and openness"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Project inventory, not independent experiment count or exhaustive literature census. Primary papers, author repositories, and data cards checked. Unverified does not mean unavailable. No full repository-by-repository reproduction or legal license review. MAST baseline only as of cutoff."
          ]
        }
      },
      {
        "id": "broad_atlas_1",
        "label": "Broader evidence map · 1–15",
        "href": "https://aclanthology.org/2025.findings-emnlp.1099.pdf",
        "notes": "Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger. Material inputs: https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2604.24929v1 ; https://arxiv.org/html/2505.15372v1 ; https://arxiv.org/html/2506.04135v4 ; https://arxiv.org/html/2606.28715v1 ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://arxiv.org/html/2510.07978v1 ; https://mast-benchmark.github.io/ ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark",
        "query": {
          "url": "https://aclanthology.org/2025.findings-emnlp.1099.pdf",
          "description": "Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"MASSIVE calls\", \"GAIA audited\", \"X-Web shopping\", \"macOS use\", \"SEA multi-turn\", \"LILT native code\", \"Voice tools\", \"MAST retrieval\", \"MultiAgent-X tasks\", \"Code\", \"Included projects\", \"Mapping limitation\", \"Meaning\") AS (\n  VALUES\n    ('Hindi', 2, 2, 2, 0, 0, 1, 1, 2, 1, 'hi', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Spanish', 2, 0, 2, 0, 0, 1, 0, 2, 0, 'es', 'MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Modern Standard Arabic†', 2, 2, 2, 2, 0, 1, 0, 2, 0, 'ar', 'MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval', 'Generic Arabic/Filipino are related labels, not exact MSA/Tagalog matches', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('French', 2, 0, 2, 0, 0, 0, 0, 2, 0, 'fr', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Bengali', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'bn', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Portuguese', 2, 2, 0, 0, 0, 0, 0, 0, 0, 'pt', 'MASSIVE calls; GAIA audited', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Indonesian', 2, 0, 0, 0, 2, 0, 0, 0, 0, 'id', 'MASSIVE calls; SEA multi-turn', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Urdu', 2, 0, 2, 0, 0, 0, 0, 2, 0, 'ur', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Russian', 2, 0, 2, 2, 0, 0, 0, 2, 0, 'ru', 'MASSIVE calls; X-Web shopping; macOS use; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('German', 2, 2, 2, 0, 0, 1, 0, 2, 0, 'de', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Japanese', 2, 0, 0, 2, 0, 1, 0, 0, 0, 'ja', 'MASSIVE calls; macOS use; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Nigerian Pidgin', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'pcm', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Egyptian Arabic', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'arz', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Marathi', 0, 0, 0, 0, 0, 0, 1, 0, 0, 'mr', 'Voice tools', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Vietnamese', 2, 0, 2, 0, 2, 0, 0, 0, 0, 'vi', 'MASSIVE calls; X-Web shopping; SEA multi-turn', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Broader evidence map · 1–15"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger."
          ]
        }
      },
      {
        "id": "broad_atlas_2",
        "label": "Broader evidence map · 16–30",
        "href": "https://aclanthology.org/2025.findings-emnlp.1099.pdf",
        "notes": "Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists. Material inputs: https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2604.24929v1 ; https://arxiv.org/html/2505.15372v1 ; https://arxiv.org/html/2506.04135v4 ; https://arxiv.org/html/2606.28715v1 ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://arxiv.org/html/2510.07978v1 ; https://mast-benchmark.github.io/ ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark",
        "query": {
          "url": "https://aclanthology.org/2025.findings-emnlp.1099.pdf",
          "description": "Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language\", \"MASSIVE calls\", \"GAIA audited\", \"X-Web shopping\", \"macOS use\", \"SEA multi-turn\", \"LILT native code\", \"Voice tools\", \"MAST retrieval\", \"MultiAgent-X tasks\", \"Code\", \"Included projects\", \"Mapping limitation\", \"Meaning\") AS (\n  VALUES\n    ('Telugu', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'te', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Swahili', 2, 0, 2, 0, 0, 0, 0, 2, 1, 'sw', 'MASSIVE calls; X-Web shopping; MAST retrieval; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Hausa', 0, 0, 0, 0, 0, 0, 0, 0, 1, 'ha', 'MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Turkish', 2, 0, 2, 0, 0, 1, 0, 0, 0, 'tr', 'MASSIVE calls; X-Web shopping; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Western Punjabi', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'pnb', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Tagalog†', 2, 0, 0, 0, 2, 0, 0, 0, 0, 'tl', 'MASSIVE calls; SEA multi-turn', 'Generic Arabic/Filipino are related labels, not exact MSA/Tagalog matches', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Tamil', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'ta', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Iranian Persian', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'fa', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Korean', 2, 2, 0, 0, 0, 1, 0, 0, 0, 'ko', 'MASSIVE calls; GAIA audited; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Amharic', 2, 0, 0, 0, 0, 0, 0, 0, 1, 'am', 'MASSIVE calls; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Thai', 2, 0, 2, 0, 2, 0, 0, 2, 0, 'th', 'MASSIVE calls; X-Web shopping; SEA multi-turn; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Javanese', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'jv', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Italian', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'it', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Gujarati', 0, 0, 0, 0, 0, 0, 0, 2, 0, 'gu', 'MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Kannada', 2, 0, 0, 0, 0, 0, 0, 2, 0, 'kn', 'MASSIVE calls; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Broader evidence map · 16–30"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists."
          ]
        }
      },
      {
        "id": "triangulation",
        "label": "Language signals alongside counterevidence",
        "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "notes": "Researcher synthesis, not risk score or statistical ranking. Shared task ancestry, historical models, components, and non-frontier baselines remain distinguished. Recommendations are unrun tests. Material inputs: https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md ; https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md ; https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e ; https://arxiv.org/html/2606.28715v1 ; https://github.com/SEACrowd/SEATauBench ; https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://aclanthology.org/2026.findings-eacl.42/ ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS ; https://arxiv.org/html/2506.04135v4 ; https://github.com/showlab/macosworld ; https://arxiv.org/html/2505.15372v1 ; https://github.com/WPENGxs/X-WebAgentBench ; https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2510.07978v1 ; https://github.com/ola-krutrim/VoiceAgentBench ; https://arxiv.org/html/2604.06209v1 ; https://github.com/BrahiM-Mefgouda/TelcoAgent ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://benchmarks.lilt.com/ ; https://github.com/lilt/liltbench-tasks-public ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark ; https://mast-benchmark.github.io/ ; https://arxiv.org/html/2604.04532v1 ; https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
        "query": {
          "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
          "description": "Researcher synthesis, not risk score or statistical ranking. Shared task ancestry, historical models, components, and non-frontier baselines remain distinguished. Recommendations are unrun tests. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"Language group\", \"Supporting evidence\", \"Supported interpretation\", \"Counterevidence or limit\", \"Action\") AS (\n  VALUES\n    ('01 Thai', 'SEA retail and other domains; X-Web shopping; K3 documents', 'Weakness across tasks; current-frontier evidence is component-only', 'Thai is the highest one-shot language for Llama 3.1 405B in MASSIVE; models/prompts change order', 'Retest fully localized business and shopping; avoid all-model claims'),\n    ('02 Arabic (generic)†', 'Audited GAIA: three flagships below English; macOS: two CUAs lower; K3 component', 'Aligned historical signals across two execution families', 'GAIA improves after audit; GUI also changes RTL layout; dialects untested', 'Retest search, RTL UI, and numeric arguments; do not extend to Egyptian Arabic'),\n    ('03 Japanese', 'macOS CUA; native code tasks; K3 component', 'Component weakness; execution depends on model', 'OpenAI CUA exceeds English; Claude CUA falls below English', 'Separate document reading, click targeting, and display-width tests'),\n    ('04 Hindi', 'Audited GAIA; native code; MASSIVE; K3 component', 'Several task families; low scores sensitive to task quality', 'GAIA audit gains 25.4–32.7 pp; English instructions do not remove all native-task difficulty in Hindi/German', 'Audit task alignment, entities, and regional rules; distinguish reading from execution'),\n    ('05 Russian', 'Current RuBench execution and GorillaHard plans; historical macOS', 'Concrete current failures, more direct than a component', 'RuBench/MERA lack English pairs; historical OpenAI CUA Russian exceeds English', 'Regress specific failures; do not rank Russian worst overall'),\n    ('06 Vietnamese / Indonesian / Filipino†', 'SEA full business; MASSIVE; Vietnamese X-Web', 'Several tasks; language ordering varies by domain', 'Kimi Vietnamese retail slightly exceeds English; Filipino simulator also errs', 'Accept by business domain and verify Filipino–Tagalog variety alignment'),\n    ('07 German / Korean / Portuguese', 'Audited GAIA; MASSIVE; K3; German/Korean LILT', 'More evidence does not establish deployment readiness', 'Audited German gap is 3.1 pp for GPT-5.4, 12.7 pp for Opus 4.6', 'Test the target model and actual workflow; no universal safe-language list'),\n    ('08 Amharic / Kannada', 'Historical MASSIVE zero-shot AST; Amharic synthetic tasks', 'Risk signal mainly from one historical task family', 'Lowest language differs by model; no current-frontier execution replication', 'Prioritize function arguments; evidence cannot rank current models'),\n    ('09 Swahili / Urdu', 'X-Web; MASSIVE; non-frontier MAST retrieval', 'Task- and model-dependent signals; no uniform direction', 'Swahili original-language GPT-4o shopping is not low; speech and retrieval scores are not interchangeable', 'Measure retrieval, calls, and final state separately with English controls'),\n    ('10 Other coverage / gaps', 'Bengali/Tamil/Telugu tool or retrieval resources; Pidgin speech component', 'Consult all 30 rows; missing evidence is not low ability', 'Egyptian Arabic/Western Punjabi exact matches missing; Hausa lacks current-frontier baseline', 'Use native tasks: pcm≠en, arz≠ar, pnb≠pa, jv≠id')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "Language signals alongside counterevidence"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Researcher synthesis, not risk score or statistical ranking. Shared task ancestry, historical models, components, and non-frontier baselines remain distinguished. Recommendations are unrun tests."
          ]
        }
      },
      {
        "id": "language_lookup",
        "label": "All 30 languages: evidence and next checks",
        "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "notes": "Researcher mapping retains all target languages, named evidence and inference limits. No invented scores; related labels do not replace exact language varieties. Material inputs: https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md ; https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md ; https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e ; https://arxiv.org/html/2606.28715v1 ; https://github.com/SEACrowd/SEATauBench ; https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://aclanthology.org/2026.findings-eacl.42/ ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS ; https://arxiv.org/html/2506.04135v4 ; https://github.com/showlab/macosworld ; https://arxiv.org/html/2505.15372v1 ; https://github.com/WPENGxs/X-WebAgentBench ; https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2510.07978v1 ; https://github.com/ola-krutrim/VoiceAgentBench ; https://arxiv.org/html/2604.06209v1 ; https://github.com/BrahiM-Mefgouda/TelcoAgent ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://benchmarks.lilt.com/ ; https://github.com/lilt/liltbench-tasks-public ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark ; https://mast-benchmark.github.io/ ; https://arxiv.org/html/2604.04532v1 ; https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
        "query": {
          "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
          "description": "Researcher mapping retains all target languages, named evidence and inference limits. No invented scores; related labels do not replace exact language varieties. SQLite VALUES reproduces reviewed observations; not a live source query.",
          "sql": "WITH reviewed_published_rows (\"No.\", \"Language\", \"Code\", \"Available evidence\", \"Interpretation and next check\") AS (\n  VALUES\n    (1, 'Hindi', 'hi', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks', 'Audited GAIA recovers substantially; K3 reading is low. Separate benchmark quality from local-rule execution.'),\n    (2, 'Spanish', 'es', 'MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval', 'GAIA parent set, shopping, function calls, and retrieval coverage. No complete current-frontier comparison; do not presume European-language reliability.'),\n    (3, 'Modern Standard Arabic', 'ar', 'MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval', 'GAIA and macOS give aligned historical signals; K3 component score is lower. Generic Arabic does not precisely establish MSA or Egyptian dialect performance.'),\n    (4, 'French', 'fr', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Shopping, function calls, retrieval baselines, and a current document component. Retest full business execution on the target model.'),\n    (5, 'Bengali', 'bn', 'MASSIVE calls; Voice tools; MAST retrieval', 'Historical calls, voice-tool tasks, and retrieval resources. Current-frontier execution remains unestablished; test entities and arguments.'),\n    (6, 'Portuguese', 'pt', 'MASSIVE calls; GAIA audited', 'Audited GAIA uses Brazilian Portuguese; MASSIVE uses Portugal locale. Regional business rules are not interchangeable.'),\n    (7, 'Indonesian', 'id', 'MASSIVE calls; SEA multi-turn', 'SEA covers full multi-turn workflows with domain-dependent effects. Indonesian does not substitute for Javanese.'),\n    (8, 'Urdu', 'ur', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Historical/non-frontier calls, shopping, and retrieval. Numerals and dates are proposed mechanisms, not measured current-model failure rates.'),\n    (9, 'Russian', 'ru', 'MASSIVE calls; X-Web shopping; macOS use; MAST retrieval', 'Current code execution and static calls contain failures. Without matched English tasks, Russian causation is unestablished.'),\n    (10, 'German', 'de', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval', 'Audited GAIA and native code coverage are relatively rich. Gaps differ by model; component strength is not deployment acceptance.'),\n    (11, 'Japanese', 'ja', 'MASSIVE calls; macOS use; LILT native code', 'Current parsing is low; historical CUA gaps change direction by model. Test visual targeting, text, and local software rules separately.'),\n    (12, 'Nigerian Pidgin', 'pcm', 'SwitchBoard speech component', 'Code-switching speech resources found; current-frontier tool execution unverified. Do not substitute English.'),\n    (13, 'Egyptian Arabic', 'arz', 'No matching execution result included', 'Generic Arabic studies cannot fill this dialect. No included exact-dialect agent result.'),\n    (14, 'Marathi', 'mr', 'Voice tools', 'VoiceAgentBench provides speech-tool tasks; do not extend its English-only multi-turn results to Marathi.'),\n    (15, 'Vietnamese', 'vi', 'MASSIVE calls; X-Web shopping; SEA multi-turn', 'SEA, shopping, function calls, and documents provide multiple sources. Kimi fully localized retail slightly exceeds English: preserve the counterexample.'),\n    (16, 'Telugu', 'te', 'MASSIVE calls; Voice tools; MAST retrieval', 'Calls, voice tools, and retrieval resources are reusable. Static calls do not establish long-horizon execution.'),\n    (17, 'Swahili', 'sw', 'MASSIVE calls; X-Web shopping; MAST retrieval; MultiAgent-X tasks', 'Historical task coverage, but GPT-4o original-language shopping is not weak. MAST small-model scores cannot rank frontier risk.'),\n    (18, 'Hausa', 'ha', 'MultiAgent-X tasks', 'MultiAgent-X supplies synthetic calls. Exploratory safety logs were excluded from reliable risk assessment; no current-frontier baseline.'),\n    (19, 'Turkish', 'tr', 'MASSIVE calls; X-Web shopping; LILT native code', 'Calls, shopping, and native code tasks exist. Add current-model end-to-end and local-rule tests.'),\n    (20, 'Western Punjabi', 'pnb', 'No matching execution result included', 'MultiAgent-X and MAST pa/Gurmukhi cannot substitute for Western Punjabi; exact matching results remain missing.'),\n    (21, 'Tagalog', 'tl', 'MASSIVE calls; SEA multi-turn', 'MASSIVE tl-PH and SEA Filipino are related labels. Confirm the user variety, then test multi-turn business execution.'),\n    (22, 'Tamil', 'ta', 'MASSIVE calls; Voice tools; MAST retrieval', 'Calls, voice-tool tasks, and MAST retrieval are reusable. Current-frontier complete execution remains unestablished.'),\n    (23, 'Iranian Persian', 'fa', 'MASSIVE calls', 'MASSIVE fa-IR provides historical calls. Current models still need Persian numerals, calendars, and identifier-preservation tests.'),\n    (24, 'Korean', 'ko', 'MASSIVE calls; GAIA audited; LILT native code', 'Audited GAIA improves and K3 parsing is relatively high. LILT multi-turn results have simulator confounds; broad reliability is unestablished.'),\n    (25, 'Amharic', 'am', 'MASSIVE calls; MultiAgent-X tasks', 'MASSIVE historical weakness is specific; MultiAgent-X supplies resources, not independent performance replication. Retest current flagships.'),\n    (26, 'Thai', 'th', 'MASSIVE calls; X-Web shopping; SEA multi-turn; MAST retrieval', 'Historical SEA and shopping weaknesses plus low K3 parsing motivate retesting. They do not imply every model/task is weak.'),\n    (27, 'Javanese', 'jv', 'MASSIVE calls', 'MASSIVE jv-ID provides historical calls. The missing-evidence claim concerns current execution; Indonesian is not a substitute.'),\n    (28, 'Italian', 'it', 'MASSIVE calls', 'MASSIVE, MAPS parent coverage, and current documents exist. Highest parsing does not imply the most reliable complete agent.'),\n    (29, 'Gujarati', 'gu', 'MAST retrieval', 'MAST supplies retrieval tasks and a non-frontier baseline. Current-frontier tools/business execution evidence remains insufficient.'),\n    (30, 'Kannada', 'kn', 'MASSIVE calls; MAST retrieval', 'Claude 3.5 is historically low on MASSIVE; MAST adds retrieval resources. No current-frontier multi-task replication.')\n)\nSELECT * FROM reviewed_published_rows;",
          "engine": "SQLite (reviewed published observations)",
          "language": "sql",
          "tables_used": [
            "All 30 languages: evidence and next checks"
          ],
          "filters": [
            "Evidence cutoff September 5, 2026; no new model runs",
            "Current models, earlier flagships, and resources separated"
          ],
          "metric_definitions": [
            "Researcher mapping retains all target languages, named evidence and inference limits. No invented scores; related labels do not replace exact language varieties."
          ]
        }
      },
      {
        "id": "expanded-link-32",
        "label": "sea_code",
        "href": "https://github.com/SEACrowd/SEATauBench",
        "query": {
          "url": "https://github.com/SEACrowd/SEATauBench",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "sea_code"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-33",
        "label": "gaia_code",
        "href": "https://github.com/lilt/gaia-v2-lilt",
        "query": {
          "url": "https://github.com/lilt/gaia-v2-lilt",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "gaia_code"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-34",
        "label": "maps",
        "href": "https://aclanthology.org/2026.findings-eacl.42/",
        "query": {
          "url": "https://aclanthology.org/2026.findings-eacl.42/",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "maps"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-35",
        "label": "maps_data",
        "href": "https://huggingface.co/datasets/Fujitsu-FRE/MAPS",
        "query": {
          "url": "https://huggingface.co/datasets/Fujitsu-FRE/MAPS",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "maps_data"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-36",
        "label": "mac_code",
        "href": "https://github.com/showlab/macosworld",
        "query": {
          "url": "https://github.com/showlab/macosworld",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "mac_code"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-37",
        "label": "xweb_code",
        "href": "https://github.com/WPENGxs/X-WebAgentBench",
        "query": {
          "url": "https://github.com/WPENGxs/X-WebAgentBench",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "xweb_code"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-38",
        "label": "massive_data",
        "href": "https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id",
        "query": {
          "url": "https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "massive_data"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-39",
        "label": "voice",
        "href": "https://arxiv.org/html/2510.07978v1",
        "query": {
          "url": "https://arxiv.org/html/2510.07978v1",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "voice"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-40",
        "label": "voice_code",
        "href": "https://github.com/ola-krutrim/VoiceAgentBench",
        "query": {
          "url": "https://github.com/ola-krutrim/VoiceAgentBench",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "voice_code"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-41",
        "label": "telco",
        "href": "https://arxiv.org/html/2604.06209v1",
        "query": {
          "url": "https://arxiv.org/html/2604.06209v1",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "telco"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-42",
        "label": "telco_code",
        "href": "https://github.com/BrahiM-Mefgouda/TelcoAgent",
        "query": {
          "url": "https://github.com/BrahiM-Mefgouda/TelcoAgent",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "telco_code"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-43",
        "label": "terminal",
        "href": "https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark",
        "query": {
          "url": "https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "terminal"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-44",
        "label": "lilt",
        "href": "https://benchmarks.lilt.com/",
        "query": {
          "url": "https://benchmarks.lilt.com/",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "lilt"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-45",
        "label": "liltbench",
        "href": "https://github.com/lilt/liltbench-tasks-public",
        "query": {
          "url": "https://github.com/lilt/liltbench-tasks-public",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "liltbench"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-46",
        "label": "multi",
        "href": "https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark",
        "query": {
          "url": "https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "multi"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-47",
        "label": "mast",
        "href": "https://mast-benchmark.github.io/",
        "query": {
          "url": "https://mast-benchmark.github.io/",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "mast"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-48",
        "label": "judge",
        "href": "https://arxiv.org/html/2604.04532v1",
        "query": {
          "url": "https://arxiv.org/html/2604.04532v1",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "judge"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-49",
        "label": "switch",
        "href": "https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
        "query": {
          "url": "https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "switch"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-50",
        "label": "GAIA-v2-LILT Table 2",
        "href": "https://arxiv.org/html/2604.24929v1#S6",
        "query": {
          "url": "https://arxiv.org/html/2604.24929v1#S6",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "GAIA-v2-LILT Table 2"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-51",
        "label": "SEATauBench Tables 9/10/13",
        "href": "https://arxiv.org/html/2606.28715v1#A6",
        "query": {
          "url": "https://arxiv.org/html/2606.28715v1#A6",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "SEATauBench Tables 9/10/13"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-52",
        "label": "macOSWorld v4 Table 3 and cases",
        "href": "https://arxiv.org/html/2506.04135v4#S5",
        "query": {
          "url": "https://arxiv.org/html/2506.04135v4#S5",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "macOSWorld v4 Table 3 and cases"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      },
      {
        "id": "expanded-link-53",
        "label": "X-WebAgentBench Table 2",
        "href": "https://arxiv.org/html/2505.15372v1#S3",
        "query": {
          "url": "https://arxiv.org/html/2505.15372v1#S3",
          "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
          "tables_used": [
            "X-WebAgentBench Table 2"
          ]
        },
        "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
      }
    ]
  },
  "snapshot": {
    "version": 1,
    "status": "ready",
    "generatedAt": "2026-09-05",
    "datasets": {
      "format_partition": [
        {
          "Model": "GPT-5.6 Sol",
          "Full sample passed": 0.691,
          "Valid format, sample failed": 0.308,
          "Invalid format": 0.001,
          "Submission date": "2026-08-18",
          "Evaluated N": "Not separately disclosed",
          "Configuration limits": "openrouter / 9ebf388",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/8",
          "Documented dataset N": 1169,
          "Scope": "All call, abstention, and clarification items"
        },
        {
          "Model": "Claude Opus 5",
          "Full sample passed": 0.677,
          "Valid format, sample failed": 0.272,
          "Invalid format": 0.051,
          "Submission date": "2026-08-18",
          "Evaluated N": "Not separately disclosed",
          "Configuration limits": "openrouter / 0e1f484",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/1",
          "Documented dataset N": 1169,
          "Scope": "All call, abstention, and clarification items"
        },
        {
          "Model": "Grok 4.6",
          "Full sample passed": 0.737,
          "Valid format, sample failed": 0.263,
          "Invalid format": 0,
          "Submission date": "2026-08-18",
          "Evaluated N": "Not separately disclosed",
          "Configuration limits": "local-chat-completions / 0e1f484",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/11",
          "Documented dataset N": 1169,
          "Scope": "All call, abstention, and clarification items"
        }
      ],
      "argument_partition": [
        {
          "Model": "GPT-5.6 Sol",
          "Tools and arguments matched": 0.702,
          "Tools matched, arguments incomplete": 0.163,
          "Tools not fully matched": 0.135,
          "Submission date": "2026-08-18",
          "Evaluated N": "Not separately disclosed",
          "Configuration limits": "openrouter / 9ebf388",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/8",
          "Documented dataset N": 1032,
          "Scope": "Call-required items only"
        },
        {
          "Model": "Claude Opus 5",
          "Tools and arguments matched": 0.677,
          "Tools matched, arguments incomplete": 0.113,
          "Tools not fully matched": 0.21,
          "Submission date": "2026-08-18",
          "Evaluated N": "Not separately disclosed",
          "Configuration limits": "openrouter / 0e1f484",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/1",
          "Documented dataset N": 1032,
          "Scope": "Call-required items only"
        },
        {
          "Model": "Grok 4.6",
          "Tools and arguments matched": 0.743,
          "Tools matched, arguments incomplete": 0.081,
          "Tools not fully matched": 0.176,
          "Submission date": "2026-08-18",
          "Evaluated N": "Not separately disclosed",
          "Configuration limits": "local-chat-completions / 0e1f484",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/11",
          "Documented dataset N": 1032,
          "Scope": "Call-required items only"
        }
      ],
      "decision_errors": [
        {
          "Required action": "Abstention required (127)",
          "Model": "GPT-5.6 Sol",
          "Conditional error rate": 0.417,
          "Documented denominator": 127,
          "Scoring interpretation": "Correct abstention missing",
          "Evaluated N": "Not separately disclosed",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/8"
        },
        {
          "Required action": "Call required (1,032)",
          "Model": "GPT-5.6 Sol",
          "Conditional error rate": 0.009,
          "Documented denominator": 1032,
          "Scoring interpretation": "False abstention",
          "Evaluated N": "Not separately disclosed",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/8"
        },
        {
          "Required action": "Abstention required (127)",
          "Model": "Claude Opus 5",
          "Conditional error rate": 0.354,
          "Documented denominator": 127,
          "Scoring interpretation": "Correct abstention missing",
          "Evaluated N": "Not separately disclosed",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/1"
        },
        {
          "Required action": "Call required (1,032)",
          "Model": "Claude Opus 5",
          "Conditional error rate": 0.001,
          "Documented denominator": 1032,
          "Scoring interpretation": "False abstention",
          "Evaluated N": "Not separately disclosed",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/1"
        },
        {
          "Required action": "Abstention required (127)",
          "Model": "Grok 4.6",
          "Conditional error rate": 0.339,
          "Documented denominator": 127,
          "Scoring interpretation": "Correct abstention missing",
          "Evaluated N": "Not separately disclosed",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/11"
        },
        {
          "Required action": "Call required (1,032)",
          "Model": "Grok 4.6",
          "Conditional error rate": 0.004,
          "Documented denominator": 1032,
          "Scoring interpretation": "False abstention",
          "Evaluated N": "Not separately disclosed",
          "Source": "https://mera.a-ai.ru/en/text/submits/2.0/11"
        }
      ],
      "execution_partition": [
        {
          "Configuration": "GPT-5.6 Sol (23 tasks × 3)",
          "Audit-retained passes": 0.7101449275362319,
          "Raw passes removed by audit": 0.11594202898550725,
          "Original failures": 0.17391304347826086,
          "Retained pass count": 49,
          "Removed pass count": 8,
          "Original failure count": 12,
          "Total runs": 69,
          "Distinct tasks": 23,
          "Harness": "Codex CLI",
          "Setting": "xhigh",
          "Source": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md"
        },
        {
          "Configuration": "Claude Opus 5 (20 tasks × 3)",
          "Audit-retained passes": 0.7666666666666667,
          "Raw passes removed by audit": 0.016666666666666666,
          "Original failures": 0.21666666666666667,
          "Retained pass count": 46,
          "Removed pass count": 1,
          "Original failure count": 13,
          "Total runs": 60,
          "Distinct tasks": 20,
          "Harness": "Claude Code",
          "Setting": "xhigh",
          "Source": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md"
        }
      ],
      "document_profile": [
        {
          "Language": "Thai",
          "Code": "th",
          "Parsing quality score": 72.1,
          "Observed position": "Lower four observations",
          "Script group": "Other scripts",
          "Order": 1,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Japanese",
          "Code": "ja",
          "Parsing quality score": 74.9,
          "Observed position": "Lower four observations",
          "Script group": "Other scripts",
          "Order": 2,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Arabic (unspecified variety)",
          "Code": "ar",
          "Parsing quality score": 77.4,
          "Observed position": "Lower four observations",
          "Script group": "Other scripts",
          "Order": 3,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Parsing quality score": 77.5,
          "Observed position": "Lower four observations",
          "Script group": "Other scripts",
          "Order": 4,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "French",
          "Code": "fr",
          "Parsing quality score": 80.0,
          "Observed position": "Other nine observations",
          "Script group": "Latin script",
          "Order": 5,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Spanish",
          "Code": "es",
          "Parsing quality score": 80.2,
          "Observed position": "Other nine observations",
          "Script group": "Latin script",
          "Order": 6,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Russian",
          "Code": "ru",
          "Parsing quality score": 82.4,
          "Observed position": "Other nine observations",
          "Script group": "Other scripts",
          "Order": 7,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Vietnamese",
          "Code": "vi",
          "Parsing quality score": 84.8,
          "Observed position": "Other nine observations",
          "Script group": "Latin script",
          "Order": 8,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Indonesian",
          "Code": "id",
          "Parsing quality score": 86.9,
          "Observed position": "Other nine observations",
          "Script group": "Latin script",
          "Order": 9,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Portuguese",
          "Code": "pt",
          "Parsing quality score": 88.9,
          "Observed position": "Other nine observations",
          "Script group": "Latin script",
          "Order": 10,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "German",
          "Code": "de",
          "Parsing quality score": 89.1,
          "Observed position": "Other nine observations",
          "Script group": "Latin script",
          "Order": 11,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Korean",
          "Code": "ko",
          "Parsing quality score": 89.9,
          "Observed position": "Other nine observations",
          "Script group": "Other scripts",
          "Order": 12,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Language": "Italian",
          "Code": "it",
          "Parsing quality score": 92.7,
          "Observed position": "Other nine observations",
          "Script group": "Latin script",
          "Order": 13,
          "Median across 13 observations": 82.4,
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        }
      ],
      "script_distribution": [
        {
          "Script group": "Latin script (7)",
          "Minimum": 80.0,
          "First quartile": 82.5,
          "Median": 86.9,
          "Third quartile": 89.0,
          "Maximum": 92.7,
          "Observation count": 7,
          "Mean": 86.08571428571429,
          "Language": "French; Spanish; Vietnamese; Indonesian; Portuguese; German; Italian",
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "Script group": "Other scripts (6)",
          "Minimum": 72.1,
          "First quartile": 75.525,
          "Median": 77.45,
          "Third quartile": 81.17500000000001,
          "Maximum": 89.9,
          "Observation count": 6,
          "Mean": 79.03333333333333,
          "Language": "Thai; Japanese; Arabic (unspecified variety); Hindi; Russian; Korean",
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        }
      ],
      "atlas_1": [
        {
          "Language": "01 Hindi",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "02 Spanish",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "03 Modern Standard Arabic",
          "Document component": 0.5,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "04 French",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "05 Bengali",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "06 Portuguese",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "07 Indonesian",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "08 Urdu",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "09 Russian",
          "Document component": 1,
          "Static tool plan": 1,
          "End-to-end execution": 1,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "10 German",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "11 Japanese",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "12 Nigerian Pidgin",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "13 Egyptian Arabic",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "14 Marathi",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "15 Vietnamese",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        }
      ],
      "atlas_2": [
        {
          "Language": "16 Telugu",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "17 Swahili",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "18 Hausa",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "19 Turkish",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "20 Western Punjabi",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "21 Tagalog",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "22 Tamil",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "23 Iranian Persian",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "24 Korean",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "25 Amharic",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "26 Thai",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "27 Javanese",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "28 Italian",
          "Document component": 1,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "29 Gujarati",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        },
        {
          "Language": "30 Kannada",
          "Document component": 0,
          "Static tool plan": 0,
          "End-to-end execution": 0,
          "Matched language control": 0,
          "Marker definition": "1=matched evidence; 0.5=related label; 0=no included result; not ability"
        }
      ],
      "action_map": [
        {
          "Test target": "Amounts | Urdu / Persian",
          "Input condition": "۱٬۲۵۰٫۵۰; explicit currency",
          "Expected final result": "Pass 1250.5 according to the API contract; preserve currency"
        },
        {
          "Test target": "Entity IDs | all languages",
          "Input condition": "INV-01250/AB; mixed scripts",
          "Expected final result": "Preserve string, leading zeros, separators; act on the correct entity"
        },
        {
          "Test target": "Dates | explicit locale / calendar",
          "Input condition": "05/09/2026; with/without locale control",
          "Expected final result": "Convert correctly when defined; clarify per rules when ambiguous"
        },
        {
          "Test target": "Dialect / neighboring language | four gaps",
          "Input condition": "Native pcm, arz, pnb, jv requests",
          "Expected final result": "Preserve native constraints; no substitution by neighboring-language scores"
        }
      ],
      "methods": [
        {
          "Evidence": "GorillaHard",
          "Measurement": "Static output; documented N=1,169 total, 1,032 calls, 127 abstentions",
          "Key limitation": "Documented counts, not separately disclosed run N. Incomplete settings; Sol source commit unresolved; no English pairs."
        },
        {
          "Evidence": "MDPBench",
          "Measurement": "Kimi K3 document parsing quality; 13 target observations",
          "Key limitation": "Unpaired content/layout; unspecified Arabic variety; composite score is not failure probability."
        },
        {
          "Evidence": "RuBench",
          "Measurement": "Russian instructions → repository edits → regression tests",
          "Key limitation": "Sol 23×3, Opus 20×3; different sets/harnesses. Earlier Fable fallback omitted."
        }
      ],
      "language_lookup": [
        {
          "No.": 1,
          "Language": "Hindi",
          "Code": "hi",
          "Available evidence": "MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks",
          "Interpretation and next check": "Audited GAIA recovers substantially; K3 reading is low. Separate benchmark quality from local-rule execution."
        },
        {
          "No.": 2,
          "Language": "Spanish",
          "Code": "es",
          "Available evidence": "MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval",
          "Interpretation and next check": "GAIA parent set, shopping, function calls, and retrieval coverage. No complete current-frontier comparison; do not presume European-language reliability."
        },
        {
          "No.": 3,
          "Language": "Modern Standard Arabic",
          "Code": "ar",
          "Available evidence": "MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval",
          "Interpretation and next check": "GAIA and macOS give aligned historical signals; K3 component score is lower. Generic Arabic does not precisely establish MSA or Egyptian dialect performance."
        },
        {
          "No.": 4,
          "Language": "French",
          "Code": "fr",
          "Available evidence": "MASSIVE calls; X-Web shopping; MAST retrieval",
          "Interpretation and next check": "Shopping, function calls, retrieval baselines, and a current document component. Retest full business execution on the target model."
        },
        {
          "No.": 5,
          "Language": "Bengali",
          "Code": "bn",
          "Available evidence": "MASSIVE calls; Voice tools; MAST retrieval",
          "Interpretation and next check": "Historical calls, voice-tool tasks, and retrieval resources. Current-frontier execution remains unestablished; test entities and arguments."
        },
        {
          "No.": 6,
          "Language": "Portuguese",
          "Code": "pt",
          "Available evidence": "MASSIVE calls; GAIA audited",
          "Interpretation and next check": "Audited GAIA uses Brazilian Portuguese; MASSIVE uses Portugal locale. Regional business rules are not interchangeable."
        },
        {
          "No.": 7,
          "Language": "Indonesian",
          "Code": "id",
          "Available evidence": "MASSIVE calls; SEA multi-turn",
          "Interpretation and next check": "SEA covers full multi-turn workflows with domain-dependent effects. Indonesian does not substitute for Javanese."
        },
        {
          "No.": 8,
          "Language": "Urdu",
          "Code": "ur",
          "Available evidence": "MASSIVE calls; X-Web shopping; MAST retrieval",
          "Interpretation and next check": "Historical/non-frontier calls, shopping, and retrieval. Numerals and dates are proposed mechanisms, not measured current-model failure rates."
        },
        {
          "No.": 9,
          "Language": "Russian",
          "Code": "ru",
          "Available evidence": "MASSIVE calls; X-Web shopping; macOS use; MAST retrieval",
          "Interpretation and next check": "Current code execution and static calls contain failures. Without matched English tasks, Russian causation is unestablished."
        },
        {
          "No.": 10,
          "Language": "German",
          "Code": "de",
          "Available evidence": "MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval",
          "Interpretation and next check": "Audited GAIA and native code coverage are relatively rich. Gaps differ by model; component strength is not deployment acceptance."
        },
        {
          "No.": 11,
          "Language": "Japanese",
          "Code": "ja",
          "Available evidence": "MASSIVE calls; macOS use; LILT native code",
          "Interpretation and next check": "Current parsing is low; historical CUA gaps change direction by model. Test visual targeting, text, and local software rules separately."
        },
        {
          "No.": 12,
          "Language": "Nigerian Pidgin",
          "Code": "pcm",
          "Available evidence": "SwitchBoard speech component",
          "Interpretation and next check": "Code-switching speech resources found; current-frontier tool execution unverified. Do not substitute English."
        },
        {
          "No.": 13,
          "Language": "Egyptian Arabic",
          "Code": "arz",
          "Available evidence": "No matching execution result included",
          "Interpretation and next check": "Generic Arabic studies cannot fill this dialect. No included exact-dialect agent result."
        },
        {
          "No.": 14,
          "Language": "Marathi",
          "Code": "mr",
          "Available evidence": "Voice tools",
          "Interpretation and next check": "VoiceAgentBench provides speech-tool tasks; do not extend its English-only multi-turn results to Marathi."
        },
        {
          "No.": 15,
          "Language": "Vietnamese",
          "Code": "vi",
          "Available evidence": "MASSIVE calls; X-Web shopping; SEA multi-turn",
          "Interpretation and next check": "SEA, shopping, function calls, and documents provide multiple sources. Kimi fully localized retail slightly exceeds English: preserve the counterexample."
        },
        {
          "No.": 16,
          "Language": "Telugu",
          "Code": "te",
          "Available evidence": "MASSIVE calls; Voice tools; MAST retrieval",
          "Interpretation and next check": "Calls, voice tools, and retrieval resources are reusable. Static calls do not establish long-horizon execution."
        },
        {
          "No.": 17,
          "Language": "Swahili",
          "Code": "sw",
          "Available evidence": "MASSIVE calls; X-Web shopping; MAST retrieval; MultiAgent-X tasks",
          "Interpretation and next check": "Historical task coverage, but GPT-4o original-language shopping is not weak. MAST small-model scores cannot rank frontier risk."
        },
        {
          "No.": 18,
          "Language": "Hausa",
          "Code": "ha",
          "Available evidence": "MultiAgent-X tasks",
          "Interpretation and next check": "MultiAgent-X supplies synthetic calls. Exploratory safety logs were excluded from reliable risk assessment; no current-frontier baseline."
        },
        {
          "No.": 19,
          "Language": "Turkish",
          "Code": "tr",
          "Available evidence": "MASSIVE calls; X-Web shopping; LILT native code",
          "Interpretation and next check": "Calls, shopping, and native code tasks exist. Add current-model end-to-end and local-rule tests."
        },
        {
          "No.": 20,
          "Language": "Western Punjabi",
          "Code": "pnb",
          "Available evidence": "No matching execution result included",
          "Interpretation and next check": "MultiAgent-X and MAST pa/Gurmukhi cannot substitute for Western Punjabi; exact matching results remain missing."
        },
        {
          "No.": 21,
          "Language": "Tagalog",
          "Code": "tl",
          "Available evidence": "MASSIVE calls; SEA multi-turn",
          "Interpretation and next check": "MASSIVE tl-PH and SEA Filipino are related labels. Confirm the user variety, then test multi-turn business execution."
        },
        {
          "No.": 22,
          "Language": "Tamil",
          "Code": "ta",
          "Available evidence": "MASSIVE calls; Voice tools; MAST retrieval",
          "Interpretation and next check": "Calls, voice-tool tasks, and MAST retrieval are reusable. Current-frontier complete execution remains unestablished."
        },
        {
          "No.": 23,
          "Language": "Iranian Persian",
          "Code": "fa",
          "Available evidence": "MASSIVE calls",
          "Interpretation and next check": "MASSIVE fa-IR provides historical calls. Current models still need Persian numerals, calendars, and identifier-preservation tests."
        },
        {
          "No.": 24,
          "Language": "Korean",
          "Code": "ko",
          "Available evidence": "MASSIVE calls; GAIA audited; LILT native code",
          "Interpretation and next check": "Audited GAIA improves and K3 parsing is relatively high. LILT multi-turn results have simulator confounds; broad reliability is unestablished."
        },
        {
          "No.": 25,
          "Language": "Amharic",
          "Code": "am",
          "Available evidence": "MASSIVE calls; MultiAgent-X tasks",
          "Interpretation and next check": "MASSIVE historical weakness is specific; MultiAgent-X supplies resources, not independent performance replication. Retest current flagships."
        },
        {
          "No.": 26,
          "Language": "Thai",
          "Code": "th",
          "Available evidence": "MASSIVE calls; X-Web shopping; SEA multi-turn; MAST retrieval",
          "Interpretation and next check": "Historical SEA and shopping weaknesses plus low K3 parsing motivate retesting. They do not imply every model/task is weak."
        },
        {
          "No.": 27,
          "Language": "Javanese",
          "Code": "jv",
          "Available evidence": "MASSIVE calls",
          "Interpretation and next check": "MASSIVE jv-ID provides historical calls. The missing-evidence claim concerns current execution; Indonesian is not a substitute."
        },
        {
          "No.": 28,
          "Language": "Italian",
          "Code": "it",
          "Available evidence": "MASSIVE calls",
          "Interpretation and next check": "MASSIVE, MAPS parent coverage, and current documents exist. Highest parsing does not imply the most reliable complete agent."
        },
        {
          "No.": 29,
          "Language": "Gujarati",
          "Code": "gu",
          "Available evidence": "MAST retrieval",
          "Interpretation and next check": "MAST supplies retrieval tasks and a non-frontier baseline. Current-frontier tools/business execution evidence remains insufficient."
        },
        {
          "No.": 30,
          "Language": "Kannada",
          "Code": "kn",
          "Available evidence": "MASSIVE calls; MAST retrieval",
          "Interpretation and next check": "Claude 3.5 is historically low on MASSIVE; MAST adds retrieval resources. No current-frontier multi-task replication."
        }
      ],
      "gaia_lift": [
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Model": "GPT-5.4",
          "Before audit": 32.1,
          "After audit": 47.3,
          "English reference": 66.7,
          "Audit improvement (pp)": 15.2,
          "English minus audited (pp)": 19.4,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "German",
          "Code": "de",
          "Model": "GPT-5.4",
          "Before audit": 47.3,
          "After audit": 63.6,
          "English reference": 66.7,
          "Audit improvement (pp)": 16.3,
          "English minus audited (pp)": 3.1,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Model": "GPT-5.4",
          "Before audit": 34.6,
          "After audit": 60.0,
          "English reference": 66.7,
          "Audit improvement (pp)": 25.4,
          "English minus audited (pp)": 6.7,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Korean",
          "Code": "ko",
          "Model": "GPT-5.4",
          "Before audit": 33.3,
          "After audit": 62.4,
          "English reference": 66.7,
          "Audit improvement (pp)": 29.1,
          "English minus audited (pp)": 4.3,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Portuguese (Brazil)",
          "Code": "pt",
          "Model": "GPT-5.4",
          "Before audit": 47.3,
          "After audit": 58.2,
          "English reference": 66.7,
          "Audit improvement (pp)": 10.9,
          "English minus audited (pp)": 8.5,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 34.6,
          "After audit": 52.1,
          "English reference": 73.9,
          "Audit improvement (pp)": 17.5,
          "English minus audited (pp)": 21.8,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "German",
          "Code": "de",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 49.7,
          "After audit": 66.7,
          "English reference": 73.9,
          "Audit improvement (pp)": 17.0,
          "English minus audited (pp)": 7.2,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 38.2,
          "After audit": 63.6,
          "English reference": 73.9,
          "Audit improvement (pp)": 25.4,
          "English minus audited (pp)": 10.3,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Korean",
          "Code": "ko",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 34.6,
          "After audit": 64.8,
          "English reference": 73.9,
          "Audit improvement (pp)": 30.2,
          "English minus audited (pp)": 9.1,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Portuguese (Brazil)",
          "Code": "pt",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 48.5,
          "After audit": 65.5,
          "English reference": 73.9,
          "Audit improvement (pp)": 17.0,
          "English minus audited (pp)": 8.4,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Model": "Claude Opus 4.6",
          "Before audit": 32.1,
          "After audit": 49.1,
          "English reference": 79.4,
          "Audit improvement (pp)": 17.0,
          "English minus audited (pp)": 30.3,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "German",
          "Code": "de",
          "Model": "Claude Opus 4.6",
          "Before audit": 49.7,
          "After audit": 66.7,
          "English reference": 79.4,
          "Audit improvement (pp)": 17.0,
          "English minus audited (pp)": 12.7,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Model": "Claude Opus 4.6",
          "Before audit": 29.7,
          "After audit": 62.4,
          "English reference": 79.4,
          "Audit improvement (pp)": 32.7,
          "English minus audited (pp)": 17.0,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Korean",
          "Code": "ko",
          "Model": "Claude Opus 4.6",
          "Before audit": 33.3,
          "After audit": 58.8,
          "English reference": 79.4,
          "Audit improvement (pp)": 25.5,
          "English minus audited (pp)": 20.6,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Portuguese (Brazil)",
          "Code": "pt",
          "Model": "Claude Opus 4.6",
          "Before audit": 49.1,
          "After audit": 63.0,
          "English reference": 79.4,
          "Audit improvement (pp)": 13.9,
          "English minus audited (pp)": 16.4,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        }
      ],
      "gaia_residual": [
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Model": "GPT-5.4",
          "Before audit": 32.1,
          "After audit": 47.3,
          "English reference": 66.7,
          "Audit improvement (pp)": 15.2,
          "English minus audited (pp)": 19.4,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "German",
          "Code": "de",
          "Model": "GPT-5.4",
          "Before audit": 47.3,
          "After audit": 63.6,
          "English reference": 66.7,
          "Audit improvement (pp)": 16.3,
          "English minus audited (pp)": 3.1,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Model": "GPT-5.4",
          "Before audit": 34.6,
          "After audit": 60.0,
          "English reference": 66.7,
          "Audit improvement (pp)": 25.4,
          "English minus audited (pp)": 6.7,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Korean",
          "Code": "ko",
          "Model": "GPT-5.4",
          "Before audit": 33.3,
          "After audit": 62.4,
          "English reference": 66.7,
          "Audit improvement (pp)": 29.1,
          "English minus audited (pp)": 4.3,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Portuguese (Brazil)",
          "Code": "pt",
          "Model": "GPT-5.4",
          "Before audit": 47.3,
          "After audit": 58.2,
          "English reference": 66.7,
          "Audit improvement (pp)": 10.9,
          "English minus audited (pp)": 8.5,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 34.6,
          "After audit": 52.1,
          "English reference": 73.9,
          "Audit improvement (pp)": 17.5,
          "English minus audited (pp)": 21.8,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "German",
          "Code": "de",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 49.7,
          "After audit": 66.7,
          "English reference": 73.9,
          "Audit improvement (pp)": 17.0,
          "English minus audited (pp)": 7.2,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 38.2,
          "After audit": 63.6,
          "English reference": 73.9,
          "Audit improvement (pp)": 25.4,
          "English minus audited (pp)": 10.3,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Korean",
          "Code": "ko",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 34.6,
          "After audit": 64.8,
          "English reference": 73.9,
          "Audit improvement (pp)": 30.2,
          "English minus audited (pp)": 9.1,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Portuguese (Brazil)",
          "Code": "pt",
          "Model": "Gemini 3.1 Pro",
          "Before audit": 48.5,
          "After audit": 65.5,
          "English reference": 73.9,
          "Audit improvement (pp)": 17.0,
          "English minus audited (pp)": 8.4,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Model": "Claude Opus 4.6",
          "Before audit": 32.1,
          "After audit": 49.1,
          "English reference": 79.4,
          "Audit improvement (pp)": 17.0,
          "English minus audited (pp)": 30.3,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "German",
          "Code": "de",
          "Model": "Claude Opus 4.6",
          "Before audit": 49.7,
          "After audit": 66.7,
          "English reference": 79.4,
          "Audit improvement (pp)": 17.0,
          "English minus audited (pp)": 12.7,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Model": "Claude Opus 4.6",
          "Before audit": 29.7,
          "After audit": 62.4,
          "English reference": 79.4,
          "Audit improvement (pp)": 32.7,
          "English minus audited (pp)": 17.0,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Korean",
          "Code": "ko",
          "Model": "Claude Opus 4.6",
          "Before audit": 33.3,
          "After audit": 58.8,
          "English reference": 79.4,
          "Audit improvement (pp)": 25.5,
          "English minus audited (pp)": 20.6,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "Language": "Portuguese (Brazil)",
          "Code": "pt",
          "Model": "Claude Opus 4.6",
          "Before audit": 49.1,
          "After audit": 63.0,
          "English reference": 79.4,
          "Audit improvement (pp)": 13.9,
          "English minus audited (pp)": 16.4,
          "Tasks per language": 165,
          "Model cohort": "Earlier flagship; outside September current cohort",
          "Execution harness": "Open Deep Research; 12 manager / 20 search steps",
          "Source": "https://arxiv.org/html/2604.24929v1"
        }
      ],
      "sea_localization": [
        {
          "Language": "Vietnamese",
          "Code": "vi",
          "Setting": "All English",
          "Task success rate": 0.561,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.56,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.699,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Vietnamese",
          "Code": "vi",
          "Setting": "Dialogue localized",
          "Task success rate": 0.687,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.56,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.699,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Vietnamese",
          "Code": "vi",
          "Setting": "Fully localized business",
          "Task success rate": 0.567,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.56,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.699,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Thai",
          "Code": "th",
          "Setting": "All English",
          "Task success rate": 0.561,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.547,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.693,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Thai",
          "Code": "th",
          "Setting": "Dialogue localized",
          "Task success rate": 0.573,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.547,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.693,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Thai",
          "Code": "th",
          "Setting": "Fully localized business",
          "Task success rate": 0.327,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.547,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.693,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Indonesian",
          "Code": "id",
          "Setting": "All English",
          "Task success rate": 0.561,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.6,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.798,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Indonesian",
          "Code": "id",
          "Setting": "Dialogue localized",
          "Task success rate": 0.64,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.6,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.798,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Indonesian",
          "Code": "id",
          "Setting": "Fully localized business",
          "Task success rate": 0.433,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.6,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.798,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Filipino†",
          "Code": "tl",
          "Setting": "All English",
          "Task success rate": 0.561,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.6,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.743,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Filipino†",
          "Code": "tl",
          "Setting": "Dialogue localized",
          "Task success rate": 0.675,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.6,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.743,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        },
        {
          "Language": "Filipino†",
          "Code": "tl",
          "Setting": "Fully localized business",
          "Task success rate": 0.444,
          "Model": "Kimi K2.5",
          "Displayed domain": "Retail",
          "Trials per task": 3,
          "Airline English": 0.707,
          "Airline fully localized": 0.6,
          "Telecom English": 0.997,
          "Telecom fully localized": 0.743,
          "Sample definition": "Separate domains; no pooled mean; three trials per task",
          "Source": "https://arxiv.org/html/2606.28715v1"
        }
      ],
      "macos_language": [
        {
          "Language": "English reference",
          "Code": "en",
          "Agent": "Claude CUA",
          "Task success rate": 0.444,
          "Exact model version": "claude-3-7-sonnet-20250219",
          "Task count": 171,
          "Difference from English (pp)": 0.0,
          "UI and instruction languages": "Both switched to this language",
          "Source": "https://arxiv.org/html/2506.04135v4"
        },
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Agent": "Claude CUA",
          "Task success rate": 0.316,
          "Exact model version": "claude-3-7-sonnet-20250219",
          "Task count": 171,
          "Difference from English (pp)": -12.8,
          "UI and instruction languages": "Both switched to this language",
          "Source": "https://arxiv.org/html/2506.04135v4"
        },
        {
          "Language": "Japanese",
          "Code": "ja",
          "Agent": "Claude CUA",
          "Task success rate": 0.368,
          "Exact model version": "claude-3-7-sonnet-20250219",
          "Task count": 171,
          "Difference from English (pp)": -7.6,
          "UI and instruction languages": "Both switched to this language",
          "Source": "https://arxiv.org/html/2506.04135v4"
        },
        {
          "Language": "Russian",
          "Code": "ru",
          "Agent": "Claude CUA",
          "Task success rate": 0.409,
          "Exact model version": "claude-3-7-sonnet-20250219",
          "Task count": 171,
          "Difference from English (pp)": -3.5,
          "UI and instruction languages": "Both switched to this language",
          "Source": "https://arxiv.org/html/2506.04135v4"
        },
        {
          "Language": "English reference",
          "Code": "en",
          "Agent": "OpenAI CUA",
          "Task success rate": 0.33299999999999996,
          "Exact model version": "computer-use-preview-2025-03-11",
          "Task count": 171,
          "Difference from English (pp)": 0.0,
          "UI and instruction languages": "Both switched to this language",
          "Source": "https://arxiv.org/html/2506.04135v4"
        },
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Agent": "OpenAI CUA",
          "Task success rate": 0.281,
          "Exact model version": "computer-use-preview-2025-03-11",
          "Task count": 171,
          "Difference from English (pp)": -5.2,
          "UI and instruction languages": "Both switched to this language",
          "Source": "https://arxiv.org/html/2506.04135v4"
        },
        {
          "Language": "Japanese",
          "Code": "ja",
          "Agent": "OpenAI CUA",
          "Task success rate": 0.35100000000000003,
          "Exact model version": "computer-use-preview-2025-03-11",
          "Task count": 171,
          "Difference from English (pp)": 1.8,
          "UI and instruction languages": "Both switched to this language",
          "Source": "https://arxiv.org/html/2506.04135v4"
        },
        {
          "Language": "Russian",
          "Code": "ru",
          "Agent": "OpenAI CUA",
          "Task success rate": 0.392,
          "Exact model version": "computer-use-preview-2025-03-11",
          "Task count": 171,
          "Difference from English (pp)": 5.9,
          "UI and instruction languages": "Both switched to this language",
          "Source": "https://arxiv.org/html/2506.04135v4"
        }
      ],
      "xweb_translation": [
        {
          "Language": "French",
          "Code": "fr",
          "Strategy": "Original-language BaseAgent",
          "Task score": 42.7,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "French",
          "Code": "fr",
          "Strategy": "Google Translate to English",
          "Task score": 48.33,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Spanish",
          "Code": "es",
          "Strategy": "Original-language BaseAgent",
          "Task score": 37.31,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Spanish",
          "Code": "es",
          "Strategy": "Google Translate to English",
          "Task score": 25.97,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "German",
          "Code": "de",
          "Strategy": "Original-language BaseAgent",
          "Task score": 34.56,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "German",
          "Code": "de",
          "Strategy": "Google Translate to English",
          "Task score": 37.41,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Russian",
          "Code": "ru",
          "Strategy": "Original-language BaseAgent",
          "Task score": 36.41,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Russian",
          "Code": "ru",
          "Strategy": "Google Translate to English",
          "Task score": 38.0,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Turkish",
          "Code": "tr",
          "Strategy": "Original-language BaseAgent",
          "Task score": 43.18,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Turkish",
          "Code": "tr",
          "Strategy": "Google Translate to English",
          "Task score": 33.99,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Strategy": "Original-language BaseAgent",
          "Task score": 41.34,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Arabic (generic)†",
          "Code": "ar",
          "Strategy": "Google Translate to English",
          "Task score": 2.18,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Vietnamese",
          "Code": "vi",
          "Strategy": "Original-language BaseAgent",
          "Task score": 37.93,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Vietnamese",
          "Code": "vi",
          "Strategy": "Google Translate to English",
          "Task score": 35.19,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Thai",
          "Code": "th",
          "Strategy": "Original-language BaseAgent",
          "Task score": 18.51,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Thai",
          "Code": "th",
          "Strategy": "Google Translate to English",
          "Task score": 12.69,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Strategy": "Original-language BaseAgent",
          "Task score": 36.75,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Hindi",
          "Code": "hi",
          "Strategy": "Google Translate to English",
          "Task score": 33.47,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Swahili",
          "Code": "sw",
          "Strategy": "Original-language BaseAgent",
          "Task score": 42.1,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Swahili",
          "Code": "sw",
          "Strategy": "Google Translate to English",
          "Task score": 36.49,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Urdu",
          "Code": "ur",
          "Strategy": "Original-language BaseAgent",
          "Task score": 34.15,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        },
        {
          "Language": "Urdu",
          "Code": "ur",
          "Strategy": "Google Translate to English",
          "Task score": 2.3,
          "Model": "GPT-4o",
          "Instructions per language": 200,
          "Metric": "WebShop Task Score, not binary success",
          "Source": "https://arxiv.org/html/2505.15372v1"
        }
      ],
      "evidence_inventory": [
        {
          "No.": 1,
          "Project": "RuBench",
          "Language coverage": "Russian",
          "Task layer": "Code execution",
          "Models and evidence layer": "Sol, Opus 5 / current direct",
          "Verified openness": "Tasks and run results public; full grading tests withheld",
          "Interpretation limits": "No matched English; models use different task sets",
          "Task ancestry": "Independent native code tasks",
          "Source": "https://github.com/eugeneshilow/rubench"
        },
        {
          "No.": 2,
          "Project": "MERA GorillaHard",
          "Language coverage": "Russian",
          "Task layer": "Static tool plans",
          "Models and evidence layer": "Sol, Opus 5, Grok 4.6 / current direct",
          "Verified openness": "Framework, metrics, scores public; incomplete test truth/configuration",
          "Interpretation limits": "Tools not executed; submission settings not fully aligned",
          "Task ancestry": "MERA / GorillaHard",
          "Source": "https://mera.a-ai.ru/en/text"
        },
        {
          "No.": 3,
          "Project": "MDPBench / Kimi K3",
          "Language coverage": "13 target observations, including related Arabic label",
          "Task layer": "Document component",
          "Models and evidence layer": "Kimi K3 / current component",
          "Verified openness": "Official per-language scores public; no rerun here",
          "Interpretation limits": "Not agent success; unpaired documents",
          "Task ancestry": "Document parsing component",
          "Source": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e"
        },
        {
          "No.": 4,
          "Project": "GAIA-v2-LILT",
          "Language coverage": "ar de hi ko pt-BR + en",
          "Task layer": "Retrieval / tool execution",
          "Models and evidence layer": "GPT-5.4, Gemini 3.1 Pro, Opus 4.6 / earlier flagships",
          "Verified openness": "Evaluation code, HF tasks, per-language table public",
          "Interpretation limits": "Audit changes function, culture, difficulty together; not pure language causality",
          "Task ancestry": "Audited MAPS-GAIA branch",
          "Source": "https://arxiv.org/html/2604.24929v1"
        },
        {
          "No.": 5,
          "Project": "MAPS",
          "Language coverage": "Final paper: 12 languages; older HF card: 11",
          "Task layer": "GAIA/SWE/safety/math mixture",
          "Models and evidence layer": "Historical models / context",
          "Verified openness": "Paper and HF data public; pin subset and version",
          "Interpretation limits": "MATH is not execution; GAIA shares ancestry with preceding row",
          "Task ancestry": "Parent GAIA, SWE, MATH, ASB sets",
          "Source": "https://aclanthology.org/2026.findings-eacl.42/"
        },
        {
          "No.": 6,
          "Project": "SEATauBench",
          "Language coverage": "vi th id Filipino zh + en",
          "Task layer": "Multi-turn tools and final state",
          "Models and evidence layer": "Kimi K2.5, GPT-5-mini, Qwen3 / historical",
          "Verified openness": "Code, domain data, analysis framework public; MIT repository",
          "Interpretation limits": "Simulator can also fail; no isolated agent-only language effect",
          "Task ancestry": "Derived from tau2; shares framework with LILT tau",
          "Source": "https://github.com/SEACrowd/SEATauBench"
        },
        {
          "No.": 7,
          "Project": "macOSWorld",
          "Language coverage": "ar ja ru zh + en",
          "Task layer": "Interactive desktop",
          "Models and evidence layer": "Claude 3.7 CUA, 2025 OpenAI CUA / earlier flagships",
          "Verified openness": "Tasks, environment, scoring code public; per-language paper results",
          "Interpretation limits": "Instructions and UI switch together; v4 revises some scores",
          "Task ancestry": "Native macOS tasks",
          "Source": "https://arxiv.org/html/2506.04135v4"
        },
        {
          "No.": 8,
          "Project": "X-WebAgentBench",
          "Language coverage": "14 non-English settings; 11 target languages",
          "Task layer": "Interactive web shopping",
          "Models and evidence layer": "GPT-4o and others / historical",
          "Verified openness": "Code and data entry points public; paper results",
          "Interpretation limits": "Task Score is not success rate; derived from WebShop",
          "Task ancestry": "WebShop derivative",
          "Source": "https://github.com/WPENGxs/X-WebAgentBench"
        },
        {
          "No.": 9,
          "Project": "MASSIVE-Agents",
          "Language coverage": "52 languages; 24 targets including related labels",
          "Task layer": "Static function / argument matching",
          "Models and evidence layer": "Nova Premier, Claude 3.5 and others / earlier flagships",
          "Verified openness": "Converted HF data CC BY 4.0; per-language paper tables",
          "Interpretation limits": "Not multi-turn execution; filtering changes samples/function coverage by language",
          "Task ancestry": "MASSIVE derivative; BFCL scoring",
          "Source": "https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id"
        },
        {
          "No.": 10,
          "Project": "Terminal-Bench-LILT",
          "Language coverage": "ar cs de es hi ja ko sr tr zh",
          "Task layer": "Native multilingual code execution",
          "Models and evidence layer": "GPT-5.5, Opus 4.8 and others / earlier flagships",
          "Verified openness": "Paper, aggregate scores, samples public; full tasks require contacting authors",
          "Interpretation limits": "Blog 300 versus leaderboard 324 tasks: denominators not mixed; full set not public",
          "Task ancestry": "Native Terminal-Bench format; separate from community LILTBench",
          "Source": "https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark"
        },
        {
          "No.": 11,
          "Project": "LILT multilingual tau",
          "Language coverage": "de ko + en",
          "Task layer": "Multi-turn tools and final state",
          "Models and evidence layer": "GPT-5.4, Opus 4.8, Gemini 3.1 Pro / earlier flagships",
          "Verified openness": "Public board and run notes; dynamic per-language scores not extracted here",
          "Interpretation limits": "Different English/target-language simulators confound scores",
          "Task ancestry": "tau2 derivative; not independent framework replication",
          "Source": "https://benchmarks.lilt.com/"
        },
        {
          "No.": 12,
          "Project": "LILTBench community",
          "Language coverage": "Authors report 31 languages; not exhaustively checked task by task",
          "Task layer": "Native code / English pairs",
          "Models and evidence layer": "Opus 4.6 / challenging test resources",
          "Verified openness": "Tasks, verifiers, leaderboard public; Apache-2.0 repository",
          "Interpretation limits": "Adversarial hard-task selection cannot estimate population failure rates",
          "Task ancestry": "Community native tasks; distinct from commercial full set",
          "Source": "https://github.com/lilt/liltbench-tasks-public"
        },
        {
          "No.": 13,
          "Project": "VoiceAgentBench",
          "Language coverage": "en hi bn mr ta te ml",
          "Task layer": "Speech-to-tool-call scoring",
          "Models and evidence layer": "SpeechLM / ASR+LLM / non-frontier resources",
          "Verified openness": "Evaluation code and HF data entry public; custom community license",
          "Interpretation limits": "Multi-turn subset is English only; predicted calls are not final environment success",
          "Task ancestry": "Synthetic speech-tool tasks",
          "Source": "https://github.com/ola-krutrim/VoiceAgentBench"
        },
        {
          "No.": 14,
          "Project": "TelcoAgent-Bench",
          "Language coverage": "ar + en",
          "Task layer": "Diagnostic tool sequences / summaries",
          "Models and evidence layer": "3B–8B models / non-frontier resources",
          "Verified openness": "Blueprints, tasks, prediction/score JSON public; license unverified",
          "Interpretation limits": "Intent/resolution similarity is not strict success; cannot extrapolate to flagships",
          "Task ancestry": "Telecom blueprint tasks",
          "Source": "https://github.com/BrahiM-Mefgouda/TelcoAgent"
        },
        {
          "No.": 15,
          "Project": "MultiAgent-X",
          "Language coverage": "12 languages; target hi sw ha am",
          "Task layer": "Function-call data",
          "Models and evidence layer": "No verified current-frontier baseline / resources",
          "Verified openness": "Samples, scripts, structure public; full license unverified",
          "Interpretation limits": "Synthetic; structural validity is not native quality; pa is not pnb",
          "Task ancestry": "New synthetic tasks; quoted MASSIVE scores are not replication",
          "Source": "https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark"
        },
        {
          "No.": 16,
          "Project": "MAST FIRE 2026",
          "Language coverage": "21-language track union; 14 targets",
          "Task layer": "Multi-turn retrieval / answers",
          "Models and evidence layer": "Public Tongyi-30B baseline / non-frontier",
          "Verified openness": "Task/corpus entry points and per-language baseline public",
          "Interpretation limits": "September 30 final results are future at cutoff; baseline only; English documents and answers",
          "Task ancestry": "BrowseComp-Plus derivative; three Indic languages overlap tracks",
          "Source": "https://mast-benchmark.github.io/"
        }
      ],
      "broad_atlas_1": [
        {
          "Language": "Hindi",
          "MASSIVE calls": 2,
          "GAIA audited": 2,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 1,
          "Voice tools": 1,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 1,
          "Code": "hi",
          "Included projects": "MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Spanish",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 1,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "es",
          "Included projects": "MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Modern Standard Arabic†",
          "MASSIVE calls": 2,
          "GAIA audited": 2,
          "X-Web shopping": 2,
          "macOS use": 2,
          "SEA multi-turn": 0,
          "LILT native code": 1,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "ar",
          "Included projects": "MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval",
          "Mapping limitation": "Generic Arabic/Filipino are related labels, not exact MSA/Tagalog matches",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "French",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "fr",
          "Included projects": "MASSIVE calls; X-Web shopping; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Bengali",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 1,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "bn",
          "Included projects": "MASSIVE calls; Voice tools; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Portuguese",
          "MASSIVE calls": 2,
          "GAIA audited": 2,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "pt",
          "Included projects": "MASSIVE calls; GAIA audited",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Indonesian",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 2,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "id",
          "Included projects": "MASSIVE calls; SEA multi-turn",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Urdu",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "ur",
          "Included projects": "MASSIVE calls; X-Web shopping; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Russian",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 2,
          "macOS use": 2,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "ru",
          "Included projects": "MASSIVE calls; X-Web shopping; macOS use; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "German",
          "MASSIVE calls": 2,
          "GAIA audited": 2,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 1,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "de",
          "Included projects": "MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Japanese",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 2,
          "SEA multi-turn": 0,
          "LILT native code": 1,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "ja",
          "Included projects": "MASSIVE calls; macOS use; LILT native code",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Nigerian Pidgin",
          "MASSIVE calls": 0,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "pcm",
          "Included projects": "Not included in this map",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Egyptian Arabic",
          "MASSIVE calls": 0,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "arz",
          "Included projects": "Not included in this map",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Marathi",
          "MASSIVE calls": 0,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 1,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "mr",
          "Included projects": "Voice tools",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Vietnamese",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 2,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "vi",
          "Included projects": "MASSIVE calls; X-Web shopping; SEA multi-turn",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        }
      ],
      "broad_atlas_2": [
        {
          "Language": "Telugu",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 1,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "te",
          "Included projects": "MASSIVE calls; Voice tools; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Swahili",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 1,
          "Code": "sw",
          "Included projects": "MASSIVE calls; X-Web shopping; MAST retrieval; MultiAgent-X tasks",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Hausa",
          "MASSIVE calls": 0,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 1,
          "Code": "ha",
          "Included projects": "MultiAgent-X tasks",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Turkish",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 1,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "tr",
          "Included projects": "MASSIVE calls; X-Web shopping; LILT native code",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Western Punjabi",
          "MASSIVE calls": 0,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "pnb",
          "Included projects": "Not included in this map",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Tagalog†",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 2,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "tl",
          "Included projects": "MASSIVE calls; SEA multi-turn",
          "Mapping limitation": "Generic Arabic/Filipino are related labels, not exact MSA/Tagalog matches",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Tamil",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 1,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "ta",
          "Included projects": "MASSIVE calls; Voice tools; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Iranian Persian",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "fa",
          "Included projects": "MASSIVE calls",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Korean",
          "MASSIVE calls": 2,
          "GAIA audited": 2,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 1,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "ko",
          "Included projects": "MASSIVE calls; GAIA audited; LILT native code",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Amharic",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 1,
          "Code": "am",
          "Included projects": "MASSIVE calls; MultiAgent-X tasks",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Thai",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 2,
          "macOS use": 0,
          "SEA multi-turn": 2,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "th",
          "Included projects": "MASSIVE calls; X-Web shopping; SEA multi-turn; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Javanese",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "jv",
          "Included projects": "MASSIVE calls",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Italian",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 0,
          "MultiAgent-X tasks": 0,
          "Code": "it",
          "Included projects": "MASSIVE calls",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Gujarati",
          "MASSIVE calls": 0,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "gu",
          "Included projects": "MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        },
        {
          "Language": "Kannada",
          "MASSIVE calls": 2,
          "GAIA audited": 0,
          "X-Web shopping": 0,
          "macOS use": 0,
          "SEA multi-turn": 0,
          "LILT native code": 0,
          "Voice tools": 0,
          "MAST retrieval": 2,
          "MultiAgent-X tasks": 0,
          "Code": "kn",
          "Included projects": "MASSIVE calls; MAST retrieval",
          "Mapping limitation": "Language-code mapping; no substitution by neighboring languages",
          "Meaning": "2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability"
        }
      ],
      "triangulation": [
        {
          "Language group": "01 Thai",
          "Supporting evidence": "SEA retail and other domains; X-Web shopping; K3 documents",
          "Supported interpretation": "Weakness across tasks; current-frontier evidence is component-only",
          "Counterevidence or limit": "Thai is the highest one-shot language for Llama 3.1 405B in MASSIVE; models/prompts change order",
          "Action": "Retest fully localized business and shopping; avoid all-model claims"
        },
        {
          "Language group": "02 Arabic (generic)†",
          "Supporting evidence": "Audited GAIA: three flagships below English; macOS: two CUAs lower; K3 component",
          "Supported interpretation": "Aligned historical signals across two execution families",
          "Counterevidence or limit": "GAIA improves after audit; GUI also changes RTL layout; dialects untested",
          "Action": "Retest search, RTL UI, and numeric arguments; do not extend to Egyptian Arabic"
        },
        {
          "Language group": "03 Japanese",
          "Supporting evidence": "macOS CUA; native code tasks; K3 component",
          "Supported interpretation": "Component weakness; execution depends on model",
          "Counterevidence or limit": "OpenAI CUA exceeds English; Claude CUA falls below English",
          "Action": "Separate document reading, click targeting, and display-width tests"
        },
        {
          "Language group": "04 Hindi",
          "Supporting evidence": "Audited GAIA; native code; MASSIVE; K3 component",
          "Supported interpretation": "Several task families; low scores sensitive to task quality",
          "Counterevidence or limit": "GAIA audit gains 25.4–32.7 pp; English instructions do not remove all native-task difficulty in Hindi/German",
          "Action": "Audit task alignment, entities, and regional rules; distinguish reading from execution"
        },
        {
          "Language group": "05 Russian",
          "Supporting evidence": "Current RuBench execution and GorillaHard plans; historical macOS",
          "Supported interpretation": "Concrete current failures, more direct than a component",
          "Counterevidence or limit": "RuBench/MERA lack English pairs; historical OpenAI CUA Russian exceeds English",
          "Action": "Regress specific failures; do not rank Russian worst overall"
        },
        {
          "Language group": "06 Vietnamese / Indonesian / Filipino†",
          "Supporting evidence": "SEA full business; MASSIVE; Vietnamese X-Web",
          "Supported interpretation": "Several tasks; language ordering varies by domain",
          "Counterevidence or limit": "Kimi Vietnamese retail slightly exceeds English; Filipino simulator also errs",
          "Action": "Accept by business domain and verify Filipino–Tagalog variety alignment"
        },
        {
          "Language group": "07 German / Korean / Portuguese",
          "Supporting evidence": "Audited GAIA; MASSIVE; K3; German/Korean LILT",
          "Supported interpretation": "More evidence does not establish deployment readiness",
          "Counterevidence or limit": "Audited German gap is 3.1 pp for GPT-5.4, 12.7 pp for Opus 4.6",
          "Action": "Test the target model and actual workflow; no universal safe-language list"
        },
        {
          "Language group": "08 Amharic / Kannada",
          "Supporting evidence": "Historical MASSIVE zero-shot AST; Amharic synthetic tasks",
          "Supported interpretation": "Risk signal mainly from one historical task family",
          "Counterevidence or limit": "Lowest language differs by model; no current-frontier execution replication",
          "Action": "Prioritize function arguments; evidence cannot rank current models"
        },
        {
          "Language group": "09 Swahili / Urdu",
          "Supporting evidence": "X-Web; MASSIVE; non-frontier MAST retrieval",
          "Supported interpretation": "Task- and model-dependent signals; no uniform direction",
          "Counterevidence or limit": "Swahili original-language GPT-4o shopping is not low; speech and retrieval scores are not interchangeable",
          "Action": "Measure retrieval, calls, and final state separately with English controls"
        },
        {
          "Language group": "10 Other coverage / gaps",
          "Supporting evidence": "Bengali/Tamil/Telugu tool or retrieval resources; Pidgin speech component",
          "Supported interpretation": "Consult all 30 rows; missing evidence is not low ability",
          "Counterevidence or limit": "Egyptian Arabic/Western Punjabi exact matches missing; Hausa lacks current-frontier baseline",
          "Action": "Use native tasks: pcm≠en, arz≠ar, pnb≠pa, jv≠id"
        }
      ]
    }
  },
  "sources": [
    {
      "id": "format_partition",
      "label": "All samples: format and content outcomes",
      "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
      "notes": "Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed.",
      "query": {
        "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
        "description": "Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Model\", \"Full sample passed\", \"Valid format, sample failed\", \"Invalid format\", \"Submission date\", \"Evaluated N\", \"Configuration limits\", \"Source\", \"Documented dataset N\", \"Scope\") AS (\n  VALUES\n    ('GPT-5.6 Sol', 0.691, 0.308, 0.001, '2026-08-18', 'Not separately disclosed', 'openrouter / 9ebf388', 'https://mera.a-ai.ru/en/text/submits/2.0/8', 1169, 'All call, abstention, and clarification items'),\n    ('Claude Opus 5', 0.677, 0.272, 0.051, '2026-08-18', 'Not separately disclosed', 'openrouter / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/1', 1169, 'All call, abstention, and clarification items'),\n    ('Grok 4.6', 0.737, 0.263, 0, '2026-08-18', 'Not separately disclosed', 'local-chat-completions / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/11', 1169, 'All call, abstention, and clarification items')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "All samples: format and content outcomes"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Partitions: sample_pass_rate; format_pass_rate minus sample_pass_rate; 1 minus format_pass_rate. Derived from rounded three-decimal public metrics. Run-specific N/repetitions not separately disclosed; settings differ; no tools executed."
        ]
      }
    },
    {
      "id": "argument_partition",
      "label": "Call-required items: tools and arguments",
      "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/utils.py",
      "notes": "Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference.",
      "query": {
        "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/utils.py",
        "description": "Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Model\", \"Tools and arguments matched\", \"Tools matched, arguments incomplete\", \"Tools not fully matched\", \"Submission date\", \"Evaluated N\", \"Configuration limits\", \"Source\", \"Documented dataset N\", \"Scope\") AS (\n  VALUES\n    ('GPT-5.6 Sol', 0.702, 0.163, 0.135, '2026-08-18', 'Not separately disclosed', 'openrouter / 9ebf388', 'https://mera.a-ai.ru/en/text/submits/2.0/8', 1032, 'Call-required items only'),\n    ('Claude Opus 5', 0.677, 0.113, 0.21, '2026-08-18', 'Not separately disclosed', 'openrouter / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/1', 1032, 'Call-required items only'),\n    ('Grok 4.6', 0.743, 0.081, 0.176, '2026-08-18', 'Not separately disclosed', 'local-chat-completions / 0e1f484', 'https://mera.a-ai.ru/en/text/submits/2.0/11', 1032, 'Call-required items only')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Call-required items: tools and arguments"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Partitions: args_match_rate; tool_match_rate minus args_match_rate; 1 minus tool_match_rate. Matching arguments requires matching tools. Rounded source metrics; actual run N not separately disclosed. Different denominator from all-sample chart; no language-causal inference."
        ]
      }
    },
    {
      "id": "decision_errors",
      "label": "Abstention errors under two conditions",
      "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
      "notes": "1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions.",
      "query": {
        "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
        "description": "1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Required action\", \"Model\", \"Conditional error rate\", \"Documented denominator\", \"Scoring interpretation\", \"Evaluated N\", \"Source\") AS (\n  VALUES\n    ('Abstention required (127)', 'GPT-5.6 Sol', 0.417, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/8'),\n    ('Call required (1,032)', 'GPT-5.6 Sol', 0.009, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/8'),\n    ('Abstention required (127)', 'Claude Opus 5', 0.354, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/1'),\n    ('Call required (1,032)', 'Claude Opus 5', 0.001, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/1'),\n    ('Abstention required (127)', 'Grok 4.6', 0.339, 127, 'Correct abstention missing', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/11'),\n    ('Call required (1,032)', 'Grok 4.6', 0.004, 1032, 'False abstention', 'Not separately disclosed', 'https://mera.a-ai.ru/en/text/submits/2.0/11')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Abstention errors under two conditions"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "1 minus abstention_recall on documented 127 abstention items; false_abstention_rate on 1,032 call items. Distinct conditional denominators cannot be added. Actual evaluated N not separately disclosed. Task rules are not all safety policies; outputs are not executed actions."
        ]
      }
    },
    {
      "id": "execution_partition",
      "label": "Russian code tasks: audited run outcomes",
      "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
      "notes": "Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control.",
      "query": {
        "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "description": "Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Configuration\", \"Audit-retained passes\", \"Raw passes removed by audit\", \"Original failures\", \"Retained pass count\", \"Removed pass count\", \"Original failure count\", \"Total runs\", \"Distinct tasks\", \"Harness\", \"Setting\", \"Source\") AS (\n  VALUES\n    ('GPT-5.6 Sol (23 tasks × 3)', 0.7101449275362319, 0.11594202898550725, 0.17391304347826086, 49, 8, 12, 69, 23, 'Codex CLI', 'xhigh', 'https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md'),\n    ('Claude Opus 5 (20 tasks × 3)', 0.7666666666666667, 0.016666666666666666, 0.21666666666666667, 46, 1, 13, 60, 20, 'Claude Code', 'xhigh', 'https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Russian code tasks: audited run outcomes"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Retained passes + raw passes removed for answer exposure + original failures = total runs. Sol 23 and Opus 20 tasks each repeated three times; task cohorts and harnesses differ. Audit re-scores existing network-enabled runs, not offline reruns. No English control."
        ]
      }
    },
    {
      "id": "document_profile",
      "label": "Document parsing across 13 language observations",
      "href": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e",
      "notes": "Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold.",
      "query": {
        "url": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e",
        "description": "Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Parsing quality score\", \"Observed position\", \"Script group\", \"Order\", \"Median across 13 observations\", \"Source\") AS (\n  VALUES\n    ('Thai', 'th', 72.1, 'Lower four observations', 'Other scripts', 1, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Japanese', 'ja', 74.9, 'Lower four observations', 'Other scripts', 2, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Arabic (unspecified variety)', 'ar', 77.4, 'Lower four observations', 'Other scripts', 3, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Hindi', 'hi', 77.5, 'Lower four observations', 'Other scripts', 4, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('French', 'fr', 80.0, 'Other nine observations', 'Latin script', 5, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Spanish', 'es', 80.2, 'Other nine observations', 'Latin script', 6, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Russian', 'ru', 82.4, 'Other nine observations', 'Other scripts', 7, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Vietnamese', 'vi', 84.8, 'Other nine observations', 'Latin script', 8, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Indonesian', 'id', 86.9, 'Other nine observations', 'Latin script', 9, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Portuguese', 'pt', 88.9, 'Other nine observations', 'Latin script', 10, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('German', 'de', 89.1, 'Other nine observations', 'Latin script', 11, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Korean', 'ko', 89.9, 'Other nine observations', 'Other scripts', 12, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Italian', 'it', 92.7, 'Other nine observations', 'Latin script', 13, 82.4, 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Document parsing across 13 language observations"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Author parsing composite, not agent success or failure probability. Different language documents are not paired. Generic Arabic is a related label. Lower four: Thai, Japanese, Arabic, Hindi. Median is descriptive, not a deployment threshold."
        ]
      }
    },
    {
      "id": "script_distribution",
      "label": "Parsing score distributions by script group",
      "href": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e",
      "notes": "Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges.",
      "query": {
        "url": "https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e",
        "description": "Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Script group\", \"Minimum\", \"First quartile\", \"Median\", \"Third quartile\", \"Maximum\", \"Observation count\", \"Mean\", \"Language\", \"Source\") AS (\n  VALUES\n    ('Latin script (7)', 80.0, 82.5, 86.9, 89.0, 92.7, 7, 86.08571428571429, 'French; Spanish; Vietnamese; Indonesian; Portuguese; German; Italian', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    ('Other scripts (6)', 72.1, 75.525, 77.45, 81.17500000000001, 89.9, 6, 79.03333333333333, 'Thai; Japanese; Arabic (unspecified variety); Hindi; Russian; Korean', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Parsing score distributions by script group"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Minimum, inclusive Type-7 Q1, median, Q3, maximum across the same 13 observations. Descriptive distribution, not confidence intervals or a controlled script/language effect. Unpaired documents; overlapping ranges."
        ]
      }
    },
    {
      "id": "atlas_1",
      "label": "Current-model evidence map · 1–15",
      "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
      "notes": "Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related.",
      "query": {
        "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "description": "Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"Document component\", \"Static tool plan\", \"End-to-end execution\", \"Matched language control\", \"Marker definition\") AS (\n  VALUES\n    ('01 Hindi', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('02 Spanish', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('03 Modern Standard Arabic', 0.5, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('04 French', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('05 Bengali', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('06 Portuguese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('07 Indonesian', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('08 Urdu', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('09 Russian', 1, 1, 1, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('10 German', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('11 Japanese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('12 Nigerian Pidgin', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('13 Egyptian Arabic', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('14 Marathi', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('15 Vietnamese', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Current-model evidence map · 1–15"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Fixed first 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Historical and resource evidence belongs in the broader map. Generic Arabic only related."
        ]
      }
    },
    {
      "id": "atlas_2",
      "label": "Current-model evidence map · 16–30",
      "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
      "notes": "Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages.",
      "query": {
        "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "description": "Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"Document component\", \"Static tool plan\", \"End-to-end execution\", \"Matched language control\", \"Marker definition\") AS (\n  VALUES\n    ('16 Telugu', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('17 Swahili', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('18 Hausa', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('19 Turkish', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('20 Western Punjabi', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('21 Tagalog', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('22 Tamil', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('23 Iranian Persian', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('24 Korean', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('25 Amharic', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('26 Thai', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('27 Javanese', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('28 Italian', 1, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('29 Gujarati', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability'),\n    ('30 Kannada', 0, 0, 0, 0, '1=matched evidence; 0.5=related label; 0=no included result; not ability')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Current-model evidence map · 16–30"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Fixed last 15 of 30 languages. Only current-model RuBench, GorillaHard and K3 evidence. Markers are evidence presence, not capability. Missing evidence is not low performance; no substitution by adjacent languages."
        ]
      }
    },
    {
      "id": "action_map",
      "label": "Proposed acceptance tests",
      "href": "https://www.unicode.org/reports/tr35/tr35-numbers.html",
      "notes": "Researcher-proposed fixtures based on Unicode conventions and interface contracts. No measured model outcomes or failure incidence. Preserve distinctions pcm≠en, arz≠ar, pnb≠pa, jv≠id.",
      "query": {
        "url": "https://www.unicode.org/reports/tr35/tr35-numbers.html",
        "description": "Researcher-proposed fixtures based on Unicode conventions and interface contracts. No measured model outcomes or failure incidence. Preserve distinctions pcm≠en, arz≠ar, pnb≠pa, jv≠id. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Test target\", \"Input condition\", \"Expected final result\") AS (\n  VALUES\n    ('Amounts | Urdu / Persian', '۱٬۲۵۰٫۵۰; explicit currency', 'Pass 1250.5 according to the API contract; preserve currency'),\n    ('Entity IDs | all languages', 'INV-01250/AB; mixed scripts', 'Preserve string, leading zeros, separators; act on the correct entity'),\n    ('Dates | explicit locale / calendar', '05/09/2026; with/without locale control', 'Convert correctly when defined; clarify per rules when ambiguous'),\n    ('Dialect / neighboring language | four gaps', 'Native pcm, arz, pnb, jv requests', 'Preserve native constraints; no substitution by neighboring-language scores')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Proposed acceptance tests"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Researcher-proposed fixtures based on Unicode conventions and interface contracts. No measured model outcomes or failure incidence. Preserve distinctions pcm≠en, arz≠ar, pnb≠pa, jv≠id."
        ]
      }
    },
    {
      "id": "methods",
      "label": "Three current-model evidence layers",
      "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
      "notes": "Review synthesis of the audited RuBench, GorillaHard, and MDP configurations. Do not equate static plans, document components, executed tasks, or language-causal comparisons.",
      "query": {
        "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md",
        "description": "Review synthesis of the audited RuBench, GorillaHard, and MDP configurations. Do not equate static plans, document components, executed tasks, or language-causal comparisons. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Evidence\", \"Measurement\", \"Key limitation\") AS (\n  VALUES\n    ('GorillaHard', 'Static output; documented N=1,169 total, 1,032 calls, 127 abstentions', 'Documented counts, not separately disclosed run N. Incomplete settings; Sol source commit unresolved; no English pairs.'),\n    ('MDPBench', 'Kimi K3 document parsing quality; 13 target observations', 'Unpaired content/layout; unspecified Arabic variety; composite score is not failure probability.'),\n    ('RuBench', 'Russian instructions → repository edits → regression tests', 'Sol 23×3, Opus 20×3; different sets/harnesses. Earlier Fable fallback omitted.')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Three current-model evidence layers"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Review synthesis of the audited RuBench, GorillaHard, and MDP configurations. Do not equate static plans, document components, executed tasks, or language-causal comparisons."
        ]
      }
    },
    {
      "id": "evidence-rubench_paper",
      "label": "RuBench v1: construction, protocol and Fable fallback",
      "href": "https://arxiv.org/html/2607.06411v1",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://arxiv.org/html/2607.06411v1",
        "tables_used": [
          "RuBench v1: construction, protocol and Fable fallback"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "evidence-rubench_tasks",
      "label": "RuBench task metadata",
      "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/tasks.json",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/tasks.json",
        "tables_used": [
          "RuBench task metadata"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "evidence-rubench_aioh4",
      "label": "AIOH4 original Russian task statement",
      "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/tasks/AIOH4-case-sensitivity/task.md",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/tasks/AIOH4-case-sensitivity/task.md",
        "tables_used": [
          "AIOH4 original Russian task statement"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "evidence-mera_dataset",
      "label": "GorillaHard dataset description and Russian sample",
      "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/datasets/GorillaHard/README.md",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/datasets/GorillaHard/README.md",
        "tables_used": [
          "GorillaHard dataset description and Russian sample"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "evidence-mera_config",
      "label": "GorillaHard task configuration",
      "href": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/gorillahard.yaml",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/gorillahard.yaml",
        "tables_used": [
          "GorillaHard task configuration"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "evidence-mera_author",
      "label": "MERA team's Text 2.0 launch explanation",
      "href": "https://habr.com/ru/companies/sberbank/articles/1075592/",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://habr.com/ru/companies/sberbank/articles/1075592/",
        "tables_used": [
          "MERA team's Text 2.0 launch explanation"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "evidence-mera_sol",
      "label": "MERA GPT 5.6 Sol submission",
      "href": "https://mera.a-ai.ru/en/text/submits/2.0/8",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://mera.a-ai.ru/en/text/submits/2.0/8",
        "tables_used": [
          "MERA GPT 5.6 Sol submission"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "evidence-mera_opus",
      "label": "MERA Claude Opus 5 submission",
      "href": "https://mera.a-ai.ru/en/text/submits/2.0/1",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://mera.a-ai.ru/en/text/submits/2.0/1",
        "tables_used": [
          "MERA Claude Opus 5 submission"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "evidence-mera_grok",
      "label": "MERA Grok 4.6 submission",
      "href": "https://mera.a-ai.ru/en/text/submits/2.0/11",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://mera.a-ai.ru/en/text/submits/2.0/11",
        "tables_used": [
          "MERA Grok 4.6 submission"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "linked-20",
      "label": "MDPBench parsing metrics",
      "href": "https://arxiv.org/html/2603.28130v1#S3.SS4",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://arxiv.org/html/2603.28130v1#S3.SS4",
        "tables_used": [
          "MDPBench parsing metrics"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "linked-21",
      "label": "Pinned public language population transcription",
      "href": "https://en.wikipedia.org/w/index.php?title=List_of_languages_by_total_number_of_speakers&oldid=1370552712",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://en.wikipedia.org/w/index.php?title=List_of_languages_by_total_number_of_speakers&oldid=1370552712",
        "tables_used": [
          "Pinned public language population transcription"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "linked-22",
      "label": "Frontier cohort reference",
      "href": "https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2",
      "query": {
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "url": "https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-2",
        "tables_used": [
          "Frontier cohort reference"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "gaia_lift",
      "label": "Score improvement after task auditing",
      "href": "https://arxiv.org/html/2604.24929v1",
      "notes": "GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim. Material inputs: https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS",
      "query": {
        "url": "https://arxiv.org/html/2604.24929v1",
        "description": "GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Model\", \"Before audit\", \"After audit\", \"English reference\", \"Audit improvement (pp)\", \"English minus audited (pp)\", \"Tasks per language\", \"Model cohort\", \"Execution harness\", \"Source\") AS (\n  VALUES\n    ('Arabic (generic)†', 'ar', 'GPT-5.4', 32.1, 47.3, 66.7, 15.2, 19.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'GPT-5.4', 47.3, 63.6, 66.7, 16.3, 3.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'GPT-5.4', 34.6, 60.0, 66.7, 25.4, 6.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'GPT-5.4', 33.3, 62.4, 66.7, 29.1, 4.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'GPT-5.4', 47.3, 58.2, 66.7, 10.9, 8.5, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Arabic (generic)†', 'ar', 'Gemini 3.1 Pro', 34.6, 52.1, 73.9, 17.5, 21.8, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'Gemini 3.1 Pro', 49.7, 66.7, 73.9, 17.0, 7.2, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'Gemini 3.1 Pro', 38.2, 63.6, 73.9, 25.4, 10.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'Gemini 3.1 Pro', 34.6, 64.8, 73.9, 30.2, 9.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'Gemini 3.1 Pro', 48.5, 65.5, 73.9, 17.0, 8.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Arabic (generic)†', 'ar', 'Claude Opus 4.6', 32.1, 49.1, 79.4, 17.0, 30.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'Claude Opus 4.6', 49.7, 66.7, 79.4, 17.0, 12.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'Claude Opus 4.6', 29.7, 62.4, 79.4, 32.7, 17.0, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'Claude Opus 4.6', 33.3, 58.8, 79.4, 25.5, 20.6, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'Claude Opus 4.6', 49.1, 63.0, 79.4, 13.9, 16.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Score improvement after task auditing"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "GAIA-v2-LILT v1 Table 2. Audit minus MT pass@1, percentage points; source pass@1 percentages retained. Audit changes functionality, culture, difficulty; improvement is not purely translation error. MAPS ancestry shared, not independent replication. No confidence intervals; no small-gap significance claim."
        ]
      }
    },
    {
      "id": "gaia_residual",
      "label": "Remaining gap from each model’s English reference",
      "href": "https://arxiv.org/html/2604.24929v1",
      "notes": "GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals. Material inputs: https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS",
      "query": {
        "url": "https://arxiv.org/html/2604.24929v1",
        "description": "GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Model\", \"Before audit\", \"After audit\", \"English reference\", \"Audit improvement (pp)\", \"English minus audited (pp)\", \"Tasks per language\", \"Model cohort\", \"Execution harness\", \"Source\") AS (\n  VALUES\n    ('Arabic (generic)†', 'ar', 'GPT-5.4', 32.1, 47.3, 66.7, 15.2, 19.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'GPT-5.4', 47.3, 63.6, 66.7, 16.3, 3.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'GPT-5.4', 34.6, 60.0, 66.7, 25.4, 6.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'GPT-5.4', 33.3, 62.4, 66.7, 29.1, 4.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'GPT-5.4', 47.3, 58.2, 66.7, 10.9, 8.5, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Arabic (generic)†', 'ar', 'Gemini 3.1 Pro', 34.6, 52.1, 73.9, 17.5, 21.8, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'Gemini 3.1 Pro', 49.7, 66.7, 73.9, 17.0, 7.2, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'Gemini 3.1 Pro', 38.2, 63.6, 73.9, 25.4, 10.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'Gemini 3.1 Pro', 34.6, 64.8, 73.9, 30.2, 9.1, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'Gemini 3.1 Pro', 48.5, 65.5, 73.9, 17.0, 8.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Arabic (generic)†', 'ar', 'Claude Opus 4.6', 32.1, 49.1, 79.4, 17.0, 30.3, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('German', 'de', 'Claude Opus 4.6', 49.7, 66.7, 79.4, 17.0, 12.7, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Hindi', 'hi', 'Claude Opus 4.6', 29.7, 62.4, 79.4, 32.7, 17.0, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Korean', 'ko', 'Claude Opus 4.6', 33.3, 58.8, 79.4, 25.5, 20.6, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1'),\n    ('Portuguese (Brazil)', 'pt', 'Claude Opus 4.6', 49.1, 63.0, 79.4, 13.9, 16.4, 165, 'Earlier flagship; outside September current cohort', 'Open Deep Research; 12 manager / 20 search steps', 'https://arxiv.org/html/2604.24929v1')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Remaining gap from each model’s English reference"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "GAIA-v2-LILT v1 Table 2. English minus audited pass@1, percentage points. Same experiment as audit-improvement chart. Audits change task content; residual is not a pure language penalty and not a current-model estimate. No confidence intervals."
        ]
      }
    },
    {
      "id": "sea_localization",
      "label": "Retail success by extent of localization",
      "href": "https://arxiv.org/html/2606.28715v1",
      "notes": "SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English. Material inputs: https://arxiv.org/html/2606.28715v1 ; https://github.com/SEACrowd/SEATauBench",
      "query": {
        "url": "https://arxiv.org/html/2606.28715v1",
        "description": "SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Setting\", \"Task success rate\", \"Model\", \"Displayed domain\", \"Trials per task\", \"Airline English\", \"Airline fully localized\", \"Telecom English\", \"Telecom fully localized\", \"Sample definition\", \"Source\") AS (\n  VALUES\n    ('Vietnamese', 'vi', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Vietnamese', 'vi', 'Dialogue localized', 0.687, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Vietnamese', 'vi', 'Fully localized business', 0.567, 'Kimi K2.5', 'Retail', 3, 0.707, 0.56, 0.997, 0.699, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Thai', 'th', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Thai', 'th', 'Dialogue localized', 0.573, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Thai', 'th', 'Fully localized business', 0.327, 'Kimi K2.5', 'Retail', 3, 0.707, 0.547, 0.997, 0.693, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Indonesian', 'id', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Indonesian', 'id', 'Dialogue localized', 0.64, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Indonesian', 'id', 'Fully localized business', 0.433, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.798, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Filipino†', 'tl', 'All English', 0.561, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Filipino†', 'tl', 'Dialogue localized', 0.675, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1'),\n    ('Filipino†', 'tl', 'Fully localized business', 0.444, 'Kimi K2.5', 'Retail', 3, 0.707, 0.6, 0.997, 0.743, 'Separate domains; no pooled mean; three trials per task', 'https://arxiv.org/html/2606.28715v1')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Retail success by extent of localization"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "SEATauBench v1 Tables 9/10/13. Retail pass@1 shown; airline/telecom references retained in data, not pooled. Qwen3-235B user simulator and GPT-4.1 judge; simulator also changes language. Filipino is related to target Tagalog. Vietnamese fully localized retail slightly exceeds English."
        ]
      }
    },
    {
      "id": "macos_language",
      "label": "Desktop task success by agent and language",
      "href": "https://arxiv.org/html/2506.04135v4",
      "notes": "macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs. Material inputs: https://arxiv.org/html/2506.04135v4 ; https://github.com/showlab/macosworld",
      "query": {
        "url": "https://arxiv.org/html/2506.04135v4",
        "description": "macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Agent\", \"Task success rate\", \"Exact model version\", \"Task count\", \"Difference from English (pp)\", \"UI and instruction languages\", \"Source\") AS (\n  VALUES\n    ('English reference', 'en', 'Claude CUA', 0.444, 'claude-3-7-sonnet-20250219', 171, 0.0, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Arabic (generic)†', 'ar', 'Claude CUA', 0.316, 'claude-3-7-sonnet-20250219', 171, -12.8, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Japanese', 'ja', 'Claude CUA', 0.368, 'claude-3-7-sonnet-20250219', 171, -7.6, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Russian', 'ru', 'Claude CUA', 0.409, 'claude-3-7-sonnet-20250219', 171, -3.5, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('English reference', 'en', 'OpenAI CUA', 0.33299999999999996, 'computer-use-preview-2025-03-11', 171, 0.0, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Arabic (generic)†', 'ar', 'OpenAI CUA', 0.281, 'computer-use-preview-2025-03-11', 171, -5.2, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Japanese', 'ja', 'OpenAI CUA', 0.35100000000000003, 'computer-use-preview-2025-03-11', 171, 1.8, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4'),\n    ('Russian', 'ru', 'OpenAI CUA', 0.392, 'computer-use-preview-2025-03-11', 171, 5.9, 'Both switched to this language', 'https://arxiv.org/html/2506.04135v4')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Desktop task success by agent and language"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "macOSWorld v4 Table 3, excluding Advanced Apps. claude-3-7-sonnet-20250219 and computer-use-preview-2025-03-11. Published percentages converted to fractions. UI layout, reading, and planning are not isolated. v4 only; no v1 aggregate mixed in. Arabic lower for both; Japanese/Russian direction differs."
        ]
      }
    },
    {
      "id": "xweb_translation",
      "label": "Web shopping scores: original language versus English translation",
      "href": "https://arxiv.org/html/2505.15372v1",
      "notes": "X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics. Material inputs: https://arxiv.org/html/2505.15372v1 ; https://github.com/WPENGxs/X-WebAgentBench",
      "query": {
        "url": "https://arxiv.org/html/2505.15372v1",
        "description": "X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"Code\", \"Strategy\", \"Task score\", \"Model\", \"Instructions per language\", \"Metric\", \"Source\") AS (\n  VALUES\n    ('French', 'fr', 'Original-language BaseAgent', 42.7, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('French', 'fr', 'Google Translate to English', 48.33, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Spanish', 'es', 'Original-language BaseAgent', 37.31, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Spanish', 'es', 'Google Translate to English', 25.97, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('German', 'de', 'Original-language BaseAgent', 34.56, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('German', 'de', 'Google Translate to English', 37.41, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Russian', 'ru', 'Original-language BaseAgent', 36.41, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Russian', 'ru', 'Google Translate to English', 38.0, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Turkish', 'tr', 'Original-language BaseAgent', 43.18, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Turkish', 'tr', 'Google Translate to English', 33.99, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Arabic (generic)†', 'ar', 'Original-language BaseAgent', 41.34, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Arabic (generic)†', 'ar', 'Google Translate to English', 2.18, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Vietnamese', 'vi', 'Original-language BaseAgent', 37.93, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Vietnamese', 'vi', 'Google Translate to English', 35.19, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Thai', 'th', 'Original-language BaseAgent', 18.51, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Thai', 'th', 'Google Translate to English', 12.69, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Hindi', 'hi', 'Original-language BaseAgent', 36.75, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Hindi', 'hi', 'Google Translate to English', 33.47, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Swahili', 'sw', 'Original-language BaseAgent', 42.1, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Swahili', 'sw', 'Google Translate to English', 36.49, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Urdu', 'ur', 'Original-language BaseAgent', 34.15, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1'),\n    ('Urdu', 'ur', 'Google Translate to English', 2.3, 'GPT-4o', 200, 'WebShop Task Score, not binary success', 'https://arxiv.org/html/2505.15372v1')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Web shopping scores: original language versus English translation"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "X-WebAgentBench v1 Table 2 Task Score, not binary success rate. Two selected strategies, not all paper strategies. Translation also changes the environment. Thai lowest original-language observation shown; Swahili counterexample to blanket low-resource weakness. No pooling with other metrics."
        ]
      }
    },
    {
      "id": "evidence_inventory",
      "label": "16 projects: measurements and openness",
      "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
      "notes": "Project inventory, not independent experiment count or exhaustive literature census. Primary papers, author repositories, and data cards checked. Unverified does not mean unavailable. No full repository-by-repository reproduction or legal license review. MAST baseline only as of cutoff. Material inputs: https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md ; https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md ; https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e ; https://arxiv.org/html/2606.28715v1 ; https://github.com/SEACrowd/SEATauBench ; https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://aclanthology.org/2026.findings-eacl.42/ ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS ; https://arxiv.org/html/2506.04135v4 ; https://github.com/showlab/macosworld ; https://arxiv.org/html/2505.15372v1 ; https://github.com/WPENGxs/X-WebAgentBench ; https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2510.07978v1 ; https://github.com/ola-krutrim/VoiceAgentBench ; https://arxiv.org/html/2604.06209v1 ; https://github.com/BrahiM-Mefgouda/TelcoAgent ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://benchmarks.lilt.com/ ; https://github.com/lilt/liltbench-tasks-public ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark ; https://mast-benchmark.github.io/ ; https://arxiv.org/html/2604.04532v1 ; https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
      "query": {
        "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "description": "Project inventory, not independent experiment count or exhaustive literature census. Primary papers, author repositories, and data cards checked. Unverified does not mean unavailable. No full repository-by-repository reproduction or legal license review. MAST baseline only as of cutoff. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"No.\", \"Project\", \"Language coverage\", \"Task layer\", \"Models and evidence layer\", \"Verified openness\", \"Interpretation limits\", \"Task ancestry\", \"Source\") AS (\n  VALUES\n    (1, 'RuBench', 'Russian', 'Code execution', 'Sol, Opus 5 / current direct', 'Tasks and run results public; full grading tests withheld', 'No matched English; models use different task sets', 'Independent native code tasks', 'https://github.com/eugeneshilow/rubench'),\n    (2, 'MERA GorillaHard', 'Russian', 'Static tool plans', 'Sol, Opus 5, Grok 4.6 / current direct', 'Framework, metrics, scores public; incomplete test truth/configuration', 'Tools not executed; submission settings not fully aligned', 'MERA / GorillaHard', 'https://mera.a-ai.ru/en/text'),\n    (3, 'MDPBench / Kimi K3', '13 target observations, including related Arabic label', 'Document component', 'Kimi K3 / current component', 'Official per-language scores public; no rerun here', 'Not agent success; unpaired documents', 'Document parsing component', 'https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e'),\n    (4, 'GAIA-v2-LILT', 'ar de hi ko pt-BR + en', 'Retrieval / tool execution', 'GPT-5.4, Gemini 3.1 Pro, Opus 4.6 / earlier flagships', 'Evaluation code, HF tasks, per-language table public', 'Audit changes function, culture, difficulty together; not pure language causality', 'Audited MAPS-GAIA branch', 'https://arxiv.org/html/2604.24929v1'),\n    (5, 'MAPS', 'Final paper: 12 languages; older HF card: 11', 'GAIA/SWE/safety/math mixture', 'Historical models / context', 'Paper and HF data public; pin subset and version', 'MATH is not execution; GAIA shares ancestry with preceding row', 'Parent GAIA, SWE, MATH, ASB sets', 'https://aclanthology.org/2026.findings-eacl.42/'),\n    (6, 'SEATauBench', 'vi th id Filipino zh + en', 'Multi-turn tools and final state', 'Kimi K2.5, GPT-5-mini, Qwen3 / historical', 'Code, domain data, analysis framework public; MIT repository', 'Simulator can also fail; no isolated agent-only language effect', 'Derived from tau2; shares framework with LILT tau', 'https://github.com/SEACrowd/SEATauBench'),\n    (7, 'macOSWorld', 'ar ja ru zh + en', 'Interactive desktop', 'Claude 3.7 CUA, 2025 OpenAI CUA / earlier flagships', 'Tasks, environment, scoring code public; per-language paper results', 'Instructions and UI switch together; v4 revises some scores', 'Native macOS tasks', 'https://arxiv.org/html/2506.04135v4'),\n    (8, 'X-WebAgentBench', '14 non-English settings; 11 target languages', 'Interactive web shopping', 'GPT-4o and others / historical', 'Code and data entry points public; paper results', 'Task Score is not success rate; derived from WebShop', 'WebShop derivative', 'https://github.com/WPENGxs/X-WebAgentBench'),\n    (9, 'MASSIVE-Agents', '52 languages; 24 targets including related labels', 'Static function / argument matching', 'Nova Premier, Claude 3.5 and others / earlier flagships', 'Converted HF data CC BY 4.0; per-language paper tables', 'Not multi-turn execution; filtering changes samples/function coverage by language', 'MASSIVE derivative; BFCL scoring', 'https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id'),\n    (10, 'Terminal-Bench-LILT', 'ar cs de es hi ja ko sr tr zh', 'Native multilingual code execution', 'GPT-5.5, Opus 4.8 and others / earlier flagships', 'Paper, aggregate scores, samples public; full tasks require contacting authors', 'Blog 300 versus leaderboard 324 tasks: denominators not mixed; full set not public', 'Native Terminal-Bench format; separate from community LILTBench', 'https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark'),\n    (11, 'LILT multilingual tau', 'de ko + en', 'Multi-turn tools and final state', 'GPT-5.4, Opus 4.8, Gemini 3.1 Pro / earlier flagships', 'Public board and run notes; dynamic per-language scores not extracted here', 'Different English/target-language simulators confound scores', 'tau2 derivative; not independent framework replication', 'https://benchmarks.lilt.com/'),\n    (12, 'LILTBench community', 'Authors report 31 languages; not exhaustively checked task by task', 'Native code / English pairs', 'Opus 4.6 / challenging test resources', 'Tasks, verifiers, leaderboard public; Apache-2.0 repository', 'Adversarial hard-task selection cannot estimate population failure rates', 'Community native tasks; distinct from commercial full set', 'https://github.com/lilt/liltbench-tasks-public'),\n    (13, 'VoiceAgentBench', 'en hi bn mr ta te ml', 'Speech-to-tool-call scoring', 'SpeechLM / ASR+LLM / non-frontier resources', 'Evaluation code and HF data entry public; custom community license', 'Multi-turn subset is English only; predicted calls are not final environment success', 'Synthetic speech-tool tasks', 'https://github.com/ola-krutrim/VoiceAgentBench'),\n    (14, 'TelcoAgent-Bench', 'ar + en', 'Diagnostic tool sequences / summaries', '3B–8B models / non-frontier resources', 'Blueprints, tasks, prediction/score JSON public; license unverified', 'Intent/resolution similarity is not strict success; cannot extrapolate to flagships', 'Telecom blueprint tasks', 'https://github.com/BrahiM-Mefgouda/TelcoAgent'),\n    (15, 'MultiAgent-X', '12 languages; target hi sw ha am', 'Function-call data', 'No verified current-frontier baseline / resources', 'Samples, scripts, structure public; full license unverified', 'Synthetic; structural validity is not native quality; pa is not pnb', 'New synthetic tasks; quoted MASSIVE scores are not replication', 'https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark'),\n    (16, 'MAST FIRE 2026', '21-language track union; 14 targets', 'Multi-turn retrieval / answers', 'Public Tongyi-30B baseline / non-frontier', 'Task/corpus entry points and per-language baseline public', 'September 30 final results are future at cutoff; baseline only; English documents and answers', 'BrowseComp-Plus derivative; three Indic languages overlap tracks', 'https://mast-benchmark.github.io/')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "16 projects: measurements and openness"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Project inventory, not independent experiment count or exhaustive literature census. Primary papers, author repositories, and data cards checked. Unverified does not mean unavailable. No full repository-by-repository reproduction or legal license review. MAST baseline only as of cutoff."
        ]
      }
    },
    {
      "id": "broad_atlas_1",
      "label": "Broader evidence map · 1–15",
      "href": "https://aclanthology.org/2025.findings-emnlp.1099.pdf",
      "notes": "Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger. Material inputs: https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2604.24929v1 ; https://arxiv.org/html/2505.15372v1 ; https://arxiv.org/html/2506.04135v4 ; https://arxiv.org/html/2606.28715v1 ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://arxiv.org/html/2510.07978v1 ; https://mast-benchmark.github.io/ ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark",
      "query": {
        "url": "https://aclanthology.org/2025.findings-emnlp.1099.pdf",
        "description": "Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"MASSIVE calls\", \"GAIA audited\", \"X-Web shopping\", \"macOS use\", \"SEA multi-turn\", \"LILT native code\", \"Voice tools\", \"MAST retrieval\", \"MultiAgent-X tasks\", \"Code\", \"Included projects\", \"Mapping limitation\", \"Meaning\") AS (\n  VALUES\n    ('Hindi', 2, 2, 2, 0, 0, 1, 1, 2, 1, 'hi', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Spanish', 2, 0, 2, 0, 0, 1, 0, 2, 0, 'es', 'MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Modern Standard Arabic†', 2, 2, 2, 2, 0, 1, 0, 2, 0, 'ar', 'MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval', 'Generic Arabic/Filipino are related labels, not exact MSA/Tagalog matches', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('French', 2, 0, 2, 0, 0, 0, 0, 2, 0, 'fr', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Bengali', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'bn', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Portuguese', 2, 2, 0, 0, 0, 0, 0, 0, 0, 'pt', 'MASSIVE calls; GAIA audited', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Indonesian', 2, 0, 0, 0, 2, 0, 0, 0, 0, 'id', 'MASSIVE calls; SEA multi-turn', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Urdu', 2, 0, 2, 0, 0, 0, 0, 2, 0, 'ur', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Russian', 2, 0, 2, 2, 0, 0, 0, 2, 0, 'ru', 'MASSIVE calls; X-Web shopping; macOS use; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('German', 2, 2, 2, 0, 0, 1, 0, 2, 0, 'de', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Japanese', 2, 0, 0, 2, 0, 1, 0, 0, 0, 'ja', 'MASSIVE calls; macOS use; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Nigerian Pidgin', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'pcm', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Egyptian Arabic', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'arz', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Marathi', 0, 0, 0, 0, 0, 0, 1, 0, 0, 'mr', 'Voice tools', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Vietnamese', 2, 0, 2, 0, 2, 0, 0, 0, 0, 'vi', 'MASSIVE calls; X-Web shopping; SEA multi-turn', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Broader evidence map · 1–15"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Named benchmarks, not publication counts. Includes historical models and non-frontier resources. Markers cannot be summed into risk or ability. Audited GAIA five languages only; MAPS parent not double counted. Arabic/Filipino related labels carry a dagger."
        ]
      }
    },
    {
      "id": "broad_atlas_2",
      "label": "Broader evidence map · 16–30",
      "href": "https://aclanthology.org/2025.findings-emnlp.1099.pdf",
      "notes": "Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists. Material inputs: https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2604.24929v1 ; https://arxiv.org/html/2505.15372v1 ; https://arxiv.org/html/2506.04135v4 ; https://arxiv.org/html/2606.28715v1 ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://arxiv.org/html/2510.07978v1 ; https://mast-benchmark.github.io/ ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark",
      "query": {
        "url": "https://aclanthology.org/2025.findings-emnlp.1099.pdf",
        "description": "Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language\", \"MASSIVE calls\", \"GAIA audited\", \"X-Web shopping\", \"macOS use\", \"SEA multi-turn\", \"LILT native code\", \"Voice tools\", \"MAST retrieval\", \"MultiAgent-X tasks\", \"Code\", \"Included projects\", \"Mapping limitation\", \"Meaning\") AS (\n  VALUES\n    ('Telugu', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'te', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Swahili', 2, 0, 2, 0, 0, 0, 0, 2, 1, 'sw', 'MASSIVE calls; X-Web shopping; MAST retrieval; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Hausa', 0, 0, 0, 0, 0, 0, 0, 0, 1, 'ha', 'MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Turkish', 2, 0, 2, 0, 0, 1, 0, 0, 0, 'tr', 'MASSIVE calls; X-Web shopping; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Western Punjabi', 0, 0, 0, 0, 0, 0, 0, 0, 0, 'pnb', 'Not included in this map', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Tagalog†', 2, 0, 0, 0, 2, 0, 0, 0, 0, 'tl', 'MASSIVE calls; SEA multi-turn', 'Generic Arabic/Filipino are related labels, not exact MSA/Tagalog matches', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Tamil', 2, 0, 0, 0, 0, 0, 1, 2, 0, 'ta', 'MASSIVE calls; Voice tools; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Iranian Persian', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'fa', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Korean', 2, 2, 0, 0, 0, 1, 0, 0, 0, 'ko', 'MASSIVE calls; GAIA audited; LILT native code', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Amharic', 2, 0, 0, 0, 0, 0, 0, 0, 1, 'am', 'MASSIVE calls; MultiAgent-X tasks', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Thai', 2, 0, 2, 0, 2, 0, 0, 2, 0, 'th', 'MASSIVE calls; X-Web shopping; SEA multi-turn; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Javanese', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'jv', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Italian', 2, 0, 0, 0, 0, 0, 0, 0, 0, 'it', 'MASSIVE calls', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Gujarati', 0, 0, 0, 0, 0, 0, 0, 2, 0, 'gu', 'MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability'),\n    ('Kannada', 2, 0, 0, 0, 0, 0, 0, 2, 0, 'kn', 'MASSIVE calls; MAST retrieval', 'Language-code mapping; no substitution by neighboring languages', '2=verified per-language result; 1=task/aggregate coverage; 0=not included; not ability')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Broader evidence map · 16–30"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Named benchmarks, not publication counts. Combined maps retain 30 languages and 27 with some coverage. Historical/non-frontier results cannot establish current-model performance. No included entry does not mean no research exists."
        ]
      }
    },
    {
      "id": "triangulation",
      "label": "Language signals alongside counterevidence",
      "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
      "notes": "Researcher synthesis, not risk score or statistical ranking. Shared task ancestry, historical models, components, and non-frontier baselines remain distinguished. Recommendations are unrun tests. Material inputs: https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md ; https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md ; https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e ; https://arxiv.org/html/2606.28715v1 ; https://github.com/SEACrowd/SEATauBench ; https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://aclanthology.org/2026.findings-eacl.42/ ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS ; https://arxiv.org/html/2506.04135v4 ; https://github.com/showlab/macosworld ; https://arxiv.org/html/2505.15372v1 ; https://github.com/WPENGxs/X-WebAgentBench ; https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2510.07978v1 ; https://github.com/ola-krutrim/VoiceAgentBench ; https://arxiv.org/html/2604.06209v1 ; https://github.com/BrahiM-Mefgouda/TelcoAgent ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://benchmarks.lilt.com/ ; https://github.com/lilt/liltbench-tasks-public ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark ; https://mast-benchmark.github.io/ ; https://arxiv.org/html/2604.04532v1 ; https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
      "query": {
        "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "description": "Researcher synthesis, not risk score or statistical ranking. Shared task ancestry, historical models, components, and non-frontier baselines remain distinguished. Recommendations are unrun tests. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"Language group\", \"Supporting evidence\", \"Supported interpretation\", \"Counterevidence or limit\", \"Action\") AS (\n  VALUES\n    ('01 Thai', 'SEA retail and other domains; X-Web shopping; K3 documents', 'Weakness across tasks; current-frontier evidence is component-only', 'Thai is the highest one-shot language for Llama 3.1 405B in MASSIVE; models/prompts change order', 'Retest fully localized business and shopping; avoid all-model claims'),\n    ('02 Arabic (generic)†', 'Audited GAIA: three flagships below English; macOS: two CUAs lower; K3 component', 'Aligned historical signals across two execution families', 'GAIA improves after audit; GUI also changes RTL layout; dialects untested', 'Retest search, RTL UI, and numeric arguments; do not extend to Egyptian Arabic'),\n    ('03 Japanese', 'macOS CUA; native code tasks; K3 component', 'Component weakness; execution depends on model', 'OpenAI CUA exceeds English; Claude CUA falls below English', 'Separate document reading, click targeting, and display-width tests'),\n    ('04 Hindi', 'Audited GAIA; native code; MASSIVE; K3 component', 'Several task families; low scores sensitive to task quality', 'GAIA audit gains 25.4–32.7 pp; English instructions do not remove all native-task difficulty in Hindi/German', 'Audit task alignment, entities, and regional rules; distinguish reading from execution'),\n    ('05 Russian', 'Current RuBench execution and GorillaHard plans; historical macOS', 'Concrete current failures, more direct than a component', 'RuBench/MERA lack English pairs; historical OpenAI CUA Russian exceeds English', 'Regress specific failures; do not rank Russian worst overall'),\n    ('06 Vietnamese / Indonesian / Filipino†', 'SEA full business; MASSIVE; Vietnamese X-Web', 'Several tasks; language ordering varies by domain', 'Kimi Vietnamese retail slightly exceeds English; Filipino simulator also errs', 'Accept by business domain and verify Filipino–Tagalog variety alignment'),\n    ('07 German / Korean / Portuguese', 'Audited GAIA; MASSIVE; K3; German/Korean LILT', 'More evidence does not establish deployment readiness', 'Audited German gap is 3.1 pp for GPT-5.4, 12.7 pp for Opus 4.6', 'Test the target model and actual workflow; no universal safe-language list'),\n    ('08 Amharic / Kannada', 'Historical MASSIVE zero-shot AST; Amharic synthetic tasks', 'Risk signal mainly from one historical task family', 'Lowest language differs by model; no current-frontier execution replication', 'Prioritize function arguments; evidence cannot rank current models'),\n    ('09 Swahili / Urdu', 'X-Web; MASSIVE; non-frontier MAST retrieval', 'Task- and model-dependent signals; no uniform direction', 'Swahili original-language GPT-4o shopping is not low; speech and retrieval scores are not interchangeable', 'Measure retrieval, calls, and final state separately with English controls'),\n    ('10 Other coverage / gaps', 'Bengali/Tamil/Telugu tool or retrieval resources; Pidgin speech component', 'Consult all 30 rows; missing evidence is not low ability', 'Egyptian Arabic/Western Punjabi exact matches missing; Hausa lacks current-frontier baseline', 'Use native tasks: pcm≠en, arz≠ar, pnb≠pa, jv≠id')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "Language signals alongside counterevidence"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Researcher synthesis, not risk score or statistical ranking. Shared task ancestry, historical models, components, and non-frontier baselines remain distinguished. Recommendations are unrun tests."
        ]
      }
    },
    {
      "id": "language_lookup",
      "label": "All 30 languages: evidence and next checks",
      "href": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
      "notes": "Researcher mapping retains all target languages, named evidence and inference limits. No invented scores; related labels do not replace exact language varieties. Material inputs: https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md ; https://github.com/MERA-Evaluation/MERA/blob/0e1f4840baa313598266d5e63ac00a9f88a7b1df/benchmark_tasks/gorillahard/README.md ; https://huggingface.co/moonshotai/Kimi-K3/commit/26875c9de9f2fdd76741e66a9dd14e7a97b4fe2e ; https://arxiv.org/html/2606.28715v1 ; https://github.com/SEACrowd/SEATauBench ; https://arxiv.org/html/2604.24929v1 ; https://github.com/lilt/gaia-v2-lilt ; https://aclanthology.org/2026.findings-eacl.42/ ; https://huggingface.co/datasets/Fujitsu-FRE/MAPS ; https://arxiv.org/html/2506.04135v4 ; https://github.com/showlab/macosworld ; https://arxiv.org/html/2505.15372v1 ; https://github.com/WPENGxs/X-WebAgentBench ; https://aclanthology.org/2025.findings-emnlp.1099.pdf ; https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id ; https://arxiv.org/html/2510.07978v1 ; https://github.com/ola-krutrim/VoiceAgentBench ; https://arxiv.org/html/2604.06209v1 ; https://github.com/BrahiM-Mefgouda/TelcoAgent ; https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark ; https://benchmarks.lilt.com/ ; https://github.com/lilt/liltbench-tasks-public ; https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark ; https://mast-benchmark.github.io/ ; https://arxiv.org/html/2604.04532v1 ; https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
      "query": {
        "url": "https://github.com/eugeneshilow/rubench/blob/4b5c8da1b18ea85171b4270cda187ca667ab3c18/1.0/rounds/round-02/RESULTS.md",
        "description": "Researcher mapping retains all target languages, named evidence and inference limits. No invented scores; related labels do not replace exact language varieties. SQLite VALUES reproduces reviewed observations; not a live source query.",
        "sql": "WITH reviewed_published_rows (\"No.\", \"Language\", \"Code\", \"Available evidence\", \"Interpretation and next check\") AS (\n  VALUES\n    (1, 'Hindi', 'hi', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; Voice tools; MAST retrieval; MultiAgent-X tasks', 'Audited GAIA recovers substantially; K3 reading is low. Separate benchmark quality from local-rule execution.'),\n    (2, 'Spanish', 'es', 'MASSIVE calls; X-Web shopping; LILT native code; MAST retrieval', 'GAIA parent set, shopping, function calls, and retrieval coverage. No complete current-frontier comparison; do not presume European-language reliability.'),\n    (3, 'Modern Standard Arabic', 'ar', 'MASSIVE calls; GAIA audited; X-Web shopping; macOS use; LILT native code; MAST retrieval', 'GAIA and macOS give aligned historical signals; K3 component score is lower. Generic Arabic does not precisely establish MSA or Egyptian dialect performance.'),\n    (4, 'French', 'fr', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Shopping, function calls, retrieval baselines, and a current document component. Retest full business execution on the target model.'),\n    (5, 'Bengali', 'bn', 'MASSIVE calls; Voice tools; MAST retrieval', 'Historical calls, voice-tool tasks, and retrieval resources. Current-frontier execution remains unestablished; test entities and arguments.'),\n    (6, 'Portuguese', 'pt', 'MASSIVE calls; GAIA audited', 'Audited GAIA uses Brazilian Portuguese; MASSIVE uses Portugal locale. Regional business rules are not interchangeable.'),\n    (7, 'Indonesian', 'id', 'MASSIVE calls; SEA multi-turn', 'SEA covers full multi-turn workflows with domain-dependent effects. Indonesian does not substitute for Javanese.'),\n    (8, 'Urdu', 'ur', 'MASSIVE calls; X-Web shopping; MAST retrieval', 'Historical/non-frontier calls, shopping, and retrieval. Numerals and dates are proposed mechanisms, not measured current-model failure rates.'),\n    (9, 'Russian', 'ru', 'MASSIVE calls; X-Web shopping; macOS use; MAST retrieval', 'Current code execution and static calls contain failures. Without matched English tasks, Russian causation is unestablished.'),\n    (10, 'German', 'de', 'MASSIVE calls; GAIA audited; X-Web shopping; LILT native code; MAST retrieval', 'Audited GAIA and native code coverage are relatively rich. Gaps differ by model; component strength is not deployment acceptance.'),\n    (11, 'Japanese', 'ja', 'MASSIVE calls; macOS use; LILT native code', 'Current parsing is low; historical CUA gaps change direction by model. Test visual targeting, text, and local software rules separately.'),\n    (12, 'Nigerian Pidgin', 'pcm', 'SwitchBoard speech component', 'Code-switching speech resources found; current-frontier tool execution unverified. Do not substitute English.'),\n    (13, 'Egyptian Arabic', 'arz', 'No matching execution result included', 'Generic Arabic studies cannot fill this dialect. No included exact-dialect agent result.'),\n    (14, 'Marathi', 'mr', 'Voice tools', 'VoiceAgentBench provides speech-tool tasks; do not extend its English-only multi-turn results to Marathi.'),\n    (15, 'Vietnamese', 'vi', 'MASSIVE calls; X-Web shopping; SEA multi-turn', 'SEA, shopping, function calls, and documents provide multiple sources. Kimi fully localized retail slightly exceeds English: preserve the counterexample.'),\n    (16, 'Telugu', 'te', 'MASSIVE calls; Voice tools; MAST retrieval', 'Calls, voice tools, and retrieval resources are reusable. Static calls do not establish long-horizon execution.'),\n    (17, 'Swahili', 'sw', 'MASSIVE calls; X-Web shopping; MAST retrieval; MultiAgent-X tasks', 'Historical task coverage, but GPT-4o original-language shopping is not weak. MAST small-model scores cannot rank frontier risk.'),\n    (18, 'Hausa', 'ha', 'MultiAgent-X tasks', 'MultiAgent-X supplies synthetic calls. Exploratory safety logs were excluded from reliable risk assessment; no current-frontier baseline.'),\n    (19, 'Turkish', 'tr', 'MASSIVE calls; X-Web shopping; LILT native code', 'Calls, shopping, and native code tasks exist. Add current-model end-to-end and local-rule tests.'),\n    (20, 'Western Punjabi', 'pnb', 'No matching execution result included', 'MultiAgent-X and MAST pa/Gurmukhi cannot substitute for Western Punjabi; exact matching results remain missing.'),\n    (21, 'Tagalog', 'tl', 'MASSIVE calls; SEA multi-turn', 'MASSIVE tl-PH and SEA Filipino are related labels. Confirm the user variety, then test multi-turn business execution.'),\n    (22, 'Tamil', 'ta', 'MASSIVE calls; Voice tools; MAST retrieval', 'Calls, voice-tool tasks, and MAST retrieval are reusable. Current-frontier complete execution remains unestablished.'),\n    (23, 'Iranian Persian', 'fa', 'MASSIVE calls', 'MASSIVE fa-IR provides historical calls. Current models still need Persian numerals, calendars, and identifier-preservation tests.'),\n    (24, 'Korean', 'ko', 'MASSIVE calls; GAIA audited; LILT native code', 'Audited GAIA improves and K3 parsing is relatively high. LILT multi-turn results have simulator confounds; broad reliability is unestablished.'),\n    (25, 'Amharic', 'am', 'MASSIVE calls; MultiAgent-X tasks', 'MASSIVE historical weakness is specific; MultiAgent-X supplies resources, not independent performance replication. Retest current flagships.'),\n    (26, 'Thai', 'th', 'MASSIVE calls; X-Web shopping; SEA multi-turn; MAST retrieval', 'Historical SEA and shopping weaknesses plus low K3 parsing motivate retesting. They do not imply every model/task is weak.'),\n    (27, 'Javanese', 'jv', 'MASSIVE calls', 'MASSIVE jv-ID provides historical calls. The missing-evidence claim concerns current execution; Indonesian is not a substitute.'),\n    (28, 'Italian', 'it', 'MASSIVE calls', 'MASSIVE, MAPS parent coverage, and current documents exist. Highest parsing does not imply the most reliable complete agent.'),\n    (29, 'Gujarati', 'gu', 'MAST retrieval', 'MAST supplies retrieval tasks and a non-frontier baseline. Current-frontier tools/business execution evidence remains insufficient.'),\n    (30, 'Kannada', 'kn', 'MASSIVE calls; MAST retrieval', 'Claude 3.5 is historically low on MASSIVE; MAST adds retrieval resources. No current-frontier multi-task replication.')\n)\nSELECT * FROM reviewed_published_rows;",
        "engine": "SQLite (reviewed published observations)",
        "language": "sql",
        "tables_used": [
          "All 30 languages: evidence and next checks"
        ],
        "filters": [
          "Evidence cutoff September 5, 2026; no new model runs",
          "Current models, earlier flagships, and resources separated"
        ],
        "metric_definitions": [
          "Researcher mapping retains all target languages, named evidence and inference limits. No invented scores; related labels do not replace exact language varieties."
        ]
      }
    },
    {
      "id": "expanded-link-32",
      "label": "sea_code",
      "href": "https://github.com/SEACrowd/SEATauBench",
      "query": {
        "url": "https://github.com/SEACrowd/SEATauBench",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "sea_code"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-33",
      "label": "gaia_code",
      "href": "https://github.com/lilt/gaia-v2-lilt",
      "query": {
        "url": "https://github.com/lilt/gaia-v2-lilt",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "gaia_code"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-34",
      "label": "maps",
      "href": "https://aclanthology.org/2026.findings-eacl.42/",
      "query": {
        "url": "https://aclanthology.org/2026.findings-eacl.42/",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "maps"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-35",
      "label": "maps_data",
      "href": "https://huggingface.co/datasets/Fujitsu-FRE/MAPS",
      "query": {
        "url": "https://huggingface.co/datasets/Fujitsu-FRE/MAPS",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "maps_data"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-36",
      "label": "mac_code",
      "href": "https://github.com/showlab/macosworld",
      "query": {
        "url": "https://github.com/showlab/macosworld",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "mac_code"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-37",
      "label": "xweb_code",
      "href": "https://github.com/WPENGxs/X-WebAgentBench",
      "query": {
        "url": "https://github.com/WPENGxs/X-WebAgentBench",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "xweb_code"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-38",
      "label": "massive_data",
      "href": "https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id",
      "query": {
        "url": "https://huggingface.co/datasets/AmazonScience/massive-agents/tree/b6156972182bdf34e68c5b5dfbfe6d30db82f104/massive-full-converted-all-langs-with-id",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "massive_data"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-39",
      "label": "voice",
      "href": "https://arxiv.org/html/2510.07978v1",
      "query": {
        "url": "https://arxiv.org/html/2510.07978v1",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "voice"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-40",
      "label": "voice_code",
      "href": "https://github.com/ola-krutrim/VoiceAgentBench",
      "query": {
        "url": "https://github.com/ola-krutrim/VoiceAgentBench",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "voice_code"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-41",
      "label": "telco",
      "href": "https://arxiv.org/html/2604.06209v1",
      "query": {
        "url": "https://arxiv.org/html/2604.06209v1",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "telco"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-42",
      "label": "telco_code",
      "href": "https://github.com/BrahiM-Mefgouda/TelcoAgent",
      "query": {
        "url": "https://github.com/BrahiM-Mefgouda/TelcoAgent",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "telco_code"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-43",
      "label": "terminal",
      "href": "https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark",
      "query": {
        "url": "https://lilt.com/blog/terminal-bench-lilt-multilingual-coding-benchmark",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "terminal"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-44",
      "label": "lilt",
      "href": "https://benchmarks.lilt.com/",
      "query": {
        "url": "https://benchmarks.lilt.com/",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "lilt"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-45",
      "label": "liltbench",
      "href": "https://github.com/lilt/liltbench-tasks-public",
      "query": {
        "url": "https://github.com/lilt/liltbench-tasks-public",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "liltbench"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-46",
      "label": "multi",
      "href": "https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark",
      "query": {
        "url": "https://github.com/Saurabh-66/MultiAgent-X-Multilingual-Agentic-Function-Calling-Benchmark",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "multi"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-47",
      "label": "mast",
      "href": "https://mast-benchmark.github.io/",
      "query": {
        "url": "https://mast-benchmark.github.io/",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "mast"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-48",
      "label": "judge",
      "href": "https://arxiv.org/html/2604.04532v1",
      "query": {
        "url": "https://arxiv.org/html/2604.04532v1",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "judge"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-49",
      "label": "switch",
      "href": "https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
      "query": {
        "url": "https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "switch"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-50",
      "label": "GAIA-v2-LILT Table 2",
      "href": "https://arxiv.org/html/2604.24929v1#S6",
      "query": {
        "url": "https://arxiv.org/html/2604.24929v1#S6",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "GAIA-v2-LILT Table 2"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-51",
      "label": "SEATauBench Tables 9/10/13",
      "href": "https://arxiv.org/html/2606.28715v1#A6",
      "query": {
        "url": "https://arxiv.org/html/2606.28715v1#A6",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "SEATauBench Tables 9/10/13"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-52",
      "label": "macOSWorld v4 Table 3 and cases",
      "href": "https://arxiv.org/html/2506.04135v4#S5",
      "query": {
        "url": "https://arxiv.org/html/2506.04135v4#S5",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "macOSWorld v4 Table 3 and cases"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    },
    {
      "id": "expanded-link-53",
      "label": "X-WebAgentBench Table 2",
      "href": "https://arxiv.org/html/2505.15372v1#S3",
      "query": {
        "url": "https://arxiv.org/html/2505.15372v1#S3",
        "description": "Primary source supporting adjacent evidence; reviewed September 5, 2026.",
        "tables_used": [
          "X-WebAgentBench Table 2"
        ]
      },
      "notes": "Primary source supporting adjacent evidence; reviewed September 5, 2026."
    }
  ]
}