Meta released OmnilingualGAIA2 on Hugging Face, a multilingual extension of the GAIA2 agentic benchmark. The dataset spans 6,400 scenarios across 10 languages, testing AI agents beyond English-centric evaluations.
Key facts
- 6,400 scenarios in OmnilingualGAIA2
- 10 languages covered
- Released on Hugging Face by Meta
- Extension of GAIA2 agentic benchmark
- Original GAIA introduced in 2023
Meta has released OmnilingualGAIA2 on Hugging Face, a multilingual extension of the GAIA2 agentic benchmark according to @HuggingPapers. The dataset comprises 6,400 scenarios across 10 languages, designed to test AI agents' ability to perform real-world tasks in non-English contexts.
The original GAIA benchmark, introduced in 2023, set a high bar for general AI assistants, with questions requiring reasoning, multi-modal understanding, and tool use. GAIA2, its successor, expands this to more complex agentic scenarios. OmnilingualGAIA2 extends this further by translating and adapting these scenarios into multiple languages, likely including major ones like Spanish, French, German, Chinese, and others, though the exact list of languages was not specified in the announcement.
The release comes amid growing recognition that AI agents are disproportionately evaluated in English, leaving gaps in performance for other languages. For example, a 2024 study by Meta AI researchers found that multilingual agents lag in non-English tasks, with performance drops of up to 20% on benchmark tasks when compared to English. This benchmark directly addresses that gap.
What the benchmark measures
OmnilingualGAIA2 retains the agentic focus of GAIA2, which requires models to not just answer questions but to plan, use tools, and synthesize information from multiple sources. The 6,400 scenarios are designed to be "human-verifiable," meaning answers can be checked objectively, a key feature of the original GAIA design. This makes it a robust test of agentic capability rather than mere knowledge recall.
The multilingual dimension adds a layer of complexity: agents must understand instructions in the target language, retrieve information from language-specific sources, and produce outputs in that language. This tests not just translation skills but deeper reasoning and cultural context, which are often lost in simplistic translation-based benchmarks.
Why it matters
The release signals a shift in how the industry evaluates AI agents. Most benchmarks, including the original GAIA, are English-centric, which can mask significant performance gaps in other languages. For enterprises deploying agents globally, this benchmark could become a standard reference point.
It also raises the bar for model evaluation: a model that excels on OmnilingualGAIA2 demonstrates robust multilingual agentic capabilities, a differentiator in a market where many models are optimized for English. The benchmark's release on Hugging Face makes it accessible for researchers and developers to test their own models, potentially accelerating improvements in multilingual agent performance.
The timing is notable: OpenAI and Anthropic have both expanded multilingual support in recent months, and this benchmark provides a common yardstick to compare their progress. However, the source did not disclose whether Meta plans to include OmnilingualGAIA2 in future model evaluations or leaderboards, leaving open the question of how it will be used beyond a research artifact.
Limitations and open questions
The announcement is brief, and several details remain unspecified. The exact list of 10 languages is not provided, nor is the methodology for translating or adapting the scenarios. It is unclear whether the scenarios are direct translations or culturally adapted versions, which could affect difficulty. The source also does not state whether Meta will release baseline scores for its own models, such as Llama 4, on this benchmark.
These gaps matter: without baseline scores, it is hard to gauge how challenging the scenarios are or how current models perform. The benchmark's utility will depend on community adoption and clear documentation, which Meta has not yet provided in the announcement.
What to watch
Watch for the release of baseline scores from Meta's own models on OmnilingualGAIA2, which would establish a reference point for the community. Also track whether the benchmark is adopted in major evaluation suites like HELM or OpenLLM, which would signal its acceptance as a standard. Finally, look for the full language list and documentation on the Hugging Face page, which will clarify the benchmark's scope and methodology.






