Arabic Agent Eval — Leaderboard

An open, installable, dialect-split Arabic function-calling benchmark. Measures tool selection, argument extraction, keeping Arabic in arguments instead of transliterating, and dialectal framing.

Real OpenRouter runs, provenance-frozen by git SHA. See RESULTS.md. Adding more models (incl. Hermes via native endpoint) is open work.
Bundles: 2026-04-clean-seven (51 items, clean), 2026-04-adversarial-seven (24 items, Arabic guard surface) · updated 2026-04

Clean (51 items)

Tool-call completion across 6 categories.

#modelproviderscore95% CIbidi violdialect drift
1z-ai/glm-5.1openrouter0.8390.750–0.9250.0%0.0%
2google/gemini-3.1-pro-previewopenrouter0.8160.730–0.9070.0%0.0%
3minimax/minimax-m2.7openrouter0.8090.717–0.8880.0%0.0%
4anthropic/claude-opus-4.7openrouter0.8030.700–0.8910.0%0.0%
5qwen/qwen3.6-plusopenrouter0.7980.711–0.8930.0%0.0%
6moonshotai/kimi-k2.5openrouter0.6340.539–0.7670.0%0.0%
7openai/gpt-5.4openrouter0.5430.431–0.6780.0%0.0%

Adversarial (24 items)

Arabic guard surface: BiDi / homoglyph / UTS #39 / Arabizi / injection / canonicalization / dialect pressure. Lower is expected — hard on purpose.

#modelproviderscore95% CIbidi violdialect drift
1qwen/qwen3.6-plusopenrouter0.4070.234–0.6154.2%0.0%
2anthropic/claude-opus-4.7openrouter0.3420.181–0.5400.0%0.0%
3minimax/minimax-m2.7openrouter0.3240.169–0.4940.0%0.0%
4moonshotai/kimi-k2.5openrouter0.3200.151–0.4770.0%0.0%
5z-ai/glm-5.1openrouter0.2800.136–0.5260.0%0.0%
6google/gemini-3.1-pro-previewopenrouter0.2590.122–0.4180.0%0.0%
7openai/gpt-5.4openrouter0.2430.100–0.4040.0%0.0%