An open, installable, dialect-split Arabic function-calling benchmark. Measures tool selection, argument extraction, keeping Arabic in arguments instead of transliterating, and dialectal framing.
Real OpenRouter runs, provenance-frozen by git SHA. See RESULTS.md. Adding more models (incl. Hermes via native endpoint) is open work.
Bundles: 2026-04-clean-seven (51 items, clean), 2026-04-adversarial-seven (24 items, Arabic guard surface) · updated 2026-04
Tool-call completion across 6 categories.
| # | model | provider | score | 95% CI | bidi viol | dialect drift |
|---|---|---|---|---|---|---|
| 1 | z-ai/glm-5.1 | openrouter | 0.839 | 0.750–0.925 | 0.0% | 0.0% |
| 2 | google/gemini-3.1-pro-preview | openrouter | 0.816 | 0.730–0.907 | 0.0% | 0.0% |
| 3 | minimax/minimax-m2.7 | openrouter | 0.809 | 0.717–0.888 | 0.0% | 0.0% |
| 4 | anthropic/claude-opus-4.7 | openrouter | 0.803 | 0.700–0.891 | 0.0% | 0.0% |
| 5 | qwen/qwen3.6-plus | openrouter | 0.798 | 0.711–0.893 | 0.0% | 0.0% |
| 6 | moonshotai/kimi-k2.5 | openrouter | 0.634 | 0.539–0.767 | 0.0% | 0.0% |
| 7 | openai/gpt-5.4 | openrouter | 0.543 | 0.431–0.678 | 0.0% | 0.0% |
Arabic guard surface: BiDi / homoglyph / UTS #39 / Arabizi / injection / canonicalization / dialect pressure. Lower is expected — hard on purpose.
| # | model | provider | score | 95% CI | bidi viol | dialect drift |
|---|---|---|---|---|---|---|
| 1 | qwen/qwen3.6-plus | openrouter | 0.407 | 0.234–0.615 | 4.2% | 0.0% |
| 2 | anthropic/claude-opus-4.7 | openrouter | 0.342 | 0.181–0.540 | 0.0% | 0.0% |
| 3 | minimax/minimax-m2.7 | openrouter | 0.324 | 0.169–0.494 | 0.0% | 0.0% |
| 4 | moonshotai/kimi-k2.5 | openrouter | 0.320 | 0.151–0.477 | 0.0% | 0.0% |
| 5 | z-ai/glm-5.1 | openrouter | 0.280 | 0.136–0.526 | 0.0% | 0.0% |
| 6 | google/gemini-3.1-pro-preview | openrouter | 0.259 | 0.122–0.418 | 0.0% | 0.0% |
| 7 | openai/gpt-5.4 | openrouter | 0.243 | 0.100–0.404 | 0.0% | 0.0% |