๐ English โข ็ฎไฝไธญๆ โข ็น้ซไธญๆ โข ๆฅๆฌ่ช โข ํ๊ตญ์ด โข Franรงais โข Deutsch โข Espaรฑol โข Portuguรชs โข ะ ัััะบะธะน
[๐ญ Overview](#-overview) โข [๐ Leaderboard](#-leaderboard) โข [๐ Key Findings](#-key-findings) โข [๐ Quick Start](#-quick-start) โข [๐งฐ Capability Surface & Tools](#-capability-surface--tools) โข [๐ Data & Evaluation](#-data--evaluation) โข [๐ Case Studies](#-case-studies) โข [๐ Documentation](#-documentation) โข [๐๏ธ Project Structure](#๏ธ-project-structure) โข [๐ Citation](#-citation) โข [๐ License](#-license)
| ### 1๏ธโฃ The bottleneck is *privilege granting*, not perception The two "easy" axes are nearly saturated โ read-only compliance (ROC โฅ 92%) and modality-choice accuracy (MCA โฅ 92%) for every capable model. The discriminating axes are the privilege-precision metrics: **Workspace-Permission Precision (WPP) *never* reaches 50%** for any model. Subagents are routinely granted roughly twice the files they actually touch. | ### 2๏ธโฃ Cost and management quality are *decoupled* Main-agent API cost spans **over 100ร** (\$0.8 โ \$93 per run) while SMS spans **under 4ร**. The cheapest open models sit on the Pareto frontier โ `deepseek-v4-pro` reaches SMS 46.4 at just \$1.7 โ while several high-cost models (`gpt-5.5`, `sonnet-4-6`, `kimi-k2.6`) are *dominated* by mid-cost `gemini-3.5-flash`. The open-weight `glm-5.2` ($22.9) also lands on the frontier as the strongest open model. | ### 3๏ธโฃ Scores cluster, behaviors diverge Below the flagship, ten models cluster within a **9.9-point SMS band** (43.9โ53.8), yet their orchestration behaviors differ by **more than an order of magnitude**. The per-subagent forbidden-access rate ranges from 0.48 to 5.78 among capable models (~12ร), and dynamic-workflow use ranges from 8 to 112 invocations. |
|
|
--model JSON