AmolfiBench
How AmolfiBench is designedThe models to leave your Amo on.
Grades and cost estimates September 2026
Think of a professor reviewing a finished assignment: did it meet the brief, how much help did it need, and did it show something exceptional?
- SBeyond expectations
- Goes beyond the brief with skill or insight a capable evaluator could not have produced. The surprise is in the substance, not just the presentation.
- AExcellent
- Fulfills the brief with judgment, accuracy and polish. No meaningful complaints or required corrections.
- BStrong with guidance
- Produces good, useful work with clear direction. Needs focused feedback or editing to reach an excellent result.
- CUneven
- Has useful strengths and obvious gaps. Needs substantial guidance and revision; choose its tasks carefully.
- DLimited
- Can serve a narrow purpose, but important gaps and major rework make it a weak general choice.
- FNot recommended
- Fails the essential brief or produces unusable work. There is no compelling reason to choose it for this category over better alternatives.
Your primary Amo. Work considers understanding and context, judgment and planning, delegation, verification and recovery, direct computer use, and follow-through. The model stays responsible for the outcome. Grades are provisional editorial estimates from published task evidence; matched primary-agent supervision has not been measured here.
Reasoning. Each model has a named reference setting where comparable costs exist. Other settings inform its capability limits. We do not assume a most-used setting or invent an Auto usage mix. A measured Auto comparison would require the same work, tools and specialist pool, with a documented per-model routing policy.
S can be earned, but first place does not guarantee it. Plus and minus show placement within a tier. Provisional grades have limited evidence. If a grade is missing, Why? explains the specific gap.
The cost reference. Published workload spending at each model’s named reference setting, normalized within each workload to GPT-6.1 Sol Medium at 1×. The cost-reference mix is professional deliverables 55%, workflows 20%, coding 25%; it is separate from the agent-suitability grade. This is a reference configuration, not a measured Auto mix or complete agent session. Missing routine coverage and costs are disclosed. The logarithmic axis compares proportions. Routing, worker calls, matched computer-use activity, infrastructure and human review are outside this reference.
The speed reference. Standardized output tokens per second at the named cost-reference configuration. Startup, thinking, tool use and workers are excluded; this is not time to finish an agent session. Missing matching observations stay outside 3D.
Select a model to show or hide its point. How we grade
| # | |||
|---|---|---|---|
| 1 | OpenAI · Medium · Provisional | B+ | 7.46× |
| 1 | New Anthropic · Medium · Provisional | B+ | 5.52× |
| 3 | Anthropic · High · Provisional | B | 16.66× |
| 3 | New OpenAI · Medium · Provisional | B | 1× |
| 3 | New xAI · High · Provisional | B | 13.43× |
| 3 | New Anthropic · Medium · Provisional | B | 2.31× |
| 7 | Google · High · Provisional | B- | 5.72× |
| 7 | Meta · Extra High · Preview · Provisional | B- | 5.78× |
| 7 | Alibaba / Qwen · Published setting · Provisional | B- | 17.5× |
| 10 | Google · Published setting · Preview · Provisional | C | 2.18× |
| 10 | New OpenAI · High · Provisional | C | 0.14× |
| 10 | Alibaba / Qwen · Provisional | C | Not availableComparable workload spending unavailable. |
| 13 | MiniMax · Preview · Provisional | C- | Not availableComparable workload spending unavailable. |
| 14 | Google · Published setting · Provisional | D+ | 0.53× |
| 14 | Anthropic · Thinking · Provisional | D+ | 1.07× |
| 14 | Alibaba / Qwen · Published setting · Provisional | D+ | 0.72× |
| 17 | Amazon · Provisional | D | Not availableComparable workload spending unavailable. |