Skip to benchmark

AmolfiBench

How AmolfiBench is designed
Benchmark category

The models to leave your Amo on.

Graph dimensions
AmolfiBench current model comparisonModels whose icons would overlap share a badge divided into one selectable logo slice per model. Zoom separates nearby placements; shared badges do not mean equal scores or costs. Lower cost is on the right. Letter bands run from F to S. Models without a grade or comparable cost appear in the table. Work cost uses a logarithmic relative scale, with GPT-6.1 Sol at 1×.GradeAMOLFIBENCHSABCDF20×10×5×1×0.2×0.1×GPT-6 AstraOpus 5.5Fable 5.1GPT-6.1 SolGrok 4.7Sonnet 5.5Gemini 3.8 FlashMuse Spark 1.3Qwen 3.8 MaxGemini 3.1 Pro PreviewGPT-6 LunaGemini 3.5 Flash-LiteHaiku 4.5Qwen 3.7 PlusCost
Select a point to inspect a model.
Reviewed October 1, 202617 current models
Grades and cost estimates September 2026

Think of a professor reviewing a finished assignment: did it meet the brief, how much help did it need, and did it show something exceptional?

SBeyond expectations
Goes beyond the brief with skill or insight a capable evaluator could not have produced. The surprise is in the substance, not just the presentation.
AExcellent
Fulfills the brief with judgment, accuracy and polish. No meaningful complaints or required corrections.
BStrong with guidance
Produces good, useful work with clear direction. Needs focused feedback or editing to reach an excellent result.
CUneven
Has useful strengths and obvious gaps. Needs substantial guidance and revision; choose its tasks carefully.
DLimited
Can serve a narrow purpose, but important gaps and major rework make it a weak general choice.
FNot recommended
Fails the essential brief or produces unusable work. There is no compelling reason to choose it for this category over better alternatives.

Your primary Amo. Work considers understanding and context, judgment and planning, delegation, verification and recovery, direct computer use, and follow-through. The model stays responsible for the outcome. Grades are provisional editorial estimates from published task evidence; matched primary-agent supervision has not been measured here.

Reasoning. Each model has a named reference setting where comparable costs exist. Other settings inform its capability limits. We do not assume a most-used setting or invent an Auto usage mix. A measured Auto comparison would require the same work, tools and specialist pool, with a documented per-model routing policy.

S can be earned, but first place does not guarantee it. Plus and minus show placement within a tier. Provisional grades have limited evidence. If a grade is missing, Why? explains the specific gap.

The cost reference. Published workload spending at each model’s named reference setting, normalized within each workload to GPT-6.1 Sol Medium at 1×. The cost-reference mix is professional deliverables 55%, workflows 20%, coding 25%; it is separate from the agent-suitability grade. This is a reference configuration, not a measured Auto mix or complete agent session. Missing routine coverage and costs are disclosed. The logarithmic axis compares proportions. Routing, worker calls, matched computer-use activity, infrastructure and human review are outside this reference.

The speed reference. Standardized output tokens per second at the named cost-reference configuration. Startup, thinking, tool use and workers are excluded; this is not time to finish an agent session. Missing matching observations stay outside 3D.

Select a model to show or hide its point. How we grade

Current models, editorial grades and estimated costs. Select a rated model with a comparable cost to show or hide its point.
#
1
OpenAI · Medium · Provisional
B+
7.46×
1
New
Anthropic · Medium · Provisional
B+
5.52×
3
Anthropic · High · Provisional
B
16.66×
3
New
OpenAI · Medium · Provisional
B
1×
3
New
xAI · High · Provisional
B
13.43×
3
New
Anthropic · Medium · Provisional
B
2.31×
7
Google · High · Provisional
B-
5.72×
7
Meta · Extra High · Preview · Provisional
B-
5.78×
7
Alibaba / Qwen · Published setting · Provisional
B-
17.5×
10
Google · Published setting · Preview · Provisional
C
2.18×
10
New
OpenAI · High · Provisional
C
0.14×
10
Alibaba / Qwen · Provisional
C
Not availableComparable workload spending unavailable.
13
MiniMax · Preview · Provisional
C-
Not availableComparable workload spending unavailable.
14
Google · Published setting · Provisional
D+
0.53×
14
Anthropic · Thinking · Provisional
D+
1.07×
14
Alibaba / Qwen · Published setting · Provisional
D+
0.72×
17
Amazon · Provisional
D
Not availableComparable workload spending unavailable.