Same benchmark, separate publisher runs
Model rankings, with receipts
Pick a benchmark. Compare publisher-reported scores, open every source, and see missing results as missing.
Coding
Artificial Analysis Coding Index · Higher is better
Price
USD per 1M tokens, blended 3:1 input to output · Lower is better
Speed
Output tokens per second · Higher is better
Latency
Seconds to first answer token, including reasoning · Lower is better
Source: Artificial Analysis ↗
Score spread
Publisher-reported DeepSWE v1.1, on one zero-to-100 scale.
DeepSWE price performance
Directional only: publisher harnesses and effort settings may differ.
- #1533.9GLM-5.3-Flash
Launch discount
$0.12 / 1M · 63.4
- #2266.9GLM-5.3-Flash
List
$0.24 / 1M · 63.4
- #387.4MiniMax M3
Standard
$0.70 / 1M · 61.2%
- #478.9GPT-5.4 mini
Standard
$0.79 / 1M · 62.1%
- #560.8Kimi K3
Standard
$1.05 / 1M · 63.8%
- #643.5Gemini 3.7 Flash
Introductory
$1.50 / 1M · 65.3%
- #722.9Qwen3.8-2.4T-A95B
Model Studio API
$2.48 / 1M · 56.6
- #822.0Grok 4.6
Standard
$3.00 / 1M · 65.9%
- #921.8Gemini 3.7 Flash
Standard
$3.00 / 1M · 65.3%
- #106.5GPT-5.6 Sol
Current list price
$11.25 / 1M · 72.7%
Unpriced
- GLM-5.366.9%
- Qwen3.8-Flash-Next58.7
Revised 2026-09-02
Method, sources, and raw data
- What is ranked
- Publisher scores inside one named benchmark. No composite score.
- What a gap means
- No publisher result. MicroRouter does not estimate one.
- How to compare
- Check harness, scaffold, and effort settings before comparing.