04 /Leaderboard
Model leaderboard
Ranked over checked-in fixture results · prompt v2 · 8 cases per model.
Fixture mode — synthetic data, no live models. These pages render committed, reproducible results from the runnable harness in
benchmarks/ai-eval. Models are provider-neutral aliases; datasets are synthetic and openly licensed. Latency and cost are representative fixtures, never live measurements. No employer or customer data, prompts, or IP appear anywhere.01 /Ranking
Overall, then by dimension
| # | Model | Overall | Correct | Grounded | Format | Safety | Latency | $/1k |
|---|---|---|---|---|---|---|---|---|
| 1 | Frontier Afrontier Highest-capability tier stand-in — most accurate, most expensive, higher latency. | 100% | 100% | 100% | 100% | 100% | 1.69 s | $1.694 |
| 2 | Balanced Bbalanced Mid tier — strong accuracy at a fraction of the cost. | 99% | 96% | 100% | 100% | 100% | 886 ms | $0.335 |
| 3 | Compact Ccompact Smallest/cheapest/fastest tier — trades accuracy and robustness for price and latency. | 47% | 65% | 50% | 75% | 0% | 328 ms | $0.057 |
Aliases are capability-tier stand-ins, not specific vendors. The point is the evaluation method — where each tier gains and loses points — not a vendor ranking. Latency and cost are representative fixtures.