Every benchmark result for every model — and what each one actually costs.
Benchmark Heaven (formerly Model Market Comparison) has its primary base URL at benchmarkheaven.com (2026-09-11). The previous host model-market-comparison.app.mintapis.com remains valid and serves the same app and API endpoints. The GitHub repo, file paths, data schema, IDs, units, Composite and settings keys did not move. Browser preferences are stored per hostname and are not automatically transferred to the new domain.
The daily refresh stages source changes and requires a cheap worker plus a different-family critic before publication. Each observation keeps its source and date; unavailable or contested candidates do not erase accepted data.
Compare open-source and frontier LLMs by capability and price in one place. Capability comes from ArtificialAnalysis (Coding, Coding Agent & Intelligence indices) and Intelligence.ai / DesignArena (Agentic Web Dev Frontend & Full-Stack Elo). Prices are aggregated across OpenRouter inference providers, AWS Bedrock, Azure AI Foundry, Google Vertex AI, Nebius, Inceptron, TensorX, Scaleway, IONOS, Mistral, Chutes, OVHcloud, STACKIT, T-Systems LLM Hub, TrustedTokens, GitHub Copilot, and the Anthropic / Claude Code list price — normalized to USD per 1M tokens.
📋 Product spec: PRD.md · 🗂️ Data schema: data/SCHEMA.md · 🔄 Data collection & refresh: data/SCRAPING.md · 🔌 Public API: API.md · 🚀 Deploy / migrate / env vars: DEPLOYMENT.md · 📜 Data/URL changes for consumers: CHANGELOG.md
Explore benchmark rankings, compare up to four exact model configurations on Compare, or use the standalone Radar. Model pages include complete benchmark sheets, source links and dates, missing coverage, and explainable profile signals. Every comparison keeps benchmark versions separate. Methodology and limitations.
JevBench is our own typed-decision benchmark for Jev-class models. The live board identifies the current release and links its aggregate results and frozen version page. Rankings, per-model pages, an alternatives guide, and a how-to-choose guide are all on the site; this repo hosts the scoring code and data.
Training on JevBench's public split is allowed and should be declared with each submission. Rankings continue to use all benchmark items. We report held-out results separately so that specialisation on public tasks is visible. Held-out means not publicly released, not guaranteed unseen: hosted systems receive these tasks during evaluation. We periodically issue fresh tasks to reduce the value of prior exposure.
- Price & provider filters (apply to price comparisons and model offers; benchmark exploration uses its own filters, persisted): selectable score, min-score, One variant for Reasoning models (collapse GPT/Claude/GLM/Kimi to one), Featured, Hide deprecated (on by default), Exclude Chinese providers, EU-hosted / approved equivalent only, Non-US provider only, TEE / confidential only, plus provider- and model-checklist filters.
- Selectable scores: Composite (seven percentile slots, including general and Software Engineering ECI, with model-mean imputation plus a dominance-safe projection, 0–100, default), ArtificialAnalysis Coding Index, Coding Agent Index (median across published harnesses for the exact model/effort variant), Intelligence Index, and DesignArena Frontend / Full-Stack Elo.
- Overview — fully sortable table with per-column filters, Has benchmark evidence for Composite (or Has score for a direct metric) / Has provider toggles, and Excel-style data bars on score & cost.
- Compare — pick two models (A/B) head-to-head; scores & cheapest price shown big with the winner highlighted, plus each model's cheapest providers side by side.
- Cost vs Capability — scatter: x = cost (cheapest 10:1 blended $/1M, axis inverted so cheaper is right), y = capability, circle = closed / square = open, with a Pareto frontier line of best-value models.
- Charts — capability leaderboard, cheapest-model ranking, open vs closed comparisons.
- Model detail — top-5 cheapest providers and all offers by platform within the active global filter scope, full benchmark breakdown, reasoning-variant comparison, current GitHub Copilot token/AI-Credit prices and legacy per-request cost.
- Providers per Model — searchable model picker → that model's providers ranked by price.
- Provider explorer — pick one provider → every model it offers; click a model to compare that provider's price against all others (ranked, with its price-rank in the field).
- Gateways — 35+ LLM gateways/routers/aggregators compared on EU routing capability, self-host/local, open-source license, HQ and pricing.
- EU & Sovereign — which providers are EU-hosted/sovereign and which SOTA models they actually serve (incl. TEE/confidentiality notes + a separate dedicated/BYOC-only list). Interactive EU filtering requires evidence on the exact model offer; provider-level dedicated/BYOC capability alone never turns a global route into an EU-hosted one. The only policy exceptions are Azure Direct Global DeepSeek V4 Pro and Kimi K2.7 Code: they qualify as company-approved equivalents without being relabeled as technically EU-resident.
- Public read-only JSON API (CORS-enabled) —
/api/dataset(full export),/api/models,/api/models/[id],/api/providers,/api/meta,/api/health. See API.md.
Featured models include GPT-6 Astra, GLM-5.3 Flash, Qwen3.8 Max (including the separately priced 0902 release), Muse Spark 1.3, and GPT-5.6 Sol/Terra/Luna, GPT-5.5 / GPT-5.4 (with Mini and low/medium/high/xhigh settings), Claude Opus 4.8 / 4.7 / 4.6, Sonnet 4.6 / 5, Fable 5, Kimi K2.5 / K2.6 / K2.7-Coding, GLM 5.1 / 5.2, MiniMax M2.5 / M2.7 / M3, Xiaomi MiMo-V2.5-Pro and DeepSeek V4 Pro.
Next.js (App Router, TypeScript) · Recharts · PostgreSQL · live at benchmarkheaven.com (Sandy/Coolify, snapshot mode); the former host model-market-comparison.app.mintapis.com remains valid and serves the same app.
The app reads from Postgres when DATABASE_URL is set (seeded from
data/dataset.json) and otherwise serves the committed snapshot, so it always
renders even without a database.
npm install
npm run dev # http://localhost:3000 (uses bundled data/dataset.json)export ARTIFICIAL_ANALYSIS_API_KEY=aa_… # for the AA fetch
npm run data:refresh # fetch live APIs + rebuild dataset.json
# with a database:
export DATABASE_URL=postgres://…
npm run db:seedSee data/SCRAPING.md for per-source details (incl. how to update the manually scraped AWS / Azure / Copilot / Claude snapshots).
Full guide incl. required env vars / secrets, Postgres setup, Docker and Azure
hosting: DEPLOYMENT.md. In short — the only runtime secret is the
optional DATABASE_URL (the app falls back to the bundled data/dataset.json snapshot
without it); ARTIFICIAL_ANALYSIS_API_KEY is needed only to refresh data, not at runtime.
A render.yaml Blueprint and a Dockerfile are included.
Benchmark Heaven is open source under the MIT licence — a hobby project, built to give the world a better tool to decide which LLM fits a job. The licence covers this repository's code only. The benchmark results, prices and other data it collects are third-party data and are not relicensed: each source keeps its own terms (Artificial Analysis — attribution required; Epoch AI — CC BY 4.0; DesignArena — two boards under a documented risk decision; OpenRouter; and the benchmark maintainers and provider catalogs named on the site's Sources & methodology page).
If Benchmark Heaven helps you choose a model, you can Support this project. Payments go to productivity-boost.com Betriebs UG (haftungsbeschränkt) & Co. KG, the one-person company behind these projects.
10:1 blended cost = (10·input + 1·output) / 11 per 1M tokens. EU residency is
audited per offer/model/region; an EU-capable provider does not make its US or global
routes EU-hosted. Separately, eu_policy_equivalent admits only Azure Direct Global
DeepSeek V4 Pro and Kimi K2.7 Code to the EU filter under this company's legal/business
classification; inference may occur outside the EU and the flag is not a technical residency
guarantee. The Composite uses seven capability slots: AA Coding,
source-matched AA Coding Agent, AA Intelligence, Epoch general ECI, Epoch Software
Engineering ECI, DesignArena Frontend and DesignArena Full-Stack. AA values are clamped to
0–100. Epoch ECI stays on its native scale in its own score view and is percentile-normalized
inside the Composite. A DesignArena board qualifies with at least
200 battles (an app minimum aligned with the source's typical preliminary/reliability threshold) and its Elo is converted to the expected score against a fixed Elo 1000
opponent: 100 / (1 + 10^((1000 − Elo) / 400)). Each observed slot is then converted
to its empirical percentile among the current catalog's unique observed values. Every
missing slot is assigned that model's mean percentile across its observed slots. The
base is therefore exactly the mean of its available percentiles, so missing slots cannot
move it. A deterministic least-squares projection then enforces shared-evidence dominance:
if one model covers every reliable slot of another measured model and is
no worse in any shared slot (strictly better in at least one), it remains at least 0.1 points
ahead. This is the smallest symmetric catalog-wide adjustment satisfying those constraints;
the base and adjustment are exposed separately in model details and the models API. A model
with no reliable observed slot receives the neutral fallback 50. Coverage
is tracked separately so this fallback is not mistaken for benchmark evidence when choosing
a family representative, applying Has benchmark evidence, or building capability charts and the
Pareto frontier. GitHub Copilot's
current token/AI-Credit rates and legacy annual-plan
request multipliers are kept separate from provider API offers.
Model identities are joined conservatively across sources; modes, releases, pricing tiers
and hosting routes remain separate unless a stable source id/repository proves equivalence.
When Intelligence.ai publishes only a bare product identity, that family-scoped result is
attached exactly once to the deterministic active representative used by collapsed views,
with an explicit provenance note that it does not identify the tested effort setting.
Figures may still change upstream —
verify before relying on them. Not affiliated with any provider.