Skip to content

Latest commit

 

History

1,759 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Benchmark Heaven

Every benchmark result for every model — and what each one actually costs.

Benchmark Heaven (formerly Model Market Comparison) has its primary base URL at benchmarkheaven.com (2026-09-11). The previous host model-market-comparison.app.mintapis.com remains valid and serves the same app and API endpoints. The GitHub repo, file paths, data schema, IDs, units, Composite and settings keys did not move. Browser preferences are stored per hostname and are not automatically transferred to the new domain.

The daily refresh stages source changes and requires a cheap worker plus a different-family critic before publication. Each observation keeps its source and date; unavailable or contested candidates do not erase accepted data.

Compare open-source and frontier LLMs by capability and price in one place. Capability comes from ArtificialAnalysis (Coding, Coding Agent & Intelligence indices) and Intelligence.ai / DesignArena (Agentic Web Dev Frontend & Full-Stack Elo). Prices are aggregated across OpenRouter inference providers, AWS Bedrock, Azure AI Foundry, Google Vertex AI, Nebius, Inceptron, TensorX, Scaleway, IONOS, Mistral, Chutes, OVHcloud, STACKIT, T-Systems LLM Hub, TrustedTokens, GitHub Copilot, and the Anthropic / Claude Code list price — normalized to USD per 1M tokens.

📋 Product spec: PRD.md · 🗂️ Data schema: data/SCHEMA.md · 🔄 Data collection & refresh: data/SCRAPING.md · 🔌 Public API: API.md · 🚀 Deploy / migrate / env vars: DEPLOYMENT.md · 📜 Data/URL changes for consumers: CHANGELOG.md

Benchmark exploration

Explore benchmark rankings, compare up to four exact model configurations on Compare, or use the standalone Radar. Model pages include complete benchmark sheets, source links and dates, missing coverage, and explainable profile signals. Every comparison keeps benchmark versions separate. Methodology and limitations.

JevBench public and held-out tasks

JevBench is our own typed-decision benchmark for Jev-class models. The live board identifies the current release and links its aggregate results and frozen version page. Rankings, per-model pages, an alternatives guide, and a how-to-choose guide are all on the site; this repo hosts the scoring code and data.

Training on JevBench's public split is allowed and should be declared with each submission. Rankings continue to use all benchmark items. We report held-out results separately so that specialisation on public tasks is visible. Held-out means not publicly released, not guaranteed unseen: hosted systems receive these tasks during evaluation. We periodically issue fresh tasks to reduce the value of prior exposure.

Features

  • Price & provider filters (apply to price comparisons and model offers; benchmark exploration uses its own filters, persisted): selectable score, min-score, One variant for Reasoning models (collapse GPT/Claude/GLM/Kimi to one), Featured, Hide deprecated (on by default), Exclude Chinese providers, EU-hosted / approved equivalent only, Non-US provider only, TEE / confidential only, plus provider- and model-checklist filters.
  • Selectable scores: Composite (seven percentile slots, including general and Software Engineering ECI, with model-mean imputation plus a dominance-safe projection, 0–100, default), ArtificialAnalysis Coding Index, Coding Agent Index (median across published harnesses for the exact model/effort variant), Intelligence Index, and DesignArena Frontend / Full-Stack Elo.
  • Overview — fully sortable table with per-column filters, Has benchmark evidence for Composite (or Has score for a direct metric) / Has provider toggles, and Excel-style data bars on score & cost.
  • Compare — pick two models (A/B) head-to-head; scores & cheapest price shown big with the winner highlighted, plus each model's cheapest providers side by side.
  • Cost vs Capability — scatter: x = cost (cheapest 10:1 blended $/1M, axis inverted so cheaper is right), y = capability, circle = closed / square = open, with a Pareto frontier line of best-value models.
  • Charts — capability leaderboard, cheapest-model ranking, open vs closed comparisons.
  • Model detail — top-5 cheapest providers and all offers by platform within the active global filter scope, full benchmark breakdown, reasoning-variant comparison, current GitHub Copilot token/AI-Credit prices and legacy per-request cost.
  • Providers per Model — searchable model picker → that model's providers ranked by price.
  • Provider explorer — pick one provider → every model it offers; click a model to compare that provider's price against all others (ranked, with its price-rank in the field).
  • Gateways — 35+ LLM gateways/routers/aggregators compared on EU routing capability, self-host/local, open-source license, HQ and pricing.
  • EU & Sovereign — which providers are EU-hosted/sovereign and which SOTA models they actually serve (incl. TEE/confidentiality notes + a separate dedicated/BYOC-only list). Interactive EU filtering requires evidence on the exact model offer; provider-level dedicated/BYOC capability alone never turns a global route into an EU-hosted one. The only policy exceptions are Azure Direct Global DeepSeek V4 Pro and Kimi K2.7 Code: they qualify as company-approved equivalents without being relabeled as technically EU-resident.
  • Public read-only JSON API (CORS-enabled) — /api/dataset (full export), /api/models, /api/models/[id], /api/providers, /api/meta, /api/health. See API.md.

Featured models include GPT-6 Astra, GLM-5.3 Flash, Qwen3.8 Max (including the separately priced 0902 release), Muse Spark 1.3, and GPT-5.6 Sol/Terra/Luna, GPT-5.5 / GPT-5.4 (with Mini and low/medium/high/xhigh settings), Claude Opus 4.8 / 4.7 / 4.6, Sonnet 4.6 / 5, Fable 5, Kimi K2.5 / K2.6 / K2.7-Coding, GLM 5.1 / 5.2, MiniMax M2.5 / M2.7 / M3, Xiaomi MiMo-V2.5-Pro and DeepSeek V4 Pro.

Stack

Next.js (App Router, TypeScript) · Recharts · PostgreSQL · live at benchmarkheaven.com (Sandy/Coolify, snapshot mode); the former host model-market-comparison.app.mintapis.com remains valid and serves the same app.

The app reads from Postgres when DATABASE_URL is set (seeded from data/dataset.json) and otherwise serves the committed snapshot, so it always renders even without a database.

Local development

npm install
npm run dev          # http://localhost:3000  (uses bundled data/dataset.json)

Refresh the data

export ARTIFICIAL_ANALYSIS_API_KEY=aa_…    # for the AA fetch
npm run data:refresh                        # fetch live APIs + rebuild dataset.json
# with a database:
export DATABASE_URL=postgres://…
npm run db:seed

See data/SCRAPING.md for per-source details (incl. how to update the manually scraped AWS / Azure / Copilot / Claude snapshots).

Deploy / migrate

Full guide incl. required env vars / secrets, Postgres setup, Docker and Azure hosting: DEPLOYMENT.md. In short — the only runtime secret is the optional DATABASE_URL (the app falls back to the bundled data/dataset.json snapshot without it); ARTIFICIAL_ANALYSIS_API_KEY is needed only to refresh data, not at runtime. A render.yaml Blueprint and a Dockerfile are included.

Licence

Benchmark Heaven is open source under the MIT licence — a hobby project, built to give the world a better tool to decide which LLM fits a job. The licence covers this repository's code only. The benchmark results, prices and other data it collects are third-party data and are not relicensed: each source keeps its own terms (Artificial Analysis — attribution required; Epoch AI — CC BY 4.0; DesignArena — two boards under a documented risk decision; OpenRouter; and the benchmark maintainers and provider catalogs named on the site's Sources & methodology page).

Support

If Benchmark Heaven helps you choose a model, you can Support this project. Payments go to productivity-boost.com Betriebs UG (haftungsbeschränkt) & Co. KG, the one-person company behind these projects.

Methodology & caveats

10:1 blended cost = (10·input + 1·output) / 11 per 1M tokens. EU residency is audited per offer/model/region; an EU-capable provider does not make its US or global routes EU-hosted. Separately, eu_policy_equivalent admits only Azure Direct Global DeepSeek V4 Pro and Kimi K2.7 Code to the EU filter under this company's legal/business classification; inference may occur outside the EU and the flag is not a technical residency guarantee. The Composite uses seven capability slots: AA Coding, source-matched AA Coding Agent, AA Intelligence, Epoch general ECI, Epoch Software Engineering ECI, DesignArena Frontend and DesignArena Full-Stack. AA values are clamped to 0–100. Epoch ECI stays on its native scale in its own score view and is percentile-normalized inside the Composite. A DesignArena board qualifies with at least 200 battles (an app minimum aligned with the source's typical preliminary/reliability threshold) and its Elo is converted to the expected score against a fixed Elo 1000 opponent: 100 / (1 + 10^((1000 − Elo) / 400)). Each observed slot is then converted to its empirical percentile among the current catalog's unique observed values. Every missing slot is assigned that model's mean percentile across its observed slots. The base is therefore exactly the mean of its available percentiles, so missing slots cannot move it. A deterministic least-squares projection then enforces shared-evidence dominance: if one model covers every reliable slot of another measured model and is no worse in any shared slot (strictly better in at least one), it remains at least 0.1 points ahead. This is the smallest symmetric catalog-wide adjustment satisfying those constraints; the base and adjustment are exposed separately in model details and the models API. A model with no reliable observed slot receives the neutral fallback 50. Coverage is tracked separately so this fallback is not mistaken for benchmark evidence when choosing a family representative, applying Has benchmark evidence, or building capability charts and the Pareto frontier. GitHub Copilot's current token/AI-Credit rates and legacy annual-plan request multipliers are kept separate from provider API offers. Model identities are joined conservatively across sources; modes, releases, pricing tiers and hosting routes remain separate unless a stable source id/repository proves equivalence. When Intelligence.ai publishes only a bare product identity, that family-scoped result is attached exactly once to the deterministic active representative used by collapsed views, with an explicit provenance note that it does not identify the tested effort setting. Figures may still change upstream — verify before relying on them. Not affiliated with any provider.

Brand: direction, usage and assets · three visual concepts.

About

Compare open-source & frontier LLM prices across providers (OpenRouter, AWS Bedrock, Azure Foundry, GitHub Copilot, Claude Code) with ArtificialAnalysis & DesignArena benchmarks

Resources

Stars

15 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages