Get the Model Frontier dataset
The Model Frontier dataset: benchmark capability against transacted price per task. models is the per-model scatter - Epoch Capabilities Index plus SWE-bench Verified / GPQA Diamond / AIME scores, priced with token-weighted OTPI dollars per million tokens as of the last benchmark refresh (costPerTask = price x tpt / 1e6). tests lists the tracked thinking-level benchmarks, efforts carries every (model x thinking level) run per test, and releaseDates maps model display names to first-release dates (YYYY-MM-DD) - join it against the efforts rows to reconstruct the frontier as of any past date. history carries real month-by-month prices where they exist (currently GPQA Diamond only); for the tracked benchmarks, past frontier states come from efforts + releaseDates. The dataset refreshes on benchmark refresh rather than daily, so responses are stable between refreshes. Requires an API key - there is no free tier.
Overview
The Model Frontier dataset: benchmark capability against transacted price per task, for every model and thinking level.models is the per-model scatter - Epoch Capabilities Index and SWE-bench Verified / GPQA Diamond / AIME scores, priced with token-weighted OTPI dollars per million tokens as of the last benchmark refresh (costPerTask = price × tpt / 1e6). tests lists the tracked thinking-level benchmarks, efforts carries every (model × thinking level) run per test, and releaseDates maps model display names to first-release dates - join it against the efforts rows to reconstruct the frontier as it stood on any past date. history carries real month-by-month prices where they exist (currently GPQA Diamond only). A model with no real score on an axis is dropped from that view, never estimated. The dataset moves on benchmark refresh, not daily. Requires an API key - there is no free tier.Authorizations
API key passed as a Bearer token. Create one in the dashboard.
Response
The full Model Frontier dataset.

