Skip to main content
GET
cURL

Overview

The Model Frontier dataset: benchmark capability against transacted price per task, for every model and thinking level. models is the per-model scatter - Epoch Capabilities Index and SWE-bench Verified / GPQA Diamond / AIME scores, priced with token-weighted OTPI dollars per million tokens as of the last benchmark refresh (costPerTask = price × tpt / 1e6). tests lists the tracked thinking-level benchmarks, efforts carries every (model × thinking level) run per test, and releaseDates maps model display names to first-release dates - join it against the efforts rows to reconstruct the frontier as it stood on any past date. history carries real month-by-month prices where they exist (currently GPQA Diamond only). A model with no real score on an axis is dropped from that view, never estimated. The dataset moves on benchmark refresh, not daily. Requires an API key - there is no free tier.

Authorizations

Authorization
string
header
required

API key passed as a Bearer token. Create one in the dashboard.

Response

The full Model Frontier dataset.