Frontier Artifact Bench

Universal Agent Evaluation via Pairwise Artifact Comparison

Frontier Artifact Bench 1.0

Win rate (%) · higher is better
GPT 6 Astra
Claude Fable 5.1
GPT 5.6 Sol
Claude Opus 5
Claude Opus 4.8
Muse Spark 1.3
Grok 4.6
GPT 5.5
Gemini 3.8 Flash
Gemini 3.7 Flash

Frontier Artifact Bench 1.0

Category-level win rates · Models grouped by overall rank

Overall ranks 1–5

Category win rates, zero to 100 percent20406080100CreativeDesignSoftwareEngineeringDataAnalyticsVideoAgenticScienceGameDevelopmentHealthcareSpatialAgentic

Select a point to inspect its category win rate.

Overall ranks 6–10

Category win rates, zero to 100 percent20406080100CreativeDesignSoftwareEngineeringDataAnalyticsVideoAgenticScienceGameDevelopmentHealthcareSpatialAgentic

Select a point to inspect its category win rate.

Frontier Artifact Bench 1.0

Cost–Performance · Win rate vs. estimated token cost per task

Pareto frontier
Frontier Artifact Bench: cost–performance Pareto frontierLogarithmic cost axis from 1 to 64 dollars per task; win rate axis from 0 to 70 percent. GPT 6 Astra: 55.8 percent at $40.42 per task; Claude Fable 5.1: 52.6 percent at $12.70 per task; GPT 5.6 Sol: 51.1 percent at $5.47 per task; Claude Opus 5: 50.2 percent at $21.63 per task; Claude Opus 4.8: 41.9 percent at $21.74 per task; Muse Spark 1.3: 40.7 percent at $3.56 per task; Grok 4.6: 32.1 percent at $8.11 per task; GPT 5.5: 26.7 percent at $5.12 per task; Gemini 3.8 Flash: 24 percent at $3.49 per task; Gemini 3.7 Flash: 23.6 percent at $2.10 per task. Exact values are also available in the leaderboard table.010203040506070$1$2$4$8$16$32$64Win rate (%)Token cost per task (USD, log scale)GPT 6 AstraClaude Fable 5.1GPT 5.6 SolClaude Opus 5Claude Opus 4.8Muse Spark 1.3Grok 4.6GPT 5.5Gemini 3.8 FlashGemini 3.7 FlashGPT 6 Astra: 55.8% · $40.42Claude Fable 5.1: 52.6% · $12.70GPT 5.6 Sol: 51.1% · $5.47Claude Opus 5: 50.2% · $21.63Claude Opus 4.8: 41.9% · $21.74Muse Spark 1.3: 40.7% · $3.56Grok 4.6: 32.1% · $8.11GPT 5.5: 26.7% · $5.12Gemini 3.8 Flash: 24% · $3.49Gemini 3.7 Flash: 23.6% · $2.10

Expert-grounded Pairwise Evaluation

FAB pairwise comparison protocol with baseline artifacts A task package is given to an evaluated agent, which produces a candidate artifact. A pairwise judge compares that artifact with a pre-prepared near-state-of-the-art baseline artifact using expert evaluation guidance and task-specific rubrics. Preferences are aggregated into win rates and Elo ratings without an all-pairs tournament. Task package 01 Instruction 02 Input files 03 Evaluation guidance Agent Instruction + files only Produces candidate artifact Evaluation guidance Pairwise judge Inspect artifacts side by side 01 Candidate vs. baseline 02 Choose candidate, baseline, or tie 03 Apply expert guidance and rubrics Baseline artifacts Built by multiple agents and humans Near-SOTA reference pool Aggregation Win rate Elo rating Fixed baselines: O(N) vs. O(N2) all-pairs
The baseline-anchored pairwise protocol. Baseline artifacts are produced by multiple agents and humans and curated to represent near-SOTA performance. Evaluation guidance is supplied to the judge, not the evaluated agent.

Comparison and aggregation

Judges select a winner or tie using expert guidance. Preferences are averaged within tasks, then equally across tasks.

Scalable task creation

Experts provide instructions, files, and guidance. Fixed baselines require O(N) comparison sets for N agents.

Economically Valuable Task Distribution

FAB 1.0 contains 240 expert-authored tasks across eight categories. Each task consists of an instruction, input files, and evaluation guidance, without a task-specific scoring program.

One representative task per category
DomainTasks
Metallography image analysis. Task input micrograph, shown instead of the artifact.
Example 01 / Science

Metallography image analysis

Input: one alloy micrograph. Output: grain-size and porosity measurements, a segmentation overlay, and an auditable report.

Calibration Against Human Judgments

Agreement with human model-level results
Agentic judgePearson correlationMean absolute errorModel orderings preserved
Muse Spark 1.30.953.1 percentage points41 / 45
GPT 5.6 Sol0.898.1 percentage points39 / 45

We use Muse Spark 1.3 as the agentic judge for the reported benchmark results. In the human-judgment calibration, it more closely reproduces human model-level results than GPT 5.6 Sol: Pearson correlation is 0.95 versus 0.89, and mean absolute error is 3.1 versus 8.1 percentage points. It also preserves 41 of 45 human model orderings, compared with 39 of 45 for GPT 5.6 Sol.