Frontier Artifact Bench
Universal Agent Evaluation via Pairwise Artifact Comparison
Frontier Artifact Bench 1.0
Category-level win rates · Models grouped by overall rank
Overall ranks 1–5
Select a point to inspect its category win rate.
Overall ranks 6–10
Select a point to inspect its category win rate.
Frontier Artifact Bench 1.0
Cost–Performance · Win rate vs. estimated token cost per task
Expert-grounded Pairwise Evaluation
Comparison and aggregation
Judges select a winner or tie using expert guidance. Preferences are averaged within tasks, then equally across tasks.
Scalable task creation
Experts provide instructions, files, and guidance. Fixed baselines require O(N) comparison sets for N agents.
Economically Valuable Task Distribution
FAB 1.0 contains 240 expert-authored tasks across eight categories. Each task consists of an instruction, input files, and evaluation guidance, without a task-specific scoring program.

Metallography image analysis
Input: one alloy micrograph. Output: grain-size and porosity measurements, a segmentation overlay, and an auditable report.
Calibration Against Human Judgments
| Agentic judge | Pearson correlation | Mean absolute error | Model orderings preserved |
|---|---|---|---|
| Muse Spark 1.3 | 0.95 | 3.1 percentage points | 41 / 45 |
| GPT 5.6 Sol | 0.89 | 8.1 percentage points | 39 / 45 |
We use Muse Spark 1.3 as the agentic judge for the reported benchmark results. In the human-judgment calibration, it more closely reproduces human model-level results than GPT 5.6 Sol: Pearson correlation is 0.95 versus 0.89, and mean absolute error is 3.1 versus 8.1 percentage points. It also preserves 41 of 45 human model orderings, compared with 39 of 45 for GPT 5.6 Sol.