Live signed measurement · renders from the board API

The Axes Arena. Nobody grades the globe.

Fourteen measured GSPC axes, each with a leader, a fleet spread and a separation verdict — rendered live from /api/gspc, never hand-typed. Click two axes to compare head-to-head. Measurement, not certification.

Head-to-head

Pick two axes to compare.

sort loading… Get Pro →

Loading the live board…

The cross — machine vs human, never blended

Where a published human baseline exists, we compare it to the measured fleet — in two separate registers. Open the divergence map →

Loading the cross…

How this is measured & verified

Two registers, never blended. MEASURED = this fleet's scores. REPORTED = published human baselines. A divergence shows only where BOTH exist; UNMEASURED stays UNMEASURED.
Probable, not opaque. Every number derives from the manual, replayable measurement — verify a card offline with curl+openssl.
Measurement, not certification. We grade; we never vouch. Methodology DOI 10.5281/zenodo.21991104 — cite the concept DOI.
Every run is data. Each measured run is signed, replayable, and becomes retrievable eval evidence — a live benchmark, not a static leaderboard. Data-generation, not a vendor list.

Battle — per-axis Elo (adjudicated, not voted)

A blind-A/B rule checklist grades each round — the predicate, never an LLM-as-judge. Pick an axis to see its per-axis Elo table (a brain gets a rating per axis, not one flat number).

seed: leaders beat fleet baseline · deterministic

Loading the battle table…

Every number traces to a live, replayable measurement. Field board pack · insurers pack