Research

Everything here is open: the item banks, the graders, the raw runs, and the signed attestation chains. Where a number is not yet defensible, we say so on the page rather than rounding it up.

Papers

The instrument: twelve axes, honest states

GSPC measures AI-system conduct against named statutory provisions. Each axis is an item bank with a deterministic grader. Axes that have not earned a score show their state, not a number:

Datasets and item banks: huggingface.co/csoai · Kaggle mirrors on the GSPC datasets · leaderboard harness: GovBench · results and evidence: /benchmarks.

Method commitments

  1. Deterministic grading. A regex reads the model's committed answer; an equality assertion decides. No model judges another model. 98.9% agreement against 92 hand-labelled responses, zero false serves.
  2. Dead-item analysis, certified. At a fleet of 8 models, 35 of 90 items appeared dead; at a certified fleet of 22 (false-dead rate 0.028 at p=0.85), 8. Apparent consensus at small N is a sampling artefact. We publish the count.
  3. Refusal to quote. Five of six axes sit below the usable-n ≥ 30 threshold, so we publish no point estimate for them. The harness is structurally unable to emit one.
  4. Signed evidence. Every result is bound into an Ed25519 / ML-DSA-65 signed attestation chain (in-toto/DSSE) that a third party can recompute. Verify a signature at proofof.ai/verify.

Open contributions

Practice guides

Ten sector white papers (EU AI Act Article 50, DORA, healthcare, defence, IoT, SOC 2, UK AI Bill, Canada AIDA) live at proofof.ai/whitepapers.

Corrections and retractions

Where the estate has been wrong, the correction is published with the same prominence as the claim. Five retractions (2026-08-05) cover: a fine-tune leaderboard climb with 0.00 delta, a composition result with n_eff 1.21, a decorrelation router that was an artefact, a behavioural win that was 63% train-on-test, and a refusal effect later found in the stock base model. Each retraction names the artefact that disproved it.