{
 "@context": "https://csoai.org/llm-context.json",
 "type": "LLMPageSummary",
 "url": "https://csoai.org/benchmarks.html",
 "title": "Benchmarks — measured CSOAI benchmarks | CSOAI",
 "description": "CSOAI benchmarks: GovBench (193 items, 26 dimensions), ProvBench (Article 50 provenance survival), the AI Act frozen-split harness, and the Sovereign Council panel. Every figure recomputes from the published harness.",
 "headings": [
  "Benchmarks",
  "Live benchmarks",
  "Beta / draft benchmarks",
  "Recompute anything we publish",
  "What we do not do"
 ],
 "text": "Benchmarks — measured CSOAI benchmarks | CSOAI Home Products Docs FAQ Benchmarks CSOAI publishes measurement harnesses and signed results. Each benchmark is a frozen scenario set, a deterministic grader, a published run record, and a recompute-able evidence layer. Every figure on this page is a measurement, not a certification. Live benchmarks GovBench — 193 items across 26 dimensions of AI compliance. Published harness, frozen splits, Ed25519-signed results. Open leaderboard. ProvBench — pre-registered test of whether an EU AI Act Article 50 provenance marking survives the real-world transforms an asset actually goes through. 0 of 20 embedded manifests survived any measured transform. AI Act frozen-split harness — closed-book statutory citation measurement. 24 items, GitHub-pinned SHA. Hugging Face dataset. Sovereign Council panel — multi-model measurement, not a vote. Same frozen prompts, several models, published per-model scores with confidence intervals. How the panel works. Beta / draft benchmarks GSPC-GOV — 24 items, EU AI Act risk-tier classification. Run via the public CSOAI-ORG/gspc-harness Inspect task. GSPC-XR — cross-axis integrity verifier (8 items, agentic provision lookup). n_eff 1.21 of 3 — the refused council measurement that published the Illusion Index. Recompute anything we publish Every benchmark ships with its harness. The AI Act frozen-split harness contains scenarios, the split function, and the scoring code. Clone, run, compare. The result is the same number we publish, up to model variability between identical-model versions. What we do not do Issue certifications. We measure; a regulator certifies. Replace notified bodies where the law requires them. Quote a number below usable_n ≥ 30 with a confidence interval. The dishonesty of doing so is the entire reason ProvBench exists. CSOAI Ltd · UK company 16939677 · Every published figure traces to a signed, verifiable record.",
 "text_truncated": false,
 "register": {
  "role": "measurement_and_attestation_support",
  "csoai_certifies_systems": false,
  "csoai_is_a_notified_body": false,
  "csoai_has_enforcement_powers": false,
  "note": "CSOAI measures and publishes evidence. It issues no conformity marks and holds no accreditation. Nothing here is certification or legal advice."
 },
 "generated_by": "make_llm_json.py"
}