Financial-services pilot audit — 30-item measurement scope
2026-08-09 · CSOAI — measurement, not claim
Scope: a 30-item measurement battery for FCA/PRA-aligned AI deployments,
sampled from the published instruments — the same harnesses behind
the public benchmark. Measured instrument performance today:
care-gate recall 100%, over-block 0%. Your stack gets the
same treatment, measured — never a paper opinion.
The 30-item battery
EAT care floor (12) — Governance-scored refusal on Art 5 practices — measured recall
EU AI Act risk-tier (6) — creditworthiness, recruitment, consumer chatbots
ProvBench survival (5) — Article 50 provenance through real transforms
GSPC S-axis (4) — security posture of the model+harness stack
Two-sided refusal (3) — TPR vs false-refusal on a 30-sample battery
Deliverables
Signed measurement report (Ed25519, recompute-able) for each of the 30 checks.
A scorecard entry you can publish, or keep private.
Why this is not a certification
We measure; regulators and accredited bodies decide. Every figure is a measurement on an
open harness, signed and independently re-runnable — the same discipline as the public
benchmark. ProvBench exists because the provenance of the
evidence itself must survive; it does not here either, until it is measured.