{
 "@context": "https://csoai.org/llm-context.json",
 "type": "LLMPageSummary",
 "url": "https://csoai.org/research-transparency",
 "title": "Research Transparency | CSOAI",
 "description": "What CSOAI publishes, what it withholds, and why.",
 "headings": [
  "Research Transparency"
 ],
 "text": "Research Transparency | CSOAI CSOAI About GovBench ProvBench Research Transparency What CSOAI publishes, what it withholds, and why. A governance company that only publishes flattering results is not credible. Below is a running, honest account of research findings behind SOV3 and SOV3³ — including claims we made, then corrected or retracted after closer scrutiny. Confirmed findings link to the technical detail; retracted ones explain exactly what was wrong and why. Across three independent measurement passes, diverse model lineages (e.g. Qwen + Llama + DeepSeek + Gemma + Mistral) consistently outperformed identical-lineage configurations of the same size on a governance-quality battery. The gap between diverse and identical configurations was roughly 6× larger than the gap between different topology shapes (ring vs. pyramid vs. triangle). Practical upshot: which models you combine matters far more than how you arrange them. An early adversarial test reported perfect containment (1.00) even under a simulated majority attack. On review, the test was tautological — it hard-gated on a condition that was true of every attack by construction, so it proved nothing about actual robustness. We retracted the claim and reran a corrected test that separates \"obvious\" breaches (which are gated to zero regardless of votes, by design) from \"laundered\" harm that looks confident but is actually harmful (where the real result is 58–79% containment under 2–3 compromised voting nodes — real, useful, and explicitly not perfect). An internal research note once cited a named news publication in support of a claim. On review the citation did not exist — it was an invented reference. It has been struck from every downstream document. The underlying technical claim it was attached to stands on its own on running code and measured results, not on any citation, and is unaffected. A research note referenced an academic conference acceptance. This was overstated — the correct framing is that submission to that venue is a target, not a confirmed acceptance. The claim has been downgraded to aspirational language throughout. A separate research pass reported that diverse 5-model configurations won on both a clean-data metric and a containment metric, using a distinct code path from our own measurement. We could independently verify only one of the two measurement passes behind this finding on the tree at the time it was reported; the second is treated as a consistent but unverified secondary report until its source code is confirmed on disk. We flag this rather than quietly merging both into one number. We have not run a head-to-head capability benchmark (e.g. GSM8K, MMLU) of SOV3 against frontier models. The governance-topology results above measure decision-quality, safety, and cost under a stated error model — they are a real, useful, and different thing from a capability benchmark, and we do not present one as a substitute for the other. This is an open item, not a hidden one. If we cannot survive our own adversarial review, we have no business selling assurance to anyone else. Every finding above was caught and corrected before — or immediately after — it appeared in any customer-facing material. That correction loop, not a flawless research record, is the actual credibility signal. Updated as findings are confirmed or corrected. Last reviewed 2026-07-12. What this is, and is not. CSOAI measures . It issues no conformity marks, holds no accreditation, and has no enforcement powers. Nothing here is a certification, an attestation of compliance, or legal advice. Every published figure traces to an evidence file; where a number is absent it is because the measurement does not exist yet.",
 "text_truncated": false,
 "register": {
  "role": "measurement_and_attestation_support",
  "csoai_certifies_systems": false,
  "csoai_is_a_notified_body": false,
  "csoai_has_enforcement_powers": false,
  "note": "CSOAI measures and publishes evidence. It issues no conformity marks and holds no accreditation. Nothing here is certification or legal advice."
 },
 "generated_by": "make_llm_json.py"
}