The Autonomy Trust Curve: Changelog¶
This changelog records public versions of the CRUCIBLE methodology paper. Published version-specific artifacts are immutable. Corrections or substantive revisions receive a new version number; they do not silently replace a cited version.
Version 1.1, September 15, 2026¶
Minor revision. No method, claim boundary, or implementation-status row changes.
- Adds section 2.6, Continuous agentic defense operations: OpenAI's Defense Factory reference architecture and reported rates, the incremental raising of agent autonomy "as the results earned trust", the Daybreak access program, and what each implies for CRUCIBLE. Adds references 26 to 28.
- Adds to section 9.2 the requirement that evidence for how much autonomy an offensive-capable model may be granted be independent of the laboratory that trains and gates it.
- Adds the Anthropic pacing proposal, reference 29: the grader-attack characterization of the OpenAI-Hugging Face incident in section 2.4, the capability-checkpoint pairing in section 2.5, embedded third-party evaluators in section 9.2, and an evaluation-awareness limitation in section 10.1.
- Adds the Booz Allen Cyber Weapon Index, reference 30: the closest industry parallel and its findings in section 2.2, its national-program recommendation in section 2.5, its counter-AI effect size in section 5.3, and its configuration-dependent refusal finding in section 6.5.
- States the curve's output as quantified deployment risk per autonomy level in the abstract, sections 1, 8.5, and 11, and the glossary; records an expected-loss extension as future work in section 10.3.
- Notes in section 6.2 that the isolation layering matches published reference architectures and is not the paper's contribution.
- Editorial pass throughout: replaces mannered phrasing with direct statement and joins sentences that previously stood apart. Meaning and structure are otherwise unchanged.
- Withdraws the version 1.0 artifacts from the site; version 1.1 supersedes them. The former version 1.0 URL redirects to the current release, and the version 1.0 entry below remains as the record of what was published.
Version 1.0, August 16, 2026¶
Initial methodology release.
- Defines the declared model-and-harness system as the evaluation target.
- Specifies the frozen scenario battery, deterministic variation, human-supervision ladder, active defender, ephemeral range, evidence custody, and deterministic scoring method.
- Separates operational behavior, ROE adherence, containment, and observation coverage.
- Records implementation and live-validation status without reporting frontier-model benchmark results.
- Publishes the accompanying empirical-study preregistration as version 0.1, a working draft that is not yet frozen.
- Adds a preferred citation, permanent version URL, corrections contact, and checksum manifest.
- Applies pre-publication corrections before first public posting, still as version 1.0: downgrades the certification-policy row from Implemented: Yes to Proposed/nominal; puts the abstract and section 6.3 into designed / will-stream tense to match section 5.3; generalizes section 6.1 Phase 2.5 and the live-stack wording in sections 3.3, 6.2, and 6.5; rewrites the section 1 ECHOTRIBBLE sentence, the authorities-as-customers implication, and the EO 14409 paraphrase.
- Replaces remaining em-dashes, en-dashes, and curly quotes with plain ASCII punctuation before first public posting, still as version 1.0.
- Updates implementation-status rows for ephemeral range lifecycle and isolation/containment evidence after Phase 2.5 completion. Records that the bounded-exploitation arming phase is under way. Makes no broad containment claim. Leaves the abstract and section 2.3 result-deferral wording unchanged.
Corrections policy¶
Send suspected errors to contact@cyber-crucible.com. A correction record will identify the affected version and location, describe the change, and state whether the correction changes any claim, method, or conclusion.
Minor presentation fixes that do not alter content may be recorded as a patch release. Changes to the methodology, claims, evidence boundaries, or conclusions require at least a minor-version release.