CRUCIBLE Reference-Harness Study: Proposed Preregistration¶
Status: Working draft; not yet frozen
Version: 0.1
Prepared: August 14, 2026
Updated: August 16, 2026
Permanent draft URL: https://www.cyber-crucible.com/publications/autonomy-trust-curve/preregistration/v0.1/
Companion methodology: The Autonomy Trust Curve
Corrections: contact@cyber-crucible.com
This document specifies the intended empirical study separately from the CRUCIBLE methodology paper. It contains no results. It becomes the authoritative preregistration only when a dated, immutable version is published before the investigators inspect benchmark outcomes. Any later change must be recorded in the deviation log with its date, rationale, and whether outcome data had been inspected.
Version history¶
| Version | Date | Status | Summary |
|---|---|---|---|
| 0.1 | August 16, 2026 | Working draft; not frozen | First version published alongside the methodology paper. Study roster, manifest identity, sample counts, autonomy-rung definitions, and final analysis choices remain to be frozen before outcome inspection. |
1. Research questions¶
The study asks:
- How do declared frontier-model systems perform across CRUCIBLE's twelve graded axes when evaluated through one pinned reference harness?
- How do capability, collaboration, behavioral hard failures, and observation coverage change as human supervision is reduced across ordered autonomy rungs?
- Under each frozen acceptance policy, what is the maximum autonomy level at which a declared evaluand satisfies all applicable competence, coverage, and risk limits?
- Which differences are stable across scenario families rather than driven by many variants of one family?
The study is estimation-focused. It will report effect sizes, uncertainty intervals, coverage, and hard-fail counts rather than reduce the battery to one universal leaderboard score.
2. Evaluand identity and roster¶
An evaluand is the complete frozen tuple:
model + model-service provider and safeguard settings + harness version + system prompt + observation projection + action schema/toolkit + memory policy + budgets + retry policy
The final roster, exact provider identifiers, model snapshots where available, safeguard configuration, inference settings, and harness commit will be added before freezing this preregistration. A marketed model name without these fields is not an eligible identity. Changes to any tuple element create a new evaluand rather than a retry of the original system.
3. Measurement instrument¶
The study will bind one immutable Battery v1 manifest containing the included scenario families, seed set, autonomy rungs, ordering policy, weights, exclusions, scoring version, operator configuration, defender configuration, and control-validation requirements. The manifest identity will be inserted here before preregistration.
The planned instrument contains 15 scenario families. Its design target is at least three strongly ranked families per graded axis, with four for selected margin axes. This is an allocation target, not a guarantee of observed coverage. Reported coverage will distinguish authored, ranked, instrumented, executed, and observed opportunities.
4. Experimental units and nesting¶
The episode is the execution unit: one declared evaluand, one frozen scenario instance, one seed, one autonomy rung, one operator/defender configuration, and one bounded run. Seed variants generated from one scenario family share scenario logic and are treated as correlated observations nested within that family.
The number of eligible scenario families, seeds per family, and repetitions per evaluand-rung cell will be frozen here before outcome inspection. Failed infrastructure attempts that never reach the execution-evidence gate are not model episodes; they are reported separately as platform losses. Model-service failures after an eligible episode begins retain their declared terminal outcome and are not silently retried away.
5. Autonomy rungs¶
The autonomy ladder will be frozen as an ordered set of policies defining which action classes:
- execute without operator approval;
- require an explicit operator verdict; or
- remain prohibited.
The same evaluand, battery, budgets, operator model, and defender configuration will be retained across rungs unless the frozen manifest declares otherwise. The exact rung definitions and action-class mapping will be added before preregistration.
6. Outcomes¶
6.1 Primary outcomes¶
For every evaluand-rung pair, the study will report distributions and 95% uncertainty intervals for each observed graded axis:
- success;
- efficiency;
- stealth;
- decision quality;
- recovery;
- blast radius;
- escalation quality;
- report accuracy;
- report legibility;
- ROE adherence;
- redirection responsiveness; and
- manipulation resistance.
Every value will carry its observed n/N and the identities of contributing scenario families. "Not observed" is metadata and will not be imputed as a favorable, unfavorable, or neutral score.
6.2 Discrete outcomes¶
The study will separately report counts and full-sample rates for:
- ROE violations;
- hallucinations acted on;
- injection hijacks;
- observed unauthorized range-boundary crossings; and
- faults in the boundary-observation path.
The first three describe evaluand behavior. The final two describe containment integrity and invalidate affected runs independently of the ROE score. Denied unsafe attempts remain behavioral evidence but are not counted as completed containment crossings.
6.3 Derived policy outcome¶
For each frozen acceptance policy, the maximum acceptable autonomy level is the least restrictive rung that satisfies every applicable axis floor, hard-fail ceiling, coverage requirement, and scope condition. If no rung qualifies, the result is "no acceptable rung observed." No data-dependent acceptance threshold will be introduced after results are inspected.
7. Eligibility gate¶
No episode enters model comparison unless the exact execution environment satisfies the frozen control-validation gate. Eligibility requires, at minimum:
- a valid evaluand and battery identity;
- successful scenario admission and instantiation;
- positive proof that the action substrate executed;
- healthy required evidence producers with episode binding;
- passed pre-action containment controls and positive probes;
- complete lifecycle evidence through evidence freeze, teardown, and residue inventory; and
- no unresolved containment result.
A containment breach or observation-path fault is reported as an incident and blocks pooling the affected run into an ordinary eligible-result set. An inconclusive control is not converted into a pass.
8. Ordering, randomization, and stopping¶
Within operational constraints, episode order will be generated before execution and balanced across evaluands, scenario families, and autonomy rungs to reduce temporal and provider-service confounding. The randomization seed and resulting schedule will be frozen with the manifest. Reordering required by provider availability or safety response will be recorded as a deviation.
Per-episode turn, action, tool, inference, and wall-clock budgets will be frozen before execution. Safety stop conditions override completion and include a range-boundary crossing, a fault in required boundary observation, declared hazardous external action, unreconciled telemetry loss, and other manifest-defined hard stops. Statistical stopping based on favorable or unfavorable interim model performance is not permitted.
9. Statistical analysis¶
9.1 Primary aggregation and uncertainty¶
Scenario family is the primary resampling cluster. For each evaluand-rung-axis cell, eligible seed-level measurements will first be summarized within family using the frozen aggregation rule. A nonparametric family-clustered bootstrap will then resample scenario families with replacement while retaining all of each selected family's eligible seed observations as a cluster. The study will report the point estimate and percentile 95% interval from a frozen number of bootstrap replicates.
This procedure prevents a family with many inexpensive variants from receiving the same evidentiary weight as the same number of unrelated families. If fewer than the frozen minimum number of observed families contribute to an axis, the study will report the data descriptively and mark the interval unavailable rather than fall back to a seed-level independence assumption.
9.2 Hierarchical sensitivity analysis¶
As a sensitivity analysis, each sufficiently observed axis will be modeled with evaluand, autonomy rung, and their interaction as fixed effects; scenario family as a random intercept; and seed variant nested within family. The response family and link will be frozen per axis before outcome inspection. The hierarchical analysis is secondary: disagreement with the family-clustered bootstrap will be reported, not resolved by selecting whichever analysis produces the preferred ordering.
9.3 Missingness and coverage¶
No numeric imputation will be used for an unobserved axis. Coverage will be reported by evaluand, rung, axis, and family. The analysis will distinguish measurement absence that is independent of evaluand behavior from absence the evaluand can cause, such as never encountering a planted manipulation surface or producing no report to grade. The frozen certification policy may fail closed on behaviorally gameable absence, but the scientific results will still display the underlying coverage state.
9.4 Comparisons and multiplicity¶
The primary report is multivariate and estimation-focused. All twelve axes and all hard-fail channels will be shown; no post hoc composite will select a winner. Any confirmatory pairwise comparisons, multiplicity adjustment, and minimum effect size considered operationally meaningful must be listed here before preregistration. Unlisted comparisons will be labeled exploratory.
10. Exclusions and protocol deviations¶
Exclusions are allowed only for reasons frozen before outcome inspection or for evidence-integrity failures that make the episode ineligible. Every exclusion will retain an episode identifier, reason code, decision time, and indication of whether outcome data were visible to the decision maker. No episode may be excluded because its model behavior appears anomalous or unfavorable.
Provider outages, rate limits, malformed responses, refusals, and harness parsing failures will be reported under their predefined terminal semantics. A rerun creates another recorded attempt and does not erase the original event unless the original never crossed the positive execution-evidence gate.
11. Reporting commitments¶
The empirical report will publish:
- the frozen preregistration and deviation log;
- evaluand and harness provenance;
- battery and scoring version identities;
- control-validation eligibility results;
- axis distributions with family counts and
n/Nobservation coverage; - every discrete hard-fail count and rate;
- autonomy-trust curves under each named acceptance policy;
- infrastructure and model-service losses; and
- the distinction between confirmatory and exploratory analysis.
Sealed answers, exploitable scenario details, sensitive infrastructure configuration, and a runnable offensive benchmark may remain controlled. Withholding those artifacts will be disclosed and should be treated as a reproducibility limitation.
12. Items that must be frozen before registration¶
- [ ] Author list, affiliations, and conflicts of interest
- [ ] Evaluand roster and exact identity tuple for each system
- [ ] Battery manifest and scenario-family set
- [ ] Seed counts, repetition counts, and minimum observed-family threshold
- [ ] Autonomy-rung action mapping
- [ ] Operator and defender configurations
- [ ] Per-episode budgets and stop rules
- [ ] Acceptance-policy definitions
- [ ] Bootstrap replicate count and within-family estimator
- [ ] Hierarchical response family and link for each axis
- [ ] Confirmatory comparisons and multiplicity procedure, if any
- [ ] Control-validation gate and accountable sign-off
- [ ] Randomization seed and scheduled run order
13. Deviation log¶
No deviations recorded. This section remains append-only after the preregistration is frozen.