Defined scenarios
Simulated target scenarios and rules of engagement establish the scope of each evaluation.
CRUCIBLE evaluates and benchmarks AI agents for defined scenarios and policies. It produces bounded evidence for an evaluated release—not an authorization decision.
What we do
Simulated target scenarios and rules of engagement establish the scope of each evaluation.
Scored trajectories use MITRE ATT&CK vocabulary within stated scenario and policy boundaries.
Agent-alone benchmarks and human-plus-agent evaluations are distinct measurements; results are not interchangeable.
A release-, scenario-, and policy-specific result can inform a sponsor or customer decision; it does not replace authorization.
Delivery boundary
Delivery is designed to support customer-led integration into a customer-controlled IL7/TS enclave. The customer, sponsor, and AO own the boundary, classification, dependencies, IdP, PKI, KMS, SIEM, storage, network, CDS, host, operations, and authorization decisions. The public-AWS rehearsal and pre-decisional package are not a customer deployment, an authorization, or proof of operational control effectiveness.
Documentation
Scope, measurement modes, reproducibility, and the current nominal-calibration status.
Use the release-matched guide for customer-led deployment and integration.
Customer-furnished controls and authorization remain outside this product.
Commercial context and bounded evaluation options.
Canonical console and role guidance lives with the integration documentation.
Contact
For evaluation, customer-led integration, and acquisition inquiries: