CRUCIBLE

Bounded evaluation for AI agents in cyber operations.

CRUCIBLE evaluates and benchmarks AI agents for defined scenarios and policies. It produces bounded evidence for an evaluated release—not an authorization decision.

Contact us

  • Defined scenarios
  • Distinct evaluation modes
  • Bounded evidence

What we do

Evaluation evidence, kept in scope.

Defined scenarios

Simulated target scenarios and rules of engagement establish the scope of each evaluation.

ATT&CK-grounded measurement

Scored trajectories use MITRE ATT&CK vocabulary within stated scenario and policy boundaries.

Separate evaluation modes

Agent-alone benchmarks and human-plus-agent evaluations are distinct measurements; results are not interchangeable.

Evidence for defined decisions

A release-, scenario-, and policy-specific result can inform a sponsor or customer decision; it does not replace authorization.

Delivery boundary

Designed for customer-led enclave integration.

Delivery is designed to support customer-led integration into a customer-controlled IL7/TS enclave. The customer, sponsor, and AO own the boundary, classification, dependencies, IdP, PKI, KMS, SIEM, storage, network, CDS, host, operations, and authorization decisions. The public-AWS rehearsal and pre-decisional package are not a customer deployment, an authorization, or proof of operational control effectiveness.

Documentation

Read the canonical guidance.

Evaluation methodology

Scope, measurement modes, reproducibility, and the current nominal-calibration status.

Deployment guidance

Use the release-matched guide for customer-led deployment and integration.

Acquisition

Commercial context and bounded evaluation options.

Role guides

Canonical console and role guidance lives with the integration documentation.

Contact

Get in touch

For evaluation, customer-led integration, and acquisition inquiries:

contact@cyber-crucible.com