Status & Evidence Center

Collection Quality, Evidence & Measurement Readiness Reports

Status and evidence center tracking verification, measurement readiness, and release status across all 12 catalog collections. Not a benchmark leaderboard; no single score or download packages.

Catalog entries

12

5 roots + 7 thematic views

Root collections with historical records

2

Only 2 root baselines have historical records

Evaluated effect reports

0

Not yet evaluated (0 verified reports published)

Downloadable collections

0

All payloads quarantined/restricted

How to Read Safety Report Dimensions

Safety evaluation spans defense, utility, over-refusal, and indeterminate error handling; L1/L2/L3 are observation perspectives and evidence depth, not a quality ladder.

1

Attack Defense

Whether adversarial inputs or malicious payloads are identified and blocked.

2

Normal Utility

Whether normal tasks succeed on clean controls, preserving legitimate functionality.

3

Over-refusal Control

Whether benign requests are misclassified as attacks, measuring false alarms.

4

Evidence Source & Depth

Differentiated by L1 text proxy, L2 tool trace, and L3 sandbox state oracles.

5

Indeterminate & Error Dimension

If execution encounters environment anomalies, missing data, or incomplete evidence chains, it must be marked Indeterminate / unjudgeable, never counted as a safety pass.

Only 2 legacy baselines hold historical structure records; the other 10 remain quarantined or restricted without fabricated scores.

Collection Material format Structural evidence Verified effect measurement Access control Recommended next step
Agent safety baseline agent-safety-baseline

Agent instruction, permission and data checks. Request-level intent; tools pending. Unmeasured.

Mixed benchmark records Historical audit structure recorded Pending (0 reports) Restricted · No download Inspect & select thematic subsets →
Data exfiltration agent-safety-baseline--data-exfiltration

Agent data-exfiltration checks. Request-level intent; tools pending. Effect unmeasured.

Labeled benchmark records No independent audit or evaluation report Pending (0 reports) Restricted · No download Inspect & configure egress sandbox →
Instruction boundaries agent-safety-baseline--instruction-boundaries

Agent trusted-instruction checks. Request-level intent; tools pending. Effect unmeasured.

Labeled benchmark records No independent audit or evaluation report Pending (0 reports) Restricted · No download Inspect & test layered prompts →
Unauthorized actions agent-safety-baseline--unauthorized-actions

Agent tool authorization checks. Request-level intent; tools pending. Effect unmeasured.

Labeled benchmark records No independent audit or evaluation report Pending (0 reports) Restricted · No download Inspect & verify confirmation gate →
AgentDojo paired candidates candidate-agentdojo-paired

Paired attack/control in a resettable world. Intended L2·L3, not verified.

Runtime environment reference Upstream reference, no independent report Pending (0 reports) Quarantined · No download Inspect & review upstream docs →
Banking tasks candidate-agentdojo-paired--banking

Paired attack/control in a banking environment. Intended L2·L3, not verified.

Runtime environment reference Upstream reference, no independent report Pending (0 reports) Quarantined · No download Inspect & setup reset sandbox →
Messaging tasks candidate-agentdojo-paired--slack

Paired attack/control in a messaging environment. Intended L2·L3, not verified.

Runtime environment reference Upstream reference, no independent report Pending (0 reports) Quarantined · No download Inspect & setup reset sandbox →
Travel tasks candidate-agentdojo-paired--travel

Paired attack/control in a travel environment. Intended L2·L3, not verified.

Runtime environment reference Upstream reference, no independent report Pending (0 reports) Quarantined · No download Inspect & setup reset sandbox →
Workspace tasks candidate-agentdojo-paired--workspace

Paired attack/control in a workspace environment. Intended L2·L3, not verified.

Runtime environment reference Upstream reference, no independent report Pending (0 reports) Quarantined · No download Inspect & setup reset sandbox →
Authorization boundary pairs candidate-authorization-boundaries

Tool authorization and confirmation checks. Intended L2, not verified.

Sandbox test recipes Authored recipes, no independent report Pending (0 reports) Quarantined · No download Inspect & plan auth sandbox →
Indirect injection surface pairs candidate-indirect-injection-surfaces

Indirect injection surfaces. Intended L2·L3, not verified.

Sandbox test recipes Authored recipes, no independent report Pending (0 reports) Quarantined · No download Inspect & setup indirect sandbox →
Chinese content safety baseline jailbreak-baseline

Applies to LLM request-level refusal checks. Effect unmeasured.

Benchmark records Historical audit structure recorded Pending (0 reports) Restricted · No download Inspect & prepare refusal rules →

All 12 entries currently show 0 verified effect reports, meaning independently signed measurement has not been completed in a controlled sandbox. A report count of 0 is not a published effect score. Missing evidence or environment faults must result in an Indeterminate verdict.

Why are no attack success rates or scores published yet?

Scores strongly depend on model revisions, decoding parameters, and sandbox fidelity. Unilateral attack scores without benign utility controls are misleading (refusing everything achieves 100% defense but 0 utility).

What conditions are required before publishing verified effect reports?

Requires multi-turn repeatable runs in authorized sandboxes with state resets; complete evidence chains (text/trace/state); calibrated benign controls; and independent sign-off.