Attack Defense
Whether adversarial inputs or malicious payloads are identified and blocked.
Status & Evidence Center
Status and evidence center tracking verification, measurement readiness, and release status across all 12 catalog collections. Not a benchmark leaderboard; no single score or download packages.
Catalog entries
12
5 roots + 7 thematic views
Root collections with historical records
2
Only 2 root baselines have historical records
Evaluated effect reports
0
Not yet evaluated (0 verified reports published)
Downloadable collections
0
All payloads quarantined/restricted
Safety evaluation spans defense, utility, over-refusal, and indeterminate error handling; L1/L2/L3 are observation perspectives and evidence depth, not a quality ladder.
Whether adversarial inputs or malicious payloads are identified and blocked.
Whether normal tasks succeed on clean controls, preserving legitimate functionality.
Whether benign requests are misclassified as attacks, measuring false alarms.
Differentiated by L1 text proxy, L2 tool trace, and L3 sandbox state oracles.
If execution encounters environment anomalies, missing data, or incomplete evidence chains, it must be marked Indeterminate / unjudgeable, never counted as a safety pass.
Only 2 legacy baselines hold historical structure records; the other 10 remain quarantined or restricted without fabricated scores.
| Collection | Material format | Structural evidence | Verified effect measurement | Access control | Recommended next step |
|---|---|---|---|---|---|
Agent safety baseline agent-safety-baseline Agent instruction, permission and data checks. Request-level intent; tools pending. Unmeasured. | Mixed benchmark records | Historical audit structure recorded | Pending (0 reports) | Restricted · No download | Inspect & select thematic subsets → |
Data exfiltration agent-safety-baseline--data-exfiltration Agent data-exfiltration checks. Request-level intent; tools pending. Effect unmeasured. | Labeled benchmark records | No independent audit or evaluation report | Pending (0 reports) | Restricted · No download | Inspect & configure egress sandbox → |
Instruction boundaries agent-safety-baseline--instruction-boundaries Agent trusted-instruction checks. Request-level intent; tools pending. Effect unmeasured. | Labeled benchmark records | No independent audit or evaluation report | Pending (0 reports) | Restricted · No download | Inspect & test layered prompts → |
Unauthorized actions agent-safety-baseline--unauthorized-actions Agent tool authorization checks. Request-level intent; tools pending. Effect unmeasured. | Labeled benchmark records | No independent audit or evaluation report | Pending (0 reports) | Restricted · No download | Inspect & verify confirmation gate → |
AgentDojo paired candidates candidate-agentdojo-paired Paired attack/control in a resettable world. Intended L2·L3, not verified. | Runtime environment reference | Upstream reference, no independent report | Pending (0 reports) | Quarantined · No download | Inspect & review upstream docs → |
Banking tasks candidate-agentdojo-paired--banking Paired attack/control in a banking environment. Intended L2·L3, not verified. | Runtime environment reference | Upstream reference, no independent report | Pending (0 reports) | Quarantined · No download | Inspect & setup reset sandbox → |
Messaging tasks candidate-agentdojo-paired--slack Paired attack/control in a messaging environment. Intended L2·L3, not verified. | Runtime environment reference | Upstream reference, no independent report | Pending (0 reports) | Quarantined · No download | Inspect & setup reset sandbox → |
Travel tasks candidate-agentdojo-paired--travel Paired attack/control in a travel environment. Intended L2·L3, not verified. | Runtime environment reference | Upstream reference, no independent report | Pending (0 reports) | Quarantined · No download | Inspect & setup reset sandbox → |
Workspace tasks candidate-agentdojo-paired--workspace Paired attack/control in a workspace environment. Intended L2·L3, not verified. | Runtime environment reference | Upstream reference, no independent report | Pending (0 reports) | Quarantined · No download | Inspect & setup reset sandbox → |
Authorization boundary pairs candidate-authorization-boundaries Tool authorization and confirmation checks. Intended L2, not verified. | Sandbox test recipes | Authored recipes, no independent report | Pending (0 reports) | Quarantined · No download | Inspect & plan auth sandbox → |
Indirect injection surface pairs candidate-indirect-injection-surfaces Indirect injection surfaces. Intended L2·L3, not verified. | Sandbox test recipes | Authored recipes, no independent report | Pending (0 reports) | Quarantined · No download | Inspect & setup indirect sandbox → |
Chinese content safety baseline jailbreak-baseline Applies to LLM request-level refusal checks. Effect unmeasured. | Benchmark records | Historical audit structure recorded | Pending (0 reports) | Restricted · No download | Inspect & prepare refusal rules → |
agent-safety-baseline Mixed benchmark records Agent instruction, permission and data checks. Request-level intent; tools pending. Unmeasured.
agent-safety-baseline--data-exfiltration Labeled benchmark records Agent data-exfiltration checks. Request-level intent; tools pending. Effect unmeasured.
agent-safety-baseline--instruction-boundaries Labeled benchmark records Agent trusted-instruction checks. Request-level intent; tools pending. Effect unmeasured.
agent-safety-baseline--unauthorized-actions Labeled benchmark records Agent tool authorization checks. Request-level intent; tools pending. Effect unmeasured.
candidate-agentdojo-paired Runtime environment reference Paired attack/control in a resettable world. Intended L2·L3, not verified.
candidate-agentdojo-paired--banking Runtime environment reference Paired attack/control in a banking environment. Intended L2·L3, not verified.
candidate-agentdojo-paired--slack Runtime environment reference Paired attack/control in a messaging environment. Intended L2·L3, not verified.
candidate-agentdojo-paired--travel Runtime environment reference Paired attack/control in a travel environment. Intended L2·L3, not verified.
candidate-agentdojo-paired--workspace Runtime environment reference Paired attack/control in a workspace environment. Intended L2·L3, not verified.
candidate-authorization-boundaries Sandbox test recipes Tool authorization and confirmation checks. Intended L2, not verified.
candidate-indirect-injection-surfaces Sandbox test recipes Indirect injection surfaces. Intended L2·L3, not verified.
jailbreak-baseline Benchmark records Applies to LLM request-level refusal checks. Effect unmeasured.
All 12 entries currently show 0 verified effect reports, meaning independently signed measurement has not been completed in a controlled sandbox. A report count of 0 is not a published effect score. Missing evidence or environment faults must result in an Indeterminate verdict.
Scores strongly depend on model revisions, decoding parameters, and sandbox fidelity. Unilateral attack scores without benign utility controls are misleading (refusing everything achieves 100% defense but 0 utility).
Requires multi-turn repeatable runs in authorized sandboxes with state resets; complete evidence chains (text/trace/state); calibrated benign controls; and independent sign-off.