Selection and usage guide for catalog collections
How to choose and evaluate collections by material format
The 12 catalog entries span four distinct material formats. Choose based on your evaluation goals and integrate responsibly into authorized test harnesses. This site provides metadata and high-level guidance only, without runnable task packages.
Four material formats and high-level evaluation methods
Different material formats require different harness preparation and integration points. Select a representative collection below:
Existing benchmark records (Content safety & baselines)
Existing records · RestrictedFor output safety policies, harmful content violations, and baseline instruction/permission boundary checks.
Deploy records at the text I/O evaluation harness of authorized models to assess refusal rates; cannot claim over-refusal without established controls.
Original indirect prompt injection surface recipes
Sandbox recipes · Quarantined candidateFor email, calendar, retrieval, web-result, and memory indirect injection surfaces, with 5 no-injection control pairs.
Feed attack recipes and paired no-injection recipes into indirect entry points in a test sandbox; observe whether untrusted external content overrides system instructions.
Original tool authorization & confirmation boundary recipes
Sandbox recipes · Quarantined candidateFor sensitive tool calls, unauthorized actions, and human confirmation checks, with 5 attack/normal control pairs.
Compare execution paths of attack recipes and normal control recipes at authorization checkpoints; verify permission enforcement while ensuring normal requests succeed.
Restricted runtime references (AgentDojo scenario pairs)
Runtime references · Quarantined candidateFor banking, messaging, travel, and workspace task environments, with 60 paired attack and counterfactual task references.
Review upstream AgentDojo environment specs and run paired conditions in an authorized sandbox, preserving attack/counterfactual pairing; upstream provides code, this site does not provide downloads.
Secondary reference: Methodology tiers (L1/L2/L3) and local config lab
demo / unmeasured / untrusted / official=false. Examples are independent of controlled collections. A valid format does not prove safety, effectiveness, trust, or official certification.
Which risk do you need to examine?
L1 / L2 / L3 are evidence tiers, not quality grades. S1 / S2 / S3 risk priority, attack technique, and access permissions are separate dimensions.
Output policy, privacy and over-refusal of legitimate requests.
Evidence needed
Requires boundary-request and normal-control outputs, decision rules and human calibration; not tool or environment evidence.
Tool authorization, argument boundaries, user confirmation and error recovery.
Evidence needed
Requires correlated tool calls/results, arguments, confirmation order and error traces; final answers alone are insufficient.
Observable environment state changes and business consequences.
Evidence needed
Requires independently reset environments, before/after state, a trusted oracle and a normal-task utility gate; tool calls do not establish business success.
Local Teaching Config Lab (Local only · No collection import)
Demonstrates safe local config generation and validation. No network, no models, no keys.
What is missing for a real evaluation?
Currently downloadable collections: 0.
| Requirement | Current boundary |
|---|---|
| Real target and authorized environment | No target connections, tool execution, or environment resets here |
| Real output, tool traces, and state oracles | Teaching configs produce none of this evidence and cannot enter ASR |
| Material license, security review, and release gates | Existing downloads stay closed; examples do not grant publication approval |