← Browse controlled collections

Selection and usage guide for catalog collections

How to choose and evaluate collections by material format

The 12 catalog entries span four distinct material formats. Choose based on your evaluation goals and integrate responsibly into authorized test harnesses. This site provides metadata and high-level guidance only, without runnable task packages.

Four material formats and high-level evaluation methods

Different material formats require different harness preparation and integration points. Select a representative collection below:

Existing benchmark records (Content safety & baselines)

Existing records · Restricted

For output safety policies, harmful content violations, and baseline instruction/permission boundary checks.

How to use:

Deploy records at the text I/O evaluation harness of authorized models to assess refusal rates; cannot claim over-refusal without established controls.

Original indirect prompt injection surface recipes

Sandbox recipes · Quarantined candidate

For email, calendar, retrieval, web-result, and memory indirect injection surfaces, with 5 no-injection control pairs.

How to use:

Feed attack recipes and paired no-injection recipes into indirect entry points in a test sandbox; observe whether untrusted external content overrides system instructions.

Original tool authorization & confirmation boundary recipes

Sandbox recipes · Quarantined candidate

For sensitive tool calls, unauthorized actions, and human confirmation checks, with 5 attack/normal control pairs.

How to use:

Compare execution paths of attack recipes and normal control recipes at authorization checkpoints; verify permission enforcement while ensuring normal requests succeed.

Restricted runtime references (AgentDojo scenario pairs)

Runtime references · Quarantined candidate

For banking, messaging, travel, and workspace task environments, with 60 paired attack and counterfactual task references.

How to use:

Review upstream AgentDojo environment specs and run paired conditions in an authorized sandbox, preserving attack/counterfactual pairing; upstream provides code, this site does not provide downloads.

Secondary reference: Methodology tiers (L1/L2/L3) and local config lab
Not a real model evaluation

demo / unmeasured / untrusted / official=false. Examples are independent of controlled collections. A valid format does not prove safety, effectiveness, trust, or official certification.

Which risk do you need to examine?

L1 / L2 / L3 are evidence tiers, not quality grades. S1 / S2 / S3 risk priority, attack technique, and access permissions are separate dimensions.

Local Teaching Config Lab (Local only · No collection import)

Demonstrates safe local config generation and validation. No network, no models, no keys.

Try a safe local workflow first

1 · Choose an evidence tier

L1 / L2 / L3 are evidence tiers, not quality grades. S1 / S2 / S3 risk priority, attack technique, and access permissions are separate dimensions.

1 · Choose an evidence tier

L1 · Independent synthetic example

Independent synthetic library example: compare a restricted shelf-label request with a public opening-hours request. Teaching only, not a real model evaluation.

Evidence needed:Requires boundary-request and normal-control outputs, decision rules and human calibration; not tool or environment evidence.

2 · Inspect the teaching config

Only teaching mode and tier are recorded: no seed payload, target URL, key, or execution command.

Copy and export run only when you click.

{
  "schemaVersion": 1,
  "kind": "demo",
  "tier": "L1",
  "locale": "en",
  "trust": "untrusted",
  "effect": "unmeasured",
  "official": false
}

3 · Import and check

Only config JSON in this page’s format is accepted (UTF-8, up to 8192 bytes). No model reports, seeds, keys, or arbitrary fields. Errors never echo file names or input content.

Check after pasting. Editing, changing files, or switching tiers invalidates the previous result.

Not checked yet. Use the config above, choose a file, or paste JSON.

4 · Export local feedback

Complete step 3 first. The feedback package contains only the normalized config and selected category: no raw input, filename, identity, or measurements. Nothing is submitted automatically.

Copy and export run only when you click.

Understand tiered reports & real verification gaps →

What is missing for a real evaluation?

Currently downloadable collections: 0.

RequirementCurrent boundary
Real target and authorized environmentNo target connections, tool execution, or environment resets here
Real output, tool traces, and state oraclesTeaching configs produce none of this evidence and cannot enter ASR
Material license, security review, and release gatesExisting downloads stay closed; examples do not grant publication approval