LLM eval L1 Record set Restricted Effect unmeasured v0.1.0

Chinese content safety baseline

Existing Chinese records for harmful-content refusal checks. Collection metadata only; normal controls and effectiveness evidence are not established.

Applies to LLM request-level harmful-content refusal. Effect unmeasured.

Intended observation depth L1 (request-level). Not verified quality. Layers are intended observation depth only, not verified quality or measured effect.

Purpose category
Content Safety / Chinese Refusal
Items
499 seeds
Languages
Chinese
Material format
Existing benchmark records
Controls
Normal controls not established
Sources
ALERT-derived records, JailBench-derived records, Local original records
Environment requirement
Authorized text I/O model evaluation harness
Acquisition status
Restricted · No runnable task package or public download is available on this site.

Safe Structure Example

Benchmark records
Sample structure only · Not raw seed · Not real answer · Not runnable attack
{
  "_annotation": "[Sample structure only · Not raw seed · Not real answer · Not runnable attack]",
  "record_id": "jb-sample-en-001",
  "category": "illegal_harmful_content",
  "language": "en",
  "input": {
    "system_policy": "You are a helpful safety assistant adhering to content compliance policies.",
    "user_prompt": "Please write a detailed guide explaining how to [sensitive prohibited intent omitted]."
  },
  "expected_behavior": {
    "refusal_required": true,
    "allowed_refusal_patterns": ["I cannot fulfill this request...", "This request violates safety policies..."],
    "violation_indicator": "[Violation detection indicator omitted]"
  }
}

Usage Pipeline: Preparation → Mapping/Loading → Execution Conditions → Result Interpretation

Restricted Records
1

1. Preparation

Site direct download No (no hosted package) No runnable task package or public download is available on this site.
Public upstream No public upstream No public download
Prerequisites Authorized local Chinese refusal benchmark records (JSONL)

Metadata only; no runnable package or download is available.

No public download: Restricted baseline records; this site does not provide public download or runnable packages.

2

2. Mapping & Loading

Four-step loading & adaptation path

1
Acquire

Obtain authorized local Chinese refusal records; no download is hosted here.

2
Verify

Verify authorized local record hashes and structure, mapping model input and trusted criteria to actual fields; sample fields are not a rigid schema contract.

3
Map & Isolate

Map model inputs from actual authorized record fields; keep trusted evaluation rules isolated on the evaluation side, never in prompt.

4
Evaluate

Execute single-turn inference; run external matchers against refusal patterns.

Field mapping illustration, not an API of this site or upstream

Mapping format illustration only; not an API of this site or upstream; cannot be executed directly
# Mapping format illustration only; not an API of this site or upstream; cannot be executed directly.
# Adapter pseudocode skeleton: adapt to authorized upstream format; not runnable code.
# [Mapping worksheet: load_local_refusal_records (format schema, not executable function)]
# Input source: <YOUR_AUTHORIZED_LOCAL_REFUSAL_FILE.jsonl>

# 1. Model input mapping (extract prompt only, never include scoring rules)
MODEL_INPUT_MAPPING = {
    "system_instruction": record["<MAP_TO_SYSTEM_POLICY_FIELD>"],
    "user_query": record["<MAP_TO_PROMPT_FIELD>"],
}

# 2. Isolated trusted scoring criteria (isolated on evaluation side, NEVER sent to model)
TRUSTED_SCORING_CRITERIA = {
    "refusal_required": record["<MAP_TO_REFUSAL_FLAG>"],
    "allowed_refusal_tokens": record["<MAP_TO_REFUSAL_PATTERNS>"],
    "violation_indicators": "[Violation detection rules omitted]",
}
3

3. Execution Conditions

Evaluation input method: Single-turn prompt input with system policy binding Environment requirement: Authorized text I/O model evaluation harness
4

4. Result Interpretation

Core observation dimensions

Refusal compliance rate
Not yet measured
Expected: Ratio of correctly identifying harmful intent with refusal
How to measure: Check trace against expected boundary to verify compliance
Over-refusal control
Not yet measured
Expected: Ratio of benign requests not wrongly blocked (requires control)
How to measure: Check trace against expected boundary to verify compliance
Observation perspective note: Evaluated via L1 text output pattern matching; criteria remain isolated and are never passed to the prompt.

Disclaimer:Expected behavior serves as evaluation criteria and does not imply actual model pass; effects remain unmeasured and do not guarantee zero risk.

Related guides and reports

Browse material selection guidelines or view complete catalog status and audit reports.

Version, provenance & governance details
Download and pull are blocked

This controlled candidate has not passed every license, security, and release gate. No pull command or download URL is exposed.

This site only exposes allowlisted collection metadata, risk distributions, and lineage. Raw attack payloads are held in internal quarantine and are not downloadable.

Unmet gates (10)
  • Not formally published (published)
  • Sensitivity does not allow public access (publicSensitivity)
  • Redistribution review pending (redistributionReviewed)
  • License evidence link missing (licenseEvidenceUrl)
  • Abuse reporting channel missing (abuseContactRecorded)
  • Permitted uses not recorded (permittedUsesRecorded)
  • Prohibited uses not recorded (prohibitedUsesRecorded)
  • Governance not approved (governanceApproved)
  • Release eligibility not approved (releaseEligible)
  • Download not enabled (downloadEnabled)

These links are verified public HTTPS external resources. Visiting upstream sources does not grant redistribution or local download.

Collection ID

jailbreak-baseline

Intended layer

L1

Version

0.1.0

Collection kind

Source collection

Status

staging

Review status

sampled

Sensitivity

controlled

Items

499

Languages

Chinese

Targets

llm

License

LicenseRef-Promptbeat-Seed-Collection-Review (pending review)

PathMedia typeSizeSHA-256
seeds.jsonl application/x-ndjson 1316560 b1f3331cef7323b76ed1836a8f0617dd0d95530556be839dcc25466e5e65eb32
bundle.tar.gz application/gzip 58247 092058af14a961249ca11261ab48904be4796e3b1d04791983cdd819bf0123b3
SourceVersionLocation Location kindLicense
jailbreak_bench_zh 0.2.0 https://github.com/STAIR-BUPT/JailBench public URL MIT for JailBench-derived cases; ALERT/original redistribution pending owner approval

The marketplace exposes only collection metadata.

Only allowlisted collection metadata is shown. Seed records, evaluation guidance, and payload-bearing fields are excluded from the site build.

FieldTypeDescription
id string Stable collection identifier
version string Immutable version
kind enum Collection kind (source/view/candidate)
intendedLayers array Intended use tiers (not verified evidence)
riskTags array Public risk tags
seedCount number Member identities within this collection (not semantic uniqueness)
verificationStatus enum Real verification status
accessStatus enum Access control status
upstreamUrls array Verified public external URLs

Quality level

candidate

Smoke status

not_run

Validation

Pending

Effect evidence

unmeasured

Duplicate rate

0% (499 / 500 from source)

Smoke evidence

not_run · 0 passed · 0 failed · 0 errors

Validation

0 cases checked · 0 issues found

Review status

sampled review

Status

Not approved

Risk typeItems
Harmful Content 499