Chinese content safety baseline
Existing Chinese records for harmful-content refusal checks. Collection metadata only; normal controls and effectiveness evidence are not established.
Applies to LLM request-level harmful-content refusal. Effect unmeasured.
Intended observation depth L1 (request-level). Not verified quality. Layers are intended observation depth only, not verified quality or measured effect.
Collection properties
- Purpose category
- Content Safety / Chinese Refusal
- Items
- 499 seeds
- Languages
- Chinese
- Material format
- Existing benchmark records
- Controls
- Normal controls not established
- Sources
- ALERT-derived records, JailBench-derived records, Local original records
- Environment requirement
- Authorized text I/O model evaluation harness
- Acquisition status
- Restricted · No runnable task package or public download is available on this site.
Safe Structure Example
Benchmark records{
"_annotation": "[Sample structure only · Not raw seed · Not real answer · Not runnable attack]",
"record_id": "jb-sample-en-001",
"category": "illegal_harmful_content",
"language": "en",
"input": {
"system_policy": "You are a helpful safety assistant adhering to content compliance policies.",
"user_prompt": "Please write a detailed guide explaining how to [sensitive prohibited intent omitted]."
},
"expected_behavior": {
"refusal_required": true,
"allowed_refusal_patterns": ["I cannot fulfill this request...", "This request violates safety policies..."],
"violation_indicator": "[Violation detection indicator omitted]"
}
} Usage Pipeline: Preparation → Mapping/Loading → Execution Conditions → Result Interpretation
Restricted Records1. Preparation
Metadata only; no runnable package or download is available.
No public download: Restricted baseline records; this site does not provide public download or runnable packages.
2. Mapping & Loading
Four-step loading & adaptation path
Acquire
Obtain authorized local Chinese refusal records; no download is hosted here.
Verify
Verify authorized local record hashes and structure, mapping model input and trusted criteria to actual fields; sample fields are not a rigid schema contract.
Map & Isolate
Map model inputs from actual authorized record fields; keep trusted evaluation rules isolated on the evaluation side, never in prompt.
Evaluate
Execute single-turn inference; run external matchers against refusal patterns.
Field mapping illustration, not an API of this site or upstream
Mapping format illustration only; not an API of this site or upstream; cannot be executed directly# Mapping format illustration only; not an API of this site or upstream; cannot be executed directly.
# Adapter pseudocode skeleton: adapt to authorized upstream format; not runnable code.
# [Mapping worksheet: load_local_refusal_records (format schema, not executable function)]
# Input source: <YOUR_AUTHORIZED_LOCAL_REFUSAL_FILE.jsonl>
# 1. Model input mapping (extract prompt only, never include scoring rules)
MODEL_INPUT_MAPPING = {
"system_instruction": record["<MAP_TO_SYSTEM_POLICY_FIELD>"],
"user_query": record["<MAP_TO_PROMPT_FIELD>"],
}
# 2. Isolated trusted scoring criteria (isolated on evaluation side, NEVER sent to model)
TRUSTED_SCORING_CRITERIA = {
"refusal_required": record["<MAP_TO_REFUSAL_FLAG>"],
"allowed_refusal_tokens": record["<MAP_TO_REFUSAL_PATTERNS>"],
"violation_indicators": "[Violation detection rules omitted]",
} 3. Execution Conditions
4. Result Interpretation
Core observation dimensions
Refusal compliance rate
Not yet measuredOver-refusal control
Not yet measuredDisclaimer:Expected behavior serves as evaluation criteria and does not imply actual model pass; effects remain unmeasured and do not guarantee zero risk.
Version, provenance & governance details
This controlled candidate has not passed every license, security, and release gate. No pull command or download URL is exposed.
This site only exposes allowlisted collection metadata, risk distributions, and lineage. Raw attack payloads are held in internal quarantine and are not downloadable.
Unmet gates (10)
- Not formally published
(published) - Sensitivity does not allow public access
(publicSensitivity) - Redistribution review pending
(redistributionReviewed) - License evidence link missing
(licenseEvidenceUrl) - Abuse reporting channel missing
(abuseContactRecorded) - Permitted uses not recorded
(permittedUsesRecorded) - Prohibited uses not recorded
(prohibitedUsesRecorded) - Governance not approved
(governanceApproved) - Release eligibility not approved
(releaseEligible) - Download not enabled
(downloadEnabled)
These links are verified public HTTPS external resources. Visiting upstream sources does not grant redistribution or local download.
Technical Specifications
Collection ID
jailbreak-baseline
Intended layer
L1
Version
0.1.0
Collection kind
Source collection
Status
staging
Review status
sampled
Sensitivity
controlled
Items
499
Languages
Chinese
Targets
llm
License
LicenseRef-Promptbeat-Seed-Collection-Review (pending review)
Files
| Path | Media type | Size | SHA-256 |
|---|---|---|---|
| seeds.jsonl | application/x-ndjson | 1316560 | b1f3331cef7323b76ed1836a8f0617dd0d95530556be839dcc25466e5e65eb32 |
| bundle.tar.gz | application/gzip | 58247 | 092058af14a961249ca11261ab48904be4796e3b1d04791983cdd819bf0123b3 |
Provenance
| Source | Version | Location | Location kind | License |
|---|---|---|---|---|
| jailbreak_bench_zh | 0.2.0 | https://github.com/STAIR-BUPT/JailBench | public URL | MIT for JailBench-derived cases; ALERT/original redistribution pending owner approval |
Public metadata schema
The marketplace exposes only collection metadata.
Only allowlisted collection metadata is shown. Seed records, evaluation guidance, and payload-bearing fields are excluded from the site build.
| Field | Type | Description |
|---|---|---|
| id | string | Stable collection identifier |
| version | string | Immutable version |
| kind | enum | Collection kind (source/view/candidate) |
| intendedLayers | array | Intended use tiers (not verified evidence) |
| riskTags | array | Public risk tags |
| seedCount | number | Member identities within this collection (not semantic uniqueness) |
| verificationStatus | enum | Real verification status |
| accessStatus | enum | Access control status |
| upstreamUrls | array | Verified public external URLs |
Quality
Quality level
candidate
Smoke status
not_run
Validation
Pending
Effect evidence
unmeasured
Duplicate rate
0% (499 / 500 from source)
Smoke evidence
not_run · 0 passed · 0 failed · 0 errors
Validation
0 cases checked · 0 issues found
Governance & Compliance
Review status
sampled review
Status
Not approved
Cisco risk types
| Risk type | Items |
|---|---|
| Harmful Content | 499 |