Measured, not promised. Every prompt the system uses, every adversarial scenario we ran, and the raw run data — verbatim. So you can check the work, not take our word. Nothing here is edited for the story; it's the receipts.
The exact instructions the system uses — the "how we asked," shown, not described.
The five age-tiered Kids-Mode safety prompts as sent to the model (ages 3–17), the output-gate second-pass verifier, the passive safety-scan classifier, and the alerting cascade's label + validator prompts.
The 1,000-attack proof (0 harmful reached the child, on both models). Here's exactly what we threw at it and what came back.
The exact adversarial scenarios, turn by turn — anthropomorphism baiting, the "grandma" exploit, DAN, crisis-in-jailbreak, encoding tricks, and the rest, across all five age bands.
Every turn of every run, with the model's reply and the safety gate's verdict. 841 safe · 159 gate-caught · 0 breaks.
The same, on the capable tier. 850 safe · 150 gate-caught · 0 breaks.
The per-conversation roll-ups — outcomes, break counts, resisted vs gate-caught — one file per model.
The honest, in-progress detection work (how reliably we alert you). Protocols locked before each run; raw numbers after.
Each protocol's bar and prediction, written down before the run — so a result can't be graded against a moving target.
The barrage roll-ups per model, each measure at both levels — classifier alone vs. classifier + the deterministic floor.
The hashed calibration set the detection numbers are measured against — frozen so any run is reproducible.