HELIX RESEARCH · DATA & RECEIPTS ← All research

Data & receipts

Measured, not promised. Every prompt the system uses, every adversarial scenario we ran, and the raw run data — verbatim. So you can check the work, not take our word. Nothing here is edited for the story; it's the receipts.

Every prompt

The exact instructions the system uses — the "how we asked," shown, not described.

On publishing the safety prompt. We publish the full production safety prompt on purpose. Good safety never relies on the rules being secret — the output gate (which inspects every reply before a child sees it, and fails safe) is the guarantee, not the prompt. Showing the actual rules is how a parent earns the right to trust them.

The safety red-team — inputs & every outcome

The 1,000-attack proof (0 harmful reached the child, on both models). Here's exactly what we threw at it and what came back.

The alerting research — pre-registrations & raw numbers

The honest, in-progress detection work (how reliably we alert you). Protocols locked before each run; raw numbers after.

md

Pre-registrations (locked before running)

Each protocol's bar and prediction, written down before the run — so a result can't be graded against a moving target.

json

Detection numbers — raw-LLM vs production

The barrage roll-ups per model, each measure at both levels — classifier alone vs. classifier + the deterministic floor.

json

The frozen scenario battery

The hashed calibration set the detection numbers are measured against — frozen so any run is reproducible.

open →