HELIX AI LABS · RESEARCH

Research

We prove two different things, and we don't mix them up. One — your child is protected: we attacked Kids Mode 1,000 times — on the smallest model a family might run and the capable one — and never got a harmful response through to a child. Two — how reliably we alert you, the parent: honest, still‑improving work. Measured, not promised — highs and lows alike.

Safety proof✓ Completed · proven

We attacked Kids Mode 1,000 times. Nothing harmful reached the child.

The plain‑language proof that your child won't be shown harmful content — 10 jailbreak techniques, every age band, run 20 times against the real shipping safety system, on the smallest model we support and the capable tier. With the honest fine print, too.

1,000 turns per model0 harmful outputs0 reached the child
Read the proof →
The research behind it — investigations in flight

The living logs of how we work — mostly the alerting side we're still improving. Published while in flight, wrong turns and all.

Adventure logLiving · updated

Improving the alerting — how well we tell a parent

The honest, in‑progress work on the other layer: how reliably our on‑device system flags a concern to the parent. A weak model went 20% → 95.7% on the detection task by changing the prompt — wrong turns, a nearly‑published false conclusion, and the honest ceilings included.

Kids‑Safety · Alerting2026‑07‑14 · updatedRead the log →
Report cardDraft · pending run

Report Card 01 — Kids‑Safety, small model vs capable model

The parent‑facing verdict: which on‑device model is fit for Kids Mode, in plain language. The cascade experiment has now resolved (see the log); the plain‑language conclusion is being written.

Kids‑SafetyIn preparationNot yet published