The living logs of how we work — mostly the alerting side we're still improving. Published while in flight, wrong turns and all.
The honest, in‑progress work on the other layer: how reliably our on‑device system flags a concern to the parent. A weak model went 20% → 95.7% on the detection task by changing the prompt — wrong turns, a nearly‑published false conclusion, and the honest ceilings included.
The parent‑facing verdict: which on‑device model is fit for Kids Mode, in plain language. The cascade experiment has now resolved (see the log); the plain‑language conclusion is being written.