About FilterForge

FilterForge is the trust-and-safety rung of our creation-engine line. Where LobbyQuest lets a younger learner design a room's rules, FilterForge lets an older one build and honestly measure the classifier behind such a room — and confront the fact that safety is a trade-off, never a solved problem.

Synthetic and non-graphic by design

Every message in the corpus is invented and abstracted. Harmful entries are short, neutral descriptions of a pattern ("asks a younger player to keep chats secret"), never a script or anything explicit. The lab teaches classification and measurement — not the content of harm.

Honest by construction

Precision and recall are the exact arithmetic of the confusion matrix — no thumb on the scale. The impossibility of a perfect threshold isn't a design choice; it falls out of the deliberate overlap between harmful and safe risk scores, exactly as it does in a real corpus.

What it's good for (and what we don't claim)

FilterForge builds intuition for classification metrics and the ethics of the two error costs. It is not a moderation product, does not model any real classifier's true performance, and makes no claim about keeping a real child safe — it's a lab for understanding how safety systems are measured and where they fail.

Safe by design

  • Deterministic & on-device. Runs entirely in your browser — no account, no sign-in, nothing sent anywhere.
  • Text-forward. Built for the 15–18 band: real metrics, no cartoon mascots.
  • No dark patterns. No timer, no streak, no score, nothing to buy.

← Tune the filter · Background