About FilterForge
FilterForge is the trust-and-safety rung of our creation-engine line. Where LobbyQuest lets a younger learner design a room's rules, FilterForge lets an older one build and honestly measure the classifier behind such a room — and confront the fact that safety is a trade-off, never a solved problem.
Synthetic and non-graphic by design
Every message in the corpus is invented and abstracted. Harmful entries are short, neutral descriptions of a pattern ("asks a younger player to keep chats secret"), never a script or anything explicit. The lab teaches classification and measurement — not the content of harm.
Honest by construction
Precision and recall are the exact arithmetic of the confusion matrix — no thumb on the scale. The impossibility of a perfect threshold isn't a design choice; it falls out of the deliberate overlap between harmful and safe risk scores, exactly as it does in a real corpus.
What it's good for (and what we don't claim)
FilterForge builds intuition for classification metrics and the ethics of the two error costs. It is not a moderation product, does not model any real classifier's true performance, and makes no claim about keeping a real child safe — it's a lab for understanding how safety systems are measured and where they fail.
Safe by design
- Deterministic & on-device. Runs entirely in your browser — no account, no sign-in, nothing sent anywhere.
- Text-forward. Built for the 15–18 band: real metrics, no cartoon mascots.
- No dark patterns. No timer, no streak, no score, nothing to buy.