Reference

Reading the evidence

Four ideas make the lab make sense. Back to the cases anytime: Open the evidence lab →

  • The one-sided p-value

    Assume the honest hypothesis is true (a fair coin, the stated drop rate). The p-value is the probability of seeing a result AT LEAST as extreme as the one you got, purely by chance. Small means "this rarely happens if the honest rate held" — surprising. It is not the probability that the hypothesis is false.

  • The binomial tail

    For n independent trials each succeeding with probability p, the count of successes follows a binomial distribution. The upper tail — P(X ≥ k) — is exactly the one-sided p-value for "k or more". The bars in the lab show that distribution; the shaded ones are the tail.

  • Evidence, not proof

    A tiny p-value says the data is unlikely by chance — never WHY. A weighted coin, a lucky streak, or a mistake could all produce it. That is why OddsForge weighs evidence and never accuses: a low p-value raises a question, it does not answer it.

  • The peeking trap

    If you test at α = 5% but keep peeking and stop the moment it "looks significant", your real false-alarm rate balloons — with 10 looks it is roughly 40%, not 5%. Deciding when to stop AFTER seeing the data manufactures significance. Set your test before you look.

Used well, a p-value sharpens a question. Used carelessly, it invents an answer. Try a case →