saved
When random is not actually random enough
Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.
gist
Austin Seipp argues that random_u64() modulo N quietly breaks uniform choice, then prefers APIs that spell out a distribution—especially relative integer weights via random_choice(). He ties the bug to thinking about sample values instead of random variables, and says East River’s Antithesis fuzz harnesses dropped residual random_u64() calls for weighted discrete choice after a review last week.
ideas
- Modulo bias is the classic trap. Clamping a uniform u64 with % n does not keep equal odds when the range is not a divisor of the input space; a 0–9 draw mapped to three objects yields 40/30/30, not 33/33/33.
- Prefer random_between and weighted choice. Correct inclusive ranges need more than raw u64; relative integer weights (Python-style) avoid float-sum fragility and make lopsided odds obvious.
- Spell the distribution. random_choice with explicit weights is a better default than low-level uniform draws because you usually want objects and non-uniform steering.
- Wrong mental domain. Operators on draws are not operators on expectations; nonlinear maps do not preserve expected value under the map.
- Fuzzer practice at ERSC. Antithesis workloads weight blob sizes and path conditions; the team moratoriumed new random_u64() after replacing remaining uses with weighted choice.
quotes
“it doesn’t preserve the underlying uniform distribution of the given random_u64() function.”
“This means choice 1 is picked 40% of the time rather than the intended 33%”
“non-linear operators do not respect the expected value of random variables”
“we’ve put a moratorium on introducing new calls to it until further notice.”