I didn't set out to test the rule. I just wanted somewhere to put requests so they'd stop living in my inbox and my memory at the same time, which is a worse filing system than it sounds. The note started as a flat list. By month two it had become a table with a tally column, because I kept noticing the same idea come back in different words and wanted to know if that was actually happening or if I was pattern matching on nothing.
Where the number five even comes from
It's not a rule anyone can point to a study for. It's the kind of thing that gets repeated in enough startup threads that it starts sounding like research. The appeal is obvious: counting feels objective in a way that "does this seem important" doesn't, and it protects you from building whatever you personally find interesting, which is a real risk for anyone who also writes the code. I liked the rule for exactly that reason before I'd tested it against anything.
What the tally actually looked like after four months
Seventy one messages, mapped onto twenty six distinct ideas. Most ideas got mentioned once and never again. Two ideas crossed five mentions. One was a home screen widget, showing your current verdict without opening the app. Six people asked for it, unprompted, across six weeks, which by the rule made it an easy yes.
It also told me almost nothing about what to actually put on the widget. Three people said "a widget would be cool." Two said they wanted to see their score. One wanted a reminder that a Prove deadline was coming up. Six votes, three different products hiding inside them.
The one that sat at two for six weeks
The other idea, meanwhile, sat at two mentions for a month and a half. Not "a widget would be cool" two. Two full paragraphs, both describing the exact same afternoon: building a Build verdict for one version of a feature, then wanting to change a single answer and see how the verdict moved, without losing the first version to compare against. One person called it duplicating a session. The other called it forking. Same request, arrived nine days apart, written by two people who had clearly never spoken to each other.
Two is below five. I left it in the note and moved on, because the rule said to, and because the widget had six votes sitting right above it looking like the obviously bigger opportunity.
What I nearly missed by waiting
I built the widget first. Three weeks of work, including a small animation on the verdict chip that I was pleased with for about a day. I added basic on-device logging afterward, no account needed for it, just a local count of how often a widget gets glanced at versus tapped into the app. After a month, most of the installed widgets were being glanced at once and never again. People wanted to see the score exist, once, and then the widget's job was done. It was a fine feature. It was not close to the thing six requests had made it look like.
While I was building it, a third message about duplicating sessions arrived, except it wasn't really a new request. It was someone mentioning, almost in passing, that they'd been taking a screenshot of their answers before changing anything, so they could re-enter them into a second session by hand if the new version came out worse. They'd been doing this for weeks. They hadn't mentioned it again after the first time, not because they'd stopped wanting it, but because they'd quietly built their own bad version of the feature and stopped expecting me to.
That's the part the count could never have shown me. A request that goes quiet doesn't mean the need went away. Sometimes it means the person worked around you and gave up on asking.
What I count instead now
Not votes. I still write every request down, but I weigh a detailed description of one real afternoon over three people saying a feature sounds nice, whatever the running total says. A specific story is a signal on its own. A vague one needs three more before I trust it, and even then I go back and ask the people who sent them what they'd actually do with it, because "score as a widget" and "a Prove deadline reminder" are not the same six months of engineering.
I shipped the duplicate-session feature five weeks after the widget, once I'd gone back and asked the screenshot person exactly what they wanted to compare. It took four days to build, a fraction of the widget's three weeks, and it's now one of the two or three actions people reach for most inside the app, going by the same local logging. The count said this one mattered less. The count was wrong, and it took a support message I nearly filed under "already covered this" to notice.
Where this leaves the five-person rule
I don't think it's useless. A number that keeps climbing on its own, without you doing anything to encourage it, is still worth noticing. What it can't do is tell you whether the people behind that number are describing the same problem or six different ones wearing the same three words, and it can't hear the person who already gave up asking. Read what people actually wrote before you count how many of them wrote it. The count is the last thing to check, not the first.
Appray walks you through the same territory in eight rounds of questions, then puts every feature you are considering on trial against what you actually said. The result is a Build, Kill or Prove verdict for each one, plus interview questions written from your own answers rather than a template.
It runs entirely on your iPhone or iPad. No account, no sign in, nothing leaves your device, and no AI is involved anywhere in it.