पाठशाला Pathshala · ग्राहक Grāhak, The customer · Lesson 11 · Build
The Sean Ellis test and where it misleads
One question, one threshold: how would you feel if you could no longer use the product? Run it properly, read the 40 percent correctly and segment before you act on it.
Pathshala, The Founder Library · 11 October 2026 · 7 min read

One question, one threshold. How would you feel if you could no longer use the product? If more than 40% of users say very disappointed, the story goes, you have product-market fit. The question is good. The story is where founders get hurt.
This lesson sets out where the test came from, whom to ask and how many, how to read the 40% honestly, why segmenting comes before acting, where the test misleads, a worked example and a quarterly habit.
The question and the threshold
The survey belongs to [Sean Ellis](/library/thinkers), a growth adviser to early-stage companies. Its core question is how a user would feel if they could no longer use the product, with three answers: very disappointed, somewhat disappointed, not disappointed. Ellis benchmarked the answers across nearly a hundred startups and found that those which struggled to grow almost always had fewer than 40% of users saying very disappointed, while most of those gaining traction had more.
The question works because it asks about loss rather than liking. Satisfaction is cheap to express and polite people express it freely. Imagining the product gone forces a comparison with the alternative, and an honest user who would simply go back to a spreadsheet will say so. It is the [Mom Test](/library/mom-test-for-indian-politeness) logic applied to a product that already exists.
The best-known use is Rahul Vohra’s at Superhuman, written up by First Round Review. He added three follow-up questions: what type of people would most benefit from the product, what is the main benefit you receive, and how can we improve it for you. The first answer sorted the respondents; the second told the team what to protect; the third told them what to fix.
Who to survey, and how many
Ask only people who have used the product enough to know what they would miss. A sensible house rule is users who have reached the core action at least twice in the last fortnight. Surveying sign-ups who never activated measures your onboarding, not your fit; surveying only your fans measures your fans. Send to every user who meets the rule, or to a random sample of them, and record how many you sent and how many answered.
Vohra’s guidance was that around forty responses give a directionally correct result. Treat forty as a floor per segment you intend to read, not for the survey as a whole. At forty responses a 40% score carries a margin of roughly fifteen points either way, which the figure below makes visible. Directional means it can tell 20% from 50%, not 38% from 42%.
Reading 40 percent correctly
Ellis himself has called the threshold a bit arbitrary, as Justin Jackson’s critique of the survey records. It is an empirical line drawn from one sample of companies, mostly software, mostly in the United States, at one period. It is useful the way a blood-pressure reference range is useful: a reading far below it says something is wrong, a reading far above it is reassuring, and a reading near it says measure again and look at other evidence.
For comparison points, Superhuman began at 22% in the summer of 2017. Hiten Shah’s open survey of 731 Slack users in 2015 found 51% would be very disappointed without it, at a time when Slack had around half a million paying users. Neither number is a target for an Indian B2B tool or a consumer app; both show that the score moves with the product and with the people asked.
Product-market fit itself is not a survey result. Marc Andreessen’s definition from 2007 is being in a good market with a product that can satisfy that market, and his description of what it feels like is about customers buying as fast as you can supply, usage growing as fast as you can add capacity and revenue piling up. The survey is a leading indicator of that state, not a substitute for the [retention curve](/library/retention-curves-and-flattening-test) that eventually proves it.
Segment before you act
The blended score is the least useful number the survey produces. Superhuman’s 22% became 33% the moment the team looked only at the kinds of people the very disappointed users resembled, before a line of code changed. Every product has users it fits and users it does not, and averaging them describes nobody.
Vohra’s method is a sequence. Use the first follow-up question to describe the people who love the product, Julie Supan’s high-expectation customer, and keep only respondents who look like them. Read what the very disappointed say is the main benefit and protect it. Then find the somewhat disappointed users for whom that same benefit is the main one, and read what they want improved: they are the users nearest to loving it. Superhuman split its roadmap half and half between deepening what the lovers loved and removing what held back that nearest group, and the score rose to 58% over three quarters.
The segment you act on must still carry enough responses. A segment of twenty-five with eleven very disappointed is 44%, but its margin runs from the mid-twenties to the low sixties. Segment by a characteristic you chose before reading the answers, not by slicing until a cell clears 40%.
Where it misleads
It asks for a prediction. Rob Fitzpatrick’s objection, quoted by Jackson, is that people are poor forecasters of their own feelings. Jackson’s own example is a calendar app he expected to miss and did not. The survey measures anticipated loss, which is close to but not the same as behaviour.

It misses habit and lock-in. A tool people find irritating but cannot easily leave, because the team, the data or the workflow lives there, can score low and retain well. Jackson notes that small teams described switching from Slack as not worth the trouble. For workflow and B2B products, read the score beside retention and expansion, never instead of them.
It flatters early adopters. The first hundred users of anything self-select for enthusiasm. A high score from them predicts little about the next thousand, especially when the next thousand are in a different city, speak a different language or buy through a distributor rather than a link.
It is easy to cut until it agrees. With enough segments one will clear 40% by chance. Decide the segments before the survey goes out.
The 40 percent line is a reference range, not a finish line. The segment that clears it, and the reason it does, are what the survey is for.
A worked example: an invoicing app with six hundred active users
A Bengaluru team runs an invoicing and collections app with about six hundred users who have sent at least two invoices in the past fortnight. They survey all of them and get 150 responses. Blended, 32% say very disappointed. The founders, who had expected forty-something, start planning a pivot.
Before the survey went out they had decided to read two segments: finance teams at companies of ten to fifty people, and freelancers and solo consultants. Sixty finance-team users responded, 48% very disappointed, with a margin of about thirteen points. Ninety freelancers responded, 22% very disappointed. The main benefit named by the very disappointed finance users is automatic GST-compliant invoices with payment reminders. The freelancers mostly want a cheaper plan and a mobile app.
The team does not pivot. It narrows. Marketing stops targeting freelancers. The roadmap splits between deepening the reminder and reconciliation features the finance teams love and fixing the two complaints from somewhat disappointed finance users: bulk invoice upload and Tally export. The next quarter’s survey of finance teams alone reads 53% on 84 responses, and the retention curve for that segment begins to flatten. That curve, not the survey, is what goes in the next board update.
The fit survey, every quarter
Run the survey once a quarter to the same definition of active user, with the same four questions and the same pre-declared segments. Report each segment’s score with its response count and margin beside it, and the blended score last. Read the main-benefit answers of the very disappointed and write the benefit in one sentence; if the sentence changes from quarter to quarter, the product does not yet know what it is for. Choose two improvements from the somewhat disappointed who share that benefit, ship them, and check next quarter whether that group moved.
Keep the score beside retention and revenue on one page. When all three agree, believe them. When the survey says yes and retention says no, believe retention.
The 40% benchmark comes from one researcher’s sample; treat it as a guide to where to look, never as a verdict on its own.
Sources
- First Round Review, How Superhuman Built an Engine to Find Product/Market Fit (Rahul Vohra), 2018: Ellis’s benchmark of nearly 100 startups, Superhuman’s 22%, 33% and 58%, Hiten Shah’s Slack survey
- Justin Jackson, Is the Product-Market Fit survey accurate? (critique, including Ellis on the threshold being “a bit arbitrary”)
- Marc Andreessen, Part 4: The only thing that matters, pmarchive, 25 June 2007
- Lenny Rachitsky, How to know if you’ve got product-market fit, Lenny’s Newsletter, 28 January 2020