Product

The MVP Canon, Checked Against the Record: Buffer, Zappos, Dropbox, ChatGPT

Every product talk retells the same three MVP stories — Dropbox's video, Zappos's shoe photos, Buffer's landing page — usually with numbers nobody in the room has ever checked. So I checked. Buffer holds up on contemporaneous evidence; Zappos is real but retroactively labeled; Dropbox's famous waitlist figure is founder-self-reported and attached to the wrong video half the time; ChatGPT's 100M is an analyst estimate wearing a fact's clothing. The discipline survives the embellishment — but a PM quoting unverified numbers in a strategy deck is committing the exact sin the MVP exists to prevent.

The MVP Canon, Checked Against the Record: Buffer, Zappos, Dropbox, ChatGPT

I have sat through the Dropbox video slide maybe a dozen times. Different companies, different decks, same beats: Drew Houston couldn’t build the product yet, so he made a screencast, and overnight the waitlist exploded. Then comes Zappos and the shoe store that didn’t have shoes, then Buffer and the pricing page that didn’t have a product behind it. The stories are told as settled fact, with numbers delivered in the confident tone people reserve for things they’ve never checked.

So at some point I checked. Not because I suspected the stories were fake — they mostly aren’t — but because I’d started noticing that the numbers drifted between tellings, and drift is a symptom. What I found is more interesting than a debunk: the canonical MVP stories sit on wildly different tiers of evidence, from contemporaneous founder blog posts you can still read today down to analyst estimates that fossilized into “facts” through repetition. Grading them by evidence quality turns out to teach the MVP lesson better than the stories themselves do. Because the whole point of a minimum viable product — the entire discipline — is refusing to treat a plausible story as evidence. It would be a strange tribute to that discipline to teach it with numbers nobody verified.

Here’s the canon, graded.

Buffer: the one that actually holds up

Start with the least famous story, because it’s the best documented. In 2010, Joel Gascoigne wanted to know whether anyone would pay for a tool that scheduled tweets. Before building it, he put up a landing page describing the product, watched whether people clicked through, then added a pricing page to see whether they’d flinch at being asked for money. Only then did he build. By his own account, the distance from idea to first paying customer was about seven weeks.

What earns Buffer the top grade isn’t that the story is more dramatic — it’s less dramatic than the others, which is rather the point. It’s that Gascoigne wrote it up contemporaneously, on Buffer’s own blog, while it was happening and shortly after, with the actual sequence of steps laid out. There’s no decade of retelling between the events and the record. No conference-stage compression. You can read the founder’s account from the era and reconstruct exactly what was tested, in what order, and what the signal was: first “do people want this at all,” then “will they pay,” each answered with the smallest artifact that could answer it.

That two-step is the thing worth stealing. Gascoigne didn’t test everything at once; he sequenced the assumptions and tested the riskiest one first with the cheapest instrument available — the same move I described in the risk-reduction vocabulary post as the Riskiest Assumption Test, here executed with two web pages and a signup form. Verdict: real, contemporaneously documented, and the safest of these stories to put in a deck.

Zappos: a real story wearing a borrowed label

The Zappos story is also real, and its founder has confirmed the substance: in 1999, Nick Swinmurn tested whether people would buy shoes online by photographing inventory in local shoe stores, listing it on a website, and — when an order came in — buying the pair at retail and shipping it himself. No warehouse, no inventory, no logistics. Just the one question that mattered, “will strangers buy shoes they can’t try on,” answered with the minimum machinery required to ask it.

The asterisk is smaller here, but it’s instructive: nobody at Zappos in 1999 was building an “MVP,” because the term and the framework it belongs to didn’t exist yet in their current form. The label was applied retroactively, years later, when the lean-startup movement went looking for ancestors and found this one waiting. That doesn’t diminish what Swinmurn did — if anything it strengthens the underlying claim, since he arrived at the technique from first principles rather than from a book. But it does mean that when the story is told as “Zappos followed the MVP playbook,” the causality is backwards. The playbook followed Zappos. Retrospective framing is a mild distortion as distortions go, but it’s worth noticing, because it’s the same move that turns every messy, contingent success into a tidy validation of whatever framework the speaker is selling. Verdict: real and founder-confirmed; the story predates its own moral.

Dropbox: two videos, one number, and a decade of conflation

Now the headliner, and the place where the evidence gets genuinely fuzzy. The core of the Dropbox story is true: Houston made screencast videos demonstrating file syncing before the product was ready for the public, and the videos generated enormous waitlist interest. But the version that circulates on conference stages has quietly merged two different videos from two different moments into one clean anecdote — an early demo aimed at technical audiences, and a later, 2008 video seeded to the Digg community, salted with in-jokes for that crowd. The famous number — the waitlist jumping from around 5,000 to around 75,000 essentially overnight — attaches to the Digg video, by Houston’s account.

That phrase is doing load-bearing work, so let me lean on it: by Houston’s account. The 75,000 figure is founder-self-reported. I don’t say that to insinuate it’s false — Houston has no obvious motive to inflate a number from before Dropbox was famous, and the broad shape of the event is corroborated by the fact that Dropbox, you know, exists. But there’s a meaningful evidentiary difference between “a founder’s recollection, told in talks and interviews years later, with the two videos routinely conflated even by careful retellers” and “a contemporaneous record you can audit.” Buffer gives you the second. Dropbox gives you the first, plus a date-fuzziness problem that most retellings don’t survive: ask the next person who cites the story which video produced the number and watch the confidence drain out of the room.

None of which touches the actual lesson, which is why the story endures. Houston’s riskiest assumption wasn’t technical — he knew he could build sync. It was whether the value would be legible: whether people who’d never articulated the file-syncing problem would recognize the solution on sight. A video tests exactly that, and nothing else, for a rounding error of the cost of building distribution-ready software. That’s the discipline. It survives even if the waitlist number turns out to be soft. Verdict: real, but self-reported and date-fuzzy; quote the mechanism, hedge the number.

ChatGPT: a research preview doing MVP work at planetary scale

The fresh entry in the canon, and the one where I can watch the fossilization happening in real time. When OpenAI released ChatGPT in November 2022, it didn’t call it a product launch. It called it a “research preview” — a label doing exactly the work “MVP” does, at a scale no landing-page test ever contemplated: we are shipping the smallest presentable version of this to find out what happens when real people touch it, and the label is our permission structure for it being rough. Reported adoption was one million users within five days, and by January 2023 the figure everyone quotes — 100 million monthly active users, fastest-growing consumer application in history — was everywhere.

Here’s what almost nobody who quotes that 100 million says: it’s an estimate, produced by UBS analysts working from Similarweb traffic data. Not an OpenAI disclosure. Not an audited figure. An analyst’s inference from third-party web-traffic panels, which is a perfectly respectable thing for an analyst to produce and a very different thing from a company-reported metric — and the caveat was shed within roughly one news cycle of the number’s birth. I’ve seen “100M MAU (Jan 2023)” sitting unhedged in strategy decks at companies whose own analytics teams would never let an internal number of that provenance through review.

The ChatGPT case matters for the canon because it shows the pattern isn’t confined to founder nostalgia. The Dropbox number softened over fifteen years of retelling; the ChatGPT number arrived pre-softened, an estimate laundered into a fact in weeks. If anything, the modern information environment fossilizes numbers faster. Verdict: the release and the five-day figure are the story; the 100M is an analyst estimate and should be labeled as one, every single time.

What survives the audit

Line the four up and the grades are: Buffer, contemporaneous and auditable. Zappos, real but retroactively framed. Dropbox, real but self-reported and chronologically smeared. ChatGPT, real event, estimate-grade headline number. Not one of these stories is fake. Every one of them is softer than the version you last heard on a stage.

And here’s the thing I find genuinely satisfying: the lesson survives the audit completely intact. Strip away every embellished number and the mechanism underneath all four stories is untouched — identify the riskiest assumption in your plan, then test it with the smallest artifact that can produce a real signal. Two web pages. A camera in a shoe store. A screencast. A chat window with a disclaimer on it. The discipline these stories teach doesn’t depend on whether the waitlist was 75,000 or 60,000, which is precisely why it’s worth teaching. It’s the same instinct that good discovery work runs on: behavior over stated intent, evidence over narrative, the cheapest test that can actually falsify what you believe.

But that’s also why the sloppy retelling grates on me more here than it would anywhere else. The MVP exists for exactly one reason: to stop you from mistaking a compelling story for evidence. A plausible narrative — “customers will obviously want this” — is the thing the landing page, the fake storefront, and the demo video were invented to interrogate. So a PM who puts “Dropbox went from 5K to 75K overnight” or “ChatGPT: 100M MAU” in a strategy deck, unhedged and unverified, is committing the precise sin the slide is supposedly warning against. They’re accepting a story because it’s well-shaped and often repeated. The audit isn’t pedantry. It’s the method, applied to its own mythology — and if the method can’t survive being applied to itself, it was never a method, just a genre of anecdote.

Put it to work

  1. Grade your next deck’s numbers by provenance before anyone else sees it. For every statistic, write one of four labels in the margin: primary record, self-reported, third-party estimate, or unsourced. Anything in the last two categories gets a hedge in the text (“analyst estimate,” “by the founder’s account”) or gets cut. If the slide’s argument collapses without the unsourced number, the argument was the problem.
  2. Retell one canonical story to your team with its verification verdict attached. Next time someone reaches for Dropbox or ChatGPT in a planning discussion, add the one-sentence asterisk — “self-reported, and the number belongs to the second video” — and watch how it sharpens the conversation. Teams that practice hedging borrowed numbers get noticeably more honest about their own.
  3. Run the Buffer sequence, not the Buffer slide, on your current riskiest assumption. The transferable part isn’t “landing pages work”; it’s the ordering — demand first, willingness to pay second, product third, each tested with the cheapest artifact that produces behavior rather than opinion. Write down your plan’s riskiest assumption and the smallest thing that could falsify it this week, then compare that to what’s actually on the sprint board.

Further reading

  • Joel Gascoigne, “Idea to Paying Customers in 7 Weeks: How We Did It” — the contemporaneous Buffer write-up, and the gold standard for what a documented MVP story looks like. Read it before you cite any of the others.
  • Drew Houston, “On Starting and Scaling Dropbox” (Y Combinator Startup Library) — his own telling of the demo-video story, worth reading as the primary account it is, noting where the two videos and the waitlist figure attach, and where the record depends entirely on his recollection.
  • Eric Ries, The Lean Startup — the book that canonized the Dropbox and Zappos stories and retroactively supplied the “MVP” label the Zappos experiment now wears; useful both for the framework and as an artifact of how the canon got assembled.
  • The UBS analysis of ChatGPT’s early growth, built on Similarweb traffic data — the actual origin of the “100 million MAU” figure. Cite it as what it is: a well-constructed estimate, not a disclosure.
About the author

Prakash Poudel Sharma

Engineering Manager · Product Owner · Varicon

Engineering Manager at Varicon, leading the Onboarding squad as Product Owner. Eleven years of building software — first as a programmer, then as a founder, now sharpening the product craft from the inside of a focused team.

Keep reading

More on this

Join the conversation0 comments

What did you take away?

Thoughts, pushback, or a story of your own? Drop a reply below — I read every one.

Comments are powered by Disqus. By posting you agree to theirterms.

0:000:00