Healthcare.gov: The Rescue That Invented a Playbook
On October 1, 2013, Healthcare.gov launched after roughly three years of work across more than fifty contractors — and collapsed on contact with its users. Reportedly six people enrolled on day one. Two months later it worked for the vast majority of visitors, and the fix wasn't a relaunch or a new contractor. It was a small team adding the things the project never had: monitoring, twice-daily contact with reality, one prioritized list, and small releases shipped continuously. The interesting question isn't why it failed. It's why the rescue worked — and how a two-month emergency became codified government practice.

On October 1, 2013, the United States government launched the most anticipated website in its history, and by the end of the day the most widely reported enrollment figure was six. Not six thousand. Six people, reportedly, successfully enrolled through Healthcare.gov on day one — out of roughly 250,000 concurrent users hitting a system planned for far fewer. Behind that number sat roughly three years of work, more than fifty-five contractors with CGI Federal in the lead, and reported spending in the hundreds of millions of dollars by launch. And behind that sat a failure pattern this series has already dissected at length: the Sentinel pattern, where documentation stays green while the product quietly doesn’t exist. If Healthcare.gov were only that story again, I wouldn’t write it. What makes it worth a post is what happened next: within about two months, a small team turned a live national disaster into a working service — and the way they did it became, literally, a government playbook. The failure is the familiar half. The rescue is the instructive one.
The launch: every deferred risk, arriving at once
The pre-launch pathologies are documented in the GAO’s post-mortem reports, and reading them is like reading a checklist of everything this series has warned about, executed simultaneously. No meaningful end-to-end testing until roughly two weeks before launch — which means that for roughly three years, “does the whole thing work” was a question no one had asked of the actual system. Integration across dozens of contractors with no single owner of the answer; each vendor could truthfully report their piece was on track while the assembled whole had never once been assembled. Late requirement changes, including a decision made close to launch to require account creation before users could even browse plans — a change that funneled every curious visitor through the most fragile part of the system and multiplied the load on it. And at launch, effectively no production monitoring: when the site fell over, the people responsible couldn’t see where, or why, or for whom.
I want to be precise about what kind of failure this is, because “the contractors were incompetent” is the lazy reading and the record doesn’t support it. Each contractor was delivering against their contract. The status reports rolling up to leadership were, as far as anyone has shown, mostly accurate reports about component-level work. What no artifact in the whole program measured was the only thing that mattered: can a person get insurance through this system, end to end, today? That question was scheduled to be answered exactly once — at launch, in front of the entire country, with the political weight of a presidency riding on it. It’s the Sentinel mechanism with the stakes turned up: a progress signal that never has to meet reality stays green right up until reality shows up in person. On October 1, reality showed up 250,000 at a time.
The surge: what do you add to a burning project?
Here’s where the story turns, and where it stops being a rerun. In mid-October, the White House assembled what got called the “tech surge” — coordinated by Jeff Zients and U.S. CTO Todd Park, with a Google site-reliability engineer named Mikey Dickerson leading the technical effort. The team was small, ad hoc, and pulled together in days. And the first thing they did is the detail I’d put on a wall: they didn’t fix anything. They built dashboards. Before touching a line of the application, they stood up monitoring so that, for the first time in the project’s life, someone could see what the system was actually doing — error rates, response times, where users were falling out. You can’t fix what you can’t see, and until mid-October 2013, nobody could see.
Think about what that choice means against the instincts of a normal enterprise crisis response. The standard moves are additive process: more oversight, a new program office, a replacement contractor, a re-plan. The surge team added none of that. What they added, piece by piece, was feedback. Monitoring first — a truthful, continuous, unfakeable signal of system health, replacing the status report as the measure of reality. Then a war room with twice-daily standups, morning and evening, so the gap between “something is wrong” and “the people who can fix it know about it” shrank from weeks to hours. Then a single, ruthlessly prioritized punch list — one list, not fifty-five contractor-shaped lists — so at any moment there was exactly one answer to “what’s the most important thing to fix next.” And then fixes shipped continuously, in small increments, dozens of changes flowing out as they were ready, rather than saved up for some big-bang relaunch that would have recreated the original sin at smaller scale.
If you’ve read the manifesto post in this series, you’ll recognize the shape immediately. Working software as the measure of progress — except here the software was live, so the measure was error rates on a dashboard instead of a sprint demo. Short feedback loops. Small releases. Individuals and interactions — a room full of the right people, twice a day — over processes and tools. Nobody in that war room, as far as I can tell, said the word “agile.” They were doing something more convincing than adopting a methodology: under maximum pressure, with no time for ideology, they independently rediscovered the same small set of moves, because those moves are what working on reality’s actual terms looks like. It reads like the manifesto with a pager.
The standup that was actually about something
One rule from those war-room meetings deserves its own section, because I think it’s the single most transferable artifact of the whole rescue. Dickerson’s reported rule for the twice-daily standup was blunt: the meeting is for solving problems, not for reporting status. If you stood up, it was to surface a blocker, make a decision, or ask for help — not to narrate your progress for the benefit of a hierarchy. Status lived on the dashboards, where it belonged, visible to everyone continuously. The humans convened for the one thing dashboards can’t do: resolve.
Sit with how much that inverts the standup most organizations actually run. The default enterprise standup is a status ritual — a round-robin of yesterday/today/no-blockers performed at a manager, which is to say a verbal status report, which is to say the exact artifact this series keeps catching in the act of lying. The Healthcare.gov war room worked precisely because it refused that mode. It could refuse it because the status question was already answered by instrumentation; the meeting’s entire budget was spent on problems. Paired with that was a cultural rule the surge reportedly enforced with equal force: blamelessness. “The people in this room are the ones who will fix it” — not the ones who broke it, not the ones to be investigated, the ones who will fix it. Blame is backward-looking status theater of its own kind, and there were only about six weeks. I’m going to take the standup apart properly in the next post in this series, because the distance between the meeting Dickerson ran and the meeting most teams run at 9:15 every morning is the whole argument of this series in fifteen minutes of ceremony. For now, hold the contrast: same meeting name, opposite meeting.
Eight weeks later
The results came fast, and they came in exactly the currency the dashboards measured. By December 1, 2013 — the surge team’s self-imposed deadline — the site reportedly worked for the vast majority of users. Error rates were reported down from around 6% to under 1%. Uptime, which had languished well below 50% in October, climbed into the 90s. None of it arrived as a relaunch; it arrived as an accumulation of small fixes, each one shipped, measured, and confirmed against live traffic before the next. By April 2014, roughly eight million people had enrolled through the system. The website that couldn’t enroll more than a reported handful on day one had, within a couple of months of feedback-driven repair, done the job it was built for.
And now the fairness note, which I think is the most underrated fact in the whole story. Most of the people who fixed Healthcare.gov were the same contractors who built it. The surge team was tiny; the hands on keyboards in November were largely the hands that had been on keyboards in September. Same engineers, same talent, radically different results — because the system around them changed. In October they had no visibility, fifty-five separate lists, weekly-at-best contact with reality, and an incentive structure organized around not being the one blamed. By November they had dashboards, one list, contact with reality twice a day, and a room where blame was explicitly off the table. If you ever need a clean argument that delivery failures are system failures rather than talent failures, this is the cleanest one on public record. The rescue didn’t swap the people. It swapped the feedback.
The one-off that became a playbook
Most rescue stories end when the fire goes out. This one is in the series because of what happened after. The surge worked well enough, and visibly enough, that the government asked the obvious question: why wait for the fire? In 2014 the administration founded the U.S. Digital Service — with Dickerson as its first administrator — an in-house team built to bring exactly this way of working to government technology before launch day, not after. Alongside it came the Digital Services Playbook: thirteen plays distilled from the rescue and from consumer-internet practice, published at playbook.cio.gov, one of which is literally titled “build the service using agile and iterative practices.” Others read as direct scar tissue from October 2013: default to modern monitoring, test with real users, assign one accountable product owner, deploy in small increments.
That trajectory — emergency improvisation, then codified practice — is worth pausing on, because it’s how every good playbook is actually born, and it carries its own warning. The plays work because each one encodes a feedback loop somebody once bled for. The risk, the moment anything is codified, is the one this series documented with the Spotify model’s many imitators: the practice gets copied, the principle gets left behind, and someone ends up running a “war-room standup” that’s a status meeting with a more dramatic name. The Healthcare.gov rescue is the best recent evidence I know for the claim underneath this whole series — that the operating core of good delivery is a handful of feedback loops, not a framework. The surge team added four of them to a burning project: see the system, meet reality twice a day, keep one list, ship small. Two months later the fire was out. That’s the whole playbook; the other twelve plays are elaboration.
Put it to work
- Before your next crisis, check whether you could even see one. Ask, for your most important service: if it degraded for real users right now, what dashboard would show it, and who’d be looking? If the honest answer is “a customer would email us,” you’re running October 1st infrastructure — and the surge team’s first move applies to you in peacetime: instrument before you optimize, because until you can see the system, every fix is a guess.
- Split status out of your meetings and see what’s left. For one week, move every status update into a written or dashboard channel, and hold your standing meetings to Dickerson’s rule: problems, decisions, and asks only. If the meeting collapses to five minutes of silence, you’ve learned it was a status ritual — and you’ve found the fifteen daily minutes to spend on the blockers nobody was raising.
- When something’s on fire, resist adding process — add feedback. The instinct under pressure is oversight: more reporting, more approvals, a recovery steering committee. Run the surge checklist instead: can everyone see the system’s real state, is there exactly one prioritized list, does the team touch reality at least daily, and are fixes shipping small and continuously? Every “no” is a faster lever than any new process — and cheaper than the meeting you were about to schedule.
Further reading
- Steven Brill, “Obama’s Trauma Team” (Time, 2014) — the definitive inside account of the surge: the war room, the punch list, the twice-daily meetings, and the people in the room; the primary narrative source for the rescue.
- The Digital Services Playbook (playbook.cio.gov) — the thirteen plays the rescue produced; read it as scar tissue, with October 2013 in mind, and every play snaps into focus.
- U.S. Government Accountability Office — the Healthcare.gov report series on the launch failures: testing, integration, requirements churn, and oversight; the sober forensic record of how the fire started.
Agile from First Principles
9 parts in this series.
A nine-part series on agile as reasoning rather than ritual — rereading the manifesto's four values and twelve principles as engineering advice, running fully agile without Scrum or Kanban, cleaning up a rotten backlog, documented product stories (the FBI's Sentinel, the Spotify model, Cyberpunk 2077, Healthcare.gov) that show the principles doing actual work, a full autopsy of the daily standup, and what agile's constraints become when AI agents make implementation cheap.
- 01The Agile Manifesto, Reread: Four Values, Twelve Principles
- 02Running Fully Agile Without Scrum or Kanban
- 03Backlog Cleanup: How to Actually Do It
- 04FBI's Sentinel: Green Status Reports vs. Working Software
- 05The Spotify Model Spotify Doesn't Use
- 06The Deadline That Moved Three Times: Cyberpunk 2077previous
- 07Healthcare.gov: The Rescue That Invented a Playbook← you are here
- 08The Standup Autopsyup next
- 09Agile in the AI Agent Era: The Ceremonies Die, the Principles Don't

What did you take away?
Thoughts, pushback, or a story of your own? Drop a reply below — I read every one.
Comments are powered by Disqus. By posting you agree to theirterms.