From July 2010, hospitals in the Canadian province of Ontario had to report publicly how consistently their operating teams were using a surgical safety checklist. The reported figures were excellent. After July 2010, compliance never fell below 98 per cent. David Urbach and colleagues then compared around 109,000 operations before the checklist was introduced with around 106,000 afterwards, across 101 hospitals. Their results appeared in the New England Journal of Medicine in 2014. Adjusted mortality was 0.71 per cent before and 0.65 per cent after; the complication rate was 3.86 and 3.82 per cent. Neither difference was statistically significant. The checklist had been introduced, by the hospitals’ own account it was almost always worked through, and for patients nothing measurable changed.
The Original
This was not what anyone had expected. Five years earlier, the same journal had published the study that made the checklist’s reputation. In eight hospitals on five continents, from Seattle and London to Manila and Ifakara in Tanzania, mortality fell from 1.5 to 0.8 per cent and the complication rate from 11 to 7 per cent after the 19-item WHO checklist was introduced. One of the authors, the surgeon Atul Gawande, made the checklist known far beyond medicine that same year with his book The Checklist Manifesto.
The WHO study was a before-and-after comparison without a control group. That does not make its result wrong, but a design of this kind cannot separate what did the work: the checklist, or everything that arrived with it in the eight hospitals. Ontario, unintentionally, offers a clue. There the checklist arrived by regulation, and much of what had accompanied it in the pilot hospitals did not come along.
How a Checklist Is Meant to Work
Gawande himself describes what an effective checklist looks like, drawing on Daniel Boorman, who developed flight deck checklists at Boeing. It is kept short, five to nine items as a rule of thumb. It specifies whether it is to be worked through step by step (“read-do”) or confirmed jointly once the work is done (“do-confirm”). And it is not written at a desk. In Gawande’s account, first drafts always fall apart; you have to watch where they fail and keep revising until the checklist holds up in real use.
There is a second purpose, and it is easy to miss. Before the first incision, the WHO checklist has everyone in the room introduce themselves by name and role. Gawande describes how team members who have spoken once at the start are more likely to speak up later when they notice something. The checklist thus creates a moment in which surgery, anaesthesia and nursing talk to each other across the hierarchy. The tick on the form only shows that this moment was supposed to happen.
What Was Missing in Ontario
Gawande responded to the Ontario study himself in March 2014. The mandate, he wrote, had come with “no team training, local adaptation of the checklist, or tracking of adoption”. Research had shown that “self-reported compliance does not remotely correlate with reality”. Measuring mortality without knowing whether teams were actually using the checklist was “like running a drug trial without knowing if the patients actually took the drug”. His conclusion: “If you don’t use it, it doesn’t work.”
What stands out is that the other side hardly disputes this. On publication, Urbach himself said: “While a greater effect of surgical safety checklists might occur with more intensive team training or better monitoring of compliance, as currently implemented, surgical safety checklists did not result in improved patient outcomes.” Both sides, then, look for the explanation in how the checklist was introduced. That explanation does not go far enough. In the eight pilot hospitals, the same research group also surveyed how surgical teams rated teamwork and safety climate. After the checklist was introduced, the scores rose slightly but significantly, from a mean of 3.91 to 4.01 on a scale of 1 to 5. And the improved ratings went hand in hand with the fall in complications and deaths. That is a correlation, not proof. But it fits what the round of introductions is for: the checklist works in good part through the culture in the room, that is, through whether someone who notices something says so, and whether they are heard. Ken Catchpole and Stephanie Russ reached a similar conclusion in 2015 for checklists in healthcare generally: anyone who wants to introduce them effectively has to attend to the sociocultural conditions they land in, and not only to the checklist itself.
A compliance rate measures whether the form was filled in. Whether the checklist set off a conversation, it does not measure.
A second case shows how much cultural work can lie behind a successful checklist. In the Michigan project, across 103 intensive care units, the mean rate of central line bloodstream infections fell from 7.7 to 1.4 per 1,000 catheter-days after eighteen months. In his book, Gawande presents the project as evidence of what a checklist can do. In 2011 the sociologist Mary Dixon-Woods and colleagues, among them the programme’s leader Peter Pronovost himself, examined why it worked. Popular accounts of the programme, they wrote, “have often been simplistic and partial, and have perpetuated the myth that the program’s achievements can be traced to a ‘simple checklist’ rather than a complex social intervention”. One of the mechanisms they describe is cultural: the programme broke the norm that such infections are inevitable in intensive care. Another was that the data collected made every unit’s rate visible and ranked the units, which created pressure.
What the Case Teaches
Ontario introduced what can be decreed: the mandate, the form and the rate. What had made the checklist work in the pilot hospitals and in Michigan cannot be decreed. One part of it is a checklist that the team has tested and adapted in use. The other is a culture in which people actually speak at the pause point, including from the bottom up, and in which feedback shows whether outcomes actually change. The 98 per cent measured the ticking of boxes, and the ticking of boxes was the only thing demonstrably introduced.
This holds far beyond the operating theatre, wherever checklists are mandated and managed through compliance rates. If you want to know whether a checklist works where you are, do not look at the rate. Go to one of the pause points and listen for whether anyone says something that would have gone unsaid without the checklist.
Thoughts on this? Take it to LinkedIn.
No comments here – but discussion is welcome, where the readership is anyway.
Join the discussion on LinkedInSources
- David R. Urbach et al. – Introduction of Surgical Safety Checklists in Ontario, Canada, New England Journal of Medicine 370(11), 2014, pp. 1029–1038
- Alex B. Haynes et al. – A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population, New England Journal of Medicine 360(5), 2009, pp. 491–499
- Atul Gawande – The Checklist Manifesto: How to Get Things Right, Metropolitan Books 2009
- Atul Gawande – When Checklists Work and When They Don’t, The Incidental Economist, 15 March 2014
- Mary Dixon-Woods et al. – Explaining Michigan: Developing an Ex Post Theory of a Quality Improvement Program, Milbank Quarterly 89(2), 2011, pp. 167–205
- Alex B. Haynes et al. – Changes in Safety Attitude and Relationship to Decreased Postoperative Morbidity and Mortality Following Implementation of a Checklist-Based Surgical Safety Intervention, BMJ Quality & Safety 20(1), 2011, pp. 102–107
- Ken Catchpole & Stephanie Russ – The Problem with Checklists, BMJ Quality & Safety 24(9), 2015, pp. 545–549
- ICES (Institute for Clinical Evaluative Sciences) – news release on the Ontario study, March 2014