A Line-by-Line Appraisal of the Surgical Safety Checklist Study: What Eight Hospitals Showed and What They Could Not
Student Name
American College of Education
NUR4053: Research Methods and Evidence-Based Practice in Nursing
Module 4 Assignment
Instructor Name
July 27, 2026
The Study and Its Purpose
Haynes et al. (2009) set out to test whether introducing a 19-item surgical safety checklist, developed through the World Health Organization's Safe Surgery Saves Lives program, would reduce complications and deaths after surgery. The checklist groups its items around three moments: before induction of anesthesia, before skin incision and before the patient leaves the operating room. It asks the team to confirm the patient's identity and procedure, review anticipated critical events, confirm antibiotic prophylaxis and count instruments and sponges, among other steps. The authors hypothesized that a tool designed to improve team communication and consistency of care would reduce harm.
The purpose is stated clearly and the hypothesis is directional, which makes the study's success or failure easy to judge. The question also matters to nursing directly. Operating room nurses often lead the checklist's sign-in and sign-out phases, so whether the tool works bears on how perioperative nurses spend minutes that are already scarce. Melnyk and Fineout-Overholt (2019) suggest that a rapid appraisal begin by asking whether a study's question matches the reader's own clinical question, and for a nurse considering how to strengthen checklist use, this study is an obvious match.
Design, Setting and Sample
The study used a before-and-after design, sometimes described as a quasi-experimental pretest and posttest design without a concurrent control group. Between October 2007 and September 2008, eight hospitals in eight cities took part: Toronto, New Delhi, Amman, Auckland, Manila, Ifakara, London and Seattle. Researchers collected data on 3,733 consecutively enrolled patients aged 16 or older having noncardiac surgery before the checklist was introduced, and then on 3,955 consecutively enrolled patients afterward.
The setting is the study's great strength and one of its complications. Including hospitals in high-, middle- and low-income countries makes the findings relevant across very different health systems, and consecutive enrollment reduces the risk that researchers chose patients likely to do well. But the hospitals differed widely in baseline complication rates, staffing and resources, so a pooled result may hide very different effects at individual sites. The design is the larger concern. Without a group of hospitals that did not receive the checklist during the same period, any change after its introduction could reflect the checklist, other improvements happening at the same time or simply the attention that comes with being studied.
The sample size was large enough to detect the differences the study reported, and the authors enrolled patients consecutively rather than by convenience, which is a real methodological strength for a study of this design. What the sample cannot do is separate the checklist's effect from the effect of time.
Intervention and Outcome Measures
The intervention was more than a piece of paper. Hospitals received the checklist along with training for local team members, and each site had to adapt it to its own practices. That matters for appraisal because the result may reflect the whole implementation effort, including training, leadership attention and data collection, rather than the checklist alone. A hospital that simply printed the checklist and placed it in its operating rooms might not see the same result.
The primary outcome was the rate of complications, including death, during hospitalization within the first 30 days after the operation. Complications were defined in advance and included surgical site infection, unplanned return to the operating room, pneumonia and other major events. Measuring inpatient complications is practical, but it misses complications that develop after discharge, which is a meaningful limitation for surgical site infections that often appear a week or more after the operation. The outcome data were collected by local staff at each site, who knew whether they were in the before or after period, which introduces a possibility of detection bias if staff looked harder, or less hard, for complications once the checklist was in use.
Results
The results were striking. The death rate fell from 1.5 percent before the checklist to 0.8 percent afterward, a difference the authors reported as statistically significant with a P value of 0.003. Inpatient complications fell from 11.0 percent to 7.0 percent, with a P value below 0.001. In absolute terms, the complication figures mean four fewer patients with a complication for every 100 operations, and the mortality figures mean about seven fewer deaths for every 1,000 operations.
Translating relative changes into absolute ones matters when judging clinical significance. A reduction in deaths of nearly half sounds dramatic; seven fewer deaths per 1,000 operations is still an important effect, but it depends heavily on the baseline death rate. In a hospital where surgical mortality is already well below 1.5 percent, the same relative reduction would save far fewer lives. The authors reported the combined results and noted that the reductions occurred across the hospitals, but the effect sizes are best read as belonging to this mix of sites.
Limitations and Later Evidence
The authors acknowledged that the before-and-after design could not rule out other causes of the improvement, including the Hawthorne effect, in which people change their behavior because they know they are being observed. They also noted that the study could not identify which of the checklist's 19 items, or which part of the implementation, produced the change. Those acknowledgments are appropriate. Two limitations received less attention: the unblinded collection of outcome data by local staff and the relatively short observation periods, which cannot show whether the effect lasted once the study ended.
Later evidence tests those concerns directly. Urbach et al. (2014) studied the introduction of surgical safety checklists across 101 hospitals in Ontario, Canada, using administrative data on more than 100,000 procedures in each of two three-month periods before and after adoption. They found no significant reduction in operative mortality, which was 0.71 percent before and 0.65 percent after, or in surgical complications. Ontario's baseline rates were much lower than those in the original study, and the checklists were introduced under a government mandate rather than a supported implementation program. Read together, the two studies suggest that a checklist is not a treatment with a fixed effect; its value depends on the setting's baseline risk and on how seriously teams use it.
Implications for Perioperative Nursing
For a nurse deciding how to use this evidence, the appraisal points to three practical conclusions. First, the checklist should be treated as a team conversation rather than a document, because the original study delivered it with training and leadership support, and the later study that relied on a mandate found no effect. Second, a unit that adopts or refreshes the checklist should measure its own baseline complication and surgical site infection rates, including infections found after discharge, so that it can judge whether its results resemble the first study or the second. Third, the circulating nurse's role in leading the sign-in and sign-out phases deserves protected time, since a checklist read aloud while the team is setting up instruments is not the intervention that was tested.
Conclusion
Haynes et al. (2009) is a well-conducted study with a clear hypothesis, a diverse international sample and results that were large enough to change surgical practice around the world. Its before-and-after design, unblinded outcome collection and short observation periods mean it cannot prove that the checklist itself caused the improvement, and a later study in a lower-risk system found no comparable effect. For a perioperative nurse, the study supports using the checklist as a tool worth doing well, with real team engagement, rather than as a form to complete. Its weight as evidence is moderate: strong enough to justify the practice, not strong enough to promise a particular result.
References
Haynes, A. B., Weiser, T. G., Berry, W. R., Lipsitz, S. R., Breizat, A.-H. S., Dellinger, E. P., Herbosa, T., Joseph, S., Kibatala, P. L., Lapitan, M. C. M., Merry, A. F., Moorthy, K., Reznick, R. K., Taylor, B., & Gawande, A. A. (2009). A surgical safety checklist to reduce morbidity and mortality in a global population. New England Journal of Medicine, 360(5), 491-499. https://doi.org/10.1056/NEJMsa0810119
Melnyk, B. M., & Fineout-Overholt, E. (2019). Evidence-based practice in nursing and healthcare: A guide to best practice (4th ed.). Wolters Kluwer.
Urbach, D. R., Govindarajan, A., Saskin, R., Wilton, A. S., & Baxter, N. N. (2014). Introduction of surgical safety checklists in Ontario, Canada. New England Journal of Medicine, 370(11), 1029-1038. https://doi.org/10.1056/NEJMsa1308261
How this NUR 4053 Module 4 example is structured
NUR 4053 Module 4 in many sections appraises one study line by line, sample through limitations; your classroom's instructions decide which appraisal tool or question set to use. This example follows the order of the article itself, which makes the critique easy to check against the source: purpose, design, setting and sample, intervention, outcome measures, results, then the limitations the authors named and the ones they did not. Each section describes what the study did before judging it. A final section weighs the study against a later, larger evaluation, because appraisal is incomplete until a study is placed beside the evidence that followed it.
NUR4053 Module 4 questions, answered
What does NUR4053 Module 4 usually ask for?
NUR4053 Module 4 in many sections asks students to appraise one research study in detail, from its purpose, design and sample through its results and limitations. Some instructors provide an appraisal tool or a set of questions; others expect a narrative critique in APA format. Your classroom's instructions decide the study type and the tool.
How do I judge clinical significance in a research critique?
Convert the reported percentages into absolute differences, such as events prevented per 100 or per 1,000 patients, and consider the baseline rate. A large relative reduction can mean a small absolute benefit when the outcome is rare. Then ask whether the difference would matter to patients and whether it would hold in your own setting.
Should a critique include studies other than the one being appraised?
Briefly, yes. Placing the study beside later or larger studies on the same question shows whether its findings have held up. Keep most of the paper on the assigned study, but a paragraph comparing it with the evidence that followed often separates a strong critique from a summary.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.