Did the Standardized Patients Change What Students Do on the Unit? A Four-Level Evaluation Plan for a Mental Health Simulation Program That Refuses to Stop at Satisfaction
Student Name
American College of Education
NUR6033: Innovation in Nursing Education
Module 6 Assignment
Instructor Name
August 12, 2030
The Question and the Framework
Over this course I have analyzed a learning need, designed a standardized patient scenario, planned its debriefing, built a gamified activity and planned the expansion of the standardized patient program to three courses. Permanent funding now depends on what this evaluation shows. The question the evaluation must answer is not whether students like the encounters, which the pilot already showed, but whether they change what students do with real patients.
Kirkpatrick's framework classifies the outcomes of training into four levels: reaction, learning, behavior and results. A systematic review of 13 studies evaluating debriefing after high-fidelity simulation in health care education found a paucity of studies at the highest levels of evaluation, identifying this as an area where research is needed (Johnston et al., 2018). Most simulation evaluations stop at reaction and learning because behavior and results are harder to measure. This plan tries to reach all four, while being honest about how firmly each can be measured. Satisfaction is where evaluation is easiest and where it matters least.
Level One: Reaction
After every encounter and its debriefing, students will complete a brief questionnaire rating relevance, realism, psychological safety and the usefulness of the debriefing on five-point scales, with one open question on what the student will do differently. Reaction data are useful for improving the program, for example by identifying a debriefer whose sessions students find unsafe, but they will not be used as evidence that the program works. A review of published simulation evaluation instruments found that many instruments existed but that evidence of their reliability and validity varied, and recommended that educators choose instruments with reported psychometric evidence (Adamson et al., 2013). For reaction I will use a brief locally developed survey, since the stakes are low, but for learning and behavior I will use instruments with reported evidence.
Level Two: Learning
Learning will be measured by performance, not by self-report. For the mental health encounters, faculty will score each student's first and second encounters against the four objectives, as in the pilot, with two raters scoring a random 20% of encounters to check agreement. I will also compare the cohort's scores on the mental health course's therapeutic communication examination items with the two cohorts before the program began. For the health assessment and leadership encounters, each course's faculty will define a comparable checklist. The main learning measure is the change between first and second encounters within each student, which shows whether practice with debriefing produces improvement, and the comparison with previous cohorts, which shows whether the program adds to what the course achieved before.
The published evidence gives a sense of what to expect. A meta-analysis of 62 studies of mental health simulation found that standardized participants produced large pretest-to-posttest gains in competence and confidence and reduced anxiety, and were particularly effective for clinical preparedness (Zhang & Wang, 2025). Most of those gains were measured within the simulation itself, so a similar within-student improvement here would confirm that the program works as others have, but it would not yet show that the improvement carries into practice. That question belongs to the next level.
Level Three: Behavior
Behavior means what students do in clinical practice, and measuring it is the heart of the plan. During the four-hour psychiatric placement that follows the encounters, the preceptor, a staff nurse, will complete a short rating of each student's interaction with a patient experiencing psychosis on the same four behaviors as the simulation objectives: introducing themselves and their role, responding to the patient's experience without disputing its content, asking about risk and setting limits calmly. Preceptors will be given a 15-minute orientation to the rating. To give the ratings a comparison, preceptors on the same unit will rate a cohort of students from the program's evening section, who will receive the standardized patient encounters one term later, during their placements this term. The two groups are similar in admission criteria and curriculum, although not randomized. A difference between the two groups in preceptor ratings would be the strongest evidence the evaluation can produce that the encounters change behavior.
Level Four: Results
Results are the effects on patients and the organization, and here the plan must be most modest. A four-hour student placement will not change patient outcomes on a psychiatric unit in any measurable way. What can be measured is a result that matters to the program and its partners. The psychiatric partner reduced student access because students were seen as a burden on a stressed unit. I will ask the unit's nurse manager to rate, on a short survey at the end of each term, whether students were helpful, neutral or burdensome, and whether the unit would accept more student hours. An increase in the partner's willingness to host students would be a real result for the program, since it addresses the original problem of shrinking placements. I will also track whether graduates who choose psychiatric nursing as their first job increase over three years, recognizing that many factors affect that choice.
Ethics and Data Handling
Evaluation data about students are educational records and must be handled accordingly. Encounter scores, preceptor ratings and survey responses will be stored in the program's secure system, identified by study number rather than name for analysis, and only the simulation committee will have access to identified data. Because the evaluation may be published, the plan will go to the university's human subjects review before any data are gathered, and students will be asked for consent to include their de-identified data in any publication; students who decline will still take part in the encounters, and their data will be used only for internal program improvement. Preceptors will be told that their ratings are for program evaluation, not for student grades, which should make them more candid, and the ratings will not be shared with course faculty in identified form. Video recordings of encounters, used in debriefing, will be deleted at the end of each term unless a student consents to their use in faculty training. These steps take time to set up, but they protect students, allow the results to be shared beyond the program and prevent the evaluation from compromising the psychological safety the encounters depend on.
Analysis, Reporting and Limits
Within-student change in encounter scores will be analyzed with paired comparisons, and differences between the program and comparison groups in preceptor ratings with independent comparisons, reporting effect sizes as well as statistical significance, since the samples will be modest. Qualitative comments from surveys will be analyzed for recurring themes. Results will be reported each term to the simulation committee and each year to the curriculum committee and the dean, and the full evaluation after two years will be prepared as a scholarly paper. The plan has limits. The comparison groups are not randomized, preceptors may know which students received the encounters, and the four-hour placement gives only a brief view of behavior. I will report these limits alongside the findings, because a program funded on overstated evidence is vulnerable when the next budget is set.
References
Adamson, K. A., Kardong-Edgren, S., & Willhaus, J. (2013). An updated review of published simulation evaluation instruments. Clinical Simulation in Nursing, 9(9), e393-e400. https://doi.org/10.1016/j.ecns.2012.09.004
Johnston, S., Coyer, F. M., & Nash, R. (2018). Kirkpatrick's evaluation of simulation and debriefing in health care education: A systematic review. Journal of Nursing Education, 57(7), 393-398. https://doi.org/10.3928/01484834-20180618-03
Zhang, X., & Wang, H. (2025). Comparative effectiveness of mental health simulation techniques in nursing education: A systematic review and meta-analysis. International Journal of Mental Health Nursing, 34(1), Article e13502. https://doi.org/10.1111/inm.13502
How this NUR 6033 Module 6 example is structured
NUR 6033 Module 6 frequently closes with an evaluation plan for an educational innovation; your classroom's instructions decide the model. This example names an evaluation framework from its source, applies it level by level with specific measures and instruments, chooses a comparison design, plans data collection and analysis, and states the plan's limits.
NUR6033 Module 6 questions, answered
What does NUR6033 Module 6 usually ask for?
NUR6033 Module 6 frequently asks for an evaluation plan for an educational innovation, using a named model and specifying measures, design and analysis. Your classroom's instructions decide the model.
What are Kirkpatrick's four levels?
Reaction, learning, behavior and results. Most simulation evaluations stop at the first two; stronger plans also measure behavior in practice and results for patients or organizations.
How do I measure behavior after simulation?
Ask clinical preceptors to rate the same behaviors the simulation targeted during real patient care, ideally with a comparison group of students who have not yet had the simulation.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.