| Course | CAP 6003 Capstone in Education Specialist |
|---|---|
| Module | Module 4 |
| Paper type | Evaluation plan |
| Length | 1,300 words, about 5 pages plus title and reference pages |
| Format | APA 7 student paper |
| School | American College of Education |
| Program | Ed.S. in Nursing Education |
| Updated | October 2026 |
Free sample paper for CAP 6003 Module 4
Seconds to the Call: An Evaluation Plan for the Call Early Curriculum Thread, With a Historical Comparison Cohort, Rubric Ratings and an Honest Account of What It Cannot Show
Student Name
American College of Education
CAP6003: Capstone in Education Specialist
Module 4 Assignment
Instructor Name
November 2, 2026
Introduction
Module 3 designed Call Early, a three-part curriculum thread intended to shorten the time between a nursing student noticing deterioration and calling for help. This paper sets out how the composite Arkansas associate degree program will judge whether it worked. It states the evaluation questions, the framework that organizes them, the comparison design, the measures and data sources, the analysis, the targets that would count as success and the limits on what any of it can prove.
Evaluation Questions
Three questions guide the plan. First, was Call Early delivered as designed, to the students it was meant for? Second, do fourth-semester students who complete it escalate sooner and more clearly in the end-of-program sepsis scenario than the fall 2025 cohort did? Third, is there any sign that the change reaches clinical practice? The first is a process question; the second and third are outcome questions.
Framework
Kirkpatrick (1998) sorted the evidence a training program can produce into four tiers: reaction, meaning how participants felt about it; learning, meaning what knowledge, skills or attitudes changed; behavior, meaning whether people act differently on the job; and results, meaning effects on the organization's outcomes. The levels are useful here because they force a plain statement of how far the evidence will reach. This plan measures reaction and learning thoroughly, behavior in simulation directly and behavior in clinical practice only roughly. It does not measure results in the form of patient outcomes, which no program-level evaluation of this size could attribute to a curriculum change.
Design
The main comparison is historical. The spring 2027 fourth-semester cohort will complete the same sepsis scenario, with the same programmed change at minute 12 and the same manikin settings, as the fall 2025 cohort, whose median time to call was nine minutes and whose rate of calling within five minutes was 39%. Prebriefing and debriefing will differ, as the design requires, but the scenario events up to the call will not.
A historical comparison cannot rule out other explanations. The spring cohort may differ in ability, clinical placements or instructors, and students may hear about the scenario from earlier cohorts. The plan reduces these threats where it can: it will compare the two cohorts' entrance examination scores and third-semester course grades to check whether they are similar, and it will keep scenario details unchanged so that any leak would have affected both cohorts. Even so, the evaluation can show whether results changed, not prove that Call Early changed them.
Measures and Data Sources
Time to escalation is the primary measure: seconds from the programmed change to the moment the student starts dialing anyone with authority to act, read from the scenario video's time stamp by a rater who did not run the session. From it come two secondary measures, the proportions of students calling within three minutes, the design's objective, and within five minutes, the baseline comparison.
To judge how good the call was, raters will use only the responding section of the rubric Lasater (2007) developed, whose four dimensions cover a calm, confident manner, clear communication, a well-planned intervention and skill. Adamson et al. (2012) summarized three studies of the rubric in simulation and reported interrater reliability ranging from an intraclass correlation of 0.889 in one study to percent agreement between 57% and 100% in another, along with evidence supporting its validity for simulated care. Because agreement varied, two trained raters will score every video independently, and their agreement will be reported.
Students' confidence will be measured with a single program-made item, "How confident are you that you would call a provider promptly about a deteriorating patient?" rated 1 to 10, asked before the workshop and after the sepsis scenario. The item has no validity evidence and will be reported as a reaction measure only. Process will be tracked through workshop attendance, the number of mock calls each student completed and a weekly tally from each clinical instructor of how many students answered the prompt-card questions. Clinical behavior will be estimated from instructors' reports of student-initiated escalations, recorded on a simple form.
Analysis
Because time-to-event data in simulation tend to be skewed, the two cohorts' times will be compared with medians and a Mann-Whitney U test, and the proportions calling within three and five minutes with chi-square tests. Effect sizes will be reported alongside p values. Rubric agreement will be summarized with an intraclass correlation. Confidence ratings will be compared before and after within the spring cohort. With roughly 84 students in each cohort the comparison has reasonable power for a large change in median time but not for a small one, so a nonsignificant result will be reported as inconclusive rather than as evidence of no effect.
Targets
Faculty agreed on targets before seeing any spring data: a median time to call under four minutes, at least 70% of students calling within five minutes and responding-phase rubric scores at the accomplished level for at least half of students. Delivery targets are attendance of 90% at the workshop and prompt-card use on at least three of every four clinical days. Missing a delivery target would not end the project, but it would change how the outcome results are read: a thread that reached only part of the cohort cannot be judged on the whole cohort's times, so results will also be reported separately for students who attended the workshop and those who did not.
Evaluation Timeline and Responsibilities
In January, before any student contact, I will pull the fall 2025 scenario videos and rescore them with the same two raters who will score the spring videos, which keeps the scoring rules the same before and after the change; the exercise also gives the raters practice and a first estimate of their agreement. Workshop attendance and mock-call counts will be logged in February by the faculty who run the workshop. Prompt-card tallies will arrive weekly from March to May from each clinical instructor through a two-question online form. The sepsis scenario runs in the second week of April, the raters score videos within two weeks and I will complete the analysis by the first week of May. The program director has agreed to review a draft of the results before they go to faculty, which gives a second reader the chance to question any conclusion that goes further than the data.
Data Handling and Ethics
All data are collected as part of routine course evaluation. The recordings will sit in a restricted folder on the college server, be scored with student names replaced by codes and deleted after the semester's course review, as current program policy requires. Results will be reported only in aggregate. The college's institutional review office has been asked to confirm that the evaluation is quality improvement, and no data will leave the program for publication without its approval.
Reporting
Findings will go to three audiences in three forms: a two-page summary for the faculty meeting in May, a slide for the curriculum committee's annual review and a short message to the partner hospitals' nurse educators, who will orient these graduates in the summer. Students in the spring cohort will receive a summary at their final post-conference.
Conclusion
The plan measures what the program can measure well, time and quality of the call in a controlled scenario, and it is honest about what it cannot, the effect on patients after graduation. If the spring cohort calls in under four minutes where the last one took nine, and the rubric shows clearer calls, the program will have good reason to keep Call Early and to test it again with a new cohort.
References
Adamson, K. A., Gubrud, P., Sideras, S., & Lasater, K. (2012). Assessing the reliability, validity, and use of the Lasater Clinical Judgment Rubric: Three approaches. Journal of Nursing Education, 51(2), 66-73. https://doi.org/10.3928/01484834-20111130-03
Kirkpatrick, D. L. (1998). The four levels of evaluation. In S. M. Brown & C. J. Seidner (Eds.), Evaluating corporate training: Models and issues (pp. 95-112). Springer. https://doi.org/10.1007/978-94-011-4850-4_5
Lasater, K. (2007). Clinical judgment development: Using simulation to create an assessment rubric. Journal of Nursing Education, 46(11), 496-503. https://doi.org/10.3928/01484834-20071101-04
CAP 6003 Module 4 instructions, in plain terms
The fourth CAP 6003 module generally asks how you will know whether your capstone worked. Prompts often call for evaluation questions, a framework or model, a design, specific measures with their data sources, an analysis plan and criteria for success, and some ask for a timeline and a plan for reporting results. Choose measures you can actually collect in your setting, and say which ones are validated and which are not. Be explicit about the design's weaknesses, especially when you compare cohorts or terms rather than randomly assigned groups. Describe how data will be protected and who will see the results, and set targets before the data arrive.
How the CAP 6003 Module 4 example is put together
The plan begins with three evaluation questions, one about delivery and two about outcomes. Kirkpatrick's four levels are then used to state how far the evidence will reach, including a plain admission that patient outcomes are out of scope. A design section sets out the historical comparison and the steps that reduce its threats. Measures follow in detail: timed calls from video, two rubric raters, a confidence item labeled as unvalidated and process tallies. The analysis section matches tests to skewed time data and proportions and treats a null result as inconclusive. Targets, data handling, reporting to three audiences and a short conclusion complete the plan. Each measure is named once with its source and its known weaknesses, so the reader never has to guess where a number will come from or how much weight it can bear.
CAP 6003 Module 4 rubric: what full marks look like
Capstone evaluation plans tend to be judged on alignment between questions and measures, the soundness of the design and the honesty of its limits. Graders look for measures tied to the intervention's objectives, data sources the writer can really reach, analysis methods suited to the data and success criteria stated in advance. Recognizing threats to a comparison and taking practical steps against them is a mark of specialist-level judgment, as is reporting reliability for rated measures. Attention to data security and to reporting results back to stakeholders shows a complete plan. APA 7 citations for the framework and instruments, including book chapters, round out the paper.
Common CAP 6003 Module 4 mistakes, and how to avoid them
Evaluation plans frequently promise to measure everything with no clear way of collecting any of it, or rely on a single satisfaction survey. If you need help choosing measures, picking a framework such as Kirkpatrick's, setting targets or explaining the limits of a cohort comparison, we can help you build it. Send the intervention you designed and the module prompt, and a Module 4 evaluation plan will come back matched to the data your program already gathers. Health education capstones can be planned the same way around attendance, knowledge checks and behavior measures. A one-page data collection calendar or rating form can also be drafted.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.
More CAP 6003 and Ed.S. in Nursing Education sample papers
- CAP 6003 Module 1: Capstone Proposal
- CAP 6003 Module 2: Research Synthesis
- CAP 6003 Module 3: Intervention Design
- CAP 6003 Module 5: Capstone Report and Portfolio
- NUR 6033 Module 4: Gamified Learning Activity Design
- NUR 6013 Module 3: Scholarship Role Analysis
- NUR 6023 Module 6: Deep Learning Assessment Plan
- NUR 6043 Module 2: Curriculum Design Comparison
CAP 6003 Module 4 questions, answered
What does CAP6003 Module 4 usually ask for?
Many CAP6003 sections give Module 4 to the evaluation plan: questions, a framework, design, measures, data sources, analysis and targets showing how you will judge whether the capstone intervention worked.
What are Kirkpatrick's four levels?
Reaction, learning, behavior and results, a way to sort training evaluation from how participants felt to whether the organization's outcomes changed.
Is a historical comparison group acceptable in an Ed.S. capstone?
Often yes, when random assignment is impossible, provided you state its weaknesses and check whether the cohorts are similar on measures you already have.
Where can I find a free CAP 6003 Module 4 sample paper?
You can read it here in full: an evaluation plan timing nursing students' calls from scenario video, rating them on the Lasater rubric and comparing two cohorts against preset targets.
Why set targets before collecting data?
Agreeing on what counts as success in advance keeps the evaluation honest; targets chosen after seeing results can be bent to fit them.