Predict, Don't Name: Writing Application-Level Items and a Mechanism Rubric for the Heart Failure Unit, Then Reading the Item Analysis From 64 Accelerated Students
Student Name
American College of Education
NUR5233: Curriculum Development, Assessment, and Evaluation in Nursing
Module 4 Assignment
Instructor Name
August 1, 2028
Why Item Writing Matters Here
The heart failure unit in Module 3 includes a weekly quiz worth 30% of the unit grade, and its outcome asks students to explain drug mechanisms and predict what nurses must monitor. If the quiz items ask only for drug names or classes, the quiz will reward recall and undo the alignment the unit was designed for. The problem is common. In one nursing department's review of 2,770 multiple-choice questions used over five years, 46.2% contained at least one violation of accepted item-writing guidelines, and more than 90% were written at low cognitive levels; questions at lower cognitive levels were more likely to be flawed (Tarrant et al., 2006). A later study of high-stakes nursing examinations found that flawed items were common and that they tended to penalize high-achieving students more than borderline ones (Tarrant & Ware, 2008). Flawed items, in other words, distort the measure for the students the program most wants to identify.
The Rules the Faculty Adopted
The faculty adopted a short set of rules drawn from a widely cited review of multiple-choice item-writing guidelines (Haladyna et al., 2002). Each item tests one important idea from the unit outcomes. The stem poses a complete question that could be answered without seeing the options. Options are plausible, similar in length and grammar, and free of clues such as absolute words or repetition of stem language in the correct answer. None of the above and all of the above are not used. Three or four options are enough; a distractor no student would choose adds reading time without adding information. And, for this unit, every item places the student in a clinical situation and asks what the student would expect, monitor or do, so that recall alone is not enough to answer it.
Two Sample Items
Item 3. A patient with chronic heart failure is started on furosemide 40 mg orally each morning. Which finding at the next day's assessment would indicate an expected effect of the drug? A. Weight 1.2 kg less than the previous morning. B. Blood pressure 12 mm Hg higher than yesterday. C. Serum potassium 0.6 mmol/L higher than yesterday. D. Heart rate 20 beats per minute slower than yesterday. The correct answer is A. The item requires the student to reason from the mechanism, a loop diuretic causing loss of sodium and water, to its measurable effect, weight loss from fluid removal. The distractors each represent a plausible misunderstanding: that the drug raises blood pressure, that it retains potassium, confusing it with a potassium-sparing diuretic, or that it slows the heart, confusing it with a beta-blocker.
Item 8, as originally written. A patient with heart failure is started on lisinopril. Which assessment is the nurse's priority before the second dose? A. Blood pressure. B. Serum potassium. C. Presence of a dry cough. D. Serum sodium. The item analysis below shows what went wrong with it.
The Mechanism Rubric
The written task asks students to account for two familiar features of heart failure, weight gain and breathlessness on lying down, and is scored with a four-level rubric whose levels differ by the quality of the causal chain, not by adjectives. At level four, the student traces the full chain: reduced cardiac output leads to reduced kidney perfusion, activation of the renin-angiotensin-aldosterone system and retention of sodium and water, raising venous pressure; lying flat increases venous return to an already overloaded heart, raising pressure in the pulmonary circulation and producing fluid in the lungs. At level three, the chain is correct but one link is missing, most often the reason lying flat worsens symptoms. At level two, the student names the right mechanisms but does not connect them causally. At level one, the student lists signs without a mechanism. Two faculty independently scored the same 20 responses before the rubric was used, agreed on 17, and revised the wording of level three to resolve the disagreements.
Reading the Item Analysis
The quiz was taken by 64 students. Difficulty was calculated as the proportion of students answering each item correctly, and discrimination as the difference between the proportion correct in the top 27% of scorers, 17 students, and the bottom 27%, another 17. Most items performed well. Four deserve comment.
Item 3 had a difficulty of 0.72 and a discrimination index of 0.47, with 16 of the top 17 and 8 of the bottom 17 answering correctly: a moderately difficult item that separates stronger from weaker students well. It will be kept.
Item 6, which asked which drug class metoprolol belongs to, had a difficulty of 0.95 and a discrimination index of 0.06. Nearly everyone answered correctly, so the item gave almost no information, and it tested recall, which the unit's outcome does not target. It will be rewritten to ask what the nurse should assess before giving the drug.
Item 8 had a difficulty of 0.31 and a negative discrimination index of minus 0.06: 5 of the top 17 answered correctly, compared with 6 of the bottom 17. Forty-one percent of all students, including 11 of the top group, chose potassium rather than the keyed answer, blood pressure. When the strongest students choose a distractor more often than the weakest, the problem is usually the item, not the students. On review, both answers are defensible, since an angiotensin-converting enzyme inhibitor can cause both first-dose hypotension and a rise in potassium. The item will be credited for both answers on this quiz and rewritten with a scenario that makes one priority clear.
Item 9 had acceptable difficulty and discrimination, 0.58 and 0.35, but one distractor was chosen by no student. That distractor will be replaced with one based on a common misconception, drawn from the wrong answers students gave in the written explanations.
Limits of a Ten-Item Quiz
A ten-item quiz is short, and its reliability as a total score is modest; its purpose is formative feedback within the week, not a high-stakes decision. Item statistics from 64 students are also unstable, so the faculty will look for patterns across two cohorts before discarding items that perform poorly once. The larger lesson for the curriculum is that item analysis should be a routine step after every quiz and examination, reviewed by the course team before grades are released, so that flawed items are caught before they affect students. The course team will keep a shared bank of items with their statistics from each administration, marked by outcome and cognitive level, so that the next cohort's quizzes are built from items with known performance and new items are piloted alongside them. Over two or three years, the bank will also show whether the proportion of items above the recall level has risen, which is one concrete sign that the curriculum's shift toward application has reached its assessments.
References
Haladyna, T. M., Downing, S. M., & Rodriguez, M. C. (2002). A review of multiple-choice item-writing guidelines for classroom assessment. Applied Measurement in Education, 15(3), 309-333. https://doi.org/10.1207/S15324818AME1503_5
Tarrant, M., & Ware, J. (2008). Impact of item-writing flaws in multiple-choice questions on student achievement in high-stakes nursing assessments. Medical Education, 42(2), 198-206. https://doi.org/10.1111/j.1365-2923.2007.02957.x
Tarrant, M., Knierim, A., Hayes, S. K., & Ware, J. (2006). The frequency of item writing flaws in multiple-choice questions used in high stakes nursing assessments. Nurse Education Today, 26(8), 662-671. https://doi.org/10.1016/j.nedt.2006.07.006
How this NUR 5233 Module 4 example is structured
NUR 5233 Module 4 in many sections writes test items or a rubric and reads an item analysis; your classroom's instructions decide whether a practice data set is provided. This example states item-writing guidelines with evidence on common flaws, presents sample items with their rationale, describes a rubric with distinct levels, reports difficulty and discrimination for problem items with the calculation shown, and makes a decision for each.
NUR5233 Module 4 questions, answered
What does NUR5233 Module 4 usually ask for?
NUR5233 Module 4 in many sections asks you to write test items or a rubric aligned to course outcomes and to interpret an item analysis, deciding which items to keep, revise or discard. Your classroom's instructions decide whether a data set is provided.
What do difficulty and discrimination mean?
Difficulty is the proportion of students who answered an item correctly. Discrimination shows whether stronger students answered it correctly more often than weaker students; a negative value usually means the item is flawed.
How do I write items above the recall level?
Place the student in a clinical situation and ask what they would expect, monitor or do, so that answering requires applying knowledge rather than remembering a fact.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.