NUR5233 Module 4 test items, rubric and item analysis example

Reviewed by Junia Fairbank, MSN, RN · American College of Education · True APA form, annotated

This page holds a complete NUR 5233 Module 4 example in true APA form: a test item, rubric and item analysis paper for American College of Education's Curriculum Development, Assessment, and Evaluation in Nursing course. For the heart failure unit designed in Module 3, it explains the item-writing rules the faculty adopted, presents two application-level items and the rubric for the written mechanism explanation, reads the item analysis of the week's ten-item quiz taken by 64 accelerated BSN students, and decides what to keep, revise or discard.

1

Predict, Don't Name: Writing Application-Level Items and a Mechanism Rubric for the Heart Failure Unit, Then Reading the Item Analysis From 64 Accelerated Students

Student Name

American College of Education

NUR5233: Curriculum Development, Assessment, and Evaluation in Nursing

Module 4 Assignment

Instructor Name

August 1, 2028

What this page is doingThe title states the principle behind the items, asking students to predict rather than recall, and names the three products of the paper. The APA 7 title page carries the course line and the module assignment as listed.
2

Why Item Writing Matters Here

The heart failure unit in Module 3 includes a weekly quiz worth 30% of the unit grade, and its outcome asks students to explain drug mechanisms and predict what nurses must monitor. If the quiz items ask only for drug names or classes, the quiz will reward recall and undo the alignment the unit was designed for. The problem is common. In one nursing department's review of 2,770 multiple-choice questions used over five years, 46.2% contained at least one violation of accepted item-writing guidelines, and more than 90% were written at low cognitive levels; questions at lower cognitive levels were more likely to be flawed (Tarrant et al., 2006). A later study of high-stakes nursing examinations found that flawed items were common and that they tended to penalize high-achieving students more than borderline ones (Tarrant & Ware, 2008). Flawed items, in other words, distort the measure for the students the program most wants to identify.

What this page is doingThe paper establishes with evidence why item quality matters for this unit specifically, connecting item writing to the alignment designed in Module 3.
3

The Rules the Faculty Adopted

The faculty adopted a short set of rules drawn from a widely cited review of multiple-choice item-writing guidelines (Haladyna et al., 2002). Each item tests one important idea from the unit outcomes. The stem poses a complete question that could be answered without seeing the options. Options are plausible, similar in length and grammar, and free of clues such as absolute words or repetition of stem language in the correct answer. None of the above and all of the above are not used. Three or four options are enough; a distractor no student would choose adds reading time without adding information. And, for this unit, every item places the student in a clinical situation and asks what the student would expect, monitor or do, so that recall alone is not enough to answer it.

What this page is doingThe rules are stated concretely and attributed to an authoritative review, with one rule added specifically to serve the unit's outcome.
4

Two Sample Items

Item 3. A patient with chronic heart failure is started on furosemide 40 mg orally each morning. Which finding at the next day's assessment would indicate an expected effect of the drug? A. Weight 1.2 kg less than the previous morning. B. Blood pressure 12 mm Hg higher than yesterday. C. Serum potassium 0.6 mmol/L higher than yesterday. D. Heart rate 20 beats per minute slower than yesterday. The correct answer is A. The item requires the student to reason from the mechanism, a loop diuretic causing loss of sodium and water, to its measurable effect, weight loss from fluid removal. The distractors each represent a plausible misunderstanding: that the drug raises blood pressure, that it retains potassium, confusing it with a potassium-sparing diuretic, or that it slows the heart, confusing it with a beta-blocker.

Item 8, as originally written. A patient with heart failure is started on lisinopril. Which assessment is the nurse's priority before the second dose? A. Blood pressure. B. Serum potassium. C. Presence of a dry cough. D. Serum sodium. The item analysis below shows what went wrong with it.

What this page is doingSample items are presented in full with their rationale and the misunderstanding each distractor targets, showing the rules applied. Including a flawed original item sets up the analysis that follows.
5

The Mechanism Rubric

The written task asks students to account for two familiar features of heart failure, weight gain and breathlessness on lying down, and is scored with a four-level rubric whose levels differ by the quality of the causal chain, not by adjectives. At level four, the student traces the full chain: reduced cardiac output leads to reduced kidney perfusion, activation of the renin-angiotensin-aldosterone system and retention of sodium and water, raising venous pressure; lying flat increases venous return to an already overloaded heart, raising pressure in the pulmonary circulation and producing fluid in the lungs. At level three, the chain is correct but one link is missing, most often the reason lying flat worsens symptoms. At level two, the student names the right mechanisms but does not connect them causally. At level one, the student lists signs without a mechanism. Two faculty independently scored the same 20 responses before the rubric was used, agreed on 17, and revised the wording of level three to resolve the disagreements.

What this page is doingThe rubric's levels are distinguished by the specific quality of reasoning, and its reliability was checked before use, which is what separates a usable rubric from a list of descriptors.
6

Reading the Item Analysis

The quiz was taken by 64 students. Difficulty was calculated as the proportion of students answering each item correctly, and discrimination as the difference between the proportion correct in the top 27% of scorers, 17 students, and the bottom 27%, another 17. Most items performed well. Four deserve comment.

Item 3 had a difficulty of 0.72 and a discrimination index of 0.47, with 16 of the top 17 and 8 of the bottom 17 answering correctly: a moderately difficult item that separates stronger from weaker students well. It will be kept.

Item 6, which asked which drug class metoprolol belongs to, had a difficulty of 0.95 and a discrimination index of 0.06. Nearly everyone answered correctly, so the item gave almost no information, and it tested recall, which the unit's outcome does not target. It will be rewritten to ask what the nurse should assess before giving the drug.

Item 8 had a difficulty of 0.31 and a negative discrimination index of minus 0.06: 5 of the top 17 answered correctly, compared with 6 of the bottom 17. Forty-one percent of all students, including 11 of the top group, chose potassium rather than the keyed answer, blood pressure. When the strongest students choose a distractor more often than the weakest, the problem is usually the item, not the students. On review, both answers are defensible, since an angiotensin-converting enzyme inhibitor can cause both first-dose hypotension and a rise in potassium. The item will be credited for both answers on this quiz and rewritten with a scenario that makes one priority clear.

Item 9 had acceptable difficulty and discrimination, 0.58 and 0.35, but one distractor was chosen by no student. That distractor will be replaced with one based on a common misconception, drawn from the wrong answers students gave in the written explanations.

What this page is doingThe analysis method is stated, the calculations are shown for each problem item, and each finding leads to a specific decision. Recognizing the negatively discriminating item as a flawed key is the key interpretive skill the module assesses.
7

Limits of a Ten-Item Quiz

A ten-item quiz is short, and its reliability as a total score is modest; its purpose is formative feedback within the week, not a high-stakes decision. Item statistics from 64 students are also unstable, so the faculty will look for patterns across two cohorts before discarding items that perform poorly once. The larger lesson for the curriculum is that item analysis should be a routine step after every quiz and examination, reviewed by the course team before grades are released, so that flawed items are caught before they affect students. The course team will keep a shared bank of items with their statistics from each administration, marked by outcome and cognitive level, so that the next cohort's quizzes are built from items with known performance and new items are piloted alongside them. Over two or three years, the bank will also show whether the proportion of items above the recall level has risen, which is one concrete sign that the curriculum's shift toward application has reached its assessments.

What this page is doingThe paper states the limits of the data and turns the exercise into a curriculum policy, linking to the evaluation plan in the next module.
8

References

Haladyna, T. M., Downing, S. M., & Rodriguez, M. C. (2002). A review of multiple-choice item-writing guidelines for classroom assessment. Applied Measurement in Education, 15(3), 309-333. https://doi.org/10.1207/S15324818AME1503_5

Tarrant, M., & Ware, J. (2008). Impact of item-writing flaws in multiple-choice questions on student achievement in high-stakes nursing assessments. Medical Education, 42(2), 198-206. https://doi.org/10.1111/j.1365-2923.2007.02957.x

Tarrant, M., Knierim, A., Hayes, S. K., & Ware, J. (2006). The frequency of item writing flaws in multiple-choice questions used in high stakes nursing assessments. Nurse Education Today, 26(8), 662-671. https://doi.org/10.1016/j.nedt.2006.07.006

How this NUR 5233 Module 4 example is structured

NUR 5233 Module 4 in many sections writes test items or a rubric and reads an item analysis; your classroom's instructions decide whether a practice data set is provided. This example states item-writing guidelines with evidence on common flaws, presents sample items with their rationale, describes a rubric with distinct levels, reports difficulty and discrimination for problem items with the calculation shown, and makes a decision for each.

NUR5233 Module 4 questions, answered

What does NUR5233 Module 4 usually ask for?

NUR5233 Module 4 in many sections asks you to write test items or a rubric aligned to course outcomes and to interpret an item analysis, deciding which items to keep, revise or discard. Your classroom's instructions decide whether a data set is provided.

What do difficulty and discrimination mean?

Difficulty is the proportion of students who answered an item correctly. Discrimination shows whether stronger students answered it correctly more often than weaker students; a negative value usually means the item is flawed.

How do I write items above the recall level?

Place the student in a clinical situation and ask what they would expect, monitor or do, so that answering requires applying knowledge rather than remembering a fact.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.