Between 27 and 35 Percent: Two Confidence Intervals From an Audit of 400 Returns, and What They Are Worth to a $60,000 Decision
Student Name
American College of Education
STAT5003: Business Statistics: Data-Informed Decision-Making
Module 3 Assignment
Instructor Name
May 26, 2025
The Sample
Module 2 showed that monthly return rates at Riverbend Home Supply, the composite retailer used in this course, cannot identify which stores have a problem, and recommended an audit of why items are returned. From the 1,700 large-item returns recorded over the last six months, 400 were drawn by random selection of return numbers. A trained reviewer read each return's notes, photographs and delivery record and assigned one primary cause. For every return classified as delivery damage, the reviewer also recorded its cost to Riverbend: the second truck trip, the crew time and any markdown or write-off on the damaged item.
Random selection matters here for a specific reason. The easiest returns to audit are the most recent ones, whose paperwork is still on the loading dock, but the last six weeks included a spring sale on outdoor furniture, which is bulky and damaged more often than most items. A sample weighted toward recent returns would have overstated damage. Drawing return numbers at random from the full six months gives every return the same chance of selection and lets the result describe the period as a whole.
Of the 400 returns, 124 were classified as delivery damage, 31.0 percent of the sample. The costs of those 124 returns had a mean of $185.50, a median of $143.50 and a standard deviation of $147.70, ranging from $20 for a scratched table that was repaired on site to $788 for a refrigerator written off. The operations team is considering a $60,000 a year program of better packaging and crew training that its vendor says would halve delivery damage. Two things need estimating: how many returns are caused by damage and how much each costs.
The Share of Returns Caused by Damage
A confidence interval for a proportion is often built by adding and subtracting 1.96 standard errors from the sample share, but Agresti and Coull (1998) showed that this simple method covers the true value less often than it claims, and that adding two successes and two failures before computing the interval gives coverage much closer to the stated level. Their adjusted interval is used here. Because the sample is large relative to the six-month total, 400 of 1,700 returns, the standard error is also multiplied by a finite population correction of about 0.875, which reflects that sampling almost a quarter of all returns leaves less uncertainty than sampling from an unlimited population.
The resulting 95 percent confidence interval for the damage share runs from 27.2 to 35.1 percent. In plain language: the audit shows that about three in ten large-item returns come from delivery damage, and the true share over this period is very likely between about 27 and 35 percent. A manager can act on a range of 27 to 35 percent; a manager should not act on 31 percent as if it were known exactly. Without the correction, the interval would run from 26.5 to 35.5 percent, slightly wider but leading to the same conclusion.
The Cost of Each Damage Return
For the mean cost, the usual interval uses the t distribution with 123 degrees of freedom: the sample mean plus or minus 1.979 standard errors, where the standard error is the standard deviation divided by the square root of 124, about $13.26. That gives a 95 percent interval of $159.30 to $211.80. The t interval assumes the sampling distribution of the mean is roughly normal. The costs themselves are clearly skewed, with the mean well above the median and a long tail of written-off appliances, so the condition was checked rather than assumed.
The check used a bootstrap, the method Efron and Tibshirani (1993) developed for estimating uncertainty directly from the data. The 124 costs were resampled with replacement 10,000 times, the mean was computed for each resample and the middle 95 percent of those means was taken as the interval. The bootstrap interval ran from $160.90 to $212.20, within about two dollars of the t interval at each end. With 124 observations, the mean's sampling distribution is close enough to normal for the t interval to be trusted despite the skew. In plain language, each damage return costs Riverbend, on average, somewhere between about $160 and $212.
From Intervals to the Size of the Problem
Cumming (2014) argued that estimates with intervals, rather than yes-or-no tests, should be the basis of most conclusions, because they show both the size of an effect and the precision with which it is known. Here, the useful quantity is the annual cost of damage returns. Riverbend handles about 3,400 large-item returns a year. At the point estimates, about 1,054 would be damage returns, costing about $196,000 a year. Using the low ends of both intervals gives about 925 damage returns at $159 each, about $147,000; using the high ends gives about 1,193 at $212, about $253,000.
That range is deliberately conservative, since both ends are unlikely to be reached together, but it is the right range for a decision that should hold even in the worst reasonable case. The annual cost of delivery damage to Riverbend is very likely between about $147,000 and $253,000, and most likely near $196,000.
Does the Program Pay?
If the program halves delivery damage as the vendor claims, it would save between about $74,000 and $126,000 a year against its $60,000 cost, so it pays for itself across the whole range of the estimate. The vendor's claim is itself uncertain, so it is more useful to ask how much the program must reduce damage to break even. At the point estimate, it must cut damage costs by about 31 percent; at the low end of the range, by about 41 percent. A program that cannot plausibly cut damage by two-fifths is a risk; one that can is a sound investment on this evidence.
The recommendation is to run the program at the South distribution center first, which Module 1 found handles the more difficult deliveries, and to repeat the audit on a fresh sample of 400 returns after six months. A second interval that does not overlap the first would show the program working, and one that overlaps heavily would show it has not yet worked.
Limits
The intervals describe the past six months, and seasonal products or a change in carriers could shift both the share and the cost. The classification depends on one reviewer's judgment; a second reviewer checking a random 50 of the 400 would show how consistent the coding is. And the costs exclude any lost future sales from customers whose deliveries arrived damaged, so the true cost of damage is probably higher than these intervals suggest.
None of these limits changes the direction of the recommendation. Each either leaves the estimates roughly where they are or suggests that damage costs Riverbend more than the intervals show, which would make the prevention program more attractive, not less.
References
Agresti, A., & Coull, B. A. (1998). Approximate is better than "exact" for interval estimation of binomial proportions. The American Statistician, 52(2), 119-126. https://doi.org/10.1080/00031305.1998.10480550
Cumming, G. (2014). The new statistics: Why and how. Psychological Science, 25(1), 7-29. https://doi.org/10.1177/0956797613504966
Efron, B., & Tibshirani, R. J. (1993). An introduction to the bootstrap. Chapman & Hall.
How this STAT 5003 Module 3 example is structured
STAT 5003 Module 3 in many sections builds confidence intervals and reads them in plain language; your classroom's instructions decide which intervals and software to use. This example describes the sample and how it was drawn, then builds each interval with the method suited to its data and states the conditions checked. Each interval is read in words a manager could repeat, and the final sections combine the intervals to bound a business quantity and show that the decision holds across the whole range.
STAT5003 Module 3 questions, answered
What does STAT5003 Module 3 usually ask for?
STAT5003 Module 3 often asks students to build confidence intervals from business data and explain them in plain language for a manager. Many sections include both a proportion and a mean. Your classroom's instructions decide the data and software.
How do I explain a confidence interval to a manager?
Say what was estimated, give the range and say that the true value is very likely inside it. Then say what the range means for the decision, which usually matters more than the interval itself.
What if my data are skewed?
Check whether the interval method's conditions still hold. With a reasonably large sample, the t interval for a mean often remains accurate; a bootstrap interval is a simple way to confirm it.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.