STAT5003 Module 2 probability and sampling analysis example

Reviewed by Cornelius Ravenhill, MBA · American College of Education · True APA form, annotated

This page holds a complete STAT 5003 Module 2 example in true APA form: a probability and sampling analysis for American College of Education's Business Statistics: Data-Informed Decision-Making course. The composite home goods retailer's regional manager ranks its eight stores each month by the share of large-item orders returned, praising the lowest and questioning the highest. Last month the best rate, 4.1 percent, and the worst, 9.3 percent, both came from the two smallest stores. Using the binomial model, standard errors and the probability of results at least that extreme, the paper shows both rates are well within what chance produces at those volumes, and it designs a random sample to find out why items are actually returned.

1

The Best Store and the Worst Store Were Both the Smallest: Probability, Sample Size and Eight Monthly Return Rates

Student Name

American College of Education

STAT5003: Business Statistics: Data-Informed Decision-Making

Module 2 Assignment

Instructor Name

May 19, 2025

What this page is doingThe title states the pattern that gives the analysis away, extremes in the smallest stores, and names the tools that explain it. The retailer, its stores and every figure are composites. The APA 7 title page carries the course line and module assignment as listed.
2

The Number Being Quoted

Riverbend Home Supply, the composite home goods retailer used in this course, tracks the share of large-item orders that customers return within 30 days. Returns are expensive, since each one requires a truck, a crew and often a markdown, so the regional manager ranks the eight stores each month by return rate and discusses the results on a call with store managers. Last month, 285 of 4,600 large-item orders were returned, a company rate of 6.2 percent. Wexford had the lowest rate, 4.1 percent, and was praised; Stillwater had the highest, 9.3 percent, and its manager was asked for an improvement plan.

Wexford and Stillwater are also the two smallest stores, with 220 and 150 large-item orders that month, against 1,100 at the largest. That pattern is the starting point of this analysis. When the best and worst performers are both the smallest units, the first suspect is not management but arithmetic.

3

A Probability Model for Returns

A simple model treats each order as an independent trial with the same probability of being returned. If every store shared the company rate of 6.2 percent, the number of returns at a store with n orders would follow a binomial distribution, and its return rate would have a standard error equal to the square root of p times the quantity one minus p, divided by n, where p is 0.062 and n is the store's number of orders. The formula matters less than what it implies, which is that the uncertainty in a rate depends on how many orders produced it. The model's assumptions are worth stating. It assumes orders are independent, which is reasonable for different customers, and that every store has the same underlying rate, which is exactly the claim being tested. It also assumes the store's product mix is similar, which is roughly true across Riverbend's stores.

The standard error shrinks as volume grows, but only with the square root. At Lakeside, with 1,100 orders, the standard error is about 0.73 percentage points, so chance alone would usually keep the monthly rate between about 4.7 and 7.7 percent, two standard errors either side of 6.2. At Stillwater, with 150 orders, the standard error is about 1.97 points, and the usual range runs from about 2.3 to 10.1 percent. The same process, with no difference in management at all, produces a range more than two and a half times as wide at the smallest store as at the largest.

What this page is doingThe model is stated with its assumptions before it is used, and the standard error is translated into a range a manager can picture at each store.
4

Testing the Two Extremes

Stillwater's 14 returns on 150 orders is a rate of 9.3 percent. Under the binomial model with p equal to 0.062, the probability of 14 or more returns in 150 orders is about 0.08. Wexford's 9 returns on 220 orders is a rate of 4.1 percent, and the probability of 9 or fewer is about 0.12. Neither result is unusual for a single store in a single month; each would occur by chance roughly once a year at a store of that size, even if its managers did nothing differently.

The ranking makes the problem worse, because it looks across eight stores at once. If every store truly shared the company rate, the chance that at least one of the eight would show a result as far into the upper tail as Stillwater's, a probability of 0.08 or less, is about one in two, treating the stores as independent. Some store will look bad almost every other month by chance. Wainer (2007) called the formula for the standard error of a mean the most dangerous equation, because ignoring it leads decision makers to see talent or failure in what is only the greater variability of small units, and he documented the error in settings from school rankings to cancer rates. Tversky and Kahneman (1971) found that even trained researchers expect small samples to resemble the population much more closely than they do, a belief they called the law of small numbers.

5

What the Ranking Is Doing

Deming (1986) distinguished variation from common causes, built into a stable process, from variation with special causes that can be traced and fixed, and warned that treating common-cause variation as if it had a special cause leads managers to blame and reward people for outcomes they did not control. The monthly ranking does exactly that. Of the eight stores last month, all eight fell inside the two-standard-error range for their size. The manager at Stillwater has been asked to explain a number that the model says could easily be chance, and the manager at Wexford may draw lessons from a good month that will not repeat.

None of this proves that the stores are identical. A store could have a real problem, such as a delivery crew that damages items, and still fall within the range in a given month. But a single month at a small store cannot separate a real problem from chance, and the ranking gives both the same weight.

6

Sampling to Find Out Why

If the rates cannot say which store has a problem, the reasons for returns might. Riverbend records a reason code, but store staff choose it quickly, and the most common code is simply customer changed mind. A better approach is to draw a random sample of returned orders from the last six months and have one trained reviewer read each order's notes, photographs and delivery record to classify the true reason. To estimate the share of returns caused by delivery damage within plus or minus 3 percentage points with 95 percent confidence, using the conservative assumption that the share is near one half, the sample needs about 1,068 returns. If the share is thought to be nearer 30 percent, about 897 are enough.

Riverbend had about 1,700 large-item returns in the last six months, so a sample of about 900 is feasible but large. A smaller sample of 400, giving a margin of about plus or minus 4.5 points at a 30 percent share, would still separate a problem affecting one return in three from one affecting one in ten. The sample should be drawn by random selection of return numbers across all stores, so that the result describes the company and, with enough returns per store, allows the largest stores to be compared.

7

What the Manager Should Do Differently

Three changes follow. First, stop ranking stores on one month's return rate. Report each store's rate against its own expected range, so that a store is flagged only when its rate falls outside what chance would produce at its volume. Second, judge small stores over longer periods: six months of orders at Stillwater, about 900, would narrow its standard error to about 0.8 points, similar to a large store in one month. Third, use the reason audit, not the rate, to decide where to act, since knowing that a store's returns come mostly from delivery damage says what to fix in a way that a rate never can.

8

References

Deming, W. E. (1986). Out of the crisis. MIT Press.

Tversky, A., & Kahneman, D. (1971). Belief in the law of small numbers. Psychological Bulletin, 76(2), 105-110. https://doi.org/10.1037/h0031322

Wainer, H. (2007). The most dangerous equation. American Scientist, 95(3), 249-256. https://doi.org/10.1511/2007.65.1026

How this STAT 5003 Module 2 example is structured

STAT 5003 Module 2 typically brings probability and sampling behind the numbers managers quote; your classroom's instructions decide which distributions and calculations to use. This example states the number managers are using, sets out a probability model and its assumptions, computes how much each store's rate would vary by chance alone and tests the two extreme stores against that model. It then turns to sampling, calculating the size of an audit sample from a stated margin of error, and ends with what the manager should do differently.

STAT5003 Module 2 questions, answered

What does STAT5003 Module 2 usually ask for?

STAT5003 Module 2 often asks students to apply probability and sampling ideas to numbers managers use, such as rates, averages or survey results. Many sections include a binomial or normal calculation and a sample size question. Your classroom's instructions decide the data and methods.

Why do small groups produce extreme rates?

The standard error of a rate falls with the square root of the number of cases, so small groups vary much more from month to month by chance alone. Rankings then place small groups at both the top and the bottom.

How do I calculate a sample size for estimating a proportion?

Use the required margin of error, the confidence level and a planning value for the proportion. Using one half gives the largest, most conservative sample; a better guess of the proportion usually reduces the size needed.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.