From $445 to $196 an Hour: What 208 Store-Weeks Say About Adding Sales Floor Staff Once Customer Traffic Is Held Constant
Student Name
American College of Education
STAT5003: Business Statistics: Data-Informed Decision-Making
Module 5 Assignment
Instructor Name
June 16, 2025
The Proposal Under Review
Each week, the store managers at Riverbend, the invented eight-store chain this course keeps returning to, set the number of sales floor hours they will staff. Its regional manager has proposed adding 40 staffed hours a week at every store, about $860 a week per store in loaded labor cost at $21.50 an hour, arguing from a chart that plots weekly large-item sales against staffed hours and shows a steep upward line. The question for this module is what an additional staffed hour is actually associated with, in dollars of weekly sales, once other influences are accounted for.
The data cover 26 weeks at each of the eight stores, 208 store-weeks in all. Each record gives the store's large-item sales for the week, the staffed sales floor hours and the door count, the number of customers entering, recorded by sensors at each entrance. Across the 208 weeks, sales averaged about $88,800, staffed hours about 239 and door counts about 3,570. Hours ranged from 133 in a quiet week at the smallest store to 397 in a busy week at the largest.
The Simple Regression
A simple regression of weekly sales on staffed hours reproduces the manager's chart. The fitted slope is $445.40 per hour, with a standard error of $11.50, and hours alone explain 88 percent of the variation in weekly sales. Read literally, each additional staffed hour goes with $445 more in weekly sales, and at a gross margin of 38 percent, that would be about $169 of gross profit for $21.50 of labor. On those figures, the proposal looks overwhelming.
The problem is what else moves with staffed hours. Store managers schedule more staff when they expect more customers, so hours and door counts rise and fall together; their correlation across the 208 weeks is 0.97. Door counts also drive sales directly, since more customers buy more. The simple regression therefore credits staffing with sales that traffic produced. A regression answers the question it is asked, and this one was asked the wrong question.
Holding Traffic Constant
Adding door count as a second predictor changes the picture. In the multiple regression, the coefficient on staffed hours falls to $195.85, with a standard error of $46.44, and the coefficient on door count is $9.63 per customer, with a standard error of $1.74. The model explains 89.5 percent of the variation, only slightly more than hours alone, and the residual standard error is about $9,040 a week.
In dollars, the coefficients say this. Among weeks with the same number of customers coming through the door, a week with one more staffed hour had, on average, about $196 more in large-item sales. Among weeks with the same staffing, each additional customer went with about $9.63 in sales. The 95 percent confidence interval for the staffing coefficient runs from about $104 to $287 per hour, much wider than the simple regression's interval, because hours and traffic move so closely together that the data contain little independent variation in staffing to learn from. Angrist and Pischke (2009) set out the omitted variable bias formula that explains the drop: the simple slope equals the true staffing effect plus the effect of the omitted variable multiplied by how strongly it moves with the included one.
Checking the Conditions
Gelman and Hill (2007) rank the assumptions of a regression by importance, with the validity of the data for the question and the linearity of the relationship ahead of the distribution of the errors, and the checks below follow that order. Four conditions were checked. Linearity: plots of residuals against each predictor showed no curve, though the data contain few weeks with very high staffing at small stores, so the linear fit is least tested where the proposal would push hardest. Equal spread: residual standard deviations were about $8,600, $9,600 and $9,000 for the lowest, middle and highest thirds of door counts, similar enough that the standard errors are trustworthy. Normality: a histogram of residuals was roughly symmetric with no extreme values. Independence: weeks at the same store are not fully independent, since each store has its own habits, and standard errors that allow for this clustering would likely be somewhat larger than those reported here, making the interval wider still.
The strong correlation between hours and door counts, 0.97, is not a violation, but it explains why the staffing coefficient is imprecise. Collinearity leaves the estimate unbiased if the model is right but makes it hard to pin down.
What the Model Cannot Show
Even the multiple regression is not a causal estimate. Freedman (1991) argued that regression on observational data rarely settles causal questions by itself, and that the careful collection of the right data, the shoe leather of his title, usually matters more than a more elaborate model. Here, managers also schedule extra staff for promotions, holiday weekends and deliveries of new stock, and those events raise sales through channels the door count does not capture. If they do, even $196 overstates what an added hour would earn in an ordinary week.
A second concern is that the model pools differences between stores with differences between weeks. The largest store has both the most staff and the most sales in almost every week, partly for reasons unrelated to staffing, such as its location beside a highway exit and its larger showroom. A model that compares weeks within each store, by including a separate intercept for every store, would ask a narrower and more relevant question: when a given store staffs more heavily than usual, do its sales rise more than its traffic would predict? That model is the natural next step, but with hours and traffic moving so closely together inside each store, it would likely produce an even wider interval, which is itself an argument for collecting better data rather than fitting more models.
The Decision
At the point estimate, an added hour would bring about $74 of gross profit for $21.50 of labor, and at the low end of the interval about $40, so the proposal may well pay. But the evidence does not justify adding 40 hours at every store at once. At the two smallest stores, 40 added hours would push staffing above anything in the data for their traffic levels, and the estimate says nothing reliable about that range. The better course is a staged test: add 40 hours in alternating weeks at four stores chosen at random, hold everything else as usual for eight weeks and compare sales in added and ordinary weeks. Random assignment of the extra hours breaks the link between staffing and anticipated demand, which is the one thing this regression cannot do, and the cost is about $13,800 in labor, four stores times 40 hours times four added weeks, a small price for an answer the chain could then apply everywhere.
References
Angrist, J. D., & Pischke, J.-S. (2009). Mostly harmless econometrics: An empiricist's companion. Princeton University Press.
Freedman, D. A. (1991). Statistical models and shoe leather. Sociological Methodology, 21, 291-313. https://doi.org/10.2307/270939
Gelman, A., & Hill, J. (2007). Data analysis using regression and multilevel/hierarchical models. Cambridge University Press.
How this STAT 5003 Module 5 example is structured
STAT 5003 Module 5 typically compares groups or fits a regression against a business question; your classroom's instructions decide which. This example states the question and data, fits the simple regression managers already rely on and then adds the variable that explains why it misleads. Coefficients are interpreted in dollars, conditions are checked with the evidence stated, and the paper separates what the model shows from what it cannot show before turning the result into a decision.
STAT5003 Module 5 questions, answered
What does STAT5003 Module 5 usually ask for?
STAT5003 Module 5 often asks students to compare groups or fit a regression to answer a business question, with coefficients interpreted in business terms and conditions checked. Your classroom's instructions decide which method and data to use.
Why did my regression coefficient shrink when I added a variable?
The first model was probably crediting your predictor with the effect of something that moves with it. Adding that variable separates the two effects, which usually gives a smaller and more honest estimate.
Can a regression on business data show cause and effect?
Rarely by itself. Observational data leave room for other explanations, so a regression is best used to size a relationship and a controlled test is used to confirm that changing one thing causes the other.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.