Does the Redesigned Checkout Raise Order Value? A Two-Sample Test of 812 Transactions at a Regional Home Goods Retailer
Student Name
American College of Education
STAT5003: Business Statistics: Data-Informed Decision-Making
Module 4 Assignment
Instructor Name
June 9, 2025
The Business Question and the Data Set
Riverbend Home Supply is a composite regional retailer written for teaching, so no real company, employee, or customer is described here. The firm sells kitchen and bath goods through eight stores and one website, and in March it launched a redesigned online checkout that removes the account creation step and shows shipping cost on the first screen. The merchandising director wants to know one thing before the design is rolled out to the mobile application: does the new checkout raise the average order value, or does it simply move the same money through a shorter path? The question is worth testing because the rollout carries a $46,000 development cost.
The data set is a random sample of 841 online orders drawn from the 9,431 orders placed between March 3 and May 30. Sampling was done by random selection of order numbers rather than by taking the most recent orders, because recency would load the sample toward the new design and toward the spring promotion. Each record carries the order total in dollars, the checkout version served, the device category, and whether the order included a promotional code. Of the 812 orders, 407 were served the legacy checkout and 405 the redesigned checkout. Twenty-nine records with returned or canceled totals were removed before analysis, leaving the 812 orders reported here, and the removal is stated rather than left silent.
The descriptive picture comes first because it sets expectations for the test. Legacy orders averaged $118.42 with a standard deviation of $61.30. Redesigned orders averaged $127.85 with a standard deviation of $64.71. The observed difference is $9.43, or about 8 percent of the legacy mean. Both distributions are right skewed, which is ordinary for order value, with a long tail of large furniture orders above $400 and a floor near $19. Median values sit below the means at $104.10 and $110.75. A difference of $9.43 on samples of about 400 each is small enough that it could easily be noise, which is the reason for a test rather than a report.
Hypotheses, Assumptions, and the Choice of Test
The null hypothesis is that the two checkout versions produce the same mean order value, and the alternative is that the redesigned version produces a higher mean. The test is one tailed because the firm will only act on an increase, and the significance level was fixed at 0.05 before the data were examined. Setting the direction and the level in advance matters here for a practical reason as much as a statistical one. An analyst who inspects the sample means first and then chooses a one-tailed test has borrowed significance from the data, and Black (2023) treats that sequence as a reporting error rather than a stylistic preference.
Four assumptions were checked rather than asserted. Independence holds by design, since orders were selected at random and each record is a separate transaction. Sample size covers the normality requirement, because with 407 and 405 observations the central limit theorem applies to the sampling distribution of the mean even though the underlying order values are skewed. Equal variance was not assumed, since the standard deviations differ and the group sizes are close but not identical, so the Welch version of the two-sample t test was used. Measurement consistency was verified by confirming that both versions record order value after discounts and before shipping.
The choice of test follows from those checks. A two-sample t test on means fits a continuous outcome compared across two independent groups, and the Welch adjustment removes the equal variance requirement at the cost of fractional degrees of freedom. A z test was rejected because the population standard deviation is unknown. A paired test was rejected because no customer appears in both groups. Comparing medians with a rank-based test was considered given the skew, and it remains a reasonable sensitivity check, but the mean is the quantity the firm budgets against, so the mean is the quantity tested. Sample size was checked against the effect the firm cares about as well: about 405 orders per group gives roughly 80 percent power to detect a difference of 0.18 standard deviations under Cohen's (1988) conventions, or about $11 per order.
Results, With the Arithmetic Shown
The standard error of the difference is the square root of 61.30 squared divided by 407 plus 64.71 squared divided by 405, which is the square root of 9.233 plus 10.339, or 4.424. The test statistic is the observed difference divided by that standard error, 9.43 divided by 4.424, giving t equal to 2.13. Welch degrees of freedom come to 806. The one-tailed p value is 0.017, below the 0.05 level set in advance, so the null hypothesis of equal means is rejected. The two-sided 95 percent confidence interval for the difference runs from $0.75 to $18.11, clearing zero by a margin thin enough to be worth saying out loud.
The interval deserves more attention than the p value. It says the data are consistent with an increase as small as $0.75 per order and as large as $18.11, a range wide enough to change the decision at one end and not at the other. Across roughly 37,700 online orders a year, the low end is worth about $28,000, which does not cover the $46,000 development cost, while the point estimate is worth about $355,000 and the upper bound about $683,000. Reporting only the significant result would hide that spread, and the spread is what the merchandising director is actually buying (Wasserstein & Lazar, 2016; Wasserstein et al., 2019).
A second test guards against the most likely confound. Promotional code use could differ between the two groups and could explain the gap on its own, so a chi-square test of independence was run on checkout version against promotional code use. Promotional codes appeared on 96 of 407 legacy orders and 103 of 405 redesigned orders. The chi-square statistic is 0.51 with 1 degree of freedom, p equal to 0.475, so promotional code use is not distinguishable between the groups and does not account for the difference in means. Device mix was compared the same way and showed no meaningful imbalance.
The Decision, and What This Test Cannot Say
The recommendation is a staged release rather than the full commitment. The evidence supports a real increase in average order value, but the interval does not establish that the increase pays for the $46,000 build, because its lower bound annualizes to roughly $28,000. Releasing the redesign to half of mobile traffic for 60 days costs a fraction of the full build, keeps a comparison group alive, and roughly doubles the sample behind the estimate, which is the pattern Kohavi et al. (2020) recommend after a positive but imprecise result. The decision rule was fixed before the data were seen and holds either way: commit to the full rollout once the lower bound of the interval clears an annualized $46,000, and hold at the staged release until it does.
Three limits belong in the same section as the recommendation, not in a footnote. The sample covers March through May, so seasonal demand and the spring promotion are inside the window and any estimate carried into a fourth quarter is an extrapolation. The comparison is observational rather than randomized, since assignment followed a release schedule instead of a coin flip, which means an unmeasured difference between the groups remains possible even though promotional code use and device mix were ruled out. Order value also says nothing about margin, and a shorter path to purchase may shift the product mix toward lower margin goods.
Two measures would close those gaps and both are already collected. Gross margin per order, tested the same way on the same sample, would tell the firm whether the extra $9.43 survives the cost of goods. Completed orders as a share of checkout starts would show whether the redesign also lifts conversion, which is the effect the design was meant to produce and the one this analysis did not test. Reporting the second study alongside the first keeps the finding honest, because a checkout that raises order value while lowering completion could easily reduce total revenue.
References
Black, K. (2023). Business statistics: For contemporary decision making (11th ed.). Wiley.
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press.
Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133.
Wasserstein, R. L., Schirm, A. L., & Lazar, N. A. (2019). Moving to a world beyond p < 0.05. The American Statistician, 73(sup1), 1-19.
How this STAT 5003 Module 4 example is structured
In many sections this STAT5003 Module 4 assignment in the Business Statistics: Data-Informed Decision-Making course asks you to test a question against real numbers and report what the test supports; your course instructions and rubric decide the exact form. The example follows the order an analyst works in. It opens with the business question and the data set, so a reader knows what was measured and how much of it there is. The second section states the hypotheses and the assumptions, because a test that hides its assumptions cannot be checked. The third section reports the arithmetic, the statistic, the p value, and the interval together. The last section turns the result into a decision, names what the study cannot claim, and says what would be measured next.
STAT5003 Module 4 questions, answered
What does STAT5003 Module 4 usually ask for?
American College of Education does not publish deliverable names module by module, so treat this as the common shape rather than a fixed name. In many sections a Module 4 assignment in Business Statistics asks you to apply a hypothesis test or a comparison of means to a business data set and report the decision it supports. Your course instructions and rubric decide the exact form.
Do I have to show the calculations, or is the software output enough?
Show enough that a reader can rebuild the number. Give the group means, the standard deviations, the sample sizes, the standard error, the statistic, the degrees of freedom, and the p value. Pasted output with no interpretation reads as unfinished work, and a stated result with no inputs cannot be checked. The example above writes the standard error out in words for that reason.
Why report a confidence interval when the p value is already significant?
Because the p value says only that the difference is unlikely to be zero, while the interval says how large the difference plausibly is. A gain of $0.75 per order and a gain of $18.11 per order lead to different business decisions, and the interval carries both. Reviewers of graduate statistics work consistently reward the writer who reports the size of an effect, not only its significance.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.