| Course | HLTH 5013 Epidemiology and Statistics |
|---|---|
| Module | Module 4 |
| Paper type | Descriptive and inferential data analysis |
| Length | 1,170 words, about 4 pages plus title and reference pages |
| Format | APA 7 student paper |
| School | American College of Education |
| Program | Master of Public Health |
| Updated | September 2026 |
Free sample paper for HLTH 5013 Module 4
Older, Heavier and Lower Starting A1c: A Descriptive and Inferential Analysis of 2,400 Adults in a County Diabetes Prevention Data Set
Student Name
American College of Education
HLTH5013: Epidemiology and Statistics
Module 4 Assignment
Instructor Name
October 26, 2026
Purpose and Data
The previous module found that adults in a composite county who enrolled in a lifestyle program for prediabetes developed diabetes less often over three years, and that part of the difference reflected enrollees' lower starting A1c. This module analyzes the full data set of 2,400 adults, adding age, body mass index, sex and insurance, to describe the two groups and to estimate the program's association with diabetes after accounting for several factors at once. The outcome is the same as before, diabetes by diagnosis or by laboratory value over the three years of follow-up; 347 adults, 14.5%, reached it.
Body mass index was missing for 142 adults, about 6%. The main analysis uses the 2,258 with complete data, and a sensitivity analysis described below checks whether the missing records change the result.
Describing the Groups
Descriptive statistics come first because they show whether the groups being compared are alike. The 620 enrolled adults had a mean age of 56.2 years with a standard deviation of 10.8, compared with 51.9 years and 12.4 among the 1,780 not enrolled. Mean body mass index was 33.1 among enrollees and 31.4 among non-enrollees. Mean baseline A1c was 5.96% and 6.03%. Women made up 71% of enrollees and 54% of non-enrollees, and 22% of enrollees had Medicaid or no insurance, compared with 38% of non-enrollees.
The profile is informative. Enrollees were older and heavier, which would raise their diabetes risk, but they had lower starting A1c and were more often women and privately insured, which may lower it. The groups differ in ways that pull the comparison in opposite directions, which is exactly why a single crude comparison cannot be trusted.
Testing Baseline Differences
Two-sample t-tests compared continuous variables. The difference in mean baseline A1c, 0.07 percentage points, has a standard error of about 0.009 using the unequal-variance formula, giving a t statistic near 7.8 and a p-value well below .001. The age difference of 4.3 years gives a t statistic of about 8.2. Chi-square tests showed that sex and insurance also differed between groups, each with a p-value below .001.
These p-values need careful reading. With groups this large, even small differences produce very small p-values. A difference of 0.07 in A1c is statistically clear but small clinically; what matters for the analysis is that it is consistent in direction and that A1c strongly predicts the outcome, so even a small baseline difference can confound the comparison. Greenland et al. (2016) list common misinterpretations of such tests, including treating a small p-value as a measure of the size or importance of an effect, which it is not.
Multivariable Logistic Regression
To estimate the program's association with diabetes while accounting for several differences at once, a logistic regression modeled the odds of diabetes as a function of enrollment, age in ten-year units, body mass index in five-unit steps, baseline A1c in tenths of a percentage point, sex and insurance. With 347 outcome events and six predictors, the model has well over the ten events per predictor commonly recommended for stable estimates.
The adjusted odds ratio for enrollment was 0.64, with a 95% confidence interval of 0.47 to 0.87. Each additional 0.1 in baseline A1c was associated with higher odds of diabetes, an odds ratio of 1.38 with an interval of 1.29 to 1.48, and each five-unit increase in body mass index with an odds ratio of 1.24, 1.12 to 1.37. Medicaid or no insurance had an odds ratio of 1.29, 1.00 to 1.66. Age, with an odds ratio of 1.08 and an interval of 0.97 to 1.20, and female sex, 0.91 with 0.71 to 1.17, had intervals that included 1.0.
Checking the Model
A regression result is only as good as the model behind it, so four checks were run before interpreting the estimates. First, continuous predictors were entered in grouped categories as well as linearly; the pattern for A1c rose steadily across tenths, supporting the linear term, while body mass index showed a steeper rise above 35, which the five-unit coding captures reasonably. Second, an interaction term between enrollment and baseline A1c tested whether the program's association differed by starting A1c; it did not clearly differ, with a p-value of about .41, consistent with the similar stratum ratios in the previous module. Third, the model's discrimination, its ability to separate people who developed diabetes from those who did not, was moderate, with a c-statistic of about 0.71. Fourth, predicted and observed risks were compared across tenths of predicted risk and agreed closely, indicating acceptable calibration. None of these checks revealed a problem that would change the interpretation of the enrollment estimate, and all four are reported so that a reader can judge the model rather than take it on trust.
Interpreting the Results
Because diabetes developed in about 16% of non-enrollees, the odds ratio overstates the relative risk. Converting it with the baseline risk gives an approximate adjusted relative risk of 0.68, close to the Mantel-Haenszel estimate of 0.69 from the previous module. Two different adjustment approaches, one stratifying on A1c alone and one modeling five covariates, arrive at nearly the same answer, which increases confidence that the measured confounders have been handled.
Two cautions apply. The interval for age includes 1.0, but Altman and Bland (1995) warn that failing to detect an effect does not show that none exists; the data are compatible with a modest effect of age as well as none. And p-values are better read as graded evidence than as a pass or fail at .05 (Sterne & Davey Smith, 2001), which is why this paper reports intervals throughout. The insurance estimate, whose lower bound sits at 1.00, is best described as suggestive rather than dismissed or overstated.
Sensitivity Analysis and Limits
Excluding the 142 adults with missing body mass index could bias the result if they differed systematically. Refitting the model without body mass index on all 2,400 adults gave an enrollment odds ratio of 0.62, close to the main estimate, so the missing values do not appear to change the conclusion. The analysis still cannot account for unmeasured factors, such as motivation, diet and physical activity before enrollment, that may differ between people who choose the program and those who do not. It also treats enrollment as attending four sessions, a threshold chosen in advance; a dose-response analysis by sessions attended would be a useful next step.
Summary for the Health Department
Enrollees and non-enrollees differed at baseline in ways that matter, but after adjusting for age, body mass index, starting A1c, sex and insurance, enrollment remained associated with about 32% lower risk of developing diabetes within three years. The finding is consistent across two methods and robust to missing data, but it is an association from observational data. The department can reasonably continue and expand the program while using the randomized clinic rollout recommended earlier to confirm the effect.
References
Altman, D. G., & Bland, J. M. (1995). Absence of evidence is not evidence of absence. BMJ, 311(7003), 485. https://doi.org/10.1136/bmj.311.7003.485
Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3
Sterne, J. A. C., & Davey Smith, G. (2001). Sifting the evidence: What's wrong with significance tests? Another comment on the role of statistical methods. BMJ, 322(7280), 226-231. https://doi.org/10.1136/bmj.322.7280.226
Reading the HLTH 5013 Module 4 instructions
In many sections, the HLTH 5013 Module 4 prompt has you take a data set from raw variables to conclusions. Prompts typically ask you to describe the variables with appropriate statistics, such as means and standard deviations or percentages, compare groups with inferential tests, sometimes fit a regression model, and interpret every result in context. Some versions supply the data in a spreadsheet or statistical package; others describe a published study's data. Graders expect you to choose tests that fit the type of variable, report test statistics and confidence intervals as well as p-values, and discuss limitations such as missing data. Check the Canvas prompt for which software is allowed and whether output must be appended.
How the HLTH 5013 Module 4 example is put together
The sample starts by stating its purpose, variables, outcome and how missing data will be handled. Descriptive statistics for both groups come next, with an interpretation of how the differences could pull the comparison in opposite directions. Baseline differences are tested with t-tests and chi-square tests, and the paper explains why large samples yield tiny p-values for small differences. A multivariable logistic regression is specified, checked for events per predictor and reported for every variable. Interpretation converts the odds ratio for comparison with earlier results and reads borderline intervals with cited cautions. A sensitivity analysis and a plain summary close the paper.
HLTH 5013 Module 4 rubric: what full marks look like
Data analysis rubrics usually weigh appropriate choice of statistics, correct calculation or reporting, correct interpretation and attention to limitations. The choice criterion rewards tests matched to variable types and study questions. Reporting earns points when results include effect sizes, confidence intervals and test statistics, not p-values alone. Interpretation carries heavy weight, and graders deduct for common misreadings, such as treating a non-significant result as proof of no effect. Handling of missing data, model checks and sensitivity analyses earn additional credit. A clear summary for a public health audience is often part of the rubric. Methods citations and APA 7 finish the score.
HLTH 5013 Module 4 help: mistakes that cost points
Analysis papers often lose marks by pasting software output without explaining it. Another frequent error is choosing a test that does not fit the data, such as a t-test on a categorical variable. Students also treat p-values as the whole story, calling results with p above .05 negative and results below it important. Describe your groups before comparing them. Report intervals for every estimate. Say what you did about missing data. Explain each result in a sentence a health officer could use. If your data set covers injuries, vaccination or birth outcomes instead, a codebook or the file itself, plus your rubric, is enough for us to set up a Module 4 analysis of those variables. Mention which software your section uses so the output matches what your instructor expects to see.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.
More HLTH 5013 and Master of Public Health sample papers
- HLTH 5013 Module 1: Prevalence and Incidence Analysis
- HLTH 5013 Module 2: Study Design Comparison
- HLTH 5013 Module 3: Measures of Association Analysis
- HLTH 5013 Module 5: Epidemiologic Research Proposal
- HLTH 5053 Module 4: Risk Communication Plan
- HLTH 5043 Module 3: Evaluation Approach Comparison
- HLTH 5033 Module 4: Capital Expenditure Evaluation
- HLTH 5053 Module 5: Campaign Plan and Evaluation
HLTH 5013 Module 4 questions, answered
What does HLTH5013 Module 4 usually ask for?
In many sections, the fourth HLTH5013 module asks you to analyze a public health data set with descriptive statistics and inferential tests, such as t-tests, chi-square tests or regression, and to interpret the results correctly. The data set is set by your own section.
Why do large samples produce very small p-values?
Because standard errors shrink as samples grow, so even small differences become statistically detectable. The p-value does not measure how large or important a difference is.
What does a confidence interval that includes 1.0 mean for an odds ratio?
That the data are compatible with no association as well as with the other values in the interval. It does not prove there is no effect.
Where can I find a free HLTH 5013 Module 4 sample paper?
The complete Module 4 analysis is on this page: descriptive statistics, t-tests, chi-square tests and a logistic regression on a county diabetes prevention data set of 2,400 adults, with a sensitivity analysis.
How many outcome events does a logistic regression need?
A common rule of thumb is at least ten outcome events per predictor variable to produce stable estimates.