HLTH 5013 Module 4 Descriptive and Inferential Data Analysis Example

Reviewed by Cornelius Ravenhill, MBA · American College of Education · Updated

Here is a complete HLTH 5013 Module 4 data analysis, in APA 7, of 2,400 adults with prediabetes in a composite county, 620 of whom joined a lifestyle program. It was prepared for American College of Education HLTH 5013, Epidemiology and Statistics, the HLTH5013 course in ACE's Master of Public Health. Descriptive statistics show enrollees were older, heavier and more often women, with lower starting A1c. Unequal-variance t-tests of about 7.8 and 8.2 and chi-square tests confirm the gaps, read with Greenland's guide to misinterpreting p-values. A logistic regression with 347 events and six predictors gives an enrollment odds ratio of 0.64, an adjusted relative risk near 0.68 that matches the earlier Mantel-Haenszel figure. Altman and Bland and Sterne frame the borderline results, and a missing-data check holds. In many sections the data set is supplied.

CourseHLTH 5013 Epidemiology and Statistics
ModuleModule 4
Paper typeDescriptive and inferential data analysis
Length1,170 words, about 4 pages plus title and reference pages
FormatAPA 7 student paper
SchoolAmerican College of Education
ProgramMaster of Public Health
UpdatedSeptember 2026

Free sample paper for HLTH 5013 Module 4

1

Older, Heavier and Lower Starting A1c: A Descriptive and Inferential Analysis of 2,400 Adults in a County Diabetes Prevention Data Set

Student Name

American College of Education

HLTH5013: Epidemiology and Statistics

Module 4 Assignment

Instructor Name

October 26, 2026

What this page is doingThe title reports how the two groups differed at baseline, which tells the grader the analysis takes group comparability seriously before estimating any effect. The APA 7 title page carries the course line and the module assignment as listed.
2

Purpose and Data

The previous module found that adults in a composite county who enrolled in a lifestyle program for prediabetes developed diabetes less often over three years, and that part of the difference reflected enrollees' lower starting A1c. This module analyzes the full data set of 2,400 adults, adding age, body mass index, sex and insurance, to describe the two groups and to estimate the program's association with diabetes after accounting for several factors at once. The outcome is the same as before, diabetes by diagnosis or by laboratory value over the three years of follow-up; 347 adults, 14.5%, reached it.

Body mass index was missing for 142 adults, about 6%. The main analysis uses the 2,258 with complete data, and a sensitivity analysis described below checks whether the missing records change the result.

What this page is doingThe analysis purpose, variables, outcome and missing data are stated at the outset, including how missingness will be handled.
3

Describing the Groups

Descriptive statistics come first because they show whether the groups being compared are alike. The 620 enrolled adults had a mean age of 56.2 years with a standard deviation of 10.8, compared with 51.9 years and 12.4 among the 1,780 not enrolled. Mean body mass index was 33.1 among enrollees and 31.4 among non-enrollees. Mean baseline A1c was 5.96% and 6.03%. Women made up 71% of enrollees and 54% of non-enrollees, and 22% of enrollees had Medicaid or no insurance, compared with 38% of non-enrollees.

The profile is informative. Enrollees were older and heavier, which would raise their diabetes risk, but they had lower starting A1c and were more often women and privately insured, which may lower it. The groups differ in ways that pull the comparison in opposite directions, which is exactly why a single crude comparison cannot be trusted.

What this page is doingMeans with standard deviations and percentages describe both groups, and the paper interprets how the differences could bias the comparison in both directions.
4

Testing Baseline Differences

Two-sample t-tests compared continuous variables. The difference in mean baseline A1c, 0.07 percentage points, has a standard error of about 0.009 using the unequal-variance formula, giving a t statistic near 7.8 and a p-value well below .001. The age difference of 4.3 years gives a t statistic of about 8.2. Chi-square tests showed that sex and insurance also differed between groups, each with a p-value below .001.

These p-values need careful reading. With groups this large, even small differences produce very small p-values. A difference of 0.07 in A1c is statistically clear but small clinically; what matters for the analysis is that it is consistent in direction and that A1c strongly predicts the outcome, so even a small baseline difference can confound the comparison. Greenland et al. (2016) list common misinterpretations of such tests, including treating a small p-value as a measure of the size or importance of an effect, which it is not.

What this page is doingTest statistics are derived and reported, and the paper explains why large samples yield small p-values for small differences, citing a guide to misinterpretation.
5

Multivariable Logistic Regression

To estimate the program's association with diabetes while accounting for several differences at once, a logistic regression modeled the odds of diabetes as a function of enrollment, age in ten-year units, body mass index in five-unit steps, baseline A1c in tenths of a percentage point, sex and insurance. With 347 outcome events and six predictors, the model has well over the ten events per predictor commonly recommended for stable estimates.

The adjusted odds ratio for enrollment was 0.64, with a 95% confidence interval of 0.47 to 0.87. Each additional 0.1 in baseline A1c was associated with higher odds of diabetes, an odds ratio of 1.38 with an interval of 1.29 to 1.48, and each five-unit increase in body mass index with an odds ratio of 1.24, 1.12 to 1.37. Medicaid or no insurance had an odds ratio of 1.29, 1.00 to 1.66. Age, with an odds ratio of 1.08 and an interval of 0.97 to 1.20, and female sex, 0.91 with 0.71 to 1.17, had intervals that included 1.0.

What this page is doingThe model specification, a check on events per predictor and the adjusted estimates with intervals are reported clearly for every variable.
6

Checking the Model

A regression result is only as good as the model behind it, so four checks were run before interpreting the estimates. First, continuous predictors were entered in grouped categories as well as linearly; the pattern for A1c rose steadily across tenths, supporting the linear term, while body mass index showed a steeper rise above 35, which the five-unit coding captures reasonably. Second, an interaction term between enrollment and baseline A1c tested whether the program's association differed by starting A1c; it did not clearly differ, with a p-value of about .41, consistent with the similar stratum ratios in the previous module. Third, the model's discrimination, its ability to separate people who developed diabetes from those who did not, was moderate, with a c-statistic of about 0.71. Fourth, predicted and observed risks were compared across tenths of predicted risk and agreed closely, indicating acceptable calibration. None of these checks revealed a problem that would change the interpretation of the enrollment estimate, and all four are reported so that a reader can judge the model rather than take it on trust.

What this page is doingModel assumptions, interaction, discrimination and calibration are checked and reported, which supports the credibility of the regression estimates.
7

Interpreting the Results

Because diabetes developed in about 16% of non-enrollees, the odds ratio overstates the relative risk. Converting it with the baseline risk gives an approximate adjusted relative risk of 0.68, close to the Mantel-Haenszel estimate of 0.69 from the previous module. Two different adjustment approaches, one stratifying on A1c alone and one modeling five covariates, arrive at nearly the same answer, which increases confidence that the measured confounders have been handled.

Two cautions apply. The interval for age includes 1.0, but Altman and Bland (1995) warn that failing to detect an effect does not show that none exists; the data are compatible with a modest effect of age as well as none. And p-values are better read as graded evidence than as a pass or fail at .05 (Sterne & Davey Smith, 2001), which is why this paper reports intervals throughout. The insurance estimate, whose lower bound sits at 1.00, is best described as suggestive rather than dismissed or overstated.

What this page is doingResults are converted for interpretation, compared with the previous estimate, and read with cited cautions about non-significant results and dichotomous thinking.
8

Sensitivity Analysis and Limits

Excluding the 142 adults with missing body mass index could bias the result if they differed systematically. Refitting the model without body mass index on all 2,400 adults gave an enrollment odds ratio of 0.62, close to the main estimate, so the missing values do not appear to change the conclusion. The analysis still cannot account for unmeasured factors, such as motivation, diet and physical activity before enrollment, that may differ between people who choose the program and those who do not. It also treats enrollment as attending four sessions, a threshold chosen in advance; a dose-response analysis by sessions attended would be a useful next step.

What this page is doingA sensitivity analysis tests the effect of missing data, and remaining limitations, including unmeasured confounding and the exposure definition, are stated plainly.
9

Summary for the Health Department

Enrollees and non-enrollees differed at baseline in ways that matter, but after adjusting for age, body mass index, starting A1c, sex and insurance, enrollment remained associated with about 32% lower risk of developing diabetes within three years. The finding is consistent across two methods and robust to missing data, but it is an association from observational data. The department can reasonably continue and expand the program while using the randomized clinic rollout recommended earlier to confirm the effect.

What this page is doingThe summary states the adjusted finding in plain terms, notes its consistency and limits, and connects it to the department's decision.
10

References

Altman, D. G., & Bland, J. M. (1995). Absence of evidence is not evidence of absence. BMJ, 311(7003), 485. https://doi.org/10.1136/bmj.311.7003.485

Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3

Sterne, J. A. C., & Davey Smith, G. (2001). Sifting the evidence: What's wrong with significance tests? Another comment on the role of statistical methods. BMJ, 322(7280), 226-231. https://doi.org/10.1136/bmj.322.7280.226

Reading the HLTH 5013 Module 4 instructions

In many sections, the HLTH 5013 Module 4 prompt has you take a data set from raw variables to conclusions. Prompts typically ask you to describe the variables with appropriate statistics, such as means and standard deviations or percentages, compare groups with inferential tests, sometimes fit a regression model, and interpret every result in context. Some versions supply the data in a spreadsheet or statistical package; others describe a published study's data. Graders expect you to choose tests that fit the type of variable, report test statistics and confidence intervals as well as p-values, and discuss limitations such as missing data. Check the Canvas prompt for which software is allowed and whether output must be appended.

How the HLTH 5013 Module 4 example is put together

The sample starts by stating its purpose, variables, outcome and how missing data will be handled. Descriptive statistics for both groups come next, with an interpretation of how the differences could pull the comparison in opposite directions. Baseline differences are tested with t-tests and chi-square tests, and the paper explains why large samples yield tiny p-values for small differences. A multivariable logistic regression is specified, checked for events per predictor and reported for every variable. Interpretation converts the odds ratio for comparison with earlier results and reads borderline intervals with cited cautions. A sensitivity analysis and a plain summary close the paper.

HLTH 5013 Module 4 rubric: what full marks look like

Data analysis rubrics usually weigh appropriate choice of statistics, correct calculation or reporting, correct interpretation and attention to limitations. The choice criterion rewards tests matched to variable types and study questions. Reporting earns points when results include effect sizes, confidence intervals and test statistics, not p-values alone. Interpretation carries heavy weight, and graders deduct for common misreadings, such as treating a non-significant result as proof of no effect. Handling of missing data, model checks and sensitivity analyses earn additional credit. A clear summary for a public health audience is often part of the rubric. Methods citations and APA 7 finish the score.

HLTH 5013 Module 4 help: mistakes that cost points

Analysis papers often lose marks by pasting software output without explaining it. Another frequent error is choosing a test that does not fit the data, such as a t-test on a categorical variable. Students also treat p-values as the whole story, calling results with p above .05 negative and results below it important. Describe your groups before comparing them. Report intervals for every estimate. Say what you did about missing data. Explain each result in a sentence a health officer could use. If your data set covers injuries, vaccination or birth outcomes instead, a codebook or the file itself, plus your rubric, is enough for us to set up a Module 4 analysis of those variables. Mention which software your section uses so the output matches what your instructor expects to see.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official American College of Education document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.

More HLTH 5013 and Master of Public Health sample papers

HLTH 5013 Module 4 questions, answered

What does HLTH5013 Module 4 usually ask for?

In many sections, the fourth HLTH5013 module asks you to analyze a public health data set with descriptive statistics and inferential tests, such as t-tests, chi-square tests or regression, and to interpret the results correctly. The data set is set by your own section.

Why do large samples produce very small p-values?

Because standard errors shrink as samples grow, so even small differences become statistically detectable. The p-value does not measure how large or important a difference is.

What does a confidence interval that includes 1.0 mean for an odds ratio?

That the data are compatible with no association as well as with the other values in the interval. It does not prove there is no effect.

Where can I find a free HLTH 5013 Module 4 sample paper?

The complete Module 4 analysis is on this page: descriptive statistics, t-tests, chi-square tests and a logistic regression on a county diabetes prevention data set of 2,400 adults, with a sensitivity analysis.

How many outcome events does a logistic regression need?

A common rule of thumb is at least ten outcome events per predictor variable to produce stable estimates.