Volume 12 · Research Methods and Statistical Literacy
Chapter 3
Statistics for Nutrition Experts
Understand the statistical concepts that underpin nutrition research: distributions, confidence intervals, hypothesis testing, effect sizes, and how to interpret risk ratios correctly.
Goal of this chapter: Learn the statistical foundations of nutrition research. You will understand populations and samples, measures of central tendency and variation, confidence intervals, hypothesis testing and p-values, effect sizes, and how to calculate and interpret relative risk, absolute risk, odds ratios, and hazard ratios. These concepts are not optional for a nutrition expert. Without them, you cannot tell whether a study finding is real, clinically important, or merely noise. Understanding statistics is understanding how to separate signal from hype in nutrition research.
In this chapter
| Lesson 3.1: Populations and Samples |
| Lesson 3.2: Mean, Median and Distribution |
| Lesson 3.3: Variance and Standard Deviation |
| Lesson 3.4: Confidence Intervals |
| Lesson 3.5: P-Values |
| Lesson 3.6: Effect Size |
| Lesson 3.7: Relative Risk |
| Lesson 3.8: Absolute Risk |
| Lesson 3.9: Odds Ratios and Hazard Ratios |
| Lesson 3.10: Statistical vs Clinical Significance |
| Lesson 3.11: Chapter Revision |
| Lesson 3.12: Statistics Interpretation Cases |
Populations and Samples
Learning goal: Understand the difference between populations and samples, and why sampling variability matters in interpreting research.
A population is the entire group you want to understand: all adults in India, all people with type 2 diabetes, all women aged 50–60. A sample is a subset of the population actually studied: 1000 adults recruited from Delhi hospitals, 500 people with diabetes from a clinic. Nutrition research always uses samples because studying entire populations is impractical. But samples differ from populations by random chance. A sample of 500 adults from Delhi may have a slightly different average diet than the true population average of all Indian adults. Understanding this sampling variability is essential for interpreting whether a research finding is real or merely due to chance.
1Why We Sample Instead of Studying Populations
Studying an entire population is usually impossible. To study "all adults in India" for nutritional status would require examining hundreds of millions of people—impractical. Instead, researchers recruit a sample, measure variables in that sample, and use statistical methods to estimate what is true in the population. The sample should be representative: selected randomly or using methods that avoid systematic bias. A sample drawn only from wealthy urban areas does not represent rural or poorer populations. Sampling bias—when the sample systematically differs from the population—undermines validity. Proper sampling (random, stratified, or systematic) ensures the sample is representative and findings can be generalized to the population.
2Sampling Variability and the Role of Chance
Even with perfect sampling, different samples from the same population will differ by random chance. Draw 100 adults at random from Delhi, measure their average daily protein intake, and get 50 g. Draw another 100 adults at random, and get 52 g. Neither measurement is "wrong"—sampling variability produces these small differences. If the true population average is 51 g, both samples are close but not identical. This variability is predictable; statistics describe it using the standard error (a measure of how much samples differ from each other and from the true population value). If a study result falls within expected sampling variability, it may be a chance finding rather than a true effect. If it falls far outside expected variability, it likely represents a real effect.
3Sample Size and Statistical Power
Large samples produce smaller sampling variability than small samples. A sample of 5000 people will have a narrower range of average estimates than a sample of 50 people. This is why large studies are more trustworthy than small studies—they reduce random variation. The "statistical power" of a study is its ability to detect a true effect if it exists. Power depends on sample size, effect size (how large the effect is), and the variability of the outcome. A study with small sample size, small effect size, or high outcome variability has low power. Low power means the study may find no effect even if one truly exists (a false negative). This is why studies claiming "no effect" are only credible if they have adequate sample size and power.
4Representativeness and Generalizability
A study conducted in one city or one hospital may not reflect nutrition patterns nationwide. A study of wealthy health-conscious participants may not reflect average populations. The question is: to what population can we generalize the findings? If a study is conducted in Bangalore hospitals with middle-class adults, findings may generalize to middle-class urban populations but not to rural or lower-income populations. Researchers should be transparent about their sample characteristics—age, sex, income, location, health status—so readers can judge whether the findings apply to them. A finding is most trustworthy when replicated in multiple samples and populations.
5Constructed Example: Sampling Variability in Vegetarian Protein Intake
Imagine the true average daily protein intake in vegetarian adults in India is 48 g. A researcher recruits 100 vegetarians and measures average protein intake of 49.5 g. Another researcher recruits a different 100 vegetarians and measures 46.8 g. Both are close to the truth (48 g) but differ from each other and from the true value due to sampling variability. If sampling variability is known to be ±2.5 g (the standard error), both samples fall within expected ranges. A third researcher recruits only organic-food-conscious vegetarians and measures 58 g—far higher. This represents sampling bias, not just variability; the sample is not representative of the general vegetarian population.
Populations are entire groups; samples are subsets studied in research. Sampling variability means different samples differ by chance. Large samples reduce variability. Representativeness ensures findings generalize to the population. Understanding sampling is the foundation for interpreting whether research findings are real or due to chance.
6Sampling across a country as varied as India
A sample is only useful if it represents the population you intend to apply it to, and India makes that unusually demanding. Diet varies fundamentally between states — rice in the south and east, wheat in the north, millets in parts of the Deccan, fish along both coasts, near-universal vegetarianism in some communities and none in others. Add urban and rural differences, income spread, and genetic variation across populations, and a study conducted in one city may not describe the country at all.
This is why national surveys such as NFHS sample by state and by urban-rural stratum rather than drawing one large convenience sample, and why a single-centre hospital study from Delhi should not be read as an Indian finding without qualification. When evaluating any study claiming to describe Indians, the first question is where and among whom it was conducted — a result from urban Kerala and one from rural Bihar can differ so much that averaging them describes nobody.
A study measures dietary calcium in 150 adults from one city clinic. The average is 800 mg/day. Why is generalization to "all Indian adults" problematic?
Answer: The sample is from one city clinic (geographic limitation) and likely includes more health-conscious people than the general population (bias). Findings may not represent rural populations, lower-income groups, or those without access to clinics. Generalization should be to the population that matches the sample characteristics, not to all Indian adults.
- Populations are entire groups; samples are subsets studied in research.
- Sampling variability means samples differ by random chance; larger samples have smaller variability.
- Representative sampling ensures findings generalize to the intended population.
- Sampling bias occurs when samples systematically differ from populations, limiting generalizability.
Next: Lesson 3.2 introduces measures of central tendency and distributions—how to describe where data clusters.
Mean, Median and Distribution
Learning goal: Understand mean and median as measures of central tendency, and understand how distributions describe data patterns.
When you collect data (heights, daily calorie intakes, blood glucose levels) from a sample, the first step is describing where the data clusters. The mean (average) and median (middle value) are two ways to describe the center of the data. The distribution is the pattern: are values clustered tightly around the center, or spread widely? Understanding means, medians, and distributions is essential for interpreting research data.
1Mean (Average) and Why It Matters
The mean is the sum of all values divided by the count. If 100 adults consume 48, 52, 50, 51, 49 g protein respectively (and so on for 100 people), the mean is the sum divided by 100. The mean is useful because it uses all data points and has nice statistical properties. But the mean is sensitive to outliers (extreme values). If one person in the sample consumed 500 g protein (an extreme outlier), the mean would be artificially inflated. The mean is most useful for symmetrical, bell-shaped distributions without extreme outliers.
2Median (Middle Value) and Distributions With Outliers
The median is the middle value when data is ranked from lowest to highest. If you rank 100 protein intakes from lowest to highest and pick the 50th value, that is the median. The median is robust to outliers: even if one person ate 500 g protein, the median is not affected. The median is more appropriate than the mean when data has outliers or is skewed (asymmetrical). For example, income data is typically skewed—most people earn moderate amounts, but a few earn millions. The median income is more representative than the mean (which is inflated by billionaires). In nutrition, energy intake data is often skewed; the median may be more representative than the mean.
3Normal (Bell-Shaped) Distributions
Many biological variables (height, weight, blood glucose in healthy people) follow a normal (Gaussian) distribution—a bell shape with most values near the center and progressively fewer values at the extremes. In a normal distribution, mean and median are equal (or very close). The normal distribution is important in statistics because many statistical tests assume normality. If data is normally distributed, the mean is representative. If data is skewed (not normally distributed), the median is more representative. Research papers should report both mean and standard deviation (a measure of spread; see Lesson 3.3) or median and interquartile range, depending on the distribution.
4Skewed Distributions: Right-Skew and Left-Skew
A right-skewed distribution has a long tail on the right side: a few high values pull the mean higher than the median. Examples: income, disease severity (most people healthy; a few very sick). A left-skewed distribution has a long tail on the left: a few low values pull the mean lower than the median. For right-skewed data, the median is below the mean and more representative. For left-skewed data, the median is above the mean and more representative. Recognizing skew is important: if a paper reports mean and standard deviation for clearly skewed data (like energy intake), the mean may overstate or understate the typical value. Ask: is the data skewed? If so, is the median reported? If not, interpret the mean cautiously.
5Constructed Example: Dietary Energy Intake in a Sample
Imagine measuring daily energy intake in 50 adults. Most consume 1800–2200 kcal, but two consume only 1200 kcal (dieters) and two consume 3500 kcal (athletes). The data is right-skewed (a few high values). Mean = 2250 kcal. Median = 2000 kcal. The median (2000 kcal) represents the "typical" person better than the mean (2250 kcal) because the mean is inflated by the athletes. If a researcher reports only the mean (2250 kcal), it overstates the typical intake. This is why reporting both mean and median, and being transparent about distribution shape, is important.
Myth: "The mean always represents the typical value." Reality: The mean represents the typical value only when data is symmetrically distributed (normal distribution). For skewed data, the median is more representative. Good research reports both and notes whether data is skewed.
A nutrition study reports mean daily sodium intake = 4500 mg, standard deviation = 2800 mg. This suggests the data is skewed. Why?
Answer: The standard deviation is 62% of the mean (2800/4500). This large ratio indicates high variability and suggests skew (a few people with very high sodium intake pulling the mean up). For skewed data, the median and interquartile range are more appropriate than mean and standard deviation.
- Mean (average) and median (middle value) describe the center of data.
- Mean is sensitive to outliers; median is robust.
- For normally distributed data, mean and median are equal; for skewed data, they differ.
- Good research reports both mean and median, and notes distribution shape.
Next: Lesson 3.3 explores variance and standard deviation—how to measure data spread.
Variance and Standard Deviation
Learning goal: Understand variance and standard deviation as measures of data spread, and use them to interpret whether values are tightly clustered or widely scattered.
Knowing the mean tells you where data clusters, but not how tightly. Two studies might find mean daily protein intake of 50 g, but in one study, everyone consumes 48–52 g (tightly clustered), and in the other, consumption ranges from 20–80 g (widely scattered). Variance and standard deviation measure this spread. Understanding spread is as important as understanding the center for interpreting research.
1Variance as Average Squared Distance From the Mean
Variance measures how far, on average, data points are from the mean. It is calculated as the average of the squared distances from the mean. If all values equal the mean (no variability), variance is zero. If values are spread far from the mean, variance is large. Variance is measured in squared units (which is why it is less intuitive than standard deviation). For protein intake in grams, variance would be in grams-squared, which is awkward to interpret. This is why standard deviation (the square root of variance) is more commonly reported.
2Standard Deviation as the Typical Distance From the Mean
Standard deviation (SD) is the square root of variance, measured in the same units as the data. If protein intake has SD = 5 g and mean = 50 g, the SD of 5 g tells you that most values are within about 5 g of the mean (roughly 45–55 g). Standard deviation is intuitive: it describes the typical scatter around the mean. The larger the SD, the more scattered the data. A small SD indicates tight clustering; a large SD indicates wide scatter. In a normal distribution, about 68% of values fall within ±1 SD of the mean, about 95% within ±2 SDs, and about 99.7% within ±3 SDs.
3Coefficient of Variation: Comparing Variability Across Different Scales
Comparing standard deviations directly is problematic when variables are on different scales. If one study reports protein intake (mean = 50 g, SD = 5 g) and another reports energy intake (mean = 2000 kcal, SD = 300 kcal), which is more variable? The coefficient of variation (CV = SD / mean × 100%) allows comparison. For protein: CV = (5/50) × 100% = 10%. For energy: CV = (300/2000) × 100% = 15%. Energy intake is more variable (15% > 10%) even though the SD is much larger in absolute terms. CV is useful for comparing variability across different measurements.
4Standard Error vs Standard Deviation: Understanding the Difference
Standard deviation (SD) describes the spread of data points in a sample. Standard error (SE) describes the uncertainty in estimating the population mean from a sample. SE = SD / √n, where n is sample size. A study with n=100 will have smaller standard error (narrower uncertainty) than a study with n=25, even if both have the same SD. Standard error is used to construct confidence intervals (Lesson 3.4). When a paper reports mean ± standard error, they are describing uncertainty in the estimate. When they report mean ± standard deviation, they are describing the spread of data. Both are useful; they answer different questions.
5Constructed Example: Two Studies With Different Variability
Study A measures daily calcium intake in 60 urban adults: mean = 800 mg, SD = 150 mg. Study B measures daily calcium intake in 60 rural adults: mean = 750 mg, SD = 350 mg. Both are reasonable samples. Study A's data is tightly clustered (most values 650–950 mg); Study B's is more scattered (values range 50–1450 mg, reflecting greater dietary diversity). The SD immediately tells you Study A has more consistent intake; Study B has greater variability. This might reflect different food availability or cultural food patterns. A researcher should comment on this difference when interpreting results.
Variance and standard deviation measure data spread. Small SD indicates tight clustering around the mean; large SD indicates wide scatter. Standard error is different from SD: SE measures uncertainty in estimating the population mean. Both are important for interpreting research; they answer different questions about data.
Study X reports mean blood glucose = 100 mg/dL, SD = 15 mg/dL (n=100). Study Y reports mean = 100 mg/dL, SD = 15 mg/dL (n=400). Why does Study Y provide stronger evidence about the population mean?
Answer: Both have the same SD (data spread), but Study Y has larger sample size (n=400). Standard error = SD/√n. Study Y's SE = 15/√400 = 0.75. Study X's SE = 15/√100 = 1.5. Study Y has half the SE, meaning less uncertainty in the estimate of the population mean. Larger samples reduce uncertainty even if data spread is identical.
- Variance measures average squared distance from the mean.
- Standard deviation (SD) is the square root of variance, measured in original units.
- Large SD indicates data spread wide from the mean; small SD indicates tight clustering.
- Standard error measures uncertainty in the sample estimate; SE = SD/√n decreases with larger sample size.
Next: Lesson 3.4 introduces confidence intervals—a way to express uncertainty about population parameters estimated from samples.
Confidence Intervals
Learning goal: Understand confidence intervals as a range of plausible values for a population parameter, and interpret what "95% confidence" means.
A study reports mean protein intake = 50 g. But is the true population mean exactly 50 g, or could it be 48 g or 52 g? A confidence interval answers this question. It provides a range of plausible values for the true population parameter (mean, difference between groups, odds ratio, etc.) estimated from sample data. A 95% confidence interval means: if this study were repeated 100 times, the confidence interval would contain the true population parameter 95 times and miss it 5 times. Confidence intervals quantify uncertainty.
1Constructing a 95% Confidence Interval Around a Mean
If a study measures protein intake (mean = 50 g, SE = 2 g, n = 100), the 95% confidence interval is approximately mean ± 1.96 × SE. This gives 50 ± (1.96 × 2) = 50 ± 3.92, or 46.08 to 53.92 g. The interpretation: based on this sample, we are 95% confident the true population mean lies between 46.08 and 53.92 g. The 1.96 is a constant from the normal distribution; it ensures that 95% of repeated samples fall within this range. If the sample size were smaller (larger SE), the interval would be wider (more uncertainty). If the sample size were larger (smaller SE), the interval would be narrower (less uncertainty).
2Interpreting Confidence Intervals: What They Mean and Don't Mean
A 95% confidence interval does NOT mean there is a 95% probability the true value is in the interval. The true population parameter is fixed; it either is or is not in the interval. Rather, a 95% CI means: the method used to construct the interval works correctly 95% of the time. If you repeated the study 100 times and calculated a CI each time, 95 of those intervals would contain the true parameter. This is a subtle distinction but important: CIs quantify the long-run frequency of the method, not the probability of a particular interval containing the truth. A wider CI indicates more uncertainty (less confidence in the estimate); a narrow CI indicates more certainty (more precision).
3Using Confidence Intervals to Test Hypotheses
If a study compares protein intake in vegetarians (mean = 45 g) vs omnivores (mean = 55 g), the difference is 10 g. But is this a real difference or chance variation? Calculate the 95% CI for the difference. If the CI is 5 to 15 g (does not include zero), the difference is statistically significant at the 0.05 level—it is unlikely to be zero (due to chance). If the CI is -2 to 22 g (includes zero), the difference is not statistically significant—it could be zero (a chance finding). This is equivalent to a hypothesis test; CIs provide more information than simple yes/no significance (see Lesson 3.5).
4Confidence Intervals for Risk Ratios and Odds Ratios
CIs are not limited to means. They apply to any estimate: differences between groups, risk ratios, odds ratios, hazard ratios (see Lessons 3.7–3.9). An odds ratio of 1.5 with 95% CI = (1.1 to 2.0) means: based on the sample, the odds are 1.5 times higher in the exposed group, and we are 95% confident the true odds ratio in the population is between 1.1 and 2.0. Because the CI does not include 1.0 (no effect), the finding is statistically significant. If the CI were (0.9 to 2.2), it would include 1.0 (no effect), and the finding would not be significant, indicating the sample did not provide strong evidence of an effect.
5Constructed Example: Two Studies With Different Precision
Study A (n=500) measures calcium intake: mean = 800 mg, 95% CI = (780 to 820 mg). Study B (n=50) measures calcium intake in the same population: mean = 800 mg, 95% CI = (720 to 880 mg). Both estimate the same mean (800 mg), but Study A's CI is much narrower (only ±20 mg), indicating high precision. Study B's CI is much wider (±80 mg), indicating low precision due to small sample size. Study A provides stronger evidence about the population mean. Both are valid, but Study A's precision makes its conclusion more trustworthy.
Confidence intervals provide a range of plausible values for a population parameter. A 95% CI means: if this study were repeated 100 times, the CI would contain the true parameter 95 times. Narrow CI indicates high precision; wide CI indicates uncertainty. CIs directly show whether a finding is statistically significant (if the CI excludes the null value, e.g., zero or 1.0).
A study compares fiber intake between two groups: difference = 5 g/day, 95% CI = (-2 to 12 g). Is the finding statistically significant?
Answer: No. The CI includes zero (falls within -2 to 12 g), meaning the true difference could be zero (no difference between groups). This is not statistically significant at the 0.05 level. If the CI were (2 to 8 g), not including zero, it would be statistically significant.
- Confidence intervals provide a range of plausible values for a population parameter.
- 95% CI means the method works correctly 95% of the time across repeated studies.
- Narrow CI indicates high precision; wide CI indicates uncertainty (from small sample size or high variability).
- If CI excludes the null value (zero, or 1.0), the finding is statistically significant.
Next: Lesson 3.5 explores p-values and hypothesis testing—another way to assess whether findings are statistically significant.
P-Values
Learning goal: Understand p-values as the probability of observing a result as extreme (or more extreme) if the null hypothesis were true; understand the limits of p-values for interpreting research.
A p-value is one of the most misunderstood statistics in research. It is NOT the probability that the null hypothesis is true, and it is NOT the probability that the finding is a false positive. Instead, a p-value is the probability of observing a result as extreme as (or more extreme than) the one observed if the null hypothesis (no effect) were true. If p < 0.05, the result is called "statistically significant"—it is unlikely (less than 5% chance) if no effect exists. But "statistically significant" does not mean "clinically important." Understanding p-values is essential for avoiding over-interpretation of research.
1Null Hypothesis and Alternative Hypothesis
A hypothesis test starts with a null hypothesis (H0): usually, that there is no effect or difference. For example, "protein supplementation does not change muscle mass" or "vegetarian diets do not affect mortality." The alternative hypothesis (H1) is that there is an effect: "protein supplementation increases muscle mass" or "vegetarian diets do affect mortality." The study collects data and tests whether the evidence supports H0 or H1. Typically, a significance level (alpha) of 0.05 is chosen: if the p-value is less than 0.05, we reject H0 (conclude an effect likely exists). If p ≥ 0.05, we fail to reject H0 (conclude the evidence does not support an effect).
2What a P-Value Is (and Is Not)
A p-value is the probability of observing the data (or more extreme data) if H0 is true. Example: a study compares protein intake between vegetarians (mean = 45 g) and omnivores (mean = 55 g), finding a difference of 10 g. The p-value might be 0.02. This means: if vegetarians and omnivores truly had the same intake (H0 true), there is a 2% probability of observing a difference as large as 10 g (or larger) by random chance. Since 2% < 5%, the result is statistically significant. A p-value is NOT the probability that H0 is true. H0 is either true or false; it is not probabilistic. A p-value is also NOT the probability that the finding is false. A p < 0.05 means the finding is unlikely if H0 is true; it does not directly tell you the probability that the result is a false positive.
3P < 0.05: Statistical Significance vs Real-World Importance
Statistical significance (p < 0.05) means the result is unlikely due to chance. It does not mean the finding is clinically or practically important. A large study might find a tiny difference (e.g., vitamin D supplementation increases bone density by 0.5%, p = 0.04) that is statistically significant but too small to matter clinically. Conversely, a small study might find a large difference (e.g., 15% improvement) that is NOT statistically significant (p = 0.10) because sample size is too small to reach p < 0.05 threshold, even though the effect might be clinically meaningful. Distinguishing statistical from clinical significance (Lesson 3.10) is crucial for interpreting research correctly.
4Multiple Testing and the Multiple Comparisons Problem
A study with 10 outcome variables might be analyzed 10 times. If each test uses p < 0.05, on average, one test will show p < 0.05 purely by chance (1 out of 20 tests, statistically speaking, even if no true effects exist). This is the multiple comparisons problem: the more tests you run, the more likely at least one will be significant by chance. To correct for this, researchers use adjusted p-value thresholds (e.g., p < 0.005 if testing 10 outcomes, or Bonferroni correction: 0.05/10 = 0.005). Without correction, studies testing many outcomes risk spurious significant findings. When reading research, check: how many outcomes were tested? Were p-values corrected for multiple testing?
5Constructed Example: Interpreting a P-Value in Context
A study (n=400) tests whether a vegetarian diet affects blood pressure. Systolic BP is 1.5 mmHg lower in vegetarians vs omnivores, p = 0.03. The result is statistically significant (p < 0.05). But is 1.5 mmHg clinically meaningful? For hypertension management, changes of 5–10 mmHg are considered clinically relevant. A 1.5 mmHg reduction, while real, is too small to guide clinical practice. The p-value tells you the difference is unlikely to be zero; effect size and clinical judgment tell you whether the difference matters. Both pieces of information are needed for correct interpretation.
Myth: "p = 0.05 means there is a 5% probability the result is a false positive." Reality: A p-value does not directly give you the probability of a false positive (that depends on prior probability and effect size). p = 0.05 means: if the null hypothesis were true, there is a 5% probability of observing data as extreme as observed due to random chance.
A study tests 20 different biomarkers and finds one with p = 0.04. Should you conclude this biomarker is affected by the intervention?
Answer: Cautiously. With 20 tests, we expect ~1 to be significant at p < 0.05 by chance alone, even if no true effects exist. This finding may be a false positive due to multiple testing. If p-values were not corrected for multiple comparisons (e.g., Bonferroni correction would be 0.05/20 = 0.0025), this result is suspect. Replication in an independent sample is needed to confirm.
- P-value is the probability of observing data as extreme as (or more) observed if null hypothesis is true.
- p < 0.05 means statistically significant—unlikely due to chance.
- Statistical significance does not equal clinical importance; both should be evaluated.
- Multiple comparisons inflate false positive risk; p-values should be corrected if testing many outcomes.
Next: Lesson 3.6 explores effect size—the magnitude of the difference or association—which complements p-values.
Effect Size
Learning goal: Understand effect size as a measure of the magnitude of an effect, independent of sample size, and use it to interpret the practical importance of research findings.
Effect size measures how large an effect is. A p-value tells you whether an effect exists (unlikely due to chance); effect size tells you how big it is. A study with large sample size might find a tiny effect that is statistically significant (p < 0.05). A study with small sample size might find a large effect that is not statistically significant (p > 0.05). Effect size is independent of sample size; it reflects the true magnitude of the phenomenon. Understanding effect size is essential for interpreting whether research findings are practically important.
1Cohen's d: Effect Size for Comparing Two Groups
Cohen's d measures the difference between two group means in standard deviation units. d = (mean1 - mean2) / pooled SD. Interpretation: d = 0.2 is small, d = 0.5 is medium, d = 0.8 is large. Example: if a protein supplementation study finds mean muscle mass gain of 2 kg in the supplement group and 0.5 kg in the control group, the raw difference is 1.5 kg. If pooled SD = 1.5 kg, then d = 1.5/1.5 = 1.0 (a large effect). This means the difference is large relative to the variability in the outcome. Cohen's d is straightforward to calculate and interpret, making it useful for any study comparing two groups.
2R-Squared: Effect Size for Correlations and Regressions
R-squared (R²) represents the proportion of variance in one variable explained by another. If a study measures height and finds R² = 0.30 for predicting weight, it means height explains 30% of weight variation. The remaining 70% is due to other factors (age, genetics, nutrition, etc.). R² ranges from 0 to 1. R² = 0 means no relationship; R² = 1 means perfect prediction. Interpretation: R² = 0.01–0.09 is small, 0.09–0.25 is medium, > 0.25 is large (these are rough guidelines). R-squared is common in nutrition research when exploring associations between variables.
3Odds Ratio and Risk Ratio as Effect Sizes
An odds ratio (OR) or risk ratio (RR) comparing two groups is itself an effect size (see Lessons 3.7–3.9 for definitions). A RR = 1.5 means one group has 1.5 times the risk of an outcome vs the other group—a 50% increase. Interpretation depends on context: a RR = 1.1 (10% increase) might be large for a small risk (e.g., rare cancer) but small for a common outcome (e.g., common disease). The absolute risk change matters more than the relative change (see Lesson 3.8). Reporting both relative and absolute effect sizes gives a complete picture.
4Combining Effect Size and Confidence Intervals
The complete picture is effect size + confidence interval. A study might report: "protein supplementation increased muscle mass by 2 kg (d = 0.8, 95% CI = 1.0 to 3.0 kg)." This tells you: the effect size is large (d = 0.8), the point estimate is 2 kg, and you can be 95% confident the true effect is between 1.0 and 3.0 kg. The narrow CI indicates precision. Compare this to: "supplementation increased muscle mass by 2 kg (d = 0.8, 95% CI = -1 to 5 kg)." Same effect size and point estimate, but a wide CI indicates uncertainty (the true effect could be zero or even negative). Effect size + CI together provide complete information about magnitude and precision.
5Constructed Example: Small P-Value, Large Effect vs Large P-Value, Small Effect
Study A (n=5000) compares vitamin D supplementation (mean increase in bone density = 1%, p = 0.001, d = 0.15). The p-value is very small (highly significant), but the effect size is small (d = 0.15). Study B (n=50) compares vitamin D supplementation (mean increase = 8%, p = 0.08, d = 0.7). The p-value is not significant (p > 0.05), but the effect size is medium-to-large (d = 0.7). Study A has statistical significance but little practical importance. Study B has practical importance but failed to reach statistical significance (due to small sample size). Both effect sizes and p-values are necessary for full interpretation; neither alone is sufficient.
Effect size measures the magnitude of an effect, independent of sample size. p-values tell you whether an effect exists; effect size tells you how large it is. Always report effect sizes alongside p-values. A statistically significant (p < 0.05) effect might be too small to matter clinically; a non-significant effect might be large but not statistically significant due to small sample size.
A study compares two dietary interventions for weight loss: intervention A reduces weight by 3 kg (p = 0.02, d = 0.3), intervention B reduces weight by 8 kg (p = 0.15, d = 0.9). Which is more important clinically?
Answer: Intervention B, despite the larger p-value. Intervention B has a large effect size (d = 0.9) and larger weight loss (8 kg), which is clinically substantial. Intervention A is statistically significant but has small effect size and small weight loss (3 kg). Intervention B is more clinically important, even though the p-value is not significant (possibly due to small sample size).
- Effect size measures magnitude of an effect, independent of sample size.
- Cohen's d, R², odds ratios, and risk ratios are common effect size measures.
- Statistical significance (p-value) does not equal clinical importance (effect size).
- Always report effect sizes alongside p-values for complete interpretation.
Next: Lesson 3.7 introduces relative risk and relative risk reduction—how to compare risks between groups.
Relative Risk
Learning goal: Understand relative risk (RR) as a comparison of risk between two groups, and distinguish it from absolute risk reduction.
Relative risk is one of the most commonly cited statistics in nutrition research and media. A headline might say: "Study shows vegetarian diet cuts heart disease risk by 50%." This refers to relative risk reduction. But what does 50% reduction mean? Relative risk can be misinterpreted, especially when absolute risks are not reported. Understanding relative risk, alongside absolute risk, is essential for accurate interpretation.
1Defining Relative Risk: Comparing Risk Between Groups
Risk is the probability of an outcome. Risk of heart disease might be 10% in one group and 5% in another. Relative risk (RR) is the ratio of risks: RR = risk in group 1 / risk in group 2. If risk is 10% in exposed and 5% in unexposed, RR = 10/5 = 2.0. This means exposure doubles the risk (or the exposed group has twice the risk). RR = 1.0 means equal risk (no effect). RR < 1.0 means lower risk in the exposed group (protective effect). RR > 1.0 means higher risk (harmful effect). A RR of 0.75 means 25% lower risk in the exposed group.
2Relative Risk Reduction vs Absolute Risk Reduction
Relative Risk Reduction (RRR) is (1 - RR) × 100%. If RR = 0.5, then RRR = (1 - 0.5) × 100% = 50%. This is the 50% reduction often cited in headlines. But it is misleading without context. If baseline risk is 1% (1 in 100 get disease), a 50% RR reduction means new risk is 0.5% (1 in 200 get disease). The absolute reduction is only 0.5 percentage points. If baseline risk is 40% (40 in 100 get disease), the same 50% RR reduction means new risk is 20%, an absolute reduction of 20 percentage points. Same 50% RRR, vastly different clinical importance. Absolute Risk Reduction (ARR) = baseline risk - new risk, expressed in percentage points. ARR is more clinically meaningful than RRR.
3Number Needed to Treat (NNT)
Number Needed to Treat (NNT) is the number of people who must receive an intervention to prevent one outcome. NNT = 1 / ARR. Example: if an intervention reduces absolute risk from 20% to 15%, ARR = 5%, NNT = 1/0.05 = 20. This means 20 people must receive the intervention to prevent one case of the outcome. A low NNT (e.g., 5) is efficient; a high NNT (e.g., 1000) means you must treat many people to benefit one. NNT directly translates relative/absolute statistics into practical clinical language. A study should always report NNT for interventions, especially preventive interventions in healthy populations.
4Constructed Example: Misleading Relative Risk Reporting
A study finds that consuming 50 grams of dietary fiber daily reduces colorectal cancer risk by 50% (RRR = 50%, RR = 0.5). Headline: "Fiber cuts cancer risk in half!" Sounds impressive. But baseline colorectal cancer risk (without fiber) might be 5% (5 in 100 people). After fiber, it becomes 2.5% (2.5 in 100 people). Absolute risk reduction is 2.5 percentage points. NNT = 1/0.025 = 40. This means 40 people must consume high-fiber diets to prevent one case of colorectal cancer. This is still worthwhile (especially given other fiber benefits), but far less dramatic than "cuts risk in half." Reporting relative and absolute together prevents misinterpretation.
5Reporting Relative Risk Responsibly
When presenting relative risk, always include absolute risk or NNT. A responsible statement: "A high-vegetable diet reduces mortality risk by 20% (RR = 0.80), corresponding to an absolute reduction of 3 percentage points (from 15% to 12% over 10 years), NNT = 33." This gives readers the relative change (20% reduction), the absolute change (3 points), and practical context (33 people must follow the diet for 10 years to prevent one death). Without this detail, statistics mislead even well-intentioned readers.
Relative risk (RR) compares risk between two groups; RR = risk1 / risk2. RR = 1.0 means no difference; RR < 1.0 means lower risk (protective); RR > 1.0 means higher risk. Relative risk reduction (RRR) is often reported but can mislead without absolute risk reduction (ARR) or number needed to treat (NNT). Always report both relative and absolute changes for complete interpretation.
6Relative risk, and how it gets used in Indian health reporting
Relative risk expresses how much more likely an outcome is in one group than another, and it is the number most often quoted precisely because it sounds largest. A headline reporting that a food raises diabetes risk by 30% is reporting a relative risk of 1.3, and tells you nothing about how common the outcome was to begin with. This matters acutely in India, where baseline risks for conditions such as type 2 diabetes are already high, so the same relative risk translates into a very different number of affected people than it would elsewhere.
Read relative risk as a signal about direction and strength, never as a statement about personal risk. The follow-up question is always what the absolute numbers were, and in which population they were measured. A relative risk derived from a European cohort applied to an Indian client is doubly uncertain: the ratio may hold while the baseline it multiplies is entirely different.
A dietary intervention reduces disease risk: RR = 0.75 (25% reduction). Baseline risk is 4%. What is absolute risk reduction and NNT?
Answer: Baseline risk = 4%. New risk = 4% × 0.75 = 3%. Absolute risk reduction (ARR) = 4% - 3% = 1 percentage point. NNT = 1/0.01 = 100. This means 100 people must follow the intervention to prevent one case. Despite the 25% relative reduction, the absolute benefit is modest.
- Relative risk (RR) compares risk between groups: RR = risk1 / risk2.
- Relative risk reduction (RRR) is often reported but misleads without absolute risk reduction (ARR).
- Absolute risk reduction (ARR) = baseline risk - new risk, in percentage points.
- Number needed to treat (NNT) = 1 / ARR; it translates statistics into practical context.
Next: Lesson 3.8 continues with absolute risk, comparing absolute vs relative risk reduction.
Absolute Risk
Learning goal: Understand absolute risk as the actual probability of an outcome in a population, and use it to evaluate the clinical importance of interventions.
Absolute risk is the actual probability of an outcome in a group, expressed as a percentage. If 10 out of 100 people get heart disease, absolute risk is 10%. Absolute risk is more intuitive than relative risk and directly answers the question: "What is my chance of this outcome?" Understanding absolute risk and how interventions change it is essential for patients and clinicians making decisions.
1Absolute Risk vs Relative Risk: Context Matters
Relative risk of 2.0 (doubled risk) sounds alarming. But context matters. If absolute risk is 1% and RR = 2.0, new absolute risk is 2%—a 1 percentage point increase, clinically modest. If absolute risk is 40% and RR = 2.0, new absolute risk is 80%—a 40 percentage point increase, clinically major. The same RR leads to vastly different conclusions depending on baseline absolute risk. This is why research must report both. A study comparing two dietary interventions should report: "The high-protein diet increased mortality risk (RR = 1.20, 95% CI = 0.95–1.51, absolute risk 5% vs 4.2%, NNT = 125)." This gives relative, absolute, and practical context.
2Number Needed to Treat and Number Needed to Harm
NNT is the number of people who must receive an intervention to benefit one person. NNH (number needed to harm) is the number who must receive it for one to be harmed. Example: a medication reduces heart attack risk by 2 percentage points (ARR = 2%), NNT = 50. But if it causes serious side effects in 1% of patients, NNH = 100. Comparing NNT and NNH helps clinicians and patients decide: is the benefit worth the risk? If NNT = 50 (50 people treated to prevent one attack) but NNH = 100 (100 treated to cause one serious side effect), the benefit-harm ratio is favorable. In nutrition (mostly low-risk interventions), NNH is typically very large; the decision usually favors the intervention.
3Cumulative Absolute Risk Over Time
Absolute risk is often expressed over a time period. "10-year cardiovascular risk" means the probability of a heart attack or stroke in the next 10 years. "Lifetime risk" means over a person's lifespan (usually to age 85 or 90). An intervention's benefit depends on the time frame. A diet that reduces 10-year heart disease risk by 1 percentage point (ARR = 1%) might reduce lifetime risk by 5 percentage points (different baseline, longer time). Always note the time frame when interpreting absolute risk. Short follow-up studies underestimate long-term benefits; long follow-up studies provide more accurate estimates.
4Absolute Risk for Prevention vs Treatment
Preventive interventions in healthy populations typically have high NNT because baseline risk is low. Preventing heart attacks in healthy adults might require NNT = 50–100 (treatment of 50–100 healthy people to prevent one attack). Treatment of diseased populations typically has lower NNT. Treating people with established heart disease to prevent a second attack might have NNT = 10–20. This difference is important: the same treatment looks less impressive in prevention (high NNT) but may be worthwhile due to low side effects. In treatment (low NNT), the benefit-harm ratio strongly favors the intervention.
5Constructed Example: Comparing Interventions by Absolute Benefit
Two dietary interventions for reducing type 2 diabetes risk in prediabetic adults: Intervention A reduces risk by 30% (RRR = 30%, RR = 0.70). Baseline risk is 15% over 5 years. New risk = 15% × 0.70 = 10.5%. ARR = 15% - 10.5% = 4.5%, NNT = 22. Intervention B reduces risk by 50% (RRR = 50%, RR = 0.50). Baseline risk is 2% over 5 years (lower-risk population). New risk = 2% × 0.50 = 1%. ARR = 2% - 1% = 1%, NNT = 100. Intervention A has lower NNT (22 vs 100) despite lower RRR, because baseline risk is higher. Absolute benefit matters more than relative reduction.
Absolute risk is the actual probability of an outcome (e.g., 10% chance of heart disease). Absolute risk reduction (ARR) is the difference in absolute risk between intervention and control groups. Number needed to treat (NNT = 1/ARR) translates this into practical terms: how many people must be treated to benefit one. Always report absolute risk and NNT alongside relative risk for complete interpretation.
6Absolute risk against Indian baseline rates
Absolute risk is what a client actually experiences, and converting to it is the single most useful thing a practitioner can do with a scary statistic. If an outcome occurs in 2 people per 1,000 and a food raises that by 50%, it now occurs in 3 per 1,000 — a relative increase that sounds alarming and an absolute increase of one person in a thousand. Presenting both is honest; presenting only the relative figure is how most nutrition scare stories are built.
The Indian dimension is that baselines here are genuinely high for some conditions, which cuts both ways. With diabetes and prediabetes affecting very large numbers of Indian adults per ICMR-INDIAB, even a modest relative reduction from a dietary change translates into a large absolute benefit across the population — a stronger argument for public-health intervention than the same relative figure would justify in a low-prevalence country. Using local baselines rather than imported ones is what makes the arithmetic meaningful.
Study A (low-risk population): baseline risk = 2%, intervention reduces risk to 1% (ARR = 1%, RR = 0.50). Study B (high-risk population): baseline risk = 20%, intervention reduces risk to 10% (ARR = 10%, RR = 0.50). Both have the same RR. Which has lower NNT and higher absolute benefit?
Answer: Study B. NNT for Study A = 1/0.01 = 100. NNT for Study B = 1/0.10 = 10. Study B has lower NNT and higher absolute benefit (10 percentage points vs 1 percentage point), despite identical RR. Baseline risk determines absolute benefit; the same intervention is more beneficial in higher-risk populations.
- Absolute risk is the actual probability of an outcome in a population.
- Absolute risk reduction (ARR) = baseline risk - new risk, in percentage points.
- Number needed to treat (NNT) = 1/ARR; it contextualizes absolute benefit in practical terms.
- The same intervention has higher absolute benefit in higher-risk populations; NNT varies with baseline risk.
Next: Lesson 3.9 introduces odds ratios and hazard ratios—effect measures for different study designs.
Odds Ratios and Hazard Ratios
Learning goal: Understand odds ratios (used in case-control studies) and hazard ratios (used in survival analysis), and distinguish them from relative risk.
Odds ratios and hazard ratios are effect measures common in observational and survival studies. They are often interpreted as risk ratios (or as approximations of them), but they have distinct meanings. Understanding them is essential for interpreting case-control and survival studies in nutrition research.
1Odds and Odds Ratios: Definitions and Uses
Odds is a different way to express probability. If risk is 20%, odds is 20/80 = 0.25. If risk is 50%, odds is 50/50 = 1.0. An odds ratio (OR) is the ratio of odds in two groups: OR = (a/b) / (c/d) where a and b are events and non-events in group 1, c and d are events and non-events in group 2. Example: a case-control study compares high-fiber diet (exposed) vs low-fiber (unexposed) among people with colorectal cancer (cases) and controls. Among cases: 30 ate high-fiber, 70 ate low-fiber. Odds of high-fiber among cases = 30/70 = 0.43. Among controls: 40 ate high-fiber, 60 ate low-fiber. Odds among controls = 40/60 = 0.67. OR = 0.43/0.67 = 0.64. This means cases are 36% less likely to have eaten high-fiber (protective effect). OR < 1.0 indicates protection; OR > 1.0 indicates increased risk.
2Why Odds Ratios Are Used: Case-Control Study Advantage
Case-control studies identify people with an outcome (cases) and without (controls), then look backward to exposure. You cannot directly calculate risk (probability of outcome in exposed vs unexposed) because you selected on outcome. But you can calculate odds of exposure in cases vs controls, which gives an odds ratio. For rare diseases, OR approximates RR: if a disease is rare, odds ≈ risk, so OR ≈ RR. For common diseases (e.g., obesity), OR and RR differ substantially. A study should be clear: is the measure an odds ratio (appropriate for case-control studies) or is it being misinterpreted as a risk ratio? An OR of 2.0 for a rare disease is approximately a 2-fold increased risk; an OR of 2.0 for a common disease overstates risk.
3Hazard Ratios: Effect Sizes for Survival Time
Hazard ratio (HR) is used in survival analysis (e.g., time to death, time to disease). HR compares the "hazard" (instantaneous rate of an event) between two groups. HR = 1.0 means equal hazard (no difference in survival). HR < 1.0 means lower hazard in exposed group (longer survival). HR > 1.0 means higher hazard (shorter survival). Example: a study comparing two diets follows participants for 10 years and records deaths. Diet A: 20 deaths in 1000 people; Diet B: 30 deaths in 1000 people. If death times are similar (not stratified by follow-up time), HR ≈ 0.67 (Diet A has about 33% lower hazard). Like OR, HR should be interpreted in context; it reflects the rate of events, not simple probabilities.
4Comparing OR, RR, and HR: Interpretation and Context
When reading research: (1) Identify the study design. RCTs report RR or risk difference. Case-control studies report OR. Cohort and survival studies report RR, HR, or both. (2) Check if OR is being misinterpreted. An OR ≈ RR only for rare outcomes (< 10%). (3) Check if HR is appropriate. HR is correct for time-to-event data; RR is not. (4) Look for clarification in text: "the odds of [outcome] were [OR] times higher" (correct language) vs "the risk of [outcome] increased [OR] times" (incorrect; confuses odds with risk). Reading carefully prevents misinterpretation.
5Constructed Example: Case-Control Interpreting an Odds Ratio
A case-control study compares vegetarian diet (exposed) and non-vegetarian (unexposed) among people with heart disease (cases = 200) and healthy controls (controls = 300). Among cases: 60 vegetarian, 140 non-vegetarian. Among controls: 80 vegetarian, 220 non-vegetarian. Odds of vegetarian among cases = 60/140 = 0.43. Odds among controls = 80/220 = 0.36. OR = 0.43/0.36 = 1.19. This means cases are 19% more likely to be vegetarian—the opposite of expected if vegetarian diets are protective. This contradicts cohort studies showing vegetarians have lower heart disease. Why? Case-control has differential accuracy in recall of diet; people with heart disease may change diet after diagnosis; bias in recruitment. The OR should not be interpreted as proof of association; context and other studies matter.
Odds ratio (OR) compares odds of exposure in cases vs controls; used in case-control studies. For rare outcomes, OR ≈ RR. Hazard ratio (HR) compares event rates over time; used in survival analysis. Both are effect measures but are specific to study designs. Misinterpreting OR as RR leads to incorrect conclusions. Always check study design and interpret accordingly.
A case-control study reports OR = 1.5 for fish consumption and heart disease. Can you interpret this as "eating fish increases heart disease risk by 50%"?
Answer: Not directly. The OR = 1.5 means cases were 1.5 times more likely to be fish consumers than controls. You cannot conclude "risk increased 50%" from case-control data (case-control cannot calculate absolute risk). For a rare disease, OR ≈ RR, so OR = 1.5 suggests ~1.5-fold increased risk (or association). But confounding, recall bias, and selection bias in case-control can distort OR. Context and other study designs (cohorts, RCTs) are needed for valid interpretation.
- Odds ratio (OR) compares odds of exposure in cases vs controls; specific to case-control studies.
- For rare outcomes, OR ≈ RR; for common outcomes, OR overstates risk.
- Hazard ratio (HR) compares event rates over time; used in survival analysis.
- Always identify study design and interpret measures accordingly; mismatching design and measure leads to error.
Next: Lesson 3.10 addresses statistical vs clinical significance—when to care about p-values vs effect sizes.
Statistical vs Clinical Significance
Learning goal: Distinguish statistical significance (a research finding is unlikely due to chance) from clinical significance (a finding is important enough to change practice).
A study might report p = 0.02 (statistically significant) but a tiny effect size (d = 0.1), or p = 0.15 (not statistically significant) but a large effect (d = 0.8). Statistical significance reflects the precision of the estimate and sample size; clinical significance reflects practical importance. Confusing the two leads to over-treatment of insignificant findings and under-appreciation of potentially important effects.
1What Statistical Significance Means (and Doesn't Mean)
Statistical significance (p < 0.05) means the result is unlikely if the null hypothesis (no effect) is true. It says nothing about how important the effect is. A huge study (n = 100,000) can find tiny effects (d = 0.1) that are statistically significant. A tiny study (n = 20) can find large effects (d = 0.9) that are not statistically significant. Sample size and variability determine p-values; true effect size (the magnitude of the phenomenon) is independent of sample size. Confusing statistical significance with importance is one of the most common errors in interpreting research.
2Clinical Significance: Does the Effect Matter in Practice?
Clinical significance asks: is the effect large enough to change clinical practice or individual behavior? Example: a study shows a vitamin supplement increases bone density by 1% (p = 0.01, statistically significant, based on n = 5000). Is 1% density increase enough to change fracture risk? Probably not; a 5–10% increase would be more meaningful. Another study shows a weight-loss diet produces 15 kg average weight loss (p = 0.15, not statistically significant, n = 30). But 15 kg is clinically substantial for most people. Clinical significance depends on the specific outcome and context. What matters for heart disease prevention may differ from what matters for weight loss or cognitive function.
3Minimum Clinically Important Difference (MCID)
MCID is the smallest effect size considered clinically meaningful. For weight loss, MCID might be 5 kg (an amount that most clinicians agree improves health). For bone density, MCID might be 2–3% (associated with meaningful fracture risk reduction). A study finding an effect smaller than MCID is not clinically significant, even if statistically significant. A study finding an effect larger than MCID is clinically significant, even if not statistically significant (due to small sample). Ideally, studies report both p-values and effect sizes in context of MCID so readers can judge clinical importance. A responsible conclusion: "The intervention increased bone density by 1.5% (p = 0.01, 95% CI = 0.5–2.5%), below the MCID of 2%, suggesting limited clinical impact despite statistical significance."
4Practical Thresholds for Interpreting Effect Sizes
Cohen's d (effect size for comparing means): d < 0.2 is trivial or negligible; 0.2–0.5 is small; 0.5–0.8 is medium; > 0.8 is large. But these are rough guidelines. A small effect (d = 0.3) on a population scale might affect millions; the same small effect in an individual's life might be unimportant. Context matters. For rare serious outcomes, even small RR matters (RR = 1.2 for a rare cancer might justify intervention). For common outcomes, large effects are needed (RR = 1.2 for a common condition might not). Clinicians must integrate effect size, baseline risk, costs, and side effects when deciding whether to adopt an intervention.
5Constructed Example: Reconciling Statistical and Clinical Significance
Two studies examine fiber supplementation and cholesterol reduction. Study A (n = 5000): Cholesterol reduced by 5 mg/dL (p = 0.001, d = 0.15, 95% CI = 3–7 mg/dL). Statistically significant but small effect. Clinical impact: a 5 mg/dL reduction is trivial; it might lower 10-year cardiovascular risk by < 1 percentage point. NNT = probably > 200. Not clinically significant. Study B (n = 100): Cholesterol reduced by 25 mg/dL (p = 0.08, d = 0.8, 95% CI = -2 to 52 mg/dL). Not statistically significant (p > 0.05) but large effect and wide CI (could be as small as -2 or as large as 52). Clinical impact: a 25 mg/dL reduction is substantial, potentially lowering 10-year risk by 3–5 percentage points. Clinically important, but statistical uncertainty (wide CI) limits confidence. Study B should be replicated in a larger sample to confirm.
Statistical significance (p < 0.05) means a result is unlikely due to chance. Clinical significance means the result is large enough to matter in practice. The two are independent: a statistically significant effect might be trivially small; a non-significant effect might be clinically substantial. Effect sizes, clinical context, and minimum clinically important differences guide interpretation better than p-values alone.
6Statistical significance versus what changes a client's life
A result can be statistically significant and clinically irrelevant, and large studies make this common: with enough participants, a 200 g difference in weight over twelve weeks reaches significance while meaning nothing to anyone. The reverse also occurs — a genuinely useful effect in a small study may miss significance because the study lacked power, which is not the same as showing the effect is absent.
For Indian practice there is a further filter after clinical significance: affordability and adherence. An intervention that produces a real, meaningful effect but requires imported supplements, unfamiliar foods, or a plate the household will not cook is clinically significant and practically useless for that client. Conversely, a smaller effect delivered by adding curd and soya to meals the family already eats will be sustained for years. The question after “is it real?” and “is it big enough to matter?” is “will this person actually do it?”
A study (n = 10,000) shows a probiotic reduces infection rate from 12% to 11% (RR = 0.92, p = 0.03). Is this clinically significant?
Answer: Probably not, despite statistical significance. The 1 percentage point absolute reduction (ARR = 1%) means NNT = 100. You must treat 100 people with probiotics to prevent one infection. For an healthy population, this is low absolute benefit. The effect is statistically significant (p = 0.03) due to large sample size (n = 10,000), which detects small effects. Clinical significance depends on whether NNT = 100 is acceptable given cost, side effects, and effort. In most contexts, it would not justify universal probiotic use in healthy people.
- Statistical significance (p < 0.05) means a finding is unlikely due to chance.
- Clinical significance means a finding is important enough to change practice.
- The two are independent; sample size, effect size, and context determine whether a finding is both statistically and clinically significant.
- Minimum clinically important difference (MCID) helps distinguish trivial from meaningful effects.
Next: Lesson 3.11 consolidates statistical concepts and their integration into nutrition interpretation.
Chapter Revision
Learning goal: Consolidate statistical concepts and understand how they integrate to interpret nutrition research.
This chapter has built from populations and samples through distributions, hypothesis testing, effect sizes, and risk measures. These concepts are not disconnected; they form an integrated framework for interpreting research. A complete interpretation requires understanding all of them together: populations and sampling (context), measures of central tendency and spread (descriptive), confidence intervals (precision), p-values (significance), effect sizes (magnitude), and risk measures (context-specific).
1The Complete Statistical Picture: From Data to Decision
When you encounter a nutrition research study, ask: (1) Is the sample representative? (population and sampling). (2) Where does the data cluster? How much does it vary? (mean, median, distribution, SD). (3) What is the estimate and its uncertainty? (CI). (4) Is the finding unlikely due to chance? (p-value). (5) How large is the effect? (effect size, RR, OR, HR). (6) Does the magnitude matter clinically? (absolute risk, NNT, MCID). Answer all six and you have a complete interpretation. Skip any one, and you risk misinterpretation.
2Integrating p-Values and Effect Sizes
A study reports: "Intervention X increased protein intake by 8 g/day (95% CI = 5–11 g, p = 0.001, d = 0.7)." This tells you: (1) The estimate is 8 g with high precision (narrow CI). (2) The effect is statistically significant (unlikely to be zero, p = 0.001). (3) The effect is medium-to-large (d = 0.7). Together, these indicate a real, sizable effect. Compare to: "Intervention Y increased protein intake by 3 g/day (95% CI = -1 to 7 g, p = 0.10, d = 0.25)." Estimate is smaller, CI is wider (includes zero), p-value is not significant, effect size is small. The two interventions are clearly different in their evidence strength. Neither p-values nor effect sizes alone suffice; both matter.
3Relative vs Absolute: Choosing the Right Frame
For evaluating research claims, ask whether absolute or relative comparisons make sense. For rare outcomes (rare cancers, rare side effects), relative measures (RR, OR) dominate—the absolute number is too small to contextualize. For common outcomes (common diseases, common symptoms), absolute measures (absolute risk, NNT) are more meaningful. A claim that a food "cuts cancer risk by 30%" (high relative reduction) might mean 0.1 percentage point absolute reduction if cancer is rare (NNT = 1000). The same "30% reduction" for a common outcome (obesity, ~30% prevalence) might mean a 9 percentage point reduction (NNT = 11). Frame your interpretation accordingly.
4Study Design and Choice of Measures
Different study designs require different statistical measures. RCTs report RR or risk differences directly. Case-control studies report OR (not RR, because they cannot directly calculate risk). Cohort studies can report RR or HR depending on whether time-to-event analysis is used. Observational studies cannot determine causation (confounding, bias), only associations. When you see an OR in a case-control study, interpret it as an odds ratio, not a risk ratio (unless the outcome is rare). When you see HR in a survival analysis, interpret it as hazard rate, not absolute risk. Matching study design to the statistical measure prevents misinterpretation.
5Confidence and Caution in Interpreting Research
The framework provides tools for rigorous interpretation, but certainty is rare. A single study, however well-designed, has limitations (sample size, populations studied, unmeasured confounding). Confidence increases with replication: multiple studies with consistent findings, across populations and study designs, strengthen conclusions. Red flags: (1) single study claimed as proof, (2) p-value without effect size or CI, (3) relative risk without absolute risk, (4) study conflicts with existing body of evidence, (5) findings do not make biological sense. Conversely, strong evidence: (1) consistent findings across multiple studies, (2) clear effect sizes and CIs, (3) plausible mechanism, (4) dose-response relationships, (5) findings hold across subgroups.
6Applying Statistics to Your Own Judgments
As a nutrition expert, you will use these concepts daily: reviewing papers, evaluating interventions, advising clients. For each claim, systematically apply the framework: What is the population studied? What is the effect size? Is it statistically significant? What is the absolute change? Does it matter clinically? What is the number needed to treat? How does this finding fit into the broader literature? Over time, these questions become automatic. You will quickly identify weak claims (small sample, tiny effect, p-value hunting) and strong claims (large effect, narrow CI, consistent with prior evidence). This statistical literacy is a core competency of a nutrition expert.
A complete statistical interpretation integrates populations and sampling, descriptive statistics (central tendency, spread), inferential statistics (CIs, p-values), effect sizes (magnitude), and context-specific risk measures. No single statistic suffices; all together provide a full picture. Statistical literacy—understanding what these measures mean and how to use them together—is essential for interpreting nutrition research.
A study reports: supplement increased vitamin D levels by 15 ng/mL (p = 0.04, 95% CI = 1–29 ng/mL, d = 0.4). Summarize what this means using at least four statistical concepts.
Answer: (1) Sample estimate is 15 ng/mL increase (point estimate). (2) 95% CI = 1–29 ng/mL, indicating precision/uncertainty (CI is somewhat wide; true effect could be as small as 1 or as large as 29). (3) p = 0.04 indicates statistical significance (unlikely to be zero due to chance, p < 0.05). (4) Effect size d = 0.4 is small-to-medium magnitude. Interpretation: The supplement likely increases vitamin D (statistically significant), the effect is modest (small-to-medium), and there is moderate precision in the estimate. Absolute clinical importance depends on baseline vitamin D levels and whether a 15 ng/mL increase is clinically meaningful in the study population.
- Complete statistical interpretation integrates populations, descriptive stats, CIs, p-values, effect sizes, and risk measures.
- p-values and effect sizes are complementary; both are needed.
- Relative and absolute risk communicate different aspects; both inform decisions.
- Study design determines appropriate statistical measures; match them correctly.
- Statistical literacy—systematic use of these concepts—is essential for interpreting research as a nutrition expert.
Next: Lesson 3.12 applies these concepts to real nutrition research case studies.
Statistics Interpretation Cases
Learning goal: Apply statistical concepts to interpret real nutrition research, distinguishing hype from evidence.
This lesson presents five case studies applying Chapter 3 concepts to published nutrition research findings.
1Case Study 1: The Fiber and Cholesterol Study—Effect Size and Clinical Importance
A randomized trial (n = 320) tests soluble fiber supplementation on cholesterol. Results: Total cholesterol decreased by 12 mg/dL in the fiber group (mean = 188 mg/dL) vs 2 mg/dL in the control group (mean = 198 mg/dL), difference = 10 mg/dL, 95% CI = 4–16 mg/dL, p = 0.003, d = 0.3. Statistical analysis: p = 0.003 (statistically significant). d = 0.3 (small effect). CI does not include zero (significant). Interpretation: The finding is real and statistically significant, but the effect size is small. A 10 mg/dL cholesterol reduction might lower 10-year cardiovascular risk by < 1 percentage point in a low-risk population (absolute benefit minimal). In a high-risk population (existing heart disease), 10 mg/dL might reduce risk by 1–2 percentage points (NNT = 50–100). The effect is statistically significant but clinically modest, and cost-effectiveness depends on population risk. Fiber's other benefits (gut health, satiety) might justify use anyway, but cholesterol reduction alone is not compelling. This is a study showing statistical significance without strong clinical impact.
2Case Study 2: The Vegetarian Mortality Study—Relative and Absolute Risk
A large cohort study (n = 60,000) follows vegetarians and meat-eaters for 10 years, recording mortality. Results: Vegetarians had 12 deaths per 1000 person-years; meat-eaters had 18 deaths per 1000 person-years. RR = 12/18 = 0.67 (33% lower risk), 95% CI = 0.54–0.82, p < 0.001. Statistical analysis: Statistically significant (p < 0.001, CI excludes 1.0). Effect size: RR = 0.67 is substantial. Absolute risk: Baseline risk is 18 per 1000 (1.8% over 1 year; higher over 10 years). New risk is 12 per 1000 (1.2% over 1 year). ARR = 0.6 per 1000 per year (or 6 per 10,000 per year). Relative reduction is 33% (impressive), but absolute reduction over 1 year is 0.6 per 1000 (modest; NNT = 1667). Over 10 years, with higher baseline mortality in older populations, absolute benefit is larger (perhaps 6 per 100 over 10 years, NNT = 17). Interpretation: The finding is real and substantial in relative terms. Absolute benefit depends on population age/risk. In younger populations, absolute benefit is modest; in older populations, larger. The study demonstrates association (vegetarianism linked to lower mortality), but causation is not certain due to confounding (vegetarians may exercise more, smoke less, have better healthcare access).
3Case Study 3: The Probiotics and Infections Study—Multiple Comparisons and P-Hacking
A trial (n = 200) tests a probiotic on various health outcomes: infection rate, diarrhea, constipation, abdominal discomfort, bloating, fatigue, sleep quality. The paper reports: infection rate reduced by 20% (p = 0.08, not significant), diarrhea reduced by 15% (p = 0.12, NS), constipation reduced by 30% (p = 0.02, significant!), abdominal discomfort reduced by 25% (p = 0.05, borderline), bloating reduced by 35% (p = 0.01, significant!), fatigue reduced by 12% (p = 0.20, NS), sleep quality improved 10% (p = 0.15, NS). Red flags: (1) Seven outcomes tested but only three are significant. (2) p-values are not corrected for multiple comparisons (Bonferroni would be 0.05/7 ≈ 0.007; only the bloating result survives). (3) Primary outcome is not pre-specified; the authors may have highlighted post-hoc significant findings. (4) Effect sizes are not reported; we don't know if the 30% and 35% reductions are large (d > 0.5) or small (d < 0.2). Interpretation: This study has high risk of false positives due to multiple testing. The results (constipation and bloating improvement) might be chance findings. The paper should have pre-specified primary outcomes and corrected p-values for multiple comparisons. Without replication in an independent sample, these results should be viewed with skepticism.
4Case Study 4: The Coffee Consumption and Cancer Study—Confounding and Observational Bias
A case-control study (n = 500 cases with lung cancer, 500 controls without) compares coffee consumption. Results: Cases consumed average 3 cups/day; controls consumed 2 cups/day. Odds of high coffee consumption (≥3 cups/day) among cases = 60/440 = 0.136. Odds among controls = 40/460 = 0.087. OR = 0.136/0.087 = 1.56, p = 0.03 (statistically significant). Interpretation: The OR = 1.56 suggests coffee drinkers have 56% higher odds of lung cancer. But confounding is likely: (1) Smokers drink more coffee than non-smokers (coffee and smoking are correlated). (2) Smoking is a strong causal risk factor for lung cancer. (3) The study does not account for smoking in the analysis or analysis stratification. (4) The association might be entirely due to smoking confounding (coffee consumption is a proxy for smoking). Case-control studies cannot prove causation; they can only show associations that may be confounded. This result should not change coffee recommendations without (a) adjustment for smoking in analysis, (b) replication in studies of non-smokers, (c) biologically plausible mechanism. Most epidemiologic and mechanistic evidence suggests coffee is not a lung cancer cause. This study highlights why observational associations mislead without careful attention to confounding.
5Case Study 5: The Supplement Efficacy Trial—Narrow CI but Small Sample
A study (n = 40 participants with low energy) tests an energy-boosting supplement. Results: Energy score improved from 35 to 48 (±15 SD) in the supplement group, and from 35 to 40 (±14 SD) in the control group. Difference = 8 points, 95% CI = 0–16 points, p = 0.05, d = 0.55. Statistical analysis: p = 0.05 (borderline significant, just at threshold). CI is narrow and just excludes zero. Effect size d = 0.55 (small-to-medium). Interpretation: The finding is statistically significant at the conventional p < 0.05 threshold but is borderline (p = 0.05 is not strongly significant; p = 0.001 would be stronger). The CI being 0–16 means the true effect could be as small as 0 (no effect, CI boundary) or as large as 16 (substantial effect). The small sample (n = 40) and modest effect size raise concern about false positives (publication bias: small studies with null results go unpublished; we see only the positive study). The result is suggestive but not confirmatory. Before recommending this supplement, one should ask: Is there replication in an independent sample? Was the study pre-registered (preventing selective reporting)? Is the mechanism biologically plausible? Without answers to these, the single positive study in a small sample remains preliminary.
- For a nutrition claim you encounter (supplement, diet, food), identify the source study or systematic review.
- Ask: Is it a single study or multiple studies? Single studies are preliminary; systematic reviews are stronger.
- For the study: identify sample size, populations studied, effect size, p-value, CI.
- Calculate or estimate: absolute risk or absolute benefit (if possible), number needed to treat (NNT).
- Judge: Is the effect statistically significant (p < 0.05)? Is it clinically significant (effect size and absolute benefit meaningful)?
- Check: Is the finding consistent with prior evidence and biological plausibility?
- Conclude: Integrate all evidence. Does it change recommendations?
6Three statistics cases
A news report states that a common cooking oil “doubles heart disease risk”. Checking the paper: a relative risk near 2.0 in an observational cohort, with an absolute difference of a few events per thousand person-years, in a non-Indian population, unadjusted for physical activity. The honest summary for a client is that the signal is weak, the absolute difference small, and the population different. A supplement claims a “statistically significant increase in strength”. The increase is 1.2 kg on a leg-press after eight weeks, with wide confidence intervals.
A client with prediabetes asks whether switching some rice for millets is worth it. The trial evidence is limited and the effect sizes modest, but the baseline risk in an Indian adult with prediabetes is high, the change costs nothing, the food is culturally familiar, and adherence is likely to be good. The correct answer weighs all four — not the p-value alone.
A study reports a diet supplement "cuts weight gain in half" (RRR = 50%). Baseline weight gain is 4 kg/year. New weight gain is 2 kg/year. What is absolute weight reduction and NNT?
Answer: Absolute weight reduction (ARR) = 4 - 2 = 2 kg/year. RR = 2/4 = 0.5. RRR = (1 - 0.5) × 100% = 50% (matches the claim). NNT cannot be calculated directly from weight change (NNT applies to binary outcomes like disease/no disease). But the 2 kg/year reduction is substantial and may be clinically meaningful for long-term weight management. The 50% relative reduction is real, and the 2 kg absolute reduction is the clinically relevant piece.
- Systematic application of statistical concepts (sample size, effect size, p-value, CI, absolute benefit) distinguishes strong from weak evidence.
- Single studies, especially with small samples or multiple outcomes, are preliminary; replication strengthens confidence.
- Confounding in observational studies can bias results; randomized trials provide stronger evidence of causation.
- Integrating statistical evidence with biological plausibility and existing literature prevents over-interpretation of single studies.
Next: Chapter 4 shifts from research interpretation (Chapters 1–3) to dietary assessment, measurement, and implementation of evidence-based nutrition practice.