An effect size tells you how large a difference or relationship is, while a p-value addresses how compatible the data are with a statistical model such as a null hypothesis. To interpret an effect size, identify the measure, find its null value, read its direction and units, inspect its confidence interval, and then judge whether the magnitude matters in the study’s real context. Never translate a number into “small” or “large” from a cutoff alone.
This guide is for university students reading quantitative research for an assignment, literature review, journal club, or thesis. You will learn a repeatable five-step method, see a worked example, and get language you can use without overstating a result.
An effect size is a numerical description of magnitude. It may express a difference in the original units, a difference standardized by variability, a relationship between variables, or a relative comparison of event rates. The American Statistical Association statement on p-values emphasizes that a p-value does not measure effect size or the importance of a result. That is why “p < .05” cannot answer “How much?”
The clearest measure is often a raw effect. If one group averages 78 points and another averages 72 points, the mean difference is 6 points. Readers can compare that difference with the assessment scale, a passing threshold, or a meaningful educational target. Sullivan and Feinn distinguish these absolute effects from standardized effects, which help when the original scale is unfamiliar or when studies use different scales.
The reading rule: magnitude first, uncertainty second, context always.
Start by locating the exact label in the results, table, or figure. Different measures answer different questions, and their null values are not all the same.
A mean difference stays in the outcome’s original units. A value of +6 points means the first named group averaged 6 points more than the comparison group, assuming the subtraction order is first group minus second group. The null value is 0. Ask whether 6 points is important on that specific scale.
Cohen’s d and related measures express a mean difference in standard-deviation units. A d of 0.50 means the two group means differ by half a standard deviation. Conventional reference points of 0.20, 0.50, and 0.80 are often labeled small, medium, and large, but the practical primer by Daniël Lakens explains why the chosen calculation and study design matter. Treat those labels as rough orientation, not verdicts.
A correlation such as Pearson’s r runs from −1 to +1. The sign gives direction; the distance from 0 gives the strength of the linear relationship. An r of −0.40 indicates that higher values of one variable tend to accompany lower values of the other. It does not, by itself, establish that either variable caused the other.
Risk ratios and odds ratios compare relative quantities. Their null value is 1, not 0. A risk ratio of 0.75 means the event risk in the numerator group is 0.75 times the risk in the comparison group, a relative reduction of 25%. Odds are not the same as probabilities, so do not rewrite an odds ratio as a risk ratio. Cochrane Handbook Chapter 6 details which effect measures suit binary, continuous, count, and time-to-event outcomes.
This context-first approach aligns with the Institute of Education Sciences effect-size guidance, which highlights that reported effect sizes are influenced by study and outcome features and must be interpreted for the decision at hand.
Imagine a fictional study comparing a six-week feedback program with standard instruction. The outcome is a writing score from 0 to 100. The paper reports: “Program students scored 6 points higher on average; standardized mean difference d = 0.48, 95% CI [0.12, 0.84], p = .009.”
A defensible sentence is: “Students offered the program scored 6 points higher on average than students receiving standard instruction. The standardized estimate was d = 0.48, but the confidence interval ranged from 0.12 to 0.84, so the data are compatible with effects from modest to fairly substantial.”
That sentence reports the original units, the standardized magnitude, and the uncertainty. It does not say the program “proved” improvement, nor does it claim that d = 0.48 is universally important. To judge importance, you would need information such as the assessment’s reliability, the score difference educators consider meaningful, participant assignment, attrition, and program cost.
The point estimate is one best estimate from this sample; the interval shows that precision is limited. Because the interval does not include the difference-measure null value of 0, it corresponds to a conventional two-sided test below .05 under the same model. More importantly, the lower and upper bounds could support different practical decisions. Precision is therefore part of interpretation, not a footnote.
Before you summarize a result, answer these questions:
For a literature review, keep one compact record for each paper: outcome, measure, estimate, interval, group order, and one sentence of contextual meaning. Snitchnotes can help turn these structured notes into practice questions, such as identifying the null value or correcting an overstated interpretation.
Use this template: “For [outcome], [group or exposure] was associated with [direction and raw magnitude]. The reported [effect-size measure] was [estimate], with a [confidence level] confidence interval from [lower] to [upper]. In this context, the effect may be [practical interpretation], although [important uncertainty or design limitation].”
Keep causal verbs such as caused, improved, and reduced for designs and analyses that justify them. For observational work, safer verbs include was associated with, was related to, or differed by. This distinction is separate from effect size: a precisely estimated association can still be noncausal.
No. An effect size describes the magnitude of a difference or relationship. Statistical significance summarizes how the observed data relate to a statistical model and is affected by factors including sample size. A small effect can have a small p-value in a large sample, while an important effect can be estimated imprecisely in a small sample.
There is no universal cutoff. Cohen’s d values of 0.20, 0.50, and 0.80 are common reference points, but their practical meaning depends on the outcome, population, costs, risks, measurement quality, and field norms. Compare the estimate with meaningful thresholds and similar high-quality studies before applying a label.
A negative sign indicates direction relative to the stated group order or variable coding; it does not automatically mean harm. A negative mean difference in pain may favor the intervention, while a negative difference in achievement may not. Check which group was subtracted from which and whether higher outcome values are desirable.
The point estimate alone hides uncertainty. A confidence interval shows the range generated by the method and data under its assumptions, helping readers see whether both trivial and important magnitudes remain compatible with the results. Wider intervals indicate less precision, so the bounds may support different practical interpretations.
Sometimes, but only with care. Standardized measures can place different scales on a common metric, yet comparisons may still be distorted by different populations, outcome reliability, follow-up times, study designs, and calculations. Verify that the effect-size definitions are compatible and use a systematic review or meta-analysis when a formal synthesis is needed.
To interpret effect sizes in research papers, identify the measure, locate its null value, decode its direction, read the confidence interval, and judge the magnitude against the study’s real context. This method is more informative than attaching a label or repeating a p-value. On your next paper, use the 60-second checklist and record the estimate, interval, units, and meaning together; Snitchnotes can then turn that record into focused review questions.
Notes, quiz, podcasts, flashcards et chat — en un seul upload.
Essaie ta première note gratuitement