Early mathematics is the strongest predictor of school success. That does not mean it is easy to teach

Six longitudinal datasets showed that mathematics at school entry predicts later achievement more strongly than reading does. But the same literature explains why the effects fade — and that knowledge is even more useful to a director.
Duncan and colleagues' 2007 paper "School readiness and later achievement" pooled six longitudinal datasets: ECLS-K (analytic sample 21,260 children), NLSY (1,756), NICHD SECCYD (approximately 907–981), IHDP (690), the Montreal Longitudinal-Experimental Preschool Study (767) and the 1970 British Birth Cohort Study (approximately 9,000–10,000). In regressions controlling for family background and prior ability, the standardised coefficients for the latest reading and mathematics outcomes were: school-entry mathematics 0.33; reading 0.13; attention 0.07. Externalising problems were 0.00 and social skills 0.01 — the socioemotional behaviour measures had essentially no predictive power. The paper's own conclusion: among behavioural measures, "only attention-related skills predicted later academic achievement with any consistency", and their coefficient was less than a quarter the size of the mathematics coefficient.
There are two common errors in interpreting this result. The first is reading it as "mathematics is what matters, the rest does not". The coefficients are correlational; they control for family background but not for unobserved ability. The second is confusing prediction with malleability. Bailey, Duncan, Odgers and Yu (2017) make exactly this point: many interventions show "initially promising but then quickly disappearing impacts", and durable skill-building must target skills that are simultaneously malleable, fundamental, and not ones that would have developed anyway. Watts, Duncan, Siegler and Davis-Kean (2014) showed that mathematical ability at 54 months predicted mathematics achievement through age 15 — but the same paper found that growth in mathematical ability between 54 months and first grade was an even stronger predictor than the 54-month level. What matters is the trajectory, not the point.
What should a setting choose? Programmes with independent effectiveness ratings exist. Building Blocks: three cluster RCTs meeting US What Works Clearinghouse standards — Clements and Sarama (2008) 202 children, 28 classrooms, effect 1.07; Clements and colleagues (2011) 1,305 children, 106 classrooms, 0.48; Hofer and colleagues (2013) 1,714 children, 139 classrooms, 0.55. Weighted summary: 3,221 children, 273 classrooms, effect 0.58, improvement index +22 percentile points. Pre-K Mathematics: five cluster RCTs, 2,913 children, effects across seven measures from 0.23 to 0.72, summary 0.38, +15 percentile points — and the WWC's highest rating, "strong evidence of effectiveness". But honesty is required here: Building Blocks' 1.07 came from a 36-classroom trial run by the developers themselves. At scale it fell to 0.48–0.55. So quote 0.58, not 1.07. Number Worlds has no What Works Clearinghouse report — there is no basis for publishing specific effect sizes for it.
The cheapest and most available instrument is the teacher's own speech. Klibanoff, Levine, Huttenlocher, Vasilyeva and Hedges (2006) studied 26 classrooms in 13 schools, recording each of 26 head teachers for one hour. Mathematical utterances in that hour ranged from 1 to 104 (mean 28.3, SD 24.2). In a three-level hierarchical model, the coefficient for teacher maths input was γ = 0.026 (SE 0.010, t = 2.508, p = 0.033) — statistically significant. Syntactic complexity of teacher speech (p = 0.855), classroom quality (p = 0.120) and socioeconomic status (p = 0.242) were all non-significant. Most importantly: the amount of teacher maths input was uncorrelated with the classroom mean maths score in the autumn (r = 0.001, p = 0.996) — teachers were not simply responding to children's starting level. This is a teacher-behaviour target, not a child-readiness issue. The caveat: this is a correlational study of 26 classrooms with one hour of speech coded per classroom, not an experiment.
Which kind of talk counts? Gunderson and Levine (2011) are precise: counting or labelling the numerosity of present, visible objects predicted children's later cardinal-number knowledge, while other types of number talk did not. Moreover, talk about visible sets of 4 to 10 predicted more robustly than talk about smaller sets. On Klibanoff's data, cardinality already accounts for 48% of teacher maths talk — so the thing to increase is set size, not frequency. For children who are behind there is a replicable protocol: Dyson, Jordan and Glutting (2013), N = 121, 8 weeks, 3 days per week, 30-minute small-group sessions, 24 sessions in total, targeting counting, comparing and manipulating sets. Gains held at immediate and delayed post-test. Finally, the highest-value action: planning the hand-over to Grade 1. Kang and colleagues (2019) showed most fadeout comes not from children forgetting but from the control group catching up; Clements and colleagues' study across 42 schools found g = 0.51 with follow-through and g = 0.28 without. Passing each child's assessed level to the receiving first-grade teacher is the single most effective step available to a director.


