Teacher–child interaction: what the CLASS instrument measures, and what it does not

CLASS is the most widely used instrument for measuring the quality of a preschool classroom. It rates three domains — and the interesting part is which domain scores low almost everywhere.
CLASS (Classroom Assessment Scoring System), developed by Robert Pianta, Karen La Paro and Bridget Hamre, has one defining feature: it does not measure materials, curriculum or ratios — it measures interaction only. The Pre-K version has three domains and ten dimensions. Emotional Support: positive climate, negative climate (reverse-scored), teacher sensitivity, regard for child perspectives. Classroom Organization: behaviour management, productivity, instructional learning formats. Instructional Support: concept development, quality of feedback, language modelling. Each dimension is scored 1–7: 1–2 low, 3–5 middling, 6–7 consistently effective.
The most important empirical fact is the shape of the scores. US Head Start national data for 2018–19 (159 grant recipients): Emotional Support 6.05; Classroom Organization 5.79; Instructional Support 2.91. By dimension: concept development 2.43; quality of feedback 2.88; language modelling 3.42. The 2019–20 figures (78 recipients) are near-identical: 6.03 / 5.78 / 2.94. The pattern is stable across years and across studies.
How should this be read? Warmth, safety and order are already in place in most kindergartens. Open-ended questioning, inviting prediction, extending rather than closing a child’s answer, connecting a concept to the child’s own experience — that is the weak link everywhere. The headroom is there, not in the emotional domain.
Now the caution. The link between CLASS scores and children’s outcomes is weak. In Perlman and colleagues’ meta-analysis (PLOS ONE, 2016; 19 studies, 15,167 children, 14 separate meta-analyses) only two associations reached significance: Classroom Organization with inhibitory control (r = 0.06) and Instructional Support with social skills (r = 0.09). The authors described the associations as "quite limited". Margaret Burchinal’s 2018 review found linear effect sizes across 53 studies ranging from −0.04 to 0.19, only ten of which differed statistically from zero. A large 2023 meta-analysis (185 studies, 229,697 children) reported process-quality correlations of r = 0.09 for literacy, 0.09 for mathematics and 0.13 for behavioural skills. In the same study, structural quality — ratios, group size, credentials — showed no significant association with any child outcome.
Practical conclusions. First: use CLASS as a coaching lens, not a verdict. Ranking teachers on scores is not supported by the predictive evidence; using dimension-level scores to choose a teacher’s next focus is. Second: direct effort at Instructional Support, specifically concept development and quality of feedback. Third: the one well-replicated route to higher scores is one-to-one coaching on the teacher’s own recorded practice. In Pianta and colleagues’ 2008 study of 113 teachers, those receiving individualised consultation and feedback showed significantly greater gains in independently rated interaction quality. Fourth: do not adopt numerical cut-offs as an internal standard — the authors of the threshold study themselves declined to recommend minimums for practice. Fifth: a better interaction rating is a means, not the outcome. Experiments have raised measured quality without moving children’s outcomes. Pair any such initiative with a direct measure of children’s language and early mathematics.


