An Appraisal of Small Area Estimation Methods to Explore Expansion of State Subgroup Reporting on NAEP
From Large Scale Assessments in Education, 14, 32 (2026).
NAEP is the only assessment that lets policymakers compare student achievement across all 50 states, including for individual student groups. But a long-standing reporting rule requires at least 62 sampled students before a state-level average can be published, a safeguard against reporting unreliable numbers that has the unintended effect of frequently suppressing results for demographic minority groups.
Rather than testing more students, the authors ask whether statistical modeling can fill the reporting gap: are Small Area Estimation (SAE) estimates built from samples below 62 more accurate than the direct estimates NAEP would actually publish from sample of exactly 62 students?
Using 2019 grade-8 NAEP mathematics data for Black students in Georgia, North Carolina, South Carolina, and Virginia, the authors treated each state's full sample as a pseudo-population and drew repeated small samples from it. Simulated sample sizes of 43 ("typical"), 20 ("challenging"), and 10 ("extreme") were chosen because 11 states missed the 62 threshold for this subgroup in 2019, with a median obtained sample of 43 and a minimum of 20.
Roughly 16 model variants were compared—Fay–Herriot and a novel Random Forest–Fay–Herriot hybrid, two unit-level Bayesian models (including a new dynamic-borrowing approach), and three Empirical Bayes shrinkage schemes—scored by mean absolute difference from NAEP's published values, against a benchmark of direct estimates from samples of 62.
Findings
- At n = 43, SAE beat the benchmark on accuracy in about 94% of comparisons, and 13 of 16 models outperformed the benchmark in every state — with uniformly greater stability.
- The strongest performer was the Random Forest–Fay–Herriot variant using prior-cycle NAEP achievement as a predictor, with average errors of roughly 0.4 to 2.2 scale points versus 3.0 to 5.3 for the benchmark, a meaningful margin given that some researchers treat 3 NAEP points as educationally significant.
- Accuracy degraded as samples shrank, but even at n = 20 the models still outperformed the benchmark; at n = 10 several statistical models still outperformed on accuracy, though reliability advantages disappeared in some cases.
The results cut two ways. They suggest NAEP could expand subgroup reporting without lowering its accuracy standard, and they imply the rule-of-62 itself may not guarantee well-powered estimates.
The authors also argue agencies still relying on classical Fay–Herriot variants should in more novel statistical modeling extensions. At the same time, they are explicit that SAE complements rather than replaces adequate sampling, and direct estimates should remain the basis for reporting whenever samples are large enough.
This article is published open access. Analyses used restricted-use NAEP data, available under license from the National Center for Education Statistics, along with publicly available state assessment proficiency data from EDFacts. Statistical code is available in a public repository.