How Superscoring and Retake Policies Distort Admissions and Amplify Wealth Inequality
When colleges use the best scores from multiple test attempts—a practice known as superscoring—they intend to provide a more accurate picture of a student's potential. However, this policy may inadvertently turn standardized testing into a game of chance that rewards those with the deepest pockets. A new study from researchers at the Georgia Institute of Technology suggests that superscoring doesn't just measure ability; it allows affluent students to "harvest" lucky results from repeated attempts. This creates a systematic bias that masks true academic talent.
Standardized testing serves as a critical signal in university admissions. It is intended to act as an exogenous measurement (a signal independent of the observer) of a student's latent ability. Traditionally, these scores were treated as static snapshots of knowledge. But as institutions move toward more flexible aggregation rules, the nature of the signal has changed. The current landscape is dominated by two philosophies: Single-Sitting (SS), which evaluates a student's best complete exam attempt, and Superscoring (SC), which independently combines the best section scores (such as Math and Verbal) across multiple attempts.
The researchers argue that these policies are not passive measurements but active incentives. Because retaking exams incurs both temporal and financial costs, students behave strategically. They allocate their limited study effort and money in response to how the university calculates their final score. This creates a tension between statistical precision and social equity that current models fail to address.
The Failure of Current Aggregation Rules
The status quo presents a fundamental trade-off between signal fidelity (the accuracy of the measurement) and applicant participation. Under the Single-Sitting rule, universities preserve the integrity of the signal by looking at holistic attempts. However, the authors find that this imposes a heavy "simultaneous preparation burden" on students. Because a student cannot "bank" a high Math score and a high Verbal score from different days, they must master all subjects at once. This requirement disproportionately harms students who lack the resources to prepare for every subject simultaneously. This may exclude high-ability candidates from the pool entirely.
Superscoring attempts to solve this by allowing students to specialize. If a student performs well in Math on Monday but poorly in Verbal, they can retake the exam to focus solely on Verbal. While this encourages broader participation, the authors report it introduces a severe "screening distortion." As shown in, wealth-dependent costs create a massive gap in effort capacity.
Low-wealth students (Q1) exert significantly less effort, particularly in second-round preparation, compared to their wealthier peers (Q5).
This leads to what the paper calls the "Safety Net Effect." Because Superscoring protects a student's best section scores, it reduces the perceived risk of a bad test day. The authors demonstrate that this safety net actually discourages precautionary effort in the first round. Students realize they can simply "fix" a weak section later through a retake.
Harvesting Noise Through Strategic Effort
The mechanism driving this distortion is a combination of strategic effort allocation and "order-statistic selection." In statistics, an order statistic is a value selected from a set, such as the maximum or minimum. Superscoring functions by selecting the maximum value from each subject's independent noise draws across multiple attempts.
The authors model this as a two-stage optimization problem. In stage one, a student decides how much to study for the initial exam. In stage two, they decide whether to retake it based on their first-round results and their remaining budget. Under Superscoring, if a student achieves a high score in Math due to a lucky, low-noise draw, they do not need to study Math again. Instead, they concentrate all remaining effort into the deficient subject. This "effort specialization" is clearly visible in the simulation data .
Crucially, this process allows students to decouple their scores from their actual proficiency. By taking multiple tests, a student isn't necessarily learning more. They are simply increasing their chances of encountering a favorable random fluctuation. The authors formalize this as "harvesting noise." Because wealthier students face lower fixed costs for retaking exams, they can afford to play this game more often. This creates the "Access Premium"—a measurable score advantage held by wealthy students that is entirely independent of their underlying academic ability, as illustrated in .
Measuring the Algorithmic Advantage
The paper quantifies these distortions using simulations calibrated to 2025 College Board data. The most striking result is the impact on fairness. The authors report that uncorrected Superscoring can drive the False Negative Rate (FNR)—the rate at which qualified students are incorrectly rejected—to a staggering 76.0% for the lowest wealth quintile (Q1) [Table 1]. This means a vast majority of qualified low-income students are missed by the algorithm.
The researchers also highlight a phenomenon they call "signal degradation." Even though Superscoring results in higher observed scores, those scores are less accurate reflections of true ability. The authors demonstrate that under Superscoring, expected submitted scores rise even as total expected learning falls. This creates a systematic upward bias where the observed score artificially overstates true ability.
To visualize the scale of this advantage, the authors present the "Access Premium" across different ability levels. shows that students with above-median wealth consistently extract higher scores than their equally qualified, below-median wealth peers. This gap is not a byproduct of better schooling alone. It is a direct result of the scoring algorithm interacting with wealth-dependent retake costs. Furthermore, shows that while Superscoring inflates scores for everyone, the premium is most aggressive for high-wealth groups.
This acts as an "inequality accelerant."
Limits of Demographic-Blind Corrections
The authors propose three interventions to mitigate these issues: Full-Information Score Correction (FISC), a Variance Tax (TAX), and Contextual Correction (CTX). However, they find that not all mathematical fixes are created equal.
The FISC and TAX methods are "demographic-blind" structural rules. FISC subtracts a penalty based on the number of attempts. TAX uses a weighted average of high and low scores to compress the impact of variance. While these methods are effective at reducing the "noise premium," the authors report they hit a structural ceiling. Because they do not account for the fact that low-wealth students have different effort capacities, they cannot fully resolve the underlying inequality. For example, under FISC, the Q1 share of the admitted class remains significantly lower than their true ability share [Table 2].
The third method, Contextual Correction (CTX), is a post-processing intervention. Instead of changing how scores are calculated, it applies group-specific penalties to equalize the False Negative Rates across wealth cohorts. This method targets "Equal Opportunity" rather than simple statistical unbiasedness. The authors report that CTX is remarkably successful. It stabilizes error rates across the population. It also helps the Q1 cohort achieve an admitted class share (6.0%) that closely matches their true ability (5.0%) [Table 2].
The Verdict: A Choice Between Precision and Equity
The research concludes that there is no perfect scoring rule. Institutions face a fundamental, unavoidable trade-off between statistical precision and fair outcomes. Superscoring maximizes participation by lowering the barrier to entry for specialized study. However, it pays for this with degraded signal accuracy and amplified wealth gaps. Single-Sitting preserves the signal but risks excluding talented, resource-constrained students.
For practitioners designing admissions algorithms, the verdict is clear. If the goal is purely to maximize the average ability of the admitted class, Superscoring may appear efficient. However, if the goal is to ensure that high-ability students from all socioeconomic backgrounds have an equal chance of admission, demographic-blind statistical corrections are insufficient. The authors suggest that "demographic-aware structural interventions," like the Contextual Correction, are necessary. These prevent the scoring algorithm from becoming a mechanism that merely validates existing wealth disparities. Code for the simulation framework is reportedly available; see the paper for the canonical link.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 95% (passed)
Claims verified: 15 / 15
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 154,787
Wall-time: 286.1s
Tokens/s: 541.0