Making AI Act Like a Real Student
Why do some students struggle with specific math concepts while others master them easily? In educational testing, answering this requires collecting massive amounts of data from real people. This process is slow, expensive, and difficult to scale. Researchers have turned to Large Language Models (LLMs) to act as "simulated examinees." These models provide instant, cheap data to help calibrate new tests. However, current AI models are often too smart and too uniform. They behave like a room full of straight-A students rather than a diverse classroom.
A new study from the University of California, Berkeley, and the University of Minnesota proposes a way to fix this. The authors report that by using a framework called Cognitive Diagnostic Profiling (CDP), they can nudge LLMs to mimic the diverse cognitive strengths and weaknesses of actual human students. This approach allows developers to test the difficulty of new exam questions using AI that finally reflects the messy reality of human learning.
Bridging the Gap Between AI and Ability
At its core, the research addresses a fundamental problem in psychometrics (the science of measuring mental attributes). To build a reliable test, developers must perform "item calibration." This involves estimating how difficult each question is and how much it distinguishes between different levels of ability. Traditionally, this requires running a "pilot" of the test on hundreds of human students.
The authors identify two primary reasons why simply asking an LLM to "take a test" fails to replicate human data. The first is "less alignment," a location bias where models act as far more capable than the target population. For instance, the paper notes that many LLMs reach the 95th to 99th percentile of human performance. They essentially act as "super-students." The second is "less variability," a compression of the spread of responses. Instead of a wide bell curve of abilities, LLM responses often cluster into a narrow band of near-perfect scores.
To solve this, the researchers move away from simple persona prompting. Instead, they move toward a structured model of cognition.
The Mechanics of Cognitive Diagnostic Profiling
The authors propose the Cognitive Diagnostic Profiling (CDP) framework. It treats an examinee not as a single number on an ability scale, but as a collection of mastered and unmastered skills. Think of it like a character sheet in a role-playing game. Instead of just having a "strength" score, a character has specific proficiencies in "climbing" or "archery."
As shown in, the CDP framework follows a six-step process.
First, the researchers decompose a test construct into discrete binary attributes. These are specific cognitive skills, such as "borrowing from a whole number" in fraction subtraction. Next, they form "mastery patterns." These are binary vectors representing every possible combination of mastered and unmastered skills. For a test with five attributes, there are 32 unique ways a student could potentially know those skills.
These patterns are then rendered into natural language. Rather than feeding the model a raw bitstring like [1, 0, 1, 1, 0], the framework creates a descriptive profile. It might say: "The student has mastered basic subtraction but lacks the ability to borrow from whole numbers."
The study tests two ways of assembling the simulated pool of students [Figure 2b]: 1. Uninformative CDP (C2): Every possible mastery pattern is sampled with equal probability. This ensures a wide spread of simulated "students." 2. Informative CDP (C3): Patterns are sampled based on how common they are in a real human population. This uses data from previous studies.
Finally, the LLM generates responses based on these specific profiles. The resulting data is analyzed using standard psychometric models.
From Uniform Peaks to Diverse Classrooms
The impact of this structured profiling is visible in how the simulated populations look. In the baseline condition where no profile is provided (C1), the LLM's ability distribution forms a single, narrow peak of high achievers .
By applying CDP, the researchers successfully forced the models to spread out.
Under the uninformative condition (C2), the simulated abilities began to separate into distinct sub-distributions based on how many skills were mastered . The "informative" condition (C3) went a step further. It shifted the entire pool to match the actual prevalence of skills seen in humans. The authors report that for the Gemini 3.1 Pro (Thinking) model, the overlap with the human distribution rose from 0.52 in the baseline to 0.81 under the informative condition.
Crucially, the framework doesn't just get the "average" right. It gets the individual profiles right. As demonstrated in, the simulated mean scores for each of the 32 mastery patterns tracked almost perfectly with human expectations.
Weighted correlations reached between 0.92 and 0.98.
This precision carries over to the test items themselves. When the models are properly profiled, they can accurately recover the difficulty of individual questions. For the Gemini 3.0 Flash (Thinking) model, the authors report that the Spearman correlation for item difficulty rose from 0.24 in the baseline to as high as 0.90 under CDP [Table 3]. This means the AI can effectively tell the difference between a "medium" and a "hard" question.
Using this method is highly economical. The authors report that costs ranged from less than \$0.03 to about \$14.11 per 100 simulated examinees. This is orders of magnitude lower than collecting data from human respondents.
Limits of the Simulation
While the results are promising, the authors highlight several boundaries. The study was conducted on a single, relatively small 15-item mathematics instrument. Its performance on complex or much longer assessments remains untested.
There is also the risk of "data contamination." Since the fraction-subtraction dataset used for validation is public, the items might have been in the LLMs' training data. This could artificially inflate their accuracy. Furthermore, the "informative" condition (C3) uses data from the same population it is being tested against. The authors admit this represents an "upper bound" on how much help population information can provide.
Finally, the authors note a distinction between psychometric and cognitive alignment. While CDP makes the results look like human results, it does not guarantee the AI uses human-like mental processes. The simulation matches the behavior, but it does not necessarily replicate the human thought process.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: explainer
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 95% (passed)
Claims verified: 12 / 12
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 107,448
Wall-time: 231.3s
Tokens/s: 464.5