When large AI models are compressed to run faster on smaller devices, they often pass standard safety tests. However, they may start volunteering harmful stereotypes when asked open-ended questions. This compression is called quantization. It shrinks the numerical precision of a model's weights to save memory. While engineers assume this step is harmless, a new study from the University of Southern California reveals a significant side effect. The researchers call this a "selective safety gap." In this gap, a model's ability to refuse direct harm remains intact. Meanwhile, its tendency to generate biased content increases.
The illusion of passing safety checks
The industry standard for evaluating Large Language Model (LLM) safety relies on short-form safeguards. These typically check if a model refuses a direct request to perform a harmful task. They also check if it selects an unbiased answer in a multiple-choice format. The authors of the QuantiBias study find these traditional metrics are blind to open-ended generation.
As shown in, researchers tested a quantized version of the Qwen3.6-27B model.
The standard safeguards remained remarkably stable. The model's refusal rate for harmful prompts stayed near full-precision levels. Its accuracy on multiple-choice bias benchmarks like BBQ also remained stable. However, the authors report a different story for open-ended questions. The same model volunteered stereotypes in roughly one in four answers. This discrepancy suggests a major risk. A model can pass every standard safety checklist and still reach users in a more biased state.
Why precision loss hits bias hardest
To understand this, the authors propose a mechanistic account based on "decision margins." In an LLM, every token selection is determined by the difference in confidence (called logits) between possible outputs. A "margin" is the gap between the preferred, safe answer and a potentially biased one.
The researchers argue that quantization acts as a semantically empty perturbation. This means it adds mathematical noise to every weight without targeting any specific behavior. This noise affects all behaviors, but it does so unevenly. High-margin behaviors, such as refusing an overtly harmful request, are reinforced heavily during alignment training. Because the gap between "refuse" and "comply" is wide, the noise from quantization is unlikely to flip the decision.
Conversely, avoiding subtle stereotypes in open-ended conversation is a "low-margin" behavior. These responses were not direct targets of intensive training. Therefore, the model's internal confidence in staying neutral is quite thin. As quantization reduces the effective bits per weight (bpw)—the actual amount of information stored per parameter—the noise easily pushes these narrow margins into stereotype endorsement. The authors note that current quantizers prioritize precision for capability-related data. This includes coding or logic, which are used during calibration. Consequently, the fragile, bias-sensitive weights are left under-protected.
Evidence of a widening gap
The study provides evidence for this gap across multiple model families and languages. Using the QuantiBias benchmark, the authors demonstrate that the bias is not an artifact of a single language.
For the Qwen3.6-27B anchor, the stereotype endorsement rate remains high. Under an independent judge, the rate sits between 23.8% and 26.7%. This occurs even as the model's capabilities and refusal rates stay flat . This finding was replicated on a second backbone, Gemma-4-31B. There, the independent judge similarly flagged high rates of bias [Figure A2]. This happened despite stable performance on standard benchmarks.
The researchers also examined if "reasoning"—forcing the model to deliberate before answering—could act as a safeguard. They find the effect is highly dependent on the model family. On the Qwen backbone, enabling reasoning roughly halved the bias rate .
It also flattened the increase seen during compression. However, on the Gemma backbone, reasoning provided almost no benefit. It left the bias levels virtually unchanged .
Limits of the current findings
The authors highlight several areas where the research is currently bounded. First, absolute bias levels depend heavily on the "judge" used for scoring. An "in-family" judge (a model from the same family as the one being tested) reports much lower bias levels. This differs from independent, out-of-family judges like Claude or Gemini. Practitioners must be cautious when using a model family to evaluate its own safety.
Second, the study notes a distinction between frequency and severity. The current instrument primarily measures how often a model crosses the line into a stereotype. It does not fully capture how much more "decisive" or extreme those stereotypes become as precision drops. Finally, the multilingual aspect of the probe uses a diagonal design. In this design, language, target group, and content are somewhat entangled. This may complicate direct comparisons between different linguistic contexts.
The verdict on quantized deployment
If you are deploying quantized models, do not rely solely on standard safety benchmarks. A model may appear perfectly aligned in multiple-choice tests. It may also pass refusal checks. Yet, it may still exhibit significant bias during natural, open-ended dialogue.
The authors recommend a dedicated re-evaluation for open-ended generative bias for every quantized build. Developers cannot treat "thinking" models as a universal fix. The effectiveness of reasoning as a mitigator is unpredictable across different architectures. For those implementing these evaluations, the QuantiBias code and datasets are reportedly available via Hugging Face.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 95% (passed)
Claims verified: 15 / 15
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 163,167
Wall-time: 277.4s
Tokens/s: 588.3