Just because an AI passes a rigorous benchmark doesn't mean it will work well in a real-world job. In the current landscape of AI evaluation, we often treat a high score on a reasoning battery (a set of diverse tests) as a direct proxy for professional competence. However, this assumes that the "intelligence" measured in a controlled setting translates perfectly to messy, multi-step workflows.
The core problem is that AI evaluation is rarely a single jump from test to task. It is a chain of inferences. We observe a benchmark score. We interpret it as a capability. We project that capability onto a specific application. Finally, we assume human oversight will mitigate any remaining errors. This paper argues that even if every individual link in that chain is scientifically sound, the chain itself can still break. The authors call this the "non-composition principle": support for adjacent links does not automatically guarantee support for their union.
The breakdown of inferential chains
Most current work in AI evaluation focuses on "construct validity" (the degree to which a test measures the specific theoretical trait it claims to). These frameworks ask whether a benchmark actually measures the capability it claims to, such as mathematical reasoning. While these frameworks are a significant improvement over treating benchmarks as black boxes, they often stop at the boundary of the benchmark itself. They tell us what the score means. They do not tell us if that meaning survives the transition to a real-world deployment.
As shown in, evaluation arguments are typically sequences of projections.
A benchmark might support a claim about a model's ability to answer supplied-text questions. However, a company might use that model in a retrieval-augmented generation (RAG) pipeline (a system that pulls external documents to ground its answers) to produce legal memos. The paper identifies a fundamental epistemic gap here. The target of the first study is not necessarily the source of the next. When these "endpoints" do not align in terms of the object, population, conditions, or outcomes, the composition of the two studies becomes undefined. You cannot mathematically multiply a "reasoning score" by a "human review success rate" if the two studies look at different populations or different types of errors.
Auditing the interfaces of projection
To solve this, the authors propose a "projectibility audit." This is a framework designed to diagnose these unsupported joins. Instead of viewing a benchmark as a monolithic proof of capability, the audit treats every step of the inference as a "projection edge" connecting "empirical nodes."
The authors define an empirical node as a structured record consisting of five specific fields: $N = \langle \text{object, population, conditions, outcome, period} \rangle$. By typing these nodes, the auditor can explicitly check for alignment. For example, if a benchmark node specifies "outcome: item correctness" but the deployment node specifies "outcome: substantive defect in a final memo," the auditor flags a mismatch. A bridge is then "owed." This means a new study must be conducted to connect these two disparate outcomes.
The audit process moves through two primary layers of verification: 1. Endpoint Alignment: Ensuring the two links actually meet. This requires continuity in the object (the system being used), the population (the types of cases encountered), the conditions (the prompts and tools), the outcome (the metric being tracked), and the period (the timeframe and model version). 2. Warrant Transmission: Even if the endpoints align, the auditor must check if the "warrant" (the evidentiary support) actually passes through the join. This involves verifying that assumptions remain compatible. It also requires checking that uncertainty and dependencies—such as shared data lineage or model training overlaps—are propagated rather than reset at the interface.
Aggregation masks the risks of deployment
A particularly subtle failure mode identified by the paper is how statistical aggregation can destroy the very evidence needed for safe deployment. In many evaluations, researchers report a single mean accuracy score. The authors argue that this is dangerously reductive. A stable mean can hide massive instabilities at the item level (the individual test questions).
Through a reanalysis of MMLU-Pro data (a more challenging version of the Massive Multitask Language Understanding benchmark), the authors demonstrate this via three hypothetical states in .
In all three states, the mean change ($L$) is zero. This suggests perfect stability on average. However, State B represents a scenario where half the items improve and half deteriorate. This indicates high instability. State C represents a "stable-poor" scenario. Here, the mean is stable, but a specific subset of items carries a massive, persistent risk of failure.
The paper reports that for certain models, specifically a reanalysis of gpt-5.4 on MMLU-Pro, a small mean decline of 2.15 percentage points can coexist with a "worst-decile degradation" (the drop in performance for the bottom 10% of items) as high as 36.80%. This discrepancy is critical. The "average" performance is a poor predictor of "tail risk." Tail risk refers to the specific, high-impact failures that a professional user, such as a lawyer, actually encounters. To combat this, the authors suggest using metrics like Item Instability (INS) and Worst-Tail Degradation (WTD) to ensure the resolution of the source report is sufficient for the downstream projection.
Limits of the projectibility framework
While the framework is theoretically robust, I notice two significant limitations. First, the paper admits that projectibility is not a complete theory of induction (the process of deriving general principles from specific observations). It relies on the evaluator's ability to identify "plausible defeaters." These are the specific conditions that might break a projection. If an evaluator lacks deep domain expertise, the audit will fail to flag the most critical gaps.
Second, the proposed audit is computationally and logistically expensive. Moving from "aggregate scores" to "projectibility declarations" requires collecting item-level outputs. It also requires maintaining version control over entire software stacks. Finally, it requires conducting targeted sampling of real-world tasks. For many developers, the cost of providing this level of transparency may currently outweigh the perceived benefit.
Verdict: A necessary shift toward granular accountability
The paper's central thesis is correct. We have been treating AI evaluation as a series of independent successes rather than a precarious chain of inferences. The non-composition principle is a vital warning. It warns against the "accumulation of coverage" fallacy. This is the idea that passing more domains automatically grants unrestricted generality.
For engineers and researchers, the takeaway is clear. Stop reporting naked means. If you are building a system where the cost of a single error is high, a benchmark score is almost useless. It must be accompanied by a projectibility declaration. This declaration must specify the exact boundaries of its reach. Code for the statistical modules and the audit framework is reportedly available. See the paper for the canonical link at https://github.com/BrettRey/benchmark-inference-composition. Use it to move away from "does it work?" toward "under exactly which conditions is this warrant preserved?"
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: lesswrong_skeptic
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 97% (passed)
Claims verified: 11 / 11
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 122,942
Wall-time: 218.5s
Tokens/s: 562.7