Dead Reckoning: Why Current Customer Count Estimates are Fundamentally Unverifiable
Companies often estimate how many customers are still "active" by guessing if they will ever return. This is difficult in non-contractual commerce—like retail—where customers do not cancel subscriptions. They simply stop buying without saying goodbye. For decades, businesses have relied on "Buy-Till-You-Die" (BTYD) models. These models calculate a "probability of being alive" ($P(\text{alive})$) for every customer. This metric populates churn-risk dashboards and drives corporate valuations.
A new study from Karl T. Ulrich at The Wharton School reveals a fundamental flaw. The paper argues that $P(\text{alive})$ is not a verifiable fact. Instead, it is an infinite-horizon extrapolation. This is a mathematical guess about a future that never arrives. Because this value assumes a customer will eventually purchase given infinite time, it relies on assumptions the data cannot prove. Consequently, companies may report customer counts that are wildly inaccurate. In one setting, estimates varied by a factor of 7.6x.
The Category Error in Churn Dashboards
The status quo treats $P(\text{alive})$ as a direct proxy for current activity. However, the authors argue that practitioners commit a "category error." In the Beta-Geometric/Negative Binomial Distribution (BG/NBD) family of models, $P(\text{alive})$ is a mathematical limit. It is the limit of the probability that a customer will return within $H$ months, as $H$ approaches infinity.
A forecast for a finite window (e.g., "will this customer buy in 12 months?") is a verifiable event. The infinite limit is an extrapolation. The paper finds this distinction leads to massive discrepancies. In a consumer segment of 31,683 customers, the authors report a huge spread in estimates. Reported counts ranged from 3,654 to 27,734 .
Even a default weighting parameter (a ridge penalty used to prevent overfitting) swung the count by 42 percent.
Most perceived "miscalibration" is actually this error in kind. The authors show that while a model's 18-month forecast might be accurate, the summed $P(\text{alive})$ is not. In the MakerStock consumer segment, summed $P(\text{alive})$ overshoots actual 18-month returners by 2.25x .
Navigating by the Stars and the Grid
To solve this, the paper proposes moving away from "dead reckoning." Dead reckoning is navigating by assuming a constant speed from a last known position. Instead, the authors suggest a "navigation" approach based on regular fixes. They propose a three-part remedy centered on observability.
First, change the target variable. Instead of reporting $P(\text{alive})$, firms should report $R_H$. This is the probability of at least one purchase within a stated, finite horizon $H$. This turns an unobservable limit into a verifiable forecast.
Second, the authors introduce a "calibration audit" using a vintage-by-horizon grid. Analysts should not check accuracy on a single holdout set. Instead, the audit grades $R_H$ forecasts against actual outcomes. It checks various points in time (vintages) and various lengths of time (horizons). This helps distinguish between model errors and market shifts.
Third, the paper introduces a "dynamic calibration layer." This borrows logic from actuarial science. It uses a "loss-development triangle" (a method to estimate future claims from current trends). It assumes that while customer volume changes, the "shape" of customer maturity remains stable. The system learns this shape from history. It then uses the freshest, partially observed data to estimate the current level. This helps the model adjust to "drift" (changes in customer behavior over time).
Results from the Field and Benchmarks
The authors validate this using a seven-year transaction panel from MakerStock and the CDNOW benchmark. The results show that current flaws are widespread. On the MakerStock data, different model specifications produced nearly identical short-term forecasts. Yet, they resulted in vastly different total customer counts .
The study also shows that "patience" only provides a floor for the truth. In a five-year audit, realized returners only push the lower bound of the count upward. They fail to provide a reliable upper bound .
The proposed dynamic calibration layer is highly effective. In a thirteen-quarter test, the layer repaired failures in earlier models. It tracked reality through regime changes. As shown in, the layer maintains high accuracy across different quarters. Meanwhile, raw $P(\text{alive})$ scores drift significantly. The layer also acts as an early warning system. The "drift statistic" ($D$) flagged a regime change two quarters before traditional validation could detect it .
Limits of the Extrapolation
The paper is transparent about its boundaries. For the BG family, the findings are mathematically exact. For other models like the Pareto/NBD, $P(\text{alive})$ acts as an upper bound.
The dynamic calibration layer relies on an empirical conjecture. It assumes customer cohorts mature along a stable shape. This holds in the tested datasets. However, it would require more testing in highly volatile industries. The study also does not compare these models against discriminative machine learning models. The authors suggest their audit framework would apply to them. Finally, the research focuses on purchase counts. It does not explore how these errors affect revenue or profit.
The Verdict: Stop Guessing, Start Auditing
If you use these numbers for budgeting, the verdict is clear. Do not report $P(\text{alive})$ as a point estimate of your customer base.
The evidence suggests $P(\text{alive})$ is "partially identified." The data tells you the minimum number of customers you have. It cannot tell you the maximum. Instead, adopt this protocol: 1. Report $R_H$: Use the probability of return within a specific, auditable window. 2. Use Intervals: If you must report a total count, report it as a range. 3. Audit the Grid: Monitor performance across a grid of timeframes. 4. Randomize Interventions: If you trigger marketing spend based on these scores, use randomized control groups. This ensures you can measure the actual effect of your actions.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 98% (passed)
Claims verified: 15 / 15
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 100,952
Wall-time: 238.6s
Tokens/s: 423.1