P-008 · 2026 · Hackathon entry
TabPFN-3.5 vs Bayesian Pareto/NBD customer models
A benchmark of when the TabPFN-3.5 foundation model beats specialised Bayesian customer models at predicting whether customers return within 90 days.
Study design
Common inputs, then a predeclared expanded feature set
Question and context
I built this benchmark for the Prior Labs TabPFN-3.5 Hackathon. It asks when TabPFN-3.5, learning in context from labelled examples, predicts repeat buying better than the Pareto/NBD model and Abe's (2009) hierarchical Bayesian extension, which need no labels. At monthly forecast origins it predicts, for every existing customer, the probability of no purchase in the next 90 days and the number of purchase occasions in that window.
Method
The untuned TabPFN-3.5 classifier and regressor are compared with base-rate and recency references, logistic and Poisson GLMs, tuned LightGBM, Pareto/NBD by maximum likelihood and Abe's model by MCMC, which is scored only if it converges. Supervised models are compared on the Pareto/NBD inputs only and on a predeclared expanded feature set. The protocol was frozen on development origins before three untouched holdout origins were scored once, with a customer-clustered paired bootstrap and learning curves from 50 to 2,000 labelled customers.
Results
On Online Retail II with the Pareto/NBD's own three inputs, the label-free Pareto/NBD beats TabPFN-3.5 on Brier score (0.169 vs 0.177) and is better calibrated. With expanded features, TabPFN-3.5 ties with the GLM, LightGBM and Pareto/NBD on Brier score, has the best PR-AUC and the lowest purchase-count error, and is the strongest supervised model at every learning-curve size. On CDNOW, the hierarchical Bayesian model converges and wins on the common inputs, while the Bayesian models keep the best count forecasts. In short, TabPFN-3.5 wins when it can use richer features and when labels are scarce.
Scope
The data is noncontractual, so 90-day inactivity is not churn, and the results rank risk rather than estimate the effect of contacting a customer. The study covers two public retail datasets and overlapping holdout origins. The repository includes a local Streamlit workbench, saved predictions and one command to regenerate every table.