ts / Timur Salakhetdinov

P-008 · 2026 · Hackathon entry

TabPFN-3.5 vs Bayesian Pareto/NBD customer models

A benchmark of when the TabPFN-3.5 foundation model beats specialised Bayesian customer models at predicting whether customers return within 90 days.

TabPFN-3.5Pareto/NBDHierarchical BayesLightGBMStreamlit

Study design

Customer purchase histories at monthly forecast origins
Pareto/NBD · Abe (2009) HBNo labels; frequency, recency and customer age
TabPFN-3.5Labelled examples in context; no tuning
GLM · LightGBMLabelled training; LightGBM grid on development data

Common inputs, then a predeclared expanded feature set

90-day inactivity risk and purchase count per customer
The protocol was frozen on development origins; three untouched holdout origins were then scored once, with a customer-clustered paired bootstrap.

Question and context

I built this benchmark for the Prior Labs TabPFN-3.5 Hackathon. It asks when TabPFN-3.5, learning in context from labelled examples, predicts repeat buying better than the Pareto/NBD model and Abe's (2009) hierarchical Bayesian extension, which need no labels. At monthly forecast origins it predicts, for every existing customer, the probability of no purchase in the next 90 days and the number of purchase occasions in that window.

Method

The untuned TabPFN-3.5 classifier and regressor are compared with base-rate and recency references, logistic and Poisson GLMs, tuned LightGBM, Pareto/NBD by maximum likelihood and Abe's model by MCMC, which is scored only if it converges. Supervised models are compared on the Pareto/NBD inputs only and on a predeclared expanded feature set. The protocol was frozen on development origins before three untouched holdout origins were scored once, with a customer-clustered paired bootstrap and learning curves from 50 to 2,000 labelled customers.

Results

On Online Retail II with the Pareto/NBD's own three inputs, the label-free Pareto/NBD beats TabPFN-3.5 on Brier score (0.169 vs 0.177) and is better calibrated. With expanded features, TabPFN-3.5 ties with the GLM, LightGBM and Pareto/NBD on Brier score, has the best PR-AUC and the lowest purchase-count error, and is the strongest supervised model at every learning-curve size. On CDNOW, the hierarchical Bayesian model converges and wins on the common inputs, while the Bayesian models keep the best count forecasts. In short, TabPFN-3.5 wins when it can use richer features and when labels are scarce.

Scope

The data is noncontractual, so 90-day inactivity is not churn, and the results rank risk rather than estimate the effect of contacting a customer. The study covers two public retail datasets and overlapping holdout origins. The repository includes a local Streamlit workbench, saved predictions and one command to regenerate every table.

Learning curve: how many labels does TabPFN-3.5 need?

TabPFN-3.5GLMLightGBMPareto/NBD (no labels)
0.170.190.210.23501002505001,0002,000Pareto/NBD, no labels: 0.169Labelled training customers (log scale)Brier score (lower is better)
Online Retail II holdout, expanded features, mean of three origins and three seeds. TabPFN-3.5 leads every supervised model at every size, with the largest gap at 50 to 250 customers, and reaches the label-free Pareto/NBD at about 2,000.

View public code and documentation on GitHub ↗

Timur Salakhetdinov ·