P-009 · 2026 · Hackathon entry
Fuel price forecasting with TabPFN-3.5 and LLM event features
A hackathon study testing whether event features extracted from a weekly energy-market narrative improve TabPFN-3.5 fuel price forecasts beyond price history and publication timing.
Study design
Same TabPFN-3.5 model, rows, context, seed and fit schedule
Question and context
I built this study for the Prior Labs TabPFN-3.5 Hackathon. It asks whether structured event features, extracted by Jev from the weekly This Week in Petroleum narrative, improve TabPFN-3.5 forecasts of U.S. retail gasoline prices beyond price history and publication timing. The data covers 29 weekly EIA price series and 1,146 narrative issues from 2002 to 2024. Each extracted label, such as market driver, stated price outlook, region or supply disruption, keeps the sentence it came from.
Method
TabPFN-3.5 forecasts price changes one to four weeks ahead, with quantiles. The same model runs on calendar and price history, then adds issue timing, Jev features or raw issue text, all with the same rows, context, seed and fit schedule. A placebo keeps the timing but shuffles the extracted content among issues already published. Features are strictly as-of each forecast date. The protocol was fixed on 2005 to 2015 before a frozen holdout from 2016 to January 2024 was scored, with a paired moving-block bootstrap over origins.
Results
The holdout has 418 weekly origins, 29 series and 48,484 forecasts per arm. Jev features lowered MASE relative to timing alone by 0.127 (95% CI 0.050 to 0.209) and helped in all 29 series. The shuffled placebo did just as well, however (difference +0.013, CI −0.036 to +0.057), so the gain cannot be attributed to what the text says. Without event columns, TabPFN-3.5 was less accurate than a random walk, and its 80% intervals covered only 67 to 73% of outcomes.
Scope
This is a retrospective study: features are as-of, but pretrained models may carry hindsight about 2002 to 2024. The narrative restates recent prices, and EIA discontinued it in January 2024, so the workbench cannot run live. Two corrections made after the holdout run, to interval scoring and to HTML left in 104 issues, are documented and leave the answer unchanged.