ts / Timur Salakhetdinov

P-009 · 2026 · Hackathon entry

Fuel price forecasting with TabPFN-3.5 and LLM event features

A hackathon study testing whether event features extracted from a weekly energy-market narrative improve TabPFN-3.5 fuel price forecasts beyond price history and publication timing.

TabPFN-3.5Jev extractionTime seriesPlacebo testEIA data

Study design

Weekly EIA gasoline prices and This Week in Petroleum issues, as of each forecast date
Issue timingWhen an issue was published, without its content
Jev event featuresDriver, outlook, region and disruption labels
PlaceboThe same features shuffled among published issues

Same TabPFN-3.5 model, rows, context, seed and fit schedule

Price forecasts one to four weeks ahead, scored by MASE
Comparing Jev features with the placebo separates what the text says from when it was published.

Question and context

I built this study for the Prior Labs TabPFN-3.5 Hackathon. It asks whether structured event features, extracted by Jev from the weekly This Week in Petroleum narrative, improve TabPFN-3.5 forecasts of U.S. retail gasoline prices beyond price history and publication timing. The data covers 29 weekly EIA price series and 1,146 narrative issues from 2002 to 2024. Each extracted label, such as market driver, stated price outlook, region or supply disruption, keeps the sentence it came from.

Method

TabPFN-3.5 forecasts price changes one to four weeks ahead, with quantiles. The same model runs on calendar and price history, then adds issue timing, Jev features or raw issue text, all with the same rows, context, seed and fit schedule. A placebo keeps the timing but shuffles the extracted content among issues already published. Features are strictly as-of each forecast date. The protocol was fixed on 2005 to 2015 before a frozen holdout from 2016 to January 2024 was scored, with a paired moving-block bootstrap over origins.

Results

The holdout has 418 weekly origins, 29 series and 48,484 forecasts per arm. Jev features lowered MASE relative to timing alone by 0.127 (95% CI 0.050 to 0.209) and helped in all 29 series. The shuffled placebo did just as well, however (difference +0.013, CI −0.036 to +0.057), so the gain cannot be attributed to what the text says. Without event columns, TabPFN-3.5 was less accurate than a random walk, and its 80% intervals covered only 67 to 73% of outcomes.

Scope

This is a retrospective study: features are as-of, but pretrained models may carry hindsight about 2002 to 2024. The narrative restates recent prices, and EIA discontinued it in January 2024, so the workbench cannot run live. Two corrections made after the holdout run, to interval scoring and to HTML left in 104 issues, are documented and leave the answer unchanged.

Holdout MASE by arm

2.32.42.52.62.7Random walk: 2.440Calendar and price history: 2.662Calendar and price history2.662Plus issue timing: 2.528Plus issue timing2.528Plus Jev features: 2.401Plus Jev features2.401Plus raw issue text: 2.388Plus raw issue text2.388Placebo: shuffled Jev features: 2.388Placebo: shuffled Jev features2.388MASE averaged over 29 series (lower is better)
Frozen holdout, 418 weekly origins from 2016 to January 2024. Event features beat timing alone, but the placebo matches them, so the improvement comes from the extra columns rather than the meaning of the text.

View public code and documentation on GitHub ↗

Timur Salakhetdinov ·