October 4, 2026
| Claude model | Training cutoff | Obs. | Blind | Blind, with outcome |
|---|---|---|---|---|
| Haiku 4.5 | 31 July 2025 | 60 | 11 | 11 |
| Sonnet 5 | 31 January 2026 | 60 | 11 | 11 |
| Fable 5.0 | 31 January 2026 | 60 | 11 | 11 |
| Opus 5 | 31 May 2026 | 60 | 8 | 8 |
| Sonnet 5.5 | 30 June 2026 | 60 | 8 | 8 |
| Opus 5.5 | 30 June 2026 | 60 | 8 | 8 |
Notes: Each model forecast 60 observations over the 3 rounds: four variables at five horizons in each round. A forecast is blind when the target's outcome was published after the model's training cutoff. Sonnet 5.5, Opus 5.5 and Opus 5 have the fewest because their cutoffs are the latest: the 2026Q1 outcomes predate them.
So far 57 forecasts have a published outcome, all from the backfilled rounds. The largest miss is CPI inflation for 2026Q2, which came in at 6.07 percent. Claude's average miss was 2.77 points and the SPF's 1.74.
Live track: forecasts against the published outcomes
| Variable | Target | Obs. | Outcome | Claude error | SPF error | Winner |
|---|---|---|---|---|---|---|
| CPI inflation | 2026Q2 | 12 | 6.07 | 2.772 | 1.735 | SPF |
| Real GDP growth | 2026Q1 | 3 | 1.99 | 0.085 | 0.660 | Claude |
| Real GDP growth | 2026Q2 | 12 | 1.50 | 0.457 | 0.601 | Claude |
| Unemployment rate | 2026Q1 | 3 | 4.33 | 0.006 | 0.067 | Claude |
| Unemployment rate | 2026Q2 | 12 | 4.27 | 0.091 | 0.153 | Claude |
| 3-month Treasury bill | 2026Q1 | 3 | 3.59 | 0.008 | 0.057 | Claude |
| 3-month Treasury bill | 2026Q2 | 12 | 3.62 | 0.094 | 0.045 | SPF |
| All pooled | 57 | 0.724 | 0.575 | SPF | ||
| By model, the same observations | ||||||
| Sonnet 5.5 | 8 | 0.594 | 0.634 | Claude | ||
| Opus 5.5 | 8 | 0.590 | 0.634 | Claude | ||
| Fable 5.0 | 11 | 0.715 | 0.532 | SPF | ||
| Opus 5 | 8 | 0.924 | 0.634 | SPF | ||
| Sonnet 5 | 11 | 0.732 | 0.532 | SPF | ||
| Haiku 4.5 | 11 | 0.771 | 0.532 | SPF | ||
Notes: Forecasts with a published outcome, from the 2026Q1 and 2026Q2 rounds. Each observation is one model's forecast of the target from one round, the three draws averaged, where the outcome was published after the model's training cutoff. Errors are mean absolute errors against the first published estimate. Claude is closer than the SPF median on 5 of the 7 targets, and pooled over the 57 observations its error is 0.724 against the SPF's 0.575.