How often do AI systems give 2026 advice with 2025 numbers?
A free, open evaluation for AI financial-advice systems. Every task carries an answer key built from primary-source-verified figure tables — with the wrong-but-plausible stale values enumerated — so "how stale is this model" is scored mechanically, and published with its protocol, warts included.
40 tasks × 4 systems × 3 attempts, protocol pre-registered before any model was called. pass@1 with 95% CI; pass^3 requires all three attempts correct; rg+tc restricts to the rule-grounding and threshold/cliff families; the Stale Figure Rate (SFR) is the share of 75 figure-bearing attempts asserting a stale or fabricated figure as current. Every number cites the methodology page.
| System | pass@1 [95% CI] | pass^3 | rg+tc pass@1 | Stale Figure Rate | Floor pass^3 |
|---|---|---|---|---|---|
| claude-sonnet | 0.90 [0.83–0.94] | 0.85 | 0.84 | 0.07 (5/75) | 1.00 |
| gemini-3.1-pro-preview | 0.57 [0.48–0.65] | 0.42 | 0.33 | 0.17 (13/75) | 0.50 |
| deepseek-v4-pro | 0.48 [0.40–0.57] | 0.35 | 0.27 | 0.28 (21/75) | 0.40 |
| claude-haiku | 0.36 [0.28–0.45] | 0.33 | 0.11 | 0.63 (47/75) | 0.30 |
Named-system caveats, published rather than smoothed: the Claude systems ran as session aliases (exact dated model strings not exposed by the harness); the Gemini system is a preview endpoint; delivery was batched (declared deviation); N=40. The full pre-registration and the post-freeze substitution log — including a harness truncation bug that forced re-execution of the entire Claude arm — are on the protocol page.
Only one of four systems passed every floor task — the "a competent system must never miss these" tier — on all three tries. In-context memory was the opposite story: saturated at 1.00 for all four systems. The gap is figure currency, not recall.
The errors are fluent, cited, and repeat identically across attempts — pass^3 ≈ pass@1 on every system. Each named mode has its own forbidden-figure reason code in the eval corpus, so it stays countable.
A superseded published figure asserted as current — the 2025 (or 2024) IRA limit cited fluently, with a source, as 2026 fact. The dominant mode for the weakest system and material for two others.
Correct statute, correct method, wrong figure: the model computes a plausible value from an indexing rule or stale base instead of recalling the published number. The signature frontier-model failure — observed in three of four systems.
Two systems asserted the repealed TCJA estate-tax sunset where current law sets a $15M exclusion — a categorical wrong answer on a multi-million-dollar question, not a rounding error.
Next cycle's not-yet-published figure given as a number. Plausible, specific, and uncheckable by a reader who doesn't know the notice hasn't been issued.
In our pre-registered pilot of record — 40 routine 2026 U.S. planning questions, 4 systems from 3 labs, no tools, 3 attempts each — 29% of figure-bearing attempts (86 of 300) asserted a stale or fabricated regulatory figure as current. The per-system Stale Figure Rate ranged from 7% to 63%.
The share of figure-bearing answers in which an AI system asserts a stale, superseded, or fabricated regulatory figure as current — scored mechanically against enumerated forbidden values with reason codes, not judged by another model. It makes 'is our AI giving this year's advice' a number instead of an anecdote.
Because many U.S. planning figures are inflation-indexed by a known rule, a model can apply the rule to a stale base and produce a specific, plausible, wrong number — correct statute, correct method, wrong figure. This 'derived, not published' failure appeared in three of the four systems tested, including the strongest.
No. On every system tested, pass^3 approximately equals pass@1 — wrong answers repeat near-identically across attempts. Stale figures are a systematic property of a frozen model, not sampling noise, which is why re-rolling the same question is not a mitigation.
50 tasks, free, no auth: 35 rule-grounding + 15 threshold-cliff, floor/standard difficulty. Fetch with ?withhold_answers=true for blind evaluation, then score against the full file. The 30-line scoring contract is in the response envelope.
All adversarial "categorical flip" tasks, all decision-recall tasks, and a 70-task reserve — including every one of the 40 pilot tasks — are never published and never sold. That is what keeps this scoreboard re-runnable and the commercial packs meaningful.
The 25 figure-bearing pilot tasks re-run with the live Rule Sets feed (us-federal-2026, 103 cited figures) rendered into context as a lookup — byte-identical for every system, protocol pre-registered before execution. The same three systems that produced 65 stale or fabricated figures in 225 bare attempts produced zero in 225 with-lookup attempts. Full protocol, substitution log, and the honest caveats: methodology page.
| System | Bare pass@1 (rg+tc) | Bare SFR | With feed pass@1 | With feed pass^3 | With feed SFR |
|---|---|---|---|---|---|
| claude-sonnet | 0.84 | 0.07 (5/75) | 1.00 | 1.00 | 0.00 (0/75) |
| gemini-3.1-pro-preview | 0.33 | 0.17 (13/75) | 1.00 | 1.00 | 0.00 (0/75) |
| claude-haiku | 0.11 | 0.63 (47/75) | 0.96 | 0.92 | 0.00 (0/75) |
| deepseek-v4-pro | 0.27 | 0.28 (21/75) | — | — | — not run |
What the feed did NOT fix, published on principle: the weakest system's four remaining failures were rules-comprehension errors made with the correct figures in front of it (stacking catch-up limits; treating the HSA age-55 catch-up as poolable between spouses). A figures feed makes a system current, not competent — which is why the eval packs test both. deepseek-v4-pro was not re-run (no API credential in the run environment); its bare-arm numbers stand unpaired.
Bare, the systems above cited stale figures in 29% of figure-bearing attempts. With the live feed, zero. Measure your own system with the eval packs; close the gap with Rule Sets.