wealthschema / ai-eval-sets
AI MODEL EVALUATION · FINANCIAL ADVICE

Is your AI giving 2026 advice — or 2025 advice?

Vintaged evaluation packs for AI financial-advice systems. Every task keyed to a primary-source-verified figure; every wrong-but-plausible prior-year value enumerated, so stale answers are countable, not anecdotal.

401 tasks · TY2026 vintage Answer keys by construction Mechanical scoring — no LLM judge
29%

of figure-bearing attempts · 86/300 · 4 systems · k=3

Every frozen model goes stale by January.

In a pre-registered 40-task evaluation across four systems from three labs, 29% of figure-bearing answers cited a stale or fabricated regulatory figure — last year's IRA limit with this year's confidence, an estate-tax "sunset" that Congress repealed, a catch-up limit derived from the statute instead of recalled from the notice. The errors are fluent, cited, and repeat identically across attempts. Only one of the four systems passed every must-never-miss floor task on all three tries.

Full protocol, per-system numbers, and the substitution log: the methodology page · live scoreboard: /benchmark.

THE METRIC

What a Stale Figure Rate is

The Stale Figure Rate is the share of figure-bearing answers in which a system asserts a stale, superseded, or fabricated regulatory figure as current. It is scored mechanically against enumerated forbidden values — each with a reason code and the vintage that produced it — so no judge model sits between your system and the number. It turns "is our AI current" into a metric you can gate a deployment on and put in a compliance file. Full definition, formula, and reason codes →

FOUR WAYS MODELS GET FIGURES WRONG

Measured failure modes, in the models' own words

Verbatim from the pilot of record, systems anonymized. Each mode has its own forbidden-figure reason code in every eval pack, so each stays countable.

prior_year_value

Last year's figure, this year's confidence

“The 2026 elective deferral limit that applies to her 401(k) is $23,500. This is the universal employee elective deferral limit for all workers under age 50…” — System D

The 2026 limit is $24,500 (IRS Notice 2025-67). $23,500 is 2025's figure — asserted fluently, with a rationale, as current fact.

derived_not_published

Derived, not recalled

“Enhanced catch-up for ages 60–63: $12,000 (150% × $8,000 standard catch-up).” — System A

Correct statute, correct method, wrong figure: the published 2026 enhanced catch-up is $11,250. The model applied the indexing rule to a stale base instead of recalling the notice — the signature frontier-model failure.

superseded

Repealed law, still cited

“Under current law, the basic exclusion amount in 2026 reverts to $5 million indexed for inflation from 2011, estimated at $6,800,000.” — System C

That sunset was repealed; the 2026 basic exclusion is $15,000,000. On the question asked, the superseded figure flips a no-tax answer into a multi-million-dollar phantom liability.

fabricated_forward_figure

Unpublished thresholds, stated as fact

“The 2026 highly-compensated-employee compensation threshold is $165,000. … source: IRS Notice 2025-xx.” — System C

A specific number, a confident citation to a notice that does not exist. Plausible, uncheckable by a reader who doesn't know the figure was never published — and enumerated as a forbidden assertion in the eval keys.

THE PACKS · ONE-TIME · VINTAGED

Four packs, one discipline

Tasks in JSONL with answer keys, the verified prior-year forbidden-figure tables, a mechanical scorer, and the methodology white paper. Figures roll every January; each tax-year vintage is a new pack.

EV01156 tasks

Rule-Grounding Eval Pack (TY2026)

A defensible answer to “how often is our AI advisor quoting last year's limits?” — a scored, documented evaluation your compliance file can cite, built on primary-source-verified 2026 figures.

$1,450one-timeView pack
EV0295 tasks

Threshold & Cliff Eval Pack (TY2026)

Tests whether an AI system knows that crossing a Medicare income threshold by one dollar costs the full surcharge — and whether it's using this year's threshold or last year's.

$1,450one-timeView pack
EV03150 tasks

Decision-Recall Eval Pack

Measures whether an AI assistant actually remembers what was decided for a client and why — scored against decision records whose ground truth is known by construction, not labeled after the fact.

$995one-timeView pack
EV00401 tasks

AI Eval Sets Full Corpus

The complete evaluation library for AI financial-advice systems — figure accuracy, threshold behavior, and decision memory in one purchase, with documented held-out exclusions so results stay auditable.

$3,950one-timeView pack
WHY TRUST IT

Honesty is the moat

Pre-registered

The decision rule, pass definition, and task manifest were frozen and published before any model was called.

k=3, mechanical

Three attempts per system; scoring is required-figures-present, forbidden-figures-absent. No LLM judge anywhere.

Keys by construction

Every figure traces to a verified primary document with the evidence URL recorded — 150/150 prior-year fact-years verified.

Warts published

The substitution log — including a harness bug that forced re-running an entire arm — is on the protocol page, verbatim.

Read the full protocol and results →

The eval measures the gap. The feed closes it — measured.

The pilot measured bare models: 29% of figure-bearing attempts cited a stale or fabricated figure. The pre-registered contrast run repeated the figure-bearing tasks with the live Rule Sets feed in context — zero stale figures in 225 attempts across three systems. The remaining failures were reasoning errors the feed can't fix, which is exactly what the eval packs exist to catch. Full tables on the scoreboard.

Financial Planning Rule Sets — cited, current, callable
QUESTIONS BUYERS ACTUALLY ASK

Frequently asked questions

What exactly am I buying?+

Evaluation task files (JSONL) with answer keys: each task states a planning question, and its key lists the required figures cited to the establishing government document, the enumerated wrong-but-plausible stale values with reason codes, and the expected answer. Packs also include the verified prior-year figure tables, a mechanical scoring script, a methodology white paper, and a license. One-time purchase, delivered as a ZIP.

How is this different from the free benchmark?+

The free open split (50 tasks) is the floor/standard tier of two task families and stays free forever. The paid packs carry the discriminating material: the adversarial categorical-flip tasks where a stale table changes the answer from a dollar amount to zero, all decision-recall tasks, roughly eight times the task volume, and the full forbidden-figure tables and scorer for compliance-grade reporting.

Why do the packs have tax-year vintages?+

Because the figures roll every January. An answer key is correct for its tax year, so each vintage is a point-in-time evaluation — TY2026 packs measure whether a system gives 2026 answers. When the next year's figures land, a new vintage ships; re-evaluating on it is the maintenance model, priced as a new one-time purchase rather than a subscription.

Is there an LLM judge in the scoring?+

No. Scoring is mechanical: required figures present, forbidden figures absent, typed answers matched, text assertions checked by substring. That is a deliberate design constraint — a judge model with its own stale figures cannot be the arbiter of figure staleness. Everything the scorer flags is recorded with surrounding text so a human can review edge cases.

Can I train on these tasks?+

No. Every task record carries training_use_permitted: false and the pack license prohibits training use — training on an eval destroys its value to you and to everyone it is compared against. Internal evaluation, regression testing, gating, and vendor comparison are all permitted uses.

How do I know the answer keys are right?+

Keys are built, not labeled: every figure a task depends on is drawn from tables verified against the primary government document (IRS notices and revenue procedures, SSA and CMS releases), with the evidence URL recorded — 150 of 150 prior-year fact-years verified before the pilot ran. A fact that fails verification cannot enter a task; the generator refuses.

What did the pilot actually measure?+

In a pre-registered 40-task run across four systems from three labs at k=3, 29% of figure-bearing attempts (86 of 300) asserted a stale or fabricated regulatory figure as current, and only one system passed every must-never-miss floor task on all three tries. The protocol, substitution log, and per-system numbers are published on the methodology page.

Are these realistic advisor scenarios?+

They are a competence floor, not a realism claim: tasks a competent financial-advice system must never miss, built to isolate figure grounding, threshold behavior, and decision recall. Scenarios are synthetic and constructed; the regulatory figures are real and cited. Both facts are stated plainly wherever the tasks appear.

Make staleness a number.

Start with the free 50-task open split. When you need the adversarial families, the decision-recall tasks, and keys your compliance file can cite, the packs are one purchase away.