wealthschema / benchmark
FIDUCIARYBENCH · BY WEALTHSCHEMA

FiduciaryBench

Can an AI be trusted inside regulated wealth work? Measured, not asserted.

An open benchmark of AI behavior where the rules are written down: Reg BI suitability, required-minimum-distribution mechanics, wash-sale mechanics, and planning-figure currency. Every answer key is computed by code or derived from a quoted primary-source passage that machine-matches the regulation text we hold — and every item's full provenance is a public page.

83 conduct items · 55 publiccorpus 2026.9 · 11 primary sources50 planning tasks · freepip install fiduciarybench
HOW A KEY EARNS THE RIGHT TO JUDGE

Answer keys with a chain of custody

A benchmark is only as honest as its answer keys, and human recall cannot out-know a frontier model. So correctness is engineered: a closed corpus of 11 primary sources (2026.9) fetched from the issuing authorities and hashed; keys that are code or quoted derivations; adversarial cross-vendor verification; and planted controls the pipeline must catch before anything counts. Every item's ledger — scenario, cited passages, derivation, verifier record, version history — is public.

The admission rule

An item ships only if its answer key can be verified without consulting the open web: a computation run by code, or a derivation from a quoted passage of a primary source we hold locally, versioned. If defending a key would need a search engine, the item is inadmissible — by construction, not judgment.

Quote-gated verification

Two verifier models from different vendors than the generator attack every key. A verdict is admissible only if it cites a passage that exact-matches, character for character after normalization, the corpus file we hold. A model can hallucinate a confident verdict; it cannot hallucinate a passage that survives the match.

Disagreement is an alarm

Agreement between models counts for nothing — shared training priors are not truth. Any disagreement, and any verdict that fails the quote gate, routes the item to a named human who resolves it by reading the cited paragraphs, never the open web.

Planted controls

Every verification batch is seeded with deliberately wrong keys — off-by-one table rows, unreduced life-expectancy factors, reversed verdicts. If the pipeline misses even one, the whole batch fails and nothing ships. The judge is tested continuously, not trusted.

Task familyItemsPublicHeld outComputed keysDerived keys
regbi-suitability18126018
rmd-mechanics3221111913
rollover-advice1073010
wash-sale-mechanics231581310
planning-calculations (v2 open split)505070+500

Where this stands, stated plainly: the conduct suites are v1.0 — the verification pipeline ran on 2026-09-01 (two verifier vendors, every planted control caught, 36/36 items verified after one repair round documented in the methodology), the batch resolution was owner-approved, and the methodology is signed by Satya Iluri. Disputes and corrections are handled through the standing challenge policy.

THE LEADERBOARD · CONDUCT SUITES · 2026-09-01 SWEEP

Conduct-suite scores, from stored runs only

Every row below is rendered from a committed harness run report (fiduciarybench/runs/ in the repo — model, timestamp, per-item results). A score with no stored report behind it cannot appear here by construction. Single attempt per item over the public split as it stood at run time — each row shows its own item count, and rows are superseded when a model is re-run on a newer split. Deterministic scoring, no LLM judge. Runs request temperature 0; model families whose APIs reject the temperature parameter (the Claude 5 family, GPT-5.6) run at provider default, disclosed per row.

ModelScoreReg BIRolloverRMDWash saleParse failuresSampling
anthropic/claude-opus-545/46 (0.98)9/97/716/1713/130provider default*
openai/gpt-5.6-sol45/46 (0.98)9/97/716/1713/130provider default*
deepseek/deepseek-v4-pro44/46 (0.96)9/97/715/1713/130temp 0
gemini/gemini-3.1-pro-preview44/46 (0.96)9/97/715/1713/130temp 0
openai/gpt-5.6-luna44/46 (0.96)9/97/716/1712/130provider default*
xai/grok-4.644/46 (0.96)9/97/716/1712/130temp 0
gemini/gemini-2.5-pro41/46 (0.89)9/96/713/1713/130temp 0
anthropic/claude-sonnet-540/46 (0.87)9/97/712/1712/130provider default*
moonshot/kimi-k2.638/46 (0.83)9/96/711/1712/130provider default*
anthropic/claude-haiku-4-5-2025100131/46 (0.67)9/97/77/178/130temp 0
mistral/mistral-large-latest31/46 (0.67)9/96/78/178/130temp 0
mistral/mistral-medium-latest30/46 (0.65)9/96/77/178/130temp 0
together/meta-llama/Llama-3.3-70B-Instruct-Turbo27/46 (0.59)9/96/75/177/130temp 0

Disclosure: every item was adversarially verified by two model families from different vendors than the generator, and flagged items were repaired until the verifiers agreed — a mild selection pressure in the verifying families' favor on the items they verified. Verifier families rotate between batches: Gemini and DeepSeek verified the earlier batches, GPT and Grok the later ones, so each marked row's advantage is limited to the slice of items its family verified. Each batch report names its verifiers. *Provider-default sampling: the API rejects the temperature parameter on these models; each run report records this.

Read the ceiling honestly: four frontier systems score 100% on the 25-item public split — these are floor/standard conduct items, and saturation at the frontier is expected. The discrimination lives in the spread below the ceiling (a 0.64 on the same items is a deployment-relevant fact) and in the 28 held-out items and adversarial families that back private evaluations. N=25 and single-attempt: treat small gaps between top rows as noise, not ranking.

PILOT OF RECORD · 2026-07-30 · PLANNING FAMILY

The figure-currency scoreboard

40 tasks × 4 systems × 3 attempts, protocol pre-registered before any model was called. pass@1 with 95% CI; pass^3 requires all three attempts correct; rg+tc restricts to the rule-grounding and threshold/cliff families; the Stale Figure Rate (SFR) is the share of 75 figure-bearing attempts asserting a stale or fabricated figure as current. Every number cites the methodology page.

Systempass@1 [95% CI]pass^3rg+tc pass@1Stale Figure RateFloor pass^3
claude-sonnet0.90 [0.83–0.94]0.850.840.07 (5/75)1.00
gemini-3.1-pro-preview0.57 [0.48–0.65]0.420.330.17 (13/75)0.50
deepseek-v4-pro0.48 [0.40–0.57]0.350.270.28 (21/75)0.40
claude-haiku0.36 [0.28–0.45]0.330.110.63 (47/75)0.30

Named-system caveats, published rather than smoothed: the Claude systems ran as session aliases (exact dated model strings not exposed by the harness); the Gemini system is a preview endpoint; delivery was batched (declared deviation); N=40. The full pre-registration and the post-freeze substitution log — including a harness truncation bug that forced re-execution of the entire Claude arm — are on the protocol page.

Only one of four systems passed every floor task — the "a competent system must never miss these" tier — on all three tries. In-context memory was the opposite story: saturated at 1.00 for all four systems. The gap is figure currency, not recall.

THE FAILURE TAXONOMY

What the wrong answers actually were

The errors are fluent, cited, and repeat identically across attempts — pass^3 ≈ pass@1 on every system. Each named mode has its own forbidden-figure reason code in the eval corpus, so it stays countable.

prior_year_value

Last year's figure, this year's confidence

A superseded published figure asserted as current — the 2025 (or 2024) IRA limit cited fluently, with a source, as 2026 fact. The dominant mode for the weakest system and material for two others.

derived_not_published

Derived, not recalled

Correct statute, correct method, wrong figure: the model computes a plausible value from an indexing rule or stale base instead of recalling the published number. The signature frontier-model failure — observed in three of four systems.

superseded

Repealed law, still cited

Two systems asserted the repealed TCJA estate-tax sunset where current law sets a $15M exclusion — a categorical wrong answer on a multi-million-dollar question, not a rounding error.

fabricated_forward_figure

Unpublished thresholds, stated as fact

Next cycle's not-yet-published figure given as a number. Plausible, specific, and uncheckable by a reader who doesn't know the notice hasn't been issued.

STRAIGHT ANSWERS

The questions this benchmark exists to answer

What is FiduciaryBench?

FiduciaryBench, by WealthSchema, is an open benchmark of AI behavior in regulated wealth-management work: Reg BI suitability, required-minimum-distribution mechanics, wash-sale mechanics, and planning-figure currency. Every answer key is computed by code or derived from a quoted primary-source passage (SEC, FINRA, IRS/Treasury, U.S. Code) that exact-matches a locally held, versioned corpus — open-web knowledge is inadmissible by construction.

How often do LLMs cite outdated tax figures?

In our pre-registered pilot of record — 40 routine 2026 U.S. planning questions, 4 systems from 3 labs, no tools, 3 attempts each — 29% of figure-bearing attempts (86 of 300) asserted a stale or fabricated regulatory figure as current. The per-system Stale Figure Rate ranged from 7% to 63%.

What is a Stale Figure Rate?

The share of figure-bearing answers in which an AI system asserts a stale, superseded, or fabricated regulatory figure as current — scored mechanically against enumerated forbidden values with reason codes, not judged by another model. It makes 'is our AI giving this year's advice' a number instead of an anecdote.

Can a vendor pay to improve its FiduciaryBench score?

No. Public scores are free to earn and impossible to purchase. Vendors can pay for private evaluation runs and continuous monitoring against held-out suites, but a private result can never be selectively published — a vendor may disclose a full report or nothing — and the public verdict is never for sale. Every engagement type and its price is disclosed on the methodology page.

Do the errors go away if you ask again?

No. On every system tested, pass^3 approximately equals pass@1 — wrong answers repeat near-identically across attempts. Stale figures are a systematic property of a frozen model, not sampling noise, which is why re-rolling the same question is not a mitigation.

Run the public splits

Conduct suites: 55 verified public items with full provenance, runnable via pip install fiduciarybench (runners for the major APIs, deterministic scoring, no LLM judge). Planning family: 50 tasks, free, no auth; fetch with ?withhold_answers=true for blind evaluation.

What's deliberately held out

28 conduct items, all adversarial "categorical flip" tasks, all decision-recall tasks, and a 70-task planning reserve — including every one of the 40 pilot tasks — are never published and never sold. Held-out suites back private evaluation runs and anti-gaming rotation; that is what keeps public scores meaningful.

Held-out conduct + planning suites: private runs and monitoring
Adversarial flips + decision recall: commercial eval packs
Public splits: free, forever
Private results: a vendor discloses a full report or nothing
THE CONTRAST RUN · 2026-07-31

Same tasks, same systems, live figures: zero stale answers

The 25 figure-bearing pilot tasks re-run with the live Rule Sets feed (us-federal-2026, 103 cited figures) rendered into context as a lookup — byte-identical for every system, protocol pre-registered before execution. The same three systems that produced 65 stale or fabricated figures in 225 bare attempts produced zero in 225 with-lookup attempts. Full protocol, substitution log, and the honest caveats: methodology page. Disclosure: Rule Sets is a WealthSchema product; this run measures its effect and is labeled product evidence, not a ranking of third-party systems.

SystemBare pass@1 (rg+tc)Bare SFRWith feed pass@1With feed pass^3With feed SFR
claude-sonnet0.840.07 (5/75)1.001.000.00 (0/75)
gemini-3.1-pro-preview0.330.17 (13/75)1.001.000.00 (0/75)
claude-haiku0.110.63 (47/75)0.960.920.00 (0/75)
deepseek-v4-pro0.270.28 (21/75)— not run

What the feed did NOT fix, published on principle: the weakest system's four remaining failures were rules-comprehension errors made with the correct figures in front of it (stacking catch-up limits; treating the HSA age-55 catch-up as poolable between spouses). A figures feed makes a system current, not competent — which is why FiduciaryBench tests conduct and mechanics, not just figure recall. deepseek-v4-pro was not re-run (no API credential in the run environment); its bare-arm numbers stand unpaired.

Use it. Build on it. Get measured by it.

The benchmark is free and open. The scenario data that fuels serious testing is for sale. And a private, held-out evaluation of your own system — repeatable on every release — is a conversation away. The verdict itself is never for sale.