How synthetic financial data gets built, where it has to hold up under compliance review, and the math that keeps a household's numbers internally consistent.
Hard-coded planning constants do not fail loudly. They rot. A cited feed keeps them current and gives every number a source and an effective date, which is the difference between a value you can defend and one you hope is right.
You cannot grade what you cannot check. Evaluating an agent's financial memory needs decision data with answers that are correct by construction, scored per dimension, on a held-out set. Eyeballing is not a benchmark.
Production data is realistic and expensive to use safely. Generic fake data is safe and not realistic enough for a rules engine. Coherent synthetic households are both, and a free sample lets you confirm that before you commit.
One memory pipeline configuration correctly identified which decision episode mattered 78% of the time — and retrieved zero regulatory citations for any of them. Attribution exact-match 0.000 is the exact failure a regulated deployment can't accept.
Claude-haiku extracts roughly 2.8 edge facts per decision episode under a temporal-knowledge-graph pipeline; Gemini-flash extracts roughly 18 from the same episodes. Same task, same model class, a 6x spread — because extraction volume is a pipeline property, not a model property.
A memory pipeline's rationale-recall score collapsed from 0.402 on one corpus to 0.038 on a second, disjoint one — while its overall score rose. The verbatim baseline showed no such swing. That asymmetry is the case for evaluating on more than one corpus.
Every DecisionSynth Bench vendor configuration runs twice — once with Gemini, once with Claude, as the model doing extraction and answer generation. The two runs disagree enough that publishing only one would have hidden which variable actually drove each score.
DecisionSynth Bench's held-out precedent-search column has 9 tasks — the rarest task type in the corpus. Publishing that column instead of hiding it, and labeling it directional rather than significant, is the harder and more honest choice.
Zep/Graphiti posts the worst overall exact-match of any system in DecisionSynth Bench's held-out scoreboard — and the best temporal-ordering score of any LLM-backed system, ahead of verbatim storage. Architecture fit beats architecture ranking.
DecisionSynth Bench's deterministic baseline — store everything verbatim, retrieve by keyword overlap — scores 0.980 exact-match on the held-out set. That's not the benchmark being too easy. It's the ceiling reference that makes every other score interpretable.
Most AML engines fail quietly — too many alerts exhausting analysts, or too few missing exactly what FinCEN advisories warned about. The miss isn't the rule logic; it's the test corpus the rules were tuned against.
The amended FTC Safeguards Rule replaced a one-paragraph written-program standard with nine specific deliverables — a Qualified Individual, MFA, encryption, a 30-day breach clock — each one the FTC can ask to see.
Insurance illustrations look simple in the regulatory imagination. In practice, they're the most edge-case-dense calculation in personal finance — MEC crossings and substandard underwriting break most validators.
A taxpayer can be domiciled in Florida and statutorily resident in New York at the same time — and owe state tax to both. Most tax engines conflate domicile and residency and get it wrong.
In wealth-tech QA, a missed bug is denominated in regulatory penalty and customer trust — and test data, not test logic, is usually the bottleneck standing between the team and a test that actually matters.
Most QSBS engines nail the five-year holding period and quietly fail the gross-asset test at issuance and the redemption look-back — exactly where a six-figure tax surprise comes from.
Most fintech synthetic-data procurements go badly because the buyer contacts vendors before defining the use case in a single paragraph — five vendors, three weeks of demos, and no signal to show for it.
Financial ML sits where training data is hardest to get, regulators scrutinize what you trained on most closely, and a model that learned the wrong thing costs eight or nine figures to unwind.
A loss in your taxable account can be permanently disallowed by a purchase inside your IRA — no basis adjustment, unlike an ordinary wash sale. Most engines don't track wash sales across that boundary.
The 2023 GLBA Safeguards Rule amendments reset what regulators consider reasonable controls, and state AGs opened a second enforcement front. The cost of a breach has quietly outgrown the cost of the controls that prevent one.
A supervisory engine tested against a corpus with no cognitive-decline households hasn't been tested for cognitive decline. Twelve fact patterns examiners actually cite, and what a test corpus needs to exercise each one.
ACATS transfers take 5–7 business days of inconsistent state — partial transfers and in-kind lots that must preserve basis and acquisition date. Mock data assumes the transfer completes in zero time; production doesn't.
An annuity isn't an account balance — it's a contract with multiple value streams, rider-defined guarantees that can diverge sharply from the account value, and surrender mechanics mock data often omits.
Crypto tax engines look like equity tax engines until a hard fork or an airdrop shows up — each is a distinct basis event with its own tax treatment, and DeFi adds many more with no securities-world analog.
A single equity grant is a multi-year tracking commitment — it vests over years, taxes differently by award type at exercise, and errors don't surface until the employee's own tax filing disagrees with the platform.
HNW households aren't mass affluent with more zeros. Trust structures, illiquid positions, and dynastic-planning needs make the corpus a different product, not a bigger version of the mass-affluent one.
A robo-advisor looks like one product — deposit money, get an allocation — but it's six modeling problems trench-coated together, and most test corpora only exercise the prettiest one.
An S-corp owner is a W-2 employee, a K-1 partner, and an entity owner all at once — three different tax positions in one person. Most platforms model one and ship surprises in the other two.
A grant that vests in one country, is exercised in a second, and is sold from a third fragments into separate sourcing rules and forms per jurisdiction — most domestic equity-comp engines assume there's only one.
Going custodian-direct isn't one integration — it's four to six, each with its own account-number formats, lot-relief defaults, and statement cycles. Mock data treats them as interchangeable; production doesn't.
A DB pension isn't an account balance — it's an accrued benefit whose lump-sum value swings with the discount rate and whose joint-and-survivor election permanently trades a higher check now for spousal protection later.
Twelve tells that a synthetic dataset is too clean — no overdrafts, no failed trades, suspiciously round cost basis, no survivorship attrition. Run the matching query before you commit to backtesting against it.
Random-walk returns are fast and blind to the fat tails that drive sequence-of-returns risk; replay is faithful to the past and silent on regimes that haven't happened yet. Regime-switching is the production answer.
An HSA is the only US tax-advantaged account that's deductible going in, tax-free growing, and tax-free coming out. Most retirement engines treat it like a checking account, missing the stealth-retirement-account math.
The same Roth IRA comes back as 'roth' from Plaid and 'ROTH_IRA' from Yodlee — each aggregator's normalization trades depth for breadth differently, and most test data never surfaces the mismatch.
A 2-for-1 split or a partial-share spinoff propagates basis, holding-period, and tax-lot changes through every system that touches the position. Mock-data tools generate static positions; none generate that propagation.
A position has a price. A non-USD position has two prices, an FX rate, and a translation rule linking them — and a platform that treats them as a single USD number ships with a structural bug.
An NQDC plan isn't a 401(k) without a contribution limit — it's an employer promise-to-pay with no creditor protection, a distribution schedule locked years in advance, and a tax regime built to make changes punitive.
A reporting platform that ships with single-period Brinson and untested multi-period linking has a known correctness gap. Synthetic test data is how you find it before an auditor does.
A US holder of a foreign mutual fund is almost certainly holding a PFIC — and under the default IRS regime, deferred gains are taxed at each year's highest bracket plus an interest charge for the deferral itself.
An aggregator feed and a custodian feed for the same account can disagree on rounding, settlement timing, and distribution character — and both still be right. The platform's job is knowing which to trust, per field.
ITIN holders, gig-economy applicants, borrowers in active forbearance, and multi-state filers show up constantly in real applications and almost never in the test corpus meant to catch how the engine handles them.
AG 49-A tightened what an IUL illustration can show, and term, whole life, VUL, RILA, and annuity products each carry their own constraints on top of it. What a test corpus needs to catch a non-compliant illustration first.
Mortgage engines ship clean on the median W-2 borrower and break on gig income, ITIN filers, and recently-bankrupt applicants — plus manufactured, mixed-use, and deed-restricted properties the standard test corpus skips.
Schema-preserving synthesis tools like Tonic and Gretel are excellent at preserving fidelity and useless at engineering coverage — and fintech compliance testing depends almost entirely on coverage.
A naive rebalancer trades to target weights; a tax-aware one has to know which lots to sell, which account to draw from, and whether the trade triggers a wash-sale or vaporizes a QSBS position's exclusion.
Pure-synthetic fraud models underperform because adversarial signal lives in the tail, and generators trained on legitimate-behavior distributions can't reproduce it. The hybrid architecture that ships in production instead.
The same Canadian dividend withholding can net out to zero after a correctly filed Form 1116 — or to double taxation if the credit isn't claimed right. The platform decides which.
Every wealth-tech platform has an integration layer, and most ship with mock data that quietly assumes it always works. Production breaks on the round trip, in the reconciliation between aggregator and custodian.
Domestic-only test data is fine for domestic-only platforms. Anything that touches cross-border — and most institutional wealth-tech eventually does — needs a different test corpus.
Monte Carlo, Roth conversion, and RMD logic are the easy parts of decumulation. The hard parts are the products mock data forgets exist — starting with annuity riders and HSA reimbursement timing.
A wrong rule applied to a household's monthly history compounds every month, until the position has hallucinated a corporate parent that no longer exists or a cost basis that disagrees with itself across reporting paths.
A test corpus can pass QA and still ship production incidents — the gaps cluster around multi-state moves, IRMAA-crossing Roth conversions, and cross-account wash sales the corpus never included.
A partial fill across multiple lots, or an ACATS transfer that lands inside a wash-sale window — these are transaction shapes most engines only learn about from a Sev-1 ticket. Twelve of them, cataloged in advance.
Assembling edge-case test data by hand takes days to weeks; querying a well-curated synthetic corpus for the same pattern takes minutes. The QA-cycle-time win is real — not just the usual privacy pitch.
Most fintech teams build a credible-looking synthetic corpus once, then stop iterating — treating it as static infrastructure instead of a living asset. Eight named mistakes explain the bug class each one ships.
Production data is the path of least resistance for ML training in finance. After CFPB scrutiny of automated credit decisions and the EU AI Act's bias-mitigation rules, it's become the path of the most regulatory risk.
Exercising ISOs creates AMT liability on phantom income — the spread between strike price and fair value — even though no cash changes hands, sometimes forcing a sale of vested shares to cover the bill.
A CRT or CLAT is a long-run actuarial bet on the donor's life expectancy, the §7520 rate, and the funding asset's return — platforms that flatten these variables out miss the planning surface entirely.
Divorce doesn't just split assets — it splits the household entity itself, and most wealth-planning platforms that treat the household as immutable produce wrong projections from the moment the QDRO is signed.
TCJA's higher standard deduction made charitable deductions disappear for most taxpayers. A donor-advised fund restores it by bunching several years of giving into one — without changing when charities get paid.
The textbook 'taxable, then tax-deferred, then Roth' withdrawal order is right only when no other constraint binds — RMDs, IRMAA premium cliffs, the NIIT surtax, and ACA subsidy phase-outs all override it.
Production fintech usually breaks on the edge case the test corpus didn't contain, not the one it was built around — across tax, retirement, insurance, equity comp, crypto, and lending compliance.
The federal estate-tax exemption is set to halve at the end of 2025, and the IRS's 2024 no-claw-back rule means pre-sunset gifts lock in the old exemption in ways most planning engines still don't compute.
A fair-lending audit doesn't care that lending records are anonymized — it cares whether protected-class proxies and credit decisions still carry the joint-distribution fingerprint of historical bias.
Fidelity, privacy, and utility aren't a pick-two trilemma. The real trade-offs live in three upstream calibration knobs, and how a team sets them determines what the resulting dataset is actually useful for.
A dynasty trust's GST exemption allocation is locked in decades before the generation-skipping tax comes due — the inclusion ratio set at funding decides whether decades of appreciation pass tax-free or not.
GLBA, GDPR, and CCPA all define regulated data by its link to a real person — synthetic households have no real person to link to, and a documentation package turns a months-long compliance review into a single meeting.
Rule-based, GAN, LLM, and hybrid generation each have a different competence band and ship a different bug class the moment a vendor uses them outside it — know which family you're buying before the bugs show up.
Indexed universal life illustrations run under NAIC's AG 49-A, and the regulation has more teeth than most engines acknowledge — cap, floor, multiplier, and bonus all have to be computed from the actual contract.
A US citizen in London with a UK pension and a Swiss brokerage account has a tax-filing footprint — FBAR, FATCA, PFIC — that most wealth-tech platforms don't even recognize exists, let alone model.
Position-level data is insufficient for any tax-aware engine. Real holders have dozens to hundreds of lots per position, and a tax-loss-harvest, wash-sale, or QSBS decision has to operate on the lot, not the position.
Monte Carlo assumptions that are harmless in a 30-day options-pricing model turn load-bearing wrong over a 30-year retirement — and the standard libraries default to exactly those assumptions.
NUA lets a retiring 401(k) participant pay ordinary tax on company stock's cost basis and long-term capital-gains tax on its appreciation — worth seven figures, decided once, in a single tax year, with no do-over.
Section 199A's QBI deduction is the most heavily-tested provision in the TCJA, and most K-1-handling engines still get it wrong. The 2025 sunset is set to compound those bugs, not fix them.
PCI DSS scope is expensive to carry, and moving development, QA, and analytics environments onto synthetic payment data is the cleanest scope-reduction pattern most fintechs overlook.
A single is_qsbs boolean is why tax-loss-harvesting engines sometimes sell a QSBS position for a token loss months before its five-year clock unlocks the full exclusion. What the data model needs to track instead.
SECURE 2.0 rewrote RMD rules in ways that broke most production engines — the rolling 73-then-75 start age and the 10-year rule for inherited IRAs both require the engine to be rewritten, not patched.
Bracket-fill Roth conversion calculators work when there's only one bracket to fill. Real households cross IRMAA tiers, the NIIT threshold, ACA subsidy cliffs, and capital-gains stacking all at once.
Section 1031 lets real-estate investors defer capital gains by exchanging into like-kind property — and a chain of exchanges is graph-shaped, not row-shaped, which is exactly what a flat property schema can't hold.
SECURE 2.0's RMD age increase (73 now, 75 by 2033) expands the Roth-conversion window — but the IRMAA lookback and Social Security taxation thresholds haven't moved, making the ladder harder to optimize, not easier.
The highest-leverage decisions in a business exit happen years before the closing table — pre-sale trust funding, state-residency planning, Section 1045 timing — and most wealth-tech shows up only at signing, too late to help.
The 'delay to 70' recommendation is right for a single retiree at average life expectancy — not for the two-spouse households claiming engines actually serve, where survivor benefits change everything.
SR 11-7 never mentions synthetic data — examiners have made clear that's not permission, it's a burden on the bank to prove out. The provenance and validation-independence gaps that draw a finding.
The Department of Education's PSLF qualifying-payment count and the borrower's own count routinely disagree, and a generic "enter your balance" calculator can't catch it or flag that refinancing quietly forfeits forgiveness.
Vendor data sheets all say 'statistically faithful' and 'production-ready' — none of it survives a query for whether assets minus liabilities equals the reported net worth. Five dimensions that actually disqualify a dataset.
Anonymized data can be re-identified, mocked data has no relationships between fields, and aggregated data can't back a single test case — synthetic is the only one of the four built to survive a compliance review.
A mutual-fund capital-gain distribution looks like income to a beneficiary and is legally principal under UPIA — one of dozens of allocation calls wealth-tech platforms flatten into "all receipts are income."
A household can be solvent on the annual average and still need a margin loan in February — annual snapshots can't see the seasonality driving it. Which months break, and what a cash-aware engine has to track monthly.
If your engine depends on within-year cash-flow timing, annual snapshots silently understate your error rate. We generate monthly data instead — chunked, after a single long LLM call proved to quietly fabricate history.
Anonymized data leaks. Synthetic data, done right, doesn't. The case for fully synthetic households as the production-ready path for fintech and wealth-tech builders.
Exercise ISOs and hold the shares past year-end, and the spread between strike and FMV becomes an AMT preference — taxed even though no cash changed hands, sometimes on a paper gain that later evaporates.
Two retirees with the same portfolio and the same average return can still land in radically different places — the one who runs out of money usually didn't run out of returns, just ran out of them at the wrong time.
A position-level snapshot tells you the holder owns VTI at some cost basis. It doesn't tell you when each lot was acquired, or whether a wash sale in a different account just disallowed part of the loss.