AI chatbots give wrong financial answers most of the time, study finds

AI chatbots give wrong financial answers most of the time, study finds
Major AI models failed to correctly answer financial questions in 57% of cases, new research shows, raising serious questions about consumer reliance on AI tools.
SEP 21, 2026

Popular AI chatbots give incorrect or incomplete answers to financial questions more often than they get them right, according to a new comprehensive benchmarking study by Saturn, an AI and technology firm serving more than 750 financial advice firms and 6,500 advisors.

The Saturn "Artificial Authority" report - the most in-depth benchmarking exercise of its kind to date - tested 18 AI models across 121 different money-related questions, with each question run five times to account for variability. In total, more than 10,000 responses were assessed. The models tested included free and paid-for versions of ChatGPT, Gemini, Claude and CoPilot, across a range of difficulty levels.

Across all 17 current models (one retired model was excluded from the headline figures) the average accuracy rate was just 43%, meaning AI chatbots made mistakes 57% of the time, according to Saturn, which is based in the UK.

The report follows a warning from Macabacus, a Microsoft 365 productivity platform for finance and professional services teams based in New York, which highlighted that most finance firms have sent clients an AI-generated error.

Hard questions exposed the sharpest failures

The gap between AI performance and consumer expectation widened significantly on more complex queries. On hard questions, defined as multi-part scenarios involving interacting tax rules and precise figures, accuracy fell to just 12%, with mistakes made 88% of the time, per the Saturn report. Even on the easiest questions, which covered general financial literacy topics requiring no calculation, accuracy averaged only 54%.

Free models performed notably worse than their paid-for counterparts. Saturn found that free models failed 63% of the time, compared with 49% for paid-for models. The worst-performing free model, Claude Haiku 4.5, produced wrong or substantially incomplete answers 82% of the time. The best-performing free model, ChatGPT-5.6 Luna (max), still returned incorrect or incomplete responses 56% of the time, according to the Saturn study.

The best overall performer was Claude Opus 5 (reasoning), a paid-for model from Anthropic, which achieved a pass rate of 61% although that means it still failed to answer correctly in almost four out of every 10 cases. On hard questions, even this top performer made mistakes 67% of the time, the Saturn report found.

The divergence between free and paid-for performance carries equity implications the report's authors found troubling. Consumers least able to afford professional financial advice, the same group most likely to rely on free AI tools, are being exposed to the least accurate guidance.

How the models got it wrong

Saturn's methodology required each model response to meet a strict all-pass standard, meaning every required figure, rule, warning and deadline had to be present. The most common failure type (accounting for 37.1% of errors) was giving an incomplete answer in ways not captured by the other error categories. Leaving out a required figure, limit or deadline accounted for a further 18.4% of failures.

More alarming were errors involving fabricated rules, which made up 6.1% of failures, and out-of-date regulatory guidance, which accounted for 2.3%.

The report highlighted that wrong answers were delivered in the same authoritative, fluent tone as correct ones, giving consumers no clear signal that the response might be unreliable.

The advice gap problem

The backdrop to Saturn's research is a growing reliance on AI as a substitute for professional financial advice, as highlighted in an EY report in April 2026 which found that 49% of respondents used AI in the previous six months to assist with saving or investing, 21% for product recommendations, and 18% using it for budgeting, household finances and trading support.

Saturn's findings suggest that gap-closing potential remains some distance from reality. As the report concluded, the most freely available models are also the most error-prone. That dynamic risks compounding financial inequality rather than narrowing it.

For financial advisors and wealth management professionals, the Saturn data offers an important frame of reference as clients increasingly arrive having already consulted AI tools. Saturn's research, covering questions spanning debt management, student loans, mortgages, pensions, inheritance tax and retirement planning, illustrates how even routine-seeming queries can produce materially wrong answers with serious financial consequences.

Latest News

Global M&A value rises 15%, with highest sentiment in financial sector
Global M&A value rises 15%, with highest sentiment in financial sector

Megadeals are surging in 2026, but small and mid-market transaction volume remains well below historical norms.

Paramount in advanced settlement talks over $81 billion Warner Bros. deal
Paramount in advanced settlement talks over $81 billion Warner Bros. deal

From a $1.5 billion California production pledge to possible cable channel sales and a CNN oversight board, the terms of a potential deal are taking shape.

US debt crisis nears tipping point, watchdog warns Congress
US debt crisis nears tipping point, watchdog warns Congress

A new report from the National Seniors Policy Center details how a US sovereign default could unfold - and who would be hurt first.

EQT raises Perpetual takeover bid for the fourth time
EQT raises Perpetual takeover bid for the fourth time

Swedish private equity giant EQT has lobbed an improved offer for Australian asset manager Perpetual, intensifying a months-long pursuit.

In the Age of AI, Trust Becomes the Advisor's Greatest Asset
In the Age of AI, Trust Becomes the Advisor's Greatest Asset

As AI makes financial information more accessible than ever, Lana Hock explains why human judgment, trust, and empathy remain the qualities clients value most in a financial advisor

SPONSORED In the Age of AI, Trust Becomes the Advisor's Greatest Asset

As AI makes financial information more accessible than ever, Lana Hock explains why human judgment, trust, and empathy remain the qualities clients value most in a financial advisor

SPONSORED Direct indexing webinar targets tax-loss harvesting amid market swings

Northern Trust’s Ken Lassner shows advisors how to convert volatility into after-tax portfolio gains