ChatGPT Is Already Advising You on What to Do With Your Money. Can You Trust It?
Can ChatGPT and other models be trusted with decisions about taxes, loans and savings when in a fresh test they got more than half of their answers wrong — and how do you use AI for finance without taking its answer on faith?
The Saturn study ran 18 models, among them ChatGPT, Claude, Gemini, Grok and Copilot, through 121 financial questions five times each — more than 10,000 answers on debts, loans, mortgages, pension savings, taxes and savings; an answer failed if it contained a factual error or missed an essential point or a necessary warning. On average the models got 57% of the answers wrong, free models 63%, paid ones 49%, 88% on the complex questions and 93% for free models on the hardest ones. By model: Claude Haiku 4.5 — 82%, Gemini 3.1 Pro — 73%, Grok 4.5 — 60%, ChatGPT 5.6 Luna — 58%, and the best result, 39%, went to Claude Opus 5 in reasoning mode. There were no $200 Pro versions in the test, but even those, by Alexander's estimate, would be wrong some 20–30% of the time. The errors cost money: Claude Haiku 4.5's advice on pension taxes could, by the authors' estimate, have led to an additional tax charge of £17,500, a model advised paying off the most expensive debt first without taking rent and taxes into account, Gemini assured a borrower that a mortgage payment holiday would not affect their credit score, and the models ignored upcoming rule changes. GPT Finance already connects accounts at more than 12,000 financial institutions, and Claude now has a system for financial advisers, but connected data will not tell a model about your exceptions, an inheritance or a sudden change in income. Alexander's conclusion: these systems are smarter than any person at working with data, but complex financial answers must not be accepted on autopilot — they need to be double-checked.
What to watch for
Key takeaways
Right before the recording ChatGPT recommended that Alexander connect his credit rating via Experian and have it analysed. GPT Finance is a separate financial part of the system that connects accounts at more than 12,000 financial institutions — banks, cards, brokers via Plaid.
Saturn asked ChatGPT, Claude, Gemini, Grok, Copilot and other models questions about debts, loans, mortgages, pension savings, taxes and savings; these are repeated runs of a small set of questions, not 10,000 different tasks.
An answer failed if it contained a factual error or missed an essential point or a necessary warning; the test included Claude Opus 5, ChatGPT 5.6 in its Luna version, Grok 4.5, Gemini 3.1 Pro and Claude Haiku 4.5.
On average all the models got 57% of the answers wrong, free ones 63%. A tax return with a 10% chance of error is already very serious, and here it is more than half.
The paid versions were cheap ones: no 5.6 Sol Pro, no 6.0 Astra Pro, no Fable 5.1, no Extra High or Max modes. The paid models were wrong 49% of the time; a $200 Pro version, by Alexander's estimate, would be wrong some 20–30% of the time.
It was this figure that made Alexander record the episode; free models were wrong 93% of the time on the hardest questions, although the 121 questions were hardly of extreme difficulty.
By the study authors' estimate, following the recommendation on pension taxes could have led to an additional tax charge of £17,500; this is a calculation of a possible consequence, not a case that actually happened.
The model ignored rent, taxes and arrears: for a person with arrears the consequences of non-payment can be more serious than the interest saved, and an unpaid cheap loan can turn into problems in the family.
Gemini wrongly reassured a borrower, and in the US such an error can cost several years of a record in the credit system; on student loans, Claude invented a rule for people who move abroad.
That is what Sam Altman said a year ago, and ChatGPT itself warns that it can make mistakes. But people stop double-checking and act on autopilot — and in finance autopilot leads to serious consequences.
Gemini 3.1 Pro was wrong 73% of the time despite all of Google's integrations and knowledge base, Grok 4.5 60%, ChatGPT 5.6 Luna 58%, and the report does not say which mode Luna ran in; the best result went to Claude Opus 5 in reasoning mode.
The models ignored upcoming changes, tax changes for example. And GPT Finance, even once connected to your accounts, will not learn about an illness, a unique exception, an inheritance or a sudden change in income — the system cannot learn this with any guarantee.
Alexander does not use apps outside ChatGPT, Claude, Gemini, Grok, Codex and Claude Code: none of them will analyse his report through GPT-6 Astra. For one of his companies he got a report from a system that Codex wrote to bypass QuickBooks, and he sees a fairly high level of reliability in it.
What this episode is about
A solo ToTheMoon episode with Alexander Volchek on whether artificial intelligence can be trusted with financial advice and with everything to do with handling money. A year or two ago such a question would have sounded like hype: people were uploading reports, accounts, contracts and tax data into AI, and the channel had plenty of cases where the systems worked better than financiers, accountants and lawyers. But now, at the end of September 2026, there is a fresh study by the company Saturn, "Artificial Authority. Should You Trust Artificial Intelligence With Financial Advice?", on fresh models — Grok 4.5, ChatGPT 5.6, Gemini 3.1 Pro, Claude Opus 5 — and its error rate made Alexander record a separate episode.
He himself advises asking models financial questions every day — about accounts, credit ratings, investments, taxes, bookkeeping — and recently told how he analysed all his accounts and reporting through Codex, so that systems such as QuickBooks may become unnecessary. And right before the recording ChatGPT itself suggested he connect his credit rating via Experian: this is GPT Finance, a financial module, or plugin, that connects accounts at more than 12,000 financial institutions via Plaid, and not only in the US. You can ask about money through such plugins and connectors, or directly — in the chat or in Codex.
Saturn took 18 models — ChatGPT, Claude, Gemini, Grok, Copilot — and asked each of them 121 financial questions five times: more than 10,000 answers, but these are repeated runs of a small set of questions about debts, loans, mortgages, pension savings, taxes, savings and student loans. According to Alexander, it was the US market that was studied, where the systems have the most integrations, but models can know well-documented countries very well. An answer failed for a factual error or for missing an essential point or a necessary warning.
The result: 57% errors on average, 63% for free models, 49% for paid ones, 88% on the complex questions and 93% for free models on the hardest ones. At the same time the paid versions were cheap ones: no 5.6 Sol Pro, no 6.0 Astra Pro, no Fable 5.1, no Extra High or Max modes; a Pro version, by Alexander's estimate, would be wrong some 20–30% of the time. To those who begrudge $20 for a subscription he points out that today this buys a few cups of coffee or a kilo of meat, and urges them to understand the risks of the tools they use.
The errors are concrete. Claude Haiku 4.5's advice on pension taxes could, by the authors' estimate, have led to an additional tax charge of £17,500; a model advised paying off the highest-rate debt first without taking rent, taxes and arrears into account; Claude invented a student-loan rule for people who move abroad; Gemini assured a borrower that a mortgage payment holiday would not affect their credit score — in the US that can cost several years of a record. Sam Altman said a year ago that a model should not be trusted without limits, and ChatGPT itself says it can make mistakes, yet people act on autopilot.
By model: Claude Haiku 4.5 — 82%, Gemini 3.1 Pro — 73% despite all of Google's integrations and knowledge base, Grok 4.5 — 60%, ChatGPT 5.6 Luna — 58%, and the report does not say which mode Luna ran in; Alexander himself does not use such models. The best result went to Claude Opus 5 in reasoning mode, 39%, but even that is a big question for a serious modern model. A separate category of errors: the models ignored upcoming changes, tax changes for example.
Connecting data does not solve everything: GPT Finance breaks down spending, builds budgets and keeps a memory — memories — but it does not know about your illness, a unique exception, an inheritance or a sudden change in income, and accountants, banks and lawyers give differing information themselves. Alexander does not like plugins because he does not fully understand which rules are set up inside them; Claude has connectors and a new system for financial advisers, but connecting data does not mean granting access to payments — the system can be given analysis alone. Third-party finance apps run on cheap versions — he himself got a report for one of his companies from a system that Codex wrote to bypass QuickBooks.
The upshot: the figure struck Alexander, and for the coming months he may be even more careful with finances through AI models, although he still considers these systems smarter than any person at working with data. What matters is which questions you ask them, how you phrase them and whether you double-check the answer.
The value of the episode is that the argument about trusting AI with money is turned into figures — 57% errors on average, 88% on the complex questions, 39% even for the best model — and into concrete scenarios where an error costs money: from an additional tax charge to several years of a record in a credit history. The host does not give up on AI in finance — he gets the reporting for one of his companies from a system Codex wrote and considers the models smarter than any person at working with data — but insists on strong models, double-checking and questions from which a model learns about you what is not in the connected accounts. The practical conclusion: connected financial data makes an answer more convenient, but it does not make it correct.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 48 segments: 48 identified, 0 mixed, 0 probable, and 0 unresolved.
Read transcript on a separate page
Loading…