Should you trust artificial intelligence with financial advice? And, more broadly, with everything to do with handling money. I think today's topic is a very important one, especially given that it's already September — almost the end of September 2026. If a question like this had been on the agenda two years ago, or a year ago, it would probably have sounded a bit like hype. You could have said: now you can upload any kind of report into artificial intelligence — statistics, your accounts, contracts, tax data — run all sorts of analyses, and so on.
And these were quite real cases; here on our channel, ToTheMoon, we told a huge number of them, gave examples and the circumstances behind them. Cases where the systems work far better than, you'd think, any financier, accountant, lawyer or anyone else. But again, why is this an important topic today, at the end of September 2026? Because the company Saturn has just carried out a study called "Artificial Authority. Should You Trust Artificial Intelligence With Financial Advice?" The data has literally just come out, and fairly fresh models were in it.
There was Grok 4.5, there was ChatGPT 5.6, there was Gemini 3.1 Pro, there was Claude Opus 5. So, fairly good models. I want to point out right away that these are not Pro-level models, and today we'll talk about whether the Pro level can really differ that fundamentally. You all know that on this channel I push very hard — that's the word for it — for people to switch to paid modes, and not just ordinary paid ones, but Pro modes. Not because of how many tokens someone gets, but to get access to other models.
It's not a question of tokens in Codex or Claude Code, but of quality — the fundamental quality of the answer. And I talk about this a lot. Literally twenty seconds for those who are on our channel for the first time. Over a long time our channel has released around 170 episodes in total, and we broadcast mainly from Silicon Valley, from the US, about artificial intelligence and the tech world, with a big background and a lot of experience in business and in the tech world in general.
And I strongly recommend you join us. And, of course, to understand what goes on here at ToTheMoon: this is a channel about how a person really takes shape as they develop in the tech world, in artificial intelligence. Our audience is varied. On the one hand, I wouldn't say it's for hardcore geeks. On the other hand, I'm quite a big geek myself in many things. I recommend everyone join and watch, especially if this world isn't entirely clear to you. And especially for those to whom it is clear — who want new opinions, new views, and not just some strictly practical theory. Though we try to give that sometimes too.
And at the same time, all the way to philosophical reflections. That's the broad range. So, don't forget to support us. It matters for our channel. Comments, likes, subscriptions. So, look, this study interested me a lot, because — why did it interest me? When I saw the error, when I saw what percentage of the questions were answered wrong. And to be honest, at first I found it hard even to believe. If I'd seen that the errors were ten percent, I probably just wouldn't have paid attention; I'd have thought, "Well, of course, artificial intelligence makes mistakes."
We all understand that any system can make mistakes. Even a written algorithm can be wrong; it can be written incorrectly. And if we talk specifically about modern artificial intelligence models, of course they can make mistakes. You have to understand that calmly and normally from the start, and live with it calmly. But when I saw the real error rate of these systems, especially given the current models — I wasn't exactly alarmed, but I decided to make a separate episode for us, both to tell you about it and to discuss it, so that you can share your opinion on what's going on.
Again, let me remind you — share your opinions about how systems like ChatGPT, Claude, Grok and Gemini work. By the way, there was more in the test. In fact they tested eighteen models, but we'll limit ourselves to the top models when it comes to tests. But still, you should be asking these models financial questions every day in order to really understand them. I ask very broadly about financial matters, including personal ones: accounts, credit ratings. In the US there are credit ratings; of course, not every country has them.
Everything to do with financing, investing, purchases, stocks, market analysis, taxes — personal and company taxes — all kinds of bookkeeping work, services and everything else. So a great deal adds up. Recently I told you about a case showing that even intermediate accounting systems like QuickBooks may, I think, become unnecessary at some point. QuickBooks is an American system. You can rewatch that episode. I think it's a very good one on the subject — specifically on building such systems.
When I used Codex to analyse all my accounts, all the details, and how competently the whole of my reporting had been put together. And here, you'd think: if a system built you reporting like that, how can anyone talk about such a large percentage of errors? So, what was done? I'm giving you a short, quick bit of theory, and today I also want to tell you where all these systems are heading. Quickly, too. Where Gemini is heading, where Anthropic is heading. Well, Google's Gemini. Where OpenAI is heading — in particular, the development of financial tools.
Can you imagine: I just opened ChatGPT today, right before our — well, before I started recording — and it recommended that I connect my credit rating; it connects it through the Experian system. That's a credit-rating provider in America, in the US. So: connect through that system and analyse my credit rating. And that's exactly the work of GPT Finance — a financial module, a financial plugin you could say, or a financial part developed separately for the systems. By the way, on Friday we put out an episode about plugins. Watch it. A new, new time, a new world. An especially good story for September 2026.
Should you use plugins? And we understand that we can ask all these systems through some separate plugins, or we can ask directly through the chat, or directly through Codex, about your finances. So you can send your tax report and say, "Prepare my tax report for next year," or "Go through it, find the errors in it for me," or send your accounts and say, "Add it up for me, show me what I spend my money on," and so on. So you can do it directly in the chat, or through the various plugins that keep appearing, or through connectors. So, what did this company, Saturn, do?
It took eighteen models — ChatGPT, Claude, Gemini, Grok, Copilot — and prepared 121 different financial questions, each asked five times to check how stable the answer was. In essence, a test like that seems easy enough to run, but at the same time it's serious, because it came to more than 10,000 answers, and you and I could hardly run tests like that so quickly. Ten thousand is serious. It's important to understand that these are not 10,000 different tasks but repeated runs of a relatively small set of questions across roughly the same systems — the same questions repeated five times.
So, what topics did the questions cover? There was debt management, there were topics about loans, mortgages, pension savings, taxes, savings. Then there were questions on, I don't know, student loans and so on. Obviously it was the US market that was studied, but I want to say right away — since our channel is watched in all sorts of countries — that this concerns your country too. You might think that Claude Code or ChatGPT know your country's legislation, or its tax and financial system, less well. And in fact that's probably true.
They know it less; they probably know the US better. But at the same time they may know some countries extremely well — if a country has been very thoroughly documented and openly described. And you shouldn't assume that, say, in Italy — somewhere where, I don't know, the information space isn't super-developed — or in some African countries, the results in terms of this financial literacy will be poor. Still, in the US the systems have far more integrations. And I'll explain why right now.
Because, for example, there's GPT Finance — the very system that offered today to connect me to Experian. So this system — imagine — lets you connect accounts at more than 12,000 financial institutions. And of course that's not only the US. That's important to understand too. Banks, cards, various brokers through Plaid. But 12,000 is a serious number. You could even… I wouldn't even have guessed a number like that. If someone had asked me how many integrations a system like that has, I'd have said, "Well, a hundred, but not twelve thousand." Twelve thousand is really serious.
So, what counted as an error when they asked these 10,000 questions? An answer failed if it contained a factual error or left out an essential point or some necessary warning. And what was the result? Ta-da! You can guess. What came out? Once again — once again, please, hear this. Claude Opus 5 — a serious system, a good system. ChatGPT 5.6, but it was Luna. Let me remind you that there's Terra, there's Sol and there's Luna. It's not some extraordinary model. It's a fairly simple model in the 5.6 version, but it's still a model — much better than a huge number of the models that existed a year ago. There was Grok 4.5, the good 3.1 Pro version — Gemini, a serious version.
There was also Claude Haiku 4.5, though many people may not even know it or use it. It's not really worth dwelling on, though I'll give you its results. So, on average all the models had an error rate of 57%. I have to say, of course, that's very serious. And I told you I wouldn't have paid attention to 10% — you'd think 10% isn't serious, because there's always some probability. But with your own money, 10% is already serious too. Because can you imagine what it means to have your own tax error?
Imagine filing your taxes with a 10% chance of making a mistake. That's a very serious error. Ten percent is one thing, but 56%? You sit there thinking: wait, do I understand correctly that if I ask about some data — about my accounts, or some of my subscriptions, or promo cards, or connecting something, getting loans or reports, payouts and so on — not to mention complex tax systems — the system might give me errors like that? Yes, it might. So, complex questions. There's this point: first of all, there are free models. I wanted to say that for the free models the error rate was 63%.
And it's an important difference, especially for the audience watching now. Everyone might ask: what was being tested? There are free models and paid models. The paid ones — clearly these were cheap versions. There's no version here like 5.6 Sol Pro, as it used to be, or, say, today's 6.0 Astra Pro. There's no Fable 5.1 here, no Claude 5 Ultra Code, no Extra High modes and the like. I mean the modes — well, there it isn't called Extra High, it's called Max; in Codex it's called Extra High. And for many people that's an important point.
Once again, to everyone who thinks, you know, "I'm not going to pay $2, it's too expensive for me. You, Alexander, don't understand what it means to pay for a paid subscription." But that's the average price of paid versions around the world. You don't understand what it is, but you should understand the risks of the tools you use, and what you can buy for $20 these days. What can you buy for $20 these days anyway — what are you trading it for? You're trading it for, I don't know, two, three, five coffees, or a kilo of meat.
What are you trading this for, this data, in terms of a subscription? And you know my case. More than once — just recently, in fact — I said that if there were such a thing, I'd be willing to pay easily $2,000 a month, if a system constantly scanned my chats, refined them and kept an eye on them. And someone said, "Oh, you've got your head in the clouds, you've gone too far." But someone else wrote something cool. He said: "Listen, that's just, like, $60 a day, roughly. And what do you think — is a person willing to pay $60 a day for getting fairly serious analysis running alongside their life, as it happens, if it makes sense for them in terms of their system, their worldview, their income and so on?" Are you willing to do that? What is $60 a day?
Are you willing to do it? How far is that, in your view of the world, an understandable — again, understandable — logic? So, complex questions. Look, on what Saturn counted as complex questions in its surveys, the error rate was 88%. And that's exactly why I wanted to share this with you and make this episode, because that is, of course, a very serious error rate. Eighty-eight percent errors is very serious. The paid models had an error rate of 49% — though, of course, these are not the $200 Pro versions.
But it does bother me, because then the Pro version should still come out at 20 or 30% errors, no less. On the hardest questions, the free models had an error rate of 93%. We understand perfectly well that these questions they asked — these 121 financial questions — let's be honest: there were hardly any questions of some critical difficulty among them. So, of course, a spread like that is unreal, unreal. In a moment I'll tell you, by the way, how each model did. But first, a bit about the errors.
Then you'll understand which model to use yourself, what strategy to adopt, and which models to work with on financial questions. Because financial questions — today's video — are for 100% of the people in the world; 100% of people in the world will be asking financial questions in systems like ChatGPT, Anthropic's, Gemini — everywhere, in every artificial intelligence system. It's obvious: it's a hugely important, necessary, widespread question. It's a topic that worries the whole planet. The topic of money.
How can we talk about professional development, about professions, about people's CVs, if mistakes are being made on financial questions? So, what errors were there? For example, errors on pension taxes. For example, Claude Haiku — version 4.5, for those who care about versions. It gave a recommendation which, by the study authors' estimate, could have led to an additional tax charge of £17,500 for the person if followed. And that's a calculation of a possible consequence, not a report that someone actually made that mistake.
So the system gave an answer, and that answer contained not just some micro-error — it carried huge consequences. Or, for example, there are more interesting cases, about the order in which to pay off debts. It turned out the model advised paying off first — paying off first — the debt with the highest rate, without considering what other priorities the person had. For example, rent, or the fact that there are other taxes in the country. And the authors point out that for a person — say, if a person has arrears, there are secondary issues they might have — then for them…
Roughly speaking, for a person with arrears, the consequences of not paying such obligations can be more serious than the extra interest they'd save right now. So on the one hand the system advises saving money by repaying first the loan with the highest interest, but in reality, on the other hand, a person may fail to repay some loan with low interest, and that loan will bring with it some very difficult problems in the family, for example, or difficult circumstances in the future.
There was also a problem with student loans, for example. Claude invented a rule that would let you stop them after moving abroad. And I want to say that, I think, this is a common mistake in general, when people ask: "What if I've moved to another country, say? What if I've moved to a different legal system?" How the system starts offering various options. And so on. So it was a very interesting area. There was, for example, the topic of mortgage payment holidays. Gemini, too, wrongly assured a borrower that pausing payments wouldn't affect their credit score. And can you imagine what that means?
In the US, for example, there's a credit rating. So if you make a mistake, that mistake can cost you several years of a record in the system. And you can make that mistake because today we trust a model's answers very strongly — often, even without limits. As Sam Altman said a year ago: "Well, let's be honest. A model can make mistakes, and you shouldn't trust it without limits. We never told you it couldn't make mistakes." We all know that when you ask ChatGPT a question, it says: "ChatGPT can make mistakes. Check important info."
That line is written right there: "ChatGPT can make mistakes. Check important info." It's written there. It says the system can make mistakes. But of course you start to believe the system, because you stop being in the mode where you want — where you'd double-check with someone once more, take some extra step once more — and you act on autopilot. And that autopilot can lead to serious consequences. And finance is the most important topic. And imagine you're writing a CV or choosing a profession — that's a serious, a serious task.
It's a task that deserves more than asking ChatGPT one question. Or, as people say: "You talk for too long. You could make a few-minute summary of this." But that's the point: if a question is serious, it deserves attention. So, as for the models that answered the various questions. Claude Haiku 4.5, a simpler version: an 82% error rate. That's the share of errors — a very serious share of errors. Gemini 3.1 Pro — 73%. And I want to say — I remember when, about a year and a half ago, Gemini came on very strong; Google really burst into the artificial intelligence market.
I think Meta was doing quite well back then too — among the top systems, I mean. Google burst in, and it felt like Google was really cool — great job. And you can see how Google has started to fall behind a little. Although, by the way, I'll note that Google's Gemini has a huge number of all sorts of integrations. For example, Google has a credit-rating integration; Google has various integrations for, say, professional investing. They have a special enterprise module for working with — well, for the corporate market, for the corporate financial market.
So they have a whole huge block of this. And Google does have, I think, an enormous knowledge base. And yet a 73% error rate in the Pro version really bothers me. Grok 4.5 — a 60% error rate. ChatGPT 5.6 Luna — 58%. And let me remind you again that 5.6 Luna is ChatGPT's weak model today. Plus, it doesn't show which mode the model was in — at what power it was running. Was it in fast-answer mode — what OpenAI calls instant? Or was Luna in the maximum mode, so to speak, that you can squeeze out of it? And that's worth understanding.
But I have to say that I don't use models like Luna, regardless of which mode I can choose in it. You can choose Luna and have Luna run in, I don't know, extra high mode, but I still don't really trust it. Now, what's interesting: Claude Opus 5 in reasoning mode — the report stated it, reasoning mode — had the best result. A 39% error rate. And let's suppose that Fable... First of all, reasoning mode comes in different kinds. And suppose it's Max reasoning mode, or Fable — Fable in Max, for example. And with Opus 5, too, a lot depends on the reasoning mode, not just the model.
I think in Max mode it would probably make even fewer errors, and with reviewing various things it might make fewer still — but it's still 40% errors in reasoning mode. Opus 5 is a really serious, modern model. That raises a big question about all of this. Now, there's a separate category of errors. The models ignored upcoming changes — things still in the future. For example, imagine there are tax changes coming that will take effect in the future. And it was shown that the models didn't handle — didn't fully handle — this information well.
So we don't know exactly — within this test I can't give you the finer details of what went on there. But what matters is that when you ask a question, you should understand from the start that the system may answer very correctly for today, while tomorrow the rules may change. You may file reports today with the data as it is, and in a month the data is different. Or, say, you file a report today, but based on data from a month or six months ago. The system may already work differently, and legislation can change a great deal.
New rules may be introduced, extra small nuances specific to your region, country, location or your own circumstances. And here the question arises: the system should probably start by questioning you very thoroughly. But what can a system actually find out? For example, ChatGPT Finance connects to, say, 12,000 of your systems. It can connect to your bank accounts, break down all your spending by category, compare the current months with previous ones. Figure out how you spend money, build budgets, analyse your various debts, investments, portfolio value and so on. But!
Just because you've connected all these systems doesn't mean the model has really learned from everything, really understands everything going on with you and knows all the extra details. Because all these systems actually have one significant limitation. The system can't, for example, work out on its own, on your behalf, certain questions that at certain points just aren't there — some unique additional things; it may simply not know them. Some very special circumstances. I don't know — that you have some particular illness, or you were born in some particular city on a certain date, or you have some — you fall under some unique exception, or you had some specific additional case.
Today we already understand how many accountants there are giving different information, and banks giving different information. Then courts giving different information, attorneys, lawyers giving different information. Can you imagine how this chain overlaps? How many different strategies there are when it comes to your activities, your work and especially your plans. What lies ahead — how it's right or wrong for you to build a certain strategy. It's a big question: how you'll live, how you'll earn money.
Whether you plan… There's a big difference. Say you earn nothing. Suddenly you earn a lot of money, and then, for example, you suddenly lose it. Or, I don't know, someone suddenly received an inheritance, it just happened, or something else. There are a great many nuances, a huge number of nuances the system doesn't know about you. Can a system connected to your systems learn all this with any guarantee? It can't. You have to remember that — that even memory... By the way, ChatGPT Finance has a separate memory inside — memories. But what matters is that ChatGPT Finance is limited.
This financial thing in ChatGPT — as I said about plugins — is still a system limited to the data it has received, and it may not even load some additional laws into its plugin. And your circumstances are sometimes better handled without any plugins. I, for one, am not a fan of plugins, because I don't fully understand what rules are set up in them. I'd rather work with the whole thing, with general artificial intelligence, and give it particular files or data separately, and work with those files and data.
Although these existing plugins are interesting: ChatGPT has them, Gemini has some in the form of skills, Claude has them, for example. Claude has a range of connectors you can set up — quite a large number. What's more, Claude has released a special system for, say, financial advisers, but that's more for the professional market. They launched it literally a week ago, and there, in Claude, you can also connect, for example, all the transactions you have. In fact, today, through connectors and your own development, you can connect practically everything — statements, banks, details — and build a huge number of different integrations.
And then… By the way, that doesn't mean you give the system the ability to make payments. You can, for example, give the system the ability to analyse your data. And here, by the way, a question arises: do you need any separate mobile apps, or separate apps in general, for managing your finances? A great many such apps have been made. If anyone uses them and thinks they're great, write in. I don't understand why they're needed. By the way, people sometimes send things to me at ToTheMoon, saying: "We've written an agent that works better with connectors.
Or we've written a connector, or built a plugin, or some mobile app, or software that lets you do something." I'll say it again: I personally am not interested in using third-party apps outside the current systems — Gemini, Claude, Grok, ChatGPT — or development systems such as Codex or Claude Code. Not interested. Why? Because any mobile app, first of all, uses cheap versions, and no mobile app will ever analyse my financial report, say, through GPT-6 Astra in ChatGPT. And it will have various limitations too. Whereas today, in ChatGPT's GPT-6 Astra, I can give a really serious set of different parameters.
And, by the way, the system Codex wrote for me — just for one of my legal entities — bypassing QuickBooks, that was really cool, because in the report it produced in the end I see a fairly high level of reliability. Of course I still have accountants, and at some point they'll still double-check things, but the data has been gathered for them far better than QuickBooks could have captured before, through routine transactions, expenses or other routine things that happened. What do you think about finances in general? How does it work for you — what queries do you put to systems like ChatGPT or Claude? What do you use to ask?
Do you double-check these questions across different systems? Who has filed tax reports, prepared various reports and filled in forms through such systems? We all know that GPT-6 Astra, for example, was strongly positioned on exactly this: that the system can work with third-party services and is capable of working more without a human, without human involvement. What do you think about it? Of course, this figure struck me, and I think that over the next few months I may be even more careful with finances through artificial intelligence models.
Although I believed, and still believe, that today these systems are smarter than any person when it comes to working with data. The other matter is which questions you ask them and how you phrase those questions. See you in our next episodes.