Where Does AI Really Help, and Where Does It Only Make Things Worse?
In which tasks does AI genuinely make a specialist stronger, and in which does using it add errors and false confidence — and how do you see that line before the mistake becomes expensive?
The reader gets a working split of tasks into two classes: those where a re-check through AI almost always wins — the arithmetic of an estimate, whether figures reconcile, spelling, the completeness of a request — and those where it destroys the result: style, individuality, taste, trust in a live person. Plus an explanation of why the mistakes of a clerk or a reviewer grow not out of models but out of a workload that went up, and what processes do about it.
What to watch for
Key takeaways
Gemini Enterprise for Legal is a corporate platform for legal work: industry integrations, agents, confidentiality, data control. The author argues that the legal domain needs exactly such a wrapper: a model that gives an opinion and can be wrong, like a lawyer, on one side, and a base of verified data on the other.
In the federal appellate case Reuters reported on, a judge's clerk used Perplexity and non-existent material ended up in the filing: wrong parties to the case, statements absent from the record, phrases that are not in the law itself. The court considered reassigning the case to a different judge.
Almost nine out of ten fresh biomedical papers show signs of AI assistance. That does not mean a model wrote them: the material may have been processed, part of the text prepared, part of the research done with its help. What matters is that this is a field where checking is taken seriously.
Ilnar Shafigullin recalls a conversation with a fellow student who was a lawyer: a mathematics thesis derives a proof, and opinions have no place in it; a legal thesis forms and justifies an opinion that may turn out right or wrong. Hence the different footing for applying models in the two fields.
Law has an enormous history of cases: what went in, what testimony, facts and evidence were gathered, what the court concluded. A model can be tuned on such a data set and can then handle cases and the legal framework. The caveat: train it on one set of judges' decisions and you get a model that behaves like those judges.
The mechanism that will keep producing such cases: work seems easier, but tasks get done faster and more is demanded of the person. Checked five or six times, the model was right — the seventh time they did not check, because a deadline was burning. The cure is not a ban but process: checks, gatekeepers, up to a model that checks a model.
Until a paper is published in a peer-reviewed journal, it is treated as not existing. Reviewers work blind through every derivation, put questions through the editors and in effect sign their name to it. Beyond that, journal rankings and reputation do the work. In mathematics a reader can re-prove the theorem themselves — at university that is done on purpose.
In software development the problem has long stopped being writing code: anyone can generate it. The problem is letting a change into the project: someone has to read it, be sure nothing breaks and vouch for it. A pull request here is exactly a paper at a publisher, and reviewers are just as short. Hand that to agents with no human and you get the clerk's story in another industry.
The other side: the author runs everything brought to him through a model — finance, marketing, sales, product analytics, infrastructure — and looks at whether the figures reconcile and whether some formula was typed in by hand wrongly. The point is not to set this against the person but to fill in the picture and see the probabilities.
Tatyana Tsvetkova says that at this stage AI does more harm in her work: a client comes for confidence in decisions, already carrying a thousand options from a husband, a mother and friends, and on top lands a model with a huge volume of analysed data — and the person ends up terrified. In medicine or mathematics, where things are sharper, it works better.
The flip side of the same thing: the author has stopped working with analysts, marketers and editors who stick to old methods, and says the same awaits lawyers and accountants. A live example — the contractors painting a veranda: had they brought a printout, a description, a slightly wider conversation, the next order would have been theirs.
Ilnar proposes a simple rule of good manners: before sending your work, ask a model whether everything is in order and whether there are errors. For an accountant, an analyst or a builder with an estimate it would take a minute and make working with the client much easier. Tatyana's objection: the model is wrong and hallucinates too, and it is one more opinion, not the truth.
Everything becomes uniform. Tatyana tried Claude Design on two urgent client decks and agreed with the viewers: they come out all alike, and by the time you have explained what you want it is faster to do it yourself in a familiar tool. The same with letters: the recipient senses the text did not come from a person, even without the em dash.
The author shows the line on Apple's built-in text checking. Proofread fixes errors inside words and places the commas. Rewrite rewrites the text the way the model thinks right, and carries the style, the meanings and the emphases away with it. A request to fix only the spelling, given to a model in a deep-reasoning mode, behaves exactly like Proofread.
Letters to doctors, to support desks and to government bodies he writes with a model one hundred percent: what matters is not whose words they are but that the request is put as completely as possible. One example is a letter demanding a shopping centre preserve its video recordings: he could not have written it without reading through state and federal law, and the answer came straight back.
Ilnar's formula, which reconciles both positions: a model raises the lower level well and does not reach the top. Good programmers still understand code better than it does; weak ones lift their floor. The same in design: the decks he makes five times a year and dislikes, Claude Design covers perfectly — Tatyana's client projects, it does not.
In interfaces, on the web and in decks Claude Design gives the author dozens of alternatives he would once have had to ask a single designer for. But he does not use image generation for interiors or for a person's appearance: it comes out unreal. The question of how to place two sofas looks silly and in fact needs a person who answers in one click.
The closing example separates the two contributions. A friend painted the garage and chose the paint entirely with a model — Tatyana, arriving, noted both the colour and the quality. But the model picked the paint and a person with an eye and a vision did everything else; someone else with the same choice would not have got there. Hence the episode's conclusion: keep the whole spectrum open, from your own decision to a designer.
What this episode is about
A three-way conversation. The starting point is two news items from the same week that point in opposite directions. Google has extended its corporate platform for legal work, Gemini Enterprise for Legal: industry integrations, agents, confidentiality, data control. At the same time Reuters describes a federal appellate case in which a judge's assistant used Perplexity and put non-existent material into a filing — wrong parties to the case, statements absent from the record, phrases that are not in the law itself. Alongside sits another number: almost nine out of ten fresh biomedical papers carry signs of AI assistance.
The hosts agree this is not a verdict on the models. Licensed work assumes checking, and no checking happened: the cases could have been requested, screenshots obtained, the codes run separately. Ilnar Shafigullin adds the mechanism that will keep producing such cases: AI seems to make work easier, but the load goes up, more is demanded of the person — and they start cutting corners. Checked it five times, the model was right, the sixth time they did not check.
From there the episode moves to its central analogy — scientific peer review. Ilnar uses mathematics to explain how trust in a journal is built: blind reviewers, working through the derivations, a signature under the publication, the journal's reputation. And he shows that exactly the same bottleneck already stands in software development: code is generated in any quantity, but the people who have to read it, vouch for it and let it into the project are as few as before. Tatyana Tsvetkova unexpectedly recognises her own field in this: interiors go to reviewers the same way, get published or not, and a designer's name depends on it.
The second half is about where re-checking helps and where it gets in the way. Alexander Volchek defends the position of sending an estimate, a report or a medical finding to a model and getting a second opinion, with his own examples up to a legal letter he could not have written himself. Tatyana pushes back from the other side: the model hallucinates too, endless re-checking flattens individuality, and decks made in Claude Design come out all alike. Ilnar draws the line with a formula: AI raises the lower level well and does not reach the upper one — it pulls the weak up a lot, and gets in the strong one's way less than it seems.
The line the episode draws runs not between professions but between kinds of task. Where there is a checkable answer — do the figures reconcile, is the formula right, are there errors in the text, is the request complete — a re-check through a model is almost always a gain, and refusing it now reads as unprofessional. Where there is no answer but taste, style and trust — an interior, a wardrobe, the tone of a letter — the model produces something uniform and confident, which is the worst possible outcome. The hosts' shared conclusion is not about the tool but about an open mind: be able to do both, and honestly decide each time which case you are in.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 144 segments: 144 identified, 0 mixed, 0 probable, and 0 unresolved.
Loading…