Skip to content
Artificial intelligence · Claude Design · GeminiEpisode 157 · 30 August 2026 · 51:03

Where Does AI Really Help, and Where Does It Only Make Things Worse?

Central question

In which tasks does AI genuinely make a specialist stronger, and in which does using it add errors and false confidence — and how do you see that line before the mistake becomes expensive?

What you take away

The reader gets a working split of tasks into two classes: those where a re-check through AI almost always wins — the arithmetic of an estimate, whether figures reconcile, spelling, the completeness of a request — and those where it destroys the result: style, individuality, taste, trust in a live person. Plus an explanation of why the mistakes of a clerk or a reviewer grow not out of models but out of a workload that went up, and what processes do about it.

Main threads

What to watch for

1Split your tasks into two lists: those whose answer has a checkable criterion and those that do not. Send the first to a model for a re-check, not the second.
2Before sending an estimate, a report or a calculation, run it through a model for reconciliation and arithmetic — cheaper than handling the client's bounce-back.
3For letters, keep two modes apart: fixing spelling and punctuation in your own text — yes; generating the letter from scratch instead of you — no.
4If you bring a specialist a comment from a model, bring it as a question, not a verdict: both sides have to allow that the model is wrong.
5Write down the team's review process for generated work: who reads it, who signs off, what may not pass without a human.
Signals to track afterwards
Whether an industry platform for legal work becomes a product of its own with its own factual base rather than a wrapper over a general model.
Whether courts and regulators answer such cases by restricting the tool instead of requiring the check.
What happens to scientific journals once reviewers run short: slower publication, more expensive review, or checks by agents.
How software development solves the review bottleneck: more people, paid reviews, or rules for agents.
Whether generation reaches the tasks of taste — interiors, clothes, the tone of a letter — where today it loses to a person.
Most useful for
Specialists in licensed professions — lawyers, doctors, auditors — who already use models in their work.Developers and engineering leads for whom review has become the bottleneck.Designers and anyone selling a result that cannot be checked with a formula.Managers who receive reports and estimates and want to know what to check themselves.Anyone who wants to settle for themselves where to use AI and where to go to a person.

Key takeaways

01:21Google Is Building a Separate Platform for Lawyers

Gemini Enterprise for Legal is a corporate platform for legal work: industry integrations, agents, confidentiality, data control. The author argues that the legal domain needs exactly such a wrapper: a model that gives an opinion and can be wrong, like a lawyer, on one side, and a base of verified data on the other.

02:15A Judge's Assistant Put Invented Material Into a Case

In the federal appellate case Reuters reported on, a judge's clerk used Perplexity and non-existent material ended up in the filing: wrong parties to the case, statements absent from the record, phrases that are not in the law itself. The court considered reassigning the case to a different judge.

02:46Nine out of Ten Biomedical Papers Carry Traces of AI

Almost nine out of ten fresh biomedical papers show signs of AI assistance. That does not mean a model wrote them: the material may have been processed, part of the text prepared, part of the research done with its help. What matters is that this is a field where checking is taken seriously.

05:39Mathematics Proves a Fact, Law Defends an Opinion

Ilnar Shafigullin recalls a conversation with a fellow student who was a lawyer: a mathematics thesis derives a proof, and opinions have no place in it; a legal thesis forms and justifies an opinion that may turn out right or wrong. Hence the different footing for applying models in the two fields.

07:16A Legal Data Set Is Laborious to Build, but Possible

Law has an enormous history of cases: what went in, what testimony, facts and evidence were gathered, what the court concluded. A model can be tuned on such a data set and can then handle cases and the legal framework. The caveat: train it on one set of judges' decisions and you get a model that behaves like those judges.

09:13With AI the Load Goes Up and People Cut Corners

The mechanism that will keep producing such cases: work seems easier, but tasks get done faster and more is demanded of the person. Checked five or six times, the model was right — the seventh time they did not check, because a deadline was burning. The cure is not a ban but process: checks, gatekeepers, up to a model that checks a model.

15:12How Trust Works in Science

Until a paper is published in a peer-reviewed journal, it is treated as not existing. Reviewers work blind through every derivation, put questions through the editors and in effect sign their name to it. Beyond that, journal rankings and reputation do the work. In mathematics a reader can re-prove the theorem themselves — at university that is done on purpose.

18:43The Bottleneck Moved From Writing to Reviewing

In software development the problem has long stopped being writing code: anyone can generate it. The problem is letting a change into the project: someone has to read it, be sure nothing breaks and vouch for it. A pull request here is exactly a paper at a publisher, and reviewers are just as short. Hand that to agents with no human and you get the clerk's story in another industry.

21:39Re-checking as a Way to Extend Your Own Understanding

The other side: the author runs everything brought to him through a model — finance, marketing, sales, product analytics, infrastructure — and looks at whether the figures reconcile and whether some formula was typed in by hand wrongly. The point is not to set this against the person but to fill in the picture and see the probabilities.

26:25The Designer's View: More Harm Than Good Right Now

Tatyana Tsvetkova says that at this stage AI does more harm in her work: a client comes for confidence in decisions, already carrying a thousand options from a husband, a mother and friends, and on top lands a model with a huge volume of analysed data — and the person ends up terrified. In medicine or mathematics, where things are sharper, it works better.

28:04The Client Leaves Whoever Does Not Adapt

The flip side of the same thing: the author has stopped working with analysts, marketers and editors who stick to old methods, and says the same awaits lawyers and accountants. A live example — the contractors painting a veranda: had they brought a printout, a description, a slightly wider conversation, the next order would have been theirs.

30:04A Sanity Check Before Sending, as the Norm

Ilnar proposes a simple rule of good manners: before sending your work, ask a model whether everything is in order and whether there are errors. For an accountant, an analyst or a builder with an estimate it would take a minute and make working with the client much easier. Tatyana's objection: the model is wrong and hallucinates too, and it is one more opinion, not the truth.

31:07Endless Re-checking Flattens Individuality

Everything becomes uniform. Tatyana tried Claude Design on two urgent client decks and agreed with the viewers: they come out all alike, and by the time you have explained what you want it is faster to do it yourself in a familiar tool. The same with letters: the recipient senses the text did not come from a person, even without the em dash.

36:04Proofread Versus Rewrite: the Difference Is One Switch

The author shows the line on Apple's built-in text checking. Proofread fixes errors inside words and places the commas. Rewrite rewrites the text the way the model thinks right, and carries the style, the meanings and the emphases away with it. A request to fix only the spelling, given to a model in a deep-reasoning mode, behaves exactly like Proofread.

38:14The Letters the Author Hands Over Entirely

Letters to doctors, to support desks and to government bodies he writes with a model one hundred percent: what matters is not whose words they are but that the request is put as completely as possible. One example is a letter demanding a shopping centre preserve its video recordings: he could not have written it without reading through state and federal law, and the answer came straight back.

44:02AI Raises the Lower Level, Not the Upper One

Ilnar's formula, which reconciles both positions: a model raises the lower level well and does not reach the top. Good programmers still understand code better than it does; weak ones lift their floor. The same in design: the decks he makes five times a year and dislikes, Claude Design covers perfectly — Tatyana's client projects, it does not.

47:05Where Generation Does Not Work: Interiors and Appearance

In interfaces, on the web and in decks Claude Design gives the author dozens of alternatives he would once have had to ask a single designer for. But he does not use image generation for interiors or for a person's appearance: it comes out unreal. The question of how to place two sofas looks silly and in fact needs a person who answers in one click.

49:01The AI Picked the Paint, a Person With Taste Did the Work

The closing example separates the two contributions. A friend painted the garage and chose the paint entirely with a model — Tatyana, arriving, noted both the colour and the quality. But the model picked the paint and a person with an eye and a vision did everything else; someone else with the same choice would not have got there. Hence the episode's conclusion: keep the whole spectrum open, from your own decision to a designer.

What this episode is about

A three-way conversation. The starting point is two news items from the same week that point in opposite directions. Google has extended its corporate platform for legal work, Gemini Enterprise for Legal: industry integrations, agents, confidentiality, data control. At the same time Reuters describes a federal appellate case in which a judge's assistant used Perplexity and put non-existent material into a filing — wrong parties to the case, statements absent from the record, phrases that are not in the law itself. Alongside sits another number: almost nine out of ten fresh biomedical papers carry signs of AI assistance.

The hosts agree this is not a verdict on the models. Licensed work assumes checking, and no checking happened: the cases could have been requested, screenshots obtained, the codes run separately. Ilnar Shafigullin adds the mechanism that will keep producing such cases: AI seems to make work easier, but the load goes up, more is demanded of the person — and they start cutting corners. Checked it five times, the model was right, the sixth time they did not check.

From there the episode moves to its central analogy — scientific peer review. Ilnar uses mathematics to explain how trust in a journal is built: blind reviewers, working through the derivations, a signature under the publication, the journal's reputation. And he shows that exactly the same bottleneck already stands in software development: code is generated in any quantity, but the people who have to read it, vouch for it and let it into the project are as few as before. Tatyana Tsvetkova unexpectedly recognises her own field in this: interiors go to reviewers the same way, get published or not, and a designer's name depends on it.

The second half is about where re-checking helps and where it gets in the way. Alexander Volchek defends the position of sending an estimate, a report or a medical finding to a model and getting a second opinion, with his own examples up to a legal letter he could not have written himself. Tatyana pushes back from the other side: the model hallucinates too, endless re-checking flattens individuality, and decks made in Claude Design come out all alike. Ilnar draws the line with a formula: AI raises the lower level well and does not reach the upper one — it pulls the weak up a lot, and gets in the strong one's way less than it seems.

The line the episode draws runs not between professions but between kinds of task. Where there is a checkable answer — do the figures reconcile, is the formula right, are there errors in the text, is the request complete — a re-check through a model is almost always a gain, and refusing it now reads as unprofessional. Where there is no answer but taste, style and trust — an interior, a wardrobe, the tone of a letter — the model produces something uniform and confident, which is the worst possible outcome. The hosts' shared conclusion is not about the tool but about an open mind: be able to do both, and honestly decide each time which case you are in.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 144 segments: 144 identified, 0 mixed, 0 probable, and 0 unresolved.

Loading…