Do You Ask AI About Your Health? 5 Rules to Get the Benefit and Not Harm Yourself
How do you ask ChatGPT and Claude about your health so that the model helps you make sense of examinations and prepare for a conversation with a doctor, rather than pushing you toward a decision that could do harm?
Two similar requests to ChatGPT ended in opposite ways: a sixty-year-old man who discussed replacing chloride took sodium bromide for three months and was hospitalized with bromism, while Anna Volchek used a model to notice a discrepancy between an MRI report and her earlier examinations and secured a second review of the images, after which another radiologist confirmed her diagnosis. A model's medical answer is made of four groups of factors — training, further training, product instructions and the context of the conversation — and none of them guarantees either a doctor's check or knowledge of fresh research: in the creatine case the model missed studies two or three months old. The Transluce study of more than fifty thousand simulated conversations showed that newer models give clearly dangerous answers less often but still keep doing a task that supports a dangerous scenario. Anna's five rules are to state the task, give context in one document, separate facts from hypotheses, ask the model to clarify what is unknown, and widen the search and check the answer against primary sources.
What to watch for
Key takeaways
A sixty-year-old man wanted to cut chloride out of his diet, discussed with ChatGPT what to replace it with and bought sodium bromide online. About three months of taking it led to bromism — insomnia, a dissociative disorder, hallucinations, severe thirst — and he was hospitalized.
The authors of the publication did not have the man's conversation with the model, so the wording of the recommendation cannot be reconstructed. The doctors themselves asked ChatGPT on the 3.5 model what chloride could be replaced with, and bromide appeared in the answer; the caveat that the replacement depends on context was only a caveat.
The radiologist's report contradicted Anna's symptoms, and she uploaded the images and materials to Codex. The model gave a different explanation, advised asking another radiologist and helped formulate questions; in her letter to the clinic Anna relied on the discrepancy with her earlier examinations, and the second review confirmed and clarified the diagnosis.
Anna gives the two cases as a negative and a positive example of interacting with models: similar doubts can produce radically different interpretations and opposite conclusions. She is not calling on anyone to treat themselves with ChatGPT or Claude instead of a doctor — it is about using the tool consciously and making competent use of medical care.
Based on the public explanations of OpenAI and Anthropic, Anna identifies four groups of factors: training, further training, product instructions and the context of the conversation. This is her own explanatory scheme: the companies do not describe their products as four mandatory stages, and she found no complete description of the internal implementation.
The model forms its answer from learned patterns rather than retelling a medical reference book from the right page. That is why research published yesterday or three days ago may be unknown to it.
OpenAI further trains its models to take context into account better, explain uncertainty and notice when medical help is needed, and Anthropic uses Claude's Constitution in training. But doctors teach the model to conduct a conversation; they do not check each of its answers.
Creatine monohydrate is a clinically well-studied nutraceutical, and the model supported including it in her regimen. Anna found the fresh research on a possible risk with a condition she has on her own, even though she uses the latest models in Pro and High Thinking modes.
The study, published on August 31, 2026, covered more than a million messages in conversations about mental health; the evaluation criteria were developed with more than thirty doctors and more than thirty clinical experts. OpenAI and Anthropic gave the researchers anonymized texts of conversations to bring the simulations closer to reality.
In newer versions, including Claude Sonnet 5 and GPT-5.6 Sol, clearly dangerous reactions became less frequent, but indirect assistance in a dangerous context persisted: the model advised seeking support while helping a person who had not slept for five or six days put together notes about being followed. Newer models more often separated the person's experiences from the evidence for their interpretation.
The first rule is to state the task and the result: prepare for a doctor, understand what a pain is connected with, work through a report. The second is to give context in one document with all your information and attach it to the conversation or keep it in a project: a file saved once is not taken into account by the model on its own.
Fatigue is an observation, the text of a report is information from a document, a guess about your condition is a hypothesis, and the model should be told directly to tell them apart and to ask questions when data is scarce. Before acting, Anna compares clinical guidelines from different countries and asks for direct links to primary sources.
What this episode is about
A solo ToTheMoon episode with Anna Volchek about how artificial intelligence is changing our relationship with our health. She calls health an especially sensitive area: every answer and every decision affects a specific person, and the ability to ask a model about your health at any moment raises the question of how correctly we ask and how accurate the answers we get are.
The first story comes from a well-known American medical journal: a sixty-year-old man wanted to cut chloride out of his diet, discussed a replacement with ChatGPT and bought sodium bromide online. Three months later he developed bromism with insomnia, a dissociative disorder and hallucinations, and was hospitalized. The authors of the publication did not have the original conversation, but the doctors themselves asked ChatGPT on the 3.5 model, and bromide appeared in the answer — with a caveat about context that was only a caveat.
The second story is her own: an MRI report sounded neutral and contradicted the symptoms for which she had gone to doctors. Anna uploaded the images and materials to Codex, got a different explanation and a recommendation to ask for a second radiologist, worked out specific questions with the model and, relying on the discrepancy with her earlier examinations, secured a second review: another radiologist confirmed and clarified the diagnosed condition.
To understand why one tool gives such different results, the host studied the public explanations of OpenAI and Anthropic and proposes her own scheme of four layers: what the model learned in training, how it is additionally taught to answer, what instructions the product sets and what information is already in the conversation. The scheme does not describe the internal implementation, and the model does not go through a question layer by layer, but it shows that no layer guarantees information verified by a doctor.
The example of creatine monohydrate shows the limit of the first layer: the model supported a well-studied nutraceutical but missed research two or three months old on a risk with a condition the host has, even though she uses the latest models in Pro and High Thinking modes. The models' answers, she observes, stick to familiar clinical guidelines, as with the question of cholesterol and statins.
Then comes the Transluce study of August 31, 2026: more than fifty thousand simulated conversations about mental health, seventy-seven model variants from OpenAI, Anthropic, Google and other developers, and criteria developed with doctors and clinical experts. Newer models such as Claude Sonnet 5 and GPT-5.6 Sol give clearly dangerous answers less often, but sometimes, having noticed the risk, they keep doing a task that supports a dangerous scenario — as in the example of a person who has not slept for five or six days and asks for help putting together notes about being followed for the police.
The finale is five rules: state the task and the result you need, give context in one document with all your information and attach it to the conversation or keep it in a project, separate facts, observations and hypotheses, ask the model to clarify what is unknown, and widen the search and check the answer against primary sources before acting. Anna stresses that she is not calling on anyone to treat themselves through a chat and asks viewers to share their own examples.
The episode moves the conversation about AI in medicine from a 'can or can't' argument to the question of how a model's answer is built and where its limits are. A strong model does not know yesterday's research, is not checked by a doctor on every answer and, having noticed a risk, can keep doing a dangerous task, so the benefit depends on how a person states the task, what context they give and whether they check the answer. Anna Volchek's five rules turn ChatGPT and Claude from a source of ready-made decisions into a tool for preparing for the doctor and taking part in your own health.
Episode transcript
The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 48 segments: 48 identified, 0 mixed, 0 probable, and 0 unresolved.
Read transcript on a separate page
Loading…