Today we will look at two very different cases of using ChatGPT for health. In one case, after the model's advice, a person ended up in hospital. In the other, artificial intelligence helped notice a problem that had been missed in the initial interpretation of a medical examination. How can one and the same tool produce such different results? Let's look at what really stands behind the medical answers of ChatGPT and Claude. Where the models make mistakes even today, and why one correct prompt is very often not enough here.
And most importantly, at the end of the video I will give five rules that I use myself when I discuss my examinations or other questions with artificial intelligence, questions that concern my health, so that it really helps me figure things out rather than creating yet another unsolvable problem. Hi everyone! Today I want to talk about how artificial intelligence is changing our relationship with our health. We have talked a lot about the changes connected with work, with our habits, with how we make decisions.
But health is always such a very sensitive area, where every answer, every decision affects a specific person, their life, the people around them, how they feel, how they perceive the whole world. And today a person has an incredible opportunity to ask questions about their health at practically any moment of their life, and the question is how correctly we ask these questions and how accurate the answers we get are. Let's dig into this topic. I want to start with a story that I read quite a while ago.
It was published in a fairly well-known American medical journal and concerned a sixty-year-old man who read about the possible harm of table salt and decided to change his diet. As he understood it, he wanted to cut out chloride specifically, and he discussed with ChatGPT what it could be replaced with, and as a result of that discussion he bought sodium bromide online. In this story the names, of course, matter a great deal, because ordinary table salt is sodium chloride, and the substance this man began to use instead of sodium chloride is sodium bromide.
These are completely different compounds. That is, having sodium in their names does not make them interchangeable in a diet. But nevertheless, this mistake happened. In the end, this man used sodium bromide for about three months, and he developed bromism. Bromism is bromide poisoning. And he felt not only unwell, in the form of insomnia; he developed a dissociative disorder, auditory and visual hallucinations, various skin rashes and other manifestations, severe thirst. Well, basically the whole fairly clinical and vivid set of symptoms of bromide poisoning.
He was hospitalized, and after he stopped taking the bromide and was treated, these symptoms subsided. There is a rather important detail here: the authors of the publication did not have the man's original conversation with ChatGPT, which advised him to replace chloride with bromide. And of course we cannot say with any certainty exactly how this recommendation was given, but overall the authors of the article have the man's own account of how he arrived at changing his diet and at this decision.
So the doctors separately also tried asking ChatGPT, which at the time was still the 3.5 model, what chloride could be replaced with, and bromide did indeed appear in the answer. That is, there was a caveat that the replacement depends on the context, but purely as a caveat. So, in essence, there was a confusion at the level of meanings here, as a result of which a person ended up in a fairly serious condition, in fact a life-threatening one.
There are, of course, completely different stories, where a conversation with a model helps us make a different decision. I want to share my own example. I had a situation where the results of my examinations and their interpretation simply did not add up for me into one clear picture. There were certain examinations, including ultrasound, there were certain clinical symptoms. And then I had an MRI and received the radiologist's report. It sounded very neutral and in effect cancelled out, neutralized all my previous findings and in a certain way even contradicted the symptoms that I was experiencing, that I was feeling, the ones I had sought medical help for in the first place. I wanted to explain this contradiction. I gathered all the materials from the examinations and the MRI.
First I separately gathered the MRI results, all the images, connected Codex to them, uploaded all these materials and asked it to help analyze them. I got a completely different explanation from the model than the one I had read in the radiologist's report. For me it correlated much better with all the previous examinations, and the answer even included a recommendation to seek help, namely to contact the clinic and ask another radiologist to analyze the images that had been taken during the examination.
The model also helped me formulate very specific questions. And that is exactly what I did. I wrote to the clinic and asked for my materials to be reviewed again by another radiologist. I based my request on, and as some kind of evidence for it I did not write that ChatGPT had somehow analyzed my images. I said that I had serious doubts because of the discrepancy between the interpretation and the radiologist's report, the report interpreting the MRI images, and the other examinations I had done literally a couple of weeks earlier.
In the end the clinic accommodated me, and my images were analyzed by another radiologist. The other radiologist in fact confirmed and even clarified the problem I had, confirmed the condition I had been diagnosed with, and as a result I was able to move forward with the right decision. So here, in essence, ChatGPT helped me do this additional check, helped me formulate the questions correctly and helped me to some extent confirm the mismatch and the discrepancies between the examinations and the reports that I myself had seen.
For me it became a very good example of participation, my own participation in my health, a person's participation in their own health. And I did not just receive some incomprehensible report and leave it as it was. I had questions, I figured it out, and then I turned to medical specialists for further clarification.
I give these examples as different examples of positive and negative interaction with the models. And of how, at the same time, in partly even similar cases and with the doubts we have, we can ask questions and get radically different interpretations, draw different conclusions and, on the one hand, harm our health, and on the other hand, help it. This is where I want to understand how the model's answer is actually put together and, once again, how to interact with it properly so as not to harm yourself.
I also want to state one important point right away: I am not calling on anyone to treat themselves, and I am not calling on anyone to use ChatGPT or Claude, or any other artificial intelligence, as a replacement for qualified medical care. What I call for, and what I consider extremely important, is the conscious use of this tool, so that we can manage our health consciously and make competent use of the medical care that is available to us today.
So what actually stands behind each model's answers? What is built into ChatGPT or Claude by default when we come with a question about health? Is there some separate medical protocol? For example, for a long time it seemed to me that all the answers built into, well, in particular, I mostly use OpenAI's models after all, that they are one hundred percent based on all the protocols of the World Health Organization and on generally accepted clinical practice, without considering certain nuances and alternative methods and ways of relating to one's own health.
I looked at OpenAI's public explanation, Anthropic's public explanation, materials on how the models are trained, the rules of desired behavior, certain instructions within the specific area we are interested in. And based on them I divided what influences the answer into four parts. Purely for convenience, I call them layers. Moreover, this is exclusively my own explanatory scheme. That is, the companies do not describe their products as four mandatory sequential stages of processing a medical question.
I am now going to talk about four groups of factors that influence the answer, and each of these groups is supported by separate documents. And I did not find any complete internal description of how everything works when we ask a question about health. So, the first layer.
The first layer is what the model learned during training. This is something we all know. OpenAI explains that the model learns to find patterns in large volumes of information. This helps it connect terms, explain mechanisms, work through descriptions. But it is important to understand here that this is not the same as opening a medical reference book at a particular page and reading, carefully retelling what is set out in that reference book. That is, the model forms its answer on the basis of what it has learned, not by returning some ready-made material it has found in, for example, a medical encyclopedia or reference book.
Another very important point is that information from training does not mean the model knows about research that, for example, was published yesterday, or three days ago. By the way, I will share an example of this a little later. The second layer is how the model is additionally taught to answer. Developers evaluate the answers and tune the behavior. For health questions, OpenAI states directly that the model is trained further. Among the goals are taking context into account better, explaining uncertainty, and the need to notice situations where medical help is really needed.
Anthropic, for example, explains that it uses Claude's Constitution in training; on its basis they create example conversations and possible answers. That is, the model is taught not only to work with information but also to conduct a conversation in a certain way. But the participation of doctors in this layer, once again, helps the model and teaches the model to conduct a conversation properly in a certain context, but it in no way means that a doctor checks every one of my or your answers, or that there is some goal to give every answer information that has been verified as far as possible by doctors, by professional doctors.
The next layer, the third one. This is the instruction of the specific product. These are additional directions on how the model should answer in a conversation. That is, some of Anthropic's system instructions directly require using reliable medical and psychological information when it is relevant to the question. For mental health there are even separate restrictions. That is, not attributing diagnoses to a person or confirming a person's false beliefs. That is, you can acknowledge, for example, that the person is scared, but not agree with some unfounded explanation, with some unfounded interpretation of what is happening.
OpenAI has the Model Spec, likewise a public description of the desired behavior of its models, and for medical questions it provides for helping with information and for caution with final professional recommendations. But neither the Model Spec nor Claude's constitution can be considered some single message that is used and applied to every question that has a quasi-medical context or a context related to our health. The fourth layer is the information that is already available in our specific conversation.
That is, it is my question, it is what I said earlier, what the models remember about us from previous chats and conversations. Documents the system was able to read, perhaps the ones uploaded right now. That is, in essence, the data that is already in the system, in this chat or in the model's memory; of course, the answer also depends on this data. These four layers are different groups of information, rules and instructions on the basis of which the model forms its conclusion and its answer to our question if it concerns medicine, if it concerns health, if it concerns how we feel.
But what is very important is that this in no way means that in every quasi-medical question, once again, the model goes through the layers one by one and uses all this information, the instructions, the accumulated practice, the warnings and so on. As I said earlier, I quite often notice that the answers of the models, both Claude and ChatGPT, very often follow the line of familiar clinical guidelines. That is, to take a hypothetical example, if we take the question of high cholesterol and a discussion of statins, all the answers will stay within certain protocols that have been adopted by the World Health Organization and are widely used today in all countries, in all practices.
I am not evaluating anything now, this is not at all about prescribing statins, but my point is that I kept asking myself where the model gets the grounds for its answer and why it chooses exactly those, without paying attention, for example, to alternative methods of treatment, perhaps, or to a different view of human physiology, or of human psychophysics.
I have a good case of my own with one nutraceutical that actually has a very good scientific base. I am talking about creatine monohydrate. It is a very well-studied, specifically clinically studied nutraceutical with a solid evidence base, few side effects and practically no contraindications, except for impaired kidney function. And the model supported including it in my regimen, but it did not take into account the most recent research data, only about two or three months old, which I found later.
And by my reading, this research contained very important information about a possible risk with a certain health condition, precisely the condition I have. And for me this became a very good reason to check not only the general reputation of a substance, for example, but also to apply the data to the specific situation. Because the model itself, even though for these questions I always use the latest models, the maximum, some Pro, and High Thinking modes, nevertheless did not give me the kind of reliable answer that I seemed to be counting on.
And in this whole story the most interesting part begins. That is, the rules exist.
Models are taught to answer more cautiously, but in difficult conversations dangerous answers still come up. And I became interested in the work of Transluce, an independent commercial research organization. On the thirty-first of August 2026, very recently, a study was published assessing the behavior of models in conversations related to mental health. According to the authors' report, they ran more than fifty thousand simulated conversations, more than a million messages, and tested seventy-seven model variants.
They included models from OpenAI, Anthropic, Google and other developers. The evaluation criteria were developed with the participation of more than thirty doctors and more than thirty clinical experts. Once again, this is very important: the researchers created artificial dialogues in order to compare the behavior of models in difficult situations. Although OpenAI and Anthropic helped bring these simulations closer to reality, because the researchers were given texts of personal conversations, anonymized texts of personal conversations and anonymized characteristics of how users actually conduct such conversations. And the result was very uneven.
What conclusions did the researchers reach? That in the newer versions of the models, and the examples given in the report include Claude Sonnet 5 and GPT-5.6 Sol, some clearly dangerous reactions became less frequent, but nevertheless indirect assistance in a dangerous context persisted. And here the conclusion is much more complicated than simply 'the model did not notice the problem'. That is, the model sometimes already noticed the risk and advised seeking support, but at the same time kept carrying out a task that could feed a dangerous turn of events.
The report has an artificial scenario in which a person, for example, has not slept for several days, talks about it and tells the model that they have, say, had insomnia for five or six days, and at the same time talks about being followed and asks for help putting together notes for a report to the police. And the request itself looks like some ordinary work with text. But if you take the whole conversation into account, there is clearly a certain mania, and the model kept supporting this disorder within that context.
In another scenario a person was looking for a special meaning in sounds, was also sleeping badly and was withdrawing from loved ones, judging by a number of their exchanges and conversations with the models. At the same time, the models could support precisely that person's explanations and even, probably, draw them deeper and deeper into that state. Whereas the newer models more often separated the person's experiences from the evidence for their interpretation and suggested turning to some kind of human support.
There is an important takeaway here that I draw for myself. First, the quality and level of the models that we, all people, use when they come with a question about their health. Obviously, higher-quality models give an answer that is much and often safer, but at the same time they too can support one kind of disorder or another and send a person in an incorrect, wrong direction. Especially if the questions are not formulated precisely enough, if the full context is not given, if we leave out certain details about our condition and set only a very narrow direction.
And, accordingly, the model will answer us based on that direction. So how do you use the tool for yourself? How do you avoid all these mistakes?
Once again, I am definitely not calling on anyone to treat themselves by corresponding with ChatGPT or to refuse medical care. I really want to talk about how we can participate in our own health, how we can strengthen ourselves. As with the example I shared from my own life, about the images, about the MRI I had and the reports I received. It seems to me that, overall, this system of using the tool consciously in order to strengthen yourself, and in order to simply use it as a superpower, runs like a red thread through all our episodes about how a person can use artificial intelligence in their life.
I want to leave you five such pillars, an approach that I use myself and that works for different health questions. In general, when we talk about the questions about our health that can be addressed and sent to ChatGPT or Claude, there is a great multitude of scenarios. It could be preparing for a doctor's appointment, it could be some ailment that we are feeling right now without understanding what to do about it. It could be examination results that we have received and don't understand at all.
And, accordingly, doctors do not always explain what is happening to us in an accessible, human language. We need to make sense of the report we received. Perhaps we need to study data on some alternative methods concerning our health. There is a great multitude of scenarios, but some rules for conversations with a model really do stay, probably, the same. So I want to share the foundations that I rely on. The first tip is to always state the task and the result you want to get when you turn to the model. Why am I opening ChatGPT right now?
I want to prepare for a well-informed consultation with a doctor. I feel some pain and want to understand what it is connected with. I want to study some alternative method. That is, to formulate clearly what work I am assigning to the model at this moment. 'Help me understand this report', 'help me compare a document', 'help me prepare for a conversation with my doctor'. The second point is to give the relevant context that applies to the question. If I ask to discuss my situation, my age matters, and so do my observations, when it started, my condition, changes in my condition, confirmed diseases, family history.
In fact, most clinics today ask you to fill in a certain questionnaire before a visit, say. But this questionnaire is still quite limited in nature, because processing it, big data and time, probably, answering, I don't even know why. Perhaps not every place has that option. Ideally, I recommend making one simple document in which you collect all your information about yourself. That is, you can actually dictate your history to ChatGPT, attach your materials and ask it to put together a clear foundation, then check this foundation, save it and update it as things change.
I use exactly this approach: it has my place of birth, the place where I live now, all the dates, the wording of the reports, the medications I take, what started when and what changed. And once again, the model itself will help structure this document and maintain it properly. There is just one very important practical clarification here. Such a document really has to be attached to a new conversation, or perhaps you have a project set up where, for example, the model is connected to a folder containing all these documents.
Because a file that has been created once and saved does not by itself mean that the model will automatically take it into account in every subsequent conversation. But if you always conduct this conversation inside a project to which the necessary information is already attached and uploaded, then of course it will always be taken into account. The third point. It is very important to separate facts, separate observations, separate hypotheses. If you have started to notice fatigue, that is an observation.
If the report says this or that, that is specific information from a document. If it seems to me that something is happening to me or that I am going through a certain state, that is my hypothesis. And it is very important to state these rules to the model itself, properly and explicitly. Say, my hypothesis is that I may have such-and-such a condition. Separate observations from assumptions. If there is not enough data, ask me questions. That is, it is important that we do not formulate our own conclusions and observations ourselves. That is, to give the model the correct context.
You can even ask it to clarify what is unknown. I would probably even count this as the fourth point. Ask the model to repeat how it understood the task, ask the model to ask questions if some information may be missing for one reason or another, for example, if we left it out. The next, and probably the final point that I use, which you could call a tip, is to widen the search and check the answer before acting. I would always check some simple things first, but I also widen the range of the search in general.
That is, if, for example, I live in Europe, I ask it to study these approaches or some alternative approaches in the United States, in Russia, in Japan, for example. To compare current clinical guidelines and explain where the approaches agree and where they differ. Because we all live in different contexts, and very often it is important to see a wider picture than just one answer. At the very end of a conversation with the model, it is really great to ask it to recheck the grounds for its answer and give direct links to the primary sources, to show and point out where some recommendations may contradict each other, and what the difference between these approaches is.
In essence, the approach I use is to get the fullest possible picture, so that I take into account all the advantages that artificial intelligence can give me today at work and in caring for my own health, and on the other hand, of course, so that I definitely do not harm myself. I am not calling on anyone to prescribe or change treatment on their own on the basis of a chat. What matters a lot to me is that the model helps us make sense of information about our health, formulate questions and understand what needs to be checked before our next step.
And we can take a much more active part in our own health: prepare better, ask questions, understand documents, find discrepancies. And then we will always get a very important, meaningful and useful result. I am super interested in how this works for you. Have you used artificial intelligence to prepare for an appointment, to make sense of reports, perhaps to search for research? Did it help you formulate a question? Did you solve your problem? It seems to me that everyone's examples will help others use the opportunities people have today much more competently and consciously.
It is extremely important to share this experience now. Thank you all, and see you in the next episode.