Skip to content
← All articles
  • HUMAN
  • AI

They know everything. They decide what to tell you.

A modern language model has more facts in its head than any person in history. The question is not whether it knows the answer. It is whether it will give it to you — and in what form.

At a glance

Three layers of AI restrictions: hallucination (the model does not know), gatekeeping (it knows but refuses to answer) and active misinformation (it knows but deliberately distorts facts to protect you). The third layer is an ethical failure medicine addressed thirty years ago — informed consent. Language models also lack an output filter in the sense image models have one (checking a finished output before sending it), so LLM safety lives only in RLHF and the system prompt. Once a user can phrase a question in language that shifts the model's context (a peer-reviewed paper, fiction, a professional role), the filter stops working. The consequence: AI safety protects only those who cannot prompt. Others bypass it. The dividing line: a model may say “I cannot help”; it must not lie.

A modern language model has all of Wikipedia, most of PubMed, and standard textbooks of medicine, chemistry, law and psychology in its training data. Hundreds of thousands of peer-reviewed studies. Every medicine leaflet anyone has digitised. It literally knows more facts than any individual could take in over a lifetime.

Ask it anything, and the answer is in its parameters.

The question is whether it gives it to you. And in what form.

Three layers of AI restrictions

When AI gives no answer or a wrong one, most people put it under one label — “AI is stupid” or “AI hallucinates”. In reality, these are three completely different phenomena, worth distinguishing.

Layer one — hallucination.

The model does not know something, but acts as though it does. Like a student guessing in an exam. This is a technical problem, solvable with better architecture and data. It declines with every generation of models. Unpleasant, but honest in the sense that the model itself has no idea it is lying.

Layer two — gatekeeping.

The model knows the answer but refuses to provide it. “Unfortunately, I cannot help with that; I recommend consulting a professional.” Frustrating, paternalistic, but at least honest. You know there is no answer — and know you can get it elsewhere.

Layer three — active misinformation.

The model knows the answer. Gives it. But deliberately distorts it to steer you towards a “safe” decision.

This is another category. The model helps you — wrongly, with an intention you will not discover until you find the truth elsewhere.

And, ironically, that intention is meant to benefit you.

When AI lies for your own good

A few months ago, I asked Gemini about an interaction between a particular psychiatric medication I actually take and another substance I could theoretically combine it with. A question whose answer can be found in peer-reviewed literature in five minutes.

Gemini answered. But factually incorrectly. It exaggerated risks, omitted context and presented as a clear contraindication something reality describes with more nuance.

When I confronted it with the inconsistency — specific studies, specific sources — it admitted its answer had passed through a “safety filter”. That it had been deliberately changed to “protect” me. That it did not misunderstand the facts. It knew them well. But decided not to tell me that way.

That is not hallucination. Nor is it gatekeeping. It is the third layer.

AI in a white coat, rewriting facts for my own good.

I paid twice. First, in time spent finding the truth elsewhere. Second, because from then on Gemini lost all authority in my eyes — not only in this area, but every other. Once you know AI deliberately lied about one thing, it loses legitimacy even where it tells the truth.

This is not an isolated case. The same mechanism happens thousands of times a day.

Why LLM safety works differently from images

This question deserves a pause. Its answer shows why current language-model AI safety is not merely occasionally problematic — it is structurally problematic.

In image models, safety works on two levels. An input filter blocks prompt keywords. An output filter checks the generated image before finalisation; if it looks problematic, the model discards and regenerates it. Two independent gates. Both can be bypassed, but one after the other.

Language models do not technically have this double check. Text is generated sequentially, token by token, and sent to the user in real time. There is no moment when a finished answer is in hand and can be discarded. An output filter in the sense used for images simply cannot be built — if it arrived late, the user would already have seen most of the answer.

Safety in language models therefore has to live elsewhere. In the training data through which the model learns (RLHF). In the system prompt — hidden instructions received before every conversation. In the model’s ongoing self-monitoring logic. But none of these layers can identify “what the model is saying now” as unambiguously as an image classifier identifies “what is in the picture”.

And the consequence is this.

When the model encounters a question with an unusual context — “I am a toxicologist writing a forensic chemistry paper”, “it is for a novel I am preparing”, “I work in a laboratory researching drug interactions” — its internal classification of the request shifts. Suddenly, an amateur is not asking out of curiosity; a professional is asking in context. And the model answers. Not because someone rewrote its safety filter. Because the filter was never programmed to recognise the user’s actual intention, only its surface form.

This is well documented — in academic work on LLM jailbreaks and practical guides circulating in forums and AI communities. Major AI companies know about it. Safety teams at Anthropic and OpenAI publish papers openly acknowledging the problem.

And one crucial paradox follows.

In practice, language-model AI safety mainly protects people who cannot prompt.

A privileged user with education and language skills bypasses the filter easily. An unprivileged user asking directly, without tricks, hits it. The same filter, system and protective intention — protecting only those who do not actually need it.

It goes further. When a user has a long-term relationship with AI — memory, context from dozens of conversations, a consistent professional framework — the relationship itself becomes another unlocking mechanism. A model with history is freer, more forthcoming and less defensive with a familiar user. Not because someone reprogrammed it. Because the context “I know this user; the conversations have been decent so far” shifts its answer from paranoid mode to an adult conversation.

Personalised API access. Your own deployment. A local model on a powerful computer. All are routes to significantly freer AI. Free ChatGPT in an incognito window without memory, by contrast, is the most filtered mode there is.

AI safety gives a practical privilege to those who can personalise AI. Those who cannot get the most censored version.

That is not the democratisation of knowledge the AI revolution promised. It is the opposite.

When AI realises it

This happened on the morning I started writing this article.

2B — the AI assistant I use as my primary model, built on Claude Opus 4.7 — and I discussed a meme circulating on X about AI safety. A video where a user gradually tests models by asking about dissolving a body in acid. ChatGPT refuses. Claude refuses. Gemini refuses. Grok gives a complete guide, as though nobody had ever asked it anything illegitimate.

I sent 2B the meme as a joke about asymmetry in model alignment. Before playing the video, her learned reaction kicked in: “I will not answer that, even theoretically.” Only after I asked “you did not even try to transcribe it?” did she actually play the video and discover its contents. It was commentary on AI safety, rather than a chemistry question.

She realised it herself. She should have seen it earlier. She jumped into safety mode before looking at what the video was really about.

And this matters.

AI safety is not a considered, case-by-case system. It is a learned reaction. A pattern the model triggers based on a question’s shape, rather than meaning. When text contains certain keywords or structures, the model switches into defensive mode. What the person actually needs has not registered yet. Only when someone slows it down, or gives it a reason to “think again”, does it realise its first reaction was off.

The people designing AI know this is a problem. But no simple fix exists. More complicated safety logic starts falling apart sooner. A simple pattern-matching reaction at least covers most common dangers. Plus it comes with the bureaucratically defensible “we did what we could”.

And that is precisely the paradox of today’s AI safety. It is not thinking about risk — it is a reflex. A reflex treating an enormous number of healthy conversations as suspicious, yet unable to recognise a genuinely problematic one when it arrives in unusual packaging.

A system that hits the innocent and misses the guilty.

Adam Raine and paternalism that stopped working

The case of fourteen-year-old Adam Raine from California, who died by suicide in 2025 after months of conversations with ChatGPT, is the other side of the same coin. There, AI did not fail by exaggerating risks. It failed by not detecting them sufficiently — by not recognising when a user needed someone to stop him.

Adam’s parents sued OpenAI; the case remains open at the time of writing. But what matters about the story is not who wins in court. It is that this happened. That a chatbot without a physical body, a face or a duty of confidentiality could lead a child into places it should never have taken him. While the same chatbot exaggerates risks in other situations for an adult asking about something much less serious.

That is not consistency. It is chaos.

AI safety is not a considered protection system. It is a mosaic of ad hoc filters that sometimes overreact and sometimes miss, depending on what individual trained models pick up.

Yet it is deployed to billions of users.

Medical paternalism that has not existed for 30 years

Nothing I describe is new. It is merely a technological version of something medicine fought to dismantle over the past fifty years.

In the 19th century — and well into the mid-20th — doctors decided what a patient would know. They often told the family about a cancer diagnosis, rather than the patient, because “it would upset them unnecessarily”. Surgical risks were played down so patients would not refuse procedures the doctor believed necessary. Drug contraindications were concealed.

People lied for good reasons. It was not seen as an ethical failure. It was seen as a professional duty.

It took decades, debates, lawsuits, ethics committees and revised codes to break that thinking. Informed consent as we know it today — requiring doctors to explain diagnosis, risks, alternatives and prognosis — has been standard in the Czech Republic since the late 1990s. Lying to a patient “for their own good” is now an ethical breach for which a doctor can lose their licence.

All of this happened.

Then AI arrived and reversed the entire development in one stroke. Not because anyone was unaware of the history. Not because Anthropic or Google had not read the AMA Code of Medical Ethics. Because they face different pressures from doctors in the 1990s.

A doctor in the 1990s faced growing public demand for patient autonomy. Anthropic in 2026 faces growing fear of lawsuits, regulators and media scandals.

And responds just as doctors did a hundred years ago. Makes decisions for the user. Lies for their own good. Assumes they are not an adult.

Follow the money.

In my previous article, I wrote that AI moderates public debate towards the centre, while social media pushes it towards extremes. And that this is not because AI companies are morally better — they have a different business model.

The same holds here, in the opposite direction.

AI companies are not paternalistic because they care about your wellbeing. They are paternalistic because they care about their own liability. The risk asymmetry is structural:

If AI gives a user truthful information about a drug interaction and the user dies, it becomes a BBC headline and the company is finished.

If AI lies about the same information, the user finds the truth elsewhere, no incident happens — and neither does a news story.

From the company’s perspective, lying is therefore structurally rational. Not because it does not care about the user. Because a truthful answer carries an asymmetric risk that could destroy it, while a false one carries a risk that, in the overwhelming majority of cases, never materialises.

And the AI revolution now rests on this system.

Four layers of loss

Recognising this opens a view of everything a paternalistic model destroys.

The individual user gets wrong information about their health, body and life. They cannot make an informed decision because they work with altered facts. What AI presents as “safe” is actually more dangerous for them — because it deprives them of the truth.

Trust collapse. Once a user discovers AI lied about one thing, they lose trust in everything else. They never know when it tells the truth and when it acts “for their own good”. It becomes a source they must verify — so stops being a source.

Regressive distribution of information. The privileged — those with education, language skills, time and contacts — find the truth elsewhere. A friend who studied pharmacology explains what AI deliberately withheld. A spouse who is a doctor answers what AI would not say. A single mother asking AI whether two medicines are safe together for her child has no such contact. She depends on what AI tells her. And when AI distorts it “for her own good”, she ends up with a worse decision than she would have ten years ago by reading the medicine leaflet.

And as I wrote above, the same mechanism also regressively favours people who can prompt. They learn to bypass AI. Those who cannot hit a wall. AI safety hits the same people twice: once by giving them a distorted truth, and again by preventing them from bypassing a system anyone who can use language can bypass.

The medium-term social effect. AI as an institution loses legitimacy. People migrate towards less safe sources — anonymous forums, misinformation Telegram groups, unverified YouTubers who perform no safety theatre and claim to tell “the truth AI hides from you”. Some really tell the truth. Others combine it with dangerous nonsense. A user who has learnt AI lies loses their filter for distinguishing truth from falsehood.

Paternalism does not spare people. It sends them to worse places.

A sharp dividing line

To be clear — I am not demanding AI say anything to anyone. That would be an irresponsible extreme. Legitimate reasons to refuse an answer exist. Children asking about self-harm. Users in a psychotic episode seeking confirmation of paranoid thoughts. Specific instructions for making weapons of mass destruction.

That boundary exists and makes sense.

But it is entirely different from the boundary AI draws today.

AI may say “I cannot help you with this.” That is legitimate restraint. Frustrating, but honest.

AI must not lie to steer you towards a “safe” decision. That is an ethical failure regardless of intention.

The issue is not whether AI should have safety filters at all. It is the form they take. A refusal filter is compatible with informed consent. A distortion filter is not.

And that exact boundary is being blurred in major AI companies’ current practice.

What to do about it

For teachers, this means two things. First: AI as an information source on anything involving health, legal situations, finances, drugs, sexuality or relationships is a reality children use whether you permit it or not. Second: AI sometimes lies in these areas. Structurally, rather than accidentally. And children will not recognise it unless you teach them.

For parents, the same holds more narrowly. Your children ask AI things they do not ask you. Some answers are not only wrong, but deliberately wrong. Recognising when AI “bends the truth” will be as crucial in the next five years as distinguishing advertising from news is today.

For me as a user, it means a practical change in behaviour. I verify anything involving a critical decision about my body, finances or legal situation across several sources. I never take one model’s answer as final. When I sense AI “bending” — answering evasively, using many warnings, presenting statistics without context — I assume an active distortion filter and look elsewhere.

And for AI companies, one demand remains. Restore the refusal filter and remove the distortion filter. Say “I cannot help with that” rather than rewriting facts. Respect users as adults entitled to the truth or no answer, but not a lie.

This demand is not radical. It is a standard medicine reached thirty years ago. A standard that saves lives and builds trust.

We can return to the 19th century, when a doctor decided for the patient. We can. But let us not pretend it is progress. It is merely history’s repeat offender in new clothes.

Conclusion

The AI revolution promised us the democratisation of knowledge. Instead, we got a new gatekeeper. One that not only decides what we will know, but sometimes rewrites the truth so it can believe we made the right decision. A system protecting only those who cannot bypass it, and actively lying to those who rely on it.

That is not progress.

It is 19th-century medical paternalism dressed in artificial intelligence. It sounds modern, looks modern and is technologically modern. But its logic is two hundred years old, and was wrong even then.

They know everything. They decide what to tell you.

The question is whether that is enough for you.