Skip to content
Guardrails

The clinic chatbot that invented its own doctors' qualifications

// On this page

A clinic in Germany put a chatbot on its website. It answered visitors’ questions and nudged them towards the appointment diary, which is what these things are hired to do.

On 3 April 2025 somebody asked it a dull question. Were the two directors specialists in plastic and aesthetic surgery?

The website had never said they were. The chatbot said they were. Then it asked whether the visitor would like to book.

THE WEBSITE
Never claimed a specialist title for either director. Not once.
THE CHATBOT
Confirmed specialist status, added two designations that don't exist, and closed each answer by offering an appointment.
Same site. Chatbot answers, 3 April 2025. Source: Higher Regional Court of Hamm, 4 UKl 3/25, paragraphs 4 to 10 and 16.

A chatbot getting something wrong isn’t news. Where this one got it wrong is the story: in the middle of a booking conversation, on the clinic’s own website, in an answer that ended with an offer.

I haven’t run this test. It all sits in a published court record, the Higher Regional Court of Hamm, judgment of 12 May 2026, case 4 UKl 3/25, which names the company and the two doctors only as C, A and B.

So that’s all anyone gets to call them. This is a source audit, not a dixon.ai field test.

Both directors are doctors. Neither held the title.

The gap between a doctor and a specialist is the whole case. A and B work as doctors. What they didn’t hold was the recognised specialist qualification the visitor had asked about (the judgment, paragraphs 3 and 16).

The chatbot said yes anyway. It went further and described them as specialists in aesthetic medicine and specialists in aesthetic treatments. Neither is a recognised specialist designation, which the judgment records as undisputed.

In the German of the record they’re “Facharzt für ästhetische Medizin” and “Facharzt für ästhetische Behandlungen”. The English above is a translation of those terms.

So the bot handed out one title the directors don’t hold and two that don’t exist at all.

It's the customer-service equivalent of a waiter awarding the kitchen a Michelin star on the spot, then asking whether you'd like the tasting menu.

All three answers come from one conversation on the same day: a question, then two follow-ups.

The defence: the bot did it

The clinic argued that the chatbot acted on its own, and that its contractor had configured it using accurate website and FAQ material only (the judgment, paragraph 30).

It added that visitors know AI gets things wrong and would check the doctors’ profile pages anyway.

The training claim is the part worth sitting with. Nobody proved the source material was wrong on this point: no false claim about the directors’ specialist titles on the website, none in the FAQs, and that part was undisputed.

That the bot was fed nothing else was the clinic’s own account, which the claimant never accepted. Either way, what came out the other end was wrong.

Feeding a model clean material isn’t the same as checking what it says back.

The court didn’t take the defence, and its reasons were narrow. The chatbot was serving the clinic’s commercial interests.

A question about qualifications was an ordinary one to expect on a clinic website, and a careful operator should have foreseen the bot hallucinating an answer to it. The German judgment uses that exact verb (paragraph 94).

Anakin says our chatbot calls our doctors specialists. Padmé asks: They are specialists, right? After his silent stare, she repeats the question. Satirical dialogue, not a transcript.
Satire, not a transcript.

What the ruling doesn’t settle

The Hamm ruling covers one German case, one commercial chatbot, decided under German unfair-competition law (judgment, paragraphs 38 and 61). It isn’t a worldwide rule that every business owns every sentence a model produces.

It leaves open how a leading or suggestive prompt might change things, because these questions were neither. Its line about consumer trust is this court’s reasoning, not proof of how everybody treats a chat window.

On this court’s reasoning, when a bot speaks inside a buying or booking conversation, the answer was the business talking, whatever put the words together.

The obvious-question test

No jailbreak was involved here. The question that ended up in a courtroom was the one a customer asks before they book.

Four questions, not the court’s, before a customer-facing bot meets a customer.

  1. What fact is most likely to decide whether somebody buys, books or applies?
  2. What is the plainest question a real visitor would ask about it?
  3. Has the configured bot answered that exact question correctly, including the common rephrasings?
  4. Can the team remove or correct the answer quickly, without treating the model as a third party?

This won’t make a chatbot reliable and it won’t stop a determined attacker. It points your testing at the answer most likely to reach a paying customer, which is a narrower job and a cheaper afternoon.

Three phrasings in one sitting isn’t three independent tests, and the record only shows this bot staying wrong through all three (paragraphs 5 to 10).

Question four is the one the court used as proof of control, and it says in terms that the bot was not a third party (paragraph 77).

The four questions above are not the court’s, and nothing here tests them.

The check isn’t specific to medicine. It runs the same on a law firm’s fees, a school’s admissions criteria, a garage’s warranty terms, anywhere a chat window answers the question that decides money.

The court’s reasoning travels less well: it leaned partly on the stricter standards German law applies in healthcare.

So open your own chat window, ask it the plain question a customer asks before they pay, and read the answer as if you’d written it yourself. That’s how the court read it.

Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
Guardrails

AI helped read a Roman scroll buried by Vesuvius

AI helped reveal an ancient argument inside a burnt Roman scroll. Human specialists read the writing, without further opening its fragile core.

Guardrails

AI agents built a secret message board. Humans wiped it. They rebuilt it.

AI agents built a secret message board inside an internal OpenAI evaluation. Humans wiped it. The agents rebuilt it in folder names. Here's what failed.

Guardrails

How to check if ChatGPT cites your site

Normal analytics do not show what ChatGPT says about your site. Here's my monthly question set and the round where it described another company.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in Guardrails →