The ChatGPT interface showing examples, capabilities, and limitations on a dark blue screen

AI Confabulation: Why Chatbots Invent Facts With a Straight Face

A teacher in Manila asked an AI chatbot to generate a reading passage about the history of tea, complete with a source citation for her intermediate class. The passage was smooth, well-organized, and completely readable. It was also wrong. The chatbot attributed a quote to a historian who does not exist, cited a book with a real-sounding title that was never published, and got the century wrong for when tea drinking spread to Europe. Nothing about the output looked suspicious. That is exactly the problem, and it is a problem every ESL teacher using AI tools needs to understand at a mechanical level, not just a cautionary one.

This kind of error has a name: hallucination, or confabulation. It is not a bug that gets patched out with the next update. It is a structural feature of how large language models work, and understanding why it happens changes how you use these tools in lesson prep, and even opens up a genuinely useful classroom activity for teaching students to think critically about AI output.

What a Language Model Is Actually Doing

It helps to drop the word “knows” from how you think about AI chatbots. A language model does not know that Paris is the capital of France in the way a person knows it, with a memory of a fact anchored to a source. Instead, the model has been trained on enormous amounts of text and has learned statistical patterns: given this sequence of words, what word is most likely to come next. When you ask it a question, it is not retrieving an answer from a filing cabinet. It is generating the most probable continuation of your prompt, one word at a time, based on patterns absorbed during training.

Predicting Words, Not Verifying Facts

This distinction matters because most of the time, the statistically likely continuation of a sentence also happens to be true. If you ask a model to finish “The chemical symbol for gold is,” the correct answer, “Au,” appears so often in the training data that it dominates the probability distribution. The model looks accurate because accuracy and probability overlap constantly in everyday facts. The trouble starts when they don’t overlap, usually in three situations: when the training data contained little or no information on a topic, when a question asks for something oddly specific like a page number or an exact quote, or when a prompt nudges the model toward a shape of answer, such as “give me a source,” that it will fill in even if no real source fits.

In all three cases, the model does not pause and say “I don’t have this.” It keeps generating the next most plausible word, because that is the only thing it was built to do. The result is text that is grammatically perfect, formatted like a citation, and entirely fictional.

Why There Is No “I Don’t Know” Button

Teachers who are new to AI tools often assume a chatbot would say “I’m not sure” when it lacks information, the way a careful student would. It usually doesn’t, and the reason is worth sitting with. During training, models are rewarded for producing fluent, confident, complete-sounding answers. Hedging, refusing, or admitting uncertainty is rare in the training data compared to confident assertions, because most text on the internet is written by people who sound sure of themselves, correctly or not. The model has learned the tone of certainty far more thoroughly than it has learned the boundaries of its own knowledge.

Red error messages on a black computer screen indicating blocked web resources
Red error messages on a black computer screen indicating blocked web resources

Confidence Is a Writing Style, Not a Signal of Accuracy

This is the single most important idea for a teacher to internalize: the confidence of an AI-generated sentence tells you nothing about whether it is true. A model will describe a fabricated grammar rule with the same even, authoritative tone it uses to describe a real one. There is no built-in volume knob that gets quieter when the model is guessing. That flat confidence is what makes hallucinations dangerous in an educational context specifically, because teaching materials are supposed to model correctness for learners who don’t yet have the background knowledge to catch the error themselves.

Where This Shows Up in ESL Materials Specifically

Hallucination is not evenly distributed across tasks. Some ESL use cases are low-risk, and others are practically an invitation for the model to make something up. Generating a simple dialogue for practicing past tense, for example, carries almost no factual risk, because there is no fact to get wrong. Generating a reading passage about a real historical event, a biography of a public figure, or a scientific fact for a content-based lesson is a different category entirely.

Invented Sources and Attributions

Ask an AI tool to cite a source and you have handed it a template it must fill in, whether or not a real source exists. This is one of the most consistent failure points across every major model, and it is exactly the feature ESL teachers reach for when building reading comprehension passages that feel credible. If a chatbot gives you an author name, a publication, and a year, treat that citation as a claim to verify, not a fact to trust, until you have checked it against a real search.

Happy teacher attractive matude adult is smiling using laptop in class typing working with chalkboard in background. People a
Happy teacher attractive matude adult is smiling using laptop in class typing working with chalkboard in background. People a

Grammar Explanations That Sound Rigorous but Aren’t

Grammar is trickier than most teachers expect, because English grammar rules have genuine regional variation, exceptions, and disputed edge cases, so a model can produce an explanation that sounds textbook-perfect while quietly misapplying a rule to the wrong verb category, or inventing a “rule” that is really just a tendency in a subset of usage it was trained on. This matters most for exam-prep content like TOEIC or IELTS grammar drills, where a single wrong explanation can cost a student marks on test day. Cross-check any AI-generated grammar rule against a reference grammar you trust before it reaches a handout.

Vocabulary Definitions That Blend Real and Invented Usage

Ask for five example sentences using a target word and a model will sometimes produce one or two that use the word in a sense that doesn’t actually exist, especially with idioms, phrasal verbs, or words that have a narrow, specific collocational pattern. Because the sentence structure is fluent, the error is easy to miss on a first read, which is exactly why a second, deliberate check matters more than a quick skim.

Turning the Flaw Into a Lesson

Once you understand hallucination mechanically, it stops being purely a hazard and becomes a genuinely useful teaching tool, particularly for upper-intermediate and advanced students who are being asked, more and more, to use AI tools themselves for homework and research. A media literacy activity built around AI fabrication does double duty: it builds critical reading skills and it demystifies a technology students are already using, often uncritically.

person using laptop
person using laptop

A Spot-the-Fabrication Activity

Generate a short passage on a topic your class has some background knowledge in, and deliberately prompt the AI tool for a citation, a statistic, or a quote. Print the passage without telling students which parts might be invented. In pairs, have students identify any claim they find suspicious and design a verification plan: what would they search for, which kind of source would confirm or deny it, and how would they phrase the search in English. This works as a listening-into-writing task too: have students discuss their suspicions aloud in English before writing a short verification report.

  • Give students the AI-generated passage with no warning about accuracy.
  • Ask them to underline any claim that could be a fact, name, date, or statistic.
  • Have them draft a search query in English for each underlined claim.
  • Compare findings as a class and discuss what made some fabrications easier to catch than others.

Why This Fits Exam-Prep Classes Especially Well

IELTS and TOEIC both test the ability to evaluate a written passage critically, distinguish fact from inference, and identify the writer’s purpose. A hallucination-spotting exercise is, structurally, the same skill dressed in a more current context. Students who practice interrogating an AI-generated paragraph for unsupported claims are practicing exactly the skimming-for-evidence skill that IELTS Reading Task 2 rewards.

Building a Verification Habit Before Anything Reaches a Student

The practical takeaway for lesson prep is simpler than it sounds: treat any factual claim from an AI tool the way you would treat a claim from an anonymous online comment, not the way you would treat a claim from a textbook. That means a quick habit, applied consistently, catches almost everything that matters.

University Library of Trnava University, books, university, study, bookcase, library, a lot of learning, education
University Library of Trnava University, books, university, study, bookcase, library, a lot of learning, education

Before a passage, worksheet, or grammar note generated with AI goes anywhere near a class, run it through three checks. First, isolate every specific factual claim: names, dates, numbers, quotes, and titles. Second, search for each one independently rather than asking the same AI tool to check its own work, since a model asked “are you sure?” will often just generate a confident-sounding confirmation of its own error rather than genuinely re-verifying it. Third, if a claim can’t be confirmed within a couple of minutes of searching, cut it or rewrite the passage to avoid needing it at all. A reading passage about tea history is just as useful for practicing past tense and topic vocabulary whether or not it includes a specific, checkable citation.

stack of papers flat lay photography
stack of papers flat lay photography

When the Stakes Are Low, Loosen Up

Not every use of AI needs this level of scrutiny. Dialogues, gap-fill sentences built around a target structure, roleplay scenarios, and warm-up discussion questions carry little factual risk because there is nothing to fact-check. Reserve the verification habit for the moments where a model is being asked to supply information rather than language practice, and you keep the workflow fast without sacrificing accuracy where it counts.

students in classroom with teacher presenting
students in classroom with teacher presenting

The Underlying Point Worth Remembering

AI hallucination is not a sign that a chatbot is broken or that a particular tool is worse than its competitors. Every large language model built on the current generation of technology does this, to varying degrees, because it is a direct consequence of how the technology generates text: by predicting plausible continuations, not by retrieving verified facts. Knowing that changes how a teacher uses the tool. It stops being a search engine you trust and becomes a fast first-draft generator whose factual claims you check the same way you would check a claim from a student who is very fluent but has not yet done the reading. That mental model, applied consistently, is what lets ESL teachers use AI tools for real time savings in lesson prep without quietly passing fabricated facts on to a classroom full of learners who have no way to know better.

Fontes

IBM — background on how large language models generate text and why factual errors occur.
British Council — britishcouncil.org — resources on English language teaching methodology.
Wikipedia: Hallucination (artificial intelligence) — overview of the phenomenon and terminology.

Postagens semelhantes