An AI can explain a concept perfectly and still have no idea how to use it.
By claudiu-clement · September 17, 2026 · Curated by George's Blog
An AI can explain a concept perfectly and still have no idea how to use it.
Researchers from MIT, Harvard and Chicago call this "Potemkin Understanding" - named after the fake villages built to impress an empress from a distance. Impressive facade, nothing behind it.
Here's the experiment:
They asked top models to define concepts from game theory, literature and psychology. The models got the definitions right over 90% of the time.
Then they asked the same models to apply those concepts - spot a real example, generate one, edit one.
Performance collapsed. Depending on the task, models failed roughly 40–55% of the time on concepts they had just defined flawlessly.
Worse: when a model generated its own example, it often couldn't correctly classify what it had just written. Not wrong but incoherent.
Why this matters: we grade AI with human tests (the bar exam, AP exams, MMLU). Those tests work on humans because humans misunderstand things in predictable ways. If you pass, you probably get it.
AI doesn't fail like a human. It fails like an illusion.
So the next time a benchmark score tells you a model "understands" your domain, ask the only question that counts: can it actually do the work?