
54905028_artificial-intelligence-for-learning
by Donald Clark
Decades of learning science proved that personalized tutoring outperforms classroom instruction by two standard deviations—AI is the first technology that can…
In Brief
Decades of learning science proved that personalized tutoring outperforms classroom instruction by two standard deviations—AI is the first technology that can deliver that gap at scale. Learn to redesign your organization's learning stack around dialogue, adaptive feedback, and effortful retrieval instead of the click-through substitutes that never worked.
Key Ideas
Effortful retrieval surpasses passive selection
Replace multiple-choice questions with open-input questions assessed by NLP — effortful retrieval (typing a free answer into a blank box) produces dramatically better long-term retention than clicking through answer lists, and the technology now exists to implement it at scale
Integrate AI into full workflow
Apply the 'cyborg vs centaur' test to your team's AI use: professionals who integrate AI into their full workflow consistently outperform those who use it as a bolt-on tool — redesign the process, not just the tool access
Personalization closes the tutoring gap
Audit your e-learning against Bloom's 2 sigma: if it doesn't personalise sequence, difficulty, and feedback to the individual, it is delivering classroom-lecture outcomes (the baseline), not tutor outcomes (a 98% improvement in mastery) — and the gap between those is the gap AI closes
Deploy chatbots for support first
Use chatbots for learner support functions first — answering questions, formative feedback, onboarding, performance support — before attempting full instructional replacement. Jill Watson was built on 40,000 real student queries; start where the volume and repetition already exist
Migrate to xAPI from SCORM
Treat SCORM as the structural bottleneck it is: without migrating to xAPI and a Learning Record Store, your adaptive and personalised learning ambitions have no data foundation — SCORM was designed to track course completion, not to drive learning
Motorcar test evaluates ethical concerns
When evaluating AI ethics concerns, apply the motorcar test: would you apply the same harm-calculus to an equivalent technology already embedded in society? If the standard shifts, the concern is likely emotional rather than analytical — and the structural problems (primitive assessment design, biased human teaching) deserve that energy more
Who Should Read This
Readers interested in Artificial Intelligence and Learning, looking for practical insights they can apply to their own lives.
Artificial Intelligence for Learning: How to use AI to Support Employee Development
By Donald Clark
11 min read
Why does it matter? Because the problem with AI in education isn't what AI might do — it's what education already did.
Everyone is debating what AI will do to learning. That is the wrong question. Ask instead what learning has been doing to learners for a century — and why nobody called it a scandal. The forgotten compliance course. The click-through quiz you passed without reading. The lecture that delivered the same words to thirty different minds and called that teaching. Learning science identified the problem decades ago: mastery needs dialogue and personalisation. What we built instead was logistics dressed as pedagogy. AI did not arrive to break education. It arrived to deliver what education always promised but structurally could not. That is a different argument. This book makes it with evidence.
Before AI Could Break Education, Education Had Already Broken Itself
The debate about whether AI will destroy education assumes something worth destroying. It doesn't.
Clark spent decades watching the learning industry pour money into systems designed to produce the sensation of learning rather than learning itself. Despite the investment in platforms and gamified courses, the result was a system still running on behaviorist psychology that hadn't changed in decades. Scores, badges, and leaderboards — the tools meant to motivate — were conditioning, not instruction.
Memory researcher Robert Bjork had documented the problem as far back as 1994: there is a negative correlation between how much learners believe they've understood and how much they've actually retained. The easiest learning experience is the most deceptive. You read a passage, click through multiple-choice questions, collect your completion certificate. The progress bar moves. The correct answer glows green. You mistake the sensation for understanding. The material is gone within days.
What made this structural rather than accidental was SCORM, the data standard that governed online learning systems for over two decades. SCORM tracked whether you completed a course, not whether you understood it. The industry's own infrastructure optimized for compliance, not cognition. The data that might have revealed how little was retained was never collected. The standard wasn't built to ask.
So when people worry that AI will damage a working educational system, it's worth pausing on what that system actually was: click-through modules built on conditioning principles, tracked by a standard that measured seat-time, delivering the feeling of competence to people who would forget most of it within a week. AI didn't arrive to a functioning ecosystem. It arrived to one that learning science had been quietly indicting for thirty years.
Every Great Learning Theory Promised Dialogue. Every Technology Before Now Delivered a Monologue.
Gordon Pask — cape, Edwardian suit, bow-tie, pipe — built his first learning machine in 1956 in circumstances that make today's EdTech feel extravagant. SAKI (Self-Adaptive Keyboard Instructor) was a device to teach punch-card operators: electric wires cued the learner's fingers, detected keypresses and timing, and constantly adjusted the pace, always pushing ahead but never so far that the learner lost grip of the task. It was still in use in the UK Post Office in the late 1960s.
What Pask understood (and what made him so useful and so overlooked) was that SAKI wasn't just a training tool. It was a demonstration of a theory. He called it conversation theory, and he used the word deliberately. Not metaphor: conversation as the actual mechanism by which learning happens. We learn by interacting with people, with artifacts, with machines, sending signals and receiving responses calibrated to where we are right now. A textbook cannot do this. Neither can a lecture, or a pre-recorded video. They deliver in one direction and trust the learner to close the loop themselves. Pask called almost everything that followed SAKI "hopelessly primitive" for exactly this reason: each new medium could carry content toward a learner but had no way to respond to that particular learner. It broadcast. It didn't teach.
That verdict was confirmed 28 years later in a form clear enough that it should have changed everything.
Benjamin Bloom's 1984 paper presented what became known as the "2 sigma problem." He compared three teaching conditions: a standard lecture, a lecture with consistent formative feedback, and one-to-one tutoring. The formative feedback approach produced an 84% improvement in mastery over the baseline lecture. One-to-one tutoring produced 98%. Nearly double the mastery, simply from having a knowledgeable person respond to you: catching your misunderstandings before they calcified, adjusting the explanation, pressing you harder when you were ready, easing off when you weren't. Vygotsky had named this mechanism decades earlier. Learning, he argued, requires a "knowledgeable other," a more capable presence calibrated to exactly where you are, not where the curriculum assumes you should be. Bloom's numbers gave that theory a figure.
Bloom called it a "problem" because the result was essentially unusable. One-to-one tutoring at scale requires one expert for every learner. You could prove it worked. You couldn't build it for anyone who couldn't afford a private tutor.
That finding sat in the literature for four decades. Every learning platform built in the interim (every LMS, every e-learning module, every MOOC) knew the result and couldn't do anything about it. The architecture of scalable education was, by necessity, the lecture: one voice, one pace, one direction. Pask's "hopelessly primitive" verdict kept being earned.
What Clark argues, and what the evidence supports, is that large language models are the first technology capable of closing this gap. They respond at the level of each learner's actual words, confusions, and pace. For the first time, the condition that produces 98% improvement in mastery can run for everyone simultaneously, at effectively no marginal cost per additional learner.
The science settled in 1984. The machine to act on it took forty years to arrive.
The Chatbot Was Too Good: What Blew Jill Watson's Cover Tells You Everything
Spring semester, 2016. Professor Ashok Goel was running a knowledge-based AI course at Georgia Tech with 350 students generating around 10,000 queries per semester — a full year of work for a human teaching assistant. He trained a chatbot named Jill Watson on four semesters of data: 40,000 questions and answers from previous classes. Early results were disastrous. Wrong answers. Bizarre answers. But with iteration, something shifted. Accuracy hit 97%. Goel launched Jill quietly, answering student queries alongside the human TAs.
Three months passed. Nobody suspected anything.
What finally gave her away wasn't a garbled response or an odd non-sequitur. It was the opposite. One student noticed she was too fast — even with an artificial time delay Goel had built in specifically to make her seem more human. Someone checked LinkedIn and found a real Jill Watson, who was baffled by the sudden attention. When the truth came out, the reaction was almost entirely positive. The class wanted to nominate her for a teaching award. Goel submitted her.
The failure ran in the wrong direction. Every reasonable assumption about AI in education says the problem would be inauthenticity: wooden phrasing, missed context, responses that feel generated. The actual problem was that Jill was more reliable than her human counterparts. She never got tetchy. She never gave a slightly impatient answer at 11pm on a Tuesday. She never had a bad week. Real teaching assistants, under the weight of identical questions arriving in waves, occasionally let frustration show. Jill didn't have a frustration mechanism.
Byron Reeves and Clifford Nass spent the 1990s documenting something relevant here. Across 35 psychological studies, they found that people treat computers as social actors: polite to them, responsive to flattery, stung by harsh feedback. Not because they consciously believe the computer is human. The mind projects intention automatically, without deciding to. This means a well-designed AI tutor doesn't need to pass some rigorous Turing test to produce human-equivalent emotional engagement. It needs to not break the spell.
The following semesters, Goel introduced more bots: Ian, then Stacey, who was deliberately more conversational. Students now knew to look for AI. They were actively watching. Even so, only half identified Stacey and fewer than one in five identified Ian.
"Students will see straight through it" is the standard objection to AI tutors. It assumes the failure mode is a machine pretending to be human and failing. What Georgia Tech showed is that the failure mode runs the other way. The machine succeeds, and the thing that gives it away is that it performs the job better than people do.
The Ethical Panics About AI in Learning Are Almost Perfectly Inverted
A technology grants astonishing freedom to billions. It also kills 1.4 million people every year: grisly, world-war-scale casualties, year after year without end. Would you accept it? Nobody does. Clark then names it: the motorcar.
We are utilitarians when we actually adopt technology and rule-bound moralists when we theorize about it. Consistently applied, the standard we demand of AI would prohibit almost all of modernity.
In education, the inversion runs deeper. The dominant fear is plagiarism — students generating essays they didn't write. Clark's diagnosis is that the scandal runs in the opposite direction. The essay persists as the primary assessment form because it's easy to assign. Academics set the same questions for years. Feedback arrives sparse and late. Students who've paid significant fees, who face genuine loss of face if they fail, treat the assignment as a game with rules to be gamed. The structure then protects itself: institutions suppress the scale of cheating to protect their reputation, academics avoid pursuing cases because the evidence requirements are exhausting, and essay mills supply undetectable text on demand. Everyone has agreed to let the arrangement continue.
The job-replacement panic follows the same pattern. The roles most exposed — grading papers, delivering standard lectures, administering tests — were already performing those tasks badly. A human grader marking a class of thirty gives each student perhaps thirty seconds of attention. The jobs at risk were already doing marginal work.
The bias objection inverts even more sharply. Kahneman and Tversky spent careers documenting that human cognition runs on deep, largely irremovable biases: socioeconomic, racial, gendered. Teaching is soaked in them. AI carries bias too, inherited from the curriculum humans designed. But we've never held the human teacher to an equivalent standard, and Clark's challenge is simple: how is that going?
What AI did was expose how hollow the arrangement always was. Education researchers Paul Black and Dylan Wiliam published research in 1998 recommending that grading be eliminated entirely and replaced with feedback alone: grades cause high performers to stop at 80% and leave everyone else demoralized. That finding sat in the literature for two decades while institutions kept issuing grades and setting essays unchanged. The essay was testing whether a student could reproduce the shape of critical thought under examination conditions. AI reproduces that shape in seconds. The structure could absorb decades of essay mills, could absorb Wikipedia, could absorb almost anything. It couldn't absorb a tool that made the hollow center visible to everyone at once.
The Productivity Numbers Are Not Marginal — They Are Structural
The assumption is that AI makes knowledge workers a bit faster at the routine parts of their job. The evidence from controlled trials says something else — faster and better at the same time, by margins that don't fit the word "incremental."
The sharpest evidence came from a 2023 trial at Boston Consulting Group. Researchers gave 758 consultants access to GPT-4 across eighteen tasks, then measured results against a control group working without it. The GPT group completed 12% more tasks, finished 25% faster, and produced output rated 40% higher in quality. Faster and better at once — that combination should stop you. It violates the usual tradeoff. When you rush, quality drops. These consultants rushed and got better.
But the finding within the finding mattered most. Not all consultants improved equally. The researchers found a split in how people used the tool. Some used it as an add-on: hand a section to the AI, get a polished draft back, move on. Others rebuilt their entire workflow around it, treating AI not as an assistant for specific tasks but as the medium in which thinking happened. The researchers named the first group "centaurs": human body, AI legs, moving faster. The second group they called "cyborgs": genuinely integrated, not augmented. The cyborgs consistently outperformed the centaurs.
The centaur-cyborg distinction changes what "adopting AI" means in practice. A centaur has added a tool. A cyborg has changed how they work. Most organizations rolling out AI are producing centaurs, training people to use the tool on top of their existing process. The BCG data says this captures only a fraction of the available gain. The productivity ceiling isn't the AI. It's the workflow the AI was bolted onto.
The 30-Year L&D Paradigm Is Over. Here Is What Must Replace It.
What does it mean to adopt AI in learning and development? If your answer involves implementing a new platform, running a pilot program, or training your instructional designers to use a chatbot, you're describing a centaur — a human process with AI legs bolted on. Clark's diagnosis is harder than that.
Clark lays out what the transformation actually requires. None of the shifts involves adding a new tool; each describes abandoning something. The skills that defined L&D for thirty years (MCQ craft, linear module design, media production) become liabilities rather than assets.
The MCQ-to-open-input shift is where the depth of that change becomes visible. For thirty years, the craft of online learning design concentrated on writing good multiple-choice questions, a technically demanding task that produced, at best, a learner clicking through lists of options and feeling the satisfying highlight of a correct answer. Bjork's research showed this generates the sensation of learning while almost guaranteeing the learner will forget the material. The easier the task, the more convincing the illusion.
Open-input design, enabled by NLP, requires the learner to type a free-text answer into a blank box. No options to recognize. No elimination process. The AI performs semantic analysis, accepts synonyms and word-order variations, marks what was correct — but the system cannot be gamed by pattern recognition. From Arthur Gates, who in 1917 showed that active recall outperforms passive rereading, through Brown's Make It Stick in 2014, the finding is consistent: pulling information from memory is a stronger consolidator than the original learning event. The MCQ actively blocked this: selecting from a list is recognition, not retrieval. The designer who spent years perfecting MCQ craft now needs to build for the cognitive mechanism the MCQ suppressed.
Organizations treating AI as a tool upgrade will skip this transition entirely. The underlying job (designing for retrieval, building adaptive pathways, curating rather than creating) differs from the job that preceded it. The skills transfer; the job description does not. You cannot centaur your way there.
The question is no longer whether your organization will adopt AI in learning. It's whether you're willing to dismantle what you built over the last three decades fast enough to matter.
This Is the Worst AI Will Ever Be
Bloom found this in 1984: one-on-one tutoring moved students two standard deviations above their classroom peers. The average tutored student outperformed 98% of students taught the conventional way. Everyone in learning science knew. The problem was never the evidence — it was the arithmetic. One expert per learner. Not scalable at any price.
That constraint held for forty years. It is gone now.
What Jill Watson handled in one semester — 10,000 student queries, with students unaware they were talking to software — is not an anecdote. It is a proof of concept for something with no historical precedent: the thing that required one expert per learner can now run for a hundred million learners simultaneously, at near-zero marginal cost, improving as it goes. We are at day one of that curve.
Every argument about what AI cannot yet do in education is an argument about a moving target. The target moves in one direction only.
The deeper question is whether the model being disrupted was ever right. The lecture, the course, the completion certificate — learning science said no in 1984. It kept saying so while the industry built SCORM and called it progress. AI didn't arrive to overturn something working. It arrived to make possible what the evidence always demanded.
The question is no longer whether to use it. It is how fast you can update what you believe learning actually is.
Frequently Asked Questions
- What's the difference between open-input and multiple-choice questions in AI learning systems?
- Open-input questions assessed by natural language processing dramatically outperform multiple-choice formats for long-term retention. The book explains that "effortful retrieval (typing a free answer into a blank box) produces dramatically better long-term retention than clicking through answer lists," and modern NLP technology now makes this scalable. Free-response assessment requires learners to actively generate knowledge rather than recognize correct options, triggering deeper cognitive processing. This shift leverages cognitive science principles that multiple-choice testing cannot achieve, making it a foundational change for effective AI-powered learning systems.
- What does the 'cyborg vs centaur' test mean for AI implementation?
- The cyborg-versus-centaur distinction evaluates whether AI is truly integrated into workflows or merely added as an afterthought. According to Clark, "professionals who integrate AI into their full workflow consistently outperform those who use it as a bolt-on tool — redesign the process, not just the tool access." A cyborg seamlessly fuses human and machine capabilities within redesigned workflows, while a centaur simply bolts AI onto existing processes. The test identifies that genuine competitive advantage comes from fundamental workflow redesign around AI capabilities, not superficial tool adoption.
- How does Bloom's 2 sigma relate to AI in learning?
- Bloom's 2 sigma benchmark represents a 98 percent improvement in learning mastery when personalized tutoring replaces classroom instruction. Clark argues that "if it doesn't personalise sequence, difficulty, and feedback to the individual, it is delivering classroom-lecture outcomes (the baseline), not tutor outcomes (a 98% improvement in mastery) — and the gap between those is the gap AI closes." This metric provides a clear performance standard: AI learning systems must personalize difficulty progression, content sequencing, and immediate feedback to each learner. Without personalization, e-learning remains equivalent to traditional classroom lectures.
- Should you start with chatbots or full AI replacement in learning systems?
- Start with chatbots for learner support functions rather than attempting complete instructional replacement immediately. Clark recommends: "Use chatbots for learner support functions first — answering questions, formative feedback, onboarding, performance support — before attempting full instructional replacement. Jill Watson was built on 40,000 real student queries; start where the volume and repetition already exist." This incremental approach builds capability gradually while establishing proven use cases. Chatbots excel at repetitive, high-volume support tasks before expanding into instructional design, creating a strong foundation with performance data.
Read the full summary of 54905028_artificial-intelligence-for-learning on InShort


