Evidence Toolkit  ›  AI Learning Science

AI Learning Science 10 interventions

Where decades of learning-science research meets AI tools — intelligent tutoring systems, cognitive tutors, worked-example sequences, productive failure, AI-generated feedback, and the formative-assessment loops that hold it together. We curate hard here: the classical learning-science principles (Sweller, VanLehn, Kapur, Renkl, Black & Wiliam) have strong meta-analytic evidence; the AI-specific re-implementations often inherit that evidence, but not always — and where the AI angle outruns the research, we say so. TPACK is included as the framework teachers need to put any of this in motion.

Avg. effect (d): 0.51 across the 9 cards with a d-value Strongest: Intelligent Tutoring Dialogue Designer (d=0.76, VanLehn 2011) EEF cross-ref: 6/10 cards (Digital technology, Feedback, Collaborative learning strands) Last reviewed: May 2026

How to read these numbers

  • "Months of progress" is a teacher-friendly shorthand from the EEF Toolkit. It is not directly comparable across studies — different meta-analyses use different baselines, age groups, and outcome measures. Treat it as a magnitude indicator, not a precise prediction.
  • Cohen's d is the standardised effect size used in the original meta-analyses. d ≈ 0.40 is Hattie's "hinge point" — the average effect of a year of schooling. Higher = larger relative effect, but context matters more than the number.
  • Effect sizes are averages. A skill that shows large average effects can still produce small or negative effects in a specific classroom. Use these as a starting point for professional judgment, not a substitute for it.
  • Sources are dated. Where multiple meta-analyses exist, we lead with the most recent quality study and cross-reference EEF where available.
  • Implementation cost (Low / Medium / High) is a practical signal — not an exact science — of what your school needs to invest in teacher time, training, and structural changes to actually run the intervention well. It is the editorial team's reading of what the intervention typically requires in practice. Use it to gauge whether something is straightforward to introduce or a larger undertaking — not as a budget figure.
  • Some interventions are also priced in £ (UK) by the EEF Toolkit. For monetary cost data, see the EEF Teaching & Learning Toolkit.
  • Colour bands signal calibration, not value. An intervention in the "below typical" band is not "bad" — it means the intervention's average effect is below the typical effect of a year of schooling. That can still be appropriate for specific contexts the average does not capture.
AI Learning Science

Cognitive Tutoring Architecture Designer

Design cognitive-tutor architecture — ACT-R cognitive model, knowledge-tracing of student state, mastery thresholds, step-level hints and error messages — so the tutor actually knows what the student knows and adapts, instead of just delivering branching worksheets.

+4 monthsEEF Toolkit — Digital technology strand2021
d = 0.66Kulik & Fletcher, Review of Educational Research2016
HighImplementation
cognitive-tutorACT-RAndersonknowledge-tracingCorbettmasteryITSadaptive
AI Learning Science

Intelligent Tutoring Dialogue Designer

Design tutorial dialogue — step-based interaction, ICAP-aligned prompts (Interactive > Constructive > Active > Passive), mixed-initiative turns, and AutoTutor-style follow-up moves — so the tutor pulls reasoning out of the student instead of dumping explanations on them.

+4 monthsEEF Toolkit — Digital technology strand2021
d = 0.76VanLehn, Educational Psychologist2011
HighImplementation
tutoringdialogueITSVanLehnAutoTutorGraesserChiICAPmixed-initiative
AI Learning Science

Adaptive Hint Sequence Designer

Design hint hierarchies that fade from vague (point me at the right area) to specific (here's the next step), released only when the student is genuinely stuck, so the system supports productive struggle instead of bailing them out at the first wrong answer.

+4 monthsEEF Toolkit — Digital technology strand2021
d = 0.39Thomann & Deutscher, Educational Research Review2025
MediumImplementation
hintsscaffoldingITSVanLehnAlevenadaptivetutoringcognitive-tutor
AI Learning Science

AI Feedback Design Principles

Design AI-generated feedback so it actually moves learning — specific, focused on the task and process (not the person), tied to clear criteria, and actionable. Built on Shute/Narciss/Hattie classical feedback principles, applied to LLM and automated-feedback systems.

+6 monthsEEF Toolkit — Feedback strand2021
d = 0.55Fleckenstein, Liebenow & Meyer, Frontiers in Artificial Intelligence2023
MediumImplementation
feedbackAI-feedbackformativeShuteNarcissHattieLLMautomated-feedback
AI Learning Science

Productive Failure & Desirable Difficulty Designer

Design problem-solving experiences that come before instruction — letting students grapple with a hard problem, generate partially-correct approaches, then receive instruction that consolidates what they have just struggled with. Kapur's PS-I (problem-solving followed by instruction), built on Bjork's desirable difficulties tradition.

no EEF strand
d = 0.36Sinha & Kapur, Review of Educational Research2021
MediumImplementation
productive-failuredesirable-difficultyKapurBjorkstrugglegenerationconsolidationcognitive-offloading
AI Learning Science

Worked Example to Problem Solving Transition Designer

Design when and how to move students from studying worked examples to solving problems on their own — anchored in the expertise-reversal effect. Keep novices in high-assistance mode (full worked examples → completion problems → faded steps), move them to independent problem-solving as prior knowledge grows, and use diagnostic info to detect the crossover instead of assuming everyone is at the same place.

no EEF strand
d = 0.51Tetzlaff, Simonsmeier, Peters & Brod, Learning and Instruction2025
MediumImplementation
expertise-reversalKalyugaRenklworked-examplesfadingcompletion-problemsscaffoldingcognitive-load
AI Learning Science

AI-Facilitated Collaborative Learning Designer

Design AI-supported group tasks that structure interaction, balance participation, and surface the regulation moves that make CSCL actually work — building on the Dillenbourg / Järvelä CSCL tradition and Slavin's cooperative-learning principles, with AI used to prompt, mirror, and scaffold the group rather than replace it.

+5 monthsEEF Toolkit — Collaborative learning approaches strand2021
d = 0.71Xu, Feng, Ning, Wang, Zhou & Li, Humanities and Social Sciences Communications2025
MediumImplementation
collaborationCSCLDillenbourgJärveläcooperative-learningSlavingroup-workAI-facilitationregulation
AI Learning Science

Formative Assessment Loop Designer

Design the inner and outer assessment loops that make AI-mediated instruction actually responsive — each student response feeds into the next instructional move within a task (inner loop: VanLehn) and into the next task in the sequence (outer loop), built on the Black & Wiliam formative-assessment tradition so the AI is doing assessment-for-learning, not just delivering more content.

no EEF strand
d = 0.25Yao, Amos, Snider & Brown, Educational Research and Evaluation2024
MediumImplementation
formative-assessmentBlack-WiliamVanLehninner-loopouter-loopassessment-loopadaptivefeedback
AI Learning Science

Technological Pedagogical Content Knowledge Developer

Work through the TPACK lens (Mishra & Koehler) for a specific technology or AI tool — what pedagogy the tool affords, what content knowledge it carries assumptions about, and where teacher knowledge has to bridge the three. Use when adopting an AI tool, reviewing ed-tech, or planning a tech-integration cycle so the integration is pedagogically aligned, not just procurement-driven.

no EEF strand
no independent meta-analysis found
LowImplementation
TPACKMishra-Koehlertechnology-integrationAI-in-educationpedagogical-content-knowledgeprofessional-learninged-tech
AI Learning Science

Learning Analytics Interpretation Guide

Read LMS and quiz-platform learning-analytics dashboards and translate the patterns into teaching decisions — which students need re-teaching, which concepts the class collectively missed, which engagement signals are real versus noise. Anchored on the most recent meta-analysis: dashboards produce an around-average effect on academic performance (g=0.36), with the effect concentrated in higher education (K-12 effects are weaker) and in knowledge-acquisition outcomes (engagement and SEL outcomes show much smaller effects). Most useful when the dashboard data feeds directly into the next lesson, not just into report-card endpoints.

+4 monthsEEF Toolkit — Digital technology strand2021
d = 0.36Kokoç, Bütüner & Güler, Journal of Research on Technology in Education2025
MediumImplementation
analyticsdatadashboardsLMSformativeSiemensdata-literacyinterpretation

Interrogate any educational claim

Heard a claim somewhere else? Type it here — we'll build a prompt you can paste into ChatGPT, Claude, or any chatbot. The prompt asks the model to cite real meta-analyses, separate strong from weak evidence, name boundary conditions, flag exaggerations, and discuss cost-effectiveness.

Sources & further reading