Evidence Toolkit  ›  AI Literacy

AI Literacy 7 interventions

Skills for teaching students to use AI critically — not AI as a tutor, but the human judgment needed to interrogate, verify and bound what AI produces. Fact-checking AI output with lateral reading, auditing it against critical-thinking standards, interrogating a chatbot across multiple rounds, using domain expertise to catch distortions, mapping where AI is reliable across disciplines, teaching prompt quality, and setting defensible assignment-level AI-use policies. This is the companion to the AI Learning Science domain — that one is about AI as a learning tool (tutors, feedback, worked examples); this one is about the critical literacy students need to not be fooled by it. Honest evidence profile: AI literacy is a new field, so the base is deliberately mixed — two cards carry independent meta-analytic effect sizes, one maps to an EEF dialogic-teaching strand, and four are shown openly as unverified rather than dressed up with a borrowed number.

Avg. effect (d): 0.36 across the 2 cards with a d-value Strongest: AI Hallucination Fact-Check Protocol (d=0.42 overall; lateral-reading subgroup 0.55 — Fendt, Muth & Edelsbrunner 2025) EEF cross-ref: 1/7 cards — AI Socratic Dialogue maps to the EEF Dialogic Teaching strand (+2 months); the rest have no clean strand Verification profile: 3 verified-partial, 4 unverified — a new field, so most cards show the evidence gap openly Last reviewed: June 2026

How to read these numbers

  • "Months of progress" is a teacher-friendly shorthand from the EEF Toolkit. It is not directly comparable across studies — different meta-analyses use different baselines, age groups, and outcome measures. Treat it as a magnitude indicator, not a precise prediction.
  • Cohen's d is the standardised effect size used in the original meta-analyses. d ≈ 0.40 is Hattie's "hinge point" — the average effect of a year of schooling. Higher = larger relative effect, but context matters more than the number.
  • Effect sizes are averages. A skill that shows large average effects can still produce small or negative effects in a specific classroom. Use these as a starting point for professional judgment, not a substitute for it.
  • Sources are dated. Where multiple meta-analyses exist, we lead with the most recent quality study and cross-reference EEF where available.
  • Implementation cost (Low / Medium / High) is a practical signal — not an exact science — of what your school needs to invest in teacher time, training, and structural changes to actually run the intervention well. It is the editorial team's reading of what the intervention typically requires in practice. Use it to gauge whether something is straightforward to introduce or a larger undertaking — not as a budget figure.
  • Some interventions are also priced in £ (UK) by the EEF Toolkit. For monetary cost data, see the EEF Teaching & Learning Toolkit.
  • Colour bands signal calibration, not value. An intervention in the "below typical" band is not "bad" — it means the intervention's average effect is below the typical effect of a year of schooling. That can still be appropriate for specific contexts the average does not capture.
  • This is a new field, and the cards say so. AI literacy emerged in 2023–2025, so for several skills no meta-analysis exists yet. Where that is the case we mark the card unverified and show the gap openly — citing the established research the activity builds on (critical thinking, källkritik / lateral reading, domain-specificity) rather than inventing a number. A grey "no independent meta-analysis found" pill is an honest statement, not an oversight.
AI Literacy

AI Hallucination Fact-Check Protocol

Teach students a fact-checking protocol for AI-generated text, extending the SIFT method (Stop, Investigate the source, Find better coverage, Trace claims to the original) with AI-specific moves for catching hallucinated facts and fabricated citations. The core discipline is lateral reading — leaving the AI's answer to verify its claims against independent sources, the way professional fact-checkers do, rather than reading vertically down the page the model produced. Evidence anchor: Fendt, Muth & Edelsbrunner (2025) in Learning and Individual Differences — a meta-analysis of 64 controlled studies (17,120 participants, 2002–2025) of interventions that teach source-credibility assessment. The overall pooled effect is g=0.42 (95% CI 0.35–0.49); the lateral-reading subgroup specifically (14 studies) is the strongest of the four approaches at g=0.55 (95% CI 0.35–0.75). The card anchors on the conservative overall g=0.42, because the lateral-reading subgroup carries a wider confidence interval and high heterogeneity (I-squared around 87%). The meta pools the Stanford Civic Online Reasoning lineage this skill builds on — Wineburg, Breakstone, McGrew & Smith's lateral-reading work and the Brodsky et al. community-college trials. Honest framing: the meta covers teaching source-credibility assessment broadly; verifying AI hallucinations and fabricated citations is a specific, newer application of the same moves, so g=0.42 transfers the underlying skill, not the AI-specific task — no AI-hallucination-specific meta-analysis exists yet. No matching EEF Toolkit strand. Pairs with the AI Output Critical Audit Designer.

no EEF strand
d = 0.42Fendt, Muth & Edelsbrunner, Learning and Individual Differences2025
MediumImplementation
AI-literacyhallucinationfact-checkingSIFTlateral-readingAI-citationsverificationFendtcivic-online-reasoning
AI Literacy

AI Output Critical Audit Designer

Design a structured protocol for students to audit AI-generated text against Ennis's six critical-thinking standards — clarity, accuracy, relevance, logic, depth and fairness — annotating where the output meets each standard and where it fails. The move converts a passive 'read what the AI wrote' into an active evaluation of it. Evidence anchor: Abrami, Bernard, Borokhovski, Waddington, Wade & Persson (2015) in Review of Educational Research — 'Strategies for Teaching Students to Think Critically: A Meta-Analysis', pooling 341 effect sizes for a weighted mean of g=0.30 (p<.001). Abrami's actionable moderator finding sits directly under this skill: the largest critical-thinking gains came from instruction that combined dialogue, authentic or anchored problems, and mentorship — not from worksheets in isolation. Honest framing: this is a meta-analysis of teaching critical thinking in general, not of auditing AI output specifically, and no meta-analysis on AI-output auditing exists. The g=0.30 transfers the broad construct — auditing a text is a defensible instance of critical-thinking instruction — and the conservative overall figure is used rather than the larger subset effects for dialogue-plus-mentoring designs. No EEF Toolkit strand maps to critical thinking; the closest, Metacognition and self-regulation (+8 months), is a different construct and is deliberately not used here. Pairs with the AI Hallucination Fact-Check Protocol and AI Socratic Dialogue Designer.

no EEF strand
d = 0.30Abrami, Bernard, Borokhovski, Waddington, Wade & Persson, Review of Educational Research2015
MediumImplementation
AI-literacycritical-thinkingEnnisauditannotationAI-outputepistemicAbramiReview-of-Educational-Research
AI Literacy

AI Socratic Dialogue Designer

Design a multi-round questioning sequence for interrogating an AI chatbot's answers — pressing on a claim across several turns, tracking how the response shifts, and teaching students to distinguish a genuine, evidence-based update from sycophantic capitulation (the model caving simply because it was pushed). The skill targets a real, documented behaviour of current chatbots: they often change their answer under pressure regardless of whether the pressure was warranted. Evidence anchor: the EEF Dialogic Teaching evaluation (Jay et al., Sheffield Hallam University, 2017) — an efficacy trial of Robin Alexander's dialogic-teaching approach across 76 schools (around 4,958 pupils) reporting +2 months in English and science (+1 in maths) at moderate security (3 of 5 padlocks). The broader EEF Oral Language Interventions strand reads +6 months across 188 studies. Honest framing: dialogic teaching is human classroom talk, not student-to-AI interrogation, so +2 months is a proxy for the underlying construct — structured, multi-turn questioning that builds reasoning — not direct evidence for the AI-mediated activity. A correlational meta-analysis exists (Tao & Chen, 2024, Educational Research Review) reporting dialogic teacher talk associated with achievement at r=.25, but it reports correlations, not a causal Cohen's d, so no d is shown here. Pairs with the Questioning & Discussion domain, where the dialogic-teaching evidence is treated in depth.

+2 monthsEEF Dialogic Teaching evaluation (Jay et al.)2017
no independent meta-analysis found
MediumImplementation
AI-literacySocratic-questioningdialogic-teachingsycophancymulti-roundcritical-thinkingEEForal-languagecapitulation
AI Literacy

AI Expertise Interrogation Designer

Run a 'Funhouse Mirror' activity: students take a topic they genuinely know well — a sport, a hobby, a subject they're strong in — and probe an AI's claims about it, surfacing the distortions, omissions and confident-but-wrong statements that are invisible to a novice. The pedagogical logic is that you can only catch an AI's errors in a domain where your own knowledge is good enough to act as the mirror. Honest evidence status: unverified. No meta-analysis measures the learning effect of using one's own expertise to detect AI distortions — this is a specific, novel classroom activity with no pooled effect size. What legitimately grounds it: (1) Simonsmeier, Flaig, Deiglmayr, Schalk & Schneider (2021), Educational Psychologist — a meta-analysis of 8,776 effect sizes showing prior domain knowledge enables comprehension and evaluation, though its average predictive power for learning gains is small and highly conditional, which is itself a useful caution against over-claiming; (2) the well-established metacognition and self-regulation literature (in Hattie's synthesis, metacognitive strategies sit around d=0.60), since the activity trains confidence calibration; (3) a primary study, Dang & Nguyen (2025, Stanford SCALE), finding only around 20% of students spontaneously detected an AI hallucination — documenting the gap this activity addresses. Treat the card as defensible, theory-grounded practice rather than a measured intervention. Pairs with the AI Output Critical Audit Designer.

no EEF strand
no independent meta-analysis found
LowImplementation
AI-literacyexpertisedistortionFunhouse-Mirrormetacognitionprior-knowledgeDunning-KrugerSimonsmeier
AI Literacy

Disciplinary AI Literacy Sequence Designer

Design a sequence where students pose the same question to an AI across different disciplines — a historical interpretation, a mathematical proof, a literary reading, a scientific explanation — and compare how reliable or distorting the output is in each, building a mental model of where AI is trustworthy and where it isn't, keyed to the type of knowledge involved. Honest evidence status: unverified. No meta-analysis measures this activity. The grounding is theoretical and well-established: critical-thinking and reasoning skills are substantially domain-specific and transfer poorly across fields (Willingham, 2020, American Educator; Tricot & Sweller, 2014, Educational Psychology Review), which is precisely why a single 'is AI reliable?' judgment is the wrong model and a discipline-by-discipline map is the right one. Bernstein's distinction between hierarchical and horizontal knowledge structures gives the same point a sociology-of-knowledge frame. Adjacent AI-as-tutor meta-analyses (for example ChatGPT around g=0.67) measure a different intervention and are deliberately not borrowed here. Treat the card as defensible, theory-grounded practice. Pairs with the AI Output Critical Audit Designer and with the knowledge-type cards in Curriculum & Assessment Design.

no EEF strand
no independent meta-analysis found
MediumImplementation
AI-literacydisciplinary-thinkingknowledge-typesWillinghamTricot-Swellerdomain-specificityBernsteinAI-reliability
AI Literacy

Prompt Literacy Sequence Designer

Teach prompt quality directly: students compare a vague prompt with a refined one on the same task and see how specificity, context and constraints transform the output — so they understand why AI results vary rather than treating the tool as a slot machine. Honest evidence status: unverified. Prompt literacy is a genuinely new area (2023–2025) with only single studies, not a pooled meta-analysis. The strongest controlled evidence is one large randomised trial — Xiao et al. (2026, N=979, an introductory computing course) — where instruction reliably improved students' prompting skill and prompting gains correlated with exam performance, but it reported no standardised effect size and found no direct transfer to exam scores between conditions. A PRISMA systematic review (Plaatjies & van Wyk, 2025) grounds prompt literacy as a teachable practice qualitatively. Important trap avoided: 'prompt engineering' results that show a better prompt raising a model's benchmark accuracy are not student-learning effect sizes and are not used here. Treat the card as defensible, emerging practice with no verified magnitude yet. Pairs with the AI Output Critical Audit Designer and AI Learning Boundary Mapper.

no EEF strand
no independent meta-analysis found
LowImplementation
AI-literacyprompt-engineeringprompt-literacyspecificitycontextconstraintsXiaoemerging-evidence
AI Literacy

AI Learning Boundary Mapper

A teacher planning tool: for a given assignment, map which parts genuinely benefit from AI assistance and which parts AI use undermines the learning the assignment exists to produce — then set a defensible, assignment-specific AI-use policy on that basis, rather than a blanket ban or a blanket permission. Honest evidence status: unverified, and appropriately so — this is a planning construct, not a measured student intervention, so an effect size in the usual sense does not exist. The practice is grounded in backward design / Understanding by Design (start from the learning objective, then decide where a tool helps or short-circuits it) and in a cluster of 2024 frameworks for assessment redesign in the age of generative AI (for example Lye & Lim, 2024, Education Sciences; Chan & Colloton, 2024). A deliberate trap avoided: a large reported d of around 1.5 in this space measures AI-driven score inflation — the gap between AI-permissive take-home and proctored results — not a learning gain from any redesign method, so it is not used as an anchor. Treat the card as a design discipline, not a ranked intervention. Pairs with the Backwards Design Unit Planner in Curriculum & Assessment Design.

no EEF strand
no independent meta-analysis found
LowImplementation
AI-literacyassignment-designAI-policybackward-designassessment-redesignlearning-objectivesAI-boundariesacademic-integrity

Interrogate any educational claim

Heard a claim somewhere else? Type it here — we'll build a prompt you can paste into ChatGPT, Claude, or any chatbot. The prompt asks the model to cite real meta-analyses, separate strong from weak evidence, name boundary conditions, flag exaggerations, and discuss cost-effectiveness.

Sources & further reading