AI for Educators Daily with Dan Fitzpatrick

Rethinking edtech evaluation

Dan Fitzpatrick

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 11:29

Send us Fan Mail

Two-thirds of teachers use AI, but only one in five edtech products has evidence of improving outcomes.

In this episode:

  • Nearly two-thirds of teachers use AI, but only 20% of AI edtech products have evidence of improving outcomes, underscoring the urgent need for better AI education research.
  • Traditional randomized controlled trials (RCTs) are often too slow and rigid for evaluating rapidly evolving AI tools, necessitating new education research methods.
  • Stacey Alicea and Meghan McCormick propose 'implementation research and development' as a robust framework for assessing AI tool effectiveness through iterative testing and refinement.
  • Three guiding principles for AI edtech evaluation are: building evidence in stages, asking 'how it works' before 'whether it works,' and letting specific research questions dictate the methodology.
  • The Research Partnership for Professional Learning's Shared Measures Toolkit demonstrates effective, iterative evaluation, building measurement infrastructure crucial for responsible AI in classrooms.

Chapters:

  • 00:00 — Cold open & welcome
  • 00:45 — The problem: AI use outpaces AI edtech evaluation
  • 01:30 — Why traditional RCTs fail for AI education research
  • 02:45 — Introducing implementation research and development (R&D) for AI tool effectiveness
  • 03:45 — National efforts embracing iterative AI edtech evaluation
  • 04:30 — Principle 1: Build AI evidence in stages (feasibility first)
  • 05:30 — Principle 2: Ask 'how it works' before 'whether it works' for AI in classrooms
  • 06:45 — Principle 3: Let research questions drive the education research methods
  • 08:00 — Implications for school leaders and the need for faster evidence
  • 09:00 — Example: Research Partnership for Professional Learning's Shared Measures Toolkit

How can we evaluate new AI tools in education more effectively?
To evaluate new AI tools effectively, educators should shift from relying solely on slow randomized controlled trials to iterative 'implementation research and development' that rapidly tests and refines tools in real-world settings.

Why are traditional education research methods not working for AI?
Traditional education research methods like randomized controlled trials are often too slow and designed for static interventions, making them unsuitable for the rapid and continuous evolution of AI tools in education.

What is implementation research and development for AI in education?
Implementation research and development (R&D) is an approach that prioritizes rapid testing, feedback, and refinement of early-stage AI products to understand their design, delivery, and real-world usage, providing initial evidence on their effects before large-scale trials.

Featuring: Dan Fitzpatrick, Stacey Alicea, Meghan McCormick, Institute of Education Sciences, Leanlab Education, Boston University's EVAL initiative, Teaching Lab, Research Partnership for Professional Learning, Shared Measures Toolkit.

Follow AI in Education with Dan Fitzpatrick for more on AI in education.

SPEAKER_00

If this episode makes you think, please let us know in the comments and support us by subscribing and leaving a review. Thank you. Today we are exploring the critical challenge of evaluating AI-based tools in education, drawing insights from an illuminating piece titled AI is Rapidly Changing Education and Research Needs to Keep Up. This article published on july twenty first, twenty twenty six by Stacy Alysia and Megan McCormick dives into why our traditional methods for gauging the effectiveness of new educational technology just aren't cutting it for the fast-paced world of artificial intelligence. And what they reveal is quite striking. Nearly two-thirds of teachers now report that they use AI in their work, yet only about one in five ed tech products and classrooms today actually has any evidence that they can improve teaching and learning outcomes. That's a huge gap, isn't it? Now first let's talk about the problem itself. It's a fantastic starting point that conversations about AI and education often begin with the right question. Is there evidence of the effects of this technology? Of course we want to know what impact an intervention will have before it lands in a classroom impacting students and teachers. But as Alisaya and McCormick point out, the next question often becomes too narrow. We tend to jump straight to what have we learned about this technology from randomized controlled trials? And the issue is randomized controlled trials, or RCTs, are notoriously difficult to conduct in education. They take a lot of time, a lot of resources, and they typically require a stable, well-defined intervention that doesn't change much. That's where the pursuit of evidence often stops because RCTs are just not readily available for the vast majority of new tools. The authors make a really compelling case that AI is a fundamentally different type of intervention from what RCTs were designed to evaluate. Think about it. An AI tool today, let's say one that helps teachers with instructional feedback or analyses student work might look completely different next month. Developers are constantly updating models, refining prompts, and the ways teachers and students interact with these tools are evolving in real time. It's an evolution, not a revolution, yes, but a very fast evolution. An RCT assumes that the thing you're testing stays pretty much the same throughout the study, but with AI that assumption just doesn't hold. So while RETs have been essential for building evidence in areas like class size reductions or the science of reading, they simply can't keep pace with the rapid innovation in AI. We need strong evidence now, not perfect evidence years later, especially when school leaders are making decisions about adopting AI tools in their systems. This is why we need to rethink our approach to AI education research. The second big idea here is that instead of the traditional RCT, Alyssea and McCormick advocate for something called implementation research and development tool, or implementation RD. This approach is much better suited to the evidence needs of AI technology because it's traditionally used for studying early stage products. It's all about rapid test and feedback and refinement. So it helps us understand how a tool is designed, delivered, and used in real-world settings before we even think about a massive large scale evaluation. It gives us that initial evidence on a product's effects on teaching and learning, and crucially if a product isn't working, it helps us figure out why. Now the authors are clear. Implementation R and D isn't meant to replace RCTs entirely. Rather it's the foundation upon which to build research that eventually proves cause and effect. It focuses on learning what works before testing whether it works at scale. I think that's a really important distinction, and it connects directly to the purpose over technology pillar of my core philosophy. We're not just asking if the tech works, we're asking if it serves its educational purpose and how it's doing that. What's fascinating is that some national efforts are already embracing this. The Institute of Education Sciences, or IES, for example, has funded generative AI RD centers that prioritize iterative development and pilot testing before any formal trials. And then you have partnerships like Lean Lab Education and Boston University's Eval Initiative, both supporting technology through these cycles of real-world testing and refinement. This iterative feedback loop is really powerful. This brings us to a really practical framework of three guiding principles that Alisaya and McCormick offer for building evidence for AI tools. The first principle is to set build evidence in stages. For early stage AI tools, the focus should be on feasibility and rapid cycle testing to assess specific design choices. Think of it like a sandbox, right? You're experimenting, trying things out, seeing what sticks, you're trying to prove the concept works. The larger, more resource-intensive RCTs should come much later, only once the intervention is stable, the implementation is consistent, and there's already some initial evidence that target teacher and student outcomes are improving. This aligns beautifully with my Act Now principle within the seven lessons for AI adoption, which encourages educators not to wait for perfect conditions, but to get started with hands-on pilots and testing. So you can't expect a polished final product if you're not allowing for early messy iteration. The second principle really resonates with me. Ask how it works before asking whether it works. This is crucial for any educational intervention, not just AI tools. Before we even consider if an AI tool improves student achievement, we should be asking if it's actually changing the behaviors it was designed to change. For example, if you're using an AI coaching tool, is it genuinely shifting how teachers plan lessons, how they instruct, or how they respond to student thinking? If those foundational changes aren't happening, then we need to address those issues first. There's no point in measuring student impact if the tool isn't even moving the needle on teacher practice. It's about understanding the process and the productive struggle, not just the eventual outcome. The real value is not in what the machine produces, but in how the student, or in this case, the teacher, responds and adapts. And the third principle is to let the research questions drive the method. AI tools give us incredible capabilities to run frequent low-cost experiments and collect really detailed usage data in real time. The field should be leveraging these strengths through things like A B testing and continuous evaluation, instead of clinging solely to methods that were designed for static, unchanging programs. An ongoing study of an AI coaching tool from Teaching Lab, for instance, compares a standard version of the tool to one that emphasizes student-centered instruction. This allows them to see if subtle design differences actually shift coaching practices. Because the tool captures interactions in real time, researchers can analyse responses to feedback as they occur, enabling faster, more responsive research. This allows researchers to ask much more nuanced questions rather than just a blunt does it work or not. It's about precision in evaluating AI tool effectiveness. So what does this mean for school leaders and really for every educator listening who is thinking about implementing AI in classrooms? The shift that Alyssea and McCormick are proposing has massive implications for decision makers. A state or a district investing in AI today cannot afford to wait years for evidence from a study on a version of an intervention that will likely no longer exist by the time the data comes in. Evidence that arrives too late simply cannot guide decisions about adoption or scaling up. Sadly, many organizations reach significant scale without even basic outcome data on teaching and learning. That's a real risk. Implementation RD offers a much more practical path to making sure that the tools we're using are actually effective before we take them to scale. It empowers organizations to generate credible early evidence, allowing them to refine their approaches and build towards more rigorous evaluation down the line. It's about starting with why, not how, and continuously checking that your how is serving your why. An excellent example of this implementation RD in practice is the Research Partnership for Professional Learning's Shared Measures Toolkit. Instead of starting with a fully developed measurement product, the RPPL worked collaboratively with professional learning organizations and researchers. They identified what aspects of high-quality instructional materials and curriculum-based professional learning practitioners really needed to understand and improve. The resulting measures are now being tested, refined, and validated across multiple contexts, ensuring they're psychometrically sound and that they generate useful information for continuous improvement and decision making. This kind of work is built in a crucial measurement infrastructure that can help districts and professional learning organizations better understand implementation quality, compare results, and pinpoint high-leverage practices that lead to stronger teacher and student outcomes. This type of iterative evaluation is critical for responsible AI ed tech evaluation. I think this approach strengthens the conditions for eventual RCTs. It doesn't lower the bar. It just sequences implementation RD before the causal testing mechanisms. It's an evolution, not a revolution, in how we do education research. If we truly want AI tools to genuinely improve teaching and learning, we absolutely need to study how those AI interventions are being deployed in near-to-real time, using iterative experimentation to move from early design to long term impact. This way, we're teaching students not to outsmart machines, but to outthink them and ensuring the tools we use truly serve their learning. The core message is clear. If we want better tools, we need better, faster research. That's all for today. Thanks for listening.