Evaluating and Comparing AI Tools for Teachers: What Works and What Doesn’t

Image
Students in a classroom connected by digital links

AI tools are rapidly becoming part of teachers’ day-to-day instructional work, influencing how they plan, teach, and manage the classroom. Many schools are interested in using AI tools, but it’s hard to know which ones really help teachers and students. There is not enough clear evidence about which tools work best, are fair, and are easy to use.

AIR is building a way to test and compare AI tools for teachers. We will create clear rules (benchmarks) for what makes a good AI tool in the classroom. We will also build a virtual testbed, which is a safe online space where we can try out these tools and see how they work in real teaching situations.
 

Our Process

We use a four-step process to make sure our work is strong and useful.

Step 1: We identify the most common and fastgrowing ways teachers are using AI in classrooms. We do this by listening to teachers, reviewing national trends, and learning what districts say they need most.
 

Step 2: We build and test clear benchmarks. These are checklists that measure how well each AI tool works. We look at things like how well the tool helps with teaching, if it fits different classrooms, if it is easy to use, and if it is fair for all students.
 

Step 3: We put these benchmarks into a virtual testbed. This is an online space where we can safely test AI tools. We use both computer scoring and human review to make sure our results are accurate.
 

Step 4: We try out our system with real schools and teachers. We collect feedback and make improvements. At each step, we share our work through easy-to-read reports, dashboards, and guides.
 

Next Steps

We will share our findings with schools, districts, and developers. Our goal is to help teachers and leaders choose AI tools that are proven to work and are easy to use. We will also make our testbed and checklists available so others can use them to test new tools in the future.