Which LLM is best for language education tasks? Well, it depends…
Teachers, schools, universities, students — many are using LLMs in some way to help them teach or learn a second language.
But are all LLMs equal when it comes to supporting language education?
This is one of the questions we have set out to answer with the L2-Bench project. Today we’re excited to share the first round of results.
What is L2-Bench?
L2-Bench is a project in Oxford to build the world’s first standardized benchmark for AI in language education. Oxford University Press, in collaboration with researchers from the University of Oxford, has developed this benchmarking tool, focusing on pedagogical value, quality and safety. For more on how we built L2-Bench, see our previous blog.
The dataset that underpins the evaluation includes:
- More than 1,000 authentic teaching tasks
- A framework of 12 core teaching competencies with 31 sub-skills
- A structured scoring system based on expert descriptors of good performance
- Validation from more than 200 experienced educators across 45+ countries
Crucially, it evaluates learning experience design, not just knowledge.
That includes tasks like:
- Planning lessons and courses
- Designing activities
- Giving feedback
- Supporting learner interaction
- Managing socio-emotional aspects of learning
The results
We evaluated these nine LLMs against all 1,000 tasks in the L2-Bench dataset, and this is an overview of how they did. There’s a link at the end of this article to a paper that describes the research methodology and results in full.
This shows AI is already capable of producing generally high-quality outputs in language education tasks, but “Strong overall” doesn’t mean “reliable in all situations”. We need to avoid treating AI capability as a single score. It varies significantly depending on the task.
AI performs much better on structured tasks than open-ended ones
The results show that AI models do well on tasks such as lesson planning and activity design, which are clearly defined, predictable, and format-driven. But they can struggle with tasks that are more ambiguous, context-sensitive or human-centred, such as conversational interaction, socio-emotional intelligence, and professional judgement.
This is important to emphasise when people—managers and teachers—are trying to work out where the collaboration between teachers and AI tools works best.
In some ways, we haven’t yet pushed the models to do what is particularly important—interacting with learners or teachers over a series of exchanges. In the next phase of development for L2-Bench, we will introduce tasks with multi-turn activities. Maintaining the pedagogical value across that kind of task will be much more demanding.
Context matters more than we think
From our early pilots, we realised we had to develop a more systematic approach to defining different learning contexts. This has helped us to identify that these LLMs performed noticeably worse on tasks where the learning context had, for example:
- Resource-constrained environments
- Students at lower proficiency levels
- Younger learners
In many cases, the AI models also implicitly assume relatively privileged learning conditions: reliable internet access, self-directed learners, abundant resources, and technologically rich environments. As an organization that works with schools all over the world, we know first-hand that this does not reflect the reality in classrooms today.
This raises a critical question: does AI work best for already advantaged learners, and struggle where support is most needed?
As AI becomes more widely adopted, questions of equity and accessibility must become as important as questions of accuracy.
AI still makes risky mistakes
All the models showed issues with certain areas which normally require professional judgement, including:
- Cultural sensitivity
- Appropriateness for learner level
- Unrealistic assumptions about resources
- Data and privacy considerations
These findings remind us that AI outputs cannot be used uncritically, especially in high-stakes or diverse contexts; an important point in the light of the discussions around ‘cognitive surrender’, the tendency to accept AI-generated outputs without sufficient critical evaluation, described by Shaw and Nave 2026 writing from the Wharton School here.
Cost vs performance
L2-Bench also allows us to examine the relationship between cost and performance. For example, we can observe that the highest performing model can cost roughly 100 times as much as a lower-cost alternatives we tested.
For educational organisations operating at scale, this matters enormously. The future success of AI in education will depend not only on capability, but also on affordability and sustainability.
This highlights the growing importance of model optimisation, fine-tuning, and high-quality educational training data. In many cases, the challenge may be less about accessing the most powerful model and more about applying the right model effectively within a specific educational context.
What can we take from the research?
The findings highlight both the potential and the limitations of today’s AI models in language education.
Key takeaways are:
- A single score for a model hides the variable performance across different types of task. Different models are stronger for different types of activity.
- The top AI models are impressive in reaching human-level responses on many tasks, but there are areas to watch out for:
- Learners that need more support
- Context-sensitivity and professional judgement
- The socio-human dimension of teacher–learner interaction
- Designers of AI tools for language education will need to make careful choices trading off quality and running costs.
Read more about our results in our paper: L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education.
You can also see our previous blog which discusses our methods.
Looking ahead
L2-Bench is more than a research project. It represents a broader effort to establish educationally grounded standards for AI at a time when the technology is evolving rapidly. Because benchmarks influence what gets built, we believe that educational organisations and institutions have a responsibility to help define what good looks like. Deciding how educational success is measured is not down to AI companies alone; those standards must also come from the educators, researchers, and institutions that understand how learning actually happens.
Get involved
This L2-Bench project is designed for practitioners, not just researchers. Oxford University Press will be making the benchmark dataset available freely as an open-source tool that can be used by anyone involved in language education, and so would love to engage with as wide a range of practitioners as possible.
If you’re curious, sceptical, excited—or all three—you’re exactly the kind of voice we want involved.
If you’d like to contribute to the next validation phase, or simply want updates as we release findings, please get in touch via the “Register Interest” form below.
AI in education is moving fast. Let’s make sure our evaluations keep up — and reflect the real work teachers do every day.