Catch up on the 2026 e-Assessment Awards Winners Series. Explore recordings from across the series and hear directly from our award-winning teams as they share the ideas, innovations and experiences behind their recognised work.
This talk presents the design, implementation, and impact of the Duolingo English Test’s award-winning Interactive Speaking task, the first AI-powered, fully adaptive speaking interview to be operationalized in a large-scale high-stakes assessment without a human examiner.
I’ll begin by outlining the core innovation: a real-time dialogue system that dynamically selects follow-up questions using prompt-specific rubrics, and a scoring system that applies item response theory (IRT) to ensure fairness and comparability across adaptive paths. I will provide an overview of the evidence-based approach we used to design, develop, and validate the task, one that integrates large-scale piloting, psychometric modeling, and iterative expert review.
The second half of the session will highlight several current research projects growing out of this work. These include: (1) enhancing interactivity through clarification and elaboration follow-ups; (2) investigating how linguistic complexity in prompts affects performance across speaking and writing; and (3) ongoing development of a listening comprehension measure based on the interactive speaking task responses.
Together, these innovations and research strands illustrate how AI can not only make assessments more accessible and efficient, but also open new possibilities for what language tests can measure. We hope this session offers a concrete model of how research, technology, and fairness can be integrated in the next generation of language assessment.
Chaeck out the paper mentioned in this presentation here!
The Duolingo English Test (DET) is an online, remotely administered high-stakes assessment of English language proficiency. A test taker can take it on their own computer, anywhere and any time, with results returned within 48 hours. Because every session is recorded and reviewed rigorously after submission, our security system has even more signals to work with than a live proctor does.
In this talk, we present our multi-layered security system that leverages human-in-the-loop AI to protect the integrity of the Duolingo English Test. It spans multiple signal sources and modalities: the responses a test taker submits, where we detect AI-generated responses; the way those responses were produced, where typing behavior separates transcription from organic writing; the patterns that connect one session to another, which reveal impersonation and collusion; and what the cameras show, where we direct proctors to the moments worth further examining. Each layer brings additional robustness to the overall security posture.
The presentation goes through a representative set of these security layers in practitioner detail: the threat that motivated the work, how the detector was built, and how it is evaluated and applied in practice, supported by our peer-reviewed research at EMNLP, AAAI, HCOMP, NCME and AIME-Con. We will also present the human oversight on top of the AI system, and how we ensure that final decisions are accurately made with human judgment in combination with AI signals.
We hope to provide insights that can inform the design of human-in-the-loop AI systems that secure the integrity of high-stakes assessments.
View the Duolingo public security report here
Further Duolingo research and publications here
Keep informed