# Publications

We are reinventing high stakes testing with innovative tools and techniques, supported by the latest research in language assessment, psychometrics, and machine learning.

* * *

\
\
Duolingo English Test: Technical Manual\
\
Ben Naismith, Ramsey Cardwell, Geoffrey T. LaFlair, Steven Nydick, Masha Kostromitina (2026)\
\
Duolingo Research Report](https://go.duolingo.com/dettechnicalmanual)  
\
\
Duolingo English Test: Demographic and Score Properties of Test Takers (July 2024-June 2025)\
\
Allison Michalowski, Ramsey Cardwell, Steven Nydick, and Ben Naismith (2025)\
\
Duolingo Research Report](https://go.duolingo.com/demographic-score)  
\
\
The Evolution of the Duolingo English Test\
\
Masha Kostromitina (2025, updated 2026)\
\
Duolingo Research Report DRR-26-03](http://go.duolingo.com/det-evolution)  
\
\
Analytics for Quality Assurance for Item Pools (AQuAP): Monitoring and Maintaining Item Bank Health in AI-Driven Assessment Systems\
\
Alina A. von Davier, Xiaowan Zhang, Yigal Attali, Yena Park, Jacqueline Church, Andrew Runge, Geoff T. LaFlair, Alexander Tsigler (2026)\
\
ArXiv Pre-Print](https://doi.org/10.48550/arXiv.2606.18536)  
\
\
A Scalable Parametric Item Calibration Engine (SPICE) for Explanatory IRT with Sparse Data\
\
Steven Nydick, Manqian Liao, J.R. Lockwood (2026)\
\
ArXiv Pre-Print](https://arxiv.org/abs/2605.21782)  
\
\
A Systematic Review of Test Component Ordering with Implications for Language Assessment\
\
Ben Naismith and Ramsey Cardwell (2026)\
\
Language Testing](https://journals.sagepub.com/doi/10.1177/02655322261432256)  
\
\
Responsible AI Standards\
\
Jill Burstein (2026)\
\
Duolingo Research Report](https://go.duolingo.com/ResponsibleAI)  
\
\
Advancing integrity in language assessment: Response to Bruce et al.\
\
Maria Kostromitina and Ramsey Cardwell (2026)\
\
Advancing integrity in language assessment: Response to Bruce et al., ELT Journal, 2026;, ccaf065, https://doi.org/10.1093/elt/ccaf065](https://academic.oup.com/eltj/advance-article/doi/10.1093/elt/ccaf065/8417073)  
\
\
A Theoretical Assessment Ecosystem for a Digital-First Assessment\
\
J. Burstein, G. T. LaFlair, and A. A. von Davier (2026)\
\
Duolingo Research Report DRR-26-02](https://go.duolingo.com/ecosystem)  
\
\
Alignment of the Duolingo English Test (DET) to the Common European Framework of Reference for Languages (CEFR): An External Validation of Individual Skill Subscores for Speaking, Writing, Reading, and Listening\
\
Kelley Stethen and Chad Buckendahl (2025)\
\
ACS Ventures, LLC Technical Report](https://go.duolingo.com/DET-SWRL-CEFR-Technical-Report)  
\
\
Where Assessment Validation and Responsible AI Meet\
\
Jill Burstein & Geoff LaFlair (2025)\
\
In C. Coombe, T. Clark, & H. Mohebbi (Eds.), Language Teaching Research Quarterly, 50, 120–137. https://doi.org/10.32038/ltrq.2025.50.09](https://api.eurokd.com/Uploads/Article/1883/ltrq.2025.50.09.pdf)  
\
\
A Concordance Study between the DET and IELTS Academic SWRL Subscores\
\
Steven Nydick, J.R. Lockwood (2025)\
\
Duolingo Research Report](https://go.duolingo.com/concordance-det-ielts-2024)  
\
\
Responsible Al for test equity and quality: The Duolingo English Test as a case study\
\
Burstein, J., LaFlair, G. T., Yancey, K., von Davier, A. A., & Dotan, R. (2025)\
\
In E. M. Tucker, E. Armour-Thomas, & E. W. Gordon (Eds.), Handbook for Assessment in the Service of Learning, Volume I: Foundations for Assessment in the Service of Learning. University of Massachusetts Amherst Library Press.](https://www.assessment-for-learning.org/volume-i)  
\
\
An Overview of Duolingo English Test Administration and Scoring\
\
Steven Nydick, J.R. Lockwood (2024, updated 2025)\
\
Duolingo Research Report](https://go.duolingo.com/Admin+Scoring)  
\
\
Exploring AI-Enabled Test Practice, Affect, and Test Outcomes in Language Assessment\
\
Jill Burstein, Ramsey Cardwell, Ping-Lin Chuang, Allison Michalowski, and Steven Nydick (2025)\
\
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con 2025), Pittsburgh, PA.) - Full Papers Volume.](https://aclanthology.org/2025.aimecon-main.7.pdf)  
\
\
When Machines Mislead: Human Review of Erroneous AI Cheating Signals\
\
Will Belzak, Chenhao Niu, and Angel Ortmann Lee (2025)\
\
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con 2025), Pittsburgh, PA.)- Works in Progress Volume.](https://aclanthology.org/2025.aimecon-wip.11/)  
\
\
Keystroke Analysis in Digital Test Security: AI Approaches for Copy-Typing Detection and Cheating Ring Identification\
\
Chenhao Niu, Yong-Siang Shih, Manqian Liao, Ruidong Liu, and Angel Ortmann Lee (2025)\
\
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con 2025), Pittsburgh, PA.)- Works in Progress Volume.](https://aclanthology.org/2025.aimecon-wip.13/)  
\
\
Span labeling with large language models: Shell vs. meat. \  
Phoebe Mulcaire and Nitin Madnani (2025)\
\
Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025) (pp. 850–859). Association for Computational Linguistics.](https://aclanthology.org/2025.bea-1.62.pdf)  
\
\
Beyond GPA and language proficiency: A systematic literature review of international students’ academic success factors.\
\
Masha Kostromitina, Ben Naismith, Jill Burstein, & Luke Plonsky (2025)\
\
Review of Education](https://bera-journals.onlinelibrary.wiley.com/doi/10.1002/rev3.70089)  
\
\
A multi-stage interactive writing task for the assessment of English language writing proficiency\
\
Andrew Runge, Sarah Goodwin, Yigal Attali, Mya Poe, Phoebe Mulcaire, Kai-Ling Lo, and Geoffrey T. LaFlair (2025)\
\
Language Testing](https://journals.sagepub.com/doi/10.1177/02655322251349908)  
\
\
Interactive Listening – The Duolingo English Test\
\
Geoffrey T. LaFlair, Andrew Runge, Yigal Attali, Yena Park, Jacqueline Church, Sarah Goodwin (2023, updated 2025)\
\
Duolingo Research Report](https://go.duolingo.com/interactive-listening-whitepaper)  
\
\
Assessing Listening on the Duolingo English Test\
\
Sarah Goodwin, Ben Naismith (2023, updated 2025)\
\
Duolingo Research Report](https://go.duolingo.com/listening-whitepaper)  
\
\
Assessing Speaking on the Duolingo English Test\
\
Yena Park, Ramsey Cardwell, Sarah Goodwin, Ben Naismith, Geoffrey T. LaFlair, Kai-Ling Lo, Kevin P. Yancey (2023, updated 2025)\
\
Duolingo Research Report](https://go.duolingo.com/speaking-whitepaper)  
\
\
Assessing Vocabulary on the Duolingo English Test\
\
Yena Park, Ramsey Cardwell, Ben Naismith (2024, updated 2025)\
\
Duolingo Research Report](http://go.duolingo.com/vocabulary)  
\
\
Duolingo English Test: Writing Construct\
\
Sarah Goodwin, Yigal Attali, Geoffrey T. LaFlair, Yena Park, Andrew Runge, Alina A. von Davier, Kevin P. Yancey, Ben Naismith, Ping-Lin Chuang (2022, updated 2025)\
\
Duolingo Research Report](https://go.duolingo.com/scored-writing)  
\
\
Developing an Automatic Pronunciation Scorer: Aligning Speech Evaluation Models and Applied Linguistics Constructs\
\
Danwei Cai, Ben Naismith, Masha Kostromitina, Zhongwei Teng, Kevin P. Yancey, Geoffrey T. LaFlair (2025)\
\
Language Learning](https://onlinelibrary.wiley.com/doi/10.1111/lang.70000)  
\
\
Guidelines for Fair Test Content: The Duolingo English Test Example\
\
Jacqueline Church, Yena Park, and Jill Burstein (2025)\
\
Duolingo Research Report DRR-25-03](https://go.duolingo.com/DETFairnessGuidelines)  
\
\
Duolingo English Test: Security and Score Integrity\
\
William Belzak, Basim Baig, Ramsey Cardwell, Rose Hastings, André Kenji Horie, Geoff LaFlair, Manqian Liao, Chenhao Niu, and Yong-Siang Shih (2025)\
\
Duolingo Research Report DRR-25-01](https://duolingo-papers.s3.us-east-1.amazonaws.com/reports/DET_Security_Report.pdf)  
\
\
The Item Factory: Intelligent Automation in Support of Test Development at Scale\
\
Alina A. von Davier, Andrew Runge, Yena Park, Yigal Attali, Jacqueline Church, & Geoff T. LaFlair (2024)\
\
In "Machine Learning, Natural Language Processing, and Psychometrics", eds. Hong Jiao & Robert W. Lissitz](https://doi.org/10.1108/979-8-88730-606-320251002)  
\
\
Where Assessment Validation and Responsible AI Meet\
\
J. Burstein and G. T. LaFlair (2024)\
\
ArXiv Pre-Print](https://arxiv.org/abs/2411.02577)  
\
\
"We would like to see ourselves in the test:" The experiences of francophone African English learners in high-stakes English proficiency testing\
\
Kadidja Koné, Paula Winke, & Matthew Gordon (2024)\
\
Open Science Framework Preprints](https://doi.org/10.31219/osf.io/tsbf5)  
\
\
Detecting LLM-Assisted Cheating on Open-Ended Writing Tasks on\
Language Proficiency Tests\
\
Chenhao Niu, Kevin Yancey, Ruidong Liu, Mirza Basim Baig, André Kenji Horie, and James Sharpnack (2024)\
\
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track](https://aclanthology.org/2024.emnlp-industry.70.pdf)  
\
\
A Generative AI-Driven Interactive Listening Assessment Task\
\
Andrew Runge, Yigal Attali, Geoffrey T. LaFlair, Yena Park, and Jacqueline Church (2024)\
\
Frontiers in Artificial Intelligence](https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2024.1474019/full)  
\
\
The impact of task duration on the scoring of independent writing responses of adult L2-English writers\
\
Ben Naismith, Yigal Attali, and Geoffrey T. LaFlair (2024)\
\
Assessing Writing](https://www.sciencedirect.com/science/article/pii/S1075293524000886)  
\
\
Duolingo English Test: Demographic and Score Properties of Test Takers (July 2023-June 2024)\
\
Allison Michalowski, Ramsey Cardwell, Steven Nydick, and Ben Naismith (2024)\
\
Duolingo Research Report](https://go.duolingo.com/demographic-score-2024)  
\
\
Estimating Test-Retest Reliability in the Presence of Self-Selection Bias and Learning/Practice Effects\
\
Will Belzak & J.R. Lockwood (2024)\
\
Applied Psychological Measurement](https://journals.sagepub.com/doi/10.1177/01466216241284585)  
\
\
BERT-IRT: Accelerating Item Piloting with BERT Embeddings and Explainable IRT Models\
\
Kevin P. Yancey, Andrew Runge, Geoffrey LaFlair, and Phoebe Mulcaire\
\
Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024)](https://aclanthology.org/2024.bea-1.35.pdf)  
\
\
Test preparation, test-taker attitudes, and high-stakes outcomes\
\
Steven Nydick, Ramsey Cardwell, Jill Burstein, Duolingo\
\
Presented at NCME 2024](https://duolingo-testcenter-prod.s3.amazonaws.com/media/resources/NCME+2024+eBoard+efficacy+study.pdf)  
\
\
From Pen to Pixel: Rethinking English Language Proficiency Admissions Assessments in the Digital Era\
\
Ramsey Cardwell, Ben Naismith, Jill Burstein, Steven Nydick, Sarah Goodwin, and Anthony Verardi (2024)\
\
CALICO Journal](https://utppublishing.com/doi/10.1558/cj.27104)  
\
\
Facilitating the Writing Process on the DET: The Interactive Writing Task\
\
Sarah Goodwin, Mya Poe, Ramsey Cardwell, Andrew Runge, Yigal Attali, Phoebe Mulcaire, Kai-Ling Lo, and Geoffrey T. LaFlair (2024)\
\
Duolingo Research Report DRR-24-02](https://duolingo-papers.s3.amazonaws.com/other/interactive-writing-whitepaper.pdf)  
\
\
Transforming Assessment: The Impacts and Implications of Large Language Models and Generative AI\
\
Jiangang Hao, Alina A. von Davier, Victoria Yaneva, Susan Lottridge, Matthias von Davier, Deborah J. Harris\
\
Educational Measurement: Issues and Practice](https://doi.org/10.1111/emip.12602)  
\
\
Measuring Variability in Proctor Decision Making on High-Stakes Assessments: Improving Test Security in the Digital Age\
\
William Belzak, JR Lockwood, & Yigal Attali (2024)\
\
Educational Measurement: Issues and Practice](https://doi.org/10.1111/emip.12591)  
\
\
Detecting careless cases in practice tests\
\
Steven Nydick (2023)\
\
Chinese/English Journal of Educational Measurement and Evaluation](https://doi.org/10.59863/LAVM1367)  
\
\
Practical considerations when building concordances between English tests\
\
Ramsey Cardwell, Steven Nydick, J.R. Lockwood, Alina A. von Davier (2023)\
\
Language Testing](https://doi.org/10.1177/02655322231195027)  
\
\
Fairness of using different English accents: The effect of shared L1s in listening tasks of the Duolingo English test\
\
Okim Kang, Xun Yan, Masha Kostromitina, Ron Thomson, Talia Isaacs (2023)\
\
Language Testing](https://doi.org/10.1177/02655322231179134)  
\
\
Ensuring Fairness of Human- and AI-Generated Test Items\
\
William C.M. Belzak, Ben Naismith, Jill Burstein (2023)\
\
Artificial Intelligence in Education](https://link.springer.com/chapter/10.1007/978-3-031-36336-8_108)  
\
\
The interactive reading task: Transformer-based automatic item generation\
\
Yigal Attali, Andrew Runge, Geoffrey T. LaFlair, Kevin Yancey, Sarah Goodwin, Yena Park, Alina A. von Davier (2023)\
\
Frontiers in Artificial Intelligence](https://doi.org/10.3389/frai.2022.903077)  
\
\
Automated Evaluation of Written Discourse Coherence Using GPT-4\
\
Ben Naismith, Phoebe Mulcaire, Jill Burstein (2023)\
\
Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications](https://aclanthology.org/2023.bea-1.32.pdf)  
\
\
Rating Short L2 Essays on the CEFR Scale with GPT-4\
\
Kevin Yancey, Geoffrey T. LaFlair, Anthony Verardi, Jill Burstein (2023)\
\
Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications](https://aclanthology.org/2023.bea-1.49.pdf)  
\
\
Incorporating test security into the validity argument of a remotely-proctored English test\
\
Ramsey Cardwell, Mancy Liao, Will Belzak, & Geoff LaFlair (2023)\
\
Presented at LTRC 2023](https://go.duolingo.com/LTRC23Security)  
\
\
Considering inter-test relationships in high-stakes admissions testing: The case of English\
\
Ramsey Cardwell, Steven Nydick, J.R. Lockwood (2023)\
\
Presented at the 8th International Conference of the Association of Language Testers in Europe ‘23](https://d23cwzsbkjbm45.cloudfront.net/media/resources/alte_23_cardwell.pdf)  
\
\
The magic of machine learning: How to develop tests in multiple languages simultaneously and at scale\
\
Sarah Goodwin, Lauren Bilsky, Phoebe Mulcaire, Burr Settles (2023)\
\
Presented at the 8th International Conference of the Association of Language Testers in Europe ‘23](https://d23cwzsbkjbm45.cloudfront.net/media/resources/alte_23_goodwin.pdf)  
\
\
Revisiting the language needs of students in English-medium universities\
\
Ramsey Cardwell, Ben Naismith (2023)\
\
Presented at the Biennial Conference of the British Association of Lecturers of English for Academic Purposes](https://d23cwzsbkjbm45.cloudfront.net/media/resources/baleap_23_cardwell.pdf)  
\
\
Fairness Auditing in a Digital-First Learning & Assessment Ecosystem\
\
Geoffrey T. LaFlair, Jill Burstein, Alina A. von Davier (2023)\
\
Presented at the Annual Meeting of the National Council on Measurement in Education](https://go.duolingo.com/NCME23Fairness)  
\
\
Sampling Considerations for Building Concordance Tables\
\
Ramsey Cardwell, Steven Nydick, J.R. Lockwood (2023)\
\
Presented at the Annual Meeting of the National Council on Measurement in Education](https://d23cwzsbkjbm45.cloudfront.net/media/resources/ncme_23_cardwell.pdf)  
\
\
Plagiarism Detection Using Human-in-the-loop AI\
\
Mancy Liao, Sinon Tan, Basim Baig (2023)\
\
Presented at the Annual Meeting of the National Council on Measurement in Education](https://d23cwzsbkjbm45.cloudfront.net/media/resources/ncme_23_liao.pdf)  
\
\
The regDIF R Package: Evaluating Complex Sources of Measurement Bias Using Regularized Differential Item Functioning\
\
William C.M. Belzak (2023)\
\
Structural Equation Modeling: A Multidisciplinary Journal](https://doi.org/10.1080/10705511.2023.2170235)  
\
\
Human-in-the-loop automated item creation for complex items in language assessments\
\
Geoffrey T. LaFlair, Yigal Attali, Andrew Runge, Kevin Yancey, Sarah Goodwin, Yena Park, Alina A. von Davier (2023)\
\
Presented at the Annual Conference of the American Association for Applied Linguistics](https://d23cwzsbkjbm45.cloudfront.net/media/resources/aaal_23_LaFlair.pdf)  
\
\
Finding time in language assessments: Maximizing measurement properties by unit of testing time\
\
Sarah Goodwin, Geoffrey T. LaFlair, J.R. Lockwood, Steven Nydick, Alina A. von Davier (2023)\
\
Presented at the Annual Conference of the American Association for Applied Linguistics](https://d23cwzsbkjbm45.cloudfront.net/media/resources/aaal_2023_goodwin.pdf)  
\
\
Maintaining and Monitoring Quality of a Continuously Administered Digital Assessment\
\
Mancy Liao, Yigal Attali, J.R. Lockwood, Alina von Davier (2022)\
\
Frontiers in Education](https://doi.org/10.3389/feduc.2022.857496)  
\
\
Duolingo English Test: Interactive Reading\
\
Y. Park, G. LaFlair, Y. Attali, A. Runge, and S. Goodwin (2022)\
\
Duolingo Research Report DRR-22-02](https://go.duolingo.com/interactive-reading)  
\
\
Digital-First Learning and Assessment Systems for the 21st Century\
\
Thomas Langenfeld, Jill Burstein, and Alina A. von Davier (2022)\
\
Frontiers in Education](https://go.duolingo.com/Digital-First_LAS)  
\
\
The Duolingo English Test: Psychometric considerations\
\
G. Maris (2020)\
\
Duolingo Research Report DRR-20-02](https://go.duolingo.com/drr-20-02)  
\
\
Digital-first assessments: A Security Framework\
\
Geoffrey T. LaFlair, Thomas Langenfeld, Basim Baig, André Kenji Horie, Yigal Attali, Alina A. von Davier (2022)\
\
Journal of Computer Assisted Learning](https://go.duolingo.com/security_framework)  
\
\
Analyzing the Duolingo English Test According to the AERA, APA, and NCME \(2014) Standards\
\
Maria Elena Oliver & Thomas E. Langenfeld (2021)](https://go.duolingo.com/Standards-Analysis-2021)  
\
\
Improving Test Validity and Accessibility with Digital-First Assessments\
\
Naomi Care & Bryan Maddox (2021)](https://go.duolingo.com/improving_test_validity)  
\
\
The Multidimensionality of Measurement Bias in High-Stakes Testing: Using Machine Learning to Evaluate Complex Sources of Differential Item Functioning\
\
William C. M. Belzak (2022)\
\
Educational Measurement Issues and Practice](https://go.duolingo.com/multidimensionality)  
\
\
Quality Assurance in Digital-First Assessments\
\
Mancy Liao, Yigal Attali, Alina von Davier, J.R. Lockwood (2022)\
\
Quantitative Psychology](https://doi.org/10.1007/978-3-031-04572-1_20)  
\
\
Jump-Starting Item Parameters for Adaptive Language Tests\
\
Arya D. McCarthy, Kevin P. Yancey, Geoffrey T. LaFlair, Jesse Egbert, Manqian Liao, and Burr Settles (2021)\
\
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing](https://go.duolingo.com/Jump-starting_item_parameters)  
\
\
AQuAA: Analytics for Quality Assurance in Assessment (poster)\
\
Manqian Liao, Yigal Attali, Alina A. von Davier (2022)\
\
Proceedings of the 14th International Conference on Educational Data Mining](https://go.duolingo.com/AQuAA_poster)  
\
\
AQuAA: Analytics for Quality Assurance in Assessment (paper)\
\
Manqian Liao, Yigal Attali, Alina A. von Davier (2022)\
\
Proceedings of the 14th International Conference on Educational Data Mining](https://go.duolingo.com/AQuAA_paper)  
\
\
Duolingo English Test: Subscores\
\
G. T. LaFlair (2020)\
\
Duolingo Research Report DRR-20-03](https://go.duolingo.com/subscorewhitepaper)  
\
\
Machine Learning Driven Language Assessment\
\
B. Settles, G. T. LaFlair, & M. Hagiwara (2020)\
\
Transactions of the Association for Computational Linguistics, 2020](https://doi.org/10.1162/tacl_a_00310)  
\
\
The Duolingo English Test: Design, Validity, and Value\
\
J. Brenzel and B. Settles (2017)\
\
Duolingo Research Report](https://go.duolingo.com/detwhitepaper)  
\
\
The Reliability of Duolingo English Test Scores\
\
B. Settles (2016)\
\
Duolingo Research Report DRR-16-02](https://go.duolingo.com/detreliability)  
\
\
The Duolingo English Test and Academic English\
\
L. Ishikawa, K. Hall, and B. Settles (2016)\
\
Duolingo Research Report DRR-16-01](https://go.duolingo.com/detacademicenglish)  
\
\
The Duolingo English Test and East Africa: Preliminary linking results with IELTS & CEFR\
\
M. Bézy and B. Settles (2015)\
\
Duolingo Research Report DRR-15-01](https://s3.amazonaws.com/duolingo-papers/reports/DRR-15-01.pdf)  
\
\
Validity, Reliability, and Concordance of the Duolingo English Test\
\
F. Ye (2014)\
\
University of Pittsburgh Technical Report](https://go.duolingo.com/dettoefl)
