Research - Duolingo
Duolingo Research
Science powers our mission to make language education free and accessible to everyone.
About Us
With more than 500 million learners, Duolingo has the world's largest collection of language-learning data at its fingertips. This allows us to build unique systems, uncover new insights about the nature of language and learning, and apply existing theories at scales never before seen. We are also committed to sharing publications and data with the broader research community.
Publications
Jump-Starting Item Parameters for Adaptive Language Tests
A.D. McCarthy, K.P. Yancey, G.T. LaFlair, J. Egbert, M. Liao, and B. Settles
EMNLP Proceedings, 2021
Mining Process Data to Detect Aberrant Test Takers
M. Liao, J. Patton, R. Yan, and H. Jiao
Measurement: Interdisciplinary Research and Perspectives, 2021
Methods for Language Learning Assessment at Scale: Duolingo Case Study
L. Portnoff, E. Gustafson, J. Rollinson and K. Bicknell
EDM Proceedings, 2021
A Sleeping, Recovering Bandit Algorithm for Optimizing Recurring Notifications
K.P. Yancey and B. Settles
KDD Proceedings, 2020
... (other publications here)...
Data & Tools
Replication data for our KDD 2020 paper, "A Sleeping, Recovering Bandit Algorithm for Optimizing Recurring Notifications." Includes 200 million examples of Duolingo practice reminder push notifications sent to Duolingo users over a 35 day period, including which template was used, whether the user converted within 2 hours, and other metadata.
Data for the 2020 Shared Task on Simultaneous Translation And Paraphrase for Language Education (STAPLE). This corpus contains more than 3 million pairs of English sentences with multiple possible translations into Portuguese, Hungarian, Japanese, Korean, and Vietnamese.
Data for the 2018 Shared Task on Second Language Acquisition Modeling (SLAM). This corpus contains 7 million words produced by learners of English, Spanish, and French. It includes user demographics, morph-syntactic metadata, response times, and longitudinal errors for 6k+ users over 30 days.
... (other datasets here)...
Our Team
We are a diverse team of experts in AI and machine learning, data science, learning sciences, UX research, linguistics, and psychometrics. We work closely with product teams to build innovative features based on world-class research.
André Horie AI + Machine Learning
Bożena Pająk Learning + Curriculum
Erin Gustafson Data Science + Analytics
Cindy Berger Learning + Curriculum
... (other team members here) ...