Research - Duolingo

Duolingo Research

Science powers our mission to make language education free and accessible to everyone.

About Us

With more than 500 million learners, Duolingo has the world's largest collection of language-learning data at its fingertips. This allows us to build unique systems, uncover new insights about the nature of language and learning, and apply existing theories at scales never before seen. We are also committed to sharing publications and data with the broader research community.

Publications

Jump-Starting Item Parameters for Adaptive Language Tests

A.D. McCarthy, K.P. Yancey, G.T. LaFlair, J. Egbert, M. Liao, and B. Settles

EMNLP Proceedings, 2021

Mining Process Data to Detect Aberrant Test Takers

M. Liao, J. Patton, R. Yan, and H. Jiao

Measurement: Interdisciplinary Research and Perspectives, 2021

Methods for Language Learning Assessment at Scale: Duolingo Case Study

L. Portnoff, E. Gustafson, J. Rollinson and K. Bicknell

EDM Proceedings, 2021

A Sleeping, Recovering Bandit Algorithm for Optimizing Recurring Notifications

K.P. Yancey and B. Settles

KDD Proceedings, 2020

... (other publications here)...

Data & Tools

Replication data for our KDD 2020 paper, "A Sleeping, Recovering Bandit Algorithm for Optimizing Recurring Notifications." Includes 200 million examples of Duolingo practice reminder push notifications sent to Duolingo users over a 35 day period, including which template was used, whether the user converted within 2 hours, and other metadata.

Data for the 2020 Shared Task on Simultaneous Translation And Paraphrase for Language Education (STAPLE). This corpus contains more than 3 million pairs of English sentences with multiple possible translations into Portuguese, Hungarian, Japanese, Korean, and Vietnamese.

Data for the 2018 Shared Task on Second Language Acquisition Modeling (SLAM). This corpus contains 7 million words produced by learners of English, Spanish, and French. It includes user demographics, morph-syntactic metadata, response times, and longitudinal errors for 6k+ users over 30 days.

... (other datasets here)...

Our Team

We are a diverse team of experts in AI and machine learning, data science, learning sciences, UX research, linguistics, and psychometrics. We work closely with product teams to build innovative features based on world-class research.

André Horie AI + Machine Learning

Bożena Pająk Learning + Curriculum

Erin Gustafson Data Science + Analytics

Cindy Berger Learning + Curriculum

... (other team members here) ...