portnoff.edm21.pdf

Methods for Language Learning Assessment at Scale:

Duolingo Case Study

Authors:
Lucy Portnoff
Duolingo
lucy@duolingo.com

Joseph Rollinson
Duolingo
joseph@duolingo.com

Klinton Bicknell
Duolingo
klinton@duolingo.com

ABSTRACT

Students using self-directed learning platforms, such as Duolingo, cannot be adequately assessed relying solely on responses to standard learning exercises due to a lack of control over learners’ choices in how to utilize the platform: for example, how learners choose to sequence their studying and how much they choose to revisit old material. To provide accurate and well-controlled measurement of learner achievement, Duolingo developed two methods for injecting test items into the platform, which combined with Educational Data Mining techniques yield insights important for product development and curriculum design. We briefly discuss the unique characteristics and advantages of these two systems - Checkpoint Quiz and Review Exercises. We then present a case study investigating how different study approaches on Duolingo relate to learning outcomes as measured by these assessments. We demonstrate some of the unique benefits of these systems and show how educational data mining approaches are central to making use of this assessment data.

Keywords


1. INTRODUCTION

Online learning platforms have at their disposal large volumes of data about how students engage with learning material, how they navigate educational software, and how the learning process unfolds over time. Using a variety of methods - machine learning, statistics, psychometrics, etc. - Educational Data Mining (EDM) and Learning Analytics (LA) researchers identify students at risk of dropout from a course, detect changes in study behavior, predict exam performance, and characterize the different learning strategies that learners adopt.
Duolingo is a learning platform that provides free language education through mobile apps and a website. With around 40 million users active on the platform each month, Duolingo may well possess the largest language learning dataset of any company or research institution. Researchers at Duolingo leverage EDM/LA methodologies to mine datasets—including internal assessment and log data—for insights that inform improvements to the learning experience, help identify opportunities for changes to curriculum design, and fuel research on second language (L2) learning more generally.

Due to the self-directed nature of the Duolingo learning platform and the desire for holistic learner assessment, we have developed two assessment systems -the Checkpoint Quiz and Review Exercises - that allow for carefully controlled measurement of learner achievement. These two assessments were designed with the challenges gamified platforms struggle with in mind, including ensuring the learning experience remains motivating and maintaining a scalable content creation process.

The utility of the Checkpoint Quiz and Review Exercises

The utility of the Checkpoint Quiz and Review Exercises for assessing learner achievement depends, at least in part, on the high volume of data collected from Duolingo learners and the EDM methodologies that can be applied to that data. By leveraging predictive modeling and natural language processing (NLP) methods, we are able to control for the various ways that learners choose to navigate through the platform. Further, these methods allow us to uncover useful insights into how this variation in user navigation relates to learning outcomes - insights that we can leverage for product development and curriculum design. In this paper, we present two of our assessment systems and a case study highlighting the importance of applying EDM methodologies to derive insights from Duolingo assessment and log data.

2. RELATED WORK

Most EDM/LA applications at Duolingo focus on pedagogy-oriented issues or computer-supported predictive analytics. Most relevant to the current work are studies focused on predicting performance on upcoming course exercises and predicting performance on an assessment.

Other studies rely both on knowledge tracing and assessment data to analyze course effectiveness and provide this more holistic view. This approach is especially useful in more self-directed learning platforms. One study used BKT (Bayesian Knowledge Tracing) to characterize learning using a digital game and used outputs from these models to predict post-test scores following a period of learning with the game. They found that mastery scores for two knowledge components had positive and significant association with post-test scores. Insights from the BKT model itself were also useful for identifying concepts that are difficult for students to master, which highlights opportunities for improving course effectiveness.

Knowledge tracing is not the only approach used for characterizing student behavior

Knowledge tracing is not the only approach used for characterizing student behavior using clickstream or log data. To make log data useful for predictive modeling, many researchers turn to methods from NLP to aggregate events. Simple methods include calculating n-grams for particular event types. For example, unigrams can capture the number of times a student completes a particular learning module and bigrams can capture the number of times students complete two modules in sequence. Such data can be used as inputs into predictive models either relying solely on raw n-gram counts or by processing the data further using unsupervised machine learning methods - such as hierarchical clustering - to identify common sequence patterns.

3. DUOLINGO ASSESSMENT SYSTEMS

3.1 Duolingo Course Structure

Duolingo courses are organized into a series of units, each of which concludes with a Checkpoint. Courses used by the majority of learners have the following structure: 25-30 skills per unit with five difficulty levels per skill and 5-6 lessons per level. Skills are designed around a particular theme (e.g., Travel). The vocabulary taught in the skill is aligned around that theme and grammatical topics tend to be consistent across lessons within a skill. Lessons typically consist of 12-15 exercises designed to teach some vocabulary and/or grammatical concept. Duolingo curriculum designers incorporate aspects of spiral curriculum to revisit familiar concepts in more complex contexts in future skills.

When learners begin a Duolingo course, not all skills in the first unit are immediately available; a row unlocks once Level 0 is complete for all skills in the prior row. For example, only the Basics 1 skill is available at first and the next set of skills in the row below Basics 1 will only unlock once Basics 1 reaches Level 1. Once skills are unlocked, learners are free to return to them to practice previously studied material and “level up” the skill. Duolingo learners are, therefore, given agency to choose their learning path. Some learners prefer to attempt only the foundational level in a skill (Level 0) before moving on to new material, while others prefer to level up all skills. Leveling up is entirely optional and learners are required to complete only the foundational level for each skill before they can move on to the next unit of content. This self-directed nature of the learning platform provides challenges for assessing learner achievement.

Other modes of learning are available to users outside of course skills. Learners can build reading and listening proficiency through the Stories feature, which reinforces unit content through interactive dialogues with exercises to check comprehension. Learners can also complete generalized practice sessions, which drill users on content they have already studied from throughout the course. Further, after learners have leveled a skill up all the way, they can return for skill practice to reinforce their knowledge. If learners find skill material too easy, they also have the option to “test out” of a level and jump to harder exercises at the next level.

We use a variety of methods to assess learner achievement and proficiency throughout a Duolingo course. In the sections below, we describe two of the core assessments in use today: Checkpoint Quiz and Review Exercises.

3.2 Checkpoint Quiz

For a subset of Duolingo’s courses, learners must complete a custom-built assessment once they finish a unit and reach a Checkpoint. The Checkpoint Quiz is an achievement test that measures the extent to which our learners have achieved the objectives for each unit of a course. Checkpoint Quiz items are independent from the items used in course skills and users are only exposed to the quiz items during the assessment. This ensures that learners do not have the opportunity to learn the items in the assessment while studying course content and is important for test validity.

Learners do not receive corrective feedback or a final grade for the assessment and may only take the quiz once. At each Checkpoint, learners complete a randomly generated quiz consisting of 15 items (sampled from a larger pool of items). Seven items are pre-test items that test the next unit of the course that the learner is about to start and another seven are post-test items that test the unit the learner just completed. This pre-test/post-test design allows us to establish a baseline level of performance so we can later assess gain in accuracy from pre-test to post-test. The final item is a self-directed writing item designed to assess the current unit (with no pre-test). The assessment tests knowledge of vocabulary, grammar, writing using separate items designed to test one of these language skills and components.

3.3 Review Exercises

Review Exercises prompt learners to review content from a skill earlier in their course. A single Review Exercise is inserted into randomly selected lessons in the foundational level of a skill for skills beyond the first five in the course. These exercises are randomly and uniformly sampled from the pool of available exercises from either three skills or five skills earlier in the course.
For example, randomly selected exercises from the Animals skill are injected into Level 0 lessons seen by learners studying the Places skill. Review Exercises come in two forms: assisted recall and translation from L1-to-L2 or vice versa.

Checkpoint Quiz Review Exercises
Slow data collection (only at Checkpoints) Fast data collection (at every skill in a course)
Tagged and calibrated by curriculum experts Not tagged or calibrated
Items siloed from course Items sampled from course
Only certain courses All courses

4. CASE STUDY

Learners on Duolingo use the platform in a variety of different ways. In this case study, we investigate how learning decisions impact outcomes, so that we could "nudge" learners to use the app more effectively. This case study demonstrates how EDM methodologies allow us to investigate the various ways that learners choose to navigate through the platform - focusing on differences in “leveling up” behavior -and how course navigation relates to learning outcomes. We show a correlation between leveling up and higher accuracy on the Checkpoint Quiz. Complementary modeling with Review Exercise data establishes a causal link between completing sessions in higher levels and accuracy on assessments.

4.1 Checkpoint Quiz

4.1.1 Data

Our work uses four months of Checkpoint Quiz data. For every learner completing at least two consecutive Checkpoint Quizzes within this timeframe, we collected the pre-test/post-test item response pairs as well as summary statistics on learners’ studying behavior in the unit the items assess. Responses to free-form writing items were not included in this analysis.

4.1.2 Methods

To isolate the impact of lessons completed at each level on Checkpoint Quiz outcomes, we built a logistic regression model to predict post-test scores for items that were answered incorrectly in the pre-test (a measure of learning gain). Primary variables of interest capture the number of lessons learners completed at a given level for each skill in the unit of interest. We found that average post-test item accuracy increases linearly with every skill-level completed.

4.1.3 Results

We also compared the magnitudes of the leveling up effects with those of other types of learning modes. Coefficients capturing leveling up behavior show dominant effects in the model; one additional skill-level has a greater impact on Checkpoint Quiz scores than one additional Story, skill practice, or generalized practice.

5. CONCLUSIONS

In a case study of the levels mechanic, wherein learners study content in increasingly difficult contexts by “leveling up”, complementary analyses of the Checkpoint Quiz and Review Exercises showed that completing sessions in higher levels leads to stronger performance on assessments. Analyzing accuracy rates on the Checkpoint Quiz by the number of skill-levels completed in the course unit revealed a strong positive trend. The Review Exercises analysis supports a causal link between leveling up and improved assessment performance, showing that completing additional levels for a skill has measurable learning value.

Together, these results directly motivated the implementation of a number of interventions that encourage learners to reach higher levels. For example, adding design elements that give learners a visual stand-in for how the levels system works.

This study focused on one type of variation in how learners navigate the platform, aiming to capture such variation, thereby improving model fit and deepening our understanding of how other types of navigational choices relate to learning outcomes. Future work will also continue to explore the utility and limitations of the Review Exercise assessment system.