Self‐Directed Learning Favors Local, Rather Than Global, Uncertainty

Self-Directed Learning Favors Local, Rather Than Global, Uncertainty

Douglas B. Markant, Burr Settles, Todd M. Gureckis

Abstract

Collecting (or “sampling”) information that one expects to be useful is a powerful way to facilitate learning. However, relatively little is known about how people decide which information is worth sampling over the course of learning. We describe several alternative models of how people might decide to collect a piece of information inspired by “active learning” research in machine learning. We additionally provide a theoretical analysis demonstrating the situations under which these models are empirically distinguishable, and we report a novel empirical study that exploits these insights. Our model-based analysis of participants’ information gathering decisions reveals that people prefer to select items which resolve uncertainty between two possibilities at a time rather than items that have high uncertainty across all relevant possibilities simultaneously. Rather than adhering to strictly normative or confirmatory conceptions of information search, people appear to prefer a “local” sampling strategy, which may reflect cognitive constraints on the process of information gathering.

Keywords: Information sampling; Self-directed learning; Active learning; Machine learning

1. Introduction

A cornerstone of many educational philosophies is that people learn more effectively when they direct their own learning experience (Boekaerts, 1997; Bruner, 1961). Although there are many ways that such control might influence learning, one important factor is the ability to choose among different sources of information, a decision-making process we refer to as self-directed information sampling (Gureckis & Markant, 2012). A canonical example of this type of decision making is a doctor deciding which diagnostic test to perform on a sick patient (e.g., either an MRI or a blood test). Given a set of illnesses that may be causing the patient’s symptoms, the physician must decide whether a particular test, whose outcome is as yet uncertain, is most likely to reveal the correct diagnosis. The doctor plays a vital role in this assessment, relying on knowledge of the patient and potential illnesses to determine the value of a new source of information.

Relative to our understanding of how people evaluate alternatives and manage uncertainty in economic decision-making (Glimcher & Rustichini, 2004; Kahneman & Tversky, 1979; von Neumann & Morgenstern, 1944), less is known about how people judge the usefulness of new sources of information during learning in order to make sampling decisions. The goal of the present paper is to advance our understanding of this aspect of human decision-making and learning.

1.1. Two views on human information sampling

Existing studies of information sampling have contributed to two seemingly contradictory theoretical positions. On one hand, a large body of work on hypothesis testing and reasoning suggests that people prefer to sample information that is consistent with their existing beliefs even though that information may be ineffective for learning (Klayman, 1995; Nickerson, 1998), a behavior we refer to as confirmatory sampling. One well-known example of this is the positive test strategy (PTS), whereby people focus on observations that are positive examples of their current hypothesis without accounting for how those observations relate to alternative hypotheses (Klayman & Ha, 1987, 1989). In general, confirmatory sampling suggests a singular focus on an existing hypothesis and the data that are likely to be observed if it were true, while alternative hypotheses are neglected, often resulting in information that is less useful for learning.

In contrast, a separate line of research has argued that people’s information sampling decisions are consistent with normative principles, according to which alternative hypotheses are integrated to determine what information is most useful for adjudicating between them (see Nelson, 2005 for review). This normative framework, based on theories of “optimal experimental design” first developed in the statistics literature (Fedorov, 1972; Lindley, 1956), has accounted for sampling decisions in a range of information search tasks (Nelson, McKenzie, Cottrell, & Sejnowski, 2010; Oaksford & Chater, 1994), including relatively open-ended or complex domains such as visual search (Najemnik & Geisler, 2005), spatial search (Gureckis & Markant, 2009; Markant & Gureckis, 2012), causal structure learning (Steyvers, Tenenbaum, Wagenmakers, & Blum, 2003), and sequence learning (Austerweil & Griffiths, 2011). As demonstrated by Nelson (2005), different normative models within this framework (e.g., probability gain or information gain) may predict distinct sampling decisions depending on the task structure and the learner’s goal, but as a group they share the principle of integrating across all possible hypotheses to determine the most useful source of information.

1.2. An intermediate position

Although these two perspectives have long been considered to be in opposition to each other, they can also be viewed as two endpoints along a continuum representing the degree to which information from multiple hypotheses contribute to behavior. Confirmatory sampling is consistent with a single hypothesis controlling behavior while normative information selection usually implies combining information from all viable hypotheses. Where people fall along this continuum may depend on their ability to consider alternative hypotheses and their relationships to possible outcomes, which may be limited in complex problems involving large numbers of hypotheses or other demands. For example, individual differences in working memory capacity predict how many alternatives influence judgments about the probability of a focal hypothesis, thereby determining how well people’s judgments correspond to normative predictions (Dougherty & Hunter, 2003; Newstead, Thompson, & Handley, 2002; Sprenger et al., 2011). Predicting how people make sampling decisions may thus require an understanding of the cognitive constraints that limit performance in any particular task.

This perspective on information sampling, with confirmatory and normative sampling representing two extreme positions, raises a number of questions. First, there is the problem of formalizing intermediate models that make use of alternative hypotheses but to a lesser extent than predicted by normative theories. This idea is closely related to recent models of approximate Bayesian inference, which link sub-optimal learning or decision-making to an impoverished representation of the hypothesis space (Griffiths, Vul, & Sanborn, 2012; Sanborn, Griffiths, & Navarro, 2010). There have been few attempts as yet, however, to define such a framework in the context of human information sampling (but see Steyvers et al., 2003, for a similar approach in a causal learning setting). In what follows, we describe one approach to formalizing intermediate models inspired by contemporary research in active machine learning (AML).

Second, it is as yet unclear what task or environmental circumstances might lead to sampling behavior that is better described by an intermediate model. For example, for certain kinds of hypothesis spaces, confirmatory sampling is consistent with normative goals, precluding the need to consider more than one alternative (Austerweil & Griffiths, 2011; Klayman & Ha, 1987; Navarro & Perfors, 2011; Nelson & Movellan, 2001). Learners who correctly account for these constraints might be expected to show more or less confirmatory sampling depending on the problem space. Similarly, previous work has shown that prior experience with a domain is associated with normative sampling, such as when dealing with familiar materials (Cox & Griggs, 1982; McKenzie, 2006) or in problems involving social information gathering (Trope & Mackie, 1987).

One relatively unexplored possibility that we consider in the present paper is that intermediate strategies are likely to manifest in sufficiently complex problems. One strategy for collecting information in complex or multivariate domains is to decompose a problem into simpler components and to reduce “local” sources of uncertainty. For example, when multiple features may be related to an outcome, a learner might hold one feature constant while varying the other across multiple samples (Avrahami et al., 1997; Rottman & Keil, 2012), which is known as the “control of variables” strategy and is essential to scientific reasoning. Isolating variables helps people to search the space of potential hypotheses (Klahr & Dunbar, 1988) and is a key part of “learning to learn” about complex concepts (Kuhn & Dean, 2005). As a result of using this strategy during self-directed learning, sampling decisions may be better described by a model that focuses on local sources of uncertainty relevant to subsets of alternative hypotheses, rather than strictly confirmatory or normative accounts.

In this paper, we present a direct test of the idea that people prefer local sources of uncertainty when making sampling decisions, using a category-learning paradigm in which they control the selection of training items. We show that in this kind of problem, uncertainty about how to classify an item is directly related to how informative it is about the true category rule, raising the possibility that people rely on their uncertainty about how to predict the outcomes of potential queries in order to decide between them. This proposal is directly inspired by research on AML, in which such uncertainty sampling is a common, computationally efficient method for selecting training data for artificial classifiers. We describe a range of sampling models that vary in the degree to which they integrate information about alternative categories to predict what information is useful to sample. We then present an experiment designed to test between these alternative accounts of information sampling. After reporting the results of our experiment and subsequent modeling, we discuss the implications of this work for our understanding of how people direct their own learning.

2. Uncertainty sampling in AML

Effectively gathering information during learning is a problem that faces machine learners just as it affects human learners. A crucial factor in machine learning is the availability of labeled training data, which often requires costly human annotation. For example, credit card companies rely on automated systems for detecting fraudulent activity, but a human reviewer needs to label such instances before they can be used for supervised training of a classifier. Given the high costs of labeling, only instances that will improve the performance of the model should be selected for review. AML research has explored how selections should be made in order to maximize the accuracy of a machine learning model (Settles, 2012), with applications in a wide range of settings, including text classification (Olsson, 2009), natural language processing (Settles & Craven, 2008), and recommendation systems (Rubens, Kaplan, & Sugiyama, 2011).

Early work on AML (e.g., Cohn, Ghahramani, & Jordan, 1996; Mackay, 1992) drew upon the same framework of optimal experimental design that has guided recent psychological research. Accordingly, the same set of normative models explored in studies of human information sampling have been applied to problems involving machine classifiers. These models are often “prospective” in that they estimate the value of an observation by simulating the effect of each of its possible outcomes on the current model. For example, given a potential training item, information gain formalizes the reduction in uncertainty that would be achieved for each possible labeling (e.g., a reviewer identifying a transaction as fraudulent or not), and the overall expected value of selecting that item is found by weighing each outcome’s value by its predicted likelihood of occurring.

However, determining the best selection strategy in this way is computationally intractable in many machine learning applications. As a result, researchers have also developed methods for sampling an item based on a model’s current uncertainty in how to classify it (a measure which does not require estimating the effect of observing that item). In comparison to prospectively evaluating the optimal decision, using the model’s current classification uncertainty to decide what to learn about is less costly and in many cases achieves similar improvements in efficiency.

In the following sections, we describe a set of simple models, originally proposed in the AML literature (Settles, 2012), which predict the value of new information based on classification uncertainty. Importantly, these models vary in the way that uncertainty about alternative categories is integrated to make a sampling decision. For a potential training item x, there is a set of possible category labels {y₁, y₂, ...} that could result from observing that item. The probability of each category label being assigned to x is given by the distribution p(y|x).

2.1. Information gain and label entropy

The first uncertainty sampling model we consider is label entropy, which is defined as the Shannon entropy over the label distribution p(y|x) for an observation x (see Settles, 2012):

$$ L E(x)=-\sum_{i}p(y_{i}|x)\ln p(y_{i}|x). $$

Shannon entropy measures the disagreement in predictions across possible labels for an item. Entropy is highest when the model predicts each label equally, and it is lowest when a certain prediction is made for a single label. Thus, label entropy quantifies the amount of predictive uncertainty the observer has for item x. Those items that the observer has difficulty predicting the class membership for are assumed to be useful to observe, as they represent some aspect of the world where the learner is uncertain and would benefit from feedback.

While label entropy is commonly used in AML systems, it is also closely related to normative models that are used in cognitive psychology (Klayman & Ha, 1987; Nelson & Movellan, 2001; Oaksford & Chater, 1994; Steyvers et al., 2003). In fact, when hypotheses are deterministic (i.e., the likelihood of an observation is either 1 or 0 for all possible hypotheses), label entropy is formally equivalent to information gain, a prospective normative model that has been used to account for sampling behavior in a number of tasks. Information gain is defined as the reduction in uncertainty resulting from a new observation x. Having observed a previous set of observations D = {hx, yi₁; hx, yi₂;...}, uncertainty is given by the Shannon entropy of the posterior distribution (now defined over a hypothesis space H with uniform prior likelihood):

$$ I(p(h|D))=-\sum_{h\in\mathcal{H}}p(h|D)\ln p(h|D) $$

where N indicates the number of hypotheses in H that are consistent with the observations in D. Information gain is the reduction in uncertainty that would occur as a result of observing that item x has label yᵢ:

$$ I G(\langle x, y_{i}\rangle) = I(p(h|D)) - I(p(h|\langle x, y_{i}\rangle , D)). $$

Thus, for deterministic hypotheses, the expected information gain of a query (evaluated one step ahead) is equivalent to the entropy measured over its possible outcomes. This is maximized when an item is “globally” uncertain such that all outcomes have equal probability. If hypotheses have uniform prior probability (as is the case here), this occurs when each outcome is predicted by an equal number of plausible hypotheses.

2.2. Margin sampling

While focusing on items that are globally uncertain or unpredictable seems intuitively useful, there is reason to expect that it may not be the sampling strategy humans use, particularly when learning in complex, multivariate environments. One natural strategy, not captured by label entropy, might be to decompose a complex task into a set of simpler problems. We can formalize the strategy of focusing on separate components in a sampling model that values uncertainty about any boundary between only two categories. Label margin predicts that the learner will prefer instances for which the likelihood of any two categories is similar, independent of any other categories. When the label distribution P(y|x) is ordered from highest to lowest probability {p₁, p₂, ...}, with p₁ indicating the highest label probability, label margin is based on the difference between the two most likely labels for x:

$$ L M(x)=1-(p_{1}-p_{2}). $$

Critically, label margin is not maximized for only those items about which the learner is globally uncertain. Instead, there is a preference for local sources of uncertainty between subsets of potential outcomes. Whereas label entropy integrates information about all possible labelings of an item, label margin relies on the two most likely outcomes, disregarding the rest of the label distribution. As discussed above, this model thus reflects an intermediate sampling model in that a subset of possible alternatives (e.g., different categories) is used to evaluate whether a potential training item is worth learning about.

2.3. Most certain

Finally, previous work on hypothesis testing suggests that people may prefer items that they can already classify with relative confidence. People have a well-documented bias toward seeking positive evidence of the hypothesis they are considering (Klayman & Ha, 1989; Wason, 1960). To quantify this strategy, we define the most certain measure as:

$$ M C(x)=\max\big(p(y|x)\big). $$

The predictions of this model directly contrast those of label entropy, with the highest value assigned to items that can already be classified with confidence. The most certain measure is one way of instantiating confirmatory sampling—it shows a preference for items for which the learner has a strong prediction about the category label.

3. Empirical studies of information sampling during category learning in humans

In a recent study, we examined the interaction of self-directed information selection and category learning (Markant & Gureckis, 2014). In this experiment, people learned about two categories of “antennas” that varied along two perceptual dimensions (circles that differed in size and the orientation of a central line segment, see Fig. 1) and received one of two television stations (CH1 or CH2). We compared a self-directed condition, in which participants designed stimuli to learn about, with a standard, passive condition in which instances were generated from predefined distributions. A main finding from this study was that for simple uni-dimensional rules, self-directed learners acquired the correct category rule faster than passive learners (see also Castro et al., 2008).

In light of evidence that self-directed sampling can speed learning, it is important to understand how people decide what data to collect. Given a potential observation, what information do people rely on to decide if it will be useful? As we have proposed above, one aspect that may explain a person’s decision to sample an item is his or her uncertainty in how to classify it. Intuitively, a self-directed learner should direct his or her attention to items that are high in uncertainty while ignoring items that can already be confidently classified or predicted. Consistent with this strategy, the pattern of stimuli sampled by self-directed learners in our previous study revealed that participants systematically directed their samples toward the category boundary as the task progressed, suggesting a preference for items that they were uncertain about how to classify.

However, from that study we were unable to identify which sampling model best accounted for people’s decisions. One reason for this is that we did not directly measure participants’ uncertainty about the items they decided to learn about, and uncertainty can’t be directly inferred from the items chosen. A given item might be associated with either high or low subjective uncertainty depending on how much a person has learned, regardless of where it falls in the stimulus space.

In addition, the binary classification task precludes a comparison of the label entropy and label margin models, both of which make highly similar predictions in that task. As seen in the top row of Fig. 2, each heatmap describes the value assigned to a potential observation depending on the learner’s uncertainty in how to classify it. For example, an item that can be confidently classified (e.g., p(y|x) = (1,0), corresponding to the left edge of the heatmap) would be assigned a low value by label entropy and label margin, but a high value by most certain. For the binary classification problem, label entropy and label margin make highly similar predictions about how items will be valued. Items close to the center of the space have the highest value, and the ranking of items is identical between both models, making it difficult to distinguish between them. Settles (2012) observed that the predictions of these models diverge when considering more complex categorization tasks. For example, in a ternary classification task (see bottom row of Fig. 2), label entropy predicts a preference for items for which all three classes are likely (e.g., near the junction of the category boundaries). In contrast, label margin assigns the maximum value to items for which one category is highly unlikely but the learner is uncertain about the other two. In short, this model predicts that samples are likely to be allocated close to any boundary between two categories.

4. Experiment

4.1. Participants

Sixty undergraduates at New York University participated in the study for psychology course credit. One participant was excluded for ending the task early. The experiment was run on standard desktop computers in a single 1-h session.

4.2. Stimuli

The category label associated with each stimulus was deterministically defined by a ternary classification rule of the form shown in the bottom row of Fig. 2. In addition to the structure that is shown, three more rules were created through different rotations (90, 180, and 270 degrees) of the same boundaries in stimulus space. Each participant was randomly assigned to one of the four rules and a random mapping of labels (“CH1,” “CH2,” “CH3”) to categories. Training stimuli were chosen by participants according to the procedure below. Stimuli for each test block were generated by subdividing the stimulus space into a grid of 36 equally sized regions and generating a random stimulus from each region.

4.3. Procedure

Participants were instructed that the stimuli in the experiment were television “loop antennas” and that each unique antenna received one of three channels (CH1, CH2, or CH3). Their goal was to learn the difference between the three types of antennas so that they could correctly classify new antennas during the test blocks. The experiment alternated between training blocks (8 trials each) and test blocks (36 trials each). Participants were told that the experiment would end when they correctly classified 34 of 36 test items (94%) in a single test block. If a participant failed to reach that criterion, the experiment ended after 16 rounds or at the end of an hour (whichever occurred first).

4.4. Results

4.4.1. Classification performance

Thirty-six participants (62%) successfully reached the accuracy criterion of 94% correct within the available time. Of those participants, the average number of blocks to criterion was 6 (SD = 3.1). For the remaining participants, the average number of blocks completed was 9.8 (SD = 3.6).

4.4.2. Probability judgments

On each training trial, the participant judged the likelihood that the stimulus they selected belonged to each of the three categories, resulting in three values between 0 and 1 based on where they clicked within the response scale. We measured the difference between the initial cursor position and the participant’s response and classified any trial that did not differ by more than 5% of the scale from the initial position as a non-response. Using this exclusion criterion, the average proportion of non-responses was .33 (SD = .17). Two participants were excluded from further analysis because their proportion of non-responses was more than three standard deviations above the group mean (83% and 94%).

4.4.3. Overall model fits

Our first goal was to assess the overall fit of the three sampling models to each participant’s set of probability judgments. Classifying participants according to the model with the highest log-likelihood, we found that 17 people (30%) were best-fit by label entropy, 32 people (56%) were best-fit by label margin, and the remaining 8 people (14%) were best-fit by most certain. Judgments made by participants, separated by the best-fitting model, are plotted within the three-category simplex. A higher density of points reflects an increased tendency (as a group) to sample stimuli in a given region of the probability judgment space.

4.4.4. Relating sampling decisions to learning

We examined whether success at learning the rule was related to the sampling strategy reflected in participants’ probability judgments. Participants who reached the learning criterion made more samples in the label margin region than those who failed to reach the criterion (t(55) = 2.04, p < .05). Therefore, successful learning in the task was associated with increased sampling of items that are most consistent with the label margin model.

5. Discussion

In many real-world contexts, people can control what information forms the basis of their learning and decision-making, and their performance often hinges on how they make sampling decisions. Evidence of margin sampling suggests a general preference for a local form of exploration, but it is not diagnostic about the exact underlying sampling process. One possibility is that, when faced with a multidimensional task, people decompose the problem into simpler components. This kind of piecemeal strategy may be more effective when it is difficult to simultaneously consider many alternatives or to process information about multiple feature dimensions.

Another limitation of the current study is our dependence on participants’ self-reported probability judgments. A goal of ongoing work is to verify the validity of margin sampling using alternative measures of subjective uncertainty.

Moreover, people might pool information about any subset of alternatives when evaluating potential samples. Notably, margin sampling is an efficient means for improving the efficiency of training. When hypotheses are deterministic, uncertainty about an item’s label directly correlates with the expected information it is likely to convey.

6. Acknowledgments

This work was supported by grant number BCS-1255538 from the National Science Foundation and the Intelligence Advanced Research Projects Activity (IARPA).

References

Austerweil, J., & Griffiths, T. (2011). Seeking confirmation is rational for deterministic hypotheses. Cognitive Science, 35, 499–526.

Cox, J., & Griggs, R. (1982). The effects of experience on performance in Wason’s selection task. Memory and Cognition, 10(5), 496–502.

Dougherty, M., & Hunter, J. (2003). Hypothesis generation, probability judgment, and individual differences in working memory capacity. Acta Psychologica, 113(3), 263–282.

Glimcher, P., & Rustichini, A. (2004). Neuroeconomics: The confluence of the brain and decision. Science, 306, 447–452.

Klayman, J. (1995). Varieties of confirmation bias. Psychology of Learning and Motivation, 32, 385–418.

Klahr, D., & Dunbar, K. (1988). Dual space search during scientific reasoning. Cognitive Science, 12, 1–48.

Nelson, J. (2005). Finding useful questions: On Bayesian diagnosticity, probability, impact, and information gain. Psychological Review, 114(3), 677.