{"paper_id":"03de0c1a-d970-4222-9287-7c987c092ead","body_text":"Selective Pupil Size Response Within direct and random exploration and exploitation Behaviors | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Selective Pupil Size Response Within direct and random exploration and exploitation Behaviors Gili Barkay, Shai Gabay, Uri Herz This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7740172/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract When making decisions, the explore - exploit dilemma represents balancing reward maximization with uncertainty reduction. While reinforcement learning models often treat exploration as stochastic variability, theories such as Adaptive Gain Theory (AGT) and Expected Value of Control (EVC) suggest that exploration and exploitation can reflect strategic control allocation. Pupillometry provides a window into the locus coeruleus–norepinephrine system, indexing cognitive effort and task engagement. The present study combined pupillometry with the Horizon Task to examine whether directed and random exploration and exploitation differentially recruit cognitive resources under varying environmental contexts. Thirty-five adults completed a task manipulating three variables: value gap (reward difference), information gap (sampling imbalance), and choice horizon (1 vs. 6 free trials). Behavioral analyses replicated established findings: small value gaps promoted exploration, unequal sampling elicited information-seeking, and long horizons increased directed exploration. However, pupillary responses diverged from behavior, showing selective sensitivity to choice horizons. Pupil size was larger in short horizon conditions compared with long-horizon, suggesting increased control engagement under conditions of heightened consequence, whereas value and information gaps did not elicit significant modulation. These findings challenge the view that exploration uniformly entails increased effort and instead highlight that strategic context determines pupil-indexed engagement. Our results provide evidence that cognitive effort allocation during exploration depends on the prospective utility and irreversibility of decisions, bridging behavioral and neurocomputational perspectives. Biological sciences/Neuroscience Biological sciences/Psychology Social science/Psychology Figures Figure 1 Figure 2 Introduction Decision-makers face a fundamental challenge: should they exploit the best familiar option or explore uncertain alternatives that might yield better long-term outcomes? 1,2 . This trade-off manifests across species and contexts, from animal foraging behaviors to human strategic planning 3 – 5 , yet questions persist regarding the psychological and neurobiological mechanisms underlying these competing drives 4 , 6 . Human exploration can take the form of directed exploration, involving deliberate information-seeking, or random exploration, reflecting stochastic variability in choice behavior 7 . The present study addresses these processes by examining pupillary responses, a reliable physiological marker of locus coeruleus–noradrenergic activity and task engagement, across different forms of exploration. Formulations of reinforcement learning (RL), such as those grounded in the multi-armed bandit framework, modeled exploration as stochastic variability in action selection, typically relying on probabilistic rules to balance exploratory and exploitative choices 8 – 10 . This framing implicitly positions exploration and exploitation as distinct and competing behavioral modes. In contrast, optimal foraging theory conceives of exploration as a purposeful, information-driven process aimed at maximizing long-term reward 4 . This conceptual divergence highlights the broader issue of whether exploration reflects random noise in choice behavior or a purposeful strategy grounded in cognitive control. A second influential framework, the Adaptive Gain Theory (AGT), links exploration–exploitation behavior to task engagement and its neurobiological substrates. AGT posits that shifts in locus coeruleus–norepinephrine (LC–NE) activity govern exploration and exploitation, with moderate tonic activity supporting exploitation (focused, reward-driven performance) and elevated tonic states reflecting disengagement and a shift toward exploration 11 , 12 . This account provides direct predictions about pupil size during exploration-exploitation decisions, with exploration corresponding to reduced cognitive control and larger pupil size. However, subsequent evidence has revealed that human exploration comprises two functionally distinct processes; directed and random exploration that challenging this dichotomy and highlighting tensions between RL and AGT frameworks 7 , 13 – 16 . Beyond this dichotomy, research has emphasized the role of environmental uncertainty in shaping exploration and exploitation. Yu and Dayan 17 distinguished between expected uncertainty, reflecting stable but noisy environments, and unexpected uncertainty, signaling change-points that demand adaptation. Expected uncertainty arises in familiar contexts with predictable variability, whereas unexpected uncertainty emerges when outcomes violate predictions, prompting active information-seeking 17 , 18 . Evidence from multi-armed bandit tasks confirms that humans actively seek uncertain options when learning is possible, with uncertainty guiding exploration in principled ways 19 . In this context, Payzan-LeNestour and Bossaerts 20 showed that unexpected uncertainty promotes novelty-seeking (directed exploration), whereas estimation uncertainty can trigger ambiguity aversion rather than random exploration. Together, these accounts suggest that exploratory behavior is not monolithic but shaped by the interplay of potential learning benefits and perceived risk 21 – 23 . These perspectives highlight the importance of understanding cognitive effort allocation during exploration and exploitation. Another theory about engagement is the Expected Value of Control framework (EVC) 24 proposes that cognitive control is allocated by weighing expected benefits against intrinsic costs. Importantly, effort’s relationship with value is more complex than simple cost models predict, as decisions are most demanding when competing options are similar in value 24 and effort itself may become intrinsically rewarding 25 . While EVC has not been directly applied to pupillometry, its principles suggest that cognitive engagement should be scaled with decision importance and strategic value rather than exploration type per se. Thus, the intensity of cognitive engagement should scale with both anticipated benefits and the learned value of effortful engagement, offering a theoretical bridge between exploratory strategies and their physiological signatures. Within this context, pupillometry provides a natural candidate for testing predictions about control allocation, as pupil size reliably reflects LC–NE activity, cognitive effort and task engagement 11 , 26 . According to AGT, exploration should correspond to the tonic LC mode, characterized by elevated baseline activity and larger pupil diameter, reflecting disengagement from the current task and search for alternatives. Early support for this relationship came from Aston-Jones and Cohen's human studies 11 , which demonstrated that baseline pupil diameter increased progressively as task utility declined in an auditory discrimination task, reaching maximum values when participants chose to abandon the current task series. Consistent with these AGT predictions, Jepma and Nieuwenhuis 12 found larger baseline pupil size preceding exploratory choices compared to exploitative choices in a multi-armed bandit task. However, these consistent findings of pupillary dilation during exploration align with broader research demonstrating that pupil dilation reliably indexes cognitively demanding processes and mental effort deployment 27 , including uncertainty processing and surprise 28 and learning rate adjustments 29 . Moreover, pupil responses may reflect the informational relevance of outcomes, scaling with their potential to update internal models 30 . However, this raises new questions about when exploration requires cognitive effort. Several investigations have found limited evidence for directed exploration behavior in humans, potentially due to methodological limitations that make it difficult to distinguish between these exploration strategies 10 , 31 , 32 . To address these methodological constraints, Wilson and colleagues 7 developed the Horizon Task, which systematically controls participants' information regarding option payoffs through forced-choice trials that decorrelate reward and information. This paradigm enables systematic manipulation of uncertainty levels, information exposure, and future decision opportunities, providing an optimal framework for examining how environmental contexts influence the cognitive demands of different exploratory strategies. This theoretical distinction has also received empirical support. Kozunova et al. (2022) 33 demonstrated that directed exploration, but not random exploration, was associated with increased pupil dilation and slower responses in a probabilistic learning task, highlighting the importance of differentiating between exploratory subtypes when interpreting pupillometric indices of cognitive engagement The current theoretical landscape presents competing predictions about pupillary responses during exploration-exploitation decisions. While AGT provides a useful framework linking LC-NE activity to exploration behavior, its original formulation does not account for the distinction between directed and random exploration. Moreover, the relationship between pupillary responses and different exploration strategies remains underexplored, particularly regarding how environmental context moderates these associations. The relationship between task engagement, exploration strategies, and pupil size remains theoretically complex. While AGT links exploration to disengagement and larger pupil size, this framework does not differentiate between directed exploration (deliberate information-seeking) and random exploration (stochastic choice variability), nor does it account for how exploitation under different contexts might engage cognitive control. Moreover, previous pupillometry studies have not systematically examined how environmental factors that promote these distinct behavioral strategies modulate pupil-linked arousal. Given that pupil size reliably indexes cognitive effort and LC-NE activity, and that directed exploration, random exploration, and exploitation likely impose different cognitive demands depending on environmental context, we hypothesize that these strategies will elicit distinct pupillary response patterns. The present study therefore aims to clarify when cognitive effort is deployed during exploration-exploitation decisions by examining pupillary responses across environmental contexts that systematically manipulate the factors known to influence exploration behavior: value differences, information asymmetries, and decision horizons. Method Participants: Thirty-five participants between ages 18–41 (24 females, 13 males; M = 25.9, SD = 5.2) took part in the study. A sensitivity power analysis (α = .05, two-tailed, 80% power) indicated that with 35 participants the study was sufficiently powered to detect medium-sized effects (f ≈ 0.25, η²p ≈ .06) 34 , 35 . All had normal or corrected-to-normal vision and no history of neurological or psychiatric disorders. Participants were recruited through volunteer sampling within the academic institution and provided written informed consent in accordance with institutional guidelines. The experimental task lasted approximately 45 minutes, and participants received monetary compensation of $ 20 USD for their participation without any performance-based bonus. The study was approved by the institutional review board (IRB Approval code 252/22). Task and Design Participants completed an adapted version of the Horizon Task 7 , designed to manipulate directed and random exploration. The task is programmed using MATLAB and Psychtoolbox 3.017. The experiment consisted of 320 games, divided into four blocks of 80 games each. Each game contained either 5 or 10 trials. On each trial, participants chose between two slot machine options, each delivering a reward between 1 and 100 points. Rewards were drawn from Gaussian distributions (SD = 8), with the mean value of one option set to either 40 or 60 and the other differing by 4, 8, 12, 20, or 30 points. These differences manipulated the expected value gap, such that smaller differences introduced greater uncertainty and encouraged exploration, while larger differences promoted exploitation. Each game began with four forced-choice trials, during which only one option was available at a time (Fig. 1 A). This phase-controlled participants’ initial information exposure prior to free choice. Three task variables were manipulated: (1) Value gap , the reward difference between options; (2) Information gap , with two conditions, unequal information [1;3], where one option was sampled once and the other three times, and equal information [2;2], where both options were sampled twice; and (3) Choice horizon , defined by the number of subsequent free-choice trials (one vs. six). Together, these manipulations determined whether information from the first free choice could guide future behavior (long horizon) or not (short horizon). Participants were instructed to maximize their total point earnings. After each choice, the reward outcome was displayed on screen, along with a history of choices and rewards for both options within the current game. Accordingly, three key task variables captured distinct exploration- exploitation related dynamics 7 (Fig. 1 B): (1) Value Gap : the absolute difference in expected reward between the two options. Larger gaps encouraged exploitation, whereas smaller gaps increased uncertainty and were hypothesized to elicit more random exploration. (2) Information Gap : the sampling imbalance during forced trials, defined by the number of prior samples from the higher-value option (+ 2 = more samples from the better option, − 2 = more from the worse option, 0 = equal). Lower exposure to the superior option was expected to promote information-seeking via directed exploration. (3) Choice Horizon : the number of remaining free-choice trials (1 vs. 6). A longer horizon was expected to promote directed exploration by increasing the prospective utility of information, while a shorter horizon was expected to encourage exploitation due to limited opportunity to benefit from newly acquired information. This variable was dummy-coded with the long-horizon condition (6 trials) serving as the reference category (coded as 0), and the short-horizon condition (1 trial) coded as 1. Procedure Participants were first provided with on-screen illustrated instructions explaining the task structure, the fixed reward distributions within each game, and the goal of maximizing total points. Each game began with a fixation cross presented at the center of the screen. Throughout the task, pupil size was recorded continuously. To isolate anticipatory cognitive effort associated with exploratory behavior, pupillometry analyses focused exclusively on the first free-choice trial of each game. Pupillometry Recording and Analysis All behavioral and pupillometry data were processed and analyzed following a standardized pipeline. Pupil size data were recorded via the EyeLink 1000 system at a sampling rate of 1000 Hz. Preprocessing included blink detection, linear interpolation, and baseline correction. Trials with missing data were excluded. To ensure measurement accuracy and control for gaze position, the task was programmed to restrict responses unless the participant’s gaze was fixed at the center of the screen at the time of choice. Pupil data were processed and visualized using Data Viewer software (version 4.1.211 Research Ltd., Ontario, Canada), which enabled blink detection verification, baseline correction, segmentation of trials, and extraction of tonic and phasic pupil measures. The primary tonic measure was defined as the average pupil size across the entire decision interval (from choice onset until the motor response) on the first free-choice trial of each game. Phasic analyses involved segmenting each trial into 100-ms bins spanning the 600 ms prior to the motor response (from Bin − 6 to Bin 0) to examine the temporal dynamics of pupil-linked arousal, to capture the decision-related phasic burst, which is known to be time-locked to the motor response 36 . The experimental task was programmed in MATLAB (version 2023a, MathWorks, Natick, MA, USA) using Psychtoolbox-3. All statistical analyses, including mixed-effects regression models for behavioral and pupillometry data, were conducted in SPSS (version 28, IBM Corp., Armonk, NY, USA). Logistic mixed-effects models were used for the behavioral data to predict choice behavior on the first free-choice trial, and linear mixed-effects models were applied to the tonic and phasic pupil measures, with fixed effects of Value Gap, Choice Horizon, and Information Balance, and random intercepts for participants. In this analysis dependent variable is the average pupil size during the period immediately preceding the motor response in the first free-choice trial of each game, reflecting anticipatory arousal. Predictor coding is consistent with the behavioral 1, except for Information gap, which reflects the symmetry of information between options (equal = 0; unequal = 1). Results Behavioral Model: To examine the factors influencing participants' choices on the first free-choice trial of each game, we conducted a logistic mixed-effects regression model. The model predicted the probability of choosing the higher-value option as a function of three fixed-effect predictors Value gap, Choice Horizon and Information gap (Fig. 1 B, see full description in methods). Table 1 Mixed-effects logistic regression results for first free-choice decisions . Predictor Coefficientβ Std. Error Exp (OR) (Coefficient) CI for OR 95% F (1,5279) P value Value Gap (Large) 0.019 0.0034 1.019 [1.012, 1.026] 30.48 < 0.001 Information Gap (Positive) -0.058 0.0223 0.943 [0.903, 0.985] 6.83 0.009 choice horizon (Short) 0.167 0.0623 1.182 [1.046, 1.335] 7.19 0.007 The model revealed significant effects for all predictors. Participants were more likely to choose the higher-value option when the reward difference between options was larger (Value Gap: β = 0.019, p < .001), indicating increased exploitation under conditions of clear expected value. A negative effect of Information Gap (β = − 0.058, p = .009) revealed that participants tended to choose the higher-value option more often when they had received more prior information about it, whereas they were more likely to explore the lower-value option when it had been sampled more frequently that demonstrating a directed exploration strategy aimed at reducing uncertainty. Choice horizon also showed a significant positive effect (β = 0.167, p = .007), indicating that participants exploited more in short-horizon conditions (1 free trial), where the immediate payoff was crucial. Since the short horizon was dummy-coded as 1, the positive coefficient reflects a greater tendency to choose the higher-value option in that condition, compared to the long-horizon condition (6 trials), where exploration could support future gains. Pupil size analysis To investigate the relationship between task features and pupil-linked arousal, we modeled pupil size as a function of three key predictors: Value Gap (the absolute reward difference between options), Choice horizon (short vs. long horizon), and Information gap (whether sampling was equal [2:2] or unequal [1:3]). Notably, while in the behavioral model the information variable was coded directionally to reflect the relative informativeness of the higher-value option, here it was recorded as non-directional (0 = equal, 1 = unequal), since pupil size was not expected to differ based on which option had been sampled more extensively. Subject ID was included as a random effect in all models. Pupil size was analyzed using two complementary approaches: average tonic pupil size during a defined pre-decision window, and time-resolved (trial-by-trial) pupil responses aligned to the decision point. Average (tonic) pupil size: For the tonic analysis, pupil size was averaged across the entire response time (RT) of the first free-choice trial in each game, from the onset of choice presentation until the motor response. This measure captures overall anticipatory arousal during the decision process. (Fig. 1 ). Table 2 Linear mixed-effects model predicting tonic pupil size Predictor Coefficientβ Std. Error CI 95% F (1,5198) P value Value Gap (Large) -0.968 0.544 [-2.03, 0.1] 3.15 0.076 Information gap (unequal) -3.89 10.073 [-23.63, 15.85] 0.15 0.7 choice horizon (Short) 21.87 10.072 [2.12,41.61] 4.71 0.03 The model indicated that choice horizon significantly predicted tonic pupil size (β = 21.87, p = .03). Tonic pupil size was larger in the short-horizon condition (1 free choice, coded as 1) compared to the long-horizon condition (6 free choices, coded as 0), suggesting increased baseline arousal when only a single decision remained and no future opportunities for information use were available. By contrast, neither value gap (β = − 0.97, p = .076) nor information gap (β = − 3.89, p = .70) predicted tonic pupil size significantly. Trial-by-trial binned pupil size : To identify the precise timing of pupillary effects, we conducted a binned-analysis of the pupil responses (i.e., phasic response) during the time leading to the decision point. Specifically, the 700 ms period preceding each motor response was segmented into seven successive 100 ms bins, aligned to the response onset (from Bin − 6 to Bin 0). A separate linear mixed-effects model, similar to the one conducted for the average pupil size in the previous section, was fitted to all participants' pupil data within each time bin, with the same predictor structure as the tonic analysis (value gap, information gap, and choice horizon). This approach yielded a time course of beta (β) coefficients for each predictor, allowing us to identify when during the decision period each factor influenced pupil result. The full statistical tables for each bin are provided in the Supplementary Materials. Value gap was negatively associated with pupil size across the decision window, with β coefficients consistently below zero. These negative values reached significance only in bin − 2 (β = − 1.20, p = .05) and 0 (β = − 1.14, p = 0.048), showing reduced pupil size when the absolute difference between option values was larger. These results do not support a phasic effect of value-gap on pupil size, but a weak and consistent negative effect, similar to the marginal effect observed in the average pupil size analysis. Choice horizon was a significant positive predictor of pupil size at bin − 6 (β = 51.32, p = 0.016) and bin − 1 (β = 21.80, p = 0.039), with a marginally significant effect at bin 0 (β = 18.25, p = 0.08). These results indicate larger pupil size in the short-horizon condition (1 free choice) compared to the long-horizon condition (6 free choices). However, the effects did not manifest as a discrete phasic peak, but rather as a more sustained modulation across the decision window. Information gap was not significantly associated with pupil size at any point in the decision window (all p > 0.05). Coefficients were consistently negative across bins (ranging from β = − 11.70 at bin − 6 to β = 0.034 at bin 0), but none reached statistical significance. These findings indicate that sampling asymmetry (equal vs. unequal information) did not elicit reliable modulation of pupil-linked arousal, and no evidence for a phasic peak was observed. The results from all time-bin analyses show that effects were somewhat more consistent immediately before the response, where they were similar in direction to the effects observed in the average pupil size analysis. This suggests that the observed pupil modulation reflects a sustained cognitive process leading up to the decision, rather than a discrete, transient response to an external stimulus. It also indicates that temporal dynamics along the time leading to decision may contribute to pupil size Discussion The present investigation examined the relationship between pupillary responses and exploration and exploitation behavior. Behavioral analyses confirmed the importance of these factors in influencing participants' choices. Participants adjusted their choices based on absolute value differences, consistently chose options with less prior information in asymmetrical-information conditions, and exhibited greater information-seeking behavior when more free choices remained (six vs. one). These patterns replicated Wilson et al.'s 7 findings and confirmed the effectiveness of the experimental manipulations in eliciting exploration behavior. However, pupillary responses diverged from this behavioral pattern. Average pupil size increased selectively in response to the choice-horizon, with larger dilation observed in short horizon conditions where no subsequent corrections were possible. In long horizons, participants could anticipate future opportunities for information use, supporting directed exploration but without increasing immediate control demands. By contrast, value-uncertainty and information-asymmetry influenced behavior but did not elicit significant pupil modulation. This reveals a dissociation between the drivers of behavioral choice and those of cognitive effort. The absence of significant pupil modulation in response to information gaps challenges the assumption that all forms of directed exploration uniformly engage cognitive control systems. Previous work using the same Horizon Task paradigm has consistently shown that value and information gaps are key drivers of exploratory choices 7 , 37 , yet our findings demonstrate that pupil-indexed effort is insensitive to these same factors. This pattern indicates that cognitive resources are not allocated in response to immediate uncertainty but rather are selectively engaged according to the strategic context and future consequences of the decision. The current findings complement prior evidence linking pupil dilation to exploration. Kozunova et al. 33 , using a probabilistic learning task, demonstrated that directed exploration was associated with larger pupil dilations and response slowing, highlighting the role of pupil-linked arousal in deliberate information seeking. In contrast, the present results obtained with the Horizon Task revealed a selective sensitivity of baseline pupil size to the choice-horizon factor. Specifically, enlarged pupils were observed under short-horizon conditions, which in this paradigm favor exploitation rather than exploration. Taken together, these findings suggest that pupil-linked arousal does not simply differentiate exploration from exploitation as categorical states, but reflects a broader sensitivity to the strategic demands of the decision context This interpretation is consistent with the Expected Value of Control framework, which explains effort allocation in terms of anticipated benefits and costs 24 , and resonates with resource-rational perspectives that view cognitive resources as adaptively allocated according to the expected future utility of information 38 . These findings extend exploration-exploitation frameworks by demonstrating the importance of strategic context. Rather than indexing exploration-specific cognitive demands, pupillary responses appear to reflect the broader strategic context of decision-making- specifically, the consequentiality and finality of choices. This pattern suggests that cognitive control allocation may be guided by decision consequentiality, with effort deployment scaling with the strategic weight (e.g., importance) of individual choices. Trial-by-trial analysis examined the temporal dynamics of pupillary responses. Unlike prior studies reporting phasic pupil dilation near the time of decision, typically associated with transient uncertainty or conflict 39 , 40 , we observed no consistent time-locked modulation. Instead, pupil size remained elevated throughout the decision period without a distinct peak at the choice moment. The stable pupil size observed here, without discrete peaks, contrasts with the phasic bursts typically found in other cognitive tasks 36 , suggesting that task complexity may determine the temporal profile of pupil responses. This interpretation finds some support in early working memory research 41 , where complex information processing similarly showed stable pupil size responses without discrete modulation. The current findings suggest that exploration under choice constraints may engage sustained cognitive processes that warrant further investigation. In conclusion, our findings address the fundamental question of how cognitive effort maps onto distinct behavioral components of exploration and exploitation. The selective pupillary response to choice horizon suggests that cognitive engagement is not uniformly distributed across exploration strategies. Rather than supporting traditional AGT predictions, our results suggest that strategic context determines pupil-indexed engagement, consistent with neurocomputational accounts linking control mechanisms to adaptive exploration under uncertainty 42 . This pattern indicates that cognitive effort allocation depends on the strategic structure of the decision environment, providing a preliminary empirical synthesis between behavioral and physiological approaches to understanding human exploration. Future investigations should examine whether this selective pupillary sensitivity generalizes across decision contexts where choice consequences vary in their strategic importance beyond immediate outcomes. Declarations Author contributions: G.B. contributed to conceptualization, methodology, investigation, formal analysis, and writing - original draft. S.G- contributed to methodology, writing – review & editing and supervision . U.H. contributed to conceptualization formal analysis, methodology, supervision, and writing – review & editing. Data availability statement: Anonymized data and are available at OSF https://osf.io/dq3ke/files/osfstorage Declaration of Interest statement: The authors declare no competing financial or proprietary interests relevant to the content of the article. Funding: The author declares that no funds, grants, or other financial support were received for the conduct of this study References Cohen, J. D., McClure, S. M. & Yu, A. J. Should I stay or should I go? How the human brain manages the trade-off between exploitation and exploration. Philos. Trans. R Soc. Lond. B Biol. Sci. 362 , 933–942. https://doi.org/10.1098/rstb.2007.2098 (2007). Cogliati Dezza, I., Yu, A. J., Cleeremans, A. & Alexander, W. Learning the value of information and reward over time when solving exploration–exploitation problems. Sci. Rep. 7 , 16919. https://doi.org/10.1038/s41598-017-17237-w (2017). Monk, C. T. et al. How ecology shapes exploitation: a framework to predict the behavioural response of human and animal foragers along exploration–exploitation trade-offs. Ecol. Lett. 21 , 779–793. https://doi.org/10.1111/ele.12949 (2018). Berger-Tal, O., Nathan, J., Meron, E. & Saltz, D. The exploration–exploitation dilemma: a multidisciplinary framework. PLoS ONE . 9 , e95693. https://doi.org/10.1371/journal.pone.0095693 (2014). Rich, P. & Gureckis, T. M. The limits of learning: exploration, generalization, and the development of learning traps. J. Exp. Psychol. Gen. 147 , 1553–1570. https://https://doi 10.1037/xge0000466 (2018). Mehlhorn, K. et al. Unpacking the exploration–exploitation tradeoff: a synthesis of human and animal literatures. Decision 2 , 191–215 . Wilson, R. C., Geana, A., White, J. M., Ludvig, E. A. & Cohen, J. D. Humans use directed and random exploration to solve the explore–exploit dilemma. J. Exp. Psychol. Gen. 143 , 2074–2081. https://doi.org/10.1037/a0038199 (2014). Lai, T. L. & Robbins, H. Asymptotically efficient adaptive allocation rules. Adv. Appl. Math. 6 , 4–22. https://doi.org/10.1016/0196-8858(85)90002-8 (1985). Carmel, D. & Markovitch, S. Exploration strategies for model-based learning in multi-agent systems: exploration strategies. Auton. Agents Multi-Agent Syst. 2 , 141–172. https://doi.org/10.1023/A:1010007108196 (1999). Daw, N. D., O’Doherty, J. P., Dayan, P., Seymour, B. & Dolan, R. J. Cortical substrates for exploratory decisions in humans. Nature 441 , 876–879. https://doi.org/10.1038/nature04766 (2006). Aston-Jones, G. & Cohen, J. D. An integrative theory of locus coeruleus–norepinephrine function: adaptive gain and optimal performance. Annu. Rev. Neurosci. 28 , 403–450. https://doi.org/10.1146/annurev.neuro.28.061604.135709 (2005). Jepma, M. & Nieuwenhuis, S. Pupil size predicts changes in the exploration–exploitation trade-off: evidence for the adaptive gain theory. J. Cogn. Neurosci. 23 , 1587–1596. https://doi.org/10.1162/jocn.2010.21548 (2011). Wilson, R. C., Bonawitz, E., Costa, V. D. & Ebitz, R. B. Balancing exploration and exploitation with information and randomization. Curr. Opin. Behav. Sci. 38 , 49–56. https://doi.org/10.1016/j.cobeha.2020.10.001 (2021). Geana, A., Wilson, R. C., Daw, N. & Cohen, J. D. Boredom, information-seeking and exploration. In Proc. 38th Annu. Meet. Cogn. Sci. Soc. 1751–1756 (2016). Zajkowski, W. K., Kossut, M. & Wilson, R. C. A causal role for right frontopolar cortex in directed, but not random, exploration. eLife 6, e27430. (2017). https://doi.org/10.7554/eLife.27430 Schwartenbeck, P. et al. Computational mechanisms of curiosity and goal-directed exploration. eLife 8 , e41703. https://doi.org/10.7554/eLife.41703 (2019). Yu, A. J. & Dayan, P. Uncertainty, neuromodulation, and attention. Neuron 46 , 681–692. https://doi.org/10.1016/j.neuron.2005.04.026 (2005). Behrens, T. E., Woolrich, M. W., Walton, M. E. & Rushworth, M. F. Learning the value of information in an uncertain world. Nat. Neurosci. 10 , 1214–1221. https://doi.org/10.1038/nn1954 (2007). Speekenbrink, M. Chasing unknown bandits: Uncertainty guidance in learning and decision making. Curr. Dir. Psychol. Sci. 31 , 419–427. https://doi.org/10.1177/09637214221105051 (2022). Payzan-LeNestour, É. & Bossaerts, P. Do not bet on the unknown versus try to find out more: estimation uncertainty and unexpected uncertainty both modulate exploration. Front. Neurosci. 6 , 150. https://doi.org/10.3389/fnins.2012.00150 (2012). Courville, A. C., Daw, N. D. & Touretzky, D. S. Bayesian theories of conditioning in a changing world. Trends Cogn. Sci. 10 , 294–300. https://doi.org/10.1016/j.tics.2006.05.004 (2006). Gershman, S. J. & Tzovaras, B. G. Dopaminergic genes are associated with both directed and random exploration. Neuropsychologia 120 , 97–104. https://doi.org/10.1016/j.neuropsychologia.2018.10.009 (2018). Gershman, S. J. Uncertainty and exploration. Decision 6 , 277. https://doi.org/10.1037/dec0000101 (2019). Shenhav, A., Botvinick, M. M. & Cohen, J. D. The expected value of control: an integrative theory of anterior cingulate cortex function. Neuron 79 , 217–240. https://doi.org/10.1016/j.neuron.2013.07.007 (2013). Inzlicht, M., Shenhav, A. & Olivola, C. Y. The effort paradox: Effort is both costly and valued. Trends Cogn. Sci. 22 , 337–349. https://doi.org/10.1016/j.tics.2018.01.007 (2018). Joshi, S., Li, Y., Kalwani, R. M. & Gold, J. I. Relationships between pupil size and neuronal activity in the locus coeruleus, colliculi, and cingulate cortex. Neuron 89 , 221–234. https://doi.org/10.1016/j.neuron.2015.11.028 (2016). Alnæs, D. et al. Pupil size signals mental effort deployed during multiple object tracking and predicts brain activity in the dorsal attention network and the locus coeruleus. J. Vis. 14 , 1. https://doi.org/10.1167/14.4.1 (2014). Preuschoff, K., ’t Hart, B. M. & Einhäuser, W. Pupil dilation signals surprise: evidence for noradrenaline's role in decision making. Front. Neurosci. 5 , 115. https://doi.org/10.3389/fnins.2011.00115 (2011). Nassar, M. R. et al. Rational regulation of learning dynamics by pupil-linked arousal systems. Nat. Neurosci. 15 , 1040–1046. https://doi.org/10.1038/nn.3130 (2012). Zénon, A. Eye pupil signals information gain. Proc. Biol. Sci. 286, 20191593. (2019). https://doi.org/10.1098/rspb.2019.1593 Payzan-LeNestour, E. & Bossaerts, P. Risk, unexpected uncertainty, and estimation uncertainty: Bayesian learning in unstable settings. PLoS Comput. Biol. 7 , e1001048. https://doi.org/10.1371/journal.pcbi.1001048 (2011). Cavanagh, J. F., Figueroa, C. M., Cohen, M. X. & Frank, M. J. Frontal theta reflects uncertainty and unexpectedness during exploration and exploitation. Cereb. Cortex . 22 , 2575–2586. https://doi.org/10.1093/cercor/bhr332 (2012). Kozunova, G. L. et al. Pupil dilation and response slowing distinguish deliberate explorative choices in the probabilistic learning task. Cogn. Affect. Behav. Neurosci. 22 , 1108–1129. https://doi.org/10.3758/s13415-022-00996-z (2022). Cohen, J. Statistical Power Analysis for the Behavioral Sciences (Routledge, 2013). Faul, F., Erdfelder, E., Buchner, A. & Lang, A. G. Statistical power analyses using G*Power 3.1: tests for correlation and regression analyses. Behav. Res. Methods . 41 , 1149–1160. https://doi.org/10.3758/BRM.41.4.1149 (2009). Gabay, S., Pertzov, Y. & Henik, A. Orienting of attention, pupil size, and the norepinephrine system. Atten. Percept. Psychophys . 73 , 123–129. https://doi.org/10.3758/s13414-010-0015-4 (2011). Sadeghiyeh, H. et al. Temporal discounting correlates with directed exploration but not with random exploration. Sci. Rep. 10 , 4020. https://doi.org/10.1038/s41598-020-60576-4 (2020). Lieder, F. & Griffiths, T. L. Resource-rational analysis: understanding human cognition as the optimal use of limited computational resources. Behav. Brain Sci. 43 , e1. https://doi.org/10.1017/S0140525X1900061X (2019). Murphy, P. R., Robertson, I. H., Balsters, J. H. & O’Connell, R. G. Pupillometry and P3 index the locus coeruleus-noradrenergic arousal function in humans. Psychophysiology 48 , 1532–1543. https://doi.org/10.1111/j.1469-8986.2011.01226.x (2011). van Steenbergen, H. & Band, G. P. Pupil dilation in the Simon task as a marker of conflict processing. Front. Hum. Neurosci. 7 , 215. https://doi.org/10.3389/fnhum.2013.00215 (2013). Granholm, E., Asarnow, R. F., Sarkin, A. J. & Dykes, K. L. Pupillary responses index cognitive resource limitations. Psychophysiology 33 , 457–461. https://doi.org/10.1111/j.1469-8986.1996.tb01071.x (1996). Badre, D., Doll, B. B., Long, N. M. & Frank, M. J. Rostrolateral prefrontal cortex and individual differences in uncertainty-driven exploration. Neuron 73 , 595–607. https://doi.org/10.1016/j.neuron.2011.12.025 (2012). Additional Declarations No competing interests reported. Supplementary Files supplementarymaterials.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-7740172\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":true,\"archivedVersions\":[],\"articleType\":\"Article\",\"associatedPublications\":[],\"authors\":[{\"id\":533051325,\"identity\":\"c568796a-59dd-457b-8db4-4416ccd9059a\",\"order_by\":0,\"name\":\"Gili Barkay\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/UlEQVRIiWNgGAWjYDACZuaGAwwMQMQD5HwAYjZ2gloYEVoYZ4C0MBO0hrGBAaaFmQdsCAEN8u6MjQd+MNyRN+85/Eza5tc2eT5mBsYPH3NwazE8zNhwsIfhmeGcs21m0rl9tw3bmBmYJWduw6OlGegXHobDjDP4GYBaem4zArWwMfMS0HLwD8Nh+xn87N+kLXtu2xPUIg8MscNAWxJn8PaYSTP8uJ1IUIsBSIuMwbPkGTxnii17G24ntzEzNuP1i3z/4cMf31TcsZ3Bk77xxo8/t23ntzcf/PARny0HwCSYzSLB2AaiwTGFxxYkaeYPDH/wKh4Fo2AUjIIRCgDn3VItRYP2YQAAAABJRU5ErkJggg==\",\"orcid\":\"\",\"institution\":\"University of Haifa\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Gili\",\"middleName\":\"\",\"lastName\":\"Barkay\",\"suffix\":\"\"},{\"id\":533051326,\"identity\":\"43c9a9f4-b9ff-4035-b6cd-93d6bddd0567\",\"order_by\":1,\"name\":\"Shai Gabay\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"University of Haifa\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Shai\",\"middleName\":\"\",\"lastName\":\"Gabay\",\"suffix\":\"\"},{\"id\":533051327,\"identity\":\"9908b8da-349b-4aa1-8818-e5c3f13044b0\",\"order_by\":2,\"name\":\"Uri Herz\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"University of Haifa\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Uri\",\"middleName\":\"\",\"lastName\":\"Herz\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2025-09-29 09:23:17\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-7740172/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-7740172/v1\",\"draftVersion\":[],\"editorialEvents\":[],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":94368896,\"identity\":\"024ba7fb-09eb-4863-889f-dcbf327a5fb6\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:19\",\"extension\":\"docx\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":765887,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"SelectivePupilSizeResponseWithindirectandrandomexplorationandexploitationBehaviorsScientificReports04102025GB.docx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/6e2a8104989b8d092eabd0aa.docx\"},{\"id\":94368973,\"identity\":\"21f8daf5-e24b-4b57-a934-dd5d8ff3100e\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:32\",\"extension\":\"json\",\"order_by\":1,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":5376,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"f42be2edb3304c4db3cde2b14220a20a.json\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/23c23bd6f011e663296bac82.json\"},{\"id\":94368988,\"identity\":\"83c3fe98-b2c2-4d7b-845e-7c9ef6eb8744\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:34\",\"extension\":\"docx\",\"order_by\":2,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":37881,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"supplementarymaterials.docx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/b61b3cd259f614e1d6a62111.docx\"},{\"id\":94369190,\"identity\":\"df4305a9-58ad-425a-97f1-7786b8483cef\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:59\",\"extension\":\"xml\",\"order_by\":3,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":108960,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"f42be2edb3304c4db3cde2b14220a20a1enriched.xml\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/cd416e4bc458a68474fc7943.xml\"},{\"id\":94368886,\"identity\":\"ee5e93a9-cd98-4406-a080-23e631a6e0d3\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:18\",\"extension\":\"png\",\"order_by\":4,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":147187,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"floatimage1.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/373bb635178213800e87786b.png\"},{\"id\":94368738,\"identity\":\"415243fb-f0b5-4849-8c38-a9085cae12c5\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:01\",\"extension\":\"png\",\"order_by\":5,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":557688,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"floatimage2.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/c9ed03616e64a5329db310de.png\"},{\"id\":94369074,\"identity\":\"dece5339-5fa1-4397-b54c-34646449651a\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:43\",\"extension\":\"png\",\"order_by\":6,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":34414,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"Onlinefloatimage1.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/9a7383680b3a9424a7176c2d.png\"},{\"id\":94368918,\"identity\":\"af452b4a-a4ab-49c4-8216-f9a4a20cb861\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:24\",\"extension\":\"png\",\"order_by\":7,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":45320,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"Onlinefloatimage2.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/553d381ef1260c5113513284.png\"},{\"id\":94489305,\"identity\":\"795e49b2-2932-48c2-9eb8-785838620139\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 17:04:08\",\"extension\":\"xml\",\"order_by\":8,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":108334,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"f42be2edb3304c4db3cde2b14220a20a1structuring.xml\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/17c6be7abdd9ab22abf39a93.xml\"},{\"id\":94368732,\"identity\":\"18d1a292-77a2-403f-9a9e-c343ca4d0106\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:15:59\",\"extension\":\"html\",\"order_by\":9,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"acdc-reference\",\"size\":119691,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"earlyproof.html\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/83b757af0bca3e6c9532a660.html\"},{\"id\":94369204,\"identity\":\"8d71b30a-6aee-4773-9b78-0a7f0193ee56\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:17:04\",\"extension\":\"png\",\"order_by\":1,\"title\":\"Figure 1\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":169764,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003e\\u003cstrong\\u003eExperimental Design\\u003c/strong\\u003e: \\u0026nbsp;A. Example trial sequence in the Horizon Task. The first four trials were forced-choice trials (left), where one choice was marked with X and the participants had to choose the other option and see its outcome. In free-choice trials a green square indicated that participants could freely choose between the two options. The selected option’s reward was then displayed (right). B. Illustration of task paradigm and variables: The paradigm included 320 mini-games, that differed in three variables: Value Gap – the absolute difference in expected reward between options, Information Gap - imbalance in the number of samples from each option during the forced-choice session, and Choice Horizon - the number of remaining free-choice trials after the forced choice session. Green boxes indicate the first free choice, which was the focus of our analysis.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"floatimage1.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/87b7be27616b4ae4e9f316b0.png\"},{\"id\":94369076,\"identity\":\"1339e956-5e5a-4b9d-978a-5f1f1d3da361\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:16:43\",\"extension\":\"png\",\"order_by\":2,\"title\":\"Figure 2\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":201950,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003e\\u003cstrong\\u003eTrial-by-Trial Analysis of Pupil Size Across Pre-Decision Time Bins. \\u003c/strong\\u003eThe beta coefficients for predictors of pupil size in the 700 ms preceding the first free-choice decision, analyzed across seven consecutive 100 ms bins (Bin –6 to Bin 0) aligned to the motor response. Beta coefficients from linear mixed-effects models are shown in circles (connected with lines) for Value Gap (blue), choice horizon (red) and information gap (green). Shaded areas represent the standard error (SE), *p\\u0026lt;0.05, ^ marginal significance p\\u0026lt;0.08.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"floatimage2.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/a5806f228c2f4037e89535ba.png\"},{\"id\":95889518,\"identity\":\"713d8608-7989-4490-81e4-6da495370aac\",\"added_by\":\"auto\",\"created_at\":\"2025-11-14 05:53:45\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":1106996,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"manuscript.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/1213ac1a-becf-4afe-a0b1-f62b59109349.pdf\"},{\"id\":94369201,\"identity\":\"a2a772d8-cf1f-45d6-826c-981afa26f858\",\"added_by\":\"auto\",\"created_at\":\"2025-10-27 13:17:03\",\"extension\":\"docx\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"supplement\",\"size\":37881,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"supplementarymaterials.docx\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7740172/v1/83ec5b784df0e33b30df3978.docx\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"Selective Pupil Size Response Within direct and random exploration and exploitation Behaviors\",\"fulltext\":[{\"header\":\"Introduction\",\"content\":\"\\u003cp\\u003eDecision-makers face a fundamental challenge: should they exploit the best familiar option or explore uncertain alternatives that might yield better long-term outcomes? \\u003csup\\u003e1,2\\u003c/sup\\u003e. This trade-off manifests across species and contexts, from animal foraging behaviors to human strategic planning \\u003csup\\u003e\\u003cspan additionalcitationids=\\\"CR4\\\" citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e\\u003c/sup\\u003e, yet questions persist regarding the psychological and neurobiological mechanisms underlying these competing drives\\u003csup\\u003e\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR6\\\" class=\\\"CitationRef\\\"\\u003e6\\u003c/span\\u003e\\u003c/sup\\u003e. Human exploration can take the form of directed exploration, involving deliberate information-seeking, or random exploration, reflecting stochastic variability in choice behavior\\u003csup\\u003e\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e\\u003c/sup\\u003e. The present study addresses these processes by examining pupillary responses, a reliable physiological marker of locus coeruleus\\u0026ndash;noradrenergic activity and task engagement, across different forms of exploration.\\u003c/p\\u003e\\u003cp\\u003eFormulations of reinforcement learning (RL), such as those grounded in the multi-armed bandit framework, modeled exploration as stochastic variability in action selection, typically relying on probabilistic rules to balance exploratory and exploitative choices\\u003csup\\u003e\\u003cspan additionalcitationids=\\\"CR9\\\" citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e\\u003c/sup\\u003e. This framing implicitly positions exploration and exploitation as distinct and competing behavioral modes. In contrast, optimal foraging theory conceives of exploration as a purposeful, information-driven process aimed at maximizing long-term reward\\u003csup\\u003e\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e\\u003c/sup\\u003e. This conceptual divergence highlights the broader issue of whether exploration reflects random noise in choice behavior or a purposeful strategy grounded in cognitive control.\\u003c/p\\u003e\\u003cp\\u003eA second influential framework, the Adaptive Gain Theory (AGT), links exploration\\u0026ndash;exploitation behavior to task engagement and its neurobiological substrates. AGT posits that shifts in locus coeruleus\\u0026ndash;norepinephrine (LC\\u0026ndash;NE) activity govern exploration and exploitation, with moderate tonic activity supporting exploitation (focused, reward-driven performance) and elevated tonic states reflecting disengagement and a shift toward exploration \\u003csup\\u003e\\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e11\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e\\u003cp\\u003eThis account provides direct predictions about pupil size during exploration-exploitation decisions, with exploration corresponding to reduced cognitive control and larger pupil size. However, subsequent evidence has revealed that human exploration comprises two functionally distinct processes; directed and random exploration that challenging this dichotomy and highlighting tensions between RL and AGT frameworks\\u003csup\\u003e\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e,\\u003cspan additionalcitationids=\\\"CR14 CR15\\\" citationid=\\\"CR13\\\" class=\\\"CitationRef\\\"\\u003e13\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR16\\\" class=\\\"CitationRef\\\"\\u003e16\\u003c/span\\u003e\\u003c/sup\\u003e. Beyond this dichotomy, research has emphasized the role of environmental uncertainty in shaping exploration and exploitation. Yu and Dayan\\u003csup\\u003e\\u003cspan citationid=\\\"CR17\\\" class=\\\"CitationRef\\\"\\u003e17\\u003c/span\\u003e\\u003c/sup\\u003e distinguished between expected uncertainty, reflecting stable but noisy environments, and unexpected uncertainty, signaling change-points that demand adaptation. Expected uncertainty arises in familiar contexts with predictable variability, whereas unexpected uncertainty emerges when outcomes violate predictions, prompting active information-seeking\\u003csup\\u003e\\u003cspan citationid=\\\"CR17\\\" class=\\\"CitationRef\\\"\\u003e17\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR18\\\" class=\\\"CitationRef\\\"\\u003e18\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e\\u003cp\\u003eEvidence from multi-armed bandit tasks confirms that humans actively seek uncertain options when learning is possible, with uncertainty guiding exploration in principled ways\\u003csup\\u003e\\u003cspan citationid=\\\"CR19\\\" class=\\\"CitationRef\\\"\\u003e19\\u003c/span\\u003e\\u003c/sup\\u003e. In this context, Payzan-LeNestour and Bossaerts\\u003csup\\u003e\\u003cspan citationid=\\\"CR20\\\" class=\\\"CitationRef\\\"\\u003e20\\u003c/span\\u003e\\u003c/sup\\u003e showed that unexpected uncertainty promotes novelty-seeking (directed exploration), whereas estimation uncertainty can trigger ambiguity aversion rather than random exploration. Together, these accounts suggest that exploratory behavior is not monolithic but shaped by the interplay of potential learning benefits and perceived risk\\u003csup\\u003e\\u003cspan additionalcitationids=\\\"CR22\\\" citationid=\\\"CR21\\\" class=\\\"CitationRef\\\"\\u003e21\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR23\\\" class=\\\"CitationRef\\\"\\u003e23\\u003c/span\\u003e\\u003c/sup\\u003e. These perspectives highlight the importance of understanding cognitive effort allocation during exploration and exploitation. Another theory about engagement is the Expected Value of Control framework (EVC) \\u003csup\\u003e\\u003cspan citationid=\\\"CR24\\\" class=\\\"CitationRef\\\"\\u003e24\\u003c/span\\u003e\\u003c/sup\\u003e proposes that cognitive control is allocated by weighing expected benefits against intrinsic costs. Importantly, effort\\u0026rsquo;s relationship with value is more complex than simple cost models predict, as decisions are most demanding when competing options are similar in value\\u003csup\\u003e\\u003cspan citationid=\\\"CR24\\\" class=\\\"CitationRef\\\"\\u003e24\\u003c/span\\u003e\\u003c/sup\\u003e and effort itself may become intrinsically rewarding\\u003csup\\u003e\\u003cspan citationid=\\\"CR25\\\" class=\\\"CitationRef\\\"\\u003e25\\u003c/span\\u003e\\u003c/sup\\u003e. While EVC has not been directly applied to pupillometry, its principles suggest that cognitive engagement should be scaled with decision importance and strategic value rather than exploration type per se.\\u003c/p\\u003e\\u003cp\\u003eThus, the intensity of cognitive engagement should scale with both anticipated benefits and the learned value of effortful engagement, offering a theoretical bridge between exploratory strategies and their physiological signatures.\\u003c/p\\u003e\\u003cp\\u003eWithin this context, pupillometry provides a natural candidate for testing predictions about control allocation, as pupil size reliably reflects LC\\u0026ndash;NE activity, cognitive effort and task engagement\\u003csup\\u003e\\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e11\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR26\\\" class=\\\"CitationRef\\\"\\u003e26\\u003c/span\\u003e\\u003c/sup\\u003e. According to AGT, exploration should correspond to the tonic LC mode, characterized by elevated baseline activity and larger pupil diameter, reflecting disengagement from the current task and search for alternatives. Early support for this relationship came from Aston-Jones and Cohen's human studies\\u003csup\\u003e\\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e11\\u003c/span\\u003e\\u003c/sup\\u003e, which demonstrated that baseline pupil diameter increased progressively as task utility declined in an auditory discrimination task, reaching maximum values when participants chose to abandon the current task series. Consistent with these AGT predictions, Jepma and Nieuwenhuis\\u003csup\\u003e\\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e\\u003c/sup\\u003e found larger baseline pupil size preceding exploratory choices compared to exploitative choices in a multi-armed bandit task. However, these consistent findings of pupillary dilation during exploration align with broader research demonstrating that pupil dilation reliably indexes cognitively demanding processes and mental effort deployment\\u003csup\\u003e\\u003cspan citationid=\\\"CR27\\\" class=\\\"CitationRef\\\"\\u003e27\\u003c/span\\u003e\\u003c/sup\\u003e, including uncertainty processing and surprise\\u003csup\\u003e\\u003cspan citationid=\\\"CR28\\\" class=\\\"CitationRef\\\"\\u003e28\\u003c/span\\u003e\\u003c/sup\\u003e and learning rate adjustments\\u003csup\\u003e\\u003cspan citationid=\\\"CR29\\\" class=\\\"CitationRef\\\"\\u003e29\\u003c/span\\u003e\\u003c/sup\\u003e. Moreover, pupil responses may reflect the informational relevance of outcomes, scaling with their potential to update internal models\\u003csup\\u003e\\u003cspan citationid=\\\"CR30\\\" class=\\\"CitationRef\\\"\\u003e30\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e\\u003cp\\u003eHowever, this raises new questions about when exploration requires cognitive effort. Several investigations have found limited evidence for directed exploration behavior in humans, potentially due to methodological limitations that make it difficult to distinguish between these exploration strategies\\u003csup\\u003e\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR31\\\" class=\\\"CitationRef\\\"\\u003e31\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR32\\\" class=\\\"CitationRef\\\"\\u003e32\\u003c/span\\u003e\\u003c/sup\\u003e. To address these methodological constraints, Wilson and colleagues\\u003csup\\u003e\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e\\u003c/sup\\u003e developed the Horizon Task, which systematically controls participants' information regarding option payoffs through forced-choice trials that decorrelate reward and information. This paradigm enables systematic manipulation of uncertainty levels, information exposure, and future decision opportunities, providing an optimal framework for examining how environmental contexts influence the cognitive demands of different exploratory strategies. This theoretical distinction has also received empirical support. Kozunova et al. (2022)\\u003csup\\u003e\\u003cspan citationid=\\\"CR33\\\" class=\\\"CitationRef\\\"\\u003e33\\u003c/span\\u003e\\u003c/sup\\u003e demonstrated that directed exploration, but not random exploration, was associated with increased pupil dilation and slower responses in a probabilistic learning task, highlighting the importance of differentiating between exploratory subtypes when interpreting pupillometric indices of cognitive engagement\\u003c/p\\u003e\\u003cp\\u003eThe current theoretical landscape presents competing predictions about pupillary responses during exploration-exploitation decisions. While AGT provides a useful framework linking LC-NE activity to exploration behavior, its original formulation does not account for the distinction between directed and random exploration. Moreover, the relationship between pupillary responses and different exploration strategies remains underexplored, particularly regarding how environmental context moderates these associations.\\u003c/p\\u003e\\u003cp\\u003eThe relationship between task engagement, exploration strategies, and pupil size remains theoretically complex. While AGT links exploration to disengagement and larger pupil size, this framework does not differentiate between directed exploration (deliberate information-seeking) and random exploration (stochastic choice variability), nor does it account for how exploitation under different contexts might engage cognitive control. Moreover, previous pupillometry studies have not systematically examined how environmental factors that promote these distinct behavioral strategies modulate pupil-linked arousal. Given that pupil size reliably indexes cognitive effort and LC-NE activity, and that directed exploration, random exploration, and exploitation likely impose different cognitive demands depending on environmental context, we hypothesize that these strategies will elicit distinct pupillary response patterns. The present study therefore aims to clarify when cognitive effort is deployed during exploration-exploitation decisions by examining pupillary responses across environmental contexts that systematically manipulate the factors known to influence exploration behavior: value differences, information asymmetries, and decision horizons.\\u003c/p\\u003e\"},{\"header\":\"Method\",\"content\":\"\\u003cdiv id=\\\"Sec3\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003eParticipants:\\u003c/h2\\u003e\\u003cp\\u003eThirty-five participants between ages 18\\u0026ndash;41 (24 females, 13 males; \\u003cem\\u003eM\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;25.9, \\u003cem\\u003eSD\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;5.2) took part in the study. A sensitivity power analysis (α\\u0026thinsp;=\\u0026thinsp;.05, two-tailed, 80% power) indicated that with 35 participants the study was sufficiently powered to detect medium-sized effects (f\\u0026thinsp;\\u0026asymp;\\u0026thinsp;0.25, η\\u0026sup2;p\\u0026thinsp;\\u0026asymp;\\u0026thinsp;.06)\\u003csup\\u003e\\u003cspan citationid=\\\"CR34\\\" class=\\\"CitationRef\\\"\\u003e34\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR35\\\" class=\\\"CitationRef\\\"\\u003e35\\u003c/span\\u003e\\u003c/sup\\u003e. All had normal or corrected-to-normal vision and no history of neurological or psychiatric disorders. Participants were recruited through volunteer sampling within the academic institution and provided written informed consent in accordance with institutional guidelines. The experimental task lasted approximately 45 minutes, and participants received monetary compensation of \\u003cspan\\u003e$\\u003c/span\\u003e20 USD for their participation without any performance-based bonus. The study was approved by the institutional review board (IRB Approval code 252/22).\\u003c/p\\u003e\\u003c/div\\u003e\\n\\u003ch3\\u003eTask and Design\\u003c/h3\\u003e\\n\\u003cp\\u003eParticipants completed an adapted version of the Horizon Task\\u003csup\\u003e\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e\\u003c/sup\\u003e, designed to manipulate directed and random exploration. The task is programmed using MATLAB and Psychtoolbox 3.017. The experiment consisted of 320 games, divided into four blocks of 80 games each. Each game contained either 5 or 10 trials. On each trial, participants chose between two slot machine options, each delivering a reward between 1 and 100 points. Rewards were drawn from Gaussian distributions (SD\\u0026thinsp;=\\u0026thinsp;8), with the mean value of one option set to either 40 or 60 and the other differing by 4, 8, 12, 20, or 30 points. These differences manipulated the expected value gap, such that smaller differences introduced greater uncertainty and encouraged exploration, while larger differences promoted exploitation.\\u003c/p\\u003e\\u003cp\\u003eEach game began with four forced-choice trials, during which only one option was available at a time (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003eA). This phase-controlled participants\\u0026rsquo; initial information exposure prior to free choice. Three task variables were manipulated: \\u003cb\\u003e(1) Value gap\\u003c/b\\u003e, the reward difference between options; \\u003cb\\u003e(2) Information gap\\u003c/b\\u003e, with two conditions, unequal information [1;3], where one option was sampled once and the other three times, and equal information [2;2], where both options were sampled twice; and \\u003cb\\u003e(3) Choice horizon\\u003c/b\\u003e, defined by the number of subsequent free-choice trials (one vs. six). Together, these manipulations determined whether information from the first free choice could guide future behavior (long horizon) or not (short horizon). Participants were instructed to maximize their total point earnings. After each choice, the reward outcome was displayed on screen, along with a history of choices and rewards for both options within the current game. Accordingly, three key task variables captured distinct exploration- exploitation related dynamics\\u003csup\\u003e\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e\\u003c/sup\\u003e (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003eB):\\u003c/p\\u003e\\u003cp\\u003e(1) \\u003cb\\u003eValue Gap\\u003c/b\\u003e: the absolute difference in expected reward between the two options. Larger gaps encouraged exploitation, whereas smaller gaps increased uncertainty and were hypothesized to elicit more random exploration.\\u003c/p\\u003e\\u003cp\\u003e(2) \\u003cb\\u003eInformation Gap\\u003c/b\\u003e: the sampling imbalance during forced trials, defined by the number of prior samples from the higher-value option (+\\u0026thinsp;2\\u0026thinsp;=\\u0026thinsp;more samples from the better option, \\u0026minus;\\u0026thinsp;2\\u0026thinsp;=\\u0026thinsp;more from the worse option, 0\\u0026thinsp;=\\u0026thinsp;equal). Lower exposure to the superior option was expected to promote information-seeking via directed exploration.\\u003c/p\\u003e\\u003cp\\u003e(3) \\u003cb\\u003eChoice Horizon\\u003c/b\\u003e: the number of remaining free-choice trials (1 vs. 6). A longer horizon was expected to promote directed exploration by increasing the prospective utility of information, while a shorter horizon was expected to encourage exploitation due to limited opportunity to benefit from newly acquired information. This variable was dummy-coded with the \\u003cem\\u003elong-horizon\\u003c/em\\u003e condition (6 trials) serving as the reference category (coded as 0), and the \\u003cem\\u003eshort-horizon\\u003c/em\\u003e condition (1 trial) coded as 1.\\u003c/p\\u003e\\u003cp\\u003e\\u003c/p\\u003e\\n\\u003ch3\\u003eProcedure\\u003c/h3\\u003e\\n\\u003cp\\u003eParticipants were first provided with on-screen illustrated instructions explaining the task structure, the fixed reward distributions within each game, and the goal of maximizing total points. Each game began with a fixation cross presented at the center of the screen. Throughout the task, pupil size was recorded continuously. To isolate anticipatory cognitive effort associated with exploratory behavior, pupillometry analyses focused exclusively on the first free-choice trial of each game.\\u003c/p\\u003e\\n\\u003ch3\\u003ePupillometry Recording and Analysis\\u003c/h3\\u003e\\n\\u003cp\\u003eAll behavioral and pupillometry data were processed and analyzed following a standardized pipeline. Pupil size data were recorded via the EyeLink 1000 system at a sampling rate of 1000 Hz. Preprocessing included blink detection, linear interpolation, and baseline correction. Trials with missing data were excluded. To ensure measurement accuracy and control for gaze position, the task was programmed to restrict responses unless the participant\\u0026rsquo;s gaze was fixed at the center of the screen at the time of choice.\\u003c/p\\u003e\\u003cp\\u003ePupil data were processed and visualized using Data Viewer software (version 4.1.211 Research Ltd., Ontario, Canada), which enabled blink detection verification, baseline correction, segmentation of trials, and extraction of tonic and phasic pupil measures. The primary tonic measure was defined as the average pupil size across the entire decision interval (from choice onset until the motor response) on the first free-choice trial of each game. Phasic analyses involved segmenting each trial into 100-ms bins spanning the 600 ms prior to the motor response (from Bin \\u0026minus;\\u0026thinsp;6 to Bin 0) to examine the temporal dynamics of pupil-linked arousal, to capture the decision-related phasic burst, which is known to be time-locked to the motor response\\u003csup\\u003e\\u003cspan citationid=\\\"CR36\\\" class=\\\"CitationRef\\\"\\u003e36\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e\\u003cp\\u003eThe experimental task was programmed in MATLAB (version 2023a, MathWorks, Natick, MA, USA) using Psychtoolbox-3. All statistical analyses, including mixed-effects regression models for behavioral and pupillometry data, were conducted in SPSS (version 28, IBM Corp., Armonk, NY, USA). Logistic mixed-effects models were used for the behavioral data to predict choice behavior on the first free-choice trial, and linear mixed-effects models were applied to the tonic and phasic pupil measures, with fixed effects of Value Gap, Choice Horizon, and Information Balance, and random intercepts for participants. In this analysis dependent variable is the average pupil size during the period immediately preceding the motor response in the first free-choice trial of each game, reflecting anticipatory arousal. Predictor coding is consistent with the behavioral 1, except for Information gap, which reflects the symmetry of information between options (equal\\u0026thinsp;=\\u0026thinsp;0; unequal\\u0026thinsp;=\\u0026thinsp;1).\\u003c/p\\u003e\"},{\"header\":\"Results\",\"content\":\"\\u003cdiv id=\\\"Sec8\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003eBehavioral Model:\\u003c/h2\\u003e\\u003cp\\u003eTo examine the factors influencing participants' choices on the first free-choice trial of each game, we conducted a logistic mixed-effects regression model. The model predicted the probability of choosing the higher-value option as a function of three fixed-effect predictors Value gap, Choice Horizon and Information gap (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003eB, see full description in methods).\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab1\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 1\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eMixed-effects logistic regression results for first free-choice decisions\\u003c/b\\u003e.\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"7\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c5\\\" colnum=\\\"5\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c6\\\" colnum=\\\"6\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c7\\\" colnum=\\\"7\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cem\\u003ePredictor\\u003c/em\\u003e\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eCoefficientβ\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eStd. Error\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eExp (OR) (Coefficient)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eCI for OR 95%\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eF (1,5279)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eP value\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eValue Gap (Large)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.019\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.0034\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e1.019\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e[1.012, 1.026]\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e30.48\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e\\u0026lt;\\u0026thinsp;0.001\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eInformation Gap (Positive)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e-0.058\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.0223\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.943\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e[0.903, 0.985]\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e6.83\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e0.009\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003echoice horizon (Short)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.167\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.0623\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e1.182\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e[1.046, 1.335]\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e7.19\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e0.007\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003eThe model revealed significant effects for all predictors. Participants were more likely to choose the higher-value option when the reward difference between options was larger (Value Gap: β\\u0026thinsp;=\\u0026thinsp;0.019, \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;\\u0026lt;\\u0026thinsp;.001), indicating increased exploitation under conditions of clear expected value. A negative effect of Information Gap (β = \\u0026minus;\\u0026thinsp;0.058, \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;.009) revealed that participants tended to choose the higher-value option more often when they had received more prior information about it, whereas they were more likely to explore the lower-value option when it had been sampled more frequently that demonstrating a directed exploration strategy aimed at reducing uncertainty. Choice horizon also showed a significant positive effect (β\\u0026thinsp;=\\u0026thinsp;0.167, \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;.007), indicating that participants exploited more in short-horizon conditions (1 free trial), where the immediate payoff was crucial. Since the short horizon was dummy-coded as 1, the positive coefficient reflects a greater tendency to choose the higher-value option in that condition, compared to the long-horizon condition (6 trials), where exploration could support future gains.\\u003c/p\\u003e\\u003c/div\\u003e\\n\\u003ch3\\u003ePupil size analysis\\u003c/h3\\u003e\\n\\u003cp\\u003eTo investigate the relationship between task features and pupil-linked arousal, we modeled pupil size as a function of three key predictors: Value Gap (the absolute reward difference between options), Choice horizon (short vs. long horizon), and Information gap (whether sampling was equal [2:2] or unequal [1:3]). Notably, while in the behavioral model the information variable was coded directionally to reflect the relative informativeness of the higher-value option, here it was recorded as non-directional (0\\u0026thinsp;=\\u0026thinsp;equal, 1\\u0026thinsp;=\\u0026thinsp;unequal), since pupil size was not expected to differ based on which option had been sampled more extensively. Subject ID was included as a random effect in all models. Pupil size was analyzed using two complementary approaches: average tonic pupil size during a defined pre-decision window, and time-resolved (trial-by-trial) pupil responses aligned to the decision point.\\u003c/p\\u003e\\n\\u003ch3\\u003eAverage (tonic) pupil size:\\u003c/h3\\u003e\\n\\u003cp\\u003eFor the tonic analysis, pupil size was averaged across the entire response time (RT) of the first free-choice trial in each game, from the onset of choice presentation until the motor response. This measure captures overall anticipatory arousal during the decision process. (Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e).\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab2\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 2\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eLinear mixed-effects model predicting tonic pupil size\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"6\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c5\\\" colnum=\\\"5\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c6\\\" colnum=\\\"6\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cem\\u003ePredictor\\u003c/em\\u003e\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eCoefficientβ\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eStd. Error\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eCI 95%\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eF (1,5198)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eP value\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eValue Gap (Large)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e-0.968\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.544\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e[-2.03, 0.1]\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e3.15\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.076\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eInformation gap (unequal)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e-3.89\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e10.073\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e[-23.63, 15.85]\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.15\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.7\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003echoice horizon (Short)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e21.87\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e10.072\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e[2.12,41.61]\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e4.71\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003e0.03\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003eThe model indicated that choice horizon significantly predicted tonic pupil size (β\\u0026thinsp;=\\u0026thinsp;21.87, p\\u0026thinsp;=\\u0026thinsp;.03). Tonic pupil size was larger in the short-horizon condition (1 free choice, coded as 1) compared to the long-horizon condition (6 free choices, coded as 0), suggesting increased baseline arousal when only a single decision remained and no future opportunities for information use were available. By contrast, neither value gap (β = \\u0026minus;\\u0026thinsp;0.97, p\\u0026thinsp;=\\u0026thinsp;.076) nor information gap (β = \\u0026minus;\\u0026thinsp;3.89, p\\u0026thinsp;=\\u0026thinsp;.70) predicted tonic pupil size significantly.\\u003c/p\\u003e\\u003cdiv id=\\\"Sec11\\\" class=\\\"Section2\\\"\\u003e\\u003ch2\\u003e\\u003cb\\u003eTrial-by-trial binned pupil size\\u003c/b\\u003e:\\u003c/h2\\u003e\\u003cp\\u003eTo identify the precise timing of pupillary effects, we conducted a binned-analysis of the pupil responses (i.e., phasic response) during the time leading to the decision point. Specifically, the 700 ms period preceding each motor response was segmented into seven successive 100 ms bins, aligned to the response onset (from Bin \\u0026minus;\\u0026thinsp;6 to Bin 0). A separate linear mixed-effects model, similar to the one conducted for the average pupil size in the previous section, was fitted to all participants' pupil data within each time bin, with the same predictor structure as the tonic analysis (value gap, information gap, and choice horizon). This approach yielded a time course of beta (β) coefficients for each predictor, allowing us to identify when during the decision period each factor influenced pupil result. The full statistical tables for each bin are provided in the Supplementary Materials.\\u003c/p\\u003e\\u003cp\\u003e\\u003c/p\\u003e\\u003cp\\u003eValue gap was negatively associated with pupil size across the decision window, with β coefficients consistently below zero. These negative values reached significance only in bin \\u0026minus;\\u0026thinsp;2 (β = \\u0026minus;\\u0026thinsp;1.20, \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;.05) and 0 (β = \\u0026minus;\\u0026thinsp;1.14, \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;0.048), showing reduced pupil size when the absolute difference between option values was larger. These results do not support a phasic effect of value-gap on pupil size, but a weak and consistent negative effect, similar to the marginal effect observed in the average pupil size analysis.\\u003c/p\\u003e\\u003cp\\u003eChoice horizon was a significant positive predictor of pupil size at bin \\u0026minus;\\u0026thinsp;6 (β\\u0026thinsp;=\\u0026thinsp;51.32, \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;0.016) and bin \\u0026minus;\\u0026thinsp;1 (β\\u0026thinsp;=\\u0026thinsp;21.80, \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;0.039), with a marginally significant effect at bin 0 (β\\u0026thinsp;=\\u0026thinsp;18.25, \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;=\\u0026thinsp;0.08). These results indicate larger pupil size in the short-horizon condition (1 free choice) compared to the long-horizon condition (6 free choices). However, the effects did not manifest as a discrete phasic peak, but rather as a more sustained modulation across the decision window.\\u003c/p\\u003e\\u003cp\\u003eInformation gap was not significantly associated with pupil size at any point in the decision window (all \\u003cem\\u003ep\\u003c/em\\u003e\\u0026thinsp;\\u0026gt;\\u0026thinsp;0.05). Coefficients were consistently negative across bins (ranging from β = \\u0026minus;\\u0026thinsp;11.70 at bin \\u0026minus;\\u0026thinsp;6 to β\\u0026thinsp;=\\u0026thinsp;0.034 at bin 0), but none reached statistical significance. These findings indicate that sampling asymmetry (equal vs. unequal information) did not elicit reliable modulation of pupil-linked arousal, and no evidence for a phasic peak was observed.\\u003c/p\\u003e\\u003cp\\u003eThe results from all time-bin analyses show that effects were somewhat more consistent immediately before the response, where they were similar in direction to the effects observed in the average pupil size analysis. This suggests that the observed pupil modulation reflects a sustained cognitive process leading up to the decision, rather than a discrete, transient response to an external stimulus. It also indicates that temporal dynamics along the time leading to decision may contribute to pupil size\\u003c/p\\u003e\\u003c/div\\u003e\"},{\"header\":\"Discussion\",\"content\":\"\\u003cp\\u003eThe present investigation examined the relationship between pupillary responses and exploration and exploitation behavior. Behavioral analyses confirmed the importance of these factors in influencing participants' choices. Participants adjusted their choices based on absolute value differences, consistently chose options with less prior information in asymmetrical-information conditions, and exhibited greater information-seeking behavior when more free choices remained (six vs. one). These patterns replicated Wilson et al.'s\\u003csup\\u003e7\\u003c/sup\\u003e findings and confirmed the effectiveness of the experimental manipulations in eliciting exploration behavior.\\u003c/p\\u003e\\u003cp\\u003eHowever, pupillary responses diverged from this behavioral pattern. Average pupil size increased selectively in response to the choice-horizon, with larger dilation observed in short horizon conditions where no subsequent corrections were possible. In long horizons, participants could anticipate future opportunities for information use, supporting directed exploration but without increasing immediate control demands. By contrast, value-uncertainty and information-asymmetry influenced behavior but did not elicit significant pupil modulation. This reveals a dissociation between the drivers of behavioral choice and those of cognitive effort.\\u003c/p\\u003e\\u003cp\\u003eThe absence of significant pupil modulation in response to information gaps challenges the assumption that all forms of directed exploration uniformly engage cognitive control systems. Previous work using the same Horizon Task paradigm has consistently shown that value and information gaps are key drivers of exploratory choices\\u003csup\\u003e\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR37\\\" class=\\\"CitationRef\\\"\\u003e37\\u003c/span\\u003e\\u003c/sup\\u003e, yet our findings demonstrate that pupil-indexed effort is insensitive to these same factors. This pattern indicates that cognitive resources are not allocated in response to immediate uncertainty but rather are selectively engaged according to the strategic context and future consequences of the decision.\\u003c/p\\u003e\\u003cp\\u003eThe current findings complement prior evidence linking pupil dilation to exploration. Kozunova et al.\\u003csup\\u003e33\\u003c/sup\\u003e, using a probabilistic learning task, demonstrated that directed exploration was associated with larger pupil dilations and response slowing, highlighting the role of pupil-linked arousal in deliberate information seeking. In contrast, the present results obtained with the Horizon Task revealed a selective sensitivity of baseline pupil size to the choice-horizon factor. Specifically, enlarged pupils were observed under short-horizon conditions, which in this paradigm favor exploitation rather than exploration. Taken together, these findings suggest that pupil-linked arousal does not simply differentiate exploration from exploitation as categorical states, but reflects a broader sensitivity to the strategic demands of the decision context\\u003c/p\\u003e\\u003cp\\u003eThis interpretation is consistent with the Expected Value of Control framework, which explains effort allocation in terms of anticipated benefits and costs\\u003csup\\u003e\\u003cspan citationid=\\\"CR24\\\" class=\\\"CitationRef\\\"\\u003e24\\u003c/span\\u003e\\u003c/sup\\u003e, and resonates with resource-rational perspectives that view cognitive resources as adaptively allocated according to the expected future utility of information\\u003csup\\u003e\\u003cspan citationid=\\\"CR38\\\" class=\\\"CitationRef\\\"\\u003e38\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e\\u003cp\\u003eThese findings extend exploration-exploitation frameworks by demonstrating the importance of strategic context. Rather than indexing exploration-specific cognitive demands, pupillary responses appear to reflect the broader strategic context of decision-making- specifically, the consequentiality and finality of choices. This pattern suggests that cognitive control allocation may be guided by decision consequentiality, with effort deployment scaling with the strategic weight (e.g., importance) of individual choices.\\u003c/p\\u003e\\u003cp\\u003eTrial-by-trial analysis examined the temporal dynamics of pupillary responses. Unlike prior studies reporting phasic pupil dilation near the time of decision, typically associated with transient uncertainty or conflict\\u003csup\\u003e\\u003cspan citationid=\\\"CR39\\\" class=\\\"CitationRef\\\"\\u003e39\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR40\\\" class=\\\"CitationRef\\\"\\u003e40\\u003c/span\\u003e\\u003c/sup\\u003e, we observed no consistent time-locked modulation. Instead, pupil size remained elevated throughout the decision period without a distinct peak at the choice moment. The stable pupil size observed here, without discrete peaks, contrasts with the phasic bursts typically found in other cognitive tasks\\u003csup\\u003e\\u003cspan citationid=\\\"CR36\\\" class=\\\"CitationRef\\\"\\u003e36\\u003c/span\\u003e\\u003c/sup\\u003e, suggesting that task complexity may determine the temporal profile of pupil responses. This interpretation finds some support in early working memory research\\u003csup\\u003e\\u003cspan citationid=\\\"CR41\\\" class=\\\"CitationRef\\\"\\u003e41\\u003c/span\\u003e\\u003c/sup\\u003e, where complex information processing similarly showed stable pupil size responses without discrete modulation. The current findings suggest that exploration under choice constraints may engage sustained cognitive processes that warrant further investigation.\\u003c/p\\u003e\\u003cp\\u003eIn conclusion, our findings address the fundamental question of how cognitive effort maps onto distinct behavioral components of exploration and exploitation. The selective pupillary response to choice horizon suggests that cognitive engagement is not uniformly distributed across exploration strategies. Rather than supporting traditional AGT predictions, our results suggest that strategic context determines pupil-indexed engagement, consistent with neurocomputational accounts linking control mechanisms to adaptive exploration under uncertainty\\u003csup\\u003e\\u003cspan citationid=\\\"CR42\\\" class=\\\"CitationRef\\\"\\u003e42\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e\\u003cp\\u003eThis pattern indicates that cognitive effort allocation depends on the strategic structure of the decision environment, providing a preliminary empirical synthesis between behavioral and physiological approaches to understanding human exploration. Future investigations should examine whether this selective pupillary sensitivity generalizes across decision contexts where choice consequences vary in their strategic importance beyond immediate outcomes.\\u003c/p\\u003e\"},{\"header\":\"Declarations\",\"content\":\"\\u003cul\\u003e\\n \\u003cli\\u003e\\u003cstrong\\u003eAuthor contributions: G.B.\\u003c/strong\\u003e contributed to conceptualization, methodology, investigation, formal analysis, and writing - original draft. \\u003cstrong\\u003eS.G-\\u0026nbsp;\\u003c/strong\\u003econtributed to methodology, writing \\u0026ndash; review \\u0026amp; editing and supervision\\u003cstrong\\u003e. \\u0026nbsp;U.H.\\u003c/strong\\u003e contributed to conceptualization formal analysis, methodology, supervision, and writing \\u0026ndash; review \\u0026amp; editing.\\u003c/li\\u003e\\n \\u003cli\\u003e\\u003cstrong\\u003eData availability statement:\\u0026nbsp;\\u003c/strong\\u003eAnonymized data and are available at OSF https://osf.io/dq3ke/files/osfstorage\\u003c/li\\u003e\\n \\u003cli\\u003e\\u003cstrong\\u003eDeclaration of Interest statement:\\u0026nbsp;\\u003c/strong\\u003eThe authors declare no competing financial or proprietary interests relevant to the content of the article.\\u003c/li\\u003e\\n \\u003cli\\u003e\\u003cstrong\\u003eFunding:\\u003c/strong\\u003e\\u003cstrong\\u003e\\u0026nbsp;\\u003c/strong\\u003eThe author declares that no funds, grants, or other financial support were received for the conduct of this study\\u003c/li\\u003e\\n\\u003c/ul\\u003e\\n\"},{\"header\":\"References\",\"content\":\"\\u003col\\u003e\\u003cli\\u003e\\u003cspan\\u003eCohen, J. D., McClure, S. M. \\u0026amp; Yu, A. J. Should I stay or should I go? How the human brain manages the trade-off between exploitation and exploration. \\u003cem\\u003ePhilos. Trans. R Soc. Lond. B Biol. Sci.\\u003c/em\\u003e \\u003cb\\u003e362\\u003c/b\\u003e, 933\\u0026ndash;942. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1098/rstb.2007.2098\\u003c/span\\u003e\\u003cspan address=\\\"10.1098/rstb.2007.2098\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2007).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eCogliati Dezza, I., Yu, A. J., Cleeremans, A. \\u0026amp; Alexander, W. Learning the value of information and reward over time when solving exploration\\u0026ndash;exploitation problems. \\u003cem\\u003eSci. Rep.\\u003c/em\\u003e \\u003cb\\u003e7\\u003c/b\\u003e, 16919. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1038/s41598-017-17237-w\\u003c/span\\u003e\\u003cspan address=\\\"10.1038/s41598-017-17237-w\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2017).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eMonk, C. T. et al. How ecology shapes exploitation: a framework to predict the behavioural response of human and animal foragers along exploration\\u0026ndash;exploitation trade-offs. \\u003cem\\u003eEcol. Lett.\\u003c/em\\u003e \\u003cb\\u003e21\\u003c/b\\u003e, 779\\u0026ndash;793. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1111/ele.12949\\u003c/span\\u003e\\u003cspan address=\\\"10.1111/ele.12949\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2018).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eBerger-Tal, O., Nathan, J., Meron, E. \\u0026amp; Saltz, D. The exploration\\u0026ndash;exploitation dilemma: a multidisciplinary framework. \\u003cem\\u003ePLoS ONE\\u003c/em\\u003e. \\u003cb\\u003e9\\u003c/b\\u003e, e95693. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1371/journal.pone.0095693\\u003c/span\\u003e\\u003cspan address=\\\"10.1371/journal.pone.0095693\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2014).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eRich, P. \\u0026amp; Gureckis, T. M. The limits of learning: exploration, generalization, and the development of learning traps. \\u003cem\\u003eJ. Exp. Psychol. Gen.\\u003c/em\\u003e \\u003cb\\u003e147\\u003c/b\\u003e, 1553\\u0026ndash;1570. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://https://doi 10.1037/xge0000466\\u003c/span\\u003e\\u003cspan address=\\\"https://https://doi 10.1037/xge0000466\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2018).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eMehlhorn, K. et al. Unpacking the exploration\\u0026ndash;exploitation tradeoff: a synthesis of human and animal literatures. \\u003cem\\u003eDecision\\u003c/em\\u003e \\u003cb\\u003e2\\u003c/b\\u003e, 191\\u0026ndash;215 .\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eWilson, R. C., Geana, A., White, J. M., Ludvig, E. A. \\u0026amp; Cohen, J. D. Humans use directed and random exploration to solve the explore\\u0026ndash;exploit dilemma. \\u003cem\\u003eJ. Exp. Psychol. Gen.\\u003c/em\\u003e \\u003cb\\u003e143\\u003c/b\\u003e, 2074\\u0026ndash;2081. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1037/a0038199\\u003c/span\\u003e\\u003cspan address=\\\"10.1037/a0038199\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2014).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eLai, T. L. \\u0026amp; Robbins, H. Asymptotically efficient adaptive allocation rules. \\u003cem\\u003eAdv. Appl. Math.\\u003c/em\\u003e \\u003cb\\u003e6\\u003c/b\\u003e, 4\\u0026ndash;22. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/0196-8858(85)90002-8\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/0196-8858(85)90002-8\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (1985).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eCarmel, D. \\u0026amp; Markovitch, S. Exploration strategies for model-based learning in multi-agent systems: exploration strategies. \\u003cem\\u003eAuton. Agents Multi-Agent Syst.\\u003c/em\\u003e \\u003cb\\u003e2\\u003c/b\\u003e, 141\\u0026ndash;172. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1023/A:1010007108196\\u003c/span\\u003e\\u003cspan address=\\\"10.1023/A:1010007108196\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (1999).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eDaw, N. D., O\\u0026rsquo;Doherty, J. P., Dayan, P., Seymour, B. \\u0026amp; Dolan, R. J. Cortical substrates for exploratory decisions in humans. \\u003cem\\u003eNature\\u003c/em\\u003e \\u003cb\\u003e441\\u003c/b\\u003e, 876\\u0026ndash;879. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1038/nature04766\\u003c/span\\u003e\\u003cspan address=\\\"10.1038/nature04766\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2006).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eAston-Jones, G. \\u0026amp; Cohen, J. D. An integrative theory of locus coeruleus\\u0026ndash;norepinephrine function: adaptive gain and optimal performance. \\u003cem\\u003eAnnu. Rev. Neurosci.\\u003c/em\\u003e \\u003cb\\u003e28\\u003c/b\\u003e, 403\\u0026ndash;450. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1146/annurev.neuro.28.061604.135709\\u003c/span\\u003e\\u003cspan address=\\\"10.1146/annurev.neuro.28.061604.135709\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2005).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eJepma, M. \\u0026amp; Nieuwenhuis, S. Pupil size predicts changes in the exploration\\u0026ndash;exploitation trade-off: evidence for the adaptive gain theory. \\u003cem\\u003eJ. Cogn. Neurosci.\\u003c/em\\u003e \\u003cb\\u003e23\\u003c/b\\u003e, 1587\\u0026ndash;1596. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1162/jocn.2010.21548\\u003c/span\\u003e\\u003cspan address=\\\"10.1162/jocn.2010.21548\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2011).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eWilson, R. C., Bonawitz, E., Costa, V. D. \\u0026amp; Ebitz, R. B. Balancing exploration and exploitation with information and randomization. \\u003cem\\u003eCurr. Opin. Behav. Sci.\\u003c/em\\u003e \\u003cb\\u003e38\\u003c/b\\u003e, 49\\u0026ndash;56. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/j.cobeha.2020.10.001\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/j.cobeha.2020.10.001\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2021).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eGeana, A., Wilson, R. C., Daw, N. \\u0026amp; Cohen, J. D. Boredom, information-seeking and exploration. In \\u003cem\\u003eProc. 38th Annu. Meet. Cogn. Sci. Soc.\\u003c/em\\u003e 1751\\u0026ndash;1756 (2016).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eZajkowski, W. K., Kossut, M. \\u0026amp; Wilson, R. C. A causal role for right frontopolar cortex in directed, but not random, exploration. \\u003cem\\u003eeLife\\u003c/em\\u003e 6, e27430. (2017). \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.7554/eLife.27430\\u003c/span\\u003e\\u003cspan address=\\\"10.7554/eLife.27430\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eSchwartenbeck, P. et al. Computational mechanisms of curiosity and goal-directed exploration. \\u003cem\\u003eeLife\\u003c/em\\u003e \\u003cb\\u003e8\\u003c/b\\u003e, e41703. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.7554/eLife.41703\\u003c/span\\u003e\\u003cspan address=\\\"10.7554/eLife.41703\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2019).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eYu, A. J. \\u0026amp; Dayan, P. Uncertainty, neuromodulation, and attention. \\u003cem\\u003eNeuron\\u003c/em\\u003e \\u003cb\\u003e46\\u003c/b\\u003e, 681\\u0026ndash;692. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/j.neuron.2005.04.026\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/j.neuron.2005.04.026\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2005).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eBehrens, T. E., Woolrich, M. W., Walton, M. E. \\u0026amp; Rushworth, M. F. Learning the value of information in an uncertain world. \\u003cem\\u003eNat. Neurosci.\\u003c/em\\u003e \\u003cb\\u003e10\\u003c/b\\u003e, 1214\\u0026ndash;1221. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1038/nn1954\\u003c/span\\u003e\\u003cspan address=\\\"10.1038/nn1954\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2007).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eSpeekenbrink, M. Chasing unknown bandits: Uncertainty guidance in learning and decision making. \\u003cem\\u003eCurr. Dir. Psychol. Sci.\\u003c/em\\u003e \\u003cb\\u003e31\\u003c/b\\u003e, 419\\u0026ndash;427. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1177/09637214221105051\\u003c/span\\u003e\\u003cspan address=\\\"10.1177/09637214221105051\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2022).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003ePayzan-LeNestour, \\u0026Eacute;. \\u0026amp; Bossaerts, P. Do not bet on the unknown versus try to find out more: estimation uncertainty and unexpected uncertainty both modulate exploration. \\u003cem\\u003eFront. Neurosci.\\u003c/em\\u003e \\u003cb\\u003e6\\u003c/b\\u003e, 150. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.3389/fnins.2012.00150\\u003c/span\\u003e\\u003cspan address=\\\"10.3389/fnins.2012.00150\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2012).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eCourville, A. C., Daw, N. D. \\u0026amp; Touretzky, D. S. Bayesian theories of conditioning in a changing world. \\u003cem\\u003eTrends Cogn. Sci.\\u003c/em\\u003e \\u003cb\\u003e10\\u003c/b\\u003e, 294\\u0026ndash;300. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/j.tics.2006.05.004\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/j.tics.2006.05.004\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2006).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eGershman, S. J. \\u0026amp; Tzovaras, B. G. Dopaminergic genes are associated with both directed and random exploration. \\u003cem\\u003eNeuropsychologia\\u003c/em\\u003e \\u003cb\\u003e120\\u003c/b\\u003e, 97\\u0026ndash;104. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/j.neuropsychologia.2018.10.009\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/j.neuropsychologia.2018.10.009\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2018).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eGershman, S. J. Uncertainty and exploration. \\u003cem\\u003eDecision\\u003c/em\\u003e \\u003cb\\u003e6\\u003c/b\\u003e, 277. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1037/dec0000101\\u003c/span\\u003e\\u003cspan address=\\\"10.1037/dec0000101\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2019).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eShenhav, A., Botvinick, M. M. \\u0026amp; Cohen, J. D. The expected value of control: an integrative theory of anterior cingulate cortex function. \\u003cem\\u003eNeuron\\u003c/em\\u003e \\u003cb\\u003e79\\u003c/b\\u003e, 217\\u0026ndash;240. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/j.neuron.2013.07.007\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/j.neuron.2013.07.007\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2013).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eInzlicht, M., Shenhav, A. \\u0026amp; Olivola, C. Y. The effort paradox: Effort is both costly and valued. \\u003cem\\u003eTrends Cogn. Sci.\\u003c/em\\u003e \\u003cb\\u003e22\\u003c/b\\u003e, 337\\u0026ndash;349. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/j.tics.2018.01.007\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/j.tics.2018.01.007\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2018).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eJoshi, S., Li, Y., Kalwani, R. M. \\u0026amp; Gold, J. I. Relationships between pupil size and neuronal activity in the locus coeruleus, colliculi, and cingulate cortex. \\u003cem\\u003eNeuron\\u003c/em\\u003e \\u003cb\\u003e89\\u003c/b\\u003e, 221\\u0026ndash;234. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/j.neuron.2015.11.028\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/j.neuron.2015.11.028\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2016).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eAln\\u0026aelig;s, D. et al. Pupil size signals mental effort deployed during multiple object tracking and predicts brain activity in the dorsal attention network and the locus coeruleus. \\u003cem\\u003eJ. Vis.\\u003c/em\\u003e \\u003cb\\u003e14\\u003c/b\\u003e, 1. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1167/14.4.1\\u003c/span\\u003e\\u003cspan address=\\\"10.1167/14.4.1\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2014).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003ePreuschoff, K., \\u0026rsquo;t Hart, B. M. \\u0026amp; Einh\\u0026auml;user, W. Pupil dilation signals surprise: evidence for noradrenaline's role in decision making. \\u003cem\\u003eFront. Neurosci.\\u003c/em\\u003e \\u003cb\\u003e5\\u003c/b\\u003e, 115. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.3389/fnins.2011.00115\\u003c/span\\u003e\\u003cspan address=\\\"10.3389/fnins.2011.00115\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2011).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eNassar, M. R. et al. Rational regulation of learning dynamics by pupil-linked arousal systems. \\u003cem\\u003eNat. Neurosci.\\u003c/em\\u003e \\u003cb\\u003e15\\u003c/b\\u003e, 1040\\u0026ndash;1046. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1038/nn.3130\\u003c/span\\u003e\\u003cspan address=\\\"10.1038/nn.3130\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2012).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eZ\\u0026eacute;non, A. Eye pupil signals information gain. \\u003cem\\u003eProc. Biol. Sci.\\u003c/em\\u003e 286, 20191593. (2019). \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1098/rspb.2019.1593\\u003c/span\\u003e\\u003cspan address=\\\"10.1098/rspb.2019.1593\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003ePayzan-LeNestour, E. \\u0026amp; Bossaerts, P. Risk, unexpected uncertainty, and estimation uncertainty: Bayesian learning in unstable settings. \\u003cem\\u003ePLoS Comput. Biol.\\u003c/em\\u003e \\u003cb\\u003e7\\u003c/b\\u003e, e1001048. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1371/journal.pcbi.1001048\\u003c/span\\u003e\\u003cspan address=\\\"10.1371/journal.pcbi.1001048\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2011).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eCavanagh, J. F., Figueroa, C. M., Cohen, M. X. \\u0026amp; Frank, M. J. Frontal theta reflects uncertainty and unexpectedness during exploration and exploitation. \\u003cem\\u003eCereb. Cortex\\u003c/em\\u003e. \\u003cb\\u003e22\\u003c/b\\u003e, 2575\\u0026ndash;2586. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1093/cercor/bhr332\\u003c/span\\u003e\\u003cspan address=\\\"10.1093/cercor/bhr332\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2012).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eKozunova, G. L. et al. Pupil dilation and response slowing distinguish deliberate explorative choices in the probabilistic learning task. \\u003cem\\u003eCogn. Affect. Behav. Neurosci.\\u003c/em\\u003e \\u003cb\\u003e22\\u003c/b\\u003e, 1108\\u0026ndash;1129. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.3758/s13415-022-00996-z\\u003c/span\\u003e\\u003cspan address=\\\"10.3758/s13415-022-00996-z\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2022).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eCohen, J. \\u003cem\\u003eStatistical Power Analysis for the Behavioral Sciences\\u003c/em\\u003e (Routledge, 2013).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eFaul, F., Erdfelder, E., Buchner, A. \\u0026amp; Lang, A. G. Statistical power analyses using G*Power 3.1: tests for correlation and regression analyses. \\u003cem\\u003eBehav. Res. Methods\\u003c/em\\u003e. \\u003cb\\u003e41\\u003c/b\\u003e, 1149\\u0026ndash;1160. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.3758/BRM.41.4.1149\\u003c/span\\u003e\\u003cspan address=\\\"10.3758/BRM.41.4.1149\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2009).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eGabay, S., Pertzov, Y. \\u0026amp; Henik, A. Orienting of attention, pupil size, and the norepinephrine system. \\u003cem\\u003eAtten. Percept. Psychophys\\u003c/em\\u003e. \\u003cb\\u003e73\\u003c/b\\u003e, 123\\u0026ndash;129. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.3758/s13414-010-0015-4\\u003c/span\\u003e\\u003cspan address=\\\"10.3758/s13414-010-0015-4\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2011).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eSadeghiyeh, H. et al. Temporal discounting correlates with directed exploration but not with random exploration. \\u003cem\\u003eSci. Rep.\\u003c/em\\u003e \\u003cb\\u003e10\\u003c/b\\u003e, 4020. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1038/s41598-020-60576-4\\u003c/span\\u003e\\u003cspan address=\\\"10.1038/s41598-020-60576-4\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2020).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eLieder, F. \\u0026amp; Griffiths, T. L. Resource-rational analysis: understanding human cognition as the optimal use of limited computational resources. \\u003cem\\u003eBehav. Brain Sci.\\u003c/em\\u003e \\u003cb\\u003e43\\u003c/b\\u003e, e1. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1017/S0140525X1900061X\\u003c/span\\u003e\\u003cspan address=\\\"10.1017/S0140525X1900061X\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2019).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eMurphy, P. R., Robertson, I. H., Balsters, J. H. \\u0026amp; O\\u0026rsquo;Connell, R. G. Pupillometry and P3 index the locus coeruleus-noradrenergic arousal function in humans. \\u003cem\\u003ePsychophysiology\\u003c/em\\u003e \\u003cb\\u003e48\\u003c/b\\u003e, 1532\\u0026ndash;1543. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1111/j.1469-8986.2011.01226.x\\u003c/span\\u003e\\u003cspan address=\\\"10.1111/j.1469-8986.2011.01226.x\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2011).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003evan Steenbergen, H. \\u0026amp; Band, G. P. Pupil dilation in the Simon task as a marker of conflict processing. \\u003cem\\u003eFront. Hum. Neurosci.\\u003c/em\\u003e \\u003cb\\u003e7\\u003c/b\\u003e, 215. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.3389/fnhum.2013.00215\\u003c/span\\u003e\\u003cspan address=\\\"10.3389/fnhum.2013.00215\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2013).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eGranholm, E., Asarnow, R. F., Sarkin, A. J. \\u0026amp; Dykes, K. L. Pupillary responses index cognitive resource limitations. \\u003cem\\u003ePsychophysiology\\u003c/em\\u003e \\u003cb\\u003e33\\u003c/b\\u003e, 457\\u0026ndash;461. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1111/j.1469-8986.1996.tb01071.x\\u003c/span\\u003e\\u003cspan address=\\\"10.1111/j.1469-8986.1996.tb01071.x\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (1996).\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eBadre, D., Doll, B. B., Long, N. M. \\u0026amp; Frank, M. J. Rostrolateral prefrontal cortex and individual differences in uncertainty-driven exploration. \\u003cem\\u003eNeuron\\u003c/em\\u003e \\u003cb\\u003e73\\u003c/b\\u003e, 595\\u0026ndash;607. \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://doi.org/10.1016/j.neuron.2011.12.025\\u003c/span\\u003e\\u003cspan address=\\\"10.1016/j.neuron.2011.12.025\\\" targettype=\\\"DOI\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e (2012).\\u003c/span\\u003e\\u003c/li\\u003e\\u003c/ol\\u003e\"}],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":true,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":true,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":false,\"isAuthorSuppliedPdf\":false,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":false,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true},\"keywords\":\"\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-7740172/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-7740172/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003cp\\u003eWhen making decisions, the explore - exploit dilemma represents balancing reward maximization with uncertainty reduction. While reinforcement learning models often treat exploration as stochastic variability, theories such as Adaptive Gain Theory (AGT) and Expected Value of Control (EVC) suggest that exploration and exploitation can reflect strategic control allocation. Pupillometry provides a window into the locus coeruleus\\u0026ndash;norepinephrine system, indexing cognitive effort and task engagement. The present study combined pupillometry with the Horizon Task to examine whether directed and random exploration and exploitation differentially recruit cognitive resources under varying environmental contexts. Thirty-five adults completed a task manipulating three variables: value gap (reward difference), information gap (sampling imbalance), and choice horizon (1 vs. 6 free trials). Behavioral analyses replicated established findings: small value gaps promoted exploration, unequal sampling elicited information-seeking, and long horizons increased directed exploration. However, pupillary responses diverged from behavior, showing selective sensitivity to choice horizons. Pupil size was larger in short horizon conditions compared with long-horizon, suggesting increased control engagement under conditions of heightened consequence, whereas value and information gaps did not elicit significant modulation. These findings challenge the view that exploration uniformly entails increased effort and instead highlight that strategic context determines pupil-indexed engagement. Our results provide evidence that cognitive effort allocation during exploration depends on the prospective utility and irreversibility of decisions, bridging behavioral and neurocomputational perspectives.\\u003c/p\\u003e\",\"manuscriptTitle\":\"Selective Pupil Size Response Within direct and random exploration and exploitation Behaviors\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2025-10-24 22:37:31\",\"doi\":\"10.21203/rs.3.rs-7740172/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"3ad1db1a-ef7e-41a4-b232-21e0a83d4772\",\"owner\":[],\"postedDate\":\"October 24th, 2025\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"posted\",\"subjectAreas\":[{\"id\":56668455,\"name\":\"Biological sciences/Neuroscience\"},{\"id\":56668456,\"name\":\"Biological sciences/Psychology\"},{\"id\":56668457,\"name\":\"Social science/Psychology\"}],\"tags\":[],\"updatedAt\":\"2025-11-14T05:53:25+00:00\",\"versionOfRecord\":[],\"versionCreatedAt\":\"2025-10-24 22:37:31\",\"video\":\"\",\"vorDoi\":\"\",\"vorDoiUrl\":\"\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-7740172\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-7740172\",\"identity\":\"rs-7740172\",\"version\":[\"v1\"]},\"buildId\":\"8U1c8b4HqxoKbykW_rLl7\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}