PMID 27241710 — Levodopa impairs probabilistic reversal learning in healthy young adults.
good_imrad R=2039w / 12¶ | figs=18 Alex
TITLE
[1] 9w Levodopa impairs probabilistic reversal learning in healthy young adults
ABSTRACT
[1] 194w Rationale Dopaminergic therapy improves some cognitive functions and worsens others in patients with Parkinson's disease (PD). These paradoxical effects are explained by the dopamine overdose hypothesis, which proposes that effects of dopaminergic therapy on a cognitive function is determined by the baseline dopamine levels in brain regions mediating that function.Objectives We directly tested this prevalent hypothesis, evaluating the effects of levodopa on stimulus-reward learning in healthy young adults, who presumably have optimal baseline dopamine levels and dopamine regulation. Methods Twenty-six healthy, young adults completed a probabilistic reversal learning task in a randomized, double-blind, placebo-controlled, crossover design. Participants completed one session on levodopa 100 mg/carbidopa 25 mg and another session on placebo. Results We found that levodopa impaired reversal learning relative to placebo. Further analyses revealed that levodopa impaired learning from both punishment and reward. Conclusions Exogenous dopamine impairs stimulusreward learning, independent of PD pathology and prior to sensitization through repeated exposure, in healthy adults with normal cognition and baseline dopamine function. Our findings support the dopamine overdose hypothesis and caution clinicians about detrimental effects of levodopa in all clinical populations (e.g., early PD, restless leg syndrome) regardless of baseline cognitive and dopaminergic system function.
INTRO
[1] 82w Dopaminergic medications, such as L-3,4-dihydroxy phenylalanine (levodopa) and dopamine receptor agonists, are commonly prescribed to improve motor abnormalities in Parkinson's disease (PD). Although motor and some cognitive symptoms are improved by this therapy, other cognitive functions appear to be worsened (Swainson et al. 2000;Cools et al. 2003;Frank et al. 2004;MacDonald et al. 2011). These paradoxical effects of dopaminergic therapy on various aspects of cognition have been explained by differences in baseline dopamine levels in underlying brain regions (Cools 2006;MacDonald and Monchi 2011).
[2] 200w In PD, motor symptoms emerge when dopamineproducing neurons in the substantia nigra degenerate, seriously restricting dopamine supply to the dorsal striatum (DS; Kish et al. 1988). The DS refers to the bulk of the caudate nucleus and putamen. By comparison, dopaminergic neurons in the ventral tegmental area (VTA) are less affected (Haber et al. 1997). VTA-innervated brain regions, such as prefrontal and limbic cortices, as well as the ventral striatum (VS), composed of the nucleus accumbens and most ventral caudate and putamen, are relatively dopamine replete. Differential baseline dopamine impairment in DS versus VTA-innervated brain regions is offered as an explanation for the disparate effect of dopaminergic therapy on functions in PD (Cools 2006). It is proposed that motor and cognitive symptoms mediated by the dopamine-deplete DS are impaired at baseline and improve with medication. In contrast, cognitive functions attributed to dopamine-replete VTA-innervated brain regions are theorized to be normal at baseline and to worsen with dopamine replacement. The latter finding is explained by the dopamine overdose hypothesis (Gotham et al. 1988;Swainson et al. 2000;Cools 2006), which postulates that exogenous dopamine, titrated to redress DS dopamine depletion and motor symptoms, causes excess dopamine levels in relatively replete, VTAinnervated brain regions.
[3] 103w Despite its prevalence and implications for treatment in PD, critical tests of the dopamine overdose hypothesis are scarce. Toward this end, we evaluated the effects of levodopa in healthy young adults on probabilistic reversal learning and learning from rewards and punishments-functions mediated by VTA-innervated brain regions (Cools et al. 2002;Fellows and Farah 2003;Cools et al. 2007) that are worsened by dopaminergic therapy in PD (Cools et al. 2001;Cools 2006;Graef et al. 2010;MacDonald et al. 2013a). Healthy young adults are expected to have optimal endogenous dopamine levels. Consequently, the dopamine overdose hypothesis predicts deficits in young adult participants comparable to those seen in PD.
RESULTS
[1] 179w We analyzed the data in terms of (i) total learning session, (ii) initial stimulus-reward acquisition period, and (iii) stimulusreward reversal learning phase (see Fig. 1). The initial stimulus-reward learning phase included the initial trial until the first reversal. All trials thereafter were considered as part of the stimulus-reward reversal learning phase. The total number of errors committed during each of these learning periods was used as a behavioral measure of stimulus-reward learning, with more errors corresponding to poorer learning. Analogous analyses were performed on response time (RT) data, as faster responding presumably indexes better learning. Because we anticipated practice effects between testing sessions, a between-subjects factor of order was included in our analyses. For each learning phase, we ran separate 2 × 2 × 2 mixed analyses of variance (ANOVAs), with order (L/P vs. P/L) and gender (male vs. female) as the between-subject factors and treatment (levodopa vs. placebo) as the withinsubject variable. We ran additional independent samples (i.e., between-group) t tests, examining error rates in session 1 only, given our expectation that there would be significant practice effects.
[2] 89w Further, we analyzed error rates and RTs to assess participants' learning from punishment and reward in the levodopa and placebo sessions separately. We computed percent accuracy scores on trials immediately following reward and punishment respectively. Mean error rates and RTs for trials following reward and punishment in the levodopa and placebo sessions are presented in Table 3. We performed a 2 × 2 × 2 repeatedmeasures ANOVA, with treatment (levodopa vs. placebo) and feedback (punishment vs. reward) as within-subject variables and order (L/P vs. P/L) as a between-subject factor.
[3] 49w We also investigated reward and punishment history on accuracy using multiple logistic regression analysis in levodopa and placebo sessions separately. Previous outcomes (i.e., reward versus punishment), RT, as well as experimental events including previous reversals or false feedback occurring at t-1, t-2, and t-3 were included as explanatory variables.
[4] 76w Results from our control measures are presented in Table 1. No participant reported adverse side effects of the drug during the experiment. Blood pressure and self-reported ratings of subjective mood did not significantly differ on levodopa versus placebo (p > 0.05). Heart rate was greater when participants were on levodopa relative to placebo (p < 0.05). Participants' judgment of whether they received drug or placebo in each testing session was at chance (53.8 % correct judgment).
[5] 24w The mean number of errors for each learning phase is presented in Fig. 2a-f. The greater number of errors was indicative of poorer learning.
[6] 373w Three 2 × 2 × 2 mixed ANOVAs were performed on the number of errors for each of the pre-defined learning phases. We found significant main effects of treatment during the stimulus-reward reversal period (F ( 1 , 2 2 ) = 5.401, MSE = 191.816, p = 0.030; Fig. 2a) and the total learning session (F (1, 22) = 4.469, MSE = 329.538, p = 0.046; Fig. 2b), respectively, though the main effect of treatment did not reach significance for the initial stimulus-reward acquisition period (F (1, 22) = 1.002, MSE = 38.244, p = 0.328; Fig. 2c). These main effects reflected more errors in the levodopa compared to the placebo sessions (Fig. 2a-c). As expected, the main effect of order was significant for stimulus-reward reversal phase (F (1, 22) = 8.798, MSE = 447.968, p = 0.007) and total learning session (F (1, 22) = 8.791, MSE = 637.615, p = 0.007), but not for initial stimulus-reward acquisition phase (F (1, 22) = 3.579, MSE = 40.828, p = 0.072). Using each of our estimates of learning, these main effects revealed significantly more errors occurring in the first relative to the second session, reflecting straightforward practice effects. There was no significant main effect of gender for stimulus-reward reversal period (F (1, 22) = 1.372, MSE = 447.968, p = 0.254), total learning session (F (1, 22) = 1.893, MSE = 637.615, p = 0.183), or initial stimulus-reward acquisition period (F (1, 22) = 2.427, MSE = 40.828, p = 0.134). Our analyses revealed significant treatment × order interactions for stimulus-reward reversal period (F (1, 24) = 5.677, MSE = 188.667, p = 0.025), total learning session (F (1, 24) = 6.390, MSE = 347.910, p = 0.018), and initial stimulus-reward acquisition phase (F (1, 24) = 4.585, MSE = 45.465, p = 0.043). In each instance, these interaction effects could be understood by the fact that there were significantly more errors in the levodopa compared to the placebo session in the L/P order (p < 0.05 for stimulusreward reversal period; total learning session; initial stimulus-reward acquisition phase) but not in the P/L order (all p > 0.05). However, there were no significant three-way interactions in any analysis (all p > 0.05).
[7] 86w In light of the significant main effects of order and the order × treatment interactions, we performed an independent sample (i.e., between-group) t test, examining error rates in session 1 only (Fig. 2d-f). This allowed us to examine the effect of levodopa on learning, independent of any practice effects. For stimulus-reward reversal period (t(24) = -2.873, p = 0.008; Fig. 2d) and total learning session (t(24) = -2.497, p = 0.020; Fig. 2e), more errors occurred for participants receiving levodopa compared to placebo in session 1.
[8] 328w Similar to our initial analysis on error rates, we performed 2 × 2 × 2 mixed ANOVAs on RTs for each learning phase (see Table 2). We found a significant main effect of order for total learning session (F (1, 22) = 4.573, MSE = 17434.219, p = 0.044) and initial stimulus-reward acquisition phase (F (1, 22) = 6.428, MSE = 139050.426, p = 0.019) but not for the stimulus-reward reversal period (F (1, 22) = 2.655, MSE = 20328.439, p = 0.117). RTs for participants in the L/P order were significantly slower than those in the P/L order. Neither the main effect of treatment nor gender reached significance in any of the analyses (all p > 0.05). There were, however, significant treatment × order interactions for stimulus-reward reversal period (F (1, 22) = 7.217, MSE = 13426.112, p = 0.013), total learning session (F (1, 22) = 10.546, MSE = 13484.811, p = 0.004), and initial stimulus-reward acquisition phase (F (1, 22) = 6.170, MSE = 134252.046, p = 0.021). Pairwise comparisons consistently revealed significantly slower responding on levodopa compared to placebo in the L/P order (all p < 0.05), but not in the P/L order (all p > 0.05; see Table 2). There were significant order × gender interactions for stimulus-reward reversal period (F (1, 22) = 4.52, MSE = 20328.439, p = 0.042) and total learning session (F (1, 22) = 4.798, MSE = 17434.216, p = 0.035), but not in the initial stimulus-reward acquisition phase (F (1, 22) = 0.003, MSE = 139050.426, p = 0.957). Participants repeated this procedure until a total of nine successful reversals were achieved or a 500-trial deadline was reached Pairwise comparisons demonstrated significantly slower responding in females compared to males in the P/L order (both p < 0.05), but not in the L/P order (both p > 0.05). Most important, there were no significant treatment × gender interactions or three-way interactions in any analysis (all p > 0.05).
[9] 305w The mean error rates and RTs for each feedback condition are presented in Table 3. The 2 × 2 × 2 ANOVAs on errors and RTs revealed the main effects of treatment (F (1,24) = 9.72, MSE = 0.015, p = 0.005 for errors and F (1,24) = 21.382, MSE = 444,762.654, p < 0.001 for RT) and feedback (F (1,24) = 5.58, MSE = 0.112, p < 0.05 for errors and F (1,24) = 12.85, MSE = 115, 485.891, p < 0.005 for RT). There was no main effect of order in terms of errors (F (1, 24) < 1), but there was a significant effect on RT (F (1, 2 4 ) = 6 . 1 3 4 , M S E = 2 2 6 , 6 5 7 . 8 2 3 , p < 0 . 0 2 5 ) . T h e treatment × order interaction was significant (F (1,24) = 6.93, MSE = 0.011, p < 0.025 for errors and F (1,24) = 18.758, MSE = 387,557.881, p < 0.001 for RT), reflecting significantly more errors and longer RT for levodopa than placebo in the L/P order, but no difference in errors in the P/L order and longer RT for placebo than levodopa in the P/L order. This result owes to the opposite effects of treatment and practice cancelling each other for errors in the P/L order. The feedback × order effect was non-significant (F (1,24) = 2.88, MSE = 0.058, p > 0.100 for errors and F (1,24) = 1.339, MSE = 12,027.962, p > 0.250 for RT). Most important, the treatment × feedback and treatment × feedback × order errors on levodopa versus placebo in session 1 only for d stimulusreward reversal phase, e total learning, and f initial stimulus-reward acquisition phase. *p < 0.05
[10] 104w Table 2 Mean reaction times for participants on levodopa versus placebo Learning Phase Order Levodopa Placebo Stimulus-reward reversal L/P 606.196 (61.161) 473.369 (27.055) P/L 415.729 (22.655) 466.653 (21.741) Total learning L/P 647.420 (56.773) 492.472 (28.275) P/L 425.494 (21.828) 490.543 (22.348) Initial stimulus-reward acquisition L/P 1072.636 (169.543) 695.270 (97.241) P/L 543.512 (27.405) 665.371 (44.005) Values reported are means (±SEM) Psychopharmacology interactions were non-significant (both F (1,24) < 1 for error and F (1,24) = 3.581, MSE = 37, 135.299, p = 0.071 and F (1, 24) = 3.182, MSE = 32,991.201, p = 0.087 for RT), reflecting equivalent effects of levodopa on reward and punishment.
[11] 94w We further constructed multiple logistic regression models to evaluate how reward and punishment affected performance accuracy in levodopa versus placebo sessions. We relied on the Akaike information criterion (AIC) to provide an index of the quality of the models (Akaike 1974). The explanatory variables were previous outcomes (i.e., reward versus punishment), RT, as well as experimental events including previous reversals or false feedback occurring at t-1, t-2, and t-3. Models incorporating explanatory variables for levodopa and placebo sessions were contrasted with a constant-only model to assess whether the explanatory variables significantly predicted trial-by-trial performance.
[12] 332w Both the levodopa and placebo models were significantly better at predicting performance accuracy than the constantonly model as expected (χ 2 (9) = 2733.07, p < 0.001, AIC = 6835.853 for levodopa and χ 2 (9) = 2748.630, p < 0.001, AIC = 6820.3 for placebo). Measures of Nagelkerke's R 2 were 0.411 and 0.413, essentially reflecting the variance accounted for by each model and constituting moderately good models. This degree of fit is anticipated given the inclusion of stimulus-reward contingency reversals as well as the provision of false feedback that make performance less predictable. All explanatory variables reached significance as measured by the Wald statistic (all p < 0.001; see Table 4). In the levodopa model, we obtained Exp (B) = 1.361 for reward and Exp (B) = 0.345 for punishment, whereas in the placebo model, we found Exp (B) = 2.329 for reward and Exp (B) = 0.467 for punishment. Participants used both reward and punishment less effectively to guide performance in the levodopa session. A single unit change in reward led to 1.361 times increase in accuracy for the levodopa session compared to 2.329 times increase in accuracy for the placebo session. A single unit change in punishment in the levodopa session was 65.5 % less likely to lead to a correct response in the levodopa session compared to 53.3 % less likely to precede a correct response in the placebo session. In this way, punishment caused more errors in the levodopa than in the placebo model. The negative effect of punishment on accuracy can be understood by our design in which punishment most frequently corresponded to unannounced changes in stimulus-reward contingency and not surprisingly decreased accuracy for performance of the trial immediately following this event. Further, using the AIC criterion, the placebo model (AIC = 6820.3) was deemed superior to the levodopa model (AIC = 6835.82), with lower AIC denoting better fit and a relative likelihood of only 0.0004 (exp((6820.3 -6835.9) / 2)) that this interpretation is incorrect.
DISCUSS
[1] 55w In a placebo-controlled experiment, we found that a first and single dose of levocarb 100/25 mg administered orally to healthy young participants impaired reversal learning. Participants made more errors when tested on levodopa compared to placebo. Further inspection revealed that this impairment reflected medication-associated deficits in learning from both punishment and reward relative to placebo.
[2] 403w Our participants performed a probabilistic reversal learning task in two sessions, once on levodopa and once on placebo. Independent of the possible effects of levodopa, we expected significant practice effects for participants who were repeatedly performing our learning task. This has been shown previously (Buelow et al. 2015). Consequently, we counterbalanced the order of levodopa and placebo sessions across participants to account for these expected practice effects. Comparing within participants, there was a main, adverse effect of levodopa compared to placebo on errors in the stimulus-reward reversal learning phase and in the total learning session. As anticipated, there were significant main effects of order with participants making significantly more errors in the L/P order compared to the P/L order. There were also order × treatment interactions explained by the fact that when participants performed on levodopa in session 1 and on placebo in session 2, there was poorer learning for the levodopa condition relative to placebo using two of our learning measures (i.e., stimulus-reward reversal and overall stimulusreward learning). We interpret that this reflects the adverse effect of levodopa on learning combined with practice effects bolstering performance in session 2. That is, both levodopa and practice effects were affecting performance in the same direction. When participants performed first on placebo and second on levodopa, there were no differences in reward learning performance across sessions. This is not at all expected, as robust practice effects were predicted. That performance did not improve in session 2 is also strong evidence that levodopa impairs stimulus-reward learning. In fact, the P/L order provides an opportunity to directly compare the strong effect of practice on learning that is clearly established in the literature (Buelow et al. 2015), with the consequence of levodopa administration on learning. The effects of levodopa and practice are acting in opposite directions in the P/L order, essentially cancelling each other. Though these within-subject effects were quite consistent in supporting the view that levodopa impairs stimulusreward and reversal learning, we next considered reward learning in session 1 only, where complex order and practice effects were precluded. We found a clear and significant impairment in reward learning for participants performing session 1 on levodopa compared to those performing session 1 on placebo, for overall learning and during stimulus-reward reversal learning. Consequently, in three separate manners, we confirm our finding of impaired stimulus-reward learning in healthy young participants from a first, single dose of levocarb 100/25 mg.
CONCL
[1] 225w Levodopa remains the most widely used and effective treatment in PD. To our knowledge, this is the first study to demonstrate the deleterious effects of levodopa on stimulusreward learning in healthy young adults. Whereas motor symptoms and some cognitive deficits are improved by levodopa in PD, other cognitive functions are worsened. The latter has been explained by the dopamine overdose hypothesis. Here, we critically tested the straightforward notion that exogenous dopamine can overwhelm mechanisms for regulating dopamine in VTA-innervated brain regions that are at baseline dopamine replete. We confirmed this hypothesis and showed that dopamine overdose effects do not owe to an interaction with PD pathology (i.e., reduced dopamine regulation at the synapse due to decreased DAT, sensitization through chronic exposure to dopaminergic therapy). In fact, by testing healthy young adults, we clearly demonstrate that learning impairment occurs as a main effect of levodopa, independent of PD pathology, severity, or cognitive reserve. We found that levodopa impairs learning from both reward and punishment, which is a somewhat controversial finding that requires further investigation. Our results suggest that in PD patients who complain of cognitive deficits, particularly difficulty in acquiring new skills or concepts, medication reductions might be considered, even in young, relatively cognitively intact patients. Finally, patients treated with levodopa for other indications, such as restless leg syndrome, are also susceptible to cognitive side effects.
METHODS
[1] 91w Twenty-six healthy young adults (mean age ± SEM = 21.08 ± 0.29 years; 17 males) participated in the present study. All participants had no history of neurological or psychiatric disorders, of current or past alcohol or drug abuse, or contraindications to taking levodopa (e.g., cardiovascular illness, therapy with monoamine oxidase (MAO) inhibitors). This study was approved by the Health Sciences Research Ethics Board of the University of Western Ontario. Participants provided informed consent to the approved protocol before beginning the experiment, in accordance with the Declaration of Helsinki (World Medical 2013).
[2] 258w All participants completed two testing sessions in a randomized, double-blind, placebo-controlled, crossover design. They received 100 mg of levodopa and 25 mg of carbidopa (i.e., levocarb 100/25 mg) in one session and an equal volume of placebo in the other session. Levodopa is a dopamine precursor that is converted to dopamine in the brain by dopaminergic neurons. Carbidopa is a peripheral decarboxylase inhibitor that, unlike levodopa, cannot cross the blood brain barrier and prevents the early conversion of levodopa to dopamine and downstream substrates outside of the brain. The dose used here was the same as has been employed in previous psychopharmacological studies (Knecht et al. 2004;Flöel et al. 2005;Onur et al. 2011). Both drug and placebo were administered orally in identical capsules to maintain doubleblindness. The drug-placebo order was counterbalanced across participants to control for practice, order, and fatigue effects. That is, half of participants were tested first on levodopa (i.e., L/P) and the other half were tested first on placebo (i.e., P/L). Participants abstained from caffeine, alcohol, and nicotine on days of testing. All participants were tested at least 1 h after a meal to be sure that protein-containing food would not interfere with levodopa absorption. Cognitive testing began approximately 45 min after capsule ingestion to allow for peak plasma levodopa levels (Olanow et al. 2000). Participants also had their blood pressure and heart rate measured and completed standardized self-reported ratings of subjective mood (Bond and Lader 1974). Testing sessions were separated by a washout period of 7 days to allow time for total drug clearance.
[3] 96w During each testing session, participants completed a probabilistic reversal learning task. Reversal learning requires a participant to adapt his/her stimulus-reward selections to changing stimulus-reward contingencies, which are acquired through trial and error. This paradigm has previously been shown to engage VTA-innervated brain regions including ventral striatum, orbitofrontal cortex, and ventromedial prefrontal cortex (Cools et al. 2002;Fellows and Farah 2003;Cools et al. 2007). Although a few studies also link dorsomedial striatum to reversal learning, its specific role appears to be limited to maintaining new reward choice patterns rather than initial discrimination learning or reversals themselves (Ragozzino 2007).
[4] 296w The experimental task required participants to choose between a pair of cards dealt from two decks: one red and the other blue. On each trial, a card from each deck was presented side-by-side in the center of the computer screen. The leftright location of the cards was randomly switched between trials to prevent a deck from becoming associated with a consistent location or key-press response. Participants were instructed that one deck contained more winning than losing cards (probabilistically favorable), whereas the other deck contained more losing than winning cards (probabilistically unfavorable), and that the probabilistically favorable deck could change at any point throughout the task without notice. They were not made aware of the exact reinforcement probabilities, however. The objective of the task was to select the card that was most likely to be the winner (i.e., from the probabilistically favorable deck) on each trial. Participants were not told which deck was associated with the more favorable outcome. The initial stimulus-reward contingency and reversals were therefore discerned through trial and error. Participants pressed the BZ^key to select the card appearing on the left and the B/^key to select the card appearing on the right. Feedback was provided at the end of each trial. Selecting the winning card resulted in a $50 increase in total pot winnings, whereas selecting the losing card resulted in a $50 decrease in total winnings. Participants were not actually financially compensated in accordance with these winnings. All trials proceeded as follows: (i) a red card and a blue card were presented, side by side, in the center of the computer screen until the participant provided a key-press response, either BZô r B/^keys; (ii) feedback, either BCorrect (+$50)^or BIncorrect (-$50),^was presented for 1000 ms; (iii) a blank screen for 500 ms separated trials.
[5] 206w At the outset, either the red or blue deck was randomly assigned to be the probabilistically favorable deck. If the card from this deck was selected, positive feedback was provided for 80 % of trials and negative feedback was given for 20 % of trials. In contrast, if a card from the probabilistically unfavorable deck was chosen, positive feedback was given on only 20 % of trials and negative feedback was provided on 80 % of trials. This increased the task difficulty, encouraged perseverative responding, and prevented participants from anticipating when an actual contingency reversal would occur. Once the card from the probabilistically favorable deck was selected on eight consecutive trials, irrespective of feedback given to the participant, a contingency reversal occurred. This resulted in a switch of stimulus-reward relations between decks, such that the previously favorable deck became unfavorable and the previously unfavorable deck became favorable. Once the card from the probabilistically favorable deck, based on the new stimulus-reward contingency, was selected on eight consecutive trials, another contingency reversal occurred. Participants continued the task until a total of nine successful reversals were achieved or until they reached a 500-trial deadline. All participants successfully completed the task before deadline, however. Figure 1 presents the experimental procedure.
UNMAPPED
[1] 184w The dopamine overdose hypothesis proposes that the effects of dopaminergic therapy on cognition in PD arise as a consequence of differential endogenous dopamine levels in underlying brain regions (Gotham et al. 1988;Swainson et al. 2000;Cools 2006). In PD, whereas functions mediated by the dopamine-depleted DS are improved by dopaminergic medication, those supported by more dopamine-replete, VTA-innervated brain regions, such as VS and orbitofrontal cortex, are worsened (Cools 2006;MacDonald and Monchi 2011). Dopaminergic therapy has been shown to worsen stimulusresponse (Vo et al. 2014;Hiebert et al. 2014), probabilistic associative (Torta et al. 2009;Jahanshahi et al. 2010), sequence (Feigin et al. 2003;Seo et al. 2010), andlist (MacDonald et al. 2013b) learning, as well as stimulusstimulus facilitation (MacDonald et al. 2011). Stimulusreward and reversal learning are mediated by VTAinnervated brain regions such as VS and ventromedial prefrontal cortex (Cools et al. 2002;Fellows and Farah 2003), and these are the functions most frequently shown to worsen with dopaminergic therapy in PD (Swainson et al. 2000;Cools et al. 2001;Cools 2006;Graef et al. 2010;MacDonald et al. 2013a). For these reasons, stimulus-reward and reversal learning were investigated in the present study.
[2] 156w Testing the straightforward predictions of the dopamine overdose hypothesis in PD patients is mired in complexity. Across and within behavioral studies of PD patients, the degree to which the VTA/SN degeneration is asymmetrical, and the extent to which dopamine-replacement is titrated to maximize SN-innervated brain function at the expense of VTAinnervated operations will be variable. Further, in PD, the mechanisms for regulating dopamine are perturbed, potentially making patients more susceptible to dopamine overdose. That is, levels of the dopamine transporter (DAT), a membrane-spanning protein that clears dopamine from the synapse (Frost et al. 1993), are reduced in PD. In fact, DAT is even decreased in healthy elderly adults (van Dyck et al. 2002). In addition, PD patients have been treated chronically with dopaminergic therapy, which results in post-synaptic dopamine receptor changes, including receptor sensitization (Bordet et al. 1997). Finally, PD patients have cognitive impairment and samples of PD patients will vary in degree of cognitive reserve.
[3] 213w A sample of healthy, young adults is presumably more consistent than typical PD samples that differ in DS dopaminergic optimization and levodopa dosages, as well as in stage and severity of disease. We surmised that evaluating the effects of levodopa administration on reward learning in healthy young adults, who have optimal endogenous dopamine levels, normal DAT expression, no previous exposure to dopaminergic therapy, and intact cognition, provided a critical and more clearly interpretable test of the dopamine overdose hypothesis. Though optimal DAT levels, lack of dopamine receptor sensitization through chronic dopaminergic therapy, and normal cognitive reserve might plausibly have safeguarded against dopamine oversaturation, the dopamine overdose hypothesis was entirely supported by our findings in healthy young adults. In line with the PD literature, we found that a first and single dose of levodopa impaired stimulus-reward learning in healthy young adults overall and during stimulusreward contingency reversals (Fig. 2a-f). Our results allow the following conclusions that could not be achieved by studies with PD patients alone: (1) levodopa-associated, reward learning impairment owes to a main effect of medication, independent of PD pathology and not arising as a function of poor cognitive reserve or a deficient dopaminergic system, and (2) levodopa overdose occurs without prior sensitization of dopaminergic pathways through repeated exposure to exogenous dopamine.
[4] 92w As a caveat, we acknowledge that levodopa in our study and in others could have effects on other neurotransmitters (Arnsten 2007). Though levodopa is a direct precursor to dopamine, dopamine in turn is a precursor to noradrenaline (Nagatsu et al. 1964). We cannot rule out the possibility that some learning-related effects could have been mediated by noradrenaline, acting either centrally or peripherally. Levodopa has been shown to specifically increase dopamine levels, however, and only small changes in the concentration of free noradrenaline following levodopa administration have been noted (Maruyama et al. 1996).
[5] 301w There have been very few investigations of dopaminergic therapy in healthy volunteers to date. The focus of these studies has frequently been to test whether dopamine supplementation ameliorates age-related dopamine decline or whether it can improve cognitive functions in healthy adults (Knecht et al. 2004;Flöel et al. 2005;Floel et al. 2008;de Vries et al. 2010). Further, most have investigated effects of dopamine agonists. Our findings add to a small but growing literature reporting dopamine therapy-related impairments in healthy participants that are analogous to those seen in PD patients. Mehta et al. (2001) found that administration of bromocriptine, a dopamine agonist, improved spatial memory span but worsened reversal learning in healthy adults. Similar impairments following treatment with dopamine agonists occur for probabilistic reward (Pizzagalli et al. 2008), associative (Breitenstein et al. 2006), and reinforcement (Santesso et al. 2009) learning. In the current study, we were specifically interested in levodopa in a participant group with optimally functioning dopaminergic system. Dopamine agonists simulate dopamine at post-synaptic receptors. In contrast, levodopa is converted to dopamine within the brain, which is released, transported, cleared, and recycled by the mechanisms that regulate endogenous dopamine. In this way, testing the effect of levodopa on learning in healthy young participants provided the most direct test of the claim that exogenous dopamine therapy might overwhelm mechanisms for regulating dopamine in the brain, producing impaired function of downstream brain regions. Here, despite efficient dopamine buffering and no dopamine receptor abnormalities related to chronic exogenous dopamine exposure in our healthy participants, we observed clear overdose effects proportional to those seen in PD patients. This suggests that the endogenous dopamine system, at least the VTA system that was tested here, is highly sensitive to overdose by exogenous dopamine. This susceptibility occurs independently of PD pathology, chronic dopamine therapy, and cognitive reserve or fragility.
[6] 127w Stimulus-reward learning and the effects of dopaminergic drugs on this cognitive function are frequently assessed using the reversal learning paradigm. In these studies, contingency changes are signaled by unexpected punishment (Swainson et al. 2000;Cools et al. 2001;Mehta et al. 2001;Cools et al. 2007;Graef et al. 2010), and therefore, it seems an ideal paradigm for investigating learning from punishment. Learning from reward, however, is also important in acquiring and maintaining the stimulus-reward associations prior to and between stimulus-reward contingency reversals. We analyzed our data in a way that allowed us to investigate the effect of levodopa in learning from punishment and reward separately. Selectively looking at trials that followed reward and punishment and using multivariate logistic regression, we found that levodopa impaired both learning from punishment and reward.
[7] 250w Extensive research has shown that dopamine neurons signal rewards and punishment by changing their firing pattern (Schultz 1997;Bayer and Glimcher 2005). These transient bursts (i.e., increases) and dips (i.e., decreases) in dopamineneuron firing are associated with reward and punishment feedback respectively, especially when this feedback is unexpected (Bayer and Glimcher 2005;Robinson et al. 2010). Impairment in learning from punishment by dopaminergic therapy has been shown in a number of studies (Swainson et al. 2000;Cools et al. 2001;Mehta et al. 2001;Frank et al. 2004;Cools et al. 2006;Cools et al. 2007;Moustafa et al. 2008;Bódi et al. 2009). Exogenous dopaminergic medication increases tonic dopamine levels within the striatum. This is hypothesized to occlude dopamine dips necessary for signaling punishment, differentially affecting punishment-versus reward-based learning (Frank 2005). Computational models of reinforcement learning predict that decreased dopamine levels improve punishment relative to reward learning, whereas this bias is reversed when baseline dopamine levels are increased (Frank 2005). This has been supported by empirical investigations (Frank et al. 2004;Cools et al. 2006;Pessiglione et al. 2006;Moustafa et al. 2008;Bódi et al. 2009;van der Schaaf et al. 2014) at odds with our finding of levodopa impairment on both reward and punishmentbased learning. These differences could relate to the population sampled (i.e., healthy volunteers here vs. PD patients in other studies; Frank et al. 2004;Cools et al. 2006;Bódi et al. 2009) or the dopaminergic therapy when only healthy volunteers were considered (i.e., levodopa here vs. dopamine agonists in other studies; Cools et al. 2009;van der Schaaf et al. 2014).
[8] 104w In line with our results, others have also found that reward learning is impaired by dopaminergic therapy (Frank and O'Reilly 2006;Pizzagalli et al. 2008;Santesso et al. 2009). For example, Pizzagalli et al. (Pizzagalli et al. 2008) reported that administration of a dopamine agonist in healthy adults blunted a bias towards selecting more rewarding stimuli compared to placebo. Presumably, enhancing baseline dopamine levels increases inhibitory autoreceptor binding on pre-synaptic dopamine neurons, decreasing phasic dopamine bursts critical for reinforcing reward-related behaviors (Pizzagalli et al. 2008). Alternatively, increasing tonic dopamine levels could decrease the dynamic range, salience, and detectability of phasic dopamine signals (Shohamy et al. 2006).
[9] 57w Currently, the issue of whether exogenous dopamine therapy preferentially impairs punishment-relative to rewardbased learning remains unresolved. Direct comparisons of PD patients and healthy controls, using levodopa versus dopamine agonists, taking into account baseline differences in dopaminergic tone, using the same experimental paradigms might reconcile discrepancies in the literature (Cools et al. 2009;van der Schaaf et al. 2014).