[1]
164w
Context congruency facilitates recognition memory (Tulving & Thomson, 1973;Morris Bransford & Franks, 1977;Craik & Kirsner, 1974;Palmeri, Goldinger & Pisoni, 1993;Bradlow, Nygaard & Pisoni, 1999;Goldinger, 1996;Campeanu, Craik, & Alain, 2013). In memory research, context refers to any number of details related to the initial study episode: location, font, voice and testing conditions are a few examples. In the auditory domain, perhaps the most important contextual cue for spoken word recognition is voice. A review by Pisoni (1993) proposed that speaker voice is encoded in a detailed perceptual representation with the episode itself. One question that emerges, then, is whether voice congruency between study and test acts to improve recognition memory in an all-or-none or in a graded fashion. For instance, voice information may be a gestalt-like "tag" with the word, such that a benefit of voice congruency occurs only for the same speaker. However, another possibility is that voice facilitates spoken word memory via a sense of familiarity, thereby making a partial congruency effect possible.
[2]
19w
In a recent study, Campeanu et al. (2013) investigated the neural correlates of voice reinstatement on spoken word memory.
[3]
108w
Participants were presented with lists of words and were instructed to remember the word as well as the voice of the speaker. During the test phases, participants indicated whether the word was old or new and whether the old words were spoken by the same or a different speaker between study and test. There were four different speakers, resulting in a combination of gender and accent (Chinese or Canadian) congruency. Participants were more accurate in recognizing old words spoken by the same speaker during study and test blocks (i.e., a same-speaker voice reinstatement benefit). There was also an indication that accent congruency conferred a partial word recognition benefit.
[4]
107w
In the present study, we used a continuous recognition paradigm to assess the impact of voice congruency on spoken word recognition. We sought to extend the characterization of a voice congruency effect by using a different paradigm than that used by Campeanu et al. (2013), but one that also allows a direct comparison with previous behavioral research (Craik & Kirsner, 1974;Palmeri et al., 1993). While our previous investigation used a block design (Campeanu et al., 2013), and emphasized attention to words and to voice at study (since both were directly tested), the current experiment asked participants to make word judgments alone, with no instruction regarding attention allocation.
[5]
133w
With respect to voice congruency, the second presentation of each word may be spoken by the same voice, a voice of different gender but the same accent, a voice of the same gender but different accent, or a voice of different gender and different accent. We anticipated overall superior performance for old words spoken in the same voice. Moreover, though voice effects might be expected to decrease over time or with the number of intervening items, a complete absence of a same-voice advantage would be expected only at delays exceeding one day (Goldinger, 1996), a far more extended time course than we investigate in the present experiment. In addition, we predicted overall memory performance to be inversely correlated with the number of items and time between the first and second presentation (i.e., lag).
[6]
163w
Prior studies (e.g., Wilding, Doyle, & Rugg 1995;Wilding & Rugg, 1996;Tendolkar et al. 2004), including some using a continuous recognition paradigm (i.e., Nielsen-Bohlman & Knight, 1994), have revealed a late positivity at parietal sites that is thought to index recollective processes. This modulation has been described by Tendolkar et al. (2004, p. 236): "[t]he first old/new effect identified during tests of recognition memory onsets approximately 400 ms post-stimulus, typically lasts around 400-600 ms, and is largest in amplitude over left temporo-parietal scalp electrodes". In our previous investigation (Campeanu et al., 2013), the voice congruency effect on memory performance coincided with the presence of this well-established ERP modulation of recollection during the word recognition test. In particular, correctly recognized old words produced a more positive deflection over parietal sites than new words, with the most positive wave occurring when old words had the voice reinstated at test. In the present investigation, then, we predict a sustained left parietal positivity for correctly recognized old words.
[7]
284w
However, recollection is only one of two processes known to affect recognition memory. Yonelinas (1994) reviewed a dualprocess model for recognition memory, in which familiarity and recollection both play a role during retrieval. Familiarity, a general strength marker that something has previously occurred, is dissociated from recollection, which usually involves retrieval of contextual details of an episode. In terms of imaging, ERP studies have typically found two distinct deflections, which were initially thought to reflect familiarity and recollection, respectively (e.g., Curran, 2000). Familiarity has been linked to a frontal positivity for old words compared to new words, between 300-500 ms (Curran, 2000;Schloerscheidt & Rugg, 2004;Curran & Dien, 2003). However, more recent research has suggested that this frontal modulation actually indexes conceptual priming rather than familiarity (Voss & Paller, 2006;Voss & Paller, 2007;Paller, Voss, & Boehm, 2007), and that the two concepts are often confounded in memory research. The present experiment does not include a behavioral measure of conceptual priming, so we cannot conclude that our participants were conceptually primed for individual words. However, it does seem reasonable to suggest that old words presented at test in the same voice as at study engage perceptual priming mechanisms. Therefore, comparing the voice congruency conditions for correctly recognized old trials presents an opportunity to investigate whether this frontal priming modulation varies as a function of perceptual congruency in a purely auditory experiment. The present investigation is the first to our knowledge to use auditory stimuli both at study and at test in an attempt to investigate this frontal deflection. If the voice congruency effect involves priming to some degree at least, we would expect to find a frontal modulation distinguishing between same-voice repetitions and different speaker trials.
[8]
82w
Since we employ a continuous recognition paradigm, a further important objective of the present study was to investigate the effect of lag on recognition memory, and more specifically on potential voice congruency effects in word recognition. In the current experiment, we manipulated lag within each voice congruency condition, so that old words were re-presented at 2, 8 or 16 words after initial presentation. This allowed us to extend our investigation to determine how long the voice congruency benefit on word memory lasts.
[9]
100w
It is important to note that lag effects may reflect trace decay over time and/or active interference from intervening items. Regardless of the mechanism, we predict a behavioral benefit of voice reinstatement at test, as well as a decrease in performance at extended lags. In terms of the ERP trace, the amplitude of the late positive parietal (old/new) effect has been found to decrease with increasing lag (Nielsen-Bohlman & Knight, 1994). Hence, we expect electrophysiological findings to show a sustained positivity over parietal sites that decreases with increasing lag and an early frontal modulation that is sensitive to voice congruency.
[1]
202w
Overall, participants correctly identified 88.0% (SD ¼ 5.4%) new words and 89.3% (SD ¼5.0%) old words. For analysis, a false alarm rate was calculated for each participant; the group mean false alarm rate was 12%. Fig. 1a shows the group mean accuracy (hit rates) for all four voice conditions at each of the three lags in the word recognition test; all accuracy analyses were conducted using hit rates. The ANOVA yielded a main effect of voice congruency, F (3,42) ¼15.63, p o0.001, η 2 ¼0.53, with pairwise comparisons indicating that performance in the same speaker condition was superior to all three different speaker conditions (po 0.001). Performance did not significantly differ among the three different speaker conditions (p¼ 0.80 for different gender/same accent vs. same gender/different accent; p ¼0.81 for different gender/same accent vs. different gender/different accent; p ¼0.98 for same gender/different accent vs. different gender/different accent). There was also a main effect of lag, F(2,28) ¼45.57, p o0.001, η 2 ¼ 0.77, with pairwise comparisons indicating that performance at all three lags were significantly different from each other (p o0.001). There was an interaction that approached significance between voice condition and lag, F(6,84) ¼ 2.42, p ¼0.058, η 2 ¼ 0.15.
[2]
472w
Since the interaction between voice condition and lag closely approached significance, separate ANOVAs were then calculated at each lag, with voice condition as the within-subject factor. At lag 2, the effect of voice condition was not significant, F(3,42)¼1.87, p¼0.16, η 2 ¼0.12. At lag 8, voice condition produced a significant effect, F(3,42)¼ 10.61, po0.001, η 2 ¼ 0.43. Pairwise comparisons indicated that performance in the same speaker condition was greater than in the other three voice conditions (all pr0.001) but that performance among the three different speaker conditions did not vary significantly (p40.35 in all cases). At lag 16, voice condition again produced a significant main effect, F(3,42)¼5.54, p¼0.004, η 2 ¼0.28. Once again, pairwise comparisons indicated that performance in the same speaker condition was greater than in the other three voice conditions (p¼ 0.006 for same speaker vs. different Fig. 1. A) Group mean accuracy and B) Group mean reaction time. Note that accuracy is represented by hits for each lag and speaker condition (i.e., correct recognition on "old" trials). Error bars represent within-subjects standard error of the mean. Fig. 2. Panel of electrodes. Waveforms at numerous scalp locations. Note that the "Old Words -Different Speakers" wave refers to an average of all the different speaker trials that were correctly recognized to be old at test. gender/same accent; p¼ 0.009 for same speaker vs. same gender/ different accent; p¼0.001 same speaker vs. different gender/different accent) but that performance among the three different speaker conditions did not vary significantly (p40.70 in all cases). Fig. 1b shows RT for correct old trials in all four voice conditions at each of the three lags. There was a main effect of voice condition, F(3,42) ¼11.44, p o0.001, η 2 ¼ 0.45. Pairwise comparisons revealed shorter RTs for the same speaker condition relative to the other three voice conditions (p ¼0.019 for same speaker vs. different gender/same accent (though this did not reach statistical significance using the Holm-Bonferroni method for multiple comparisons); po 0.001 for same speaker vs. same gender/different accent; p ¼0.002 for same speaker vs. different gender/different accent). Also, RTs on the different gender/same accent trials trended toward being faster than those on the same gender/ different accent (p ¼0.013, which did not reach statistical significance using the Holm-Bonferroni method for multiple comparisons) and the different accent/different gender trials (p ¼0.049, which did not reach statistical significance using the Holm-Bonferroni method for multiple comparisons). In addition, the ANOVA indicated a main effect of lag, F(2,28) ¼113.93, p o0.001, η 2 ¼0.89, with pairwise comparisons showing that participants responded faster at lag 2 than at lags 8 and 16 (both p o0.001). There was no significant difference in RT between lags 8 and 16, p ¼0.51. Finally, there was no significant interaction between lag and voice condition, F(6,84) ¼0.29, p ¼ 0.86, η 2 ¼0.02.
[3]
69w
In summary, behavioral results indicated that the same speaker condition generally produced the best accuracy and the shortest RT among the voice conditions. However, the benefit to accuracy of reinstating voice was significant only at longer lags (8 and 16), but not at the short lag (2). RTs also indicated faster recognition when accent was congruent between study and test. Overall, accuracy decreased and RT increased with increasing lag.
[4]
41w
The ERPs comprised N1-P2 modulations and a late positive complex over parietal sites. Waveforms at multiple electrodes are shown in a panel graph (see Fig. 2). The exogenous ERPs (i.e., P1, N1, P2) were not affected by voice condition or lag.
[5]
99w
The first analysis using BESA Statistics 1.0 focused on old and new words. It revealed a significant difference at parietal and frontal sites (p o0.001; see Fig. 3). At parietal and occipital sites, this modulation was characterized by an increased positivity for old words. This old/new effect began at about 380 and peaked at about 700 ms after word onset. The polarity of this modulation was inverted at frontal sites. Of note, there was no early frontal positivity associated with old words (as opposed to new words), a modulation traditionally thought to reflect an old/new familiarity effect (Curran, 2000).
[6]
301w
In a second analysis, we examined the effect of voice reinstatement on correctly identified old trials. Here, we collapsed across all lags and used BESA Statistics 1.0 to compare same speaker and (collapsed) different speaker trials, since we were looking for an indicator of how repetition of the same surface form affects recognition. The analysis revealed a significant difference between the same speaker condition and the collapsed different speaker conditions. The same speaker condition was consistently more negative over left frontal sites (F7, F5, F3, AF7, AF3, FC5, F1, FP1, FT9, F11) between approximately 200 and 650 ms post-stimulus (po0.001; Fig. 4). This deflection appeared to inflect over central and right parietal sites, causing an inversion (where same speaker trials were more positive than different speaker trials) between approximately 400-500 ms post-stimulus over CP1, CPz, CP2, P1, Pz, P2, PO3, PO4 and Oz (po0.001). Since the data-driven analysis produced significant differences between same speaker and collapsed different speaker conditions as early as 200 ms over the anticipated frontal area (see Curran, 2000;Curran & Dien, 2003), we then conducted ANOVAs to look for variations between the three different speaker conditions at the timeframes and sites indicated. Two ANOVAs were conducted. A comparison of the mean amplitude for 200-650 ms interval, over the left frontal scalp region (F7, F5, F3, AF7, AF3, FC5, F1, FP1, FT9 and F11) between the different gender/same accent, same gender/different accent and different gender/different accent conditions revealed no main effect of voice, F(2,28)¼1.04, p¼ 0.37, η 2 ¼0.07. For 400-500 ms interval, a comparison of the mean amplitude for the inversion at central-parietal and parieto-occipital scalp region (CP1, CPz, CP2, P1, Pz, P2, PO3, PO4 and Oz) between the three different speaker conditions did not yield a main effect of voice, F(2,28)¼2.66, p¼ 0.089, η 2 ¼ 0.16.
[7]
87w
Since behavioral results indicated a main effect of lag, specifically that there was a significant difference between lags 2, 8 and 16, we collapsed across voice conditions to look at where and when this main effect would manifest over the scalp. Prior research has indicated that short lag trials generated shorter latency for a positive deflection over central and parietal sites at 550-650 ms (Nielsen-Bohlman & Knight, 1994); therefore, we were particularly interested in lag effects as a function of peak latency of the old/new parietal effect.
[8]
109w
We measured peak latency of the old/new parietal effect during the 600-800 ms interval over the left (PO3, P5, P3, P1) and right parietal scalp region (PO4, P6, P4, P2). The main effect of lag was significant, F(2,28) ¼8.29, p ¼0.003, η 2 ¼ 0.37, with pairwise comparisons indicating that lag 2 peaked earlier than lags 8 (p¼ 0.002) and 16 (p ¼0.019; see Fig. 5). The peak amplitude also produced a main effect of lag, F(2,28) ¼ 14.09, p o0.001, η 2 ¼ 0.50, with pairwise comparisons indicating that lag 2 generated a more positive deflection than both lags 8 and 16 (p ¼0.002 and p o0.001, respectively).
[9]
238w
By collapsing trials at lags 8 and 16, we were able to compare old trials on short (lag 2) vs. long (lags 8/16) lag conditions. This analysis was conducted because behavioral findings indicated that voice congruency effects differed between short and long lags. Specifically, a main effect of voice congruency was found for accuracy at long lags, but not at short lags. This led us to believe that voice congruency effects might manifest differently in the ERP trace, based on grouping of correctly recognized words at short vs. long lags. Indeed, a preliminary analysis using BESA Statistics 1.0 indicated that short lag trials produced a more positive modulation over parietal sites from about 370-770 ms post- Fig. 6. Panel of electrodes with old trials separated by lag. Waveforms at numerous scalp locations. Note that "Lag2 -Different Speakers" and "Lags8and16 -Different Speakers" waves both refer to an average of all the different speaker trials (at the corresponding lag(s)) that were correctly recognized to be old at test. stimulus, peaking at approximately 650 ms (p o0.001), which was inflected over frontal sites around 400-700 ms (po 0.001). As such, we conducted separate analyses on short and on long lags using BESA Statistics 1.0, looking at the same speaker vs. collapsed different speaker conditions each time. These voice reinstatement contrasts at short and at long lags are shown with correctly identified new words over a panel of electrodes in Fig. 6.
[10]
130w
For the short lag, there was a significant difference between the two voice conditions at approximately 450-700 ms, with same speaker trials producing a more negative modulation than different speaker trials at left frontal-central sites (FC1, FCz, FC2, Fz, F1, F3, F5, F7, FC5, AF3, AF7), p o0.001; see Fig. 7. This inflected over central parieto-occipital sites, where same speaker trials produced a more positive wave than different speaker trials at 450-700 ms at P1, P4, P6, P7, P8, PO3, POz, PO4, O1, Oz, O2, and Iz electrodes (p o0.001). The same speaker trials at lag 2 also produced a more negative modulation at 800-1150 ms, over fronto-central, central, and centro-parietal sites (F1, Fz, FC1, FCz, FC2, C1, Cz, C2, C4, CP1, CPz, CP2, Pz; p o0.001; see Fig. 8).
[11]
103w
Once again, in an effort to discern pairwise comparisons that would indicate more fine-grained voice congruency differences, we followed the data-driven analysis with ANOVA comparing mean amplitudes for three different voice conditions, at the times and sites indicated. First, an ANOVA comparing the three different speaker conditions at 450-700 ms over the left frontal-central scalp region (FC1, FCz, FC2, Fz, F1, F3, F5, F7, FC5, AF3, AF7) did not yield a main effect of voice condition, F(2,28) ¼1.25, p ¼0.30, η 2 ¼0.08. The inversion at parieto-occipital sites (P1, P4, P6, P7, P8, PO3, POz, PO4, O1, Oz, O2, Iz) approached significance, F
[12]
81w
(2,28) ¼3.02, p¼ 0.067, η 2 ¼0.18, and pairwise comparisons did indicate that the different gender/same accent condition produced a more positive deflection than the same gender/different accent condition (p ¼ 0.044). Finally, an ANOVA comparing the three different speaker conditions for the 800-1150 ms interval over fronto-central, central and parietal scalp region (F1, Fz, FC1, FCz, FC2, C1, Cz, C2, C4, CP1, CPz, CP2 and Pz) produced no main effect of voice condition, F(2,28) ¼2.34, p ¼0.13, η 2 ¼0.14.
[13]
144w
For the long lags, the data-driven analysis showed a significant difference wherein the same speaker condition produced a more positive modulation than the collapsed different speaker condition for 230-290 ms interval at PO3, P1, Pz, P2, CP1, CPz, CP2, C1, C2, C4, C6 and CP6 electrode, and again for 400-500 ms interval at CP1, CPz, CP2, P1, Pz, P2, P5, PO3 and PO4 electrodes (both p o0.001; see Fig. 9). This effect was inverted, and the same speaker condition generated a more negative modulation, over frontal sites. This was manifested as two significant differences: at 200-500 ms over left frontal sites (F7, F5, F3, FC5, AF3, AF7, FP1; po 0.001; see Fig. 4, since this effect was almost identical to the one collapsed across all lags); and at 200-300 ms over bilateral frontal sites (FT9, F7, AF4, F11, FP2; p ¼0.003; see Fig. 10).
[14]
71w
Again, we compared three different speaker conditions using ANOVAs for each of the indicated (significant) times and sites. At 230-290 ms, there was no main effect of voice condition, F(2,28) ¼ 0.25, p¼ 0.73, η 2 ¼0.02. At 400-500 ms, there was no main effect of voice condition, F(2,28) ¼1.81, p ¼0.19, η 2 ¼0.12. At 200-500 ms over left frontal sites, there was no main effect of voice condition, F
[15]
34w
(2,28) ¼ 0.22, p ¼0.77, η 2 ¼ 0.02. Lastly, at 200-300 ms over bilateral frontal sites, there was no significant main effect of voice condition, F(2,28) ¼ 1.43, p ¼0.26, η 2 ¼0.09.
[16]
163w
For old words, the same speaker condition produced a higher accuracy score and a shorter reaction time than other three voice conditions. While accuracy between the three different speaker conditions differed in the study of Campeanu et al. (2013), there were no significant differences in accuracy of the different speaker conditions in the present experiment. There are two possible explanations for this. First, there were differences in task instructions; in our previous study, participants were told to attend to words as well as to which speaker uttered each word, since they would be tested on both aspects. In the present study, participants were told to judge each word as old or new regardless of whether the voice was the same as the first time they had heard the word. Hence, the task instructions in the present study placed less emphasis on remembering the speaker voice. This suggests that attention to speaker voice may be important for the gradient in voice congruency to emerge.
[17]
300w
Indeed, past research has looked at the role of attention in auditory recognition. According to Goldinger (1996Goldinger ( , p.1177)), "it appears that the content of study words' episodic trace is influenced by the focus of attention at study". By this account, one would expect voice congruency effects when participants are told to pay attention to both word and voice at study, but perhaps to a lesser extent, or not at all, when they are not instructed to pay attention to the voice. Palmeri et al. (1993) has previously used a continuous recognition paradigm in which voice was not attended and found a same speaker advantage but no benefit when old words were presented in voices of the same gender as at study. Goldinger (1996), on the other hand, did find a partial same gender voice congruency benefit, both when voice was specifically attended at study and even in some conditions when participants were not asked explicitly to attend to voice. However, this difference might be explained by different designs, whereby participants in the latter experiment typed out the entire word during a study phase prior to a surprise recognition test. The more thorough encoding of each word (i.e., typing it out versus simply listening) may have facilitated a stronger association between voice and word at study in the Goldinger study, resulting in finegrained voice congruency effects even when voice was not explicitly attended at study. The procedure involved in the present study was more similar to the Palmeri et al. study. Based on our results and previous research, it seems reasonable to suggest that a same-voice reinstatement benefit occurs even in the absence of attention being allocated to voice, while a more fine-grained voice congruency benefit requires attention to voice or some other form of emphasizing voice at study.
[18]
125w
It is important to note that study instructions did not indicate a need to respond as quickly as possible, yet there were still significant effects of lag and voice condition on RT. Participants generally responded faster in the same speaker condition, and there was also some indication of a speed benefit when accent was congruent between study and test. It should, however, be noted that since overall accuracy was much higher at lag 2, the potential differences among voice conditions at this lag were very likely masked by a ceiling effect. In addition, the relatively small sample size in the current experiment might have limited potential voice effects, especially since there is a graphic trend to a voice effect at lag 2 (see Fig. 1a).
[19]
58w
Therefore, though behavioral results provide indications of a same-voice facilitation at all lags, further investigations are needed to better characterize potential voice congruency benefits at short lags in particular. Further research is also needed to assess whether directing attention to an aspect of voice (i.e., gender or accent), or familiarity training with a voice, influences this facilitation pattern.
[20]
153w
The parietal old/new effect showed more positive waves for the old word conditions and a most negative wave for the new word condition, as expected. However, no voice congruency effects were found (and no systematic gradient between the different voice conditions of the old word trials). This parietal old/new effect has previously been interpreted as a combination of two effects, reflective of recollection and of retrieval demands (Tendolkar et al., 2004). Given that task instructions did not stress attention to voice, it is perhaps a reasonable finding that the strength or specificity of recollection was not reflected in voice congruency in the parietal old/new effect. However, since we did not manipulate attention allocation in the present study, further research is needed to clarify whether attention is, indeed, a pivotal factor. Alternatively, it may be that the parietal modulation deals with amodal, lexical information in recognition memory and is largely insensitive to perceptual attributes.
[21]
406w
Conversely, same-voice reinstatement did modulate ERP activity over frontal and parietal regions. Though frontal deflections have been associated with conceptual priming, the present findings suggest that perceptual priming may also be indexed by an early frontal modulation. When collapsed across all lags, voice reinstatement was associated with a frontal negativity and a parietal positivity beginning as early as 200 ms. The presence of this effect suggests that voice reinstatement modulates recognition memory even when it is not expressly attended, perhaps via priming mechanisms. It seems reasonable to argue that the perceptual fluency associated with same-voice repetition may yield a form of implicit priming that does not support the experience of conscious recollection. Though we observed electrophysiological markers of voice reinstatement, no more fine-grained voice congruency effects were reliable. This could be related to the fact that voice was not explicitly attended or more extensively processed in the current paradigm. It is interesting to note that our findings over frontal sites indicate a most negative deflection for the same speaker condition. Since the present experiment utilized only auditory stimuli and the analysis conducted investigated perceptual priming for old words only, it is possible that the sources underlying any neural activity observed are different than those previously found, thereby resulting in a potentially different orientation and/or location of ERP modulations. It seems plausible that, since previous studies that investigated modulations similar in time course and scalp topography were qualitatively very different from the current investigation, e.g., in terms of modality, the task itself, and method of analysis (see Curran, 2000;Schloerscheidt & Rugg, 2004;Curran & Dien, 2003), this type of priming might simply generate different sources of activity for the present task in the auditory domain. The positioning and orientation of the sources underlying this purely auditory, perceptual effect require further study. This continuous paradigm meant that test words were successfully encoded 2-, 8-or 16-back (that is, with 1, 7 or 15 words intervening between first and second presentations). There was a main effect of lag both behaviorally and in the ERP trace. The significant behavioral results showed that participants were better and faster with old/new judgments at lag 2, and this was mirrored in peak latency measurements of the parietal old/new effect, where lag 2 peaked higher and significantly earlier. Comparing the short and long lag conditions, ERP results indicated that the use of recollection in recognition memory is stronger (i.e., more positive parietal deflection) at short lags.
[22]
141w
Since behavioral analysis indicated that voice congruency behaved differently at lag 2 as opposed to lags 8 and 16, we conducted further analyses on the ERP data for short (lag 2) vs. long (lags 8 and 16) lag conditions. Results indicated that a voice reinstatement benefit was present at both short and long lags, but that more fine-grained voice congruency effects were absent using the current paradigm. At long lags, there was a voice reinstatement benefit associated with frontal negativity and parietal positivity as early as 230 ms. Interestingly, at the short lag only, the same pattern was observed but beginning only around 450 ms. We speculatively suggest that voice reinstatement plays little part at lag 2, since words are still in one's conscious awareness at this time and have not yet been removed from working memory due to intervening items.
[1]
115w
Nineteen participants provided written informed consent according to the guidelines set out by the Baycrest Centre and the University of Toronto. EEG data from three participants were excluded because of excessive muscle artifacts and/or eye movements during recording. Data from one participant was excluded because he did not follow task instructions. The final sample of 15 participants comprised 6 males and 9 females aged between 22 and 33 (M ¼ 24.47, SD ¼ 3.38 years). All participants learned English as their first language, were right-handed and had pure-tone thresholds within normal limits for frequencies ranging from 250 to 8000 Hz (both ears), which was verified by an audiogram performed at the beginning of the session.
[2]
289w
The word list consisted of 336 high-frequency (30 þ from Kucera & Francis, 1967), two-syllable nouns, taken from the MRC psycholinguistic database (Wilson, 1988). All words were recorded by four speakers: one native-English female, one native-English male, one Chinese-accented female and one Chinese-accented male in continuous streams, at 32,000 Hz sampling rate, mono, with 16-bit resolution. The native English-speaking female was 31 years old at the time of recording, and the native English-speaking male was 27 years old. The Chinese-accented female was 57 years old and had learned English at 27 years of age. The Chinese-accented male was 44 years old and had learned English at 19 years of age. With a Shure KSM44 microphone and an USBPre preamplifier with digitizer, speakers recorded the words using Adobe Audition 1.5 on a Dell laptop. All words were then spliced into individual files, using a MATLAB (The MathWorks, Inc., http://www.math works.com/, version 5.3) script, and their amplitudes were normalized to a standard dB level (the average of all words). The MATLAB script used to splice the words specified thresholds, duration and pre-stimulus intervals, which could be varied as necessary for different groups of words. Words were inspected for background noise, clipping, duration (1 s) and clarity. Editing was done as appropriate, either manually or automatically, using batch files. To check for intelligibility, two young adults with English as a first language each listened to three blocks of words and indicated which words were difficult to understand. Reshuffling and processing of new words occurred as necessary to ensure that potential participantsyoung adults with English as a first languagewould have no trouble understanding the spoken words. Once all words were judged adequate, they were normalized to their average loudness, at À 32.44 dB.
[3]
243w
The experiment was programmed using Presentation software (Neurobehavioral Systems, http://www.neurobs.com/, version 14.5) and custom MATLAB code (version 7.1). On each trial, participants heard one word, presented binaurally through insert earphones (EAR-TONE 3a), at 70 dB SPL. Words (digitized as .wav files) were converted to analog using a computer soundcard (16 bit, stereo, with a sampling rate of 44,100 Hz). The analog output was fed into a 10 kHz filter (Tucker Davis Technologies (TDT, Alachua, FL), FT6-2), and then to a GSI 61 audiometer. The inter-stimulus interval (ISI) was jittered between 2.8 and 3.2 s (33 or 34 ms steps, rectangular distribution). Participants indicated whether or not they had previously heard the current word, by making a button press corresponding to "old" or "new" with their left or right forefinger, respectively. They were asked to judge a word as "old" if they had heard it earlier, irrespective of whether the test word was presented in the same or a different voice as its previous occurrence. They were instructed to respond as accurately as possible on each trial. As such, it is important to note that the voice manipulation was incidental in this experiment; the participant's task was simply to judge whether the current word was old or new. The experiment consisted of two 23-min blocks, each comprising 348 trials (one word per trial). No words were repeated across blocks. Participants were given a mandatory break, lasting at least two minutes, between the two blocks.
[4]
167w
In this experiment, we manipulated two factors: voice condition and lag. There were four voice conditions: same speaker, different gender/same accent, same gender/different accent, and different gender/different accent. The three lag conditions denoted the interval between pairs of words, in which the first word was presented 2-(i.e., 7.6-8.4 s delay), 8-(30.4-33.6 s delay), or 16-back (60.8-67.2 s delay). Since these various lag conditions were randomly interleaved within each block, we created a custom MATLAB program to determine the fewest number of extra "filler" trials that would allow for 144 word pairs of interest (48 per lag, within each block) to be presented at each lag. This process generated 10 possible lag templates. The 348 trials per block consisted of 288 "new"/"old" words of interest (144 pairs), 24 early "filler" trials (12 pairs) used to build up a memory store, and 36 "filler" trials throughout (14 pairs separated by lengthy lags and 8 unpaired words). The probabilities that a word was old or new were therefore roughly equal.
[5]
430w
Custom MATLAB code was used to create unique stimulus lists for each participant and block, based on two of the ten lag templates (one per block), which were chosen randomly for each participant. The word pairs were randomly ordered within each lag for each participant. The number of trials per lag and voice condition were counterbalanced as maximally as possible. In each block, within 144 word pairs of interest, there were exactly 36 pairs within each voice condition per block (72 total pairs for each participant) collapsed across all three lags, and approximately 48 pairs (range: 47 to 49) within each lag per block (approximately 96 total pairs [ranging from 94 to 98 pairs across participants], collapsed across the four voice conditions). Within each voice condition, there were 23 to 25 pairs per lag, averaging to approximately 24 pairs per lag across participants. We initially collapsed over lag, since the number of observations per lag was small (at each lag, the number of trials per speaker condition was a maximum of 25 trials). However, visual inspection of ERP traces allowed us to notice that lag 2 often behaved differently than lags 8 and 16, which were more similar. In addition, behavioral analysis of hits at separate lags (see Section 3) indicated that performance at lag 2 followed a different voice congruency pattern than at lags 8 and 16. As such, we conducted further analyses on ERP data to see if a difference between short (2) and long (8, 16) lags emerged for different voice conditions. Moreover, since performance was superior at lag 2, we were able to get a reasonable number of trials to examine the interaction between voice condition and lag, in this way. At lag 2, the mean number of hits for the same speaker condition was 23.47 (SD ¼ 0.83). It was 23.33 (SD ¼1.05) for the different gender/same accent condition, 22.87 (SD¼ 0.92) for the same gender/different accent condition and 22.60 (SD ¼1.24) for the different gender/different accent condition. At lag 8, the mean number of hits for the same speaker condition was 23.33 (SD¼ 1.18). It was 20.53 (SD¼ 2.36) for the different gender/same accent condition, 20.40 (SD¼ 2.59) for the same gender/ different accent condition and 20.40 (SD ¼2.23) for the different gender/different accent condition. At lag 16, the mean number of hits for the same speaker condition was 21.40 (SD¼ 2.53). It was 19.27 (SD ¼ 3.15) for the different gender/same accent condition, 19.53 (SD ¼ 2.61) for the same gender/different accent condition and 19.60 (SD ¼1.50) for the different gender/different accent condition.
[6]
156w
Hit rates were calculated for the four possible voice conditionssame speaker, different gender/same accent, same gender/different accent, different gender/ different accentwithin each lag (2-, 8-, or 16-back). The false alarm rate reported is a group mean rate, averaged from individual common rates (i.e., across all new trials)based on the proportion of new words incorrectly judged as oldfor each participant. Since voice conditions represented levels of congruency rather than individual voices, it was not possible to calculate false alarm rates per condition; instead, a common rate was calculated for each individual. As such, using hits minus false alarm rates in the analysis of voice congruency effects would yield no additional information compared to using hit rates alone. Therefore, accuracy analyses were calculated based on hit rates in the current experiment. In addition to hit rates, we computed response times (RTs; ms relative to word onset, using only correct old trials) for each voice condition within each lag.
[7]
54w
The effect of context congruency on continuous word recognition was assessed using repeated measures ANOVAs with voice condition and lag as the withinsubject factors, using IBM SPSS Statistics (version 20). The Greenhouse-Geisser p-values are reported if the sphericity assumption was violated. Post-hoc pairwise least significant difference (LSD) comparisons were done following significant main effects.
[8]
93w
The electroencephalogram (EEG) was digitized continuously (sampling rate 500 Hz; bandpass filter of 0.05-100 Hz) during the study phase and the test phases using NeuroScan Synamps2 (Compumedics, El Paso, TX, USA). The ERPs were sampled at 64 scalp locations that include electrodes placed at the outer canthi and at the inferior orbits to monitor eye movements. During recording, all electrodes were referenced to the Cz electrode; for off-line data analysis, they were rereferenced to an average reference. The analysis epoch consisted of 200 ms of prestimulus activity and 1300 ms of post-stimulus activity.
[9]
85w
For each participant, a set of ocular movements was obtained prior to and after the experiment (Picton et al. 2000). A MATLAB program was used to calculate averaged eye movements for both lateral and vertical eye movements as well as for eye-blinks. A principal component analysis of these averaged recordings provided a set of components that best explained the eye movements. The scalp projections of these components were then subtracted from the experimental ERPs to minimize ocular contamination, using Brain Electrical Source Analysis (BESA 5.2).
[10]
68w
After correcting for eye movements, all experimental files for each participant were then scanned for artifacts; epochs including deflections exceeding 100 mV were marked and excluded from the analysis. The remaining epochs were averaged according to electrode position and trial type, using BESA 5.2. Each average was baseline-corrected with respect to the pre-stimulus interval and digitally low-pass filtered at 20 Hz (zero phase, 24 dB/oct), using BESA software.
[11]
273w
BESA Statistics 1.0 was the primary means of ERP analysis. Two conditions could be compared at a time, over all scalp regions and up to 1300 ms poststimulus. A two-stage analysis first computed a series of t-tests that compared the ERP amplitude between the two conditions at every time point. This identified clusters in time (adjacent time points) and space (adjacent electrodes) where the ERPs differed between the conditions. In the second stage of this analysis, permutation tests were performed on these clusters. The permutation test used a bootstrapping technique to determine the probability values for differences between conditions in each cluster. The final probability value computed was based on the proportion of permutations that were significant for each cluster, and implicitly corrected for multiple comparisons. In each of the current analyses, we used a cluster alpha of 0.05, 1000 permutations and clusters defined using a channel distance of 4 cm, which resulted in an average of about 5.08 neighbors per channel. Using this technique, we computed 4 comparisons. First, we compared all correctly identified "old" trials with all correctly identified "new" trials. Then we sought to investigate the role of voice reinstatement in word recognition. To that end, we compared the same speaker condition (voice reinstatement) with the collapsed different speaker condition (which was an unweighted combination of the different gender/same accent, same gender/different accent and different gender/different accent conditions), always using words that were correctly identified as old. Lastly, we investigated the role of voice reinstatement at short (lag 2) and then at long lags (lags 8 and 16), both times comparing the same speaker and the collapsed different speaker conditions.
[12]
123w
When the results of the data-driven analysis indicated voice reinstatement effects, we then exported mean amplitude measurements and analyzed three different speaker conditions using repeated measures ANOVAs with voice condition as the within-subject factor. This was done to identify whether significant voice reinstatement effects reflected more fine-grained voice congruency effects between three different speaker conditions. In addition, when looking at the effect of lag, trials were collapsed across voice conditions and then peak amplitude and latency measurements of the parietal old/new effect were exported and analyzed using repeated measures ANOVA with lag as the within-subject factor. Again, all deflections were measured using correct trials only. For all ANOVAs reported, results of the pairwise comparisons were corrected for multiple comparisons using the Holm-Bonferroni method.