Retesting intelligence tests in quick succession causes score inflation due to practice effects.
the verdict
SUPPORTED
the evidence backs this
refutedsupported
the weight of evidence
5 sources for · 0 against
Peer-reviewed meta-analytic and empirical research indicates that repeating cognitive tests in quick succession produces practice effects that result in score improvements or inflation.
As retest effects in cognitive ability tests have been investigated by various primary and meta-analytic studies, most studies from this area focus on score gains as a result of retesting. To the best of our knowledge, no meta-analytic study has been reported that provides sizable estimates of response time (RT) reductions due to retesting. This multilevel meta-analysis focuses on mental speed tasks, for which outcome measures often consist of RTs. The size of RT reduction due to retesting in mental speed tasks for up to four test administrations was analyzed based on 36 studies including 49 samples and 212 outcomes for a total sample size of 21,810. Significant RT reductions were found, which increased with the number of test administrations, without reaching a plateau. Larger RT reductions were observed in more complex mental speed tasks compared to simple ones, whereas age and test-retest interval mostly did not moderate the size of the effect. Although a high heterogeneity of effects exists, retest effects were shown to occur for mental speed tasks regarding RT outcomes and should thus be more thoroughly accounted for in applied and research settings.
Background The Wechsler Memory Scale-Fourth Edition (WMS-IV) has been widely used to assess memory function in people with dementia. The older adult battery of the WMS-IV includes four indices and seven subtests. The aims of this study were to examine the practice effect and test–retest reliability and calculate the reliable change index modified for practice (RCIp) for the indices and subtests of the older adult battery of the WMS-IV for people with dementia. Methods Fifty-six participants completed the WMS-IV twice, two weeks apart. The practice effect was investigated using effect size (Cohen’s d ) and bootstrapping mixed design analysis of variance while considering the severity of dementia. The test–retest reliability was estimated using intraclass correlation coefficient (ICC). Results The results showed non-significant practice effects with Cohen’s d < 0.20 in different severities of dementia on two indices and five subtests. The ICC values of these indices and subtests were 0.82–0.85 and 0.57–1.00, respectively. The other two indices (i.e., auditory memory and immediate memory) and two subtests (i.e., logical memory delayed recall and visual reproduction immediate recall) demonstrated small to moderate practice effect ( d = 0.46–0.74) for people with mild severity of dementia. Conclusion On the whole, the WMS-IV has no to moderate practice effects and moderate to excellent test–retest reliability in people with dementia. The values of the RCIp with 95% confidence interval for the indices and subtests were provided in this study, which are useful to clinicians and researchers for interpreting the real score change in persons with dementia. The two indices (i.e., auditory memory and immediate memory) and two subtests (i.e., logical memory delayed recall and visual reproduction immediate recall) with noticeable practice effect should be used with caution when assessing memory function repeatedly in people with mild severity of dementia.
Abstract Objective We conducted two empirical studies (in a cross-sectional and a longitudinal design) with the aim at establishing normative data (including norms for strategy use [i.e., clustering and switching strategies] and performance over time), and examining the convergent validity, the test–retest reliability (3–4 wks interval) and the changes in performance with practice (1 year interval) of the different verbal fluency (VF) quantitative and qualitative scores in Spanish-speaking children and adolescents. Method In S1 (n = 620 6- to 15-year-old Spanish-speaking children and adolescents), MANCOVA and Pearson’s correlations were employed. In S2 (n = 148 6- to 12-year-old Spanish-speaking children), intraclass correlation coefficient (ICC), paired t-tests, and Confirmatory Factor Analysis (CFA) were used. Results S1 results showed an age effect on all VF measures (quantitative and qualitative). The number of switches/clusters was more related to total word productivity and to executive functions (EF) than the mean cluster size. In S2, a significant increase in phonological VF performance was observed on number of switches and word productivity scores from baseline (Time 1) to repeat testing at Time 2. Practice effects were observed at Time 3 on all measures except for semantic and phonological mean cluster size. Test–retest reliability coefficients at Time 2 for number of clusters and switches, but not for mean cluster size, fell in the moderate range, ranging from ICCs .61 to ICCs .81. Test–retest reliability coefficients for total word productivity were higher (ICCs above .80) and stronger when testing as a unity with CFA methods (ϕ=.94, p < .001). Conclusions These data may be relevant for informing the neuropsychological assessment of spontaneous cognitive flexibility in children with typical development (TD) and those with developmental or acquired disorders.
Previous studies have indicated that as many as 25% to 50% of applicants in organizational and educational settings are retested with measures of cognitive ability. Researchers have shown that practice effects are found across measurement occasions such that scores improve when these applicants retest. In this study, the authors used meta-analysis to summarize the results of 50 studies of practice effects for tests of cognitive ability. Results from 107 samples and 134,436 participants revealed an adjusted overall effect size of .26. Moderator analyses indicated that effects were larger when practice was accompanied by test coaching and when identical forms were used. Additional research is needed to understand the impact of retesting on the validity inferences drawn from test scores.
Although the subjects in the test-retest and combined reassess and memory conditions reported recalling previous answers for 20-25% of the items on the second test, it was concluded that conscious repetition of specific responses did not seriously inflate the estimate of test-retest reliability. Published in The Journal of general psychology (1992)
Everything we examined (5)
This check searched the claim as stated. It did not run a separate search for evidence against it.