What Is a Work Sample Test?
In short
- A work sample test asks a candidate to actually perform a piece of the job under realistic conditions, then scores the result against a rubric instead of a self-report
- The widely cited validity figure of 0.54 traced back to Hunter and Hunter's 1984 review, which a 2005 meta-analysis of 54 studies and 10,469 people found could not be reproduced
- Roth, Bobko and McFarland (2005) found an observed correlation of 0.26 with job performance, rising to 0.33 once measurement error in the performance rating is corrected for
- That 0.33 lands in the same range as structured interviews and other top-tier selection methods, and the same paper reports work samples showing smaller ethnic-group score gaps than cognitive ability tests
Contents
A work sample test is a hiring assessment that asks a candidate to actually perform a piece of the job, a short, realistic task done under timed and observed conditions, and scores the result against a rubric instead of a self-report or an interviewer's gut read. A customer service candidate answers a mock support ticket. A data analyst cleans a messy spreadsheet. The test does not ask what someone knows about the job. It watches them do a slice of it.
For most of the last four decades, one number followed the format everywhere it was discussed: a validity of 0.54 against job performance, high enough to rank work samples among the best predictors personnel psychology had. That number traces to Hunter and Hunter's influential 1984 review. It also turned out to rest on a narrow, decades-old set of studies that nobody could fully reconstruct.
The work sample test number that got corrected
Roth, Bobko and McFarland (2005), published in Personnel Psychology, went back through the accumulated research rather than taking the 1984 figure on faith, pooling 54 studies and 10,469 people to find an observed correlation between work sample scores and job performance of 0.26, rising to 0.33 with an 80% credibility interval of 0.24 to 0.42 once the standard correction for unreliability in the performance ratings is applied, and to 0.39 once error on both sides of the equation is corrected for, with neither corrected estimate's credibility interval reaching the widely cited 0.54.
The gap is not a rounding difference. The authors call it roughly a third lower than the figure the field had been citing for twenty years, and trace it to newer studies, many from service-sector jobs that were barely represented in the original 1984 data, pulling the pooled estimate down.
See how a scored, recorded task compares to a self-rated skill
Twenty four scenarios, seven minutes, one behaviorally anchored score per domain instead of a resume adjective.
Why it still holds up
A corrected validity of 0.33 sounds like a step down, but it lands in respectable company. Sackett, Zhang, Berry and Lievens' 2022 correction to decades of personnel-selection meta-analyses puts structured interviews at 0.42 and cognitive ability tests at 0.31, with work samples again landing at 0.33, the same figure Roth and colleagues reached seventeen years earlier through an entirely separate route. Two independent corrections, run on different literatures at different times, converge on the same number.
The same 2005 paper also summarizes something a raw validity coefficient does not capture: work sample tests are reported in the literature it reviews to show smaller score gaps between ethnic groups than cognitive ability tests, and candidates tend to rate the format as fairer than a generic interview or a personality inventory, likely because the task in front of them plainly resembles the job itself.
What this means for a candidate
A work sample test is hard to talk your way through, because there is no separate step where a claim gets converted into a number. The task and the score are the same event. That is the same logic behind a situational judgement test, which presents a realistic scenario instead of a real task, and behind scoring either one against a behaviorally anchored rubric rather than a rater's impression. It is also why a self-rated personality trait loses much of its predictive value once an applicant has a motive to look good: a rating can be adjusted before it is given, but a recorded response to a real task cannot be un-performed after the fact. The six-domain check behind this site follows the same rule, scoring what someone actually does with a scenario instead of what they say about themselves.
FAQ
What is a work sample test, in one sentence?
Is a work sample test more accurate than a structured interview?
Why did the accepted validity number for work samples fall so far?
Can a work sample test be gamed the way a resume or a personality inventory can?
Find out where you actually stand.
Six domains, twenty four items, one score. It takes about seven minutes and tells you which parts of your work AI is closest to, and which parts it is not.
Take the checkFree · about 7 minutes · no account