How AI Job Risk Scores Are Calculated, and Where They Break
In short
- Job level risk scores come from breaking an occupation into tasks and rating each task for automatability
- The famous headline numbers are probabilities about categories, not predictions about you
- A score for your job title cannot see how you personally do the work
Contents
An AI job risk score is almost never a measurement of you. It is a measurement of your job title, produced by breaking that title into a list of tasks, rating how automatable each task looks, and adding the ratings back up. Understanding that one sentence tells you most of what these tools are good for, and exactly where they stop being useful.
The tools are everywhere now. Type in an occupation, receive a confident percentage, feel briefly unwell. Very few of them show their working. The underlying method is not a secret though, and it is worth knowing, because the difference between a score about a category and a score about a person is the whole game.
Where the numbers come from
Nearly every public risk score traces back to the same starting material: a structured database of what occupations actually involve. In the United States that is O*NET, which decomposes thousands of occupations into tasks, work activities, required knowledge, and skills. It is a genuinely impressive piece of public infrastructure, and it is the raw feedstock for this entire genre.
From there the method is consistent. Take the task list for an occupation. Decide, for each task, how susceptible it is to automation given some assumption about what the technology can do. Aggregate. Publish a number.
The interesting part is step two, because that is where the judgement lives. The influential 2013 Oxford Martin School paper by Frey and Osborne had experts hand-label a subset of occupations as automatable or not, then used those labels to train a classifier that extended the estimate across the rest of the labour market. That is the paper behind the figure that roughly 47 percent of total US employment fell into their high risk category, a result you have seen quoted for over a decade, usually without its method attached.
A more recent approach asks the question at task level directly. In GPTs are GPTs, the authors rated whether access to a large language model would meaningfully reduce the time needed to complete each task, using both human annotators and a model, and reported exposure rather than replacement.
A risk score is a statement about the tasks in a category. It was never designed to be a statement about a career.
What an AI job risk score actually measures
Read the methods and the honest framing becomes obvious. These studies measure exposure, which means the share of an occupation's tasks that a technology could plausibly touch. Exposure is not the same as displacement. Whether exposure becomes job loss depends on cost, regulation, liability, organisational inertia, whether the remaining tasks expand, and whether demand for the output grows when the price drops.
That gap gets lost the moment a number reaches a headline. A probability that a category of tasks is technically automatable becomes "half of all jobs will disappear", and the qualifier that made the research defensible falls off in transit. The forward-looking surveys have the same problem: the World Economic Forum's Future of Jobs work is a survey of employer expectations, which is useful evidence about what employers currently believe and is not a measurement of what will happen.
Get a score that is about you, not your job title
Six domains of judgement under uncertainty, measured from your own answers rather than inferred from an occupation code.
The two things a job level score cannot see
The first is variance inside the job title. One person spends the week producing documents from a clear brief. Their colleague, same title, spends it deciding which brief is even correct, managing a client who has changed their mind twice, and noticing that a number in the forecast stopped making sense. A task list averaged across an occupation cannot distinguish them, yet their positions are not remotely similar. Our checklist for spotting exposure in your own week is built around that distinction.
The second is the part of the job that never appears in a task list at all. Task inventories describe work that has already been specified. They are poor at capturing the labour of figuring out what the task should be, of holding a shared picture across several people, of choosing under time pressure with an asymmetric cost of being wrong. Those are the six domains that resist automation, under-recorded because they are hard to write down as a task.
So a job level score is genuinely useful for one purpose. It tells you the weather in your occupation. It will not tell you whether you personally are dressed for it.
What a person level measurement adds
Measuring the person requires a different instrument. The established approach combines two methods, both decades older than this debate.
A situational judgement test presents a realistic scenario and several plausible courses of action, and grades the choice against practice experts agree on. It works because the scenario has a defensible better answer, so the response can be scored rather than merely collected.
A behaviourally anchored rating scale replaces vague self-rating with described levels of observable behaviour, and asks which description matches what you actually did. The anchoring is what stops everyone from placing themselves comfortably above average.
Neither is flawless. But together they produce a result that varies between honest people and can be compared over time, which is more than any percentage derived from your job title can offer.
That is what the check does. It takes about seven minutes, it scores six domains separately, and it is explicit that the result is orientation rather than a prediction about your employment.
FAQ
Is the 47 percent figure wrong?
Should I ignore job level risk scores entirely?
Why measure judgement instead of technical skill?
Does the check predict whether I will lose my job?
Find out where you actually stand.
Six domains, twenty four items, one score. It takes about seven minutes and tells you which parts of your work AI is closest to, and which parts it is not.
Take the checkFree · about 7 minutes · no account