How testing works

How AI Job Risk Scores Are Calculated, and Where They Break

Published 25 July 2026 5 min read All articles
In short
  • Job level risk scores come from breaking an occupation into tasks and rating each task for automatability
  • The famous headline numbers are probabilities about categories, not predictions about you
  • A score for your job title cannot see how you personally do the work
Contents

An AI job risk score is almost never a measurement of you. It is a measurement of your job title, produced by breaking that title into a list of tasks, rating how automatable each task looks, and adding the ratings back up. Understanding that one sentence tells you most of what these tools are good for, and exactly where they stop being useful.

The tools are everywhere now. Type in an occupation, receive a confident percentage, feel briefly unwell. Very few of them show their working. The underlying method is not a secret though, and it is worth knowing, because the difference between a score about a category and a score about a person is the whole game.

Where the numbers come from

Nearly every public risk score traces back to the same starting material: a structured database of what occupations actually involve. In the United States that is O*NET, which decomposes thousands of occupations into tasks, work activities, required knowledge, and skills. It is a genuinely impressive piece of public infrastructure, and it is the raw feedstock for this entire genre.

From there the method is consistent. Take the task list for an occupation. Decide, for each task, how susceptible it is to automation given some assumption about what the technology can do. Aggregate. Publish a number.

The interesting part is step two, because that is where the judgement lives. The influential 2013 Oxford Martin School paper by Frey and Osborne had experts hand-label a subset of occupations as automatable or not, then used those labels to train a classifier that extended the estimate across the rest of the labour market. That is the paper behind the figure that roughly 47 percent of total US employment fell into their high risk category, a result you have seen quoted for over a decade, usually without its method attached.

A more recent approach asks the question at task level directly. In GPTs are GPTs, the authors rated whether access to a large language model would meaningfully reduce the time needed to complete each task, using both human annotators and a model, and reported exposure rather than replacement.

A risk score is a statement about the tasks in a category. It was never designed to be a statement about a career.

What an AI job risk score actually measures

Read the methods and the honest framing becomes obvious. These studies measure exposure, which means the share of an occupation's tasks that a technology could plausibly touch. Exposure is not the same as displacement. Whether exposure becomes job loss depends on cost, regulation, liability, organisational inertia, whether the remaining tasks expand, and whether demand for the output grows when the price drops.

That gap gets lost the moment a number reaches a headline. A probability that a category of tasks is technically automatable becomes "half of all jobs will disappear", and the qualifier that made the research defensible falls off in transit. The forward-looking surveys have the same problem: the World Economic Forum's Future of Jobs work is a survey of employer expectations, which is useful evidence about what employers currently believe and is not a measurement of what will happen.

Get a score that is about you, not your job title

Six domains of judgement under uncertainty, measured from your own answers rather than inferred from an occupation code.

Take the check

The two things a job level score cannot see

The first is variance inside the job title. One person spends the week producing documents from a clear brief. Their colleague, same title, spends it deciding which brief is even correct, managing a client who has changed their mind twice, and noticing that a number in the forecast stopped making sense. A task list averaged across an occupation cannot distinguish them, yet their positions are not remotely similar. Our checklist for spotting exposure in your own week is built around that distinction.

The second is the part of the job that never appears in a task list at all. Task inventories describe work that has already been specified. They are poor at capturing the labour of figuring out what the task should be, of holding a shared picture across several people, of choosing under time pressure with an asymmetric cost of being wrong. Those are the six domains that resist automation, under-recorded because they are hard to write down as a task.

So a job level score is genuinely useful for one purpose. It tells you the weather in your occupation. It will not tell you whether you personally are dressed for it.

What a person level measurement adds

Measuring the person requires a different instrument. The established approach combines two methods, both decades older than this debate.

A situational judgement test presents a realistic scenario and several plausible courses of action, and grades the choice against practice experts agree on. It works because the scenario has a defensible better answer, so the response can be scored rather than merely collected.

A behaviourally anchored rating scale replaces vague self-rating with described levels of observable behaviour, and asks which description matches what you actually did. The anchoring is what stops everyone from placing themselves comfortably above average.

Neither is flawless. But together they produce a result that varies between honest people and can be compared over time, which is more than any percentage derived from your job title can offer.

That is what the check does. It takes about seven minutes, it scores six domains separately, and it is explicit that the result is orientation rather than a prediction about your employment.

FAQ

Is the 47 percent figure wrong?
It is widely misquoted rather than wrong. The original paper estimated the share of US employment in a high risk category over an unspecified period of one to two decades, based on technical automatability. It was never a forecast that half of all jobs would vanish by a date, and the authors did not claim it was.
Should I ignore job level risk scores entirely?
No. They are a reasonable signal about the direction of travel in an occupation, and the task level ones are better than the older job level ones. Treat them as weather, not as a diagnosis.
Why measure judgement instead of technical skill?
Because technical skill is the part that is moving fastest. Judgement under incomplete information is the slowest moving component of most roles, which makes it the more stable thing to measure and the more sensible thing to build.
Does the check predict whether I will lose my job?
No, and it does not claim to. It gives you a picture of where your judgement sits across six domains, alongside an exposure estimate. Anything promising a reliable individual employment prediction is overselling.
How much of you can AI replace?

Find out where you actually stand.

Six domains, twenty four items, one score. It takes about seven minutes and tells you which parts of your work AI is closest to, and which parts it is not.

Take the check

Free · about 7 minutes · no account