← Back to job risk analyzer

For researchers: the willitreplace.me dataset

WillItReplace.me publishes a task-level AI automation-exposure dataset covering 535 occupations in 30 industries, with 2,172 individual task scores. Each occupation is broken into core tasks; every task receives a 0-100 score for how likely current or near-future AI (LLMs, agents, computer vision, robotics) can perform it. The role-level score is the aggregation of its task scores.

The dataset is intended as a transparent, citable indicator of task-level AI exposure - useful for exploration, teaching, journalism, and comparison against peer-reviewed studies - not as a forecast of job losses.

Download the dataset

Free to use with attribution. JSON preserves full detail (tasks, timelines, strategies); CSV is a flat task-level table with one row per task.

  • jobs-535.json - full dataset, 535 occupations (JSON, ~1.1 MB)
  • jobs-535.csv - flat table: job_name, category, overall_risk_pct, task_name, task_ai_probability_pct (~2,172 rows)

Methodology

  1. Each occupation is decomposed into its core tasks (typically 3-6 per role).
  2. Each task is scored 0-100 for AI automability, weighing: current AI capability, near-future trajectory (5-10 years), task complexity (physical dexterity, emotional intelligence, creative judgment), regulatory barriers, and real-world adoption.
  3. The role-level risk score aggregates task scores; full details on the methodology page.

The task-based approach follows Frey & Osborne (2013) and McKinsey Global Institute's work-activity framework, re-scored for generative AI. Scores are produced by structured model-assisted assessment against these rubrics; they are not survey measurements or official statistics.

Frey, C. B. & Osborne, M. A. (2013). "The Future of Employment: How Susceptible Are Jobs to Computerisation?"

Oxford Martin School, University of Oxford

The foundational task-based framework: probability of computerisation for 702 US occupations. WillItReplace.me adapts its task-level approach to current AI capabilities rather than reusing its 702-occupation probabilities directly.

PDF (oxfordmartin.ox.ac.uk)

McKinsey Global Institute (2017). "Jobs Lost, Jobs Gained: Workforce Transitions in a Time of Automation."

McKinsey & Company

Work-activity-level (not job-level) analysis of automation potential across 46 countries; found roughly half of work activities could be automated with existing technology.

Report page (mckinsey.com)

McKinsey Global Institute (2023). "The Economic Potential of Generative AI: The Next Productivity Frontier."

McKinsey & Company

Updates automation estimates for generative AI; estimates current technology could automate work hours absorbing 60-70% of employee time, with large uncertainty ranges.

Report page (mckinsey.com)

Goldman Sachs (2023). "The Potentially Large Effects of Artificial Intelligence on Economic Growth."

Goldman Sachs Research

Widely cited estimate that roughly 300 million full-time jobs globally are exposed to automation by generative AI, and ~25% of US work hours could be automated. Exposure here means task overlap, not job elimination.

Article (goldmansachs.com)

World Economic Forum (2025). "Future of Jobs Report 2025."

World Economic Forum

Employer-survey-based: 170M roles created vs 92M displaced by 2030 (net +78M); 41% of employers plan workforce reductions where AI automates work.

Report page (weforum.org)

US Bureau of Labor Statistics — Occupational Outlook Handbook

US Department of Labor

Used for occupation definitions and salary benchmarks, not for risk scoring.

bls.gov/ooh

Limitations

  • Task-level, not job-level. The scores measure exposure of tasks. A high score does not mean the job disappears; history shows automation typically restructures roles (see WEF 2025 for net job-change estimates).
  • An indicator, not a prediction. Scores express capability-based exposure under stated assumptions. They are not probabilities of job loss and carry no confidence intervals.
  • Model-assisted scoring. Task scores are produced by structured AI-assisted assessment against the rubric, then human-reviewed. They have not been validated against realized labor-market outcomes.
  • US-centric framing. Occupation definitions and salary context are US-based (BLS); regulatory and adoption barriers vary by country.
  • Snapshots age quickly. AI capability moves faster than annual surveys; treat scores as of their review date.

Citing this dataset

If you use the data, please cite the site (no formal authors are published; credit the site as publisher and include your access date). The underlying studies above have their own citations and should be cited separately.

BibTeX

@dataset{willitreplace2026jobs535,
  title  = {willitreplace.me AI Job Automation Dataset: 535 Occupations, Task-Level Exposure Scores},
  author = {{WillItReplace.me}},
  year   = {2026},
  url    = {https://willitreplace.me/research},
  note   = {Accessed: <your access date>}
}

Plain text

WillItReplace.me (2026). willitreplace.me AI Job Automation Dataset:
535 occupations with task-level AI exposure scores.
https://willitreplace.me/research (accessed <your access date>).

Questions or corrections: see the About and Terms pages.