How it works
The method, every source with its date and license, and what the numbers do not say. Sources checked September 29, 2026. Every figure that is an estimate is labeled as one wherever it appears.
What the numbers mean
The main number for a job is the share of its working time that is : time spent on desk tasks short enough for the best AI models to finish on their own, at a chosen reliability, at a chosen date. It is a statement about what models could do if the work were set up for them, with the instructions, files and tools in front of them. It is not a forecast of jobs lost, of savings, or of how fast any company will change. Those depend on decisions no data set can see.
Next to it, every job shows : Anthropic's measure of how much of the job people already hand to Claude at work. The two answer different questions. Within reach counts whole instances of a task a model could finish alone. Observed use counts real use, including help on part of a task, which is why it can run higher.
Sources
- Tasks. ® 31.0 Database, U.S. Department of Labor, Employment and Training Administration: 923 occupations with task data and 18,838 task statements, with how often each is done, how important it is and the share of workers who do it. CC BY 4.0.
- Task lengths. A Stratus estimate. Three models from three companies (Gemini 3.8 Flash, GPT-5.6 Luna and Claude Haiku 4.5) each estimated a simple, a typical and a demanding instance of all 17,578 distinct task statements under one written definition. The typical length is the median of the three. Checked against GDPval, OpenAI's set of 220 real professional tasks whose length the professionals recorded: the estimates ran long by a factor of 1.2 at the median, so every length is divided by 1.2.
- Desk or body. A Stratus estimate: whether each task could be done entirely through a computer, judged task by task by a model. An unjudged task counts as body work, never as desk work.
- Model ability. METR, Task-Completion Time Horizons of Frontier AI Models (Time Horizon 1.1), read September 19, 2026 and checked again September 29, 2026. metr.org
- Observed use. Anthropic Economic Index, labor market impacts files (published March 5, 2026, from Claude usage in August and November 2025). CC BY 4.0.
- Pay and headcount. U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, national estimates, May 2025. Cost of an hour from BLS Employer Costs for Employee Compensation, June 2026.
- Industries. The same BLS survey's national industry-specific estimates, May 2025: which occupations each sector, subsector and industry group employs, how many people and at what pay.
- States and metro areas. The same survey's state, metropolitan and nonmetropolitan area estimates, May 2025, across all industries.
How a job's time is split across its tasks
O*NET asks workers how often they do each task, in seven answers from "yearly or less" to "hourly or more", and what share of them do it at all. Each answer is turned into times a year (1, 5, 25, 100, 250, 750 and 2,000), averaged over the answers, multiplied by the share of workers who do the task and by its typical length, and capped at one working year. Each task's is its part of the total. This is an estimate: O*NET does not record hours, and the answers are ranges.
When a task comes within reach
METR measures the of the best models: the length of task, timed by how long a skilled person takes, that a model finishes half the time () or four times in five (). Its latest measurement is Claude Mythos Preview (early), April 2026: 2.2 work days at even odds and 3.1 h four in five. Nothing released since has been measured as of September 29, 2026.
carries METR's own recent trend (frontier models since 2023, doubling every 129 days) forward to the date checked: 3.1 work days at even odds and 4.4 h four in five. The four-in-five length is the even-odds length divided by 5.62, the ratio in METR's latest measurement.
From today the length grows at one of three : the (doubling every 188 days, METR's whole record since 2019, the default), the (every 129 days) and a (once a year).
Instances of a task vary. A task's typical length sits between a simple and a demanding instance, read as the 10th and 90th percentiles of a log-normal spread, taken from the median model. A task comes within reach a piece at a time: its share within reach is the part of its time spent on instances no longer than the horizon. Long instances carry more of the time, so this counts time, not instances.
| Model | Released | Even odds | Four in five |
|---|---|---|---|
| Claude Mythos Preview (early) | April 2026 | 2.2 work days | 3.1 h |
| Claude Opus 4.6 | February 2026 | 1.5 work days | 1.2 h |
| GPT-5.2 | December 2025 | 5.9 h | 1.1 h |
| Claude Opus 4.5 | November 2025 | 4.9 h | 49 min |
| Gemini 3 Pro | November 2025 | 3.7 h | 54 min |
| GPT-5 | August 2025 | 3.4 h | 38 min |
| o3 | April 2025 | 2 h | 30 min |
| Claude 3.7 Sonnet | February 2025 | 1 h | 12 min |
| o1 | December 2024 | 39 min | 7 min |
| Claude 3.5 Sonnet (Oct 2024) | October 2024 | 21 min | 3 min |
| o1-preview | September 2024 | 20 min | 4 min |
| Claude 3.5 Sonnet (June 2024) | June 2024 | 11 min | 2 min |
| GPT-4o | May 2024 | 7 min | 1 min |
| GPT-4 Turbo (Nov 2023) | November 2023 | 4 min | 1 min |
| GPT-4 | March 2023 | 4 min | 1 min |
| GPT-3.5 Turbo Instruct | March 2022 | 1 min | 0 min |
| GPT-3 (davinci-002) | May 2020 | 0 min | 0 min |
| GPT-2 | February 2019 | 0 min | 0 min |
Three kinds of work
needs hands, a place or a presence in the room. is desk work whose main action is dealing with people: a task counts as people work when one of its leading verbs is on this list:
advise, arbitrate, coach, collaborate, comfort, confer, console, consult, counsel, delegate, demonstrate, direct, discuss, encourage, entertain, escort, facilitate, greet, hire, host, instruct, interview, lead, liaise, lobby, mediate, meet, mentor, moderate, motivate, negotiate, persuade, preach, present, recruit, represent, sell, solicit, supervise, teach, testify, train, tutor, visit, welcome.
Neither counts as within reach here, whatever its length. The rule is deliberately simple so anyone can check it; it will count some tasks as people work that a model could help with, and miss some people work that opens with another verb. Everything else is .
Hours and cost
A company scan multiplies each role's share within reach by its (2,080 per person a year) and headcount. Value uses the company's pay where given, otherwise the BLS median for the matched occupation, times the : BLS reports that wages were 70.0 percent of what private employers spent per hour worked in June 2026, so the default multiple is 1.43. The is the cost of that work, not a saving.
Titles are matched to occupations in the browser, against O*NET's lists of titles people actually use. Equal matches go to the larger occupation; a match with a close rival in a different occupation is marked Check.
Industries
BLS counts the people in each occupation in each industry, for 20 sectors, 85 subsectors and 247 industry groups (). An industry's share within reach is its occupations' shares weighted by those counts; its hours and value follow the company scan, using the pay the industry pays each occupation. The same counts, scaled to a headcount, make the typical company each industry page can scan.
BLS publishes a few occupations only combined, such as buyers and purchasing agents (13-1020) for O*NET's three kinds of buyer. One rule places every such case: within a group of related codes, the O*NET occupations BLS does not publish belong to the one combined code BLS publishes there, and where that leaves a choice, nothing is guessed. Twelve combined codes are placed this way. The people BLS counts only in "all other" groups, which have no task lists, are left out, and each industry page says how much of its employment the task data covers: 98% for the median industry.
splits today's share within reach by whether Anthropic saw people bring each task to Claude at work in 2025. It counts one AI product, so it shows where use has begun, not where it has not.
States and metro areas
The same survey counts the people in each occupation in every state and territory, 393 metropolitan areas and 137 nonmetropolitan areas, across all industries. An area's numbers are built exactly as an industry's, from those counts and the pay each area pays. BLS leaves out small counts it cannot publish, so task data covers a little less of an area's people than of an industry's; each area page says how much.
How the numbers move
Every job and industry page shows how far its share within reach has come since GPT-4: at each frontier measurement METR has published since March 2023, the share within reach at four in five using the best model measured by that date and today's task data, then today's estimate. It is a backcast. It shows how fast the models have moved, not what anyone forecast at the time.
From this release on, every data build also keeps a dated record of every job's and industry's numbers, including the projections for the end of 2027, 2028 and 2030. As new measurements arrive, those records will show how each estimate changed, and in time how the projections compared with what the models actually did.
Limits
- METR's tasks are mostly software, research and reasoning work. In July 2025 METR found time horizons 40 to 100 times shorter for work done by operating a computer screen, such as clicking through web pages. Desk tasks that depend on operating software a model cannot reach may come within reach later than shown.
- Within reach assumes the work is set up for the model: context, files, systems access and a way to check the result. Most work is not set up that way yet.
- Task lengths, time shares, the desk-or-body judgment, the people rule and today's frontier are estimates. Ranges are shown where they matter.
- An occupation is an average. Any one company's version of a job can differ a lot from it.
- An industry is an average too: a typical company is the industry's staffing, not any real company's.
- Nothing here is advice about any person's job or any company's staffing decisions.
Terms
- Within reach
- Work a model could finish on its own at the chosen reliability, if the work is set up for it: the instructions, files and tools it needs are in front of it. Within reach is not the same as done. Whether a company hands the work over, and how fast, is a separate question.
- Four in five
- The length of task the best models finish four times out of five. It is the default here, because work a company hands over has to come back right most of the time. Today's estimate: 4.4 h.
- Even odds
- The length of task the best models finish half the time. It runs about 5.6 times the four-in-five length. Today's estimate: 3.1 work days.
- Pace
- How fast the length of task models can finish keeps growing from today's estimate. Three are offered: the long-run pace (doubling every 188 days, the default), the recent pace (every 129 days) and a slowdown (once a year).
- Reliability
- How often a model has to finish a task for the task to count as within reach. Four in five is the default; even odds is the rate METR headlines.
- Time horizon
- METR's measure of AI ability: the length of task, timed by how long it takes a skilled person, that a model finishes at a given rate. It has doubled every four to seven months since 2019.
- Today's estimate
- METR's latest measurement is Claude Mythos Preview (early), April 2026: 2.2 work days at even odds and 3.1 h four times in five. Nothing newer has been measured, so today's figure carries METR's recent trend forward to the date checked: 3.1 work days at even odds, 4.4 h four in five. It is an estimate.
- Long-run pace
- The length models can finish doubles every 188 days, the pace of METR's whole record since 2019. The default here.
- Recent pace
- Doubles every 129 days, the pace METR measures for frontier models since 2023.
- Slowdown
- Doubles once a year, about a third of the recent pace. It shows the line if the gains from reasoning and tools run out.
- Range
- The low end assumes the slowdown pace and the high end the recent pace. The middle figure assumes the long-run pace.
- Desk work
- Tasks a person could do entirely through a computer. The only kind of work this site counts as within reach.
- People work
- Tasks whose main action is dealing with people: conferring, negotiating, supervising, teaching, selling and the like. The full list of verbs is on the method page. Counted here as staying with people.
- Body work
- Tasks that need hands, a place or a presence in the room. Length says nothing about these, so they never count as within reach.
- Observed use
- Anthropic's measure of how much of a job's work people already hand to Claude at work, from Claude usage in August and November 2025, published March 2026. It counts help with part of a task, which is why it can be higher than the share within reach.
- Task length
- How long a person takes to do one instance of the task. Estimated by three AI models from three companies and checked against GDPval, where professionals recorded how long real work took. Shown as a typical instance, with a simple and a demanding one either side.
- Paid hours
- 2,080 a year for a full-time person, the basis BLS uses to turn hourly pay into annual pay.
- Hours within reach
- Paid hours times the share of each role's time within reach, times headcount, added up over the team.
- Cost of an hour
- Pay plus benefits and payroll taxes. BLS reports that wages were 70.0 percent of what private employers spent per hour worked in June 2026, so pay is multiplied by 1.43. You can change it.
- Value within reach
- What the hours within reach cost the company at the cost of an hour. It is the cost of the work, not a saving: what a company does with the time is its own decision.
- Occupation
- One of the 923 jobs O*NET data describes task by task, from the US Department of Labor's standard list.
- O*NET data
- The US Department of Labor's database of occupations, the tasks that make them up, and how often and how widely each task is done. This site uses version 31.0.
- Match
- The occupation a job title was matched to, using O*NET's lists of about 60,000 titles people actually use. High means an exact or near-exact title. Check anything marked Check or Pick one.
- Core task
- A task O*NET marks as central to the job rather than supplemental.
- Median pay
- Half of the people in the occupation earn more and half less. From the BLS Occupational Employment and Wage Statistics survey, May 2025.
- NAICS
- The North American Industry Classification System: the codes the US government uses to sort businesses into industries. Two digits name a sector (42 is wholesale trade), three a subsector, four an industry group. BLS publishes a few groups combined, with codes like 4230A1.
- Seen in Claude use
- The part of today's share within reach that sits in tasks Anthropic saw people bring to Claude at work in 2025 (Anthropic Economic Index). The rest is within reach in tasks not yet seen there. Other AI tools are not counted, so this understates use; it shows where use has started, not where it has not.
- MCP
- Model Context Protocol: the open standard ChatGPT, Claude and other assistants use to call tools outside themselves. An MCP address is where the assistant sends those calls.
- Connector
- Claude's name for an outside service it can call through MCP. ChatGPT calls the same thing an app.
Changes
- September 29, 2026. First release. METR checked: no measurement newer than April 2026.
- September 29, 2026, later. Every state and metro area. How far each job, industry and area has come since GPT-4, and a dated record of every estimate from here on. Industries added: every sector, subsector and industry group, with a typical company to scan. Twelve occupations BLS publishes only combined now take the combined code's pay instead of a broader group's; their shares within reach are unchanged.
Questions about the method: info@stratussc.com. Back to the start