Workers Reject Capable AI Tasks—Automation Readiness Needs Two Scores

Some workplace tasks that AI experts consider technically automatable still fall below workers’ threshold for desired automation. Managers therefore need two separate readiness scores: worker desire and demonstrated technical capability. Together they distinguish welcomed delegation from resisted automation, unmet demand and low-priority work.
WORKBank formalizes that distinction using responses from 1,500 U.S. workers across 104 occupations and assessments from 52 AI experts covering 844 tasks. Its worker-centered audit found positive attitudes toward automation for 46.1% of tasks and divided the desire-capability landscape into four zones. Capability shows what may be technically feasible; it does not establish that the people performing a task want it delegated.
Why capability alone gives the wrong answer
A conventional assessment asks whether software can complete a workflow accurately, quickly and economically. That can identify feasible projects, but it cannot distinguish a useful service from an unwanted substitution. Workers may resist delegating a capable task because it carries valued judgment, relationships, learning opportunities or responsibility.
The reverse mismatch matters too. Employees may want relief from a burdensome task that current agents cannot complete reliably. A single readiness score collapses that demand into “not ready,” concealing an opportunity for research, narrower assistance or a human-agent workflow.
The unit of analysis should be a specific recurring task, not a job title or broad responsibility. “Support customers” combines retrieval, communication, judgment and exception handling. “Retrieve the customer’s order history before a representative responds” is narrow enough for workers to assess and for a technical team to evaluate.
Build the two-axis worksheet

Create one row for each task and record its inputs, expected output, frequency, current owner and the person accountable when the result is wrong. Collect the two scores independently so a persuasive product demonstration does not shape the worker-preference rating.
- Worker-desire score: Ask people who currently perform the task how much of it they want delegated. A consistent five-point scale makes responses comparable, but retain the distribution and reasons rather than reducing disagreement to an unexplained average.
- Capability score: Evaluate the proposed agent on representative cases and exceptions under the organization’s actual tools, permissions, data and quality requirements. The question is whether that configured system can perform the defined task, not whether a general model can produce a plausible sample.
- Evidence confidence: Record the number and range of worker responses, evaluation coverage and assessment date. Two interviews or a polished demonstration should not carry the same weight as broad feedback and repeatable tests.
- Human role: Specify who sets goals, supplies context, verifies output, handles exceptions and authorizes consequential actions. This field converts a quadrant label into an operating design.
Choose thresholds before reviewing candidate projects. Scores near a boundary belong in a review band rather than being forced into a confident category. Reassess both axes when the task, workforce or deployed system changes.
Use four zones, not one ranked list

Green light combines high desire with high capability. Red light combines low desire with high capability. R&D opportunity combines high desire with low capability, while low priority means both scores are low. These labels prioritize investigation; they are not safety certifications or deployment approvals.
The paper’s published desire-capability landscape illustrates all four outcomes. Scheduling client appointments for tax preparers appears in the green-light zone. Researching hardware or software products for computer network support specialists appears in the red-light zone; creating production schedules and prototyping goals for video game designers is an R&D opportunity; and tracing lost or delayed baggage for ticket agents appears as low priority.
A green-light classification supports controlled evaluation, not immediate autonomy. For a red-light task, investigate what workers want to retain before redesigning the task or considering a smaller delegated component. An R&D opportunity calls for a constrained assistance role or further technical work, while a low-priority task should normally stay off the near-term roadmap unless either score changes.
Augmentation requires a separate control decision

The two scores identify the type of mismatch, but they do not determine how humans and agents should divide the work. A red-light task may contain a welcomed administrative subtask. An R&D opportunity may become useful when an agent gathers evidence or drafts options while a worker reviews the result and makes the decision.
The Stanford WORKBank project addresses this question with a five-level Human Agency Scale, from no human involvement to essential human involvement. Equal partnership was the dominant worker preference in 47 of the 104 occupations studied, and workers preferred more human agency than experts considered technically necessary for 47.5% of the 844 tasks.
For a local worksheet, divide control into goal setting, execution, verification, exception handling and final authority. An agent might execute a routine step while a worker chooses the objective, reviews uncertain cases and authorizes an external action. That is a defined augmentation model, not a vague midpoint between manual work and autonomy.
Convert the scores into a deployment gate
Review high-desire, high-capability tasks first, but require local evidence before granting autonomy. The gate should include representative evaluation cases, named accountability, permission limits, monitoring, an override path and a scheduled worker-feedback review.
High-capability, low-desire tasks should pause for a job-design and acceptance review. High-desire, low-capability tasks belong in bounded experiments with human verification and explicit failure limits. Low scores on both axes justify deferral rather than implementation effort.
The practical rule is straightforward: capability defines the available technical options; worker desire shows whether full delegation is wanted. Keeping both scores and the intended human role visible prevents a technically impressive agent from being mistaken for a workplace-ready intervention.
Also read:
- Opportunities for Malaysian Remote Workers and Freelancers Expanded for Cryptocurrency Earnings via Quasa Connect
- Spain: A Land of Boundless Crypto Earning Opportunities for Remote Workers via Quasa Connect
- Neuralink Unveils the Future of Gaming: Two Players Dominate Call of Duty Using Brain-Controlled Interface
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.