Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
Creator Economy

AI Agents Ran 26 Minutes per Session—but That Is Not a Productivity Guarantee

|Updated: |Author: QUASA Editorial Team|6 min read| 673
AI Agents Ran 26 Minutes per Session—but That Is Not a Productivity Guarantee

The strongest evidence for autonomous agents remains specific rather than universal. The June 2026 Perplexity–Harvard technical report found that Perplexity Computer performed an average of 26 minutes of machine execution per matched session, compared with 33 seconds for Search. That is compelling evidence of a different workflow, but it does not prove that every agent, worker or assignment becomes more productive.

What has become clearer since the study appeared is that autonomy depends on both the task and the product surface. Anthropic’s June 2026 Economic Index found higher autonomy across most output categories in Claude Code than in chat or Cowork, even after comparing sessions served by the same model. The practical shift is therefore not simply from an older model to a smarter one: software design increasingly determines how much work can be delegated and how much human judgment remains necessary.

What the 26-minute result actually measures

The Perplexity study compared Search, a conversational answer engine, with Computer, a system designed to plan and execute longer workflows. Researchers paired sessions from the same users when their opening requests were nearly identical, treating those pairs as approximations of the same underlying assignment attempted through two different products during February–May 2026.

Computer’s longer execution time is meaningful because it represents work that the user did not have to perform step by step: searching, browsing, running code and assembling deliverables could continue inside the agent workflow. Follow-up requests also moved toward checking and extending results rather than directing every intermediate action. In this narrow sense, the agent changed the user’s role from operator to supervisor.

The study’s productivity estimates require more caution. Its comparison of 269 minutes for a Search-plus-human workflow with 36 minutes for Computer plus human oversight was constructed from estimated human-equivalent time for tool actions, model costs and an assumed allowance for supervision. It was not a randomized stopwatch comparison in which workers completed every assignment under both conditions. The estimates are useful for modeling the economics of long tasks, but they should not be treated as a guaranteed 87% reduction for an arbitrary workplace.

Autonomy and productivity are different variables

An agent can work for longer without human intervention and still produce something that takes too long to inspect, revise or integrate. Autonomous runtime measures how much execution the system absorbs; productivity depends on whether the resulting artifact is correct, usable and worth the combined cost of prompting, waiting and review.

A contrasting result illustrates why the distinction matters. METR’s randomized developer trial assigned 246 real tasks to 16 experienced open-source contributors with or without early-2025 AI tools. In that particular setting, allowing AI increased completion time by 19%, even though participants believed it had made them faster. The experiment concerned familiar, mature codebases and primarily conversational coding tools—not Perplexity Computer—so it does not contradict the later agent study. It shows that results cannot be transferred across products, task environments and measurement methods without qualification.

The distinction is especially important in creative and knowledge work, where a plausible draft can conceal weak evidence, inconsistent decisions or unusable details. Producing more words, slides, code or campaign concepts is not itself the outcome. The relevant unit is a deliverable that survives editorial, technical or commercial review.

Which work should move from chat to an agent?

Conversational assistance remains a good fit when the user is still discovering the problem. Comparing interpretations, questioning assumptions and deciding what should be made are interactive activities; rapid dialogue is often more valuable than prolonged autonomous execution.

Agent delegation becomes more attractive when the destination is clear but reaching it requires many mechanical or research-heavy steps. A suitable assignment usually has four properties:

  • A defined artifact: the requested output is a research table, working prototype, edited document or other inspectable object.
  • Bounded inputs: the agent knows which files, websites, datasets or tools it may use.
  • Testable acceptance criteria: the worker can determine whether required fields, sources, functions or constraints are present.
  • Recoverable actions: mistakes can be reviewed before publication, payment, deletion or communication with another person.

A creator planning a sponsorship package, for example, could delegate the collection of public campaign references, normalization of rate-card inputs and assembly of a first presentation. The creator should retain decisions about positioning, factual claims, audience fit and the final commercial offer. That boundary gives the agent substantial execution work without quietly outsourcing accountability.

The task brief becomes part of the work product

Agents impose a larger upfront cost than a quick question because they need an operational definition of success. A useful brief identifies the intended audience, the final format, permitted evidence, excluded actions and the checks required before delivery. If those constraints remain implicit, the agent must guess—and the reviewer inherits the cost of locating every guess.

For long assignments, checkpoints are more useful than constant micromanagement. The worker can require an outline before full production, a source table before synthesis or a test report before code is merged. These gates preserve the advantage of asynchronous execution while exposing errors before they propagate through the entire workflow.

Verification should also be designed before the agent starts. Numerical work needs reproducible calculations; research needs links that directly support claims; code needs tests and human review; audience-facing material needs editorial and legal checks appropriate to its risk. The more expensive or irreversible the consequence, the less sensible it is to accept a polished output as evidence of correctness.

How to measure whether delegation is working

Teams should compare complete workflows, not model activity. The useful baseline includes the time a person previously spent gathering inputs, executing steps and correcting mistakes. The agent workflow must include specification, machine cost, waiting, review, revision and integration into the final system.

Three measurements are usually enough to expose false efficiency: elapsed time to an accepted deliverable, active human minutes and the share of outputs that pass review without substantial rework. An agent may reduce active labor while increasing elapsed time, which can still be valuable when jobs run in parallel. It may also generate quickly but demand so much correction that conversational assistance would have been cheaper.

The current evidence supports a disciplined conclusion. Agent products can absorb much more multi-step execution than chat interfaces, and product design appears to influence autonomy independently of the underlying model. The durable advantage belongs not to the workflow with the longest autonomous run, but to the one that turns delegation into a verified result with less total human effort.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0