Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
Creator Economy

Foundation Models Outran Their Benchmarks: The 2025 Report Revisited

|Updated: |Author: QUASA Editorial Team|6 min read| 5896
Foundation Models Outran Their Benchmarks: The 2025 Report Revisited

Innovation Endeavors published State of Foundation Models 2025 on June 24, 2025, presenting it as a snapshot of technical, organizational and economic changes surrounding generative AI. The firm’s official report page identifies Davis Treybig as the author and foregrounds three claims: monthly AI use by one in eight workers, billions in annual revenue for AI-native applications and rapid gains in model cost, speed and capability.

The report remains useful as a record of the industry’s position in mid-2025, but it is not a current scoreboard. Evidence published in 2026 strengthens its central argument that capabilities are moving quickly while also exposing a crucial limitation: benchmark victories do not automatically produce dependable autonomous work. The most valuable update is therefore not another catalogue of model releases, but a clearer distinction between measured performance, operational reliability and business durability.

What the 2025 report was really arguing

The report’s many observations support one broad thesis: foundation models were becoming cheaper and more capable at the same time that reasoning techniques, tool use and intense competition were changing how products should be built. In that framing, the model itself was becoming a rapidly depreciating component rather than a permanent competitive advantage.

That distinction matters. If rival systems converge quickly on benchmark scores, a product cannot rely indefinitely on privileged access to one model. Its defensible value has to come from the surrounding workflow: proprietary context obtained legitimately, dependable integrations, evaluation against a specific job, user trust and the ability to switch models without rebuilding the entire product.

The report also treated reasoning as a new scaling route. Instead of expecting every improvement to come from a larger pretrained model, developers could spend more computation at inference time, apply reinforcement learning during post-training and let a model work through intermediate steps. The direction was important, but phrases such as “thinking” should be read as shorthand for a computational process—not evidence that a system understands a problem as a person does.

The acceleration thesis survived, but the scoreboard became less stable

Later evidence supports the report’s claim that technical progress continued. The 2026 Stanford AI Index technical review says frontier performance rose by 30 percentage points in one year on Humanity’s Last Exam and that OSWorld agent accuracy increased from roughly 12% to 66.3%. Yet the same review reports that agents still fail about one in three attempts on structured tasks, widely used evaluations contain invalid-question rates as high as 42%, and four leading companies were separated by fewer than 25 Arena Elo points in March 2026.

These findings sharpen the original report rather than simply confirming it. Models are improving, but benchmarks can saturate, contain flawed items or reward adaptation to a particular evaluation environment. Close leaderboard positions also mean that a small score difference may be less important to a buyer than latency, price, consistency, privacy controls or performance on the buyer’s own material.

Stanford’s results also complicate the idea that open models were on a one-way path toward parity. Its 2026 review found that the gap between the strongest closed and open models had widened to 3.3% by March 2026 after narrowing to 0.5% in August 2024. That does not establish permanent closed-model dominance; it shows why a temporary benchmark gap should not be presented as a settled market structure.

Autonomy still needs a precise definition

One of the report’s most memorable themes was that the length of tasks models could complete was doubling about every seven months. A subsequent methodological update preserved the long-run trend but demonstrated how sensitive the result is to task selection and evaluation infrastructure.

In January 2026, METR’s Time Horizon 1.1 update expanded its suite from 170 to 228 tasks and increased the number estimated to take a human at least eight hours from 14 to 31. The combined long-run estimate remained about 196 days, or seven months, while the estimate using post-2023 data shortened from 165 days under the previous setup to 131 days under the revised one. METR cautioned that the newer suite reflects a somewhat different distribution of difficulty and that confidence intervals remain wide.

The metric measures a specific quantity: the duration a human expert would need for tasks that an evaluated model-and-scaffold configuration can complete with 50% reliability. It is not a promise that an unmonitored agent can safely perform every job of that duration. Prompts, tools, scoring rules, task families and the surrounding scaffold are part of the tested system, so the resulting capability should not be attributed to the base model alone.

Why cheaper intelligence does not guarantee a durable business

The report was strongest when it treated falling model costs and rising frontier expenditure as simultaneous realities. These trends operate at different layers. A customer may pay less for a fixed level of inference performance even while frontier developers spend more on training, data, infrastructure and serving growing demand.

For application companies and independent creators, lower inference prices reduce the cost of experimentation. They do not remove expenses associated with human review, rights management, retrieval systems, security, storage or failed generations. A workflow that looks inexpensive per token can still be costly if it requires repeated retries or extensive correction.

Rapid model turnover creates another constraint. Building around one provider’s temporary benchmark lead exposes a product to pricing changes, deprecations and a competitor’s next release. A more durable architecture separates the user experience and business logic from the model interface, records quality and cost by task, and keeps a controlled path for replacing the underlying system.

What creators and product teams should take from it now

The refreshed evidence suggests three practical conclusions. First, use public benchmarks to identify candidates, not to make the final purchasing decision. A smaller evaluation set drawn from the actual workflow—including difficult and failure-prone examples—can reveal more than a broad average score.

Second, measure completed work rather than fluent output. For a creator, that may mean factual accuracy, revision time, consistency across a series and the proportion of material that can be published after review. For a software team, it may mean successful tool calls, recoverable errors and the frequency of human intervention. These are editorial recommendations, not universal benchmark definitions.

Third, preserve human approval wherever mistakes carry legal, financial, reputational or safety consequences. A model’s ability to finish longer evaluation tasks is meaningful progress, but 50% success is not an acceptable service level for many real operations. Reliability thresholds should follow the consequence of failure, not the excitement surrounding a leaderboard gain.

The updated verdict

State of Foundation Models 2025 correctly captured an industry shifting from simple chat interfaces toward reasoning systems, tools and longer workflows. Subsequent research supports the direction of travel: capability continued to rise rapidly, leading systems clustered more closely and agents completed substantially more structured computer tasks.

What changed is the confidence that should be placed in any single metric. Updated task suites move estimates, benchmarks can become saturated or reveal faulty questions, and impressive aggregate results coexist with conspicuous failures. The report is therefore best read as a strategic snapshot whose central tension has intensified: foundation models keep improving, but evaluating and productizing that improvement remains the harder, more durable work.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0