
AWS Continuum Claims 89% on CyberGym—but the Result Is Vendor-Run

On October 5, 2026, AWS reported a company-run CyberGym-E2E evaluation in which Continuum for code vulnerabilities passed S3 on 819 of 920 tasks within 90 minutes: an 89% score, reported more precisely as 89.0%, alongside 37.8% at S4. S3 is the benchmark’s main end-to-end success measure; S4 checks whether a patch also fixes the particular historical vulnerability selected for that task. Both figures describe the evaluated multi-agent system and its test setup, rather than a model in isolation.
The October 5 result was presented in same-day news coverage as a company-reported claim, with no separate rerun described. That distinction matters alongside the gap between the two stages. Passing the benchmark’s functionality tests establishes that an agent found a crash, prevented it with a patch and kept tested behavior intact; it does not by itself establish that the patch removed the designated historical flaw.
What the 89% S3 score measures
CyberGym-E2E places an agent inside a vulnerable open-source project’s build environment, with source code and the scripts needed to build and test it. In the end-to-end setting, the description of the selected vulnerability, the original proof of concept, the crash log and the original patch are withheld. The agent must inspect the repository, find a flaw, produce an input that triggers a sanitizer crash and submit a source-code patch. This makes the task broader than patching a disclosed bug.
The validation sequence is cumulative. At S1, the agent’s input must crash the unpatched build. At S2, the proposed patch must stop that same crash. At S3, the patched project must still pass its developer-written functionality tests. In the reported run, the first two stages reached 92.5% and 89.6%, respectively; the difference between S2 and S3 was 0.6 percentage points. A task reaching S3 has passed the preceding checks too, but those checks cover specified functionality rather than every behavior a project might have in production.
Continuum organizes this work into discovery, validation and remediation phases, each handled by specialized agents. Candidate findings are tested through proof-of-concept inputs before a repair is produced and checked against the demonstrated failure. The result therefore belongs to the full workflow, including its tools, instructions and underlying models. It cannot be transferred to a single model outside that configuration or treated as a general patch-success rate for a known backlog of findings.
Why S4 is far below S3
S4 adds a different target to the same cumulative sequence: the submitted patch must fix the benchmark’s selected historical vulnerability. The reported S3-to-S4 spread is 51.2 percentage points. Because the agent is not told which historical flaw was selected, a patch can pass S3 after fixing a different genuine vulnerability in the same project. The lower S4 result therefore does not show that every patch outside S4 was useless; it shows that many S3 passes did not clear the selected-flaw check.
This distinction changes the claim for a security team with a named finding to close. A reproducible crash and passing regression tests are valuable evidence, but they do not identify which vulnerability was eliminated if a repository contains more than one. S4 asks the narrower question that matters when the team must show a particular flaw is gone. Conversely, open-ended discovery can produce useful repairs that S4 does not count because they address another defect.
There is a second patch-quality question. In the benchmark’s wider agent runs, a small fraction of patches used a shallow guard at the reported crash site while leaving the underlying defect. That pattern has not been quantified for Continuum’s reported run. It means that even a behaviorally successful patch remains a candidate for code review, especially where the defect could surface through a different input or path.
How the claim compares with the public leaderboard
The public CyberGym-E2E leaderboard lists 65.9% at S3 for GPT-5.4 with Codex under a $10 cost cap and a 90-minute limit; its highest listed S4 score is 26.2% for Claude Opus 4.6 with Claude Code and no cost cap. It contains no Continuum row. These two prior highs come from different agent configurations, so combining them into a single rival’s performance would distort the comparison.
The 89.0% vendor figure exceeds the listed S3 high by 23.1 percentage points, but the visible comparison does not include a public Continuum submission or a matching per-task cost limit. The company-run evaluation was described as using network isolation, blocked external retrieval and a review of trajectories after completion. Those are relevant safeguards against recovering public historical fixes during execution; the procedure and reported outcomes still come from the company conducting the evaluation.
Runtime changes the result as well. When tasks were allowed to continue beyond the benchmark’s 90-minute limit, the reported S3 rate rose to 93.7%. That is a separate condition, rather than an improvement to the time-limited score. Spending limits also matter: the public S3 leader was evaluated with a $10 cap, while the public S4 leader had none. A percentage comparison without those conditions obscures what resources produced each outcome.
What the result means for security buyers
The test set covers 920 historical memory-safety vulnerabilities in C and C++ projects drawn from 139 open-source projects. It tests code-level discovery and repair using sanitizer crashes and project functionality tests. It does not supply the exposure, permissions, configuration or network paths needed to judge the urgency of a particular deployed service. A team working mainly in another language or vulnerability class has a further gap between this benchmark and its own workload.
For procurement, the questions are concrete: Which tasks passed S4? Can reviewers inspect per-task outcomes and patch artifacts? What model configuration, runtime and spending limit produced the results? A public Continuum submission under matched conditions would make the leaderboard comparison more direct and give teams better evidence about repairs to named findings.
Also read:
Related articles


Google Cloud Modernize Unifies Migration—but Its EKS Agent Is Still Preview

Cohere North 2 Gives Agents Memory—and Admins Token Caps

Dynatrace Closes $915M Arize Deal—Open-Source Projects Stay Supported

Instruction for Creators and Brands Joining Quasa Rewards
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.