
Palo Alto’s AI Defense Uses Multiple Models—None Found More Than 40%

In its September 22, 2026, launch release, Palo Alto Networks announced Unit 42 Continuous Frontier AI Defense as available worldwide by annual subscription. The service routes offensive security work across multiple AI models to find and validate enterprise exposures, then recommend fixes. In the company’s own evaluation, no single model found more than 40% of vulnerabilities in a complex environment, according to Axios’s launch coverage.
The newly available service combines gated models, including Anthropic Claude Mythos 5 and OpenAI GPT-5.6-Cyber, with open-weight models and Unit 42 offensive security specialists. It begins with a baseline assessment and is designed to keep testing as a customer’s environment changes. The model result helps explain that design; it is a finding about individual models in vendor-run tests, not a measured detection rate for the combined subscription.
Why Unit 42 routes work across models
The central design choice is to send a testing task to a model suited to it instead of relying on one model for every asset and weakness. In its service account, Unit 42 puts the overlap between exposures identified by Claude Mythos 5 and GPT-5.6-Cyber below 10% in its evaluation and describes a zero-data-retention architecture for customer source code and telemetry. The overlap concerns findings from those named models; it does not measure the coverage of every model offered through the service.
Different findings can give specialists more candidates to investigate, while routing a task to a suitable model can help control the cost of using frontier systems. Open-weight models broaden the available mix, though the public description does not identify the models or assign specific testing jobs to them. Nor does it set out when the same task is run through several models rather than one.
The percentages leave a crucial denominator unresolved. To calculate how many vulnerabilities the full service detects, an evaluation would need a defined set of known weaknesses and an account of both finds and misses under consistent conditions. Complementary model outputs make a case for testing a combination; they cannot establish that the combination finds every weakness or prevents an attack.
What continuous testing is designed to validate
The service’s baseline scan is followed by testing as an organization’s environment changes. Its stated scope covers first- and third-party web applications, APIs, cloud infrastructure, source code repositories, and network assets. Those categories describe potential reach, not automatic access to every asset in a customer’s estate. The assets included, the permissions granted, and any exclusions determine what can actually be examined.
Finding a possible flaw is only the discovery stage. The service is designed to test whether a suspected exposure can be exploited and whether separate weaknesses form an end-to-end attack path. A validated path gives a security team a more specific basis for prioritizing a fix than an unverified model output. It still represents what testing established within the accessible environment and the rules of that engagement.
“Continuous” describes an operating approach rather than a published time guarantee for every change. The launch materials do not specify how quickly a changed asset is retested, which events trigger another pass, or whether all asset types receive the same cadence. Those details matter because a subscription can be ongoing while particular systems remain outside scope or await their next test.
Recommendations, patches and human authority
The remediation component delivers prioritized fixes, code-level guidance and virtual patch recommendations. A separate Frontier Virtual Patching offering can be paired with the service to implement virtual patches. That distinction matters: identifying an attack path and recommending a control are different from changing production code or deploying a patch. The service description does not establish independent authority to make either change in a customer environment.
Human offensive security specialists are part of the described workflow, alongside model-driven discovery and validation. Public descriptions do not define an approval point for each intrusive test or proposed fix. A customer’s operating agreement therefore has to determine who authorizes testing, who accepts a finding as validated, and who approves or can reverse a remediation change. A model roster cannot answer those governance questions.
The zero-data-retention description addresses customer source code and telemetry, including the assertion that customer data is not retained or used to train public models. It does not provide a route-by-route account of which model receives particular material, what records are kept to substantiate a finding, or the contractual terms governing each provider. Those details affect how an enterprise can assess data exposure while obtaining enough evidence to act on a result.
What the launch evidence can establish
The public model figures come from the vendor’s evaluation across enterprise codebases and live environments. The available description does not provide a complete test corpus, a breakdown by model and asset type, or a count of missed vulnerabilities for the combined service. Earlier internal deployment and customer assessments also took place in settings distinct from a new customer’s continuing subscription. Their outcomes cannot be converted into a guaranteed result for another environment.
For a comparable assessment, a buyer would need the tested asset scope, a sample of validated findings with supporting evidence, the method used to determine exploitable attack paths, and a way to account for false positives and misses. The same record would need to show which work was automated, where specialists intervened, and which proposed changes required customer approval. These are boundaries of the advertised capabilities, not extra performance figures supplied by the launch.
Unit 42 Continuous Frontier AI Defense is available as an ongoing offensive security subscription with a multi-model testing design and remediation recommendations. The vendor’s finding about individual models explains why combining them may be useful. How much of a particular customer’s environment the service covers, what its combined detection rate is there, and who may act on its findings remain matters for evidence and agreed operating scope.
Also read:
Related articles


AI’s Cyber Window Is Closing—but 100 Signatories Promise No Deadlines

100+ Tech Firms Warn of AI Cyberattacks—but Make No Binding Commitments

Artifactory Auth Bypass Is Exploited—Only Self-Hosted Users Must Patch

Cisco ISE Zero-Day Is Exploited—Patching Cannot Prove a Clean Network

Factory Triples to a $5B Valuation—Proof Still Depends on Enterprise ROI
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.