The White House’s AI Test Rules Exclude Open Models—and Stay Secret

At staff-level White House meetings on August 4, 2026, industry representatives reviewed a completed voluntary framework for evaluating advanced AI systems before release. According to Axios’s August 4 account, it covers closed-source frontier models with state-of-the-art capabilities and national-security risks, excludes open models and will not be made public.
The framework exists, but its precise eligibility rules remain unavailable outside the government and participating companies. Axios’s August 3 report says the White House had completed the voluntary process by its deadline, after receiving draft feedback from Anthropic, OpenAI and Google, but would not disclose its contents or say when companies would begin using it.
Which AI systems appear to be covered
The reported definition creates a narrow coverage test. A model appears to qualify only when it is closed-source, represents state-of-the-art capability and poses national-security risks under government thresholds. Failing any one of those conditions would apparently place a system outside the framework.
- Leading closed model: potentially covered if it crosses both the undisclosed capability and risk thresholds.
- Less capable closed model: apparently outside the intended scope, although no outsider can apply the secret thresholds conclusively.
- Open-weight model: excluded under the reported closed-source requirement, even if its weights are released under conditions that limit other components or uses.
- Fully open model: also outside the reported scope, but the unpublished framework cannot be checked for any distinction between open weights and broader access to code, data or documentation.
This test does not identify any specific model as covered, cleared or rejected. No public list names participating developers, qualifying model versions, evaluation scores or systems that have completed the process. It is therefore possible to describe the apparent categories, but not to classify a particular forthcoming model independently.
Why open models fall outside the process
The direct reason is the framework’s reported definition: a covered frontier model must be closed-source. It also states that its provisions should not be interpreted as restricting open models after release, placing those systems beyond the prerelease channel described to industry.
The operational design helps explain that boundary, although the White House has not publicly supplied this rationale. The process contemplates temporary government access to an unreleased model in a high-security environment, limits on employee access and detailed access logs. Companies were encouraged to provide systems close to public release rather than early-stage versions.
That controlled arrangement fits proprietary models whose developer can decide who receives access before launch. Once downloadable weights are released for redistribution and modification, the same prerelease controls no longer describe the model’s circulation. This is a structural inference from the disclosed mechanics, not a verified statement of why policymakers chose the closed-source boundary.
Important edge cases remain unresolved. The public record does not show whether a developer planning an open-weight release could submit the still-private version voluntarily, how the framework defines the moment of release, or whether licensing restrictions affect classification. It also does not establish whether “open model” means downloadable weights alone or a system whose weights, code, training information and licensing are all open.
What secrecy prevents outsiders from assessing

The public cannot independently determine what counts as state of the art or which capabilities rise to the level of a national-security risk. The underlying cyber-capability benchmark and the threshold used to identify covered models are classified, while the separate voluntary framework has not been published.
That leaves both substantive and procedural questions unanswered. Outsiders cannot inspect the tests, evidence requirements, scoring rules or treatment of comparable systems from different developers. They also cannot establish which administration officials participate in an evaluation, what findings a developer receives or whether an adverse result changes a planned release.
The access rules are similarly opaque. The framework is meant to address confidentiality, cybersecurity, insider risk, intellectual-property protection, permitted use and nondisclosure when the government obtains prerelease access for up to 30 days. Yet the practical standards for storage, personnel access, logging and enforcement cannot be compared because the governing text is unavailable.
Nor is there a public definition of a “trusted partner” that could receive early access. That uncertainty matters to developers outside the initial discussions, researchers, policymakers and US allies because participation is voluntary but the process may still shape which organizations gain privileged access to advanced systems.
How the framework compares with OpenAI’s published proposal
The federal process is narrower and less transparent than a recent industry proposal addressing the same policy field. In its June 3 frontier-safety blueprint, OpenAI advocates a durable national framework, a stronger role for the Center for AI Standards and Innovation and a government-wide resilience plan for national-security and public-safety risks.
The two approaches overlap at a high level: both assign the federal government a continuing role in the evaluation and governance of increasingly capable AI. The White House framework, as described publicly, instead focuses on a specific operational channel through which developers of certain unreleased closed models can provide temporary government access.
The comparison has clear limits. OpenAI is an interested developer proposing a policy structure, not the author of the federal rules, and its blueprint does not reveal the government’s classified benchmark. Publication nevertheless allows outsiders to examine its institutional recommendations; the White House framework’s definitions and procedures cannot receive equivalent scrutiny.
The framework’s boundary remains unverifiable
As of August 9, the clearest available account is that participation is voluntary, coverage is aimed at state-of-the-art closed systems presenting national-security risks, and open models are excluded. Those descriptions establish the framework’s apparent direction, but not a model-by-model roster or a reproducible eligibility test.
The next meaningful development would be publication of the framework, an unclassified explanation of its definitions, or evidence that named developers and models have entered the process. Until then, a proprietary frontier system may face federal prerelease review while an open-weight system with seemingly comparable capabilities does not—and outsiders cannot independently establish where or why the government draws that line.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.