
DeepMind Locks Benchmarks in a Cryptographic Box to Stop Test Leakage
On August 27, Google DeepMind and its partners disclosed a proof-of-concept evaluation that kept private test prompts and Gemini 2.5 Flash-Lite’s weights apart.

On August 27, Google DeepMind and its partners disclosed a proof-of-concept evaluation that kept private test prompts and Gemini 2.5 Flash-Lite’s weights apart.

Double-blind execution can hide reserved prompts from a model developer and proprietary weights from evaluators, reducing leakage without proving that the benchmark measures what buyers need.

Publication rejected: Google introduced Gemini 3.5 Transcribe on August 26, 2026, with 85+ supported languages, but developer access remains in public preview—not GA.

OpenAI classified Astra at its Critical cybersecurity threshold on September 1. Its strongest cyber capabilities will be restricted, and safeguards may interrupt legitimate defense work.

On September 1, 2026, Microsoft published a revised responsible-AI framework extending oversight of agents across permissions, memory, runtime behavior and post-release monitoring.

Govern an AI agent by defining its purpose, authority and human controls before measuring performance. The resulting tests then reflect the system that will actually be deployed.

As of August 28, every PaperCut NG and MF version is potentially affected by active exploitation. Emergency patches are available only for versions 25 and 26.

OpenAI’s August 24 benchmark claim puts Terra’s successful-task cost in Kiro roughly 82% lower, but the vendor-run test omits its baseline and full methodology.

Google’s August 26 Gemini Live rollout adds voice-controlled Gmail actions, spoken Daily Briefs and Spark delegation, but plans, market availability and app connections determine access.

Approve an AI agent only when its exact tools, memory paths, injection defenses, and high-impact actions have passed documented abuse tests—and material changes trigger retesting.

More than 100 organizations backed an August 27 cyber-defense appeal with specific requests for industry and government—but no binding investments, owners or deadlines.

Choose Claude Pro for repository coding evidence; choose ChatGPT Plus for terminal-agent work and broader model access. For API workloads, GPT-5.6 Sol currently costs less.

Gemini Daily Brief can read Gmail, Calendar, and earlier Gemini chats for a personalized summary. Source access, chat memory, and retained activity require separate controls.

On August 26, Google expanded Gemini Enterprise billing with pay-as-you-go pricing and hard project caps that temporarily pause agent API calls when spending reaches the limit.

On August 21, OpenAI cut GPT-5.6 Sol standard API pricing to $4 per million input tokens and $20 per million output tokens for three months; ChatGPT subscriptions did not change.