OpenAI’s “Strawberry” Is Deprecated, but Simple Puzzles Still Expose the Gap

OpenAI’s o1-preview, the reasoning model introduced under the “Strawberry” code name, is no longer a current frontier product. As of August 13, 2026, the official o1 model catalog identifies both o1 and its September 2024 preview snapshot as deprecated.
That change resolves one part of the original controversy: “Strawberry” should not be treated as a new release today. The more durable lesson is that extra reasoning time can produce striking results on difficult tests without guaranteeing reliability on an easy-looking problem, particularly when a familiar question is modified in a small but meaningful way.
What OpenAI actually released in September 2024
OpenAI launched o1-preview on September 12, 2024, as the first public version of a model family trained to deliberate before answering. Contemporary reporting on the Strawberry launch documented the code name, the staged rollout to ChatGPT users and OpenAI’s claim that the model performed similarly to PhD students on demanding physics, chemistry and biology benchmarks.
The important qualification was embedded in the product name: this was a preview, not a general declaration that machine reasoning had been solved. Its deliberate process was designed to explore alternatives and correct mistakes before presenting an answer. That mechanism improved performance on several structured tasks, but it did not turn every generated conclusion into a verified result.
Early examples involving misspelled words, counting, chess positions and riddles attracted attention because the errors looked absurd beside the model’s advanced test scores. Individual social-media prompts, however, were weak evidence by themselves. A response could vary between runs, the prompt might omit a necessary condition, and an isolated screenshot revealed little about the model’s overall error rate.
Hard benchmarks and elementary mistakes can coexist
There is no logical contradiction between solving many specialist questions and failing a short puzzle. A benchmark measures performance over a particular distribution of tasks under defined scoring conditions. It does not prove that the same system will transfer the underlying rule correctly to every new wording, format or context.
A reasoning model can also produce a fluent path toward the wrong destination. More generated steps create opportunities to notice an error, but they can equally reinforce a mistaken assumption introduced near the beginning. Length and confidence therefore provide no dependable substitute for checking whether each premise matches the user’s actual problem.
This distinction matters more than the famous question about counting letters in “strawberry.” A misspelled word or a deliberately tricky riddle has little practical value on its own. The relevant failure appears when a model recognizes the surface form of a familiar problem, retrieves a standard solution pattern and neglects the condition that makes the new version different.
Later research turned anecdotes into a controlled warning
Evidence published after the launch gave the concern a firmer basis. The peer-reviewed 2025 RoR-Bench study tested elementary arithmetic and reasoning questions with subtly shifted conditions. Its authors reported that changing one phrase could produce a 60% performance loss in leading systems such as OpenAI o1 and DeepSeek-R1.
The study did not establish that o1 was uniformly bad at reasoning. It tested a narrower proposition: whether models that performed well on familiar problem forms would preserve that performance after a consequential variation. The resulting decline indicates that some apparent reasoning success can depend heavily on recognition of a learned solution template.
This is more informative than collecting amusing failures. Controlled variants hold most of a question constant and change the detail that should alter the answer. When performance collapses under that treatment, evaluators can identify a generalization problem rather than merely observing that a stochastic system occasionally makes a mistake.
What the result means for creators
For creators, o1’s deprecation does not make the lesson obsolete. Reasoning systems can help build outlines, compare alternatives, transform data and inspect arguments, but polished intermediate logic should be treated as generated material rather than an audit trail. The cost of an error depends on where that material will be published and what readers may do with it.
- Separate drafting from verification. Use the model to propose an explanation, then check names, dates, calculations and quoted claims against the underlying material.
- Test a conclusion with a nearby variant. Change one relevant assumption, quantity or constraint and ask what must change in the answer. A response that repeats the original conclusion may be following a template.
- Request checkable outputs. Tables with explicit inputs, formulas with substituted values and claims tied to supplied documents are easier to inspect than a long narrative about how the model reasoned.
- Keep human review proportional to consequence. A brainstorming error is inexpensive; an incorrect financial comparison, health claim or contractual interpretation can survive editing and mislead an audience.
Repeated prompting can reveal instability, but agreement across several runs is not proof: the same learned shortcut may recur each time. For publication work, the strongest check remains evidence outside the model, combined with a reviewer who understands which condition determines the answer.
The real legacy of “Strawberry”
o1-preview was a meaningful product milestone because it made extended inference a visible part of mainstream AI use. Its benchmark gains and its elementary failures were not mutually exclusive; together they exposed how uneven machine reasoning could be across different task shapes.
The model itself has moved into OpenAI’s deprecated catalog, so its old launch limitations should not be projected automatically onto every successor. What remains current is the evaluation principle revealed by later research: difficult benchmark success does not eliminate the need to test whether a system notices a small change that should produce a different answer.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.