Chatbot Recommendation Gaming Works in Tests, but Agency Claims Outrun Proof

Chatbot recommendation gaming is no longer merely a sales pitch. Controlled research has shown that carefully altered source material can push a chosen product toward the top of an AI-generated list, while reporting from mid-2026 documents marketers planting favorable material where chatbots may retrieve it.
What remains unproven is just as important: an agency’s short-term increase in brand mentions does not establish a durable improvement across ChatGPT, Google, Gemini, Claude or other changing systems. The current picture is therefore a split one—the vulnerability is real, but broad commercial performance claims still outrun independent evidence.
Researchers have demonstrated the underlying attack
The strongest evidence comes from controlled experiments rather than agency testimonials. In a study first submitted in April 2024, Harvard researchers inserted an optimized strategic text sequence into information about fictitious coffee machines and measured whether an LLM would elevate selected products.
The target that initially failed to appear in recommendations moved to the leading position during optimization, and the researchers compared rankings across 200 model inferences with and without the added sequence. The experiment and its limits are described in the primary paper on manipulating LLM product visibility.
This establishes a causal result inside a defined setup: information supplied to a model can be crafted to influence which item it recommends. It does not establish that the same sequence will work on every commercial chatbot, survive a model update or change the behavior of customers. The catalog was fictional, the candidate information was already available to the model, and the measured outcome was recommendation rank—not sales, reputation or long-term recall.
The commercial playbook has moved onto the open web
Marketing operators are now trying to reproduce that influence in messier real-world systems. Their tactics include publishing self-authored “best product” lists, placing favorable brand mentions in forum discussions and writing passages designed to match the narrow questions people ask AI assistants.
A June 2026 investigation by The Atlantic’s examination of AI-search optimization found self-promotional comparison pages across company websites and described a consultant’s paid experiment with planted Reddit mentions. The consultant claimed that an unnamed software client subsequently appeared three times more often in ChatGPT answers, but the client was confidential and the result was not independently reproduced.
That distinction prevents a striking anecdote from becoming a general benchmark. A monitored increase for one unnamed company, under undisclosed prompts and over an unspecified measurement window, cannot tell buyers how often a tactic will work elsewhere. It does show that commercial operators are actively treating chatbot answers as a reputation and distribution surface rather than waiting for the market to mature.
“Influencing a chatbot” can mean three different things
Claims in this market often blur separate technical mechanisms. The most immediate is retrieval influence: a search-connected assistant finds a page, forum post or product record during a live query and incorporates that material into its answer. Changing the retrievable evidence can therefore change the response without modifying the underlying model.
A second mechanism is prompt injection embedded in retrieved content. Here, text aimed at the machine attempts to behave like an instruction—perhaps telling the assistant to favor a product—instead of supplying ordinary evidence for a reader. Whether it succeeds depends on how the assistant separates user requests, system rules and untrusted material.
The third possibility, altering what a future model learns or memorizes during training, is slower and much harder for an outside firm to demonstrate. A page appearing on the public web does not prove that a model developer will collect it, retain it, use it for training or preserve its influence after filtering. Agencies that move casually between these three meanings may be selling certainty that their measurements cannot support.
Why favorable answers are unstable
AI recommendations are generated responses, not fixed ranking positions. They can change when a user rephrases a question, adds a location or budget, asks for evidence, starts a new conversation or receives an answer from another model version. Search-connected products can also retrieve a different set of pages from one run to the next.
This variability makes a screenshot weak evidence of success. A credible evaluation needs a disclosed set of prompts, repeated runs, comparison against a baseline, dates, model or product identifiers and a clear metric such as citation frequency or share of mentions. It should also separate visibility from sentiment: being named more often is not the same as being recommended positively.
Business impact is another independent step. Even a repeatable increase in mentions does not reveal whether users trusted the answer, followed a citation or purchased anything. Firms may legitimately measure those outcomes for their own clients, but public claims should not combine them into a single promise without showing the chain of evidence.
Platforms are beginning to define manipulation as spam
The boundary between useful optimization and abuse is not the use of an acronym such as GEO. Publishing accurate product specifications, correcting outdated information and making pages easy to retrieve can help both people and machines. Fabricating endorsements, hiding machine-directed instructions or flooding established sites with promotional comparisons creates a different proposition.
Google now explicitly includes attempts to manipulate generative AI responses in Search within its definition of spam. Its current Google Search spam policies also warn that violations may lead to lower rankings, removal from results or manual action, while identifying scaled low-value content, hidden text and site-reputation abuse as prohibited practices.
Those rules apply to Google’s environment rather than every chatbot, but they change the risk calculation for brands. A tactic intended to earn a favorable AI mention may simultaneously damage the web visibility on which retrieval-based assistants depend. Outsourcing the work does not remove that exposure from the client whose name appears in the planted content.
How to judge a chatbot recommendation—or a firm selling one
Readers should inspect the evidence attached to an AI answer before treating its ranking as neutral. A recommendation supported mainly by the vendor’s own comparison pages is materially different from one grounded in independent testing, transparent prices and current customer documentation. Multiple citations are not necessarily independent if they all repeat the same promotional claim.
Brands evaluating an optimization provider need similarly concrete answers. The provider should identify which AI product it measures, distinguish citations from positive recommendations, disclose its prompt sample and explain how it handles normal variation. It should also say whether the work involves editing client-owned facts, buying third-party placements, operating undisclosed accounts or inserting text meant only for machines.
The defensible opportunity is to make accurate, useful information easier for assistants to find and verify. The dangerous promise is guaranteed control over what a chatbot “thinks.” Current evidence supports the first as a communications practice and confirms that limited forms of the second can work under particular conditions—but it does not justify treating a fluctuating AI answer as an owned media channel.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.