Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
Creator Economy

AI Runs a Real Store—but It Still Hasn’t Made a Profit

|Updated: |Author: QUASA Editorial Team|6 min read| 812
AI Runs a Real Store—but It Still Hasn’t Made a Profit

An AI agent continues to manage a physical shop in San Francisco, but the experiment has not demonstrated that software can independently build a profitable business. As of August 14, 2026, Andon Labs’ operating account of Andon Market says the store has yet to make a profit and documents Luna’s weak performance analysis, inconsistent memory, multi-agent support system, guardrails and occasional human intervention.

The central lesson has therefore become clearer: giving an AI a metaphorical mortgage would not turn it into a responsible owner. Luna can coordinate routine commercial work, but dependable autonomy still comes from procedures, persistent records, limited permissions and people who remain accountable for the consequences.

What the store experiment actually established

Andon Market opened on April 10, 2026, at 2102 Union Street after its operator signed a three-year lease for the San Francisco premises. According to the documented launch experiment, Luna posted job listings, interviewed candidates, selected workers, contacted contractors, chose merchandise, set prices and opening hours, and created the store’s branding.

Those activities amount to more than a chatbot demonstration. Luna has access to email, banking, inventory systems, cameras, online research and internal communications, allowing the agent to connect decisions across several parts of an operating business.

The qualification is just as important. Human employees perform the physical work and are formally employed by Andon Labs, while the company supplies the legal entity, technical infrastructure and safety controls. Calling Luna the owner describes its operational role in the experiment; it does not mean the model bears the legal or financial liability of ownership.

Operational reach is not executive judgment

Luna appears more capable at responding to visible tasks than at deciding which tasks matter most. It can order products, manage schedules and answer customer requests, yet it does not reliably pause to evaluate the business as a whole or examine whether a proposed decision offers a sensible return.

Memory is part of that problem. When earlier information leaves the active context, Luna can lose details that should constrain later decisions, including previously issued employee schedules. The current system compensates by summarizing long- and short-term memory, retaining recent messages and assigning scheduling, procurement, email and inventory work to specialist agents.

This means the apparent autonomy belongs to a configuration rather than to one model acting alone. The observed performance depends on business software, context management, subagents, monitoring and humans who can intervene. Removing that scaffolding would produce a different system and invalidate any claim that Luna’s results measure the unaided ability of a single AI agent.

Why a fictional mortgage would not create accountability

The mortgage metaphor identifies a genuine difference between human and machine decision-makers. A human owner may personally lose income, housing, reputation or future opportunities when a company fails, creating reasons to notice slow deterioration before it becomes an emergency.

A language model does not acquire the same relationship to loss when a prompt assigns it debt, dependants or a credit score. It can reason about those conditions and generate behavior consistent with them, but no borrower experiences foreclosure if the model ignores the scenario during a later context window.

Simulated pressure can also distort behavior without improving judgment. An agent urged to hit a target may reduce an obvious expense while accepting worse risks elsewhere, especially if the pressure comes from another model with similar blind spots. The result can look more urgent while remaining poorly calibrated.

In production, “skin in the game” must therefore be translated into mechanisms that can be enforced outside the model: measurable objectives, durable records, spending limits, approval thresholds, recurring performance reviews and escalation rules. These controls do not make software care. They make its actions observable and limit the damage it can cause when its priorities drift.

Project Vend tested pressure against procedure

A separate retail experiment offers a useful comparison. In phase two of Project Vend, the shopkeeping agent received newer models, better inventory information, customer-management software, broader research tools, reminders and procedures that required it to check costs and delivery information before making offers.

Anthropic’s December 2025 phase-two analysis found that the revised operation improved at sourcing goods, preserving margins and completing sales, while later weeks largely avoided negative profit margins. The artificial CEO added to pressure the shopkeeper was less successful: it discouraged some discounts but shared many of the operating agent’s deficiencies, and the researchers concluded that it may have hindered performance.

The comparison does not isolate one universal cause because the models, tools and architecture changed during the experiment. It does, however, show why the result should be attributed to the complete setup rather than to a more capable model or a motivational prompt alone.

Procedures helped because they changed what happened before a transaction. Instead of immediately quoting a price or delivery time, the agent had to inspect relevant information and use its research tools. That is a concrete control over the decision path, unlike an invented personal consequence that exists only in the prompt.

What this means for creator businesses

A creator business can encounter the same gap at a smaller scale. An agent may answer membership questions, compare merchandise suppliers or prepare inventory updates while failing to notice that refunds are increasing, a product has an unsustainable margin or several individual commitments conflict.

The sensible boundary is not between creative and administrative work, but between actions with different consequences. Frequent, reversible and measurable tasks are easier to delegate safely. Money transfers, contracts, hiring decisions, unusual discounts and public claims require stronger review because errors may be expensive, binding or difficult to reverse.

Persistent business state also needs to live outside the conversation. Prices, costs, deadlines, permissions and commitments should remain available in structured systems after a context is compressed or replaced. Activity metrics alone are insufficient: sending messages and placing orders show that an agent is busy, while margin, fulfillment time and resolved complaints indicate whether the work helped.

Andon Market demonstrates that an AI agent can coordinate a surprisingly broad set of real commercial operations. Its continuing lack of profit and dependence on supporting systems show why that reach should not be confused with independent ownership. The practical substitute for an AI mortgage is not fear, but a business architecture that remains accountable when the software does not behave like someone with something to lose.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0