Together Link Swaps Coding Models—but Its Savings Claim Needs Testing

Together AI put Together Link into beta on October 5, 2026, connecting familiar coding agents to open models hosted by a different provider, as MarkTechPost’s launch report records. The supported integrations cover Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode and Pi. The beta installs on macOS or Linux and requires a Together API key.
Together AI’s launch post pitches “Same agent, same workflow, a fraction of the bill,” claims savings of more than 50%, and describes Auto selecting a model from the opening task for the whole session. The tool itself is free, but inference is metered. A lower price for a model call does not establish a lower cost for a coding change that must pass review and tests.
The harness stays while inference moves
Link changes the model connection underneath an agent, rather than replacing the agent’s interface. Developers still work in their existing Claude Code, Codex or OpenCode environment and keep the repository context, commands and tool workflow supplied by that harness. Model requests go through Together’s hosted gateway. Open weights describe the models available through the service; they do not mean that those requests run on the developer’s own computer.
This distinction matters for teams with existing agent setups. The familiar interface can remain while answers, tool-use choices and token consumption change with the selected model. Repository context sent for inference is handled by a remote provider, an operational consideration that does not disappear because the model is open weight. The beta offers provider substitution within an existing workflow, without guaranteeing identical results from different models.
Installation and the configurations it touches
The installer is a shell command and expects Bash and curl, an installed supported agent, and a Together API key; Linux setup also requires unzip when installing Bun. The launcher can store the key through its configuration flow or read it from the shell environment. A team can launch a named terminal agent through Together Link instead of changing its usual entry point. That makes a trial a session choice, although paid model usage begins when requests are served.
Integration details differ by harness. Terminal agents receive settings for the launched run, while desktop apps use separate profiles. Claude Code keeps its login and existing settings; Codex keeps its normal configuration, history and sandbox settings. In OpenCode, the provider and router appear alongside existing providers, and Pi retains its extensions and sessions. Those preserved assets matter because a model switch need not force a team to rebuild the agent environment it already uses.
Auto routing creates a moving comparison
The current integration documentation describes Auto choosing a model for each request and lists GLM 5.3 at $1.40 per million input tokens and $4.40 per million output tokens, against $0.30 and $1.20, respectively, for DeepSeek V4.1 Flash. Its per-request account differs from the launch description of a route held for the session. For teams expecting stable routing and prompt caching, that difference affects how they interpret a session receipt; the exact route deserves attention while the product is in beta.
Claude Code and Claude Desktop have a further branch: with an optional Anthropic API key, difficult requests can go to Claude Opus and be billed to the Anthropic account behind that key. Without it, those requests stay with Together-hosted models. Codex, OpenCode, Pi and ChatGPT Desktop stay on Together models under the documented Auto behavior. A team wanting a controlled model comparison can select a particular model, separating route choice from the agent interface used to perform the task.
Receipts show spend, not accepted work
Each proxied session prints its token and dollar totals on exit, and Together Link has a usage command for recent spending. The launch comparison sets that spend beside an estimate of what the session would have cost using Opus. That second figure is a counterfactual price calculation, not a second run that produced an independently checked result. An Anthropic route can also create charges in another account, so the Together total alone may not capture the full model bill for a mixed session.
The vendor’s saving is plausible at the token-price level for some model choices, yet workload cost depends on how much work the selected model needs to finish an accepted change. Retries, longer conversations, extra tool calls and human repair can eat into a lower rate. Conversely, a cheaper model that completes an ordinary task with the same effort could reduce spend. Neither the launch figures nor the receipt establishes the claimed saving across a team’s own mix of tasks and quality requirements.
Returning to the original route
Terminal integrations leave normal agent configuration available after the launched session ends. Codex receives gateway settings for that run while its normal configuration remains untouched; Claude Desktop and ChatGPT Desktop use separate profiles. Turning off a desktop integration returns the app to its usual profile. A provider trial therefore need not become a permanent change to the agent setup.
The desktop profile can persist after the integration is turned off. ChatGPT Desktop’s Together profile has its own local tasks and settings, and desktop profiles store the Together API key locally. Its reset command removes the Together profile and those local artifacts while leaving the normal profile intact. Teams ending a beta trial can thus distinguish returning to their usual provider from clearing the credentials and work retained in the trial profile.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.