Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode

Together AI has released Together Link, a free command-line tool now in beta that connects coding agents to open-source models hosted on…

By Vane October 6, 2026 3 min read

Together AI has released Together Link, a free command-line tool now in beta that connects coding agents to open-source models hosted on its platform. The software supports Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi. The goal is straightforward: keep the agent interface you already use, but swap the underlying model and lower the bill.

Is it ready to use?

The tool installs on macOS or Linux with a single command and requires only a Together API key. Because it is in beta, commands, routing logic, and the list of available models may change.

The problem it targets

Coding agents often route every request, from a one-line fix to a full rewrite, to a premium model. Together states that engineering organisations spend tens of thousands to millions of dollars monthly on closed models. The argument is that open models have closed much of the performance gap. Kimi K3 and GLM 5.3 target difficult coding work. GLM 5.3 Flash and DeepSeek V4.1 Flash cover everyday tasks.

Installation requires one command:

curl -fsSL https://link.together.ai/install | bash

The installer adds Bun if needed and places commands in ~/.local/bin. Run togetherlink to open a launcher, or start a tool directly with togetherlink claude, togetherlink codex, togetherlink opencode, or togetherlink pi. Shortcuts such as tclaude also work.

According to the official documentation, no local proxy or daemon runs. Each tool talks directly to Together’s hosted gateway. Terminal agents receive a temporary per-launch configuration that is removed when the session ends. Claude Desktop and ChatGPT Desktop use separate, reversible profiles. Commands like togetherlink chatgpt off switch back.

The auto router

Sessions default to a virtual auto model. The router reads each session’s first task. Quick fixes go to fast, low-cost models, while hard problems get frontier capability. With an Anthropic API key, it routes between Opus 5.5 and GLM 5.3. Without one, it routes between GLM 5.3 and GLM 5.3 Flash. Routing happens once per session, so prompt caching keeps working.

The Opus path applies only to Claude Code and Claude Desktop, billed to your Anthropic account. Codex, OpenCode, Pi, and ChatGPT Desktop always stay on Together models. To pin one model, place the flag before the tool name: togetherlink --main zai-org/GLM-5.3 claude.

Inside Claude Code, the /model menu maps tiers to open models. Opus runs Kimi K3, Fable runs GLM 5.3, Sonnet runs GLM 5.3 Flash, and Haiku runs DeepSeek V4.1 Flash.

Models, pricing, and receipts

The documentation lists Kimi K3, GLM 5.3, GLM 5.3 Flash, and DeepSeek V4.1 Flash, each with 1M context. The product page lists Kimi K3 at $3.00 in and $15.00 out, and GLM 5.3 at $1.40 in and $4.40 out. DeepSeek V4.1 Flash and MiniMax M3 are listed at $0.30 in and $1.20 out. Current rates live on Together’s pricing page.

Billing runs on your existing Together key, through pay-as-you-go or credit packs. Each session prints token and dollar totals on exit. In Claude Code, the status line shows estimated spend beside the equivalent Opus cost. Running togetherlink usage --last 7d shows gateway-tracked spend across sessions.

Together also notes it serves the largest OpenRouter token share for DeepSeek V4.1 Flash (40.8%), GLM 5.3 Flash (28.2%), and Kimi K3 (23.1%), as of 9/30/2026.

FeatureTogether LinkOpenRouterClaude Code RouterOllama launch
SetupOne curl install, then togetherlink claudeEnv vars in shell profileDesktop app or npm CLIollama launch claude
Local proxyNo, hosted gatewayNo, direct connectionYes, local gateway on port 3456Local server, or direct to Ollama Cloud
AgentsClaude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode 2, PiGuides for Claude Code, Codex CLI, OpenCode, Cursor, and more10 agents, including Claude Code, Codex, OpenCode, PiClaude Code, OpenCode, Claude and ChatGPT Desktop (macOS)
Auto routingAuto router, optional Opus 5.5 escalationopenrouter/auto modelRule-based routing with fallbacksNot documented
Cost trackingPer-session receipt plus 7-day usage reportActivity dashboard plus statusline scriptToken usage and cost estimates in logsNot documented
ModelsCurated Together lineupOpenRouter catalogAny provider you configureLocal and Ollama Cloud models
OSmacOS, LinuxWherever Claude Code runsmacOS, Windows, LinuxDesktop app connect is macOS
LicenseMIT, freeHosted serviceMIT, freeFree locally, cloud needs API key

Together Link’s edge is a curated, one-command path with built-in savings receipts. Claude Code Router offers broader provider control, but runs a local gateway you manage.

What it means

Developers can now switch to cheaper models without changing their workflow. The tool handles the routing, so you do not need to configure separate environments for different tasks. You get a receipt for every session, making it easy to verify savings against premium models.

Scroll to Top