GitHub Copilot is reducing token costs by optimising the outcome of a task rather than minimising the number of tokens in a single tool response.
In this article
Counting tokens in an individual interaction does not measure efficiency. A brief answer that omits necessary details forces the agent to make extra calls, which increases the total cost and time. The focus has shifted to optimising the final result instead of the tool call itself.
This post outlines four changes in GitHub Copilot that apply this principle:
- Keeping useful context while cutting repetitive output.
- Removing formatting that adds no value.
- Shortening instructions without altering behaviour.
- Providing completed background work without an extra retrieval step.
Proposed changes were tested offline using agentic coding benchmarks. The most promising ones were validated through controlled online experiments before release. The examples below come from GitHub Copilot CLI. Other products, including the GitHub Copilot app and code review, use the same underlying system and have also improved.

The local metric trap
Reducing agent costs often involves shortening the output from each tool call. RTK (Rust Token Killer) is a utility that compresses shell output before an agent reads it. GitHub Copilot was evaluated using RTK with agentic coding benchmarks.
In the harness and benchmark configuration, RTK shortened some responses. However, when omitted text was important, the model sometimes reopened the original output or reran the command to recover what it needed.
Those recovery steps added turns and carried more context forward. The individual tool response was shorter, but the task used more tokens and took longer overall. Tokens were saved locally but spent more globally.

This result applies to the integration and workloads tested, not to every RTK configuration or output compression in general. Tokens per tool call is the wrong objective. An efficiency change must be evaluated across the complete task, from the user’s request through the final result.
It was more useful to look at what can be removed without making the model repeat work.
Compress noise, preserve useful information
The goal was to shorten repetitive output while preserving the context an agent needs to complete its task without retracing steps.
Analysis of benchmark runs showed that install, build, test, and lint output often contains repetitive noise, while source-like output and arbitrary command results are more likely to contain the information an agent needs. That analysis informed a selective output compressor, informed in part by RTK and similar approaches.
The prototype was evaluated on agentic coding benchmarks and a range of open source repositories, exercising their build, test, and lint systems.
Early versions were too aggressive. They made the model repeat work or read the full saved output, increasing end-to-end cost and reducing task success. For example, the system initially compressed git diff output but removed that filter after benchmark tasks showed agents reopening the original output to recover missing information.
Those early failures led to a three-part policy:
- Preserve source-like and arbitrary output. Commands such as cat, git diff, git show, and arbitrary scripts are returned unchanged.
- Reorganize search results without dropping content. Matches and file lists from tools such as grep can be grouped more efficiently while retaining every result.
- Compress repetitive noise selectively. Install, build, test, and progress output is compressed only when the savings are substantial.
The shipped version emerged through repeated evaluation and refinement. It is conservative not because the goal was to build a conservative compressor, but because that is what the evaluations supported.
When output is compressed, the agent can still retrieve the complete original through a direct recovery path.

That recovery path is both a safety mechanism and an evaluation signal. We tracked whether the agent opened the saved original, reran commands, repeated exploration, narrowed its searches, or took additional turns. Frequent recovery would indicate that the compressor had removed something valuable.
On offline tasks where output compression triggered, no statistically significant task-success regression was detected, and agents extremely rarely opened the saved originals. In the online experiment, average cost decreased slightly with no material regression detected in the tracked quality metrics.
Remove formatting before removing information
One clean token optimization came from the view tool, which agents use to read file contents into context.
Previously, view prefixed every line with a number before showing the contents to the model. Earlier file-editing tools used those numbers to target changes, but current tools instead match surrounding code and do not use line numbers. The line-number prefixes remained even though the normal workflow no longer used them.
Each prefix was small. Repeated across every line and every file read, however, that unused formatting accumulated throughout a session. So, we removed it.

Line numbers remain useful in diffs and short snippets. They were wasteful here because they were attached to every file read without serving the current editing workflow.
Removing them caused model-inference cost to fall by roughly 5% in offline agentic coding benchmarks. Success rates stayed within the expected run-to-run variance, and edit failures did not increase.
We then tested the change with Copilot CLI users. The online experiment reduced average daily model-inference cost per user by about 3%, with no material regression detected in the quality or satisfaction metrics we tracked.
For developers, that means more of the context window is available for the work itself rather than formatting the agent does not use.
This was the ideal change: no new instructions for the model, no source of information to recover, and no additional decision to make. The file contents reached the model unchanged.
Compress prompts without compressing intent
Prompts carry instructions that shape how an agent works, and they are sent to the model on every turn. Shortening them only improves efficiency if the agent keeps the behaviours developers depend on.
In GitHub Copilot, the task tool launches specialised agents for parallel work. Its guidance had accumulated across tool descriptions, schemas, agent definitions, system instructions, and companion tools.
A meta-prompting loop, in which Copilot iteratively wrote its own prompt, reduced that prompt by roughly half. Copilot produced and refined smaller candidates, and targeted behavioural tests checked the requirements we wanted to preserve.
The first online experiment found a regression that the initial offline evaluations had missed. The meta-prompting loop had rewritten cautious parallelism guidance into a hard scheduling policy, causing independent custom agents to run sequentially.
We stopped the experiment. Before changing the prompt again, we wrote a regression evaluation for the behaviour users had exposed. The eventual fix replaced an explicit allowlist and denylist with one sentence:
Independent agents can run in parallel.




