Writer introduces new AI model and upgraded harness to contain token costs

Users across the AI industry are paying more attention to deployment costs and feeling a new urgency to cut them. While open-source…

By Vane August 13, 2026 2 min read
Writer introduces new AI model and upgraded harness to contain token costs

Users across the AI industry are paying more attention to deployment costs and feeling a new urgency to cut them. While open-source models offer lower per-token rates, finding the right one for a specific job remains difficult. On Thursday, Writer, which provides AI tools and agents for marketers, launched a new flagship model called Palmyra X6 to solve this problem. Built as a post-training variation on Z.ai’s open-source model GLM-5.2, Writer says the new system should provide deployment-ready capabilities at a much lower price. The company estimates the new model, combined with changes to the company’s harness infrastructure, will cut costs for its customers by as much as 50 percent for basic tasks.

Alongside the new model, the company also released significant upgrades to its standard agentic harness. Both features will be available to Writer clients starting Thursday.

“I think the enterprise is absolutely sick of chasing the next benchmark,” CEO May Habib told TechCrunch. “They want flattening cost, and it seems like nobody can deliver that.”

The new approach puts particular emphasis on complex, multi-step tasks, executed faster and with fewer tokens. Writer sees harness optimisation as a crucial lever towards making that happen.

A recent paper from Writer researchers lends credence to this approach, testing small changes in harness efficiency across multiple different models. The research found that, in many cases, changes in the harness were a more reliable way to reduce costs than model choice, with costs falling an average of 40% across their testing.

“The harness is the one component whose efficiency multiplies across every model an organization runs—present and future,” the researchers wrote.

For Writer’s clients, the experience is still model-agnostic: Palmyra X6 will sit alongside other Writer models or outside models imported through Azure or Amazon Bedrock. But Habib also sees the push to cut costs as driving a broader distrust towards major AI labs, who have a financial incentive to drive up token use.

“The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs,” Habib told TechCrunch, adding that the AI labs “don’t deeply understand right how to help an enterprise get benefit from AI.”

What it means

For people using these tools, the shift means they can run complex workflows using fewer tokens. The focus moves away from searching for the perfect model and towards using a system that manages how prompts are handled more efficiently. This allows teams to keep spending lower while maintaining performance on multi-step tasks.

Scroll to Top