Writer Slashes AI Agent Costs by 52% with Palmyra X6 Model and Upgraded Harness

5 Sources

Share

Writer launched Palmyra X6, its new flagship AI model built on Z.ai's open-source GLM-5.2, alongside an upgraded harness infrastructure. The enterprise AI firm claims the combined system cuts costs by up to 52% while delivering 48% faster speeds and 10% quality improvements, directly addressing mounting concerns over token consumption as agentic AI deployments multiply enterprise expenses.

News article

Writer Tackles Enterprise Token Cost Crisis with Palmyra X6

Writer, the enterprise AI firm serving Fortune 500 companies including Accenture, Uber, and Vanguard, launched Palmyra X6 on Thursday alongside significant upgrades to its agent orchestration infrastructure

1

. The new flagship AI model and upgraded harness aim to address what CEO May Habib calls an unprecedented cost explosion in enterprise AI deployments

1

. Writer claims the combined system operates at 52% lower cost on average, with a 48% improvement in speed and a 10% improvement in quality when paired with Palmyra X6

3

. The enterprise AI firm estimates basic tasks could cost up to 50% less under the new setup, directly targeting the mounting pressure on IT budgets as agentic AI deployments multiply token consumption across organizations

1

5

.

Building on Z.ai's Open-Source GLM-5.2 Foundation

Palmyra X6 represents a strategic departure from training models from scratch. Writer developed the model through post-training work on Z.ai's open-source GLM-5.2, a mixture-of-experts architecture with 744 billion parameters and roughly 40 billion active parameters per token

3

. The company applied anchored supervised fine-tuning to just 626 curated synthetic agentic trajectories, trained for a single epoch at a low learning rate

3

. Writer's director of product management, Matan-Paul Shetrit, emphasized that the model runs entirely on U.S. infrastructure and is disconnected from its original developers

3

. The model scored an average of 0.87 across nine evaluations at $2 per million input tokens and $8 per million output tokens, outperforming Claude Opus 4.8 at 0.86, GPT-5.5 at 0.80, and Gemini 3.1 at 0.77

4

. Writer also highlighted that Palmyra X6 is the least politically biased model it evaluated, which matters for marketing and revenue teams producing neutral content customizable to brand guidelines

4

.

Harness Optimization as the Hidden Cost Lever

Writer positions its upgraded harness as equally important as the model itself for achieving cost efficiency in AI. The orchestration layer wraps around models to manage how multi-step workflows execute each request, and Writer argues that optimizing this component multiplies efficiency gains across every model an organization runs

1

. A recent paper from Writer researchers found that harness-efficiency changes alone trimmed costs by roughly 40% on average across testing, suggesting that orchestration improvements can be more reliable than model choice for reducing token costs

1

2

. The harness is model-agnostic, working with Writer's own models as well as external ones served through Microsoft Azure and Amazon Bedrock, ensuring efficiency gains are not locked to a single vendor's roadmap

2

. By trimming redundant calls and bloated context that agents accumulate during multi-step workflows, the upgraded harness reduces token consumption without touching the underlying model

2

.

Why Agentic AI Is Exploding Enterprise Budgets

Unlike chatbots that generate one answer per user request, AI agents turn single requests into repeated rounds of planning, retrieval, tool calls, validation, and retries, with every loop consuming metered tokens

3

. Goldman Sachs forecasts that token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month, driven by always-on enterprise agents rather than more people asking questions

3

. The same analysis warned that falling per-token prices do not guarantee falling bills: if an agentic task draws 20 times more tokens while unit prices fall 75%, total charges still rise fivefold

3

. CTO and co-founder Waseem AlShikh stated that enterprises want token consumption to explode because it signals adoption, but they need costs to flatten

3

. Shetrit framed cost as the primary obstacle to enterprise AI adoption, noting that in most cases, the alternative to AI is human labor, making the total addressable market dependent on achieving cost parity

3

.

Distrust of Major AI Labs Drives Strategic Shift

May Habib expressed blunt criticism of major AI labs, arguing that enterprises are increasingly distrustful of providers with financial incentives to drive up token use

1

. "The enterprise is absolutely sick of chasing the next benchmark," Habib told TechCrunch. "They want flattening cost, and it seems like nobody can deliver that"

1

5

. She added that the cost explosion is unprecedented and that CIOs are giving up on the labs because they "don't deeply understand how to help an enterprise get benefit from AI"

1

. This positioning aligns Writer with a wider thrift-maximizing trend where buyers reach for cheaper Chinese models, unsettling the valuations underpinning eventual OpenAI and Anthropic IPOs

2

. The company's bet is that lower per-task costs unlock workflows enterprises would otherwise never automate, expanding the total addressable market rather than cannibalizing revenue

3

.

New Governance Tools and Long-Task Capabilities

Writer launched major upgrades to reporting and governance tools alongside Palmyra X6, providing administrators clear visibility into activity across teams

4

. Employees can now build and share Playbooks and Skills, which are repeatable agent workflows and instructions surfaced in a centralized dashboard

4

. Administrators can filter and compare usage over periods to understand adoption, performance, and spend, with consumption levels and alert limits set to specific thresholds

4

. Palmyra X6 completes tasks in 26 seconds on average, produces 82 tokens per second, and can work unattended toward a long-term goal for up to eight hours, sustaining coherent reasoning across long, multistage objectives without drifting off task

4

. These capabilities enable complex, multi-step workflows at scale while giving leadership control over runaway spending

4

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved