Writer Launches Palmyra X6 Model, Slashing AI Agent Costs by 52% as Enterprise Token Spending Surges

2 Sources

Share

Writer unveiled its Palmyra X6 flagship model alongside rebuilt orchestration tools, delivering 52% lower AI agent costs and 48% faster performance. The enterprise AI agent platform now includes governance controls to manage exploding token consumption as Fortune 500 companies scale agentic workflows.

News article

Writer Tackles Enterprise AI's Biggest Cost Challenge

Writer, the enterprise AI agent platform serving Fortune 500 companies including Accenture, Uber, and Vanguard, released its Palmyra X6 flagship model with performance metrics that directly address the economics crisis facing agentic AI deployments

1

. The company reports its AI agent platform now operates at 52% lower AI agent costs, with a 48% improvement in speed and a 10% improvement in quality when paired with Palmyra X6

2

. These agentic AI improvements arrive as Goldman Sachs forecasts token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month

1

.

Why Token Consumption Explodes With AI Agents

Unlike chatbots that generate one answer per request, an AI agent transforms a single query into repeated rounds of planning, retrieval, tool calls, validation, and retries. Every loop consumes metered tokens. Users see one answer; invoices reflect the entire computational loop. Matan-Paul Shetrit, Writer's director of product management, identified cost as the primary barrier to enterprise AI adoption: "The biggest barrier today to enterprise expansion using AI is actually not model capabilities in most cases; it's actually the cost around them. The reality today is, in most cases, the alternative for AI is not another AI, it is human labor"

1

. Co-founder and CTO Waseem AlShikh stated: "The enterprise wants token consumption to explode -- it means adoption is happening -- but they need costs to flatten"

2

.

Technical Foundation Built on Chinese Open-Source Model

Palmyra X6 is a 744-billion-parameter mixture-of-experts model post-trained on GLM-5.2, the open-weight model from Beijing-based Z.ai, formerly Zhipu AI

1

. Writer openly discloses this foundation in its technical report, placing the San Francisco company at the center of debates about American enterprises building on Chinese open-source foundations. Shetrit emphasized: "This model is in no way, shape, or form connected to any of its original developers. It is fully run on our U.S. infrastructure"

1

. Dan Bikel, who leads Writer's AI research, stated: "It's very much a Palmyra model, and we just happen to grab the floating point numbers as the starting point, and train from there"

1

.

Conservative Training Approach Yields Enterprise-Ready Performance

Writer applied anchored supervised fine-tuning (ASFT) to just 626 curated synthetic agentic trajectories, trained for a single epoch at a low learning rate

1

. The tiny dataset represents a deliberate strategy. ASFT pairs token-weighting with a KL-divergence anchor that prevents the fine-tuned model from drifting too far from the base model, teaching new tool-use behaviors without eroding general capabilities. Bikel referenced research showing "small, extremely high quality data sets go a really long way"

1

. On evaluation, Palmyra X6 scored an average of 0.87 across nine evaluations at $2 per million input tokens and $8 per million output tokens, ahead of Claude Opus 4.8 at 0.86, GPT-5.5 at 0.80, and Gemini 3.1 at 0.77

2

.

Governance Tools Address Runaway Spending

Writer launched governance tools providing administrators visibility into activity across teams

2

. Employees can now build and share Playbooks and Skills, which are repeatable multistep workflows and instructions. A centralized dashboard surfaces the most common use cases, allowing admins to filter and compare over periods to understand adoption, performance, and spend. A dedicated analytics section for Playbooks and Skills lets admins monitor performance and token costs at individual task or skill levels. Consumption levels and alert limits can be set to specific thresholds to help leadership control spending across teams

2

. The model completes tasks in 26 seconds on average, produces 82 tokens per second, and can work unattended toward long-term goals for up to eight hours, sustaining coherent reasoning across multistage objectives without drifting

2

.

Expanding the Total Addressable Market Through Lower Unit Economics

Shetrit rejected the notion that cutting customers' token consumption would cannibalize Writer's per-token revenue: "Reducing the cost is not hurting my bottom line. It's actually expanding it, because it's expanding the TAM of opportunity within an organization." He argued that lower per-task costs unlock workflows enterprises would otherwise never automate

1

. This echoes a pattern from the cloud era, where unit prices fell for a decade while total bills rose as consumption expanded. Writer is betting this dynamic will repeat with AI agents. The company also highlighted that Palmyra X6 is the least politically biased model Writer evaluated, important for marketing and revenue teams to reach audiences and produce neutral, brand-aligned content that can be customized easily to brand guidelines

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved