2 Sources
[1]
Transcription is 10 cents an hour. The audio is the moat.
SpaceXAI says its new transcriber is twice as accurate as its last one at the same price. Its evaluation sets are drawn from production traffic, including one of spoken account codes and email addresses. Grok Voice Transcribe 2.0 was released on 18 September, priced at $0.10 per hour of audio for
[2]
Grok releases Voice Transcribe 2.0 with improved accuracy By Investing.com
Investing.com -- Grok launched Voice Transcribe 2.0 on Friday, a speech-to-text model that delivers twice the accuracy of its predecessor at the same price point. The company said the new model ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard. The
Share
Copy Link
SpaceXAI released Grok Voice Transcribe 2.0, claiming twice the accuracy of its predecessor at unchanged pricing of $0.10 per hour for batch work. The transcription model ranks first among 32 streaming models on the Artificial Analysis leaderboard, with notable gains in processing telephony audio, multilingual commands, and spoken credentials.
SpaceXAI launched Grok Voice Transcribe 2.0 on September 18, positioning the speech-to-text model as twice as accurate as its predecessor while maintaining the same $0.10 per hour pricing for batch transcription and $0.20 per hour for streaming
1
. The transcription model now ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard, though the company carefully describes itself as "one of the most accurate" rather than claiming outright supremacy1
. The release comes from SpaceXAI, the name xAI has carried since July following SpaceX's acquisition in February, though industry coverage continues to reference the company as xAI1
.The new model demonstrates substantial improvements in processing real-world audio conditions including poor phone connections, multiple speakers, regional accents, and spoken credentials
2
. In telephony tests using 8 kHz audio from customer-support calls, word error rate dropped from 10.6% to 7.1%2
. For conversational audio, the error rate fell from 8.7% to 3.3%, while transcription of credentials such as phone numbers and email addresses saw improvements from 7.2% to 3.2%2
. The most dramatic gains appeared in multilingual processing, where short phrases across 19 languages showed word error rate reductions from 20.6% to 6.8%2
.SpaceXAI explicitly attributes its edge to training on what it describes as a unique dataset of live, noisy, multilingual audio recorded across diverse environments
1
. The company's audio foundation model powers tens of thousands of customer-support calls daily, transcribes millions of hours of video narration, and operates the Grok assistant in Tesla vehicles1
2
. This production pipeline gives SpaceXAI access to telephony, accents, crosstalk, and Tesla in-car commands that competitors lack1
. The company measures word error rate on four internal evaluation sets drawn directly from production traffic: customer-support telephony, conversations with Grok, short multilingual voice commands, and spoken credentials including account codes, phone numbers, email addresses, and physical addresses1
.Holding price while claiming doubled accuracy represents a statement about market dynamics rather than just model performance
1
. Features like speaker diarization, word-level timestamps, and key term biasing are bundled at no extra cost, despite being billable features until recently1
2
. The commoditization pressure is industry-wide, with OpenAI releasing new voice API models and Microsoft benchmarking its in-house transcription model against Whisper, Gemini, and ElevenLabs across 25 languages1
. At $0.10 per hour, transcription is being positioned as a commodity input rather than a premium service1
.Related Stories
Atlassian Loom switched to Grok Voice Transcribe 2.0 after finding it more accurate than its existing supplier, now using it to transcribe all video recordings
1
2
. "We've always believed the best way to move work forward is to capture context once and let it flow everywhere," said Sanchan Saxena, SVP of Teamwork Collection at Atlassian2
. This represents a concrete customer switch by a named enterprise buyer, carrying more weight than benchmark charts alone1
.The model supports dozens of languages with automatic detection and can follow mid-recording language switches
2
. Short commands give models minimal context to identify language, which directly addresses the Tesla in-car problem where drivers issue brief voice commands1
. The largest reported improvement in the 19-language voice-assistant utterance set lands precisely on the use case SpaceXAI owns through Tesla1
. The Speech-to-Text API will default to version 2.0 in coming weeks, with version 1.0 being deprecated within weeks and customers moved to the new model regardless of whether they requested the transition1
2
.Summarized by
Navi
[1]
1
Technology

2
Technology

3
Science and Research
