2 Sources
[1]
Two undergrads built an AI speech model to rival NotebookLM | TechCrunch
A pair of undergrads, neither with extensive AI expertise, say that they've created an openly available AI model that can generate podcast-style clips similar to Google's NotebookLM. The market for synthetic speech tools is vast and growing. ElevenLabs is one of the largest players, but there's no
[2]
A new, open source text-to-speech model called Dia has arrived to challenge ElevenLabs, OpenAI and more
Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More A two-person startup by the name of Nari Labs has introduced Dia, a 1.6 billion parameter text-to-speech (TTS) model designed to produce naturalistic dialogue directly
Share
Copy Link
Two undergraduate students with limited AI expertise have developed Dia, an open-source AI speech model that challenges established players like Google's NotebookLM and ElevenLabs.

In a surprising turn of events, two undergraduate students with limited AI expertise have developed an open-source AI speech model that rivals industry giants. Toby Kim and his co-founder, operating under the name Nari Labs, have created Dia, a 1.6 billion parameter text-to-speech (TTS) model designed to produce naturalistic dialogue from text prompts
1
2
.Dia offers advanced features that set it apart from existing models:
The model runs on PyTorch 2.0 and CUDA 12.0, requiring about 10GB of VRAM. It can generate approximately 40 tokens per second on enterprise-grade GPUs like the NVIDIA A4000
2
.The creators of Dia leveraged Google's TPU Research Cloud program, which provided free access to the company's TPU AI chips for training. This resource was crucial in enabling the undergraduates to compete with well-funded companies in the AI space
1
.Nari Labs claims that Dia outperforms competing proprietary offerings from ElevenLabs, Google's NotebookLM, and potentially even OpenAI's recent gpt-4-0-mini-tts
2
. The company provides side-by-side comparisons on their website, demonstrating Dia's superior handling of:Dia is fully open-source, distributed under the Apache 2.0 license, allowing for commercial use. The model is available for download from Hugging Face and GitHub, and can run on most modern PCs with at least 10GB of VRAM
1
2
.Related Stories
The flexibility of Dia opens up various use cases, including:
Nari Labs is developing a consumer version of Dia for casual users interested in remixing or sharing generated conversations. They also plan to release a technical report and expand language support beyond English
1
2
.While Dia offers impressive capabilities, it also raises concerns about potential misuse. The model currently lacks robust safeguards against the creation of disinformation or scam recordings. Nari Labs discourages abuse but states they are not responsible for misuse
1
.Additionally, questions arise about the data used to train Dia, as it may include copyrighted content. This issue reflects a broader debate in the AI industry about the legality and ethics of training models on copyrighted materials
1
.As Dia enters the market, it represents both the democratization of AI technology and the need for careful consideration of its implications and responsible deployment in the rapidly evolving field of synthetic speech.
Summarized by
Navi
1
Technology

2
Technology

3
Policy and Regulation
