4 Sources
[1]
Meta introduces new SAM AI able to isolate and edit audio
No mention of protections to stop it being used to snoop on people Want to hear just the guitar riff from a song? How about cutting out the train noise from a voice recording? Meta says its new SAM Audio model can separate and edit sounds using simple prompts, cutting down on the manual work
[2]
Meta's new open-source AI tool helps you clean up noisy recordings just by typing
It supports text, visual, and time-based prompts for precise sound separation. Cleaning up audio usually means scrubbing timelines and tweaking filters, but Meta thinks it should be as easy as describing the sound you want. The company has released a new open-source AI model called SAM Audio that
[3]
Meta Platforms transforms audio editing with prompt-based sound separation - SiliconANGLE
Meta Platforms transforms audio editing with prompt-based sound separation Meta Platforms Inc. is bringing prompt-based editing to the world of sound with a new model called SAM Audio that can segment individual sounds from complex audio recordings. The new model, available today through Meta's
[4]
Meta's New AI Model Will Let You Isolate Any Sound in an Audio File
Meta says the model can be used for noise filtering and isolating sounds Meta has released another new artificial intelligence (AI) model in the Segment Anything Model (SAM) family. On Tuesday, the Menlo Park-based tech giant released SAM Audio, a large language model (LLM) that can identify,
Share
Copy Link
Meta has launched SAM Audio, an open-source AI model that simplifies audio editing by isolating specific sounds from complex recordings using text, visual, or time-based prompts. The unified multimodal model can separate voices, instruments, and background noise without manual editing, though it raises questions about privacy safeguards and struggles with similar overlapping sounds.
Meta has released SAM Audio, a new AI model designed to isolate and edit audio through simple prompts, marking the company's expansion of its Segment Anything Model family into the audio domain
1
. The open-source AI model is now available through Meta's Segment Anything Playground and for download via the company's website, GitHub, and Hugging Face4
. Meta describes SAM Audio as "the first unified multimodal model for audio separation," capable of interpreting text prompts, visual selections in video, and time-segment markings to isolate specific sounds1
.
Source: SiliconANGLE
SAM Audio supports three distinct prompting methods that can be used individually or combined for precise control. Users can describe sounds using text prompts like "drum beat" or "background noise," click on people or objects in videos to visually identify sounds, or mark time spans where specific sounds first appear
4
. The core technology relies on the Perception Encoder Audiovisual engine, built on Meta's open-source Perception Founder model released earlier this year3
. This engine acts as the model's "ears," allowing it to comprehend described sounds, isolate them in audio files, and extract them without affecting other audio elements3
.
Source: Gadgets 360
The AI model addresses use cases across multiple industries, including music production, podcasting, film and television, and scientific research
2
. Creators can clean up noisy recordings by removing traffic sounds from podcasts, isolate vocals from band recordings, or delete unwanted barking dogs from video presentations3
. Meta has partnered with US hearing aid manufacturer Starkey to explore potential integrations and is working with 2gether-International, an accelerator for disabled startup founders, to develop accessibility solutions1
. The model operates faster than real-time with RTF ≈ 0.7, processing audio efficiently at scale from 500M to 3B parameters3
.
Source: Digital Trends
Related Stories
Questions about safety features have surfaced given the model's ability to isolate specific sounds based on user prompts, potentially creating avenues for surveillance. When asked about safeguards, Meta only stated that use of SAM Audio must comply with applicable laws and regulations, including data protection laws, without detailing built-in protections
1
. The company acknowledged several limitations: SAM Audio cannot perform complete audio separation without prompting, does not support audio-based prompts, and struggles with highly similar audio events like isolating one voice from a choir or a single instrument from an orchestra1
2
.To advance the field of audio separation, Meta introduced SAM Audio-Bench, a benchmark covering speech, music, and sound effects across text, visual, and span-prompt types
3
. The company also released SAM Audio Judge, which evaluates how natural and accurate separated audio sounds to human listeners without requiring reference tracks2
. Performance evaluations show SAM Audio achieves state-of-the-art results in modality-specific tasks, with mixed-modality prompting delivering stronger outcomes than single-modality approaches3
. The launch connects to Meta's broader AI strategy, including improving voice clarity on AI-powered glasses for noisy environments and developing conversational AI to rival ChatGPT2
.Summarized by
Navi
[1]
[3]
31 Jul 2024

20 Nov 2025•Technology

20 Oct 2024•Technology

1
Technology

2
Technology

3
Policy and Regulation
