Back to News
ProductAI Understanding briefing

AudioShake launches The Refinery for AI training data

AudioShake has launched The Refinery, a system that converts existing mixed audio recordings into structured, speaker-separated tracks for AI training, reporting a 9.17% word error rate on the LibriCSS benchmark.

4 min readRead the linked source
Source-provided image accompanying AudioShake launches The Refinery for AI training data
Source referenceSource recorded
Publisher
unite.ai
Source type
Linked source — primary-source status has not been established.
ContextUnderstand this in 60 seconds

Key terms

API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Synthetic Data
Artificially generated data used to augment, simulate, or protect sensitive training data.
Benchmark
A standardized test or dataset used to measure and compare model performance.

What happened

AudioShake launched The Refinery, a product designed to process existing audio recordings into structured datasets for AI training. The system separates overlapping conversations into individual speaker tracks while distinguishing dialogue from background noise, without generating new speech. According to the company, the tool has processed over 100 million minutes of audio and is available via API or on-premises deployment.

AudioShake announced the launch of The Refinery, a system that converts existing audio recordings into structured, speaker-separated data for AI training. Unlike generation, The Refinery processes real recordings to extract individual voice tracks and separate them from music or background noise. The company states that the system does not generate or reconstruct speech but rather exposes the components already present in the original recording.

The product distinguishes between speaker diarization, which identifies when different speakers are active, and separation, which produces individual audio signals. It also differentiates between the confidence of speaker assignment and the quality of the separation itself. This allows developers to inspect specific aspects of the data, such as speaker identity or track clarity, independently. The system supports various sample rates and capture conditions and is described as acoustic rather than dependent on a language model.

AudioShake reports that The Refinery has processed more than 100 million minutes of audio and has been deployed privately with frontier AI labs over the past year. Named customers include Luel and Rime, alongside unnamed labs and data marketplaces. The company also lists media clients such as ESPN, Universal Music Group, and Warner Bros. Studios, indicating its background in audio separation for production tasks.

In terms of performance, AudioShake’s Multi-Speaker 2.0 technical evaluation reports a 9.17% concatenated minimum-permutation word error rate on the LibriCSS . This is compared to 37.75% for a tested MERL TF-Locoformer checkpoint, with both evaluated through the same Whisper large-v3 transcription pipeline. The company notes that this is an off-the-shelf comparison and that accuracy becomes harder to maintain as sustained overlap and speaker count increase.

Source details: unite.ai ↗

Why it matters

The launch addresses a specific bottleneck in voice AI development: the difficulty of obtaining clean, labeled data from natural, overlapping conversations. By automating the separation of speakers and background sounds from existing recordings, The Refinery allows developers to utilize real-world conversational data that retains the timing and complexity of human interaction. This is significant because synthetic or staged data often lacks the nuance of interruptions and background noise that real users encounter. The reported performance metrics suggest a substantial improvement over off-the-shelf baselines, potentially reducing the manual labor required to prepare high-quality training corpora for speech models.

Voice AI models require training data that reflects the messiness of real human conversation, including interruptions, background noise, and overlapping speech. Staged or synthetic recordings often fail to capture these nuances, leading to models that struggle in real-world environments. The Refinery aims to bridge this gap by making existing, rights-cleared recordings usable for training by structuring them into aligned speaker tracks.

The operational efficiency of the system is a key factor. By scoring outputs for quality and confidence, The Refinery allows organizations to sort large collections into usable, fixable, or unsuitable material. This directs human review toward questionable segments, such as crowded conversations with quiet speakers, rather than requiring manual inspection of every minute of audio. This scalability is critical for organizations with substantial audio archives.

The distinction between separation quality and downstream model performance is important. While a cleaner training corpus is valuable, the launch does not establish how much a specific conversational model will improve after learning from it. Developers must still determine how the tracks, transcripts, and metadata fit their specific learning objectives. The reported metrics support testing the technology on representative audio but do not guarantee universal performance improvements.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Model Parameter Size:8B Parameters
VRAM Required5.5 GBGPU memory footprint
Target HardwareMacBook / Single GPUDeployment tier
Privacy100% Air-GappedLocal device capability
Core takeaway: Small, quantized models (3B–8B) now run directly inside smartphones and laptops with complete data privacy, while mammoth 400B+ models remain the domain of datacenter clusters.
Interactive Concept Check+10 Points
AI Training Quiz

Which training-log entry describes one completed epoch?

What to watch next

Independent verification of the reported word error rates on diverse, non-standard audio datasets. The extent to which major AI labs integrate this specific pipeline into their training workflows. Whether the on-premises deployment option leads to broader adoption in industries with strict data privacy requirements, such as healthcare or finance.

Independent benchmarks are needed to verify the reported 9.17% word error rate on datasets that differ from LibriCSS, particularly those with higher levels of sustained overlap or more than two speakers. The company’s own evaluation notes that accuracy drops as complexity increases, so real-world performance may vary significantly.

Adoption by major AI labs will indicate the practical utility of the tool. While AudioShake mentions private deployments with frontier labs, specific details on how The Refinery is integrated into their training pipelines are not public. Monitoring for case studies or technical papers from these labs could provide deeper insight into its impact.

The availability of on-premises deployment may drive adoption in sectors with strict data privacy regulations. If organizations in healthcare, finance, or government begin using The Refinery to process sensitive audio data, it could signal a broader shift in how proprietary audio archives are utilized for AI development.

Related guides & quizzes

Found this useful?