AudioShake Launches The Refinery to Teach AI How People Actually Talk

Article Featured Image

AudioShake, an audio separation technology provider, today launched The Refinery, a system that turns raw, real-world audio into structured, training-ready data at scale.

AudioShake's The Refinery can separate people speaking simultaneously into individual speaker tracks, while also isolating dialogue, music, and background sound from finished recordings. That allows AI developers to extract usable training data from conversations that more closely resemble how people actually communicate, including interjections, overlapping speech, and background noise.

The Refinery can process an audio corpus of any size, splitting conversations into individual speaker tracks, separating dialogue from music and background sound, and isolating speech from complex real-world environments directly from recorded audio.

The system also scores its outputs for quality and confidence, allowing large corpora to be sorted into data that is usable, fixable, or unusable. The Refinery does not generate or reconstruct speech. Each separated voice and its corresponding frequencies come from the original recording, preserving real interruptions, overlap, accents, background conditions, and other characteristics of the source audio.

The Refinery can process data through AudioShake's API or be deployed on-premises, allowing organizations to process sensitive or proprietary audio without it leaving their environment. Customers retain full ownership of their data and outputs, and the system is designed to work across large audio corpora and a wide range of recording conditions.

"AI labs need models that understand the way people actually talk: interrupting each other, talking over each other, speaking in noisy rooms," said Jessica Powell, CEO and co-founder of AudioShake, in a statement. "The strange thing is that enormous amounts of that conversation already exist. The challenge has been turning it into usable training data. The Refinery lets labs recover that complexity from audio they already have rather than having to recreate the real world from scratch."

For frontier and voice-AI labs, The Refinery turns existing audio corpora into training data for applications including automatic speech recognition, diarization, speaker identification, text-to-speech, and conversational AI