FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation
Yi Yuan1, Xubo Liu2, Haohe Liu2, Xiyuan Kang1, Mark D. Plumbley3, Wenwu Wang1
1School of Computer Science and Electronic Engineering, University of Surrey, Guildford, UK
2Superintelligence Labs, Meta, USA
3Department of Informatics, King's College London, London, UK
AbstractLanguage-queried audio source separation (LASS) aims to extract target sources from audio mixtures according to natural language descriptions, offering a flexible and scalable interface for audio source separation. However, most existing LASS methods rely on discriminative, mask-based models, which estimate masks from the input mixture. These methods often over-suppress target sounds or fail to fully separate them, especially when multiple sound events strongly overlap in complex acoustic scenes. In this work, we propose FlowSep2, a text-conditioned flow-matching generative model for LASS. Instead of directly predicting a separation mask, FlowSep2 learns to generate the target source representation from Gaussian noise in a latent space, conditioned on both the mixture representation and the text query. Specifically, we employ rectified flow matching with a Diffusion Transformer backbone. We further incorporate Self-Flow, a self-supervised flow-matching paradigm, into our LASS framework. By encouraging semantically structured latent representations under the generative objective, Self-Flow improves the model’s ability to separate target sources according to text queries. Experiments on multiple LASS benchmarks show that FlowSep2 achieves state-of-the-art performance and demonstrates enhanced sound separation results in challenging scenarios with overlapping sound events. |
|---|
FlowSep2FlowSep2 generates the target source in a continuous latent space instead of predicting a time-frequency mask. A FLAN-T5 encoder embeds the text query, a Stable-Audio VAE encoder maps the mixture waveform to a latent zm, and a Diffusion Transformer (DiT) learns a rectified-flow velocity field that transports Gaussian noise z0 to the target latent z1, conditioned on [zλ, zm] and the text embedding. On top of the flow-matching loss, Self-Flow adds a self-supervised representation objective: the student sees a noisy, mixed-timestep view of the latent while an EMA teacher sees a cleaner, uniform-timestep view, and an intermediate student layer l is aligned to a deeper teacher layer k by cosine loss. At inference an ODE solver integrates from λ = 0 to λ = 1 using the student only, and the VAE decoder reconstructs the waveform.
|
|---|
Demos on AudioCaps
Each row shares one dB reference, taken from the mixture, so panel brightness is comparable across systems and an output that collapsed to near-silence correctly renders as black. CLAPA is the cosine similarity between the CLAP audio embedding of each system's output and that of the reference.
CLAPA mixture 0.76 → FlowSep2 0.84 | FlowSep 0.54 | SAM-Audio 0.40 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.34 → FlowSep2 0.81 | FlowSep 0.52 | SAM-Audio 0.01 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.44 → FlowSep2 0.90 | FlowSep 0.63 | SAM-Audio 0.19 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.80 → FlowSep2 0.90 | FlowSep 0.47 | SAM-Audio 0.65 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.60 → FlowSep2 0.75 | FlowSep 0.59 | SAM-Audio 0.52 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.80 → FlowSep2 0.92 | FlowSep 0.77 | SAM-Audio 0.53 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.73 → FlowSep2 0.91 | FlowSep 0.78 | SAM-Audio 0.78 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.64 → FlowSep2 0.90 | FlowSep 0.79 | SAM-Audio 0.79 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.66 → FlowSep2 0.89 | FlowSep 0.80 | SAM-Audio 0.63 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.76 → FlowSep2 0.90 | FlowSep 0.59 | SAM-Audio 0.81 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.57 → FlowSep2 0.80 | FlowSep 0.62 | SAM-Audio 0.75 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.77 → FlowSep2 0.86 | FlowSep 0.81 | SAM-Audio 0.82 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.86 → FlowSep2 0.93 | FlowSep 0.82 | SAM-Audio 0.90 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.73 → FlowSep2 0.82 | FlowSep 0.80 | SAM-Audio 0.72 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.50 → FlowSep2 0.68 | FlowSep 0.60 | SAM-Audio 0.67 |
![]() |
![]() |
![]() |
![]() |
![]() |
CLAPA mixture 0.85 → FlowSep2 0.90 | FlowSep 0.90 | SAM-Audio 0.39 |
![]() |
![]() |
![]() |
![]() |
![]() |
Zero-shot music separation on MUSDB18
FlowSep2 is trained only on general sound datasets (VGGSound, AudioCaps, AudioSetCaps, WavCaps) and has seen no music stems, so MUSDB18 is entirely zero-shot for it. SAM-Audio is trained on large-scale music data and is the stronger system on this benchmark overall — the examples below are chosen to show both sides: cases where FlowSep2 matches or beats it, cases where SAM-Audio collapses on the query, and cases where SAM-Audio is clearly ahead. FlowSep is not included here, as it was not evaluated on MUSDB18.
The Mountaineering Club - Mallory — vocals stem SI-SDRi FlowSep2 +10.1 dB | SAM-Audio +10.9 dB |
![]() |
![]() |
![]() |
![]() |
Lyndsey Ollard - Catching Up — vocals stem SI-SDRi FlowSep2 +5.6 dB | SAM-Audio +4.7 dB |
![]() |
![]() |
![]() |
![]() |
Detsky Sad - Walkie Talkie — drums stem SI-SDRi FlowSep2 +2.0 dB | SAM-Audio +1.1 dB |
![]() |
![]() |
![]() |
![]() |
BKS - Too Much — bass stem SI-SDRi FlowSep2 +3.8 dB | SAM-Audio -44.9 dB |
![]() |
![]() |
![]() |
![]() |
Sambasevam Shanmugam - Kaathaadi — bass stem SI-SDRi FlowSep2 +3.7 dB | SAM-Audio +9.3 dB |
![]() |
![]() |
![]() |
![]() |
Hollow Ground - Ill Fate — bass stem SI-SDRi FlowSep2 +0.4 dB | SAM-Audio -72.5 dB |
![]() |
![]() |
![]() |
![]() |
Cristina Vane - So Easy — drums stem SI-SDRi FlowSep2 +3.6 dB | SAM-Audio +13.1 dB |
![]() |
![]() |
![]() |
![]() |
Triviul feat. The Fiend - Widow — bass stem SI-SDRi FlowSep2 +5.7 dB | SAM-Audio +17.0 dB |
![]() |
![]() |
![]() |
![]() |
Objective results
| Model | FAD ↓ | CLAP Score ↑ | CLAPA Score ↑ | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AC | DE-S | VGG | ESC | AC | DE-S | DE-R | VGG | ESC | AC | DE-S | VGG | ESC | MUC | |
| Unprocessed | – | – | – | – | 31.7 | 42.7 | 41.7 | 33.6 | 38.6 | 54.4 | 66.3 | 62.7 | 66.8 | 50.2 |
| LASS-Net | 5.09 | 1.83 | 3.09 | 3.28 | 33.9 | 44.4 | 44.8 | 36.4 | 40.5 | 65.2 | 72.1 | 64.5 | 75.6 | – |
| AudioSep | 4.38 | 1.21 | 2.30 | 1.93 | 33.6 | 45.1 | 49.7 | 38.5 | 40.2 | 65.1 | 74.4 | 67.4 | 76.0 | 50.4 |
| FlowSep | 2.86 | 0.90 | 2.06 | 1.49 | 41.9 | 46.4 | 50.3 | 39.5 | 42.2 | 76.7 | 75.6 | 69.2 | 75.7 | 54.1 |
| SAM-Audio | 1.58 | 0.98 | 1.86 | 1.49 | 37.8 | 43.4 | 45.8 | 39.1 | 40.9 | 64.1 | 71.3 | 70.5 | 78.6 | 62.8 |
| FlowSep2-M | 0.88 | 0.72 | 1.32 | 1.15 | 44.9 | 47.5 | 51.0 | 42.0 | 44.5 | 80.5 | 78.5 | 72.0 | 80.5 | 58.5 |
AC = AudioCaps, DE-S = DCASE-Synth, DE-R = DCASE-Real, VGG = VGGSound, ESC = ESC-50, MUC = MUSDB18. FAD is not defined for the unprocessed mixture, since the target and the mixture share the same audio events.
Subjective and human-aligned results
| Model | SAJ ↑ | AA ↑ | REL ↑ | OVL ↑ | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AC | MUC | DE-S | AC | MUC | DE-S | DE-R | AC | MUC | DE-S | DE-R | AC | MUC | DE-S | DE-R | |
| LASS-Net | 2.57 | 2.03 | 2.68 | 4.86 | 5.27 | 5.41 | 4.89 | 3.12 | 2.87 | 2.96 | 3.59 | 2.16 | 2.45 | 2.84 | 3.88 |
| AudioSep | 2.84 | 2.97 | 2.93 | 5.18 | 5.94 | 6.03 | 5.57 | 3.66 | 3.99 | 3.24 | 3.93 | 2.69 | 3.85 | 3.53 | 4.02 |
| SAM-Audio | 2.58 | 3.76 | 2.91 | 5.56 | 6.48 | 6.14 | 5.79 | 3.54 | 4.35 | 3.24 | 3.85 | 2.69 | 4.44 | 3.75 | 4.26 |
| FlowSep | 2.73 | 3.08 | 3.02 | 5.36 | 6.07 | 5.83 | 5.28 | 4.08 | 3.71 | 3.62 | 4.11 | 3.98 | 3.84 | 3.72 | 4.26 |
| FlowSep2-M | 2.96 | 3.39 | 3.47 | 5.49 | 6.67 | 6.08 | 5.86 | 4.13 | 3.91 | 3.62 | 4.16 | 4.09 | 4.02 | 3.88 | 4.28 |
SAJ = SAM Audio Judge, AA = AudioBox Aesthetics, REL = audio-text relation, OVL = overall impression. SAM-Audio is trained on large-scale music data and is the stronger system on MUSDB (MUC); FlowSep2 is evaluated there zero-shot.
Acknowledgement
This research was partly supported by a research scholarship from the China Scholarship Council (CSC), funded by British Broadcasting Corporation Research and Development (BBC R&D), Engineering and Physical Sciences Research Council (EPSRC) Grant EP/T019751/1 'AI for Sound', and a PhD scholarship from the Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey. For the purpose of open access, the authors have applied a Creative Commons Attribution (CC BY) license to any Author Accepted Manuscript version arising.
Page updated on 22 Aug 2026
















































































































