FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search

Poster IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026) · 2026 · Barcelona, Spain

Accepted at ICASSP 2026. FlowSynth appeared in the AASP-P3: Music Generation I poster session in Barcelona, Spain.

Overview

FlowSynth addresses a central challenge in virtual instrument generation: maintaining a consistent timbre across different pitches and velocities. It combines distributional flow matching with test-time search to produce high-quality, playable instruments while preserving note-level control.

Key contributions

  • Distributional flow matching models the velocity field as a learned probability distribution, capturing uncertainty instead of producing only a deterministic estimate.
  • Confidence-weighted test-time sampling explores multiple generation trajectories and selects outputs that maximize timbre consistency.
  • Music-specific search objectives improve consistency across an instrument’s pitch range while preserving prompt alignment and audio quality.

This work was completed during an internship at Smule.

Qihui Yang, Randal Leistikow, Yongyi Zang.