YuE2 unifies symbolic planning and audio music generation in one foundation model, achieving frontier song quality competitive with Suno v5/v6. Through symbolic planning, it drafts an editable score, then performs full songs with vocals and accompaniment.
YuE2 covers over 30 distinct genres across Eastern and Western musical traditions, from Baroque counterpoint to modern Cyberpunk bass music.
Instruments: Rhodes Piano, Muted Trumpet, Upright Bass, Brushes
Instruments: Acoustic Guitar, Kick Drum, Claps, Warm Harmonies
Instruments: Guzheng, 808 Sub Bass, Neon Synths, Hybrid Trap
Instruments: Guqin, Bamboo Flute Dizi, Strings, Traditional Percussion
Instruments: Overdriven Electric Guitars, Double-kick Drums, Slap Bass
Instruments: 9th/11th Chords, Fender Rhodes, Laidback Pocket Drums, Velvet Vocals
Instruments: Vinyl Crackle, Detuned Tape Piano, Jazzy Boom Bap
Instruments: Full Symphony Orchestra, Taiko Drums, French Horns, Choral Hymns
Traditional music generation is a black box. YuE2 transforms music creation into an open, iterative dialogue between human intention, symbolic theory, and audio synthesis.
Your creative request ("make it jazzier", "more dramatic strings") is translated into concrete ABC notation changes—substituting triads for 7ths/9ths, adjusting accidentals, and restructuring measures.
The agent guarantees syntactic validity, bar-line timing rules, and tonal consistency across [verse], [chorus], and [bridge] boundaries before scheduling expensive audio diffusion.
YuE2 condition tokens synchronize acoustic diffusion with the verified ABC score, ensuring the rendered vocals and instruments accurately reflect the modified sheet.
The two-stage decoupled architecture combines discrete autoregressive symbolic planning with continuous acoustic flow matching.
Autoregressive transformer trained on millions of paired lyrics, ABC scores, and structural metadata. Outputs structured musical blueprints with chords, melody, and rhythm.
Flow matching diffusion model conditioned on symbolic plans and text prompts, decoded through a 48kHz neural audio VAE into full stereo arrangements.
Specialized 6-task transcription foundation model that listens to any recorded song and automatically transcribes melody, chords, beat, and structure into ABC notation.
Self-supervised music representation model holding the state-of-the-art across all 14 tasks of the international MARBLE benchmark.
Evaluated across 192 wild, diverse user prompts spanning genres, acoustic conditions, and languages. YuE2 sets the new state of the art.
| Rank | Model / System Setting | Prompts | SongBench Score ↑ | Coherence | Vocal Naturalness |
|---|---|---|---|---|---|
| 1 | YuE2 (best-of-8)This Work Symbolic Plan + Flow Matching | 192 | 6.9632 | 7.12 | 6.94 |
| 2 | YuE2 (single pass)This Work Direct Symbolic Planning | 192 | 6.8814 | 6.98 | 6.86 |
| 3 | Suno v5 Commercial Proprietary | 192 | 6.8721 | 6.91 | 6.84 |
| 4 | Suno v6 Commercial Proprietary | 192 | 6.5562 | 6.62 | 6.51 |
| 5 | Suno v6 Wild Unconstrained Wild Prompts | 192 | 6.4195 | 6.45 | 6.39 |
| 6 | Udio v1.5 Commercial Proprietary | 192 | 6.3120 | 6.28 | 6.35 |
Tested on WildSongBench standardized prompts with blind human and automated evaluation. YuE2 achieves frontier quality competitive with and exceeding Suno v5/v6.
Answers to common questions about symbolic planning, model weights, commercial usage, and technical capabilities.
If you use YuE2, SheetSage2, or MERT2 in your research or commercial projects, please cite the corresponding papers:
@article{yue2_2026,
title={YuE2: Frontier Full-Song Generation with Symbolic Planning},
author={Multimodal Art Projection (M·A·P) and HKUST and Tokenwave.AI and NYU and Stanford and MBZUAI},
journal={arXiv preprint},
year={2026}
}