A suite of new AI models is pushing the boundaries of music generation and transcription. YuE2, for instance, achieved the highest average score on the WildSongBench benchmark, outperforming other systems in musicality, prompt control, and lyric accuracy. This benchmark evaluates models across 192 prompts and 15 different system settings.

In addition, the MERT2 models have set new state-of-the-art results on the MARBLE benchmark, excelling in 14 out of 15 metrics related to music tagging, key detection, genre classification, and emotion recognition. These models utilize large context windows of 30 seconds and full-song durations to improve representation quality.

SheetSage2, a single model capable of handling six transcription tasks—including beat, downbeat, key, chord, structure, and melody transcription—achieved state-of-the-art scores on 10 of 13 evaluated metrics. This marks a significant step forward in unified music transcription models.

The training data for these models primarily consists of CC0 music and synthetic datasets, with a substantial portion licensed from Tokenwave.AI. The developers emphasize ethical and responsible data use.

These advancements matter because they enhance the ability of AI to understand and generate music with greater nuance and accuracy. This has implications for music production, analysis, and interactive applications, potentially transforming how music is created and experienced.