A suite of AI models has recently achieved state-of-the-art results across multiple music-related tasks, including transcription, genre classification, and song generation. The YuE2 model, evaluated on the WildSongBench dataset with 192 prompts, achieved the highest mean score of 6.9632, outperforming competitors such as Suno v5. This benchmark assesses musicality, prompt control, and lyric accuracy by selecting the best output from multiple generations.

In addition, the MERT2 models, which utilize full-context audio representations, have set new standards on the MARBLE benchmark. They lead in 14 out of 15 metrics covering tagging, key detection, genre classification, and emotion recognition. MERT2-30s uses a 30-second training context, while MERT2-FS employs a longer 300-second context, both featuring 632 million parameters.

SheetSage2, a single model handling six transcription tasks—including beat, downbeat, key, chord, structure, and melody recognition—has also reached state-of-the-art performance on 10 of 13 evaluated metrics. This model surpasses previous systems such as SheetSage1 and Madmom, demonstrating high accuracy in vocal melody and chord recognition.

These models are primarily trained on publicly available CC0 music and synthetic datasets, with Tokenwave.AI supplying a significant portion of the synthetic data under license. The developers emphasize ethical and responsible data use.

The advancements in these AI systems provide musicians, producers, and researchers with powerful tools for exploring and editing music. By enabling detailed symbolic analysis and high-quality song generation, these technologies could influence music creation workflows and music information retrieval research.