Génération de musique et de sons
Générer une musique ou un effet sonore à partir d’une description
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes a text description, optionally with lyrics or a style reference, and outputs audio — a full song, an ambient bed or a sound effect. It must handle melody, harmony, arrangement and timbre at once, adding a layer of musical structure on top of speech synthesis. Unlike text-to-speech the output is not language but rhythmically and tonally organised audio.
Comment c'est fait
The mainstream uses an audio latent representation with diffusion or autoregressive generation: audio is first compressed into discrete or continuous tokens, generated under conditioning, and decoded back to a waveform. Structural tags during training teach the model the order of intro, verse and chorus. Lyrics and vocals are produced by separate alignment and synthesis modules and then mixed with the accompaniment into a finished track.
Produits représentatifs
3Suno
2023Génère des chansons complètes avec voix à partir d’une seule description
Lyria
2023Génère de la musique instrumentale et chantée à partir de consignes textuelles
ElevenLabs
2022Un service multilingue de synthèse vocale aux voix naturelles et clonables
Organisations concernées
Usages typiques
- Background music for short video and podcasts
- Sound effects for games and apps
- Scores for ads and promotional films
- Creative demos and arrangement ideas
Comment on l'évalue
- FAD
- Distribution distance between generated and real music; lower is better
- MOS
- Mean opinion score for listening quality
- Human preference
- Pairwise judgement of melodic and arrangement appeal
Limites et points difficiles
- Long-form structure collapses: chorus returns and arrangement layers fail to hold together
- Lyrics and melody misalign and enunciation blurs, with mispronounced or swallowed syllables
- Imitating the style of a living artist under copyright carries risk and needs checking before commercial use
Concepts sous-jacents
Modèles de diffusion
Apprenez mille petits pas de débruitage et vous créerez une image à partir de bruit pur
Vue d’ensemble des modèles génératifs
Les modèles discriminatifs répondent « qu’est-ce que c’est », les génératifs « à quoi cela devrait ressembler »
Autoencodeurs et VAE
Comprimer l’information dans un goulot, puis la laisser repousser