Generación de música y efectos
Generar música o un efecto sonoro a partir de una descripción
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes a text description, optionally with lyrics or a style reference, and outputs audio — a full song, an ambient bed or a sound effect. It must handle melody, harmony, arrangement and timbre at once, adding a layer of musical structure on top of speech synthesis. Unlike text-to-speech the output is not language but rhythmically and tonally organised audio.
Cómo se consigue técnicamente
The mainstream uses an audio latent representation with diffusion or autoregressive generation: audio is first compressed into discrete or continuous tokens, generated under conditioning, and decoded back to a waveform. Structural tags during training teach the model the order of intro, verse and chorus. Lyrics and vocals are produced by separate alignment and synthesis modules and then mixed with the accompaniment into a finished track.
Productos representativos
3Suno
2023Genera canciones completas con voz a partir de una sola descripción
Lyria
2023Genera música instrumental y vocal a partir de indicaciones de texto
ElevenLabs
2022Un servicio multilingüe de texto a voz con voces naturales y clonables
Organizaciones relacionadas
Usos típicos
- Background music for short video and podcasts
- Sound effects for games and apps
- Scores for ads and promotional films
- Creative demos and arrangement ideas
Cómo se evalúa
- FAD
- Distribution distance between generated and real music; lower is better
- MOS
- Mean opinion score for listening quality
- Human preference
- Pairwise judgement of melodic and arrangement appeal
Límites y dificultades
- Long-form structure collapses: chorus returns and arrangement layers fail to hold together
- Lyrics and melody misalign and enunciation blurs, with mispronounced or swallowed syllables
- Imitating the style of a living artist under copyright carries risk and needs checking before commercial use
Conceptos detrás
Modelos de difusión
Aprende mil pequeños pasos de eliminación de ruido y podrás crear una imagen desde ruido puro
Panorama de los modelos generativos
Los modelos discriminativos responden "qué es esto"; los generativos, "cómo debería verse"
Autoencoders y VAE
Comprime la información por un cuello de botella y deja que vuelva a crecer