Генерация музыки и звуков
Создать музыку или звуковой эффект по описанию
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes a text description, optionally with lyrics or a style reference, and outputs audio — a full song, an ambient bed or a sound effect. It must handle melody, harmony, arrangement and timbre at once, adding a layer of musical structure on top of speech synthesis. Unlike text-to-speech the output is not language but rhythmically and tonally organised audio.
Как это устроено
The mainstream uses an audio latent representation with diffusion or autoregressive generation: audio is first compressed into discrete or continuous tokens, generated under conditioning, and decoded back to a waveform. Structural tags during training teach the model the order of intro, verse and chorus. Lyrics and vocals are produced by separate alignment and synthesis modules and then mixed with the accompaniment into a finished track.
Примеры продуктов
3Suno
2023Создаёт полноценные песни с вокалом по одному описанию
Lyria
2023Создаёт инструментальную и вокальную музыку по текстовому описанию
ElevenLabs
2022Многоязычный сервис синтеза речи с естественными и клонируемыми голосами
Связанные организации
Типичное применение
- Background music for short video and podcasts
- Sound effects for games and apps
- Scores for ads and promotional films
- Creative demos and arrangement ideas
Как её оценивают
- FAD
- Distribution distance between generated and real music; lower is better
- MOS
- Mean opinion score for listening quality
- Human preference
- Pairwise judgement of melodic and arrangement appeal
Границы и трудности
- Long-form structure collapses: chorus returns and arrangement layers fail to hold together
- Lyrics and melody misalign and enunciation blurs, with mispronounced or swallowed syllables
- Imitating the style of a living artist under copyright carries risk and needs checking before commercial use
Концепции в основе
Диффузионные модели
Научитесь тысяче мелких шагов удаления шума — и соберёте изображение из чистого шума
Обзор порождающих моделей
Дискриминативные модели отвечают «что это», порождающие — «как это должно выглядеть»
Автоэнкодеры и вариационные автоэнкодеры
Сжать информацию через «бутылочное горлышко», а затем восстановить её