本文へスキップ
AI図鑑

声のクローンと変換

わずかなサンプルから声質を再現する

音声と音楽中級 #29
入力音声テキスト音声

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

この能力とは何か

Takes a reference clip (the target timbre) and the text to be spoken, and outputs that text in that voice. It splits into two kinds: zero-shot cloning needs only a few seconds of audio, while voice conversion keeps the content and timing of the source and merely swaps the timbre. Unlike speech synthesis the timbre comes from a sample rather than a preset.

技術的にどう実現するか

The speaker timbre is encoded into a vector disentangled from content and injected as a condition into the acoustic model; voice conversion substitutes the target timbre features into the source audio and rebuilds the waveform. Zero-shot systems train on large multi-speaker data so the timbre encoder generalises to unseen speakers. For compliance they are usually paired with consent checks, watermarking and synthetic-audio markers to limit misuse.

代表的な製品

3

関連する組織

代表的な用途

  • A consistent narrator voice across content
  • Keeping the original voice in multilingual dubbing
  • Accessibility and voice restoration for those who lost speech
  • Character voices in games and animation

どう評価するか

Speaker similarity (SECS)
Cosine similarity between output and target speaker embeddings
MOS
Mean opinion score for naturalness
Equal error rate (EER)
Error rate where false accept and false reject meet, gauging impersonation risk

限界と難しさ

  • With only seconds of reference audio the timbre is unstable, sounding like two speakers across sentences
  • Cross-lingual or cross-emotion transfer drifts, so timbre shifts with the language
  • It can be used to impersonate people, so consent and watermarking are required or the social-engineering and fraud risk is high

背景にある概念