본문으로 건너뛰기
AI 도감

음성 복제와 변환

적은 샘플로 한 사람의 음색을 재현한다

음성과 음악중급 #29
입력오디오텍스트오디오

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

이 능력이 뜻하는 것

Takes a reference clip (the target timbre) and the text to be spoken, and outputs that text in that voice. It splits into two kinds: zero-shot cloning needs only a few seconds of audio, while voice conversion keeps the content and timing of the source and merely swaps the timbre. Unlike speech synthesis the timbre comes from a sample rather than a preset.

기술적으로 구현하는 방법

The speaker timbre is encoded into a vector disentangled from content and injected as a condition into the acoustic model; voice conversion substitutes the target timbre features into the source audio and rebuilds the waveform. Zero-shot systems train on large multi-speaker data so the timbre encoder generalises to unseen speakers. For compliance they are usually paired with consent checks, watermarking and synthetic-audio markers to limit misuse.

대표 제품

3

관련 기관

대표적 용도

  • A consistent narrator voice across content
  • Keeping the original voice in multilingual dubbing
  • Accessibility and voice restoration for those who lost speech
  • Character voices in games and animation

성능을 평가하는 방법

Speaker similarity (SECS)
Cosine similarity between output and target speaker embeddings
MOS
Mean opinion score for naturalness
Equal error rate (EER)
Error rate where false accept and false reject meet, gauging impersonation risk

경계와 난점

  • With only seconds of reference audio the timbre is unstable, sounding like two speakers across sentences
  • Cross-lingual or cross-emotion transfer drifts, so timbre shifts with the language
  • It can be used to impersonate people, so consent and watermarking are required or the social-engineering and fraud risk is high

뒤에 있는 개념