We present a usability-centered end-to-end voice cloning pipeline that enables AI non-experts to apply voice cloning in practical settings. As generative AI and zero-shot voice cloning rapidly expand across diverse fields, non-specialists still face difficulties in selecting appropriate models and lack reproducible procedures for real-world use. To address this issue, this paper selected F5-TTS, XTTS-V2, and Chatterbox Turbo-TTS, and designed a standardized workflow consisting of reference voice collection, preprocessing, synthesis, automatic evaluation, reporting, and deployment. Sample sentences were recorded in the researcher's own voice, and the generated outputs of each model were automatically evaluated using similarity, aesthetics, and accuracy metrics by comparing original voice files, synthesized voice files, and text scripts. The results showed that F5-TTS was effective for long-form narration and story-oriented generation, although its performance varied by condition. Chatterbox Turbo-TTS demonstrated strong speaker similarity, stable generation quality, and responsible-use advantages through watermarking-based safeguards. XTTS-V2 achieved the best performance in cross-lingual applicability and accuracy, recording a WER of 0.00% in the English pronunciation comparison scenario. This paper presents practical guidelines for applying voice cloning to real-world tasks using WebUI tools and provides a foundation for future research on educational curriculum development and model selection strategies for diverse storytelling contexts.
한국어
본 논문은 AI 비전공자도 음성 복제 기술을 실무에 적용할 수 있도록 사용성 중심의 종단간(end-to-end) 보이스 클로닝 파이프라인을 제안한다. 생성형 AI와 Zero-shot 보이스 클로닝 기술이 다양한 분야로 빠르게 확산되고 있지만, 비전공자는 적절한 모델을 선택하는 데 어려움을 겪고 있으며, 실제 활용을 위한 재현 가능한 절차도 부족한 상황이다. 이러한 문제를 해결하기 위해 본 연구는 F5-TTS, XTTS-V2, Chatterbox Turbo-TTS를 선정하고, 레퍼런스 음성 수집, 전처리, 합성, 자동 평가, 리포트 작성, 배포로 이어지는 표준화된 작업 흐름을 설계하였다. 예시 문장은 연구자의 실제 목소리로 녹음하였으며, 각 모델이 생성한 결과물은 원본 음성 파일, 생성 음성 파일, 텍스트 스크립트를 비교하여 유사성, 심미성, 정확성 지표를 통해 자동 평가하였다. 실험 결과, F5-TTS는 장문 낭독과 이야기형 생성에 효과적이었으나 조건에 따라 성능 편차가 나타났다. Chatterbox Turbo-TTS는 높은 화자 유사도, 안정적인 생성 품질이 강점이었다. XTTS-V2는 Cross-lingual 활용성과 정확성에서 가장 우수한 성능을 보였으며, 영어 발음 비교 시나리오에서 WER 0.00%를 기록하였다. 본 논문은 WebUI 도구를 활용해 음성 복제를 실무에 적용할 수 있는 실질적인 기준을 제시하며, 향후 교육 커리큘럼 개발과 다양한 스토리텔링 맥락에 적합한 모델 선택 전략 연구를 위한 기반을 제공한다.
국제문화기술진흥원 [The International Promotion Agency of Culture Technology]
설립연도
2009
분야
공학>공학일반
소개
본 진흥원은 문화기술(Culture Technology) 관련 산·학·연·관으로 구성된 비영리 단체이다. 문화기술(CT)은 정보통신기술(ICT), 문화적 사고 기반의 예술, 인문학, 디자인, 사회과학기술이 접목된 신융합기술(New Convergence Technology, NCT)로 정의한다. 인간의 삶의 질을 향상시키고, 진보된 방향으로 변화시키고, 문화기술 관련 분야의 학술 및 기술의 발전과 진흥에 공헌하기 위하여, 제3조의 필요한 사업을 행함을 그 목적으로 한다.
간행물
간행물명
The Journal of the Convergence on Culture Technology (JCCT) [문화기술의 융합]
간기
격월간
pISSN
2384-0358
eISSN
2384-0366
수록기간
2015~2026
등재여부
KCI 등재
십진분류
KDC 600DDC 700
이 권호 내 다른 논문 / The Journal of the Convergence on Culture Technology (JCCT) Vol.12 No.3