In this paper, we propose the design and development of an AI-driven language learning application that leverages YouTube videos to provide immersive, personalized practice environments. The system empowers learners to choose any speaker from a YouTube video, then automatically identifies, isolates, and archives segments in which that individual appears. Once captured, these audio segments are converted into written transcripts, which serve as the foundation for interactive exercises including targeted listening comprehension, oral practice, and pronunciation refinement. A distinctive capability of the platform involves pausing playback precisely when the selected speaker talks, allowing learners to step into that character's role and participate in conversational exchanges with other figures in the video, thus strengthening their communicative abilities. Additionally, a specialized speaker-matching analysis module has been integrated to measure the acoustic similarity between learner output and the target speaker's voice. Through the extraction and comparison of acoustic characteristics—including vocal quality, articulation, tone, and melodic patterns—the system generates a numerical similarity rating on a 100-point scale. This comprehensive approach not only enhances multimodal interaction within second language learning environments but also establishes innovative pathways for tailored, immersive education grounded in genuine multimedia materials.
목차
Abstract 1. Introduction 2. Related Studies 3. Proposed AI-driven Target Speaker Dialogue for Korean Learning 4. Experimental Results 4.1 Accuracy of specific person presence detection 4.2 Testing of the increase in similarity score of voice 4.3 Testing of the Word Error Rate (WER) 4.4 Testing of the Phoneme Accuracy 5. Conclusion Acknowledgement References