Earticle

현재 위치 Home

기술 융합(TC)

마이크로프로세서와 AI 모델을 활용한 몰입형 음성 인식 게임 컨트롤러 설계
Design of an Immersive Voice-Controlled Gamepad Using Microprocessor and AI Model

첫 페이지 보기
  • 발행기관
    국제문화기술진흥원 바로가기
  • 간행물
    The Journal of the Convergence on Culture Technology (JCCT) KCI 등재 바로가기
  • 통권
    Vol.11 No.6 (2025.11)바로가기
  • 페이지
    pp.973-986
  • 저자
    신성현, 이강희
  • 언어
    한국어(KOR)
  • URL
    https://www.earticle.net/Article/A486609

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

원문정보

초록

영어
Conventional game control methods—centered on physical input devices such as keyboards, mice, and controllers—have enabled intuitive and precise gameplay, yet the potential of auditory-based interaction has remained relatively underexplored. This study aims to complement existing input paradigms by integrating voice recognition into gameplay, thereby exploring a new form of immersive interaction that connects users’ speech directly to in-game control. To this end, a real-time voice-command-based game interface system was designed and implemented by integrating OpenAI’s Whisper model with an Arduino Leonardo board. The system records the user’s voice in a Python environment, converts it into text through Whisper, compares it with predefined commands, and generates corresponding HID inputs on the Arduino board to reflect in-game actions. The performance was evaluated by measuring recognition accuracy and latency across different Whisper model sizes and by testing in-game responsiveness in Overwatch and StarCraft II environments. The results showed that the proposed system achieved high recognition accuracy and low latency, operating stably in both controller-based and keyboard-only game settings. This demonstrates that the proposed system possesses versatility and scalability, making it applicable across various game genres and platform environments. By introducing a voice-driven, auditory-centered form of interaction, this study presents a multimodal interface design that complements traditional visual and haptic input systems.
한국어
기존의 전통적 게임 조작 방식은 키보드, 마우스, 컨트롤러 등 물리적 입력 장치 중심으로 발전해왔으며, 이는 직관적이고 정밀한 조작을 가능하게 하였으나 청각적 상호작용 요소의 활용은 상대적으로 제한적이었다. 이에 본 연구는 음성 인식을 활용하여 사용자의 발화를 직접 게임 조작에 연결함으로써, 기존 입력 방식의 한계를 보완하고 새로운 몰입형 상호작용 방식을 탐색하고자 한다. 이를 위해 OpenAI의 Whisper 음성 인식 모델과 아두이노 레오나르도 보드를 결합한 실시간 음성 명령 기반 게임 인터페이스 시스템을 설계·구현하였다. 시스템은 사용자의 음성을 Python 환경에서 실시간으로 녹음하고 Whisper 모델을 통해 텍스트로 변환한 후, 사전 정의된 명령어와 비교하여 아두이노 보드에서 해당 명령에 따른 HID 입력을 생성함으로써 게임 내 조작으로 반영된다. Whisper 모델의 크기별 인식 정확도와 처리 시간을 측정하고, Overwatch와 StarCraft II 환경에서 실제 조작 반응성을 평가하였다. 제안된 시스템은 높은 인식 정확도와 낮은 지연 시간을 보였으며, 컨트롤러 기반 및 키보드 전용 게임 환경 모두에서 안정적으로 작동하였다. 이로써 제안된 시스템은 다양한 게임 장르 및 플랫폼 환경에서 효과적으로 활용될 수 있는 범용성과 확장성을 지님을 확인하였다. 본 연구는 음성 인식을 활용한 청각 중심의 몰입형 상호작용 경로를 제시함으로써, 기존 시각·촉각 중심의 입력 방식을 보완하는 다중감각적 인터페이스 설계를 제안한다.

목차

요약
Abstract
Ⅰ. 서론
Ⅱ. 관련 연구
1. 몰입형 인터페이스 연구
2. 음성 인식 기반 인터페이스 연구
Ⅲ. 시스템 설계 및 구현
1. 전체 시스템 구성 개요
2. 하드웨어 구성
3. 소프트웨어 구조
Ⅳ. 실험 및 결과
1. Whisper 음성 인식 모델 성능 비교 실험
2. 게임 환경에서의 컨트롤러 작동 실험
Ⅴ. 결론
References

저자

  • 신성현 [ Sung-Hyun Shin | 준회원, 숭실대학교 글로벌미디어 학부 학사 과정 ] 제1저자
  • 이강희 [ Kang-Hee Lee | 정회원, 숭실대학교 글로벌미디어 학부 교수 ] 교신저자

참고문헌

자료제공 : 네이버학술정보

간행물 정보

발행기관

  • 발행기관명
    국제문화기술진흥원 [The International Promotion Agency of Culture Technology]
  • 설립연도
    2009
  • 분야
    공학>공학일반
  • 소개
    본 진흥원은 문화기술(Culture Technology) 관련 산·학·연·관으로 구성된 비영리 단체이다. 문화기술(CT)은 정보통신기술(ICT), 문화적 사고 기반의 예술, 인문학, 디자인, 사회과학기술이 접목된 신융합기술(New Convergence Technology, NCT)로 정의한다. 인간의 삶의 질을 향상시키고, 진보된 방향으로 변화시키고, 문화기술 관련 분야의 학술 및 기술의 발전과 진흥에 공헌하기 위하여, 제3조의 필요한 사업을 행함을 그 목적으로 한다.

간행물

  • 간행물명
    The Journal of the Convergence on Culture Technology (JCCT) [문화기술의 융합]
  • 간기
    격월간
  • pISSN
    2384-0358
  • eISSN
    2384-0366
  • 수록기간
    2015~2026
  • 등재여부
    KCI 등재
  • 십진분류
    KDC 600 DDC 700

이 권호 내 다른 논문 / The Journal of the Convergence on Culture Technology (JCCT) Vol.11 No.6

    피인용수 : 0(자료제공 : 네이버학술정보)

    함께 이용한 논문 이 논문을 다운로드한 분들이 이용한 다른 논문입니다.

      페이지 저장