According to statistics from the National Forensic Service (NFS), the number of telephone financial fraud cases has more than doubled since onset of COVID-19 compared to the previous year. Telephone financial fraud, including false emergency calls to numbers 112 or 119, involves using telecommunication networks to threaten victims or extort money and valuables. The amount lost to voice phishing, in particular, has been steadily increasing annually posing a significant social problem. In such cases of telephone financial fraud, voice recordings serve as a crucial piece of evidence. This paper researches the forensic application of the latest deep learning-based speaker recognition technology for analyzing an individual from voice evidence. It explains the operational principles and usage of VOCALISE, a commercial speaker recognition engine developed by Oxford Wave Research. Notably, VOCALISE, widely used in forensic analysis organizations worldwide, is adept at creating sample groups from speaker datasets, measuring their statistical characteristics, and analyzing specific voice signals. To evaluate performance on korean speech data, we used 736 voice files collected from 159 individuals involved in actual voice crime cases to calculate the Equal Error Rate (EER) and Cost of Log-Likelihood Ratio (CLLR). Based on these real audio recordings when actual voice evidence is submitted by investigative agencies the paper explains how to use VOCALISE for a globally accepted likelihood ratio-based verbal expression method.
목차
Abstract Ⅰ. 서론 Ⅱ. VOCALISE의 x-vector 뉴럴 네트워크 구조 1. 훈련(Training) 과정 2. 추론(Inference) 과정 Ⅲ. VOCALISE의 x-vector 기반 화자인식 과정 1. 전처리 2. 음성 파일에 대한 특징벡터 추출 3. 추출된 음성 특징벡터 간의 스코어링 Ⅳ. VOCALISE를 활용한 법과학 화자 인식 절차 수립 1. 평가 데이터 수집 2. 성능 평가 (i-vector vs. x-vector) 3. 화자 인식 결과 해석 Ⅴ. 고찰 및 향후 계획 Ⅵ. 사사 Ⅶ. 참고문헌
법과학 분야는 사회정의 구현에 있어 크나큰 가치가 있음에도 불구하고 우리나라에서는 이 분야에 대한 인식이 미흡하여 선진 외국에 비해 침체되어 있는 실정이다. 이에 우리나라에서도 법과학 분야와 관련 있는 학계, 연구기관, 수사기관 등 유관 단체들로 구성된 한국 법과학회를 창립하여 이 분야를 활성화 시켜 과학수사를 한층 더 발전시키기 위함을 목적으로 한다.