년 - 년
Speech and emotion recognition using a multi-output model
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 8th International Conference on Next Generation Computing 2022 2022.10 pp.218-221
Voice language, the primary way of human communication, delivers not only verbal information but also emotional information through various characteristics such as voice intonation, height, and surrounding environment. Currently, many studies focus on grasping emotion and speech recognition on voice for human-computer interaction and are developing deep neural network models by extracting various frequency characteristics of speech. Representatives of these speech-based deep learning algorithms include speech recognition, namely speech-to-text or automatic speech recognition, and speech emotion recognition. The development of these two algorithms has been developed for a long time, but multi-output algorithms that process them in parallel at the same time are rare. This paper introduces a multi-output model that recognizes speech and emotion in one voice, thinking that simultaneously understanding language and emotion, which are the most critical information in a human voice, will significantly help human-computer interaction. This model confirmed that there was no significant difference between the training of the language and emotion recognition models separately, with a word error rate of 6.59% in the speech recognition section and an accuracy of 79.67% on average in the emotion recognition section.
기계 학습을 활용한 음성 데이터 자동 라벨링 시스템 설계
한국혁신산업학회 혁신산업기술논문지 제3권 제3호 2025.09 pp.119-125
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
음성 데이터 라벨링은 음성 인식 및 자연어 처리 기술의 성능을 결정하는 중요한 과정이다. 그러나 전통적인 수작업 라벨링 방식은 대규모 데이터 처리에 있어 많은 시간과 비용이 소요되며, 라벨링 오류의 가능성이 높다. 본 논문에서 는 이러한 문제를 해결하기 위해 효율적인 음성 데이터 라벨링 자동화 시스템을 설계하였다. 제안된 시스템은 음성 인식 기술과 기계 학습 알고리즘을 통합하여 음성 데이터를 자동으로 분석하고 라벨링하도록 설계되었다. 이 시스템은 다양한 음성 데이터 유형을 처리할 수 있으며, 사용자 인터페이스를 통해 라벨링 작업의 직관성과 효율성을 향상시킨다. 본 설계는 음성 데이터 처리의 효율성을 높이고, 다양한 응용 분야에서 음성 인식 기술의 활용을 촉진할 수 있는 잠재력을 가진다.
Speech data labeling is a crucial process that determines the performance of speech recognition and natural language processing technologies. However, traditional manual labeling methods are time-consuming, costly, and prone to labeling errors, especially when dealing with large-scale data. This paper presents the design of an efficient automated speech data labeling system to address these challenges. The proposed system integrates the latest speech recognition technologies and machine learning algorithms to automatically analyze and label speech data. It is designed to handle various types of speech data and enhances the intuitiveness and efficiency of the labeling process through a user-friendly interface. This design improves the efficiency of speech data processing and has the potential to facilitate the application of speech recognition technologies across various domains.
잡음에 강인한 음성인식을 위한 Generalized Gamma 분포기반과 Spectral Gain Floor를 결합한 음성향상기법 KCI 등재
한국ITS학회 한국ITS학회논문지 제8권 제3호 통권23호 2009.06 pp.64-70
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 잡음에 강인한 음성인식 성능을 획득하기 위해 generalized Gamma 분포기반의 음성향상 기법을 제안한다. 우수한 음성향상을 위해서 제안된 방식에서는 generalized Gamma분포와 spectral gain floor를 이용한 음성추적 기법에 스펙트럼 최소잡음성분에 의한 희귀적인 평균 스펙트럼 값으로부터 유도되는 잡음추정을 결합하여 음질을 향상시켜 음성인식에 적용하였다. Spectral component, spectral amplitude 그리고 log spectral amplitude에 기반하여 제안된 음성향상 기법을 잡음환경에서의 음성인식에 적용하여 그 성능을 측정하였다.
This paper presents a speech enhancement technique based on generalized Gamma distribution in order to obtain robust speech recognition performance. For robust speech enhancement, the noise estimation based on a spectral noise floor controled recursive averaging spectral values is applied to speech estimation under the generalized Gamma distribution and spectral gain floor. The proposed speech enhancement technique is based on spectral component, spectral amplitude, and log spectral amplitude. The performance of three different methods is measured by recognition accuracy of automatic speech recognition (ASR).
텔레메틱스 단말용 음성 인식을 위한 음성향상 알고리듬 및 칩 구현 KCI 등재후보
한국ITS학회 한국ITS학회논문지 제7권 제5호 통권19호 2008.10 pp.90-96
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 텔레메틱스 단말용 음성인식을 위한 음성향상 단일 칩 알고리듬을 제시한다. 제안된 방법은 잡음제거와 에코제거의 두 단계로 구성되어 있으며, 첫 단계로 크로스 스펙트럼 추정에 기반한 적응필터를 통해 에코를 제거하고, 두번째 단계로 Generalized Gamma분포기반의 LSA 음성추정 방식 추정을 통해 외부 배경잡음을 제거하여 음성의 음질을 향상시킨다. 적은 계산량이 요구되는 제안된 알고리즘을 토대로 구현된 단일 칩의 성능은 다양한 잡음환경에서 신호 대잡음비율과 음성인식 평가에서 기존의 방법보다 향상된 결과를 나타내었다.
This paper presents an algorithm of a single chip acoustic speech enhancement for telematics device. The algorithm consists of two stages, i.e. noise reduction and echo cancellation. An adaptive filter based on cross spectral estimation is used to cancel echo. The external background noise is eliminated and the clear speech is estimated by using MMSE log-spectral magnitude estimation. To be suitable for use in consumer electronics, we also design a low cost, high speed and flexible hardware architecture. The performance of the proposed speech enhancement algorithms were measured both by the signal-to-noise ratio(SNR) and recognition accuracy of an automatic speech recognition(ASR) and yields better results compared with the conventional methods.
군사적 환경에서 음성인식 모델의 취약성에 관한 연구 KCI 등재
한국융합보안학회 융합보안논문지 제24권 제2호 2024.06 pp.201-207
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
목소리는 인간의 의사소통에서 중요한 요소로, 음성인식 모델의 발전은 인공지능의 중요한 성과 중 하나이며 최근 인간의 생활에 다방면으로 사용되고 있다. 음성인식 모델의 활용은 군사분야에서도 피해갈 수 없는 과제이다. 하지만 인공지능 모델 의 군사적 활용 이전에 모델의 취약성에 대한 연구가 필요하다. 본 연구에서는 다국적 음성인식 모델인 Whisper의 군사적 활 용 가능성을 알아보기 위해, 전장소음, 잡음, 적대적 공격에 대한 취약성을 평가하였다. 전장소음을 포함하는 실험에서는 Whisper의 성능 저하가 크게 나타났으며, 평균 72.4%의 문자 오류율(CER)을 기록하여 군사적 활용에 어려움이 있는 것으로 나타났다. 또한, 잡음을 포함하는 실험에서는 낮은 강도의 잡음에 대해 Whisper가 강건하였으나, 높은 강도의 잡음에서는 성 능이 저하되었고, 적대적 공격 실험에서는 특정 입실론 값에서 취약성이 드러났다. 따라서 Whisper 모델을 군사적 환경에서 사용하기 위해서는 파인튜닝, 적대적 훈련 등을 통해 개선이 필요하다는 것을 시사한다.
Voice is a critical element of human communication, and the development of speech recognition models is one of the si gnificant achievements in artificial intelligence, which has recently been applied in various aspects of human life. The appli cation of speech recognition models in the military field is also inevitable. However, before artificial intelligence models ca n be applied in the military, it is necessary to research their vulnerabilities. In this study, we evaluates the military applica bility of the multilingual speech recognition model "Whisper" by examining its vulnerabilities to battlefield noise, white noi se, and adversarial attacks. In experiments involving battlefield noise, Whisper showed significant performance degradation with an average Character Error Rate (CER) of 72.4%, indicating difficulties in military applications. In experiments with white noise, Whisper was robust to low-intensity noise but showed performance degradation under high-intensity noise. A dversarial attack experiments revealed vulnerabilities at specific epsilon values. Therefore, the Whisper model requires impr ovements through fine-tuning, adversarial training, and other methods.
디지털 환경 변화에 따른 통역 교육 ― 자동음성인식 활용을 중심으로 ― KCI 등재
한국일본학회 일본학보 제133권 2022.11 pp.95-114
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
디지털 환경의 변화에 따라 컴퓨터 보조 통역(CAI)과 컴퓨터 보조 통역 교육(CAIT) 에 활용되는 기술이 다양해진 가운데 한국과 일본도 2020년을 기점으로 기술의 활용과 관련된 통역 연구가 증가하고 있다. CAI 및 CAIT는 유럽을 중심으로 다양한 툴이 개발 되고 있으며, 전문용어 관리 등 통역의 사전준비와 원격동시통역 시스템에서 음성인식 으로 통역을 보조하기도 한다. 본 논문에서는 최근 부상하고 있는 자동음성인식(ASR) 을 활용한 통역과 통역 교육의 현황을 조사하여 통역 교육에 활용할 수 있는 툴을 정리 했다. 먼저 현재 사용 가능한 다국어 음성인식 범용 애플리케이션을 기능별로 분류했 다. 그리고 원문과 통역문에 각각 활용할 수 있는 앱으로 나누어 제시했다. 다음으로 통역 실무에서 활용하기 위한 실시간 음성인식 앱, 통역 훈련에 사용 가능한 녹음 후 전사하는 기능의 앱으로도 나누었다. 실시간 음성인식 앱은 타임라인 표시의 유무, 사전에 설정해 놓을 수 있는 언어의 수에 따라 나뉜다. 녹음 후에 전사되는 앱은 통역 과 제물을 첨삭하거나 자가 학습에 활용할 수 있다는 점에서 유용하다. 마지막으로 현재 개발 중인 통역 퍼포먼스의 자동평가 시스템의 프로토타입을 소개했다. ASR 결과로부 터 통역에서 발생할 수 있는 침묵, 필러, 백트래킹을 추출해 통계 처리해 주는 시스템이 다. 향후 음성인식을 통역 교육에 도입하는 방안이 다각도로 모색되어 훈련 과정에서 충분히 적응함으로써 초급에서 고급 및 실무에 이르기까지 유용하게 활용되기를 바란다.
With changes in the digital environment, the technology used for Computer Assisted Interpretation (CAI) and Computer Assisted Interpretation Training (CAIT) has diversified and studies related to the use of technology for interpreting is increasing since 2020. This paper focuses on the use of Automatic Speech Recognition (ASR) for interpreting and interpreter training, which is an emerging field of study, by summarizing the current state and trend in terms of tools that can be used for interpreter training. The currently available all-purpose multilingual ASR applications (apps) were classified by function and they were then categorized according to whether they can be used for transcribing original texts or translated texts. Next, these were divided into apps with real-time ASR functions that can be used for interpretation training, and apps with recording and transcribing functions that can be used for post-interpretation correction and reflection. The real-time ASR apps were further categorized according to timeline display availability and the number of preset languages. The present work also introduces an automatic interpretation performance evaluation system prototype that is currently under development. It is a platform that extracts from ASR results every silence, filler and backtracking which occur during interpretation to provide processed statistics. It is expected that various ways for introducing ASR into interpretation education will be explored in the future, from beginner to advanced level, following sufficient practical adaption through training.
음성 길이 및 발화 방법에 따른 자동화자식별 시스템 성능 확인에 관한 연구
한국법과학회 한국법과학회 학술대회 2002 국립과학수사연구소 학술 발표 대회 및 워크샵 / 제6회 한국법과학회 추계 학술 발표 대회 2002.12 p.120
Automatic Speech Recognition Technique for Bangla Words
보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology Vol.50 2013.01 pp.51-60
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Automatic recognition of spoken words is one of the most challenging tasks in the field of speech recognition. The difficulty of this task is due to the acoustic similarity of many of the words and their syllabi. Accurate recognition requires the system to perform fine phonetic distinctions. This paper presents a technique for recognizing spoken words in Bangla. In this study we first derive feature from spoken words. This paper presents some technique for recognizing spoken words in Bangla. In this work we use MFCC, LPC, GMM and DTW.
Robust Automatic Speech recognition System Implemented in a Hybrid Design DSP-FPGA
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.6 No.5 2013.10 pp.333-342
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The aim of this work is to reduce the burden task on the DSP processor by transferring a parallel computation part on a configurable circuits FPGA, in automatic speech recognition module design, signal pre-processing, feature selection and optimization, models construction and finally classification phase are necessary. LMS filter algorithm that contains more parallelism and more MACs (multiply and Accumulate) operations is implemented on FPGA Virtex 5 by Xilings , MFCCs features extraction and DTW( dynamic time wrapping ) method is used as a classifier. Major contribution of this work are hybrid solution DSP and FPGA in real time speech recognition system design, the optimization of number of MAC-core within the FPGA this result is obtained by sharing MAC resources between two operation phases: computation of output filter and updating LMS filter coefficients. The paper also provides a hardware solution of the filter with detailed description of asynchronous interface of FPGA circuit and TMS320C6713-EMIF component. The results of simulation shows an improvement in time computation and by optimizing the implementation on the FPGA a gain in space consumption is obtained.
A Review on Automatic Speech Recognition Architecture and Approaches
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.9 No.4 2016.04 pp.393-404
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Speech is the most natural communication mode for human beings. The task of speech recognition is to convert speech into a sequence of words by a computer program. Speech recognition applications enable people to use speech as another input mode to interact with applications with ease and effectively. Speech recognition interfaces in native language will enable the illiterate/semi-literate people to use the technology to greater extent without the knowledge of operating with computer keyboard or stylus. For more than three decades, a great amount of research was carried out on various aspects of speech recognition and its applications. Today many products have been developed that successfully utilize automatic speech recognition for communication between human and machines. Performance of speech recognition applications deteriorates in the presence of reverberation and even low levels of ambient noise. Robustness to noise, reverberation and characteristics of the transducer is still an unsolved problem that makes the research in the area of speech recognition still very active. A detailed study on automatic speech recognition is carried out and presented in this paper that covers the architecture, speech parameterization, methodologies, characteristics, issues, databases, tools and applications.
Comparative Analysis of Arabic Vowels using Formants and an Automatic Speech Recognition System
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition vol.3 no.2 2010.06 pp.11-22
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Arabic, the world’s second most spoken language in terms of number of speakers, has not received much attention from the traditional speech processing research community. This study is specifically concerned with the analysis of vowels in modern standard Arabic dialect. The first and second formant values in these vowels are investigated and the differences and similarities between the vowels explored using consonant-vowels-consonant (CVC) utterances. For this purpose, a Hidden Markov Model (HMM) based recognizer is built to classify the vowels and the performance of the recognizer analyzed to help understand the similarities and dissimilarities between the phonetic features of vowels. The vowels are also analyzed in both time and frequency domains, and the consistent findings of the analysis are expected to enable future Arabic speech processing tasks such as vowel and speech recognition and classification.
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition vol.2 no.2 2009.06 pp.67-74
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Environmental robustness is an important area of research in speech recognition. Mismatch between trained speech models and actual speech to be recognized is due to factors like background noise. It can cause severe degradation in the accuracy of recognizers which are based on commonly used features like mel-frequency cepstral co-efficient (MFCC) and linear predictive coding (LPC). It is well understood that all previous auditory based feature extraction methods perform extremely well in terms of robustness due to the dominantfrequency information present in them. But these methods suffer from high computational cost. Another method called sub-band spectral centroid histograms (SSCH) integrates dominant-frequency information with sub-band power information. This method is based on sub-band spectral centroids (SSC) which are closely related to spectral peaks for both clean and noisy speech. Since SSC can be computed efficiently from short-term speech power spectrum estimate, SSCH method is quite robust to background additive noise at a lower computational cost. It has been noted that MFCC method outperforms SSCH method in the case of clean speech. However in the case of speech with additive noise, MFCC method degrades substantially. In this paper, both MFCC and SSCH feature extraction have been implemented in Carnegie Melon University (CMU) Sphinx 4.0 and trained and tested on AN4 database for clean and noisy speech. Finally, a robust speech recognizer which automatically employs either MFCC or SSCH feature extraction methods based on the variance of shortterm power of the input utterance is suggested.
음성인식과 자연어 처리 딥러닝을 통한 전자의무기록 자동 생성 시스템 KCI 등재
국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.9 No.3 2023.05 pp.731-736
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
최근 의료 현장은 전자의무기록, 전자건강기록 등의 의료 기록을 전산화하여 저장하고 관리하는 시스템이 의무적으로 적 용되거나 전체 의료 현장에 보급되어 환자 개개인의 과거 의료 기록을 추가적인 의료 행위에 활용하고 있다. 그러나 일반적인 의료 문진 및 상담 간 발생하는 의료진과 환자 간의 대화는 별도로 기록되거나 저장되지 않고 있어 추가적인 환자의 주요 정보 는 효율적으로 활용되지 못하고 있다. 이에 따라, 의료 문진 현장에서 발생하는 의료진과 환자와의 대화를 저장하고 이를 텍스 트 데이터로 변환하여 주요한 문진 내용만 자동으로 추출, 요약하여 정보화하는 음성인식과 자연어 처리 딥러닝을 통한 의료 상담 요약문을 자동으로 생성하는 전자의무기록 시스템을 제안한다. 본 시스템은 의료 종사자와 환자의 의료 상담 내용의 인식 과정을 거쳐서 텍스트 정보를 획득한다. 이렇게 획득된 텍스트를 복수의 문장으로 구분하고, 생성된 문장에 포함된 복수 키워드 의 중요도를 산출한다. 산출된 중요도를 기반으로 복수의 문장에 순위를 매기고, 순위를 기반으로 문장들을 요약하여 최종 전자 의무기록 데이터를 생성한다. 제안하는 시스템 성능은 정량적 분석을 통하여 우수함을 확인한다.
Recently, the medical field has been applying mandatory Electronic Medical Records (EMRs) and Electronic Health Records (EHRs) systems that computerize and manage medical records, and distributing them throughout the entire medical industry to utilize patients' past medical records for additional medical procedures. However, the conversations between medical professionals and patients that occur during general medical consultations and counseling sessions are not separately recorded or stored, so additional important patient information cannot be efficiently utilized. Therefore, we propose an electronic medical record system that uses speech recognition and natural language processing deep learning to store conversations between medical professionals and patients in text form, automatically extracts and summarizes important medical consultation information, and generates electronic medical records. The system acquires text information through the recognition process of medical professionals and patients' medical consultation content. The acquired text is then divided into multiple sentences, and the importance of multiple keywords included in the generated sentences is calculated. Based on the calculated importance, the system ranks multiple sentences and summarizes them to create the final electronic medical record data. The proposed system's performance is verified to be excellent through quantitative analysis.
고령층의 디지털 소외 방지를 위한 ASR(Automatic Speech Recognition, 음성 인식 기술) 기반 복지 정보 검색 모델 연구
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2023 pp.771-772
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
복지 정보와 인터넷 사용에 대한 이해도가 낮은 고령층의 디지털 소외 문제를 해결하고자, 고령층 친화 UI/UX 및 음성 인식 기술 등의 기술을 활용한 <고령층의 디지털 소외 방지를 위한 ASR 기반 복지 정보 검색 모델>의 개발을 제안한다.
A Single Channel Speech Enhancement for Automatic Speech Recognition
[Kisti 연계] 한국방송공학회 한국방송공학회 학술대회논문집 2011 pp.85-88
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper describes a single channel speech enhancement as the pre-processor of automatic speech recognition system. The improvements are based on using optimally modified log-spectra (OM-LSA) gain function with a non-causal a priori signal-to-noise ratio (SNR) estimation. Experimental results show that the proposed method gives better perceptual evaluation of speech quality score (PESQ) and lower log-spectral distance, and also better word accuracy. In the enhancement system, parameters was turned for automatic speech recognition.
Automatic Segmentation of Korean Oral Stops for Automatic Speech Recognition
[NRF 연계] 한국외국어대학교 언어연구소 언어와 언어학 Vol.34 2004.09 pp.135-150
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
한국어 폐쇄음의 음성인식에 사용될 폐쇄 구간과 Voice Onset Time (VOT) 구간을 얻어내기 위하여 자동으로 폐쇄음의 폐쇄구간과 폐쇄음 뒤 모음 시작점을 분할하는 방법을 제안한다. 한국어 폐쇄음은 같은 조음위치에서 세 가지 Phonation type을 갖는다. 경음은 연음이나 격음보다 긴 폐쇄 구간을 가지고 있고, 격음은 경음이나 연음보다 긴 VOT를 가지고 있다는 음성학적 연구 성과에 비추어, 먼저 연속음성 데이터베이스인 KAIST 음성데이터에서 628개의 문장을 hand labelling 한 후, 통계 실험을 하였다. 이 실험은 문맥의 다양성에도 불구하고 폐쇄구간과 VOT의 특성이 phonation type에 따라 기존 음성학 연구에서의 결과가 같은지를 알아보기 위함이다. 실험 결과, 문맥의 다양성에도 불구하고 기존 연구 결과와 일치 함을 보였다, 이러한 결과는 음성인식에서 폐쇄음을 구분하기 위하여 폐쇄구간과 VOT의 두 가지 음향단서를 사용할 수 있음을 입증하는 것이다. 음성인식기에서 폐쇄음의 phonation type을 구별하기 위한 폐쇄구간과 VOT를 자동으로 추출할 수 있는 자동 분할 방법을 개발하였다. 클러스터링과 두 연속된 창에서의 로그 파워의 차이값의 local peak를 계산함에 의해서 이러한 구간들의 분할을 할 수 있다.
[Kisti 연계] 한국전자통신연구원 ETRI journal Vol.35 No.1 2013 pp.100-108
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this paper, a feature extraction (FE) method is proposed that is comparable to the traditional FE methods used in automatic speech recognition systems. Unlike the conventional spectral-based FE methods, the proposed method evaluates the similarities between an embedded speech signal and a set of predefined speech attractor models in the reconstructed phase space (RPS) domain. In the first step, a set of Gaussian mixture models is trained to represent the speech attractors in the RPS. Next, for a new input speech frame, a posterior-probability-based feature vector is evaluated, which represents the similarity between the embedded frame and the learned speech attractors. We conduct experiments for a speech recognition task utilizing a toolkit based on hidden Markov models, over FARSDAT, a well-known Persian speech corpus. Through the proposed FE method, we gain 3.11% absolute phoneme error rate improvement in comparison to the baseline system, which exploits the mel-frequency cepstral coefficient FE method.
[NRF 연계] 한국청각언어재활학회 Audiology and Speech Research Vol.20 No.4 2024.10 pp.253-262
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Purpose: The purpose of this study is to examine recent research trends regarding automatic speech recognition (ASR), which is used in the evaluation and intervention of speech disorders. Methods: Through a search engine, articles published in domestic journals were searched. A total of 27 papers were selected from the searched documents and analyzed according to the year, research subject, speech task, and ASR system. Results: The years with the most research was done in 2019~2021. The subjects who most frequently underwent speech evaluation and treatment using ASR system were those with dysarthria. The speech production tasks used to utilize ASR were at the word and sentence level, and commercialized and non-commercialized ASR systems were used similarly. Conclusion: These results might be used for the fundamental data to establish the most suitable evaluation methods and intervention plans for patients with speech disorders in clinical and research fields.
[Kisti 연계] 한국음향학회 한국음향학회지 Vol.29 No.e2 2010 pp.86-99
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper analyzes the performance of various single channel speech enhancement algorithms when they are applied to automatic speech recognition (ASR) systems as a preprocessor. The functional modules of speech enhancement systems are first divided into four major modules such as a gain estimator, a noise power spectrum estimator, a priori signal to noise ratio (SNR) estimator, and a speech absence probability (SAP) estimator. We investigate the relationship between speech recognition accuracy and the roles of each module. Simulation results show that the Wiener filter outperforms other gain functions such as minimum mean square error-short time spectral amplitude (MMSE-STSA) and minimum mean square error-log spectral amplitude (MMSE-LSA) estimators when a perfect noise estimator is applied. When the performance of the noise estimator degrades, however, MMSE methods including the decision directed module to estimate a priori SNR and the SAP estimation module helps to improve the performance of the enhancement algorithm for speech recognition systems.
HMM-Based Automatic Speech Recognition using EMG Signal
[Kisti 연계] 대한의용생체공학회 Journal of biomedical engineering research : the official journal of the Korean Society of Medical & Biological Engineering Vol.27 No.3 2006 pp.101-109
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
It has been known that there is strong relationship between human voices and the movements of the articulatory facial muscles. In this paper, we utilize this knowledge to implement an automatic speech recognition scheme which uses solely surface electromyogram (EMG) signals. The EMG signals were acquired from three articulatory facial muscles. Preliminary, 10 Korean digits were used as recognition variables. The various feature parameters including filter bank outputs, linear predictive coefficients and cepstrum coefficients were evaluated to find the appropriate parameters for EMG-based speech recognition. The sequence of the EMG signals for each word is modelled by a hidden Markov model (HMM) framework. A continuous word recognition approach was investigated in this work. Hence, the model for each word is obtained by concatenating the subword models and the embedded re-estimation techniques were employed in the training stage. The findings indicate that such a system may have a capacity to recognize speech signals with an accuracy of up to 90%, in case when mel-filter bank output was used as the feature parameters for recognition.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.