년 - 년
통신신호 탐지를 위한 딥러닝 기반 스펙트럼 센싱 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제9권 11호 2025.11 pp.2853-2862
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
군 통신을 중심으로 적의 탐지를 최소화하기 위해 저피탐(low probability of detection, LPD) 통신 시스템의 개발이 관심을 받고 있으며, 다양한 환경에서 높은 탐지 확률을 보장하는 신호 탐지 기술에 대한 필요성이 대두되고 있다. 따라서 본 논문은 스펙트럼 센싱을 이용하여 신호 존재 여부를 판단하는 통신 신호 탐지 기법을 제안한다. 수집 신호는 고속 샘플링된 후 fast fourier transform (FFT)를 통해 주파수 스펙트럼으로 변환된다. 변환된 스펙트럼은 시간 축으로 누적되어 2차원 스펙트로그램을 형성하며, 이를 딥러닝 모델의 입력으로 사용한다. 일차적으로 합성곱 신경망(convolutional neural network, CNN)을 이용해 신호 존재 여부를 판단하는 기법을 제안한다. 컴퓨터 모의실험 결과, 제안하는 CNN 기법은 기존 임계값 기반 기법 대비 약 2dB의 성능 향상을 보인다. 특히, signal to noise ratio (SNR) = –6dB 환경에서 약 35% 높은 탐지 정확도를 보인다. 또한, you only look once (YOLO)를 활용하여 신호의 시간-주파수 위치를 인식하는 기법을 제안한다. 실험 결과, YOLO 기반 시간-주파수 위치 인식은 SNR = 1dB부터 80% 이상의 intersection over union (IOU) 성능을 달성한다.
In military communications, low probability of detection (LPD) systems are increasingly important to minimize enemy interception, driving demand for robust detection techniques. This paper proposes a spectrum sensing–based method to determine signal presence. Collected signals are sampled, transformed by fast fourier transform (FFT), and accumulated into spectrograms for deep learning input. A convolutional neural network (CNN) identifies signal existence, achieving approximately 2 dB gain and 35% higher accuracy at signal to noise ratio (SNR) = –6 dB. A you only look once (YOLO) model then localizes time– frequency positions, reaching over 80% Intersection Over Union (IOU) from SNR = 1 dB onward.
서포트 벡터 머신(SVM)을 활용한 드론과 조류의 마이크로 도플러 신호의 분석 정확도 분석 연구 KCI 등재후보
중소기업융합학회 산업과 과학 제4권 제4호 2025.07 pp.49-55
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문에서는 서포트 벡터 머신(Support Vector Machine: SVM) 기법을 사용하여 드론과 조류의 움직임을 고려하여 생성된 마이크로 도플러 신호를 분석하여 드론과 조류를 구별하는 시뮬레이션 연구를 제시하였다. 두 개체의 스펙트로그램에서 얻은 시간-주파수 특징은 분류를 위해 평균 스펙트럼 프로파일로 변환된다. SVM 분류기는 신호의 스펙 트로그램을 활용하여 학습되며, 드론과 조류의 동작 매개변수(드론 속도, 드론 블레이드 회전, 조류 날갯짓)의 변동성을 종합적으로 고려하여 분석되었다. 제안된 방법은 드론 블레이드 회전 속도 2,000rpm(±25% 잡음)과 본체 속도 8m/s(±50% 잡음)를 기반으로 한 학습 신호를 사용하여 블레이드 회전 속도가 1,600rpm 정도로 낮은 경우에도 거의 100%의 분류 정확도를 보였다. 블레이드 회전으로 인한 도플러 주파수가 본체 이동에 의한 도플러 주파수보다 상당히 높기 때문에 본체 속도가 분류 성능에 미치는 영향은 무시할 수 있다. 본 연구 결과는 소형 드론을 구별하는 데 있어 SVM 기반 분류 기술의 효과를 제시하였으며, 이 방법의 계산 효율성 측면에서 실시간 감시 시나리오에 매우 적합할 것으로 판단된다.
In this paper, a simulation study for distinguishing between drones and birds by analyzing simulated radar micro-Doppler signatures using a Support Vector Machine (SVM) classifier is presented. Time-frequency features obtained from spectrograms are condensed into average spectral profiles for classification purposes. The SVM classifier is trained using these spectrogram-based features, incorporating variability in drone and bird motion parameters to improve its generalization capability. Results show that with training signals based on a drone rotor speed of 2,000rpm(±25% noise) and body speed of 8m/s(±50% noise), the classifier achieved nearly 100% accuracy for test cases with rotor speeds as low as 1,600rpm. Since the Doppler frequency generated by wing rotation is significantly higher than that produced by body translation, the effect of body velocity on classification performance is negligible. The results confirm the effectiveness of the SVM-based technique in small drone discrimination, and its computational efficiency makes it well-suited for real-time surveillance scenarios.
트라우마 초기 진단을 위한 음성 기반 감정 분류 방법 KCI 등재후보
국제차세대융합기술학회 차세대융합기술학회논문지 제4권 5호 2020.10 pp.509-515
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 현대인들은 다양한 원인으로 인해 트라우마 증상을 겪고 있다. 트라우마를 겪으면 감정 조절에 문 제가 생기며 불안해하고 안정감을 찾기 어려워한다. 이런 트라우마를 진단하기 위해서 병원을 방문하는 것이 필수 적이지만 사회적 편견과 자신이 가지고 있는 트라우마에 대해 인지하지 못해 진단 및 치료를 놓치는 경우가 많다. 이를 위해 본 논문에서는 합성곱 신경망 모델을 통해 음성을 기반으로 트라우마를 초기 진단하는 방법을 연구하 였다. 기본 감정 6가지 중 fear, sad, neutral, happy 4가지 감정을 트라우마 초기 진단에 활용하였다. 음성에서 fear 와 sad 감정이 나타나면 트라우마일 확률이 높고, neutral과 happy 감정이 나타나면 트라우마가 아닐 확률이 높 다고 가정하였다. 음성의 특징을 추출하기 위해 전처리 과정을 거쳐 푸리에변환 한 spectrogram 이미지로 만들고 학습하였다. 그 결과, VGG-13 모델에서 98.96%로 높은 정확도를 나타내 트라우마 진단에 좋은 결과를 얻을 수 있었다.
Recently, modern people are experiencing various types of trauma and have difficulty in controlling their emotions and stabilizing. It is essential to visit a hospital to diagnose trauma but lots of people often miss diagnosis and treatment because of being unaware of it. In this paper, we propose a method for screening trauma based on audio data using convolutional neural networks. Among the six basic emotions, four emotions were used for screening trauma: fear, sad, neutral, and happy. It was assumed that when fear and sad emotions appeared in the audio, the probability of trauma was high and when neutral and happy emotions appeared, the probability was low. To extract the features from audio, the audio was converted into spectrogram images through pre-processing and used to train convolutional neural networks. As a result, VGG-13 model showed the highest performance(98.96%) for screening trauma among others.
반향 소리를 이용한 기계 학습 기반 수박의 당도 예측 KCI 등재
한국융합학회 한국융합학회논문지 제11권 제8호 2020.08 pp.1-6
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
수박의 맛을 평가하는 다양한 방식이 있으나, 기존의 방법들은 주관적 방식, 평가 비용, 대상의 손상 등과 같은 평가 방식의 한계점이 있다. 최근에는 이러한 단점들을 해소하기 위해 소리를 이용하여 수박을 평가하는 연구들이 진행 되고 있다. 본 연구에서는 수박을 두드렸을 때 나는 반향 소리를 AI기반의 기계 학습을 이용하여 수박의 당도를 예측하 는 모델을 개발 하였다. 수박의 당도가 높을수록 높은 주파수 성분이 특이점으로 나타나며, 따라서 반향소리 시간-주파 수 특이점에 기반 하여 기계 학습 방법을 개발하였다. 2개의 수박 당도별 그룹을 구분 시에 83.2%, 3개의 그룹을 구분 시에 59.6%의 정확도로 당도를 예측 할 수 있었다.
There are various approaches to evaluate a watermelon sweetness. However, there are some limitations to evaluating cost, watermelon damage, and subjective issue. In this study, we developed a novel approach to predict a watermelon sweetness using reflected sound and the machine learning algorithm. It was observed that higher brix watermelon produced higher spectral power is reflected sound. Based on the spectral-temporal features of reflected sound, the machine learning algorithms could accurately predict the sweetness group at a rate of 83.2 and 59.6 % in 2-groups and 3-groups classification, respectively.
스펙트로그램을 이용한 근위축성측삭경화증 여성 화자의 모음 포먼트, 음성강도, 기본주파수의 변화 KCI 등재
한국융합학회 한국융합학회논문지 제10권 제9호 2019.09 pp.193-198
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 근위축성측삭경화증(amyotrophic lateral sclerosis, ALS)으로 진단된 여성을 대상으로 음향음성학 적 스펙트로그램 분석을 이용하여 11개월 동안 모음과 이중모음의 포먼트 변화(vowel formant variation)를 분석하였 다. 검사어는 단모음 /a, i, u/와 이중모음 /h + ja + da/, /h + wi + da/, /h +ɰi+ da/를 이용하였다. 발화자료는 ‘Alvin’프로그램을 이용하여 모니터에 제시된 단어읽기과제를 통해 수집되었고, 녹음환경은 nyquist frequency는 5,500Hz, sampling rate는 11,000Hz으로 설정하였다. 녹음자료는 스펙트로그램을 이용하여 강도, 음도와 이중모음의 포먼트를 분석하였다. 분석결과, ALS의 진행과정에서 기본주파수와 강도가 저하되었고, 단모음에서의 포먼트 변화보다 는 이중모음의 포먼트 기울기의 감소가 특징으로 확인되었다. 이 결과는 병의 진행에 따른 ALS의 모음왜곡이 혀와 턱의 협응력 감소에 기인함을 시사한다.
This study analyzed the changes of vowel formant, voice intensity, and fundamental frequency of vowels for 11 months using acoustochemical spectrogram analysis of women diagnosed with amyotrophic lateral sclerosis (ALS). The test word was a vowel /a, i, u/ and a diphthong /h + ja + da/, /h + wi + da/, and /h +ɰi+ da/. Speech data were collected through the word reading task presented on the monitor using 'Alvin' program, and the recording environment was set to 5,500 Hz for the nyquist frequency and 11,000 Hz for the sampling rate. The records were analyzed by using spectrograms to vowel formants, voice intensity, and fundamental frequency. As a result of analysis, the fundamental frequency and intensity of the ALS process were decreased and the formant slope of the diphthong was decreased rather than the formant change in the vowel. This result suggests that the vowel distortion of ALS due to disease progression is due to the decrease of tongue and jaw co morbidity.
교육산업 활성화를 위한 영어발음 연구 KCI 등재
중소기업융합학회 융합정보논문지(구 중소기업융합학회논문지) 제8권 제1호 2018.02 pp.37-42
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 다섯 개의 문장을 선정하여 그 안에 들어있는 다섯 개의 단어와 그 단어 내에 들어있는 모음들을 대상으 로 하여 한국인들의 모음발음 길이를 Praat 프로그램을 이용하여 측정한 후에 원어민의 실험결과와 비교분석한 연구이다. 20명의 한국인 피 실험자들이 실험을 위해 실험문장을 읽고 녹음하였으며, 음향적 특질들은 스펙트로그램 등을 활용한 Praat Software 프로그램을 이용하여 측정하여 그 결과를 얻었다. 이러한 과정을 통해 얻어진 결과들은 통계분석 하였다. 분석결과, 단어단위에서는 3개의 단어에서 두 집단 간 차이가 있었으며, 5개의 모음들은 모두 그 차이가 유의미하였다. 또한 두 실험집단 간 단어와 모음발음크기 비율에 관한 비교에서는 모든 단어와 모음들의 발음크기 비율이 두 집단 간 유의미한 차이를 보여주었다. 이러한 실험결과를 통해 영어전설저모음인 /æ/와 영어이중모음/ai/, 그리고 첫 번째 음절에 영어강세모 음이 오는 경우에 그 발음 길이크기를 분석해보면 한국인 피 실험자집단에서는 외국인어투가 존재함을 다시 확인할 수 있었다.
This study focuses on investigating and comparing the lengths of the five words, vowels, and the ratio of the length of vowels to that of words among the Korean college students with the English native speaker. English sentences were read and recorded by Korean subjects to do this experiment. The vowel lengths were measured from a sound spectrogram, the Praat software program, and these data were analyzed through statistical analysis. I could easily tell that there were differences between the groups and they were significant. In the English front low vowel /æ/, I was able to find out that native subjects pronounced differently from Korean subjects, and the differences were significant. However, the pronunciation of the English diphthong /ai/, native subjects pronounced significantly shorter than Korean subjects.
영어강세음절의 외국인어투에 관한 연구 KCI 등재후보
중소기업융합학회 융합정보논문지(구 중소기업융합학회논문지) 제6권 제4호 2016.12 pp.51-57
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 강세음절이 있는 8개의 단어를 선정하여 원어민과 한국 대학생들 사이의 모음발음 길이를 스펙트 로그램을 이용하여 측정한 후에 비교분석한 실험적 연구이다. 이 실험을 위하여 20명의 한국인 피 실험자들이 8개의 단어들이 들어있는 문장들을 발화하고 녹음하였으며, 음향적 특질들은 Praat 소프트웨어 프로그램을 이용하여 측정 하였으며 그 결과를 통계분석 하였다. 분석결과, 8개의 강세모음에서 두 집단 간 차이가 있었으며, 7개의 강세모음에 서는 그 차이가 유의미하였다. 두 실험집단 간 실험결과를 보면, 제1음절에 강세가 있는 모음들은 모두 집단 간 유의 미한 차이를 보여주었다. 그 중에서 wonderful과 glasses의 강세음절에서는 유의미성이 크게 나타나고 있었는데, 특 히 영어저모음 /æ/의 발음에서는 원어민이 한국인집단보다 훨씬 큰 길이로 발음하는 것을 알 수 있었다. 이러한 실험 결과는 영어교육현장에서 외국인어투의 개선을 위한 수업자료로 활용할 수 있으리라 판단된다.
This study aims at investigating and comparing the vowel lengths of the eight stressed syllable vowels among the Korean college students with the English native speakers. To do this English sentences were uttered and recorded by twenty Korean subjects. Acoustic features were measured from a sound spectrogram with the help of the Praat software program and analyzed through statistical analysis. From the results of the experiment, I was able to find out that the differences of the lengths of the first syllable stressed vowels were significant. Especially in the pronunciation of the English front low vowel /æ/, native subjects pronounced significantly longer than Korean subjects, and this result could be used as a teaching material in pronunciation class.
5,700원
The main purpose of this thesis is to analyze the types of errors by Korean college students in pronouncing English fricatives. The results of the tests are as follows: In production test, the highest 78.33% errors came out for /θ-ð/, followed by 50.42% for /f-v/, 34.58% for /s-z/ and 23.4% for /ʃ-s/. The minimal pair with the /θ-ð/ came out with the highest error rate. Especially, Korean college students mainly chose /s/ rather than /t/ or /d/ for English /θ/ and /ð/. This result shows that within the phonological structure of Korean the difference between a sibilant such as /s/ and a fricative such as /θ/ has less phonemic "force" than the difference between a stop and a fricative
음성 기반 치매 조기 진단 연구의 확장 : Mel-Spectrogram 중심 접근의 한계와 언어적 특징 기반 다중모달 분석 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제10권 3호 2026.03 pp.740-754
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
고령화 사회의 가속화로 치매 조기 선별의 중요성이 강조되는 가운데, 음성 기반 인공지능은 비침습적·저비 용 대안으로 주목받고 있다. 선행 연구에서는 음성 신호를 멜 스펙트로그램(Mel-Spectrogram)으로 변환하여 CNN, ViT 모델 등에 적용하였으나, 분류 정확도가 약 61~62% 수준에 머물며 음향적 특징만으로는 치매 특유의 인지 저 하를 포착하는 데 구조적 한계가 있음을 확인하였다. 본 연구는 이러한 한계를 극복하기 위해 ADRESS-2020 데이 터셋을 기반으로 전사 텍스트에서 추출한 언어적 특징을 결합한 다중모달 분석 접근을 제안하였다. 연구 결과, 어휘 다양성, 문장 복잡도, 의미적 응집성 등 14개의 언어적 변수만으로도 교차검증 정확도 76.8%를 달성하며 선행 연구 의 음향 기반 모델 성능을 크게 상회하였다. 특히 특징 수 대비 성능 효율성 측면에서 언어적 특징은 고차원 딥러닝 특징보다 월등히 높은 수치를 기록하여, 소규모 의료 데이터 환경에서 특징의 질적 설계가 중요함을 입증하였다. 음 향과 언어 특징의 결합은 안정성과 분산 측면에서 가장 균형 잡힌 결과를 나타냈으나, 모든 특징을 결합한 고차원 환경에서는 차원의 저주로 인한 성능 저하가 관찰되었다. 모델 비교에서는 로지스틱 회귀가 가장 우수한 일반화 성 능을 보였으며, 이는 실제 임상 현장에서 해석 가능하고 단순한 모델의 실용성이 높음을 시사한다. 본 연구는 언어적 특징 중심의 다중모달 분석이 치매 조기 선별의 정확성과 신뢰성을 높이는 핵심 전략임을 실증하였다.
As the acceleration of population aging intensifies the importance of early dementia screening, voice-based artificial intelligence is gaining significant attention as a non-invasive and cost-effective alternative. A previous study utilized voice signals converted into Mel-spectrograms and applied them to models such as CNN and ViT, but found that classification accuracy remained at approximately 61–62%, confirming structural limitations in capturing dementia-specific cognitive decline using only acoustic features . To overcome these limitations, this study proposes a multimodal analysis approach that integrates linguistic features extracted from transcribed text using the ADRESS-2020 dataset. The experimental results demonstrated that just 14 linguistic variables—including lexical diversity, syntactic complexity, and semantic coherence—achieved a cross-validation accuracy of 76.8%, significantly outperforming the acoustic-based models from the previous study. In terms of performance efficiency relative to the number of features, linguistic features recorded substantially higher values than high-dimensional deep learning features, proving that qualitative feature engineering is crucial in small-scale medical data environments. While the combination of acoustic and linguistic features yielded the most balanced results in terms of stability and variance, a performance decline due to the "curse of dimensionality" was observed when all high-dimensional features were combined. In model comparisons, logistic regression exhibited the most superior generalization performance, suggesting that simple, interpretable models are more practical for real-world clinical settings. This study empirically validates that multimodal analysis centered on linguistic features is a core strategy for enhancing the accuracy and reliability of early dementia screening.
Mel-Spectrogram을 활용한 CNN 기반의 퇴행성뇌질환자 분류 모델 개발에 관한 연구
한국경영정보학회 한국경영정보학회 정기 학술대회 디지털플랫폼 성공을 위한 경영정보학의 역할 2023.06 pp.1055-1062
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
사운드 스펙트로그램과 심층신경망 기법을 사용한 패셔너블한 웨어러블 제품 디자인을 위한 테크놀로지 KCI 등재
한국패션디자인학회 한국패션디자인학회지 vol.18 no.2 2018.06 pp.109-124
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
디지털 문화가 보급되면서 다양한 문화 양식과 사조를 공유하게 되었고, 새로운 장르의 스타일을 창조하 는 경향을 낳았다. 이러한 현상으로 패션과 기술이 융합된 디자인 개발도 폭넓게 전개되고 있다. 기존의 창작 과 미에 중점을 둔 스타일과는 다른 새로운 장르인 패션이 가미된 웨어러블 제품은 차세대 상품으로 많은 기대와 관심 속에 시장 도입이 되었다. 그러나 최근 웨어러블 제품의 시장 동향은 성장 없이 혼돈 속에 있으 며, 기술적 기능만을 중시한 상품 도입이 여러 부진 원인중의 하나로 분석되고 있다. 시장을 활성화하기 위한 방안의 하나로 웨어러블 제품은 외부로 드러나는 착용형의 상품이므로, 사용자의 정체성을 표현할 패션이라 는 요소가 융합되고 자동 반응하는, 독자적인 가치를 갖는 디자인을 구현해야 한다. 본 논문에서는 패션과 인공지능 기술을 융합하여 웨어러블 제품에 장착된 디스플레이가 사용자가 자각하지 않은 상태에서 주위 환경에 따라 자동으로 반응하여 사전 디자인에 따라서 패셔너블하게 디스플레이 되도록, 패션이 가미되어 디자인된 디스플레이를 갖는, 패션과 인공지능이 융합하여 자동으로 반응하는 웨어러블 제품에 사용할 테크 놀로지를 제안한다. 사운드 스펙트로그램으로 주위환경에서 입력된 소리의 특징을 추출하고 DNN의 Keras 기법으로 특성을 분류하여 레이블링한 후, 자동으로 반응하여 설정된 패션에 맞게 음향을 패션영상으로 변환 하여 패셔너블하게 디스플레이 하도록 하는 것이다. 제안된 실험 시스템으로 채집된 소리 샘플을 실험한 결 과 –10 dB이내의 노이즈 환경에서 우수한 인식률을 나타내어 웨어러블 제품의 주위 환경에서 채집 가능한 소리를 인식하는 것이 가능함을 확인하였다. 사용자의 사전 디자인에 따라 디스플레이가 자동 반응하여 사용 자의 아이덴터티를 표현할 패셔너블한 웨어러블 제품의 디자인에 사용할 음향-패션영상 변환 맵을 제안하였 다. 실험에서 분류된 레이블로 음향을 패션영상으로 변환하여 패셔너블하게 디스플레이 하는 것이 가능함을 검증하였다.
As digital culture became popular, various cultural styles and customs were shared. This has created a tendency to create new genre styles. As a result of this phenomenon, The development of the design which fuses fashion and technology is widely spread. Wearable products in new genres have been introduced into the market with much expectation. However, market of wearable products is in chaos. The introduction of products with an emphasis on technical functions has been analyzed as one of the reasons for the sluggishness. One of the ways to revitalize the market, due to the characteristic of wearable products, worn as an outer wear, amalgamation of the fashion element, which expresses an identity of a user, and the unique value of design that enables the automatic reaction of artificial intelligence should be implemented. This paper proposes a technological system with a transformation map from sound to graphic image that can be used for designing auto-responsive and fashionable wearable products that amalgamate fashion and artificial intelligence in the stylishly designed display. The display will automatically react to environmental changes and show predefined fashionable display designs without users’ awareness. After the sound spectrogram extracts the characteristics of sound input from the surrounding environment to classify and label them through Keras method of DNN, it automatically reacts to covert the sound input into stylish graphic images and display them fashionably according to the user’s preset design. The result of the experiment with the proposed sound system has proved that the system is capable of recognizing and characterizing the sound inputs around the wearable products when the noise level of the environment is under –10 dB. The experiment has verified that inputted sound can be transformed into a graphic image according to classified labels and it is successfully put on display.
[NRF 연계] 한국장애인개발원 장애인복지연구 Vol.12 No.1 2021.06 pp.35-54
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
이 연구는 말더듬을 보이는 학령전기 및 학령기 남아가 갖는 비유창의 특징 중 ʻ막힘ʼ 구간에서의 길이와 피치와 에너지의 형태가 어떤 양상을 보이는가를 비교 분석해 보는 데 목적이 있다. 그리고 ʻ막힘ʼ 구간에서 분석 된 파라미터 값 중 말더듬의 중증도를 변별해 줄 수 있는 파라미터가 있는가를 확인해 보고자 한다. 연구방법은 학령전기와 학령기 남아 각각 10명씩을 대상으로, P-FA-Ⅱ 구어평가 과업 중의 하나인 필수 과제를 녹음 후 스펙트로그램 상에서 길이, 에너지 형태, 피치 형태를 비교 분석하였다. 연구 결과 첫째, 학령전기 및 학령기 말더듬 아동에게 나타난 총 막힘 빈도는 119회로 평균 길이는 1.33sec(±.952)로 측정되었다. 둘째, 말더듬 아동에게 ʻ막힘ʼ은 세 가지 피치 형태로 나타났다. 에너지 형태의 경우 두 집단에서 다른 형태를 보였다. 마지막으로 중증도에 따라서 ʻ막힘ʼ에 유의한 차이가 있는지를 파악하기 위해서 Schéffe 사후분석을 실시한 결과 길이파라미터에서만 ʻ약함ʼ과 ʻ심함ʼ에서는 차이가 없었으나, ʻ중간ʼ의 중증도에서는 나머지 중증도와 통계적으로 유의한 수준에서 차이가 나타남을 알 수 있었다. 이 연구를 통해 말더듬을 지닌 학령전기 및 학령기 아동의 ʻ막힘ʼ의 길이에서 두 집단 간 차이가 있고 피치 형태와 에너지 형태 역시 다양하게 나타났다. 그리고 비유창성의 대표적인 특징인 ʻ막힘ʼ에서 길이 파라미터가 말더듬의 중증도를 예측할 수 있는 지표로서 역할을 할 수 있음을 확인할 수 있었다.
The purpose of this study is to compare and analyze how length, pitch, and energy contour in the 'blocking' section may be among the characteristics of dysfluency in preschool and school-age boys who stutter. Additionally, among the parameter values analyzed in the ?blockage? section, we would like to check whether there is a parameter that can discriminate the severity of stuttering. As a research method, comparative analysis was performed after recording the required task. One of the P-FA-II oral language assessment tasks was performed on ten boys. These boys were of the preschool age and the school-age. The duration, energy contour, and pitch contour are displayed on the spectrogram. The results are as follows. First, the total frequency of blockage in preschool and school-age stuttering children was 119, and the average length was 1.33 seconds(±.952). Second, in children who stutter 'blocking' appeared in three types of pitch. In the case of energy contour, the two groups showed different results. Finally, Scheffe's post-analysis was performed. As a result, there was no difference in weak and severe. However, in 'medium' severity there was a difference between the remaining severity and at a statistically significant level. In conclusion, through this study, there was a difference between the two groups in the length of 'blocking' of preschool and school-age children with stuttering. The pitch and energy forms were also diverse. In addition, it was confirmed that the length parameter can play a role as an index predicting the severity of stuttering in 'blockage', a representative characteristic of dysfluency.
스마트폰 음성 녹음 파일 위변조 검출을 위한 스펙트로그램 분석의 한계점 KCI 등재
국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.9 No.2 2023.03 pp.545-551
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
오늘날 누구나 디지털 정보를 용이하게 활용할 수 있게 됨에 따라 디지털 증거의 채택이 증가되고 있다. 하지 만 다양한 음성 파일 편집 도구를 보급과 함께 정교한 편집 과정을 거친 음성 녹음 파일의 경우 위변조 진위 여부를 판 단하는 것은 사실상 불가능하다. 본 연구는 음성 녹음 파일에 삽입, 삭제, 연결 및 합성 편집 기술을 활용해 원본 파 일과 구별하기 어려운 위변조가 가능함을 증명하고자 한다. 본 연구는 위변조 된 음성 파일을 원본과 동일한 확장자 로 인코딩하는 작업을 통해 위변조 검출의 어려움을 제시한다. 또한 특징점이 발생한 실험에 한 하여 추가적으로 천 이대역의 삭제 및 2차 인코딩 작업을 수행할 경우 위변조 검출은 불가능함을 나타냈다. 이를 통해 본 연구는 음성 녹 음 파일을 디지털 증거로 채택하기 위한 더 엄격한 증거능력 판단 기준 수립에 공헌할 것으로 기대된다.
As digital information is readily available to everyone today, the adoption of digital evidence is increasing. However, it is virtually impossible to determine the authenticity of forgery in the case of a voice recording file that has gone through a sophisticated editing process along with the spread of various voice file editing tools. This study aims to prove that forgery, which is difficult to distinguish from the original file, is possible by using insertion, deletion, linking, and synthetic editing technologies in voice recording files. This study presents the difficulty of detecting forgery by encoding a forged voice file with the same extension as the original. In addition, it was shown that forgery detection is impossible if additional transition band deletion and secondary encoding are performed only for experiments in which features occurred. Through this, this study is expected to contribute to the establishment of more stringent evidence admissibility criteria for adopting voice recording files as digital evidence.
Speech Signal Analysis Using Concentrated Spectrogram Method
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.8 No.5 2015.05 pp.127-132
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
A new method of energy distribution estimation in the joint time-frequency domain using the Channelized Instantaneous Frequency (CIF) and Local Group Delay (LGD) is proposed. The signal energy distribution is estimated by discarding and displacement of energy parts. The signal energy leads to high concentrated distribution in the time-frequency domain due to the relocation of the CIF and LGD values. In addition to this, a channelized instantaneous bandwidth and local group duration are used to remove undesired energy part. The channelized instantaneous bandwidth and local group duration express a local stretching of the signal in frequency and time respectively. This method is being used for speech signal analysis.
[NRF 연계] 대한용접·접합학회 대한용접·접합학회지 Vol.39 No.2 2021.04 pp.198-205
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
An automated welding system is essential to ensure a stable and good welding quality and improve productivity in the gas metal arc welding (GMAW) process. Therefore, various studies have been conducted on the establishment of smart factories and the demand for good weldability in the fields of production and manufacturing. In shipbuilding welding and pipe welding, the uniformly generated back-bead is an important criterion for judging the mechanical properties and weldability of the welded structure, and is also an important factor that enables the realization of an automated welding system. Therefore, in this study, the welding current signal measured in real-time in the GMAW process was pre-processed by a short time Fourier transform (STFT) to obtain a time-frequency domain feature im?age (spectrogram). Based on this, a back-bead generation detection algorithm was developed. To accelerate the training speed of the proposed convolution neural network (CNN) model, we used non-saturating neurons and a highly efficient GPU implementation of the convolution operation. As a result of applying the proposed detection model to actual welding process, the detection accuracies with and without the back-bead regions were 95.8% and 94.2%, respectively, which confirmed the excellent classification performance for back-bead generation.
Determining Key Features of Recognition Korean Traditional Music Using Spectrogram
[Kisti 연계] 한국음향학회 한국음향학회지 Vol.24 No.e2 2005 pp.67-70
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
To realize a traditional music recognition system, some characteristics pertinent to Far East Asian music should be found. Using Spectrogram, some distinct attributes of Korean traditional music are surveyed. Frequency distribution, beat cycle and frequency energy intensity within samples have distinct characteristics of their own. Experiment is done for pre-experimentation to realize Korean traditional music recognition system. Using characteristics of Korean traditional music, $94.5\%$ of classification accuracy is acquired. As Korea, Japan and China have the same musical roots, both in instruments and playing style, analyzing Korean traditional music can be helpful in the understanding of Far East Asian traditional music.
Open and Short Circuit Switches Fault Detection of Voltage Source Inverter Using Spectrogram
[Kisti 연계] 대한전기학회 분과 Journal of international Conference on Electrical Machines and Systems Vol.3 No.2 2014 pp.190-199
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In the last years, fault problem in power electronics has been more and more investigated both from theoretical and practical point of view. The fault problem can cause equipment failure, data and economical losses. And the analyze system require to ensure fault problem and also rectify failures. The current errors on these faults are applied for identified type of faults. This paper presents technique to detection and identification faults in three-phase voltage source inverter (VSI) by using time-frequency distribution (TFD). TFD capable represent time frequency representation (TFR) in temporal and spectral information. Based on TFR, signal parameters are calculated such as instantaneous average current, instantaneous root mean square current, instantaneous fundamental root mean square current and, instantaneous total current waveform distortion. From on results, the detection of VSI faults could be determined based on characteristic of parameter estimation. And also concluded that the fault detection is capable of identifying the type of inverter fault and can reduce cost maintenance.
RTO Rotary Motor의 고장 예측을 위한 스펙트로그램 기반 학습 알고리즘 구현 - 진동 분석을 통한 알고리즘 연구 -
[NRF 연계] 사단법인 안전문화포럼 안전문화연구 Vol.44 2025.07 pp.13-23
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구의 목적은 RTO(Regenerative Thermal Oxidizer, 축열식 연소산화장치) 장비에서 사용되는 Rotary Motor의 고장을 정확하게 예측할 수 있는 AI 기반 진단 시스템을 개발하는 것이다. 이 시스템은 기계 상태 모니터링에 중요한 진동 데이터를 활용하였다. 이를 위해 RC 필터를 사용하여 컷오프 주파수가 1kHz인 진동 및 모터 각상 전류 데이터를 수집하였으며, 데이터 샘플링 주파수는 1kHz로 설정하였다. 수집된 데이터는 FFT(고속 푸리에 변환)를 통해 스펙트로그램으로 변환되어 모델의 입력으로 사용된다. 이 스펙트로그램을 이용해 ResNet 기반의 딥러닝 모델을 설계하고 학습시켜 모터의 상태를 예측한다. 모델의 성능은 정확도와 손실 등의 지표를 통해 평가된다. 실험 결과, 제안된 알고리즘은 진동 데이터를 효과적으로 분석하고 Rotary Motor의 상태를 예측할 수 있음을 보여주었으며, 이는 RTO 장비의 실시간 모니터링 및 예방 유지보수에 AI 기반 진단 시스템을 적용할 가능성을 시사하며, 800℃ 이상의 고온을 이용하는 RTO 장비의 인입, 배출 흐름을 제어하는 Rotary 모터의 고장을 예지하여 기존 현장 대비 안전성을 향상시킬 수 있다.
The purpose of this study is to develop an AI-based diagnostic system capable of accurately predicting failures of the rotary motor used in regenerative thermal oxidizer (RTO) equipment. The system utilizes vibration data, which is critical for machinery condition monitoring. To achieve this, vibration and phase current data of the motor were collected using an RC low-pass filter with a cutoff frequency of 1?kHz, and the data were sampled at a rate of 1?kHz. The collected data were transformed into spectrograms via Fast Fourier Transform (FFT) and used as input for the model. A deep learning model based on a ResNet architecture was designed and trained to predict motor conditions using the spectrograms. The model’s performance was evaluated using key metrics such as accuracy and loss. Experimental results demonstrate that the proposed algorithm effectively analyzes vibration signals and accurately predicts the rotary motor’s condition. This outcome suggests the potential application of AI-based diagnostic systems in real-time monitoring and preventive maintenance of RTO equipment. In particular, by enabling early fault detection of rotary motors?which control the inflow and outflow of gases in environments exceeding 800°C?the proposed system significantly improves safety compared to conventional field operations.
음성 감정 인식에서의 어텐션 노이즈 감소를 위한 CNN 기반의 Log-Mel 스펙트로그램 이미지 압축 기법
[Kisti 연계] 한국전기전자학회 Journal of IKEEE Vol.28 No.4 2024 pp.600-606
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문은 음성 감정 인식에서 log-Mel 스펙트로그램을 기반으로 한 이미지 압축 기법을 제안하고, 이 기법이 어텐션 메커니즘을 활용한 vision transformer 모델에서 성능 향상에 기여할 수 있음을 보인다. 특히, log-Mel 스펙트로그램은 음성 신호의 주파수 특성을 잘 포착하여 음성 감정 인식에 유용하게 사용되는데, 본 연구에서는 이 스펙트로그램을 이미지 형태로 처리하면서 발생할 수 있는 어텐션 노이즈를 효과적으로 감소시키는 방법을 제시한다. 핵심적인 아이디어는 CNN을 수평 커널로 사용하여 log-Mel 스펙트로그램 이미지의 해상도를 압축하고, 이를 통해 vision transformer 모델에서 중요한 패턴을 보다 효과적으로 학습하도록 돕는 것이다. 제안된 기법은 기존의 log-Mel 스펙트로그램을 128×1001 크기로 처리하고, 이 이미지를 128×129로 고정된 크기로 압축하면서 임의의 이미지 보간이 수행되도록 설계되었다. 이러한 전처리 과정은 모델이 음성 감정 인식에서 유용한 특징을 보다 잘 추출할 수 있도록 돕는다. 본 논문에서는 log-Mel 스펙트로그램의 주어진 특성에 맞게 CNN 기반의 압축 기법을 사용하여 스펙트로그램의 중요 정보를 보존하면서, vision transformer 모델의 어텐션 메커니즘에서 발생할 수 있는 노이즈를 최소화하는 방법을 제안한다. Crowd Sourced Emotional Multimodal Actors(CREMA) 데이터셋을 이용한 실험을 통해, 제안하는 기법이 86.83%의 정확도를 나타내어 기존의 방법들보다 음성 감정 인식에서 더 뛰어난 성능을 보임을 확인하였다.
This paper proposes convolutional neural networks (CNN) based log-Mel spectrogram image compression method for attention noise reduction in speech emotion recognition (SER) and demonstrates how this method can contribute to improved performance in vision transformer models utilizing attention mechanisms. log-Mel spectrograms, which effectively capture the frequency characteristics of speech signals, are commonly used in SER tasks. In this study, we present a method to reduce attention noise that may arise when processing these spectrograms as images. The core idea is to use a CNN with horizontal kernels to compress the resolution of log-Mel spectrogram images, thereby facilitating the vision transformer model's ability to learn important patterns more effectively. The proposed approach processes the original log-Mel spectrograms at a size of 128×1001 and compresses them into a fixed 128×129 resolution while performing random image interpolation. This preprocessing step aids the model in better extracting relevant features for emotion recognition. This paper propose how the CNN-based compression method preserves essential information from the log-Mel spectrograms while minimizing attention noise in the vision transformer model's attention mechanism. Through experiments using the Crowd Sourced Emotional Multimodal Actors (CREMA) dataset, the proposed method achieved an accuracy of 86.83%, demonstrating superior performance in speech emotion recognition compared to existing methods.
음질 및 속도 향상을 위한 선형 스펙트로그램 활용 Text-to-speech
[Kisti 연계] 한국음성학회 말소리와 음성과학 Vol.13 No.3 2021 pp.71-78
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
인공신경망에 기반한 대부분의 음성 합성 모델은 고음질의 자연스러운 발화를 생성하기 위해 보코더 모델을 사용한다. 보코더 모델은 멜 스펙트로그램 예측 모델과 결합하여 멜 스펙트로그램을 음성으로 변환한다. 그러나 보코더 모델을 사용할 경우에는 많은 양의 컴퓨터 메모리와 훈련 시간이 필요하며, GPU가 제공되지 않는 실제 서비스 환경에서 음성 합성이 오래 걸린다는 단점이 있다. 기존의 선형 스펙트로그램 예측 모델에서는 보코더 모델을 사용하지 않으므로 이 문제가 발생하지 않지만, 대신에 고품질의 음성을 생성하지 못한다. 본 논문은 뉴럴넷 기반 보코더를 사용하지 않으면서도 양질의 음성을 생성하는 Tacotron 2 & Transformer 기반의 선형 스펙트로그램 예측 모델을 제시한다. 본 모델의 성능과 속도 측정 실험을 진행한 결과, 보코더 기반 모델에 비해 성능과 속도 면에서 조금 더 우세한 점을 보였으며, 따라서 고품질의 음성을 빠른 속도로 생성하는 음성 합성 모델 연구의 발판 역할을 할 것으로 기대한다.
Most neural-network-based speech synthesis models utilize neural vocoders to convert mel-scaled spectrograms into high-quality, human-like voices. However, neural vocoders combined with mel-scaled spectrogram prediction models demand considerable computer memory and time during the training phase and are subject to slow inference speeds in an environment where GPU is not used. This problem does not arise in linear spectrogram prediction models, as they do not use neural vocoders, but these models suffer from low voice quality. As a solution, this paper proposes a Tacotron 2 and Transformer-based linear spectrogram prediction model that produces high-quality speech and does not use neural vocoders. Experiments suggest that this model can serve as the foundation of a high-quality text-to-speech model with fast inference speed.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.