년 - 년
An Infant Audio Classification Using Deep Learning Technology
대한산업경영학회 International Journal of Intelligent Technologies and Innovative Practices Vol. 1 No. 1 2026.01 pp.25-31
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
The integration of deep learning techniques in the field of audio signal processing has marked a significant leap forward in the capability to analyze and classify complex sounds, including the nuanced and information-rich cries of infants. Deep learning's promise in this domain lies in its potential to decipher the subtle cues contained within these cries, offering insights into an infant's health, emotional state, and developmental needs. This potential application stands at the intersection of technology and healthcare, promising to enhance our understanding and response to the needs of the youngest members of society. This study experimentally demonstrates that a convolutional neural network–based audio classification model effectively learns discriminative spectral and temporal features from audio signals. Experimental results show that the proposed convolutional neural networks architecture achieves significantly higher classification accuracy than traditional machine-learning baselines, particularly when trained on spectrogram-based representations. The findings confirm that deep learning models not only improve overall performance but also provide robust generalization across different audio classes and noisy conditions.
음성 기반 치매 조기 진단 연구의 확장 : Mel-Spectrogram 중심 접근의 한계와 언어적 특징 기반 다중모달 분석 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제10권 3호 2026.03 pp.740-754
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
고령화 사회의 가속화로 치매 조기 선별의 중요성이 강조되는 가운데, 음성 기반 인공지능은 비침습적·저비 용 대안으로 주목받고 있다. 선행 연구에서는 음성 신호를 멜 스펙트로그램(Mel-Spectrogram)으로 변환하여 CNN, ViT 모델 등에 적용하였으나, 분류 정확도가 약 61~62% 수준에 머물며 음향적 특징만으로는 치매 특유의 인지 저 하를 포착하는 데 구조적 한계가 있음을 확인하였다. 본 연구는 이러한 한계를 극복하기 위해 ADRESS-2020 데이 터셋을 기반으로 전사 텍스트에서 추출한 언어적 특징을 결합한 다중모달 분석 접근을 제안하였다. 연구 결과, 어휘 다양성, 문장 복잡도, 의미적 응집성 등 14개의 언어적 변수만으로도 교차검증 정확도 76.8%를 달성하며 선행 연구 의 음향 기반 모델 성능을 크게 상회하였다. 특히 특징 수 대비 성능 효율성 측면에서 언어적 특징은 고차원 딥러닝 특징보다 월등히 높은 수치를 기록하여, 소규모 의료 데이터 환경에서 특징의 질적 설계가 중요함을 입증하였다. 음 향과 언어 특징의 결합은 안정성과 분산 측면에서 가장 균형 잡힌 결과를 나타냈으나, 모든 특징을 결합한 고차원 환경에서는 차원의 저주로 인한 성능 저하가 관찰되었다. 모델 비교에서는 로지스틱 회귀가 가장 우수한 일반화 성 능을 보였으며, 이는 실제 임상 현장에서 해석 가능하고 단순한 모델의 실용성이 높음을 시사한다. 본 연구는 언어적 특징 중심의 다중모달 분석이 치매 조기 선별의 정확성과 신뢰성을 높이는 핵심 전략임을 실증하였다.
As the acceleration of population aging intensifies the importance of early dementia screening, voice-based artificial intelligence is gaining significant attention as a non-invasive and cost-effective alternative. A previous study utilized voice signals converted into Mel-spectrograms and applied them to models such as CNN and ViT, but found that classification accuracy remained at approximately 61–62%, confirming structural limitations in capturing dementia-specific cognitive decline using only acoustic features . To overcome these limitations, this study proposes a multimodal analysis approach that integrates linguistic features extracted from transcribed text using the ADRESS-2020 dataset. The experimental results demonstrated that just 14 linguistic variables—including lexical diversity, syntactic complexity, and semantic coherence—achieved a cross-validation accuracy of 76.8%, significantly outperforming the acoustic-based models from the previous study. In terms of performance efficiency relative to the number of features, linguistic features recorded substantially higher values than high-dimensional deep learning features, proving that qualitative feature engineering is crucial in small-scale medical data environments. While the combination of acoustic and linguistic features yielded the most balanced results in terms of stability and variance, a performance decline due to the "curse of dimensionality" was observed when all high-dimensional features were combined. In model comparisons, logistic regression exhibited the most superior generalization performance, suggesting that simple, interpretable models are more practical for real-world clinical settings. This study empirically validates that multimodal analysis centered on linguistic features is a core strategy for enhancing the accuracy and reliability of early dementia screening.
Speech Signal Analysis Using Concentrated Spectrogram Method
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.8 No.5 2015.05 pp.127-132
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
A new method of energy distribution estimation in the joint time-frequency domain using the Channelized Instantaneous Frequency (CIF) and Local Group Delay (LGD) is proposed. The signal energy distribution is estimated by discarding and displacement of energy parts. The signal energy leads to high concentrated distribution in the time-frequency domain due to the relocation of the CIF and LGD values. In addition to this, a channelized instantaneous bandwidth and local group duration are used to remove undesired energy part. The channelized instantaneous bandwidth and local group duration express a local stretching of the signal in frequency and time respectively. This method is being used for speech signal analysis.
스마트폰 음성 녹음 파일 위변조 검출을 위한 스펙트로그램 분석의 한계점 KCI 등재
국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.9 No.2 2023.03 pp.545-551
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
오늘날 누구나 디지털 정보를 용이하게 활용할 수 있게 됨에 따라 디지털 증거의 채택이 증가되고 있다. 하지 만 다양한 음성 파일 편집 도구를 보급과 함께 정교한 편집 과정을 거친 음성 녹음 파일의 경우 위변조 진위 여부를 판 단하는 것은 사실상 불가능하다. 본 연구는 음성 녹음 파일에 삽입, 삭제, 연결 및 합성 편집 기술을 활용해 원본 파 일과 구별하기 어려운 위변조가 가능함을 증명하고자 한다. 본 연구는 위변조 된 음성 파일을 원본과 동일한 확장자 로 인코딩하는 작업을 통해 위변조 검출의 어려움을 제시한다. 또한 특징점이 발생한 실험에 한 하여 추가적으로 천 이대역의 삭제 및 2차 인코딩 작업을 수행할 경우 위변조 검출은 불가능함을 나타냈다. 이를 통해 본 연구는 음성 녹 음 파일을 디지털 증거로 채택하기 위한 더 엄격한 증거능력 판단 기준 수립에 공헌할 것으로 기대된다.
As digital information is readily available to everyone today, the adoption of digital evidence is increasing. However, it is virtually impossible to determine the authenticity of forgery in the case of a voice recording file that has gone through a sophisticated editing process along with the spread of various voice file editing tools. This study aims to prove that forgery, which is difficult to distinguish from the original file, is possible by using insertion, deletion, linking, and synthetic editing technologies in voice recording files. This study presents the difficulty of detecting forgery by encoding a forged voice file with the same extension as the original. In addition, it was shown that forgery detection is impossible if additional transition band deletion and secondary encoding are performed only for experiments in which features occurred. Through this, this study is expected to contribute to the establishment of more stringent evidence admissibility criteria for adopting voice recording files as digital evidence.
韓國人 日本語學習者의 撥音/ɴ/에 대한 音聲實現 - 스펙트로그램에 의한 分析 -
[NRF 연계] 한국일본어학회 일본어학연구 Vol.76 2023.06 pp.117-132
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this study, the substance and trend of pronunciation were examined from an acoustic phonetics point of view on how Korean Japanese learners generate various nasal sounds for Japanese Hatsuon /?/. Results are summarized as follows. First, in the case of native Japanese speakers, it was confirmed that the characteristics of each nasal sound for Japanese Hatsuon /?/ were basically in three Formants that were similar to vowels. In particular, looking at the results of the F2, F3, and F4, we could see a difference in the frequency for each nasal sound depending on the articulation position of the subsequent sound, but the deviation was not as large as that of Korean Japanese learners. This is thought to be the result of proving that the articulation position is not an absolute condition for pronunciation to produce nasal sounds. Some of the Korean Japanese learners were also able to get a glimpse of the case in which a nasal sound Formant similar to the results of the native Japanese speaker was realized. However, on the other hand, there were many cases of nasal sounds with acoustic characteristics without the F2 or F3 under the influence of Korean language. The absence of the F2 is presumed to be the result of the phonetic characteristic that the preceding vowel is a narrow vowel ([i]), as can be seen in‘シンブン([?imb??])’ or ‘シンパイ([?impai])’. From these results, it can be pointed out that the speech realization of nasal sounds by Korean Japanese learners is somewhat different from that of native Japanese speakers. It is thought that the cause of the difference is not in the articulation position of the subsequent sound of the Hatsuon /?/, but in the phonetic characteristics of the syllable (Mora) that is immediately preceded and subsequent. Field : Phonetics
광역 스펙트로그램과 심층신경망에 기반한 중첩된 소리의 인식과 영향 분석
[Kisti 연계] 한국방송공학회 방송공학회논문지 Vol.23 No.3 2018 pp.421-430
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
많은 음성인식 시스템들은 MFCC와 HMM등의 분류 기법을 사용하여 사람의 음성을 인식한다. 그러나 이러한 음성인식 시스템은 단일 음성신호를 인식하는 것을 목적으로 설계되어, 인간과 기계사이의 일대일 음성 인식에는 적합하나, 애완동물 소리와 실내 소리같은 음성보다 다양하고 넓은 주파수의 소리 군으로 중첩된 음향 속에서 설정된 소리를 인식하기에는 제한이 있다. 중첩된 소리들의 주파수는 사람의 목소리보다 높은 최대 20 kHz까지 넓은 주파수 범위로 구성된다. 본 논문에서는 광역 사운드 스펙트로그램과 DNN에 기반한 케라스 시?셜 모델 기법을 활용하여 인지 주파수 범위를 넓게 확대하는 새로운 인식방법을 제안한다. 광역 사운드 스펙트로그램이 본 논문에서 설계된 특징 추출 및 분류 시스템과 같이 넓은 주파수 범위의 다양한 소리를 분석하고 실험하도록 채택되었다. 소리 인식률을 개선하기 위하여, 케라스 시?셜 모델이 사운드 스펙트로그램에 의하여 생성되어 추출된 특징을 사용하여 패턴인식을 수행하기 위한 방법으로 채용되었다. 제안된 특징 추출 및 분류 시스템이 광역 사운드 스펙트로그램과 케라스 시?셜 모델을 채용하여 애완동물 소리와 실내 소리같은 다양한 주파수들로 구성되어 중첩된 음향 속에서 설정된 소리를 우수하게 분류하는 것을 확인하였다. 그리고 중첩된 소리의 크기에 비례하여 인식에 미치는 특성과 영향을 단계별로 비교 분석하였다.
Many voice recognition systems use methods such as MFCC, HMM to acknowledge human voice. This recognition method is designed to analyze only a targeted sound which normally appears between a human and a device one. However, the recognition capability is limited when there is a group sound formed with diversity in wider frequency range such as dog barking and indoor sounds. The frequency of overlapped sound resides in a wide range, up to 20KHz, which is higher than a voice. This paper proposes the new recognition method which provides wider frequency range by conjugating the Wideband Sound Spectrogram and the Keras Sequential Model based on DNN. The wideband sound spectrogram is adopted to analyze and verify diverse sounds from wide frequency range as it is designed to extract features and also classify as explained. The KSM is employed for the pattern recognition using extracted features from the WSS to improve sound recognition quality. The experiment verified that the proposed WSS and KSM excellently classified the targeted sound among noisy environment; overlapped sounds such as dog barking and indoor sounds. Furthermore, the paper shows a stage by stage analyzation and comparison of the factors' influences on the recognition and its characteristics according to various levels of noise.
RTO Rotary Motor의 고장 예측을 위한 스펙트로그램 기반 학습 알고리즘 구현 - 진동 분석을 통한 알고리즘 연구 -
[NRF 연계] 사단법인 안전문화포럼 안전문화연구 Vol.44 2025.07 pp.13-23
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구의 목적은 RTO(Regenerative Thermal Oxidizer, 축열식 연소산화장치) 장비에서 사용되는 Rotary Motor의 고장을 정확하게 예측할 수 있는 AI 기반 진단 시스템을 개발하는 것이다. 이 시스템은 기계 상태 모니터링에 중요한 진동 데이터를 활용하였다. 이를 위해 RC 필터를 사용하여 컷오프 주파수가 1kHz인 진동 및 모터 각상 전류 데이터를 수집하였으며, 데이터 샘플링 주파수는 1kHz로 설정하였다. 수집된 데이터는 FFT(고속 푸리에 변환)를 통해 스펙트로그램으로 변환되어 모델의 입력으로 사용된다. 이 스펙트로그램을 이용해 ResNet 기반의 딥러닝 모델을 설계하고 학습시켜 모터의 상태를 예측한다. 모델의 성능은 정확도와 손실 등의 지표를 통해 평가된다. 실험 결과, 제안된 알고리즘은 진동 데이터를 효과적으로 분석하고 Rotary Motor의 상태를 예측할 수 있음을 보여주었으며, 이는 RTO 장비의 실시간 모니터링 및 예방 유지보수에 AI 기반 진단 시스템을 적용할 가능성을 시사하며, 800℃ 이상의 고온을 이용하는 RTO 장비의 인입, 배출 흐름을 제어하는 Rotary 모터의 고장을 예지하여 기존 현장 대비 안전성을 향상시킬 수 있다.
The purpose of this study is to develop an AI-based diagnostic system capable of accurately predicting failures of the rotary motor used in regenerative thermal oxidizer (RTO) equipment. The system utilizes vibration data, which is critical for machinery condition monitoring. To achieve this, vibration and phase current data of the motor were collected using an RC low-pass filter with a cutoff frequency of 1?kHz, and the data were sampled at a rate of 1?kHz. The collected data were transformed into spectrograms via Fast Fourier Transform (FFT) and used as input for the model. A deep learning model based on a ResNet architecture was designed and trained to predict motor conditions using the spectrograms. The model’s performance was evaluated using key metrics such as accuracy and loss. Experimental results demonstrate that the proposed algorithm effectively analyzes vibration signals and accurately predicts the rotary motor’s condition. This outcome suggests the potential application of AI-based diagnostic systems in real-time monitoring and preventive maintenance of RTO equipment. In particular, by enabling early fault detection of rotary motors?which control the inflow and outflow of gases in environments exceeding 800°C?the proposed system significantly improves safety compared to conventional field operations.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.