Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 17
No
1

환경음의 가시화를 통한 도시공간의 기술 KCI 등재

민건희, 윤철재

대한건축학회지회연합회 대한건축학회연합논문집 제14권 제1호 통권 49호 2012.03 pp.143-150

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

건축 및 도시분야에서 시각정보는 지배적인 공간개념으로 다루어지는 반면 그밖의 감각들은 비교적 경시되어 왔다. 그러나 실제적인 공간체험은 시각이외에도 다양한 감각을 통해 이루어지며 특히 청각정보는 정보의 근원이 되는 공간성격의 파악에 큰 영향을 끼친다. 본 논문은 청각정보가 지니는 실시간적, 확산적 성질에 착안, 실제 도시공간의 다양한 소리를 가시화함으로써 도시공간의 이질적이고 다양한 현상을 소리라는 매체로 표현하고 그 분포의 변화를 통해 시간적으로 변화하는 공간의 양상을 동적으로 파악하는 기술법의 제안을 목표로 했다. 본 논문은 4단계로 구성된다. 1)분석대상지를 대상으로 한 시간대별 녹음조사 2)녹음된 청각데이터로부터 구별할 수 있는 모든 소리에 대해 의미론적으로 분류 3)청각정보의 분포를 가시화시켜주는 소프트웨어의 제작 및 가시화된 소리의 분포와 그 변화의 구현 4)작성된 사운드 이벤트 맵을 토대로 대상지역의 동적으로 변화하는 공간의 양상을 분석

The visual information dominates as a primary spatial concept in architectural and urban field while other senses have been relatively neglected. However in reality, human being experiences the space through various senses besides visual sense. Among all the other senses, particularly, auditory information has significant influences on the understanding environment where it comes from. This paper draws attention to the real-time and extensive feature of the sound in the space, and visualizes a variety of sounds in actual urban environment. This enables to express heterogeneous and diverse phenomena with the sound as a integrated parameter. Also the study aims at the suggestion of the method understanding the ever-changing condition of urban environment through sound event distribution change. To achieve the goal the study was done in four steps as follows; 1) Sound recording at Kichijoji and Inokashira Park neighborhood in different date and time. 2) Semantic classification of all distinguishable sounds from the recorded data 3) Programming software which visualizes the distribution and indicating visualized sound distribution and its change 4) Analyzing dynamically changing condition of the site based on drawn sound event map.

2

4,500원

본 연구에서는 CNN-DNN 기반의 음향인지 알고리즘을 제안하였다. 입력 볼륨에 따른 영향을 줄이기 위하여 Peak normalization을 적용하였으며, Log mel filter bank 특징벡터에 CNN을 적용하여 음향 분류에 적합한 특성을 추출하고 음향 데이터의 시계열적 특성을 고려할 수 있도록 연속되는 프레임들에서의 CNN 출력을 결합하여 DNN으로 입력하는 방법을 사용하였다. 공개 데이터가 아닌 실제 아파트 환경에서 샷 건 마이크를 이용하여 직접 수집한 데이터를 기반으로 실험을 진행하였으며, 음향의 종류는 홈 IoT(Internet of Things) 환경에서 발생할 수 있는 초인종, 음악, 도어락 등 13가지 음향을 고려하였다. 컴퓨터 시뮬레이션 평가 결과 90.76% 정확도를 얻었다.

In this study, we proposed a CNN-DNN based sound event detection in home IoT environments. To reduce the impact of input volume variation, we applied peak normalization, and extracted acoustic feature named Log mel filter bank. Log-mel filter bank is very popular acoustic feature based on mel filter which is powerful for speech recognition and sound event detection. Then, we used CNN-DNN model for classification. CNN outputs of sequential 32 frames were used as DNN input for considering time-series characteristic of the sound. Data were collected in real apartment environment. We used 13 sounds as target such as doorbell, babycry, vacuum, and so on. We evaluated our method using computer simulation, as a result, the accuracy of the proposed sound event detection algorithm was 90.76%.

3

CNN based Sound Event Detection Method using NMF Preprocessing in Background Noise Environment KCI 등재

Bumsuk Jang, Sang-Hyun Lee

국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 9 Number 2 2020.06 pp.20-27

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Sound event detection in real-world environments suffers from the interference of non-stationary and time-varying noise. This paper presents an adaptive noise reduction method for sound event detection based on non-negative matrix factorization (NMF). In this paper, we proposed a deep learning model that integrates Convolution Neural Network (CNN) with Non-Negative Matrix Factorization (NMF). To improve the separation quality of the NMF, it includes noise update technique that learns and adapts the characteristics of the current noise in real time. The noise update technique analyzes the sparsity and activity of the noise bias at the present time and decides the update training based on the noise candidate group obtained every frame in the previous noise reduction stage. Noise bias ranks selected as candidates for update training are updated in real time with discrimination NMF training. This NMF was applied to CNN and Hidden Markov Model(HMM) to achieve improvement for performance of sound event detection. Since CNN has a more obvious performance improvement effect, it can be widely used in sound source based CNN algorithm.

4

향상된 음향 신호 기반의 음향 이벤트 분류

최용주, 이종욱, 박대희, 정용화

[Kisti 연계] 한국정보처리학회 정보처리학회논문지/소프트웨어 및 데이터 공학 Vol.8 No.5 2019 pp.193-204

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

센서 기술과 컴퓨팅 성능의 향상으로 인한 데이터의 폭증은 산업 현장의 상황을 분석하기 위한 토대가 되었으며, 이와 같은 데이터를 기반으로 현장에서 발생하는 다양한 이벤트를 탐지 및 분류하려는 시도들이 최근 증가하고 있다. 특히 음향 센서는 상대적으로 저가의 가격으로 현장 정보를 왜곡 없이 음향 신호를 수집할 수 있다는 큰 장점을 기반으로 다양한 분야에 설치되고 있다. 그러나 소리 취득 시 발생하는 잡음을 효과적으로 제어하지 못한다면 산업 현장의 이벤트를 안정적으로 분류할 수 없으며, 분류하지 못한 이벤트가 이상 상황이라면 이로 인한 피해는 막대해질 수 있다. 본 연구에서는 잡음 상황에서도 강인한 시스템을 보장하기 위하여, 딥러닝 알고리즘을 기반으로 잡음의 영향을 개선 시킨 음향 신호를 생성한 후, 해당 음향 이벤트를 분류할 수 있는 시스템을 제안한다. 특히, GAN을 기반으로 VAE 기술을 적용한 SEGAN을 활용하여 아날로그 음향 신호 자체에서 잡음이 제거된 신호를 생성하였으며, 향상된 음향 신호를 데이터 변환과정 없이 CNN 구조의 입력 데이터로 활용한 후 음향 이벤트에 대한 식별까지도 가능하도록 end-to-end 기반의 음향 이벤트 분류 시스템을 설계하였다. 산업 현장에서 취득한 음향 데이터를 활용하여 제안하는 시스템의 성능을 실험적으로 검증한바, 99.29%(철도산업)와 97.80%(축산업)의 안정적인 분류 성능을 확인하였다.

The explosion of data due to the improvement of sensor technology and computing performance has become the basis for analyzing the situation in the industrial fields, and various attempts to detect events based on such data are increasing recently. In particular, sound signals collected from sensors are used as important information to classify events in various application fields as an advantage of efficiently collecting field information at a relatively low cost. However, the performance of sound-event classification in the field cannot be guaranteed if noise can not be removed. That is, in order to implement a system that can be practically applied, robust performance should be guaranteed even in various noise conditions. In this study, we propose a system that can classify the sound event after generating the enhanced sound signal based on the deep learning algorithm. Especially, to remove noise from the sound signal itself, the enhanced sound data against the noise is generated using SEGAN applied to the GAN with a VAE technique. Then, an end-to-end based sound-event classification system is designed to classify the sound events using the enhanced sound signal as input data of CNN structure without a data conversion process. The performance of the proposed method was verified experimentally using sound data obtained from the industrial field, and the f1 score of 99.29% (railway industry) and 97.80% (livestock industry) was confirmed.

5

잡음 학생 모델 기반의 자가 학습을 활용한 음향 사건 검지

김남균, 박창수, 김홍국, 허진욱, 임정은

[Kisti 연계] 한국음향학회 한국음향학회지 Vol.40 No.5 2021 pp.479-487

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 잡음 학생 모델 기반의 자가 학습을 활용한 음향 사건 검지 기법을 제안한다. 제안된 음향 사건 검지 모델은 두 단계로 구성된다. 첫 번째 단계에서는 잔차 합성곱 순환 신경망(Residual Convolutional Recurrent Neural Network, RCRNN)을 훈련하여 레이블이 지정되지 않은 비표기 데이터셋의 레이블 예측에 활용한다. 두 번째 단계에서는 세 가지 잡음 종류를 적용한 잡음 학생 모델을 자가학습 기법으로 반복하여 학습한다. 여기서 잡음 학생 모델은 SpecAugment, Mixup, 시간-주파수 이동을 활용한 특징 잡음, 드롭아웃을 활용한 모델 잡음, 그리고 semi-supervised loss function을 적용한 레이블 잡음을 활용하여 학습된다. 제안된 음향 사건 검지 모델의 성능은 Detection and Classification of Acoustic Scenes and Events(DCASE) 2020 Challenge Task 4의 validation set으로 평가하였다. DCASE 2020 챌린지 데이터셋의 baseline 및 최상위 랭크된 모델과 이벤트 단위 F1 점수 성능을 비교한 결과, 제안된 음향 사건 검지 모델이 단일 모델과 앙상블 모델에서 최상위 모델 대비 F1 점수를 각각 4.6 %와 3.4 % 향상시켰다.

In this paper, we propose an Sound Event Detection (SED) model using self-training based on a noisy student model. The proposed SED model consists of two stages. In the first stage, a mean-teacher model based on an Residual Convolutional Recurrent Neural Network (RCRNN) is constructed to provide target labels regarding weakly labeled or unlabeled data. In the second stage, a self-training-based noisy student model is constructed by applying different noise types. That is, feature noises, such as time-frequency shift, mixup, SpecAugment, and dropout-based model noise are used here. In addition, a semi-supervised loss function is applied to train the noisy student model, which acts as label noise injection. The performance of the proposed SED model is evaluated on the validation set of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge Task 4. The experiments show that the single model and ensemble model of the proposed SED based on the noisy student model improve F1-score by 4.6 % and 3.4 % compared to the top-ranked model in DCASE 2020 challenge Task 4, respectively.

6

딥 뉴럴네트워크 기반의 소리 이벤트 검출

정석환, 정용주

[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.14 No.2 2019 pp.389-396

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 다양한 구조의 딥 뉴럴 네트워크를 소리 이벤트 검출을 위하여 적용하였으며 공통의 오디오 데이터베이스를 이용하여 그들 간의 성능을 비교하였다. FNN, CNN, RNN 그리고 CRNN이 주어진 오디오데이터베이스 및 딥 뉴럴 네트워크의 구조에 최적화된 하이퍼파라미터 값을 이용하여 구현되었다. 구현된 방식 중에서 CRNN이 모든 테스트 환경에서 가장 좋은 성능을 보였으며 그 다음으로 CNN의 성능이 우수함을 알 수 있었다. RNN은 오디오 신호에서의 시간 상관관계를 잘 추적하는 장점에도 불구하고 CNN 과 CRNN에 비해서 저조한 성능을 보임을 확인할 수 있었다.

In this paper, various architectures of deep neural networks were applied for sound event detection and their performances were compared using a common audio database. The FNN, CNN, RNN and CRNN were implemented using hyper-parameters optimized for the database as well as the architecture of each neural network. Among the implemented deep neural networks, CRNN performed best at all testing conditions and CNN followed CRNN in performance. Although RNN has a merit in tracking the time-correlations in audio signals, it showed poor performance compared with CNN and CRNN.

7

깊은 신경망 기반의 전이학습을 이용한 사운드 이벤트 분류

임형준, 김명종, 김회린

[Kisti 연계] 한국음향학회 한국음향학회지 Vol.35 No.2 2016 pp.143-148

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

깊은 신경망은 데이터의 특성을 효과적으로 나타낼 수 있는 방법으로 최근 많은 응용 분야에서 활용되고 있다. 하지만, 제한적인 양의 데이터베이스는 깊은 신경망을 훈련하는 과정에서 과적합 문제를 야기할 수 있다. 본 논문에서는 풍부한 양의 음성 혹은 음악 데이터를 이용한 전이학습을 통해 제한적인 양의 사운드 이벤트에 대한 깊은 신경망을 효과적으로 훈련하는 방법을 제안한다. 일련의 실험을 통해 제안하는 방법이 적은 양의 사운드 이벤트 데이터만으로 훈련된 깊은 신경망에 비해 현저한 성능 향상이 있음을 확인하였다.

Deep neural network that effectively capture the characteristics of data has been widely used in various applications. However, the amount of sound database is often insufficient for learning the deep neural network properly, so resulting in overfitting problems. In this paper, we propose a transfer learning framework that can effectively train the deep neural network even with insufficient sound event data by employing rich speech or music data. A series of experimental results verify that proposed method performs significantly better than the baseline deep neural network that was trained only with small sound event data.

8

다채널 오디오 특징값 및 게이트형 순환 신경망을 사용한 다성 사운드 이벤트 검출

고상선, 조혜승, 김형국

[Kisti 연계] 한국음향학회 한국음향학회지 Vol.36 No.4 2017 pp.267-272

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 다채널 오디오 특징값을 게이트형 순환 신경망(Gated Recurrent Neural Networks, GRNN)에 적용한 효과적인 다성 사운드 이벤트 검출 방식을 제안한다. 실생활의 사운드는 여러 사운드 이벤트가 겹쳐있는 다성사운드로, 기존의 단일 채널 오디오 특징값으로는 다성 사운드에서 개별적인 이벤트의 검출이 어렵다는 한계가 있다. 이에 본 논문에서는 다채널 오디오 신호를 기반으로 추출된 특징값을 사용하여 다성 사운드 이벤트 검출에 적용하였다. 또한 본 논문에서는 현재 순환 신경망에서 가장 높은 성능을 보이는 장단기 기억 신경망(Long Short Term Memory, LSTM) 보다 간단한 GRNN을 분류에 적용하여 다성 사운드 이벤트 검출의 성능을 더욱 향상시키고자 하였다. 실험결과는 본 논문에서 제안한 방식이 기존의 방식보다 성능이 더 뛰어나다는 것을 보인다.

In this paper, we propose an effective method of applying multichannel-audio feature values to GRNNs (Gated Recurrent Neural Networks) in polyphonic sound event detection. Real life sounds are often overlapped with each other, so that it is difficult to distinguish them by using a mono-channel audio features. In the proposed method, we tried to improve the performance of polyphonic sound event detection by using multi-channel audio features. In addition, we also tried to improve the performance of polyphonic sound event detection by applying a gated recurrent neural network which is simpler than LSTM (Long Short Term Memory), which shows the highest performance among the current recurrent neural networks. The experimental results show that the proposed method achieves better sound event detection performance than other existing methods.

9

K-SVD 기반 사전 훈련과 비음수 행렬 분해 기법을 이용한 중첩음향이벤트 검출

최현식, 금민석, 고한석

[Kisti 연계] 한국음향학회 한국음향학회지 Vol.34 No.3 2015 pp.234-239

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

비음수 행렬 분해(Nonnegative Matrix Factorization, NMF) 기법은 사전행렬과 크기성분을 번갈아 가며 업데이트 하면서 구하는 방법이며 직관적 해석 및 구현의 용이성으로 인해 중첩음향이벤트 분리 및 검출방법으로 널리 활용되었다. 하지만 비음수 행렬 분해의 고유한 특성인 부분기반표현(part-based representation)으로 인해 하나의 음향 이벤트를 구성 하는 사전(dictionary)의 파편화 현상이 발생하고, 다른 음향이벤트와 중복되는 사전이 생성되어 결과적으로 분리, 검출 성능의 저하 문제가 발생한다. 본 논문에서는 사전 획득 단계의 부분기반표현에 의한 문제를 해소하기 위해 K-Singular Value Decomposition(K-SVD)을 사용하여 사전을 획득하고, 음향이벤트 검출 단계 에서는 기존 비음수 행렬 분해 기법을 이용하여 크기를 획득 한다. 제안하는 방식을 통해 비음수 행렬 분해 기반의 사전을 사용하는 경우보다 중첩음향이벤트 검출 성능이 개선되는 것을 확인하였다.

Non-Negative Matrix Factorization (NMF) is a method for updating dictionary and gain in alternating manner. Due to ease of implementation and intuitive interpretation, NMF is widely used to detect and separate overlapping sound events. However, NMF that utilizes non-negativity constraints generates parts-based representation and this distinct property leads to a dictionary containing fragmented acoustic events. As a result, the presence of shared basis results in performance degradation in both separation and detection tasks of overlapping sound events. In this paper, we propose a new method that utilizes K-Singular Value Decomposition (K-SVD) based dictionary to address and mitigate the part-based representation issue during the dictionary learning step. Subsequently, we calculate the gain using NMF in sound event detection step. We evaluate and confirm that overlapping sound event detection performance of the proposed method is better than the conventional method that utilizes NMF based dictionary.

10

심층신경망을 이용한 시간 영역 음향 이벤트 검출 알고리즘

김범준, 문현기, 박성욱, 정영호, 박영철

[Kisti 연계] 한국방송공학회 방송공학회논문지 Vol.24 No.3 2019 pp.472-484

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 심층신경망을 이용한 시간 영역 음향 이벤트 검출 알고리즘을 제시한다. 본 시스템에서는 주파수 영역으로 변환되지 않은 시간 영역의 음향 데이터를 심층신경망의 입력으로 사용한다. 전반적인 구조는 CRNN 구조를 사용하였으며, GLU, ResNet, Squeeze-and-excitation 블럭을 적용하였다. 그리고 여러 계층에서 추출된 특징을 함께 고려하는 구조를 제안하였다. 또한 본 연구에서는 강한 라벨이 있는 훈련 데이터를 확보하는 것이 현실적으로 어렵다는 전제 아래에서 약한 라벨이 있는 훈련 데이터 약간 그리고 다수의 라벨이 없는 훈련 데이터를 활용하여 훈련을 수행하였다. 적은 수의 훈련 데이터를 효과적으로 사용하기 위해 타임 스트레칭, 피치 변화, 동적 영역 압축, 블럭 혼합 등의 데이터 증강 방법을 적용하였다. 라벨이 없는 데이터에는 의사 라벨을 붙여 부족한 훈련 데이터를 보완하였다. 본 논문에서 제안한 신경망과 데이터 증강 방법을 사용하는 경우, 종래의 방식으로 CRNN 구조의 신경망을 훈련하여 사용하는 경우보다, 음향 이벤트 검출 성능이 약 6 % (f-score 기준)가 개선되었다.

This paper proposes a time-domain sound event detection algorithm using DNN (Deep Neural Network). In this system, time domain sound waveform data which is not converted into the frequency domain is used as input to the DNN. The overall structure uses CRNN structure, and GLU, ResNet, and Squeeze-and-excitation blocks are applied. And proposed structure uses structure that considers features extracted from several layers together. In addition, under the assumption that it is practically difficult to obtain training data with strong labels, this study conducted training using a small number of weakly labeled training data and a large number of unlabeled training data. To efficiently use a small number of training data, the training data applied data augmentation methods such as time stretching, pitch change, DRC (dynamic range compression), and block mixing. Unlabeled data was supplemented with insufficient training data by attaching a pseudo-label. In the case of using the neural network and the data augmentation method proposed in this paper, the sound event detection performance is improved by about 6 %(based on the f-score), compared with the case where the neural network of the CRNN structure is used by training in the conventional method.

11

실생활 음향 데이터 기반 이중 CNN 구조를 특징으로 하는 음향 이벤트 인식 알고리즘

서상원, 임우택, 정영호, 이태진, 김휘용

[Kisti 연계] 한국방송공학회 방송공학회논문지 Vol.23 No.6 2018 pp.855-865

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

음향 이벤트 인식은 다수의 음향 이벤트가 발생하는 환경에서 이를 인식하고 각각의 발생과 소멸 시점을 판단하는 기술로써 인간의 청각적 인지 특성을 모델화하는 연구다. 음향 장면 및 이벤트 인식 연구 그룹인 DCASE는 연구자들의 참여 유도와 더불어 음향 인식 연구의 활성화를 위해 챌린지를 진행하고 있다. 그러나 DCASE 챌린지에서 제공하는 데이터 세트는 이미지 인식 분야의 대표적인 데이터 세트인 이미지넷에 비해 상대적으로 작은 규모이며, 이 외에 공개된 음향 데이터 세트는 많지 않아 알고리즘 개발에 어려움이 있다. 본 연구에서는 음향 이벤트 인식 기술 개발을 위해 실내외에서 발생할 수 있는 이벤트를 정의하고 수집을 진행하였으며, 보다 큰 규모의 데이터 세트를 확보하였다. 또한, 인식 성능 개선을 위해 음향 이벤트 존재 여부를 판단하는 보조 신경망을 추가한 이중 CNN 구조의 알고리즘을 개발하였고, 2016년과 2017년의 DCASE 챌린지 기준 시스템과 성능 비교 실험을 진행하였다.

Sound event detection is one of the research areas to model human auditory cognitive characteristics by recognizing events in an environment with multiple acoustic events and determining the onset and offset time for each event. DCASE, a research group on acoustic scene classification and sound event detection, is proceeding challenges to encourage participation of researchers and to activate sound event detection research. However, the size of the dataset provided by the DCASE Challenge is relatively small compared to ImageNet, which is a representative dataset for visual object recognition, and there are not many open sources for the acoustic dataset. In this study, the sound events that can occur in indoor and outdoor are collected on a larger scale and annotated for dataset construction. Furthermore, to improve the performance of the sound event detection task, we developed a dual CNN structured sound event detection system by adding a supplementary neural network to a convolutional neural network to determine the presence of sound events. Finally, we conducted a comparative experiment with both baseline systems of the DCASE 2016 and 2017.

12

음향 이벤트 검출을 위한 누적 특징 추출 네트워크

박상원, 박상욱

[Kisti 연계] 한국음향학회 한국음향학회지 Vol.44 No.5 2025 pp.489-495

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

음향 이벤트 검출은 오디오 신호에서 음향의 종류와 발생 지점과 끝점을 검출하는 기술로 모니터링 시스템, 자율주행 자동차 등 다양한 분야에 쓰이고 있다. 음향 이벤트 검출은 음향 신호 분석에 관한 국제 경연대회(Detection and Classification of Acoustic Scenes and Events, DCASE)를 통해 음향 이벤트 검출 성능을 향상시키기 위한 다양한 방법들이 소개되고 있다. 본 논문은 기존 음향 분석 모델의 하위 계층에서 시간-주파수 정보가 손실되는 문제를 완화하기 위해 누적 특성 추출 신경망(AccNet)을 제안한다. 제안하는 모델은 DCASE 2023 task4 테스트 베드를 활용한 실험에서, DCASE 2023 Baseline과 동일한 파라미터 수를 유지하면서 F1 점수에서 44.76 ± 0.51[%]를 기록하였고, 비교대상으로 고려된 CRNN, 다중해상도 합성곱 모델, 잔차 경로 기반 모델들에 비해 가장 우수한 성능을 보여준다. 또한, 제안하는 모델은 잔차 경로 기반 모델에 비해 Blender와 Electric shaver를 제외한 관심 음향에서 향상된 F1 점수를 보여준다.

Sound event detection is a technology that detects the type, onset, and offset of sound events in audio signals and it is used in various fields such as monitoring systems and autonomous vehicles. Through the international competition (Detection and Classification of Acoustic Scenes and Events, DCASE) on acoustic signal analysis , various methods have been introduced to improve the performance of sound event detection. In this paper, we propose AccNet to solve the loss of spectro-temporal information in low layers of conventional acoustic model. In experiments performed on the DCASE 2023 Task 4 testbed, while the proposed model is comparable to the DCASE 2023 baseline in model complexity, it achieved 44.76 ± 0.51 % in event based f1 score, the best performance compared to the other models such as CRNN, multi-resolutional convolution based model, and residual path based model. Also the proposed model demonstrates improved f1 scores for target sound event except Blender and Electric shaver, compared to the residual path based model.

13

평균-교사 합성곱 순환 신경망 모델을 이용한 약지도 음향 이벤트 검출 시스템의 성능 분석

이석진

[Kisti 연계] 한국음향학회 한국음향학회지 Vol.40 No.2 2021 pp.139-147

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문은 데이터의 일부만 레이블링이 되어있는 약지도 학습을 기반으로 하는 음향 이벤트 검출 시스템을 소개 및 구현하고, 시뮬레이션을 통해 각 파라미터가 성능에 미치는 영향을 분석하였다. 음향 이벤트 검출 시스템은 음향 신호 내에 존재하는 이벤트의 종류, 시작/종료 시점을 추정하는 시스템으로, 이를 학습시키기 위해서는 음향 이벤트 신호와 그 종류, 시작/종료 시점에 대한 모든 정보가 제공되어야 한다. 하지만 이를 모두 표기하여 학습데이터를 만드는 것은 매우 큰 비용이 들어가며, 특히 시작/종료 시점을 정확히 표기하는 것은 매우 어렵다. 따라서 본 논문에서 다루는 약지도 학습 문제에서는 이벤트의 종류와 시작/종료 시점이 모두 표기된 "강하게 표기된 데이터"와, 이벤트의 종류만 표기된 "약하게 표기된 데이터", 그리고 아무런 표기가 되어 있지 않은 "미표기 데이터"를 이용하여 음향 이벤트 검출 시스템을 학습시킨다. 최근 이러한 문제에서는 평균-교사 모델을 이용한 음향 이벤트 검출 시스템의 성능이 우수하며, 따라서 널리 사용되고 있다. 다만, 평균-교사 모델은 많은 파라미터를 가지고 있고, 이는 성능에 영향을 다소 미칠 수 있으므로 신중하게 선택되어야 한다. 본 논문에서는 DCASE 2020 Task 4의 데이터를 이용하여 특징 값의 종류, 이동 평균 파라미터, 일관성 비용함수의 가중치, 램프-업 길이, 그리고 최대 학습율 등 5가지의 값에 대해 성능 분석을 진행하였으며, 각 파라미터에 대한 영향 및 최적 값에 대해 고찰하였다.

This paper introduces and implements a Sound Event Detection (SED) system based on weakly-supervised learning where only part of the data is labeled, and analyzes the effect of parameters. The SED system estimates the classes and onset/offset times of events in the acoustic signal. In order to train the model, all information on the event class and onset/offset times must be provided. Unfortunately, the onset/offset times are hard to be labeled exactly. Therefore, in the weakly-supervised task, the SED model is trained by "strongly labeled data" including the event class and activations, "weakly labeled data" including the event class, and "unlabeled data" without any label. Recently, the SED systems using the mean-teacher model are widely used for the task with several parameters. These parameters should be chosen carefully because they may affect the performance. In this paper, performance analysis was performed on parameters, such as the feature, moving average parameter, weight of the consistency cost function, ramp-up length, and maximum learning rate, using the data of DCASE 2020 Task 4. Effects and the optimal values of the parameters were discussed.

14

ResNet 모델을 이용한 일상생활 소리 예측 및 알림 애플리케이션

박유진, 정은이, 신지혜, 박태정, 양회석

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2020 pp.1004-1007

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 청각 장애인이 가정에서 듣지 못해 발생하는 낭비와 위험을 미리 예방하기 위하여 가정에서 현재 발생하고 있는 소리를 알려주는 시스템을 구현하였다. 무지향성 마이크로 일상 소리 감지 후 음향 데이터에서 Mel-Spectogram 특징 벡터를 추출하여 Convolutional Neural Network(CNN) 모델의 Resnet 알고리즘을 진행한다. 서버에서 소리에 대한 분석을 진행한 후 그 결과를 안드로이드에서 실시간으로 5 초마다 확인하여 사용자에게 알림 서비스를 제공한다. 이를 통해 낭비를 줄이고 위험에 대처할 수 있게 한다. 청각 장애인의 소리에 대한 접근성을 다양한 측면으로 고려해야 한다는 사회적 인식을 확산시키고자 한다.

15

청각장애인을 위한 사운드 이벤트 검출 기반 홈 모니터링 시스템

김지연, 신승수, 김형국

[Kisti 연계] 한국음향학회 한국음향학회지 Vol.38 No.4 2019 pp.427-432

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 청각장애인을 위해 양방향 게이트 순환 신경망을 이용한 사운드 이벤트 검출 기반의 홈 모니터링 시스템을 제안한다. 제안된 시스템에서는 우선적으로 효과적인 사운드 이벤트 검출을 위해 패킷손실 은닉을 이용하여 무선 센서 네트워크로 인해 손실된 신호를 복원하고, 멀티채널 상호 상관관계 계수를 이용하여 신뢰할 수 있는 채널을 선택한다. 선택된 채널의 사운드는 이벤트 검출을 위해 두 개의 오디오 채널을 사용하는 양방향 게이트 순환신경망에 적용된다. 검출된 사운드 이벤트는 텍스트로 변환되며, 이와 함께 하모닉/퍼커시브 음원 분리 방식을 통해 햅틱 신호로 변환되어 청각장애인에게 제공된다. 실험결과는 제안한 사운드 검출기반의 성능이 기존 방식보다 더 우수하다는 것과 음원 분리 방식을 통해 사운드를 세밀한 햅틱 신호로 표현할 수 있음을 보인다.

In this paper, we propose a home monitoring system using sound event detection based on a bidirectional gated recurrent neural network for the hard-of-hearing. First, in the proposed system, packet loss concealment is used to recover a lost signal captured through wireless sensor networks, and reliable channels are selected using multi-channel cross correlation coefficient for effective sound event detection. The detected sound event is converted into the text and haptic signal through a harmonic/percussive sound source separation method to be provided to hearing impaired people. Experimental results show that the performance of the proposed sound event detection method is superior to the conventional methods and the sound can be expressed into detailed haptic signal using the source separation.

16

음향 이벤트 검출을 위한 DenseNet-Recurrent Neural Network 학습 방법에 관한 연구

차현진, 박상욱

[Kisti 연계] 한국음향학회 한국음향학회지 Vol.42 No.5 2023 pp.395-401

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

음향 이벤트 검출(Sound Event Detection, SED)은 음향 신호에서 관심 있는 음향의 종류와 발생 구간을 검출하는 기술로, 음향 감시 시스템 및 모니터링 시스템 등 다양한 분야에서 활용되고 있다. 최근 음향 신호 분석에 관한 국제 경연 대회(Detection and Classification of Acoustic Scenes and Events, DCASE) Task 4를 통해 다양한 방법이 소개되고 있다. 본 연구는 다양한 영역에서 성능 향상을 이끌고 있는 Dense Convolutional Networks(DenseNet)을 음향 이벤트 검출에 적용하기 위해 설계 변수에 따른 성능 변화를 비교 및 분석한다. 실험에서는 DenseNet with Bottleneck and Compression(DenseNet-BC)와 순환신경망(Recurrent Neural Network, RNN)의 한 종류인 양방향 게이트 순환 유닛(Bidirectional Gated Recurrent Unit, Bi-GRU)을 결합한 DenseRNN 모델을 설계하고, 평균 교사 모델(Mean Teacher Model)을 통해 모델을 학습한다. DCASE task4의 성능 평가 기준에 따라 이벤트 기반 f-score를 바탕으로 설계 변수에 따른 DenseRNN의 성능 변화를 분석한다. 실험 결과에서 DenseRNN의 복잡도가 높을수록 성능이 향상되지만 일정 수준에 도달하면 유사한 성능을 보임을 확인할 수 있다. 또한, 학습과정에서 중도탈락을 적용하지 않는 경우, 모델이 효과적으로 학습됨을 확인할 수 있다.

Sound Event Detection (SED) aims to identify not only sound category but also time interval for target sounds in an audio waveform. It is a critical technique in field of acoustic surveillance system and monitoring system. Recently, various models have introduced through Detection and Classification of Acoustic Scenes and Events (DCASE) Task 4. This paper explored how to design optimal parameters of DenseNet based model, which has led to outstanding performance in other recognition system. In experiment, DenseRNN as an SED model consists of DensNet-BC and bi-directional Gated Recurrent Units (GRU). This model is trained with Mean teacher model. With an event-based f-score, evaluation is performed depending on parameters, related to model architecture as well as model training, under the assessment protocol of DCASE task4. Experimental result shows that the performance goes up and has been saturated to near the best. Also, DenseRNN would be trained more effectively without dropout technique.

17

다큐멘터리극의 상상력: 하비 밀크 사건을 향한 청각기억의 확장

김태형

[NRF 연계] 문학과영상학회 문학과영상 Vol.12 No.3 2011.09 pp.719-750

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The essay examines the ways in which three documentary dramas, in both filmic and theatrical representations, investigate the assassination of San Francisco mayor George Moscone and supervisor Harvey Milk in 1978. There are gaps in time and methodological differences among the three, but the analysis focuses on their approaches to the event, particularly the claims of authenticity in terms of personal memories and historical recordings. Using as much evidence as possible, whether official or unofficial, these documentary dramas concern the re-creation of the death of a gay icon, the courtroom drama surrounding his assassin, and the following social debates on justice, both for the victims and their perpetrator alike. Upon the possible (mis)representations of the protagonists of historical and theatrical events, this essay goes into the aesthetics of documentary drama, especially the sound that promotes audience empathy with the victims. Here, sound means the voices of the victims and the perpetrator, recorded or enacted, the narrations of the films, the represented vocal innuendos of Milk, and the musical effects that dramatize his personality. The narratives of the authors are not limited to the ‘found and selected’ facts in this search for justice, but expand the range of themes beyond the so-called historical truth. What is questioned is the imaginary (dramatic invention) in documentary drama, which often endows fantasy as an author/historian digs for evidence. Even one-sided stories, however, as we may notice in these three dramas, can have their own truth-claim, which is the characteristic or paradox of documentary. The function and aesthetic of documentary lies in this very slippage between the real and the represented, which in turn enriches the genre’s narrative strategy and structural formation.

 
페이지 저장