Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 91
No
1

This paper analyzed the required specification for implementing a real-time object detection system using the YOLO network. The YOLO network is one of the object detection schemes that strike a suitable balance between performance and required resources, making it widely applicable across various domains. In this paper, we measured the performance based on different activations for an efficient implementation of the YOLO network. Also, we conducted a comparative analysis of the YOLO network's performance concerning input characteristics, aiming to analyze the appropriate application method for real-time object detection.

2

Implementation of an Autostereoscopic Virtual 3D Button in Non-contact Manner Using Simple Deep Learning Network

You, Sang-Hee, Hwang, Min, Kim, Ki-Hoon, Cho, Chang-Suk

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.17 No.3 2021 pp.505-517

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This research presented an implementation of autostereoscopic virtual three-dimensional (3D) button device as non-contact style. The proposed device has several characteristics about visible feature, non-contact use and artificial intelligence (AI) engine. The device was designed to be contactless to prevent virus contamination and consists of 3D buttons in a virtual stereoscopic view. To specify the button pressed virtually by fingertip pointing, a simple deep learning network having two stages without convolution filters was designed. As confirmed in the experiment, if the input data composition is clearly designed, the deep learning network does not need to be configured so complexly. As the results of testing and evaluation by the certification institute, the proposed button device shows high reliability and stability.

3

Deep learning-driven methods for network-based intrusion detection systems: A systematic review

Ramya Chinnasamy, Malliga Subramanian, Sathishkumar Veerappampalayam Easwaramoorthy, Jaehyuk Cho

[NRF 연계] 한국통신학회 ICT Express Vol.11 No.1 2025.02 pp.181-215

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper presents a systematic review of deep learning (DL) techniques for Network-based Intrusion Detection Systems (NIDS) based on Preferred Reporting Items for Systematic Reviews and Meta-Analyses: (PRISMA2020) guidelines. It explores recent advancements in data preparation, DL architectures, and performance evaluation metrics for NIDS. The review provides insights into various datasets and tools used in the field, highlighting the effectiveness of DL in improving NIDS performance. Additionally, it discusses the applications of NIDS across different industries and identifies emerging research trends, offering a comprehensive resource for researchers and practitioners in cybersecurity.

4

Text Classification on Social Network Platforms Based on Deep Learning Models

YA, Chen, Tan, Juan, Hoekyung, Jung

[Kisti 연계] 한국정보통신학회 Journal of information and communication convergence engineering Vol.21 No.1 2023 pp.9-16

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The natural language on social network platforms has a certain front-to-back dependency in structure, and the direct conversion of Chinese text into a vector makes the dimensionality very high, thereby resulting in the low accuracy of existing text classification methods. To this end, this study establishes a deep learning model that combines a big data ultra-deep convolutional neural network (UDCNN) and long short-term memory network (LSTM). The deep structure of UDCNN is used to extract the features of text vector classification. The LSTM stores historical information to extract the context dependency of long texts, and word embedding is introduced to convert the text into low-dimensional vectors. Experiments are conducted on the social network platforms Sogou corpus and the University HowNet Chinese corpus. The research results show that compared with CNN + rand, LSTM, and other models, the neural network deep learning hybrid model can effectively improve the accuracy of text classification.

5

A High-Precision Feature Extraction Network of Fatigue Speech from Air Traffic Controller Radiotelephony Based on Improved Deep Learning

Zhiyuan Shen, Yitao Wei

[NRF 연계] 한국통신학회 ICT Express Vol.7 No.4 2021.12 pp.403-413

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Air traffic controller (ATC) fatigue is receiving considerable attention in recent studies because it represents a major cause of air traffic incidences. Research has revealed that the presence of fatigue can be detected by analysing speech utterances. However, constructing a complete labelled fatigue data set is very time-consuming. Moreover, a manually constructed speech collection will often contain only little key information to be used effectively in fatigue recognition, while multilevel deep models based on such speech materials often have overfitting problems due to an explosive increase of model parameters. To address these problems, a novel deep learning framework is proposed in this study to integrate active learning (AL) into complex speech features selected from a large set of unlabelled speech data in order to overcome the loss of information. A shallow feature set is first extracted using stacked sparse autoencoder networks, in which fatigue state challenge features from a manually selected speaker set of are exploited as the input vector. A densely connected convolutional autoencoder (DCAE) is then proposed to learn advanced features automatically from spectrograms of the selected data to supplement the fatigue features. The network can be effectively trained using a relatively small number of labelled samples with the help of AL sampling strategies, and the addition of a dense block to the convolutional automatic encoder can decrease the number of parameters and make the model easier to fit. Finally, the two above-mentioned features are combined using multiple kernel learning with a support-vector-machine classifier. A series of comparative experiments using the Civil Aviation Administration of China radiotelephony corpus demonstrates that the proposed method provides a significant improvement in the detection precision compared to current state-of-the-art approaches.

6

Deep reinforcement learning-based sum rate maximization for RIS-assisted ISAC-UAV network

Moon Sangmi, 이창건, Liu Huaping, 황인태

[NRF 연계] 한국통신학회 ICT Express Vol.10 No.5 2024.10 pp.1174-1178

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Recent advances in communications technologies have paved the way for integrating communication and sensing functionalities into unmanned aerial vehicle (UAV) networks by using reconfigurable intelligent surfaces (RIS). In this paper, we propose a novel approach to maximize the sum rate of RIS-assisted UAV networks by using an integrated sensing and communications (ISAC) network in conjunction with deep reinforcement learning (DRL). The integration of UAVs with ISAC networks results in dynamic and unpredictable channel conditions, which reduces the effectiveness of traditional optimization techniques. To address this challenge, we develop a DRL-based sum-rate maximization algorithm that adaptively configures the beamforming matrix and RIS phase shifts to optimize the communication performance while achieving the signal-to-noise ratio required for sensing. Our simulation results indicate that the proposed algorithm significantly outperforms the existing methods in terms of sum rate while accommodating the dynamic nature of the ISAC-UAV network.

7

Enhancing network function parallelism in mobile edge computing using Deep Reinforcement Learning

DongYu Lu, Shirong Long

[NRF 연계] 한국통신학회 ICT Express Vol.11 No.1 2025.02 pp.41-46

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper introduces a Deep Reinforcement Learning (DRL)-based framework to enhance Network Function Parallelism (NFP) in Mobile Edge Computing (MEC). Leveraging Network Function Virtualization (NFV), the proposed framework optimizes service delay by solving a fairness-aware throughput maximization problem for service function chain placement. It aims to maximize the long-term cumulative reward while satisfying Quality of Service (QoS) requirements. The framework also preserves resources for future requests by efficiently managing the initialized network functions distribution. Simulation results demonstrate the superior performance of the proposed framework across various metrics. Specifically, our framework improves the average delay and deployment rate by 1.2% and 2.4% compared to the existing best method.

8

Applying multi-agent deep reinforcement learning for contention window optimization to enhance wireless network performance

Ke Chih-Heng, Astuti Lia

[NRF 연계] 한국통신학회 ICT Express Vol.9 No.5 2023.10 pp.776-782

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper investigates the Contention Window (CW) optimization problem in multi-agent scenarios, where the fully cooperative among mobile stations is considered. A partially observable environment is employed to model and analyze the CW optimization problem, and Smart Exponential-Threshold-Linear with Deep Q-learning Network (SETL-DQN) Multi-Agent (MA) algorithm is proposed to obtain the optimal system throughput through the CW Threshold optimization. In the determined scenarios, SETL-DQN(MA) can effectively cope with the mutual interaction among mobile stations. The simulation results show that our proposed method is superior from both static and dynamic scenarios and has the highest optimum packet transmission efficiency.

9

Vehicle Image Recognition Using Deep Convolution Neural Network and Compressed Dictionary Learning

Zhou, Yanyan

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.17 No.2 2021 pp.411-425

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, a vehicle recognition algorithm based on deep convolutional neural network and compression dictionary is proposed. Firstly, the network structure of fine vehicle recognition based on convolutional neural network is introduced. Then, a vehicle recognition system based on multi-scale pyramid convolutional neural network is constructed. The contribution of different networks to the recognition results is adjusted by the adaptive fusion method that adjusts the network according to the recognition accuracy of a single network. The proportion of output in the network output of the entire multiscale network. Then, the compressed dictionary learning and the data dimension reduction are carried out using the effective block structure method combined with very sparse random projection matrix, which solves the computational complexity caused by high-dimensional features and shortens the dictionary learning time. Finally, the sparse representation classification method is used to realize vehicle type recognition. The experimental results show that the detection effect of the proposed algorithm is stable in sunny, cloudy and rainy weather, and it has strong adaptability to typical application scenarios such as occlusion and blurring, with an average recognition rate of more than 95%.

11

Motion prediction using brain waves based on artificial intelligence deep learning recurrent neural network SCOPUS KCI 등재

Kyoung-Seok Yoo

한국운동재활학회 JER Vol.19 No.4 2023.08 pp.219-227

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

Electroencephalogram (EEG) research has gained widespread use in various research domains due to its valuable insights into human body movements. In this study, we investigated the optimization of motion discrimination prediction by employing an artificial intelligence deep learning recurrent neural network (gated recurrent unit, GRU) on unique EEG data generated from specific movement types among EEG signals. The experiment involved participants categorized into five difficulty lev-els of postural control, targeting gymnasts in their twenties and college students majoring in physical education (n=10). Machine learning tech-niques were applied to extract brain-motor patterns from the collected EEG data, which consisted of 32 channels. The EEG data underwent spectrum analysis using fast Fourier transform conversion, and the GRU model network was utilized for machine learning on each EEG frequen-cy domain, thereby improving the performance index of the learning operation process. Through the development of the GRU network algo-rithm, the performance index achieved up to a 15.92% improvement com-pared to the accuracy of existing models, resulting in motion recognition accuracy ranging from a minimum of 94.67% to a maximum of 99.15% between actual and predicted values. These optimization outcomes are attributed to the enhanced accuracy and cost function of the GRU net-work algorithm’s hidden layers. By implementing motion identification optimization based on artificial intelligence machine learning results from EEG signals, this study contributes to the emerging field of exercise re-habilitation, presenting an innovative paradigm that reveals the inter-connectedness between the brain and the science of exercise.

14

부채널 분석을 이용한 딥러닝 네트워크 신규 내부 비밀정보 복원 방법 연구

박수진, 이주헌, 김희석

[Kisti 연계] 한국정보보호학회 정보보호학회논문지 Vol.32 No.5 2022 pp.855-867

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

IoT 장비의 발달로 딥러닝 가속기의 필요성이 증대됨에 따라 이에 탑재되는 딥러닝 가속기의 구현 및 안전성 검증에 대한 연구가 활발히 진행 중이다. 본 논문에서는 Usenix 2019에 발표된 딥러닝 네트워크 복원 논문의 한계점을 극복한 내부 비밀정보 신규 부채널 분석 방법론에 대해 제안한다. 기존 연구에서 네트워크 내부 가중치의 범위를 제한하며 32비트 가중치의 16비트만 복원한 단점이 있다, 제안하는 신규 가중치 복원 방법으로 상관전력분석을 이용하여 IEEE754 32비트 단정밀도 가중치를 99% 정확도로 복원할 수 있음을 보인다. 또한 특정 입력값에 대해서만 활성함수 복원이 가능한 기존 연구의 제약을 극복하고, 딥러닝을 이용한 신규 활성함수 복원 방법으로 입력값에 대한 조건 없이 99% 정확도로 활성함수를 복원한다. 이를 통해 기존 연구가 가지는 한계점들을 극복했을 뿐만 아니라 제안하는 신규 방법론이 효과적이라는 것을 입증한다.

As the need for a deep learning accelerator increases with the development of IoT equipment, research on the implementation and safety verification of the deep learning accelerator is actively. In this paper, we propose a new side channel analysis methodology for secret information that overcomes the limitations of the previous study in Usenix 2019. We overcome the disadvantage of limiting the range of weights and restoring only a portion of the weights in the previous work, and restore the IEEE754 32bit single-precision with 99% accuracy with a new method using CPA. In addition, it overcomes the limitations of existing studies that can reverse activation functions only for specific inputs. Using deep learning, we reverse activation functions with 99% accuracy without conditions for input values with a new method. This paper not only overcomes the limitations of previous studies, but also proves that the proposed new methodology is effective.

15

다변수 LSTM 순환신경망 딥러닝 모형을 이용한 미술품 가격 예측에 관한 실증연구

이지인, 송정석

[Kisti 연계] 한국콘텐츠학회 한국콘텐츠학회논문지 Vol.21 No.6 2021 pp.552-560

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

새로운 미술품 유통방식의 발달로 미술품의 미적 효용을 넘어 투자재로서 바라보는 시각이 활성화되고 있다. 미술품의 가격은 주식이나 채권 등과 달리 객관적 요소와 주관적 요소들이 모두 반영되어 결정되는 이질적 특성이 있기 때문에 가격 예측에 있어서 그 불확실성이 높다. 본 연구에서는 LSTM(장단기 기억) 순환신경망 딥러닝 모형을 활용하여 낙찰총액 순위 1위부터 10위까지의 한국 작가의 회화 작품을 대상으로 작가의 특성, 작품의 물리적 특성, 판매적 특성 등을 입력으로 하여 경매 낙찰가의 예측을 시도하였다. 연구 결과, 모델에 의한 예측 가격과 실제 낙찰 가격의 차이를 설명하는 RMSE 값이 0.064 수준이었으며 작가별로는 이대원 작가의 예측력이 가장 높았고, 이중섭 작가의 예측력이 가장 낮았다. 투자재로서 미술품 시장이 더욱 활성화되고 경매 낙찰 가격의 예측 수요가 높아지면서 본 연구의 결과가 활용될 수 있을 것이다.

With the recent development of the art distribution system, interest in art investment is increasing rather than seeing art as an object of aesthetic utility. Unlike stocks and bonds, the price of artworks has a heterogeneous characteristic that is determined by reflecting both objective and subjective factors, so the uncertainty in price prediction is high. In this study, we used LSTM Recurrent Neural Network deep learning model to predict the auction winning price by inputting the artist, physical and sales charateristics of the Korean artist. According to the result, the RMSE value, which explains the difference between the predicted and actual price by model, was 0.064. Painter Lee Dae Won had the highest predictive power, and Lee Joong Seop had the lowest. The results suggest the art market becomes more active as investment goods and demand for auction winning price increases.

16

딥러닝을 활용한 경량 블록 암호의 신경망 구분자 성능 개선 기술 연구

오채린, 임우상, 정수용, 김현일, 서창호

[Kisti 연계] 한국정보보호학회 정보보호학회논문지 Vol.34 No.6 2024 pp.1171-1188

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

디지털 기술의 빠른 발전으로 보안의 중요성이 더욱 강조되면서, 민감한 데이터를 보호하기 위한 암호화 기술의 필요성이 커지고 있다. 그중, 경량 블록 암호 알고리즘은 자원 제약이 큰 IoT 기기와 같은 환경에서 주로 사용되며, 이러한 알고리즘의 안전성 평가를 위해 딥러닝 기술을 활용한 연구가 주목받고 있다. 특히, 차분 분석 공격에 신경망을 적용한 신경망 구분자(neural distinguisher)가 제안되면서 암호의 취약점을 효과적으로 식별하려는 시도가 늘고 있다. 그러나 암호 라운드가 증가할수록 성능이 저하되는 한계가 존재하게 되었고, 이를 해결하기 위해 입력 데이터 형식 및 가중치 함수 적용 등 다양한 성능 개선 기술이 제안되고 있다. 최근에는 암호문 쌍의 차분 뿐만 아니라 키 쌍의 차분을 이용한 연관키 신경망 구분자(related-key neural distinguisher)가 제안되면서 구분자의 성능이 향상됨을 보였다. 그러나 연관키 신경망 구분자의 성능 개선을 위한 연구는 아직 미흡한 수준이다. 이에 본 논문에서는 신경망 구분자의 성능을 향상시키기 위해 제안된 입력 데이터 형식과 가중치 함수의 영향을 분석하고, 구분자의 출력 값 분포 특성 차이가 미치는 영향에 대해 탐구한다. 이를 통해 Simeck 32/64 알고리즘에 적합한 가중치 함수를 선정하며, 연관키 신경망 구분자의 성능 향상 방안을 도출하여 구분자의 종류에 따라 적합한 가중치 함수가 달라질 수 있음을 새로운 분석 관점으로 제시한다.

The need for secure encryption is growing as digital technology advances. Lightweight block cipher algorithms, essential for resource-limited environments like IoT, are increasingly analyzed using deep learning. Neural distinguishers applying neural networks to differential crypanalysis have been proposed to identify vulnerabilities effectively, though they suffer performance drops with more encryption rounds. To address this, enhancements like modifications in input data formats and the use of weight functions have been proposed. Recently, related-key neural distinguishers, which utilize differences in both ciphertext pairs and key pairs, have demonstrated improved performance. Despite these advancements, studies focusing on performance optimization for related-key neural distinguishers are still limited. In this paper, we analyze the impact of input data formats and weight functions proposed to improve the performance of neural distinguishers and explore how differences in output value distribution characteristics affect performance. Through this analysis, we identify a suitable weight function for the Simeck 32/64 algorithm and propose methods to enhance the performance of related-key neural distinguishers, presenting a new analytical perspective that the optimal weight function may vary depending on the type of distinguisher.

17

딥러닝을 이용한 증강현실 얼굴감정스티커 기반의 다중채널네트워크 플랫폼 구현

김대진

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.19 No.7 2018 pp.1349-1355

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

최근 인터넷을 통한 다양한 콘텐츠 서비스가 일반화 되고 있으며, 그 중에서 다중채널네트워크 플랫폼 서비스는 스마트 폰의 일반화와 함께 인기가 높아지고 있다. 다중채널네트워크 플랫폼은 스트리밍을 기본으로 하면서, 서비스 향상을 위하여 다양한 요소를 추가하고 있다. 그중 얼굴인식을 이용한 증강현실 스티커 서비스가 많이 이용되고 있다. 본 논문에서는 기존 서비스보다 흥미 요소를 더욱 증가 시킬 목적으로 얼굴 감정인식을 통하여 증강현실 스티커를 얼굴에 마스킹 하는 다중채널네트워크 플랫폼을 구현하였다. 얼굴감정인식을 위해 딥러닝 기술을 이용하여 7가지 얼굴의 감정을 분석하였고, 이를 기반으로 감정 스티커를 얼굴에 적용하여, 사용자들의 만족도를 더욱 높일 수 있었다. 제안하는 다중채널네트워크 플랫폼 구현을 위해서 클라이언트에 감정 스티커를 적용하였고, 서버에서 스트리밍 서비스 할 수 있는 여러 가지 서버들을 설계하였다.

Recently, a variety of contents services over the internet are becoming popular, among which MCN(Multi Channel Network) platform services have become popular with the generalization of smart phones. The MCN platform is based on streaming, and various factors are added to improve the service. Among them, augmented reality sticker service using face recognition is widely used. In this paper, we implemented the MCN platform that masks the augmented reality sticker on the face through facial emotion recognition in order to further increase the interest factor. We analyzed seven facial emotions using deep learning technology for facial emotion recognition, and applied the emotional sticker to the face based on it. To implement the proposed MCN platform, emotional stickers were applied to the clients and various servers that can stream the servers were designed.

18

A crucial component of designing intelligent and ecologically friendly environments nowadays is electricity consumption forecasting. The generation of energy can be enhanced to effectively meet the population's rising requirements by using the prediction of future electricity consumption. Due to the broad variety of consumption patterns, it is difficult to anticipate the energy requirements of buildings. Therefore, this work uses a dual-steam approach with multi-head attention to anticipate the power consumption of the building to address this issue and produce precise predictions. The proposed network concurrently learns temporal representations through a Bidirectional Gated Recurrent Unit (BGRU) and spatial patterns through Atrous Convolutional Neural Network (ACNN). The obtained features are combined to create a single feature vector that is used as the input for the multi-head attention, which finds the features that are most suited to forecasting the electricity consumption of a building. Finally, the dense layer receives the effective features and uses them to forecast short-term power consumption. In this paper, the proposed dual-stream network with attention outperforms competing models, achieving the lowest error value for hourly building power consumption prediction, according to experimentation on the household electricity consumption dataset.

19

고령화 사회에서 치매 환자가 빠르게 증가함에 따라 조기 진단 및 예측의 중요성이 커지고 있다. 본 연구는 2013년 부터 2025년까지 발표된 치매 관련 머신러닝 및 딥러닝 연구를 대상으로 시계열·네트워크 분석을 통해 기술 확산과 발전 경로를 규명하였다. PubMed 데이터베이스에서 수집한 8,422편의 문헌을 기반으로 키워드 빈도, 클러스터, 연결중심성 분석을 수행하여 연구의 구조와 핵심 주제를 도출하였다. 분석 결과, ‘humans’, ‘alzheimer disease’, ‘machine learning’은 전 기간에 걸쳐 높은 중심성을 유지하였으며, 시기별로는 초기 영상처리 기반 연구에서 분자 생물학 연구를 거쳐 환자 중심의 임상·공중보건 연구로 발전하는 경향을 보였다. 4개의 주요 연구 클러스터(의료영 상 분석, 분자생물학·유전학, 역학·임상 데이터 분석, 노인 돌봄·기술 활용)가 식별되었으며, 각 클러스터는 독립적 이면서도 상호 연계된 구조를 형성하였다. 이러한 결과는 치매 연구가 AI 기반 기초·응용 기술의 확산과 함께 융합 적·실용적 연구로 진화하고 있음을 보여주며, 향후 치매 연구의 전략적 방향을 설정하는 데 유용한 시사점을 제공할 것으로 기대된다.

As the population ages, the importance of early diagnosis and prediction of dementia is growing. This study investigated the diffusion and development trajectory of machine learning and deep learning technologies in dementia research from 2013 to 2025 through time-series and network analyses. Based on 8,422 publications retrieved from the PubMed database, keyword frequency, clustering, and degree centrality analyses were conducted to identify the structural and thematic landscape of the field. Results showed that ‘humans,’ ‘Alzheimer disease,’ and ‘machine learning’ consistently held high centrality, with research trends evolving from early image-processing-based studies to molecular biology and, more recently, to patient-centered clinical and public health applications. Four main research clusters were identified—medical image analysis, molecular biology/genetics, epidemiology and clinical data analysis, and elderly care/technology utilization—forming an interconnected yet distinct structure. These results indicate that dementia research is evolving into integrative and practical studies alongside the spread of AI-based basic and applied technologies, and are expected to provide useful insights for setting strategic directions for future dementia research.

20

딥러닝 기반의 Multi Scale Attention을 적용한 개선된 Pyramid Scene Parsing Network KCI 등재

김준혁, 이상훈, 한현호

한국융합학회 한국융합학회논문지 제12권 제11호 2021.11 pp.45-51

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

딥러닝의 발전으로 인하여 의미론적 분할 방법은 다양한 분야에서 연구되고 있다. 의료 영상 분석과 같이 정확성을 요구하는 분야에서 분할 정확도가 떨어지는 문제가 있다. 본 논문은 의미론적 분할 시 특징 손실을 최소화하기 위해 딥러닝 기반 분할 방법인 PSPNet을 개선하였다. 기존 딥러닝 기반의 분할 방법은 특징 추출 및 압축 과정에서 해상도가 낮아져 객체에 대한 특징 손실이 발생한다. 이러한 손실로 윤곽선이나 객체 내부 정보에 손실이 발생하여 객체 분류 시 정확도가 낮아지는 문제가 있다. 이러한 문제를 해결하기 위해 의미론적 분할 모델인 PSPNet을 개선하였다. 기존 PSPNet에 제안하 는 multi scale attention을 추가하여 객체의 특징 손실을 방지하였다. 기존 PPM 모듈에 attention 방법을 적용하여 특징 정제 과정을 수행하였다. 불필요한 특징 정보를 억제함으로써 윤곽선 및 질감 정보가 개선되었다. 제안하는 방법은 Cityscapes 데이터 셋으로 학습하였으며, 정량적 평가를 위해 분할 지표인 MIoU를 사용하였다. 실험을 통해 기존 PSPNet 대비 분할 정확도가 약 1.5% 향상되었다.

With the development of deep learning, semantic segmentation methods are being studied in various fields. There is a problem that segmenation accuracy drops in fields that require accuracy such as medical image analysis. In this paper, we improved PSPNet, which is a deep learning based segmentation method to minimized the loss of features during semantic segmentation. Conventional deep learning based segmentation methods result in lower resolution and loss of object features during feature extraction and compression. Due to these losses, the edge and the internal information of the object are lost, and there is a problem that the accuracy at the time of object segmentation is lowered. To solve these problems, we improved PSPNet, which is a semantic segmentation model. The multi-scale attention proposed to the conventional PSPNet was added to prevent feature loss of objects. The feature purification process was performed by applying the attention method to the conventional PPM module. By suppressing unnecessary feature information, eadg and texture information was improved. The proposed method trained on the Cityscapes dataset and use the segmentation index MIoU for quantitative evaluation. As a result of the experiment, the segmentation accuracy was improved by about 1.5% compared to the conventional PSPNet.

 
1 2 3 4 5
페이지 저장