년 - 년
TAL-CoLR : Temporal Action Localization using Contrastive Learning Representation
한국ITS학회 한국ITS학회 학술대회 ITS와 함께하는 미래 스마트 시티 2022.06 pp.851-854
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Attentive Transfer Learning via Self-supervised Learning for Cervical Dysplasia Diagnosis
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.17 No.3 2021 pp.453-461
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Many deep learning approaches have been studied for image classification in computer vision. However, there are not enough data to generate accurate models in medical fields, and many datasets are not annotated. This study presents a new method that can use both unlabeled and labeled data. The proposed method is applied to classify cervix images into normal versus cancerous, and we demonstrate the results. First, we use a patch self-supervised learning for training the global context of the image using an unlabeled image dataset. Second, we generate a classifier model by using the transferred knowledge from self-supervised learning. We also apply attention learning to capture the local features of the image. The combined method provides better performance than state-of-the-art approaches in accuracy and sensitivity.
[Kisti 연계] 한국정보통신학회 Journal of information and communication convergence engineering Vol.22 No.2 2024 pp.165-171
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this study, we present a novel approach for enhancing chest X-ray image classification (normal, Covid-19, edema, mass nodules, and pneumothorax) by combining contrastive learning and machine learning algorithms. A vast amount of unlabeled data was leveraged to learn representations so that data efficiency is improved as a means of addressing the limited availability of labeled data in X-ray images. Our approach involves training classification algorithms using the extracted features from a linear fine-tuned Momentum Contrast (MoCo) model. The MoCo architecture with a Resnet34, Resnet50, or Resnet101 backbone is trained to learn features from unlabeled data. Instead of only fine-tuning the linear classifier layer on the MoCopretrained model, we propose training nonlinear classifiers as substitutes for softmax in deep networks. The empirical results show that while the linear fine-tuned ImageNet-pretrained models achieved the highest accuracy of only 82.9% and the linear fine-tuned MoCo-pretrained models an increased highest accuracy of 84.8%, our proposed method offered a significant improvement and achieved the highest accuracy of 87.9%.
Robust cross-dataset deepfake detection with multitask self-supervised learning
[NRF 연계] 한국통신학회 ICT Express Vol.11 No.5 2025.10 pp.858-862
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Deepfake detection is increasingly critical due to the rise of manipulated media. Existing methods often require extensive datasets and struggle with interpretability issues. To address these issues, this study introduces a novel one-class approach for detecting and localizing deepfake artifacts in videos, using authentic images to generate manipulated data for training. By integrating segmentation and leveraging convolutional neural networks with visual transformers, the method predicts both the presence and location of the generated manipulations. Experiments on seven deepfake datasets and emerging diffusion-based manipulations show that our approach consistently outperforms existing methods, demonstrating superior accuracy and localization capabilities.
경량화된 비전-언어 모델의 효율적 학습을 위한 자기지도학습 설계
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 2025 한국차세대컴퓨팅학회 춘계학술대회 2025.05 pp.94-96
본 논문은 연안 해역의 CCTV 영상 데이터를 분석하여 위험 상황(예: 고립된 요구조자, 산불, 태풍 등)을 인지하고, 이를 자연어로 설명할 수 있는 경량 비전-언어 모델(VLM) 개발을 목표로 한다. 대규 모 라벨링 없이도 학습 가능한 자기지도학습(Self-Supervised Learning) 기법을 적용하여, 해양 환경 특화 영상 표현을 학습하고, 이후 생성형 언어모델을 결합해 장면을 기술하는 시스템을 제안한다. 특 히, MoCo, DINOv2 등 최신 대조학습 기반 자기지도 모델과 BLIP, Flamingo 등 멀티모달 학습 기법 을 분석하고, 이를 경량화 전략(지식 증류, 양자화 등)과 연계하여 실시간 추론이 가능한 구조를 설계 한다. 문헌 기반 실험 고찰을 통해, 제안된 방법이 적은 라벨로도 높은 설명 성능을 보일 가능성이 높 으며, 실제 연안 감시 시스템에 적용할 경우 위험 탐지 신뢰성과 맥락 이해력이 향상될 것으로 기대된 다. 향후 실제 구현과 도메인 특화 데이터 확보, 시간적 서술 확장 등 과제를 논의하며 본 연구의 실 용성과 확장성을 제시한다.
Growing inventions in deep learning have made a huge impact in remote sensing applications such as disaster management, smart city applications and environment monitoring. Unmanned Aerial Vehicles (UAVs) have emerged as a cost-effective resource among remote sensing platforms offering higher-resolution and detail-oriented observations. Semantic segmentation of such detailed images can lead to many potential applications such as urban planning and recent advancements in segmentation frameworks have served this purpose well. However, these frameworks rely largely on annotated data which is a costly and time-consuming process. Availability of limited labelled data pose challenges as well. To address these issues, a self-supervised learning architecture is proposed in this paper using redundancy reduction principle. A multiple sample redundancy reduction loss-based encoder-decoder architecture is presented. This loss is used to pre-train the encoder in an unsupervised way to learn data representations. These learned representations are used to initialize the encoder for the downstream task. And the encoder-decoder structure is fine-tuned for image segmentation tasks. The efficacy of the proposed network is validated on Urban Drone Dataset achieving 64.90% intersection over union (IoU), 83.06% overall accuracy and 78.37% kappa value while outperforming other dual sample-based loss architectures.
경량 깊이완성기술을 위한 효율적인 자기지도학습 기법 연구 KCI 등재
한국ITS학회 한국ITS학회논문지 제20권 제6호 통권98호 2021.12 pp.313-330
※ 기관로그인 시 무료 이용이 가능합니다.
5,200원
카메라와 라이다가 탑재된 자율주행 시스템에서 깊이완성기술을 통해 조밀한 깊이추정을 할 수 있다. 특히, 자기지도학습을 이용하면 깊이정답이 없는 주행데이터로도 깊이완성 네트워 크의 학습이 가능하다. 실제 자율주행환경에서 이러한 깊이완성의 출력은 다른 알고리즘들의 입력으로 사용되므로 매우 빠른 지연속도를 요구한다. 그래서 본 논문에서는 종래의 연구들처 럼 네트워크를 고도화하여 정확도를 높이기보단 추론속도를 극대화한 형태의 깊이완성 네트 워크를 사용한다. GPU 연산에 최적화된 RegNet 인코더를 사용하고 네트워크의 병렬성을 고려 한 U-Net 형태의 네트워크를 설계한다. 대신, 본 논문에서는 자기지도학습 과정에서 정확도를 높일 수 있는 몇 가지 기법들을 제시한다. 제시하는 기법들은 신뢰할 수 없는 라이다 입력에 대한 강인함을 높이고 사전에 추출한 시맨틱 정보를 바탕으로 에지와 하늘 영역에 대한 깊이 추정 품질을 향상시킨다. 실험을 통해 우리의 모델은 매우 경량임에도 (2.42ms at 1280x480) 노 이즈에 강하며 최신 연구들과 대등한 정확도를 보임을 확인한다.
In an autonomous driving system equipped with a camera and lidar, depth completion techniques enable dense depth estimation. In particular, using self-supervised learning it is possible to train the depth completion network even without ground truth. In actual autonomous driving, such depth completion should have very short latency as it is the input of other algorithms. So, rather than complicate the network structure to increase the accuracy like previous studies, this paper focuses on network latency. We design a U-Net type network with RegNet encoders optimized for GPU computation. Instead, this paper presents several techniques that can increase accuracy during the process of self-supervised learning. The proposed techniques increase the robustness to unreliable lidar inputs. Also, they improve the depth quality for edge and sky regions based on the semantic information extracted in advance. Our experiments confirm that our model is very lightweight (2.42 ms at 1280x480) but resistant to noise and has qualities close to the latest studies.
자기 지도 학습 기반 도로 노면 포트홀 이상 탐지에 관한 연구
한국ITS학회 한국ITS학회 학술대회 ITS, Connected World 2024.10 pp.555-558
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
도로 노면의 포트홀은 ITS분야 중 교통 안전 및 차량 성능에 중대한 영향을 미치는 문제이다. 따라서 본 연구에서는 포트홀 이상 탐지를 위한 새로운 딥러닝 방법론을 제안한다. 제안된 방법론은 비지도 학습 기반의 특징 추출 방법과 Mahalanobis 거리 측정을 활용하여 포트홀을 효과적으로 탐지한다. 카메라와 드론을 활용하여 다양한 도로 조건에서 이미지를 획득한 오픈 데이터 세트를 사용하였으며, 이러한 이미지에서 특징을 학습한다. 성능 평가 결과, 우리의 방법은 기존 기술에 비해 우수한 탐지 성능을 달성하였다. 특히, 다양한 도로 환경에서의 실험을 통해 제안된 방법론에 대한 평가를 진행하였다. 향후, 실시간 데이터 처리와 포트홀 종류별 특성을 파악하여 유지보수 시스템에 적합하도록 발전시킬 계획에 있으며 우리의 접근법은 안전한 도로 유지관리에 기여할 수 있는 가능성을 보여준다.
중복 이미지 탐지(near-duplicate image detection)는 중복에 대하여 미리 정의된 특정 조건을 만족하는 이미지 쌍을 찾는 문제이다. 최근 연구에서는 중복 이미지 쌍을 찾기 위해 이미지 분석을 위한 딥러닝 기술이 활용되고 있 는데, 기존의 딥러닝 기법은 주로 분류 문제를 중심으로 한 지도 학습 방식을 사용하므로 클래스 라벨과 같은 교사 신호가 필요하다. 또한 중복 이미지 탐지와 같이 주어지는 입력 이미지의 도메인이 객체·풍경·인물 등과 같은 특정 그룹에 국한되지 않는 일반적인 경우에는 충분한 일반화 성능을 확보하기 힘들다. 이 문제를 해결하기 위해 본 논문 은 최근 주목받는 자기 지도 학습 모델(self-supervised model)을 이용하여 이미지의 특징을 추출하고, 다른 이미 지의 특징과 거리를 비교하는 방법으로 중복 이미지 탐지 하는 방법을 제안한다. 본 논문은 California-ND 데이터 와 MIR-Flickr Near Duplicate(MFND) 데이터를 사용하여 다른 모델과 비교 실험을 수행하였고, 제안하는 방 법이 기존 방법보다 우수한 성능을 내는 것을 확인하였다.
Near-duplicate image detection is a find image pairs that satisfy some predefined conditions on image duplication. To solve this problem, deep learning techniques for image analysis are being used actively. The conventional approaches mainly use supervised learning methods for classification tasks, which needs appropriate supervised signals such as class labels. Also, it is difficult to have sufficient generalization performance when the domain of input images is not limited to a specific group. Considering these difficulties, this paper proposes a method for near duplicate image detection by extracting image features using a self-supervised model, which has recently been attracting attention, and comparing distances between the features. We conduct comparative experiments with other models using California-ND data and MIR-Flickr Near Duplicate (MFND) data, and confirm that our method achieves the state-of-the-art performance.
Self-supervised Graph Learning을 통한 멀티모달 기상관측 융합
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2023 pp.589-591
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
현재 수치예보 시스템은 항공기, 위성 등 다양한 센서에서 얻은 다종 관측 데이터를 동화하여 대기 상태를 추정하고 있지만, 관측변수 또는 물리량이 서로 다른 관측들을 처리하기 위한 계산 복잡도가 매우 높다. 본 연구에서 기존 시스템의 계산 효율성을 개선하여 관측을 평가하거나 전처리하는 데에 효율적으로 활용하기 위해, 각 관측의 특성을 고려한 자기 지도학습 방법을 통해 멀티모달 기상관측으로부터 실제 대기 상태를 추정하는 방법론을 제안하고자 한다. 비균질적으로 수집되는 멀티모달 기상관측 데이터를 융합하기 위해, (i) 기상관측의 heterogeneous network를 구축하여 개별 관측의 위상정보를 표현하고, (ii) pretext task 기반의 self-supervised learning을 바탕으로 개별 관측의 특성을 표현한다. (iii) Graph neural network 기반의 예측 모델을 통해 실제에 가까운 대기 상태를 추정한다. 제안하는 모델은 대규모 수치 시뮬레이션 시스템으로 수행되는 기존 기술의 한계점을 개선함으로써, 이상 관측 탐지, 관측의 편차 보정, 관측영향 평가 등 관측 전처리 기술로 활용할 수 있다.
[Kisti 연계] 한국전자통신연구원 ETRI journal Vol.46 No.6 2024 pp.1020-1029
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Self-supervised learning is a method that learns the data representation through unlabeled data. It is efficient because it learns from large-scale unlabeled data and through continuous research, performance comparable to supervised learning has been reached. Contrastive learning, a type of self-supervised learning algorithm, utilizes data similarity to perform instance-level learning within an embedding space. However, it suffers from the problem of false-negatives, which are the misclassification of data class during training the data representation. They result in loss of information and deteriorate the performance of the model. This study employed cosine similarity and temperature simultaneously to identify false-negatives and mitigate their impact to improve the performance of the contrastive learning model. The proposed method exhibited a performance improvement of up to 2.7% compared with the existing algorithm on the CIFAR-100 dataset. Improved performance on other datasets such as CIFAR-10 and ImageNet was also observed.
[NRF 연계] 대한배뇨장애요실금학회 International Neurourology Journal Vol.29 2025.07 pp.90-94
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Purpose: Recent studies on urolithiasis detection using deep learning have demonstrated promising accuracy; however, most rely on large-scale labeled imaging datasets. In clinical practice, only limited and partially labeled computed tomography (CT) scans are typically available, restricting the generalizability of conventional supervised models. This study aimed to propose a data-efficient framework for accurate stone detection from a small CT dataset by integrating self-supervised learning (SSL) and transfer learning (TL). Methods: A total of 100 abdominal CT scans were analyzed and labeled as stone present or normal by expert radiologists. To learn generalizable feature representations from limited data, a SimCLR-based SSL framework with a ResNet50 backbone was employed. During the SSL stage, the model learned from augmented image pairs without labels to maximize similarity between positive pairs and minimize similarity between negatives. The pretrained encoder was subsequently fine-tuned using labeled data in the TL stage, with the lower layers frozen and higher blocks optimized using a linear classifier. Model training was performed with 5-fold cross-validation, and performance was evaluated using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC). Results: The proposed SSL+TL model achieved the best performance (AUC, 0.95; F1-score, 0.91), significantly outperforming both the random initialization and TL-only models. These findings indicate that SSL pretraining effectively learns robust and transferable representations even with limited data. Conclusions: The proposed framework demonstrates the feasibility of artificial intelligence-based urolithiasis detection in small-data clinical environments. Combining SSL and TL alleviates data scarcity and provides a foundation for developing generalizable and resource-efficient diagnostic models for urological imaging.
Korean Text to Gloss: Self-Supervised Learning approach
[Kisti 연계] 한국스마트미디어학회 스마트미디어저널 Vol.12 No.1 2023 pp.32-46
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Natural Language Processing (NLP) has grown tremendously in recent years. Typically, bilingual, and multilingual translation models have been deployed widely in machine translation and gained vast attention from the research community. On the contrary, few studies have focused on translating between spoken and sign languages, especially non-English languages. Prior works on Sign Language Translation (SLT) have shown that a mid-level sign gloss representation enhances translation performance. Therefore, this study presents a new large-scale Korean sign language dataset, the Museum-Commentary Korean Sign Gloss (MCKSG) dataset, including 3828 pairs of Korean sentences and their corresponding sign glosses used in Museum-Commentary contexts. In addition, we propose a translation framework based on self-supervised learning, where the pretext task is a text-to-text from a Korean sentence to its back-translation versions, then the pre-trained network will be fine-tuned on the MCKSG dataset. Using self-supervised learning help to overcome the drawback of a shortage of sign language data. Through experimental results, our proposed model outperforms a baseline BERT model by 6.22%.
Facial Expression Recognition through Self-supervised Learning for Predicting Face Image Sequence
[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.27 No.9 2022 pp.41-47
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 자동표정인식을 위하여 얼굴 이미지 배열의 가운데 이미지를 예측하는 새롭고 간단한 자기주도학습 방법을 제안한다. 자동표정인식은 딥러닝 모델을 통해 높은 성능을 달성할 수 있으나 일반적으로 큰 비용과 시간이 투자된 대용량의 데이터 세트가 필요하고, 데이터 세트의 크기와 알고리즘의 성능이 비례한다. 제안하는 방법은 추가적인 데이터 세트 구축 없이 기존의 데이터 세트를 활용하여 자기주도학습을 통해 얼굴의 잠재적인 심층표현방법을 학습하고 학습된 파라미터를 전이시켜 자동표정인식의 성능을 향상한다. 제안한 방법은 CK+와 AFEW 8.0 두가지 데이터 세트에 대하여 높은 성능 향상을 보여주었고, 간단한 방법으로 큰 효과를 얻을 수 있음을 보여주었다.
In this paper, we propose a new and simple self-supervised learning method that predicts the middle image of a face image sequence for automatic expression recognition. Automatic facial expression recognition can achieve high performance through deep learning methods, however, generally requires a expensive large data set. The size of the data set and the performance of the algorithm are tend to be proportional. The proposed method learns latent deep representation of a face through self-supervised learning using an existing dataset without constructing an additional dataset. Then it transfers the learned parameter to new facial expression reorganization model for improving the performance of automatic expression recognition. The proposed method showed high performance improvement for two datasets, CK+ and AFEW 8.0, and showed that the proposed method can achieve a great effect.
Binding Affinity Prediction Using Self-supervised Learning and QP-Ensemble
[NRF 연계] 계명대학교 자연과학연구소 Quantitative Bio-Science Vol.42 No.1 2023.05 pp.57-64
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The drug discovery and optimization of candidate compounds are key initial stages in drug development. Predicting the affinity between drugs and proteins, known as drug-target binding affinity (DTA), is a significant problem in the fields of biochemistry, structural biology, and artificial intelligence. In this study, we propose the utilization of ensemble learning with convolutional neural networks (CNNs) using quadratic programming (QP) to improve prediction accuracy. Additionally, we suggest incorporating self-supervised and semi-supervised learning with data augmentation. Experimental results using the Davis dataset demonstrated that the proposed self-supervised learning model outperformed other learning methods in predicting DTA across all four metrics, including mean squared error (MSE).
비원어민 한국어 발음 평가를 위한 자기 지도 학습 기반 한국어 음소 인식
[Kisti 연계] 한국음성학회 말소리와 음성과학 Vol.17 No.1 2025 pp.51-61
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
비원어민의 한국어 발음 평가를 위해서는 음소 인식뿐만 아니라 발음 오류를 정확하게 탐지할 수 있는 모델이 필요하다. 자기 지도 학습(self-supervised learning) 기반 음성 인식에서 대량의 음성 자료를 통해 구축된 사전 학습(pre-trained) 모델은 정밀한 음성 인식을 가능하게 하는 것으로 알려져 있다. 특히 Wav2Vec2.0 및 Whisper와 같은 모델들은 여러 연구에서 우수한 한국어 음성 인식 성능을 보였으며, Whisper 모델은 특히 뛰어난 성능을 나타냈다. SUPERB 벤치마크를 통해 다양한 사전 학습 모델을 비교한 결과, 음소 인식 분야에서의 성과도 입증되었다. 그러나 비원어민 화자의 발화 특성을 반영하여 실제 발화에 정교하게 맞춘 레이블 데이터의 부족으로 비원어민의 한국어 음소 인식 성능을 높이는 데 한계가 존재한다. 따라서, 비원어민의 한국어 발음 평가 모델을 구축하기 위해서는 정확한 레이블을 갖춘 데이터의 확보가 중요하다. 본 연구에서는 사전 학습된 Whisper 모델을 활용하여 비원어민의 한국어 발음 평가를 위한 한국어 음소 인식 모델을 개발한다. AIHub에서는 아시아, 중국, 일본, 유럽, 영어권 비원어민의 한국어 교육용 음성 데이터를 대량으로 제공하고 있어 이를 모델의 미세 조정(fine-tuning)을 위한 데이터로 활용한다. 그러나, 제공된 음소 레이블의 정확도가 매우 떨어지는 문제가 있어, 한국인이 명료하게 발음한 "뉴스 대본 및 앵커 음성 데이터"를 추가로 활용하여 정확한 한국어 음소 발음을 학습시킨다. 이 두 데이터로부터 구축한 모델을 통해 비원어민 한국어 음성 데이터의 음소 레이블을 실제 발음에 맞게 수정하고, 이를 다시 미세 조정에 적용하여 비원어민 한국어 음소 인식 모델을 구축한다. 이 같은 과정을 몇 차례 단계별로 수행하여 미세 조정 모델을 지속적으로 갱신한다. 최종적으로 구축한 모델의 유효성을 평가하기 위해 비원어민의 한국어 발화 음성과 원어민의 한국어 음성을 대상으로 음소 인식 실험을 진행한 결과, 기본 모델에 비해 음소 인식 성능이 유의미하게 향상되었음을 확인하였다.
To evaluate the Korean pronunciation of non-native speakers, it is essential to develop models capable of recognizing Korean phonemes and detecting pronunciation errors at the phoneme level. Self-supervised learning models, such as Wav2Vec2.0 and Whisper, which were trained on large-scale speech data, have demonstrated strong performance in Korean speech recognition. However, their phoneme recognition accuracy for non-native speakers may be limited because of the lack of labeled data reflecting the unique characteristics of non-native speech. In this study, we developed a Korean phoneme recognition model tailored for non-native speakers by fine-tuning the pretrained Whisper model with Korean language education data from AIHub. This dataset includes speech samples from non-native speakers of various nationalities. In particular, to address the issue of the low phoneme label accuracy in this corpus, we proposed a method to improve label quality by incorporating news data clearly articulated by native Korean news anchors with the AIHub data. The refined dataset was then used for further fine-tuning, resulting in improved phoneme recognition performance. Experiments on Korean phoneme recognition with non-native speakers showed a significant increase in accuracy compared to models trained without the refined data.
단순한 합성데이터 생성 방식을 활용한 gMLP 기반 자기 지도 학습 이상탐지 기법
[Kisti 연계] 한국정보통신학회 한국정보통신학회논문지 Vol.27 No.1 2023 pp.8-14
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
기존의 자기지도 학습 기반의 CutPaste 기법은 정상 이미지에서 특정 패치를 자르고 붙이는 방법으로 합성 데이터를 생성한 뒤 이상탐지를 수행하였다. 그러나 이런 방식으로 생성된 합성데이터는 패치의 경계에 뚜렷한 차이가 나타나는 문제가 발생된다. 이러한 문제를 해결하기 위한 NSA 기법은 Poisson Blending을 통해 자연스러운 합성 데이터를 생성하여 더 높은 이상탐지 성능을 달성하였다. 그러나 NSA 기법은 클래스마다 조정해야하는 하이퍼 파라미터가 많은 단점을 가지고 있다. 본 논문에서는 합성 패치의 크기를 매우 작게 하는 단순한 방법으로 정상과 유사한 합성 데이터를 생성하였다. 이 때 패치가 매우 지역적으로 합성되기 때문에, 지역적인 특징을 학습하는 모델을 사용하면 합성 데이터에 쉽게 과적합 될 수 있다. 따라서 전역적인 특징을 학습하는 gMLP를 사용하여 이상탐지를 수행하였고, 단순한 합성 방법으로도 기존 자기 지도 학습 기법보다 더 높은 성능을 달성할 수 있었다.
The existing self-supervised learning-based CutPaste generated synthetic data by cutting and attaching specific patches from normal images and then performed anomaly detection. However, this method has a problem in that there is a clear difference in the boundary of the patch. NSA for solving these problems have achieved higher anomaly detection performance by generating natural synthetic data through Poisson Blending. However, NSA has the disadvantage of having many hyperparameters that need to be adjusted for each class. In this paper, synthetic data similar to normal were generated by a simple method of making the size of the synthetic patch very small. At this time, since the patches are so locally synthesized, models that learn local features can easily overfit synthetic data. Therefore, we performed anomaly detection using gMLP, which learns global features, and even with simple synthesis methods, we were able to achieve higher performance than conventional self-supervised learning techniques.
자기지도 학습 기반 Deep DIC를 활용한 인장시험의 전체 변형률장 추정
[Kisti 연계] 한국스마트미디어학회 스마트미디어저널 Vol.15 No.6 2026 pp.46-54
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
기존의 딥러닝 기반 디지털 이미지 상관법(Digital Image Correlation, DIC) 방법들은 주로 합성 스페클 데이터로 학습되어 실제 실험 환경에서 도메인 차이로 인한 성능저하가 발생할 수 있다. 본 연구에서는 Deep DIC를 기반으로, 변위-변형률 간 물리적 일관성을 학습과정에 통합한 자기지도 파인튜닝 프레임워크를 제안한다. 핵심 기여는 StrainNet의 예측 변형률과 DisplacementNet 수치미분 변형률 간 일치를 강제하는 변형률 일관성 손실의 도입으로, 두 네트워크 간 상호 제약 매커니즘을 통해 물리적으로 일관된 변형률 추정을 가능하게 한다. 1,094장의 실제 인장 시험 이미지에 대해 ZEISS 상용 소프트웨어를 기준으로 정량 평가를 수행한 결과, 합성 데이터만으로 학습된 기준 모델 대비 평균절대오차(MAE)가 14.9%에서 2.85%로, 평균제곱근오차(RMSE) 또한 16.52%에서 3.42%로 줄어들고 상관계수 0.9917을 달성하였다. 이는 변형률 일관성 제약이 자기지도 파인튜닝의 정확도를 크게 향상시키며, 합성-실험 도메인 차이를 효과적으로 완화함을 보여준다.
Deep learning-based DIC methods trained on synthetic speckle data often suffer performance degradation on real experimental environments due to the domain gap. This study proposes a self-supervised fine-tuning framework based on Deep DIC, introducing a strain consistency loss that enforces agreement between StrainNet predictions and numerically differentiated strains from DisplacementNet. Quantitative evaluation on 1,094 real tensile test images shows that the proposed method reduces MAE from 14.9% to 2.85% and RMSE from 16.52% to 3.42%, achieving a correlation coefficient of 0.9917 against ZEISS commercial software, demonstrating that strain consistency constraint effectively mitigates the synthetic-to-real domain gap.
자기지도학습 표현을 적용한 심층 신경망 기반 음성 향상 기법 연구
[Kisti 연계] 한국음향학회 한국음향학회지 Vol.44 No.1 2025 pp.58-65
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
자기지도학습(Self-Supervised Learning, SSL) 표현을 음성 처리 기법에 적용하는 다양한 연구가 진행되어 왔다. 자기지도학습은 비지도 학습 중 하나로 대량의 데이터에 대해 스스로 라벨을 생성하여 학습을 수행하는 방식을 말한다. 이 과정에서 모델은 입력된 대량의 데이터에서 잠재된 특징을 학습할 수 있게 된다. 본 논문에서는 자기지도학습에서 추출된 잠재된 음성 표현을 마스크 추정 기반 U-Net 모델에 적용한 음성 향상 연구를 제안한다. U-Net 모델의 Encoder-Decoder 구조는 입력 신호를 압축하고 skip-connection을 통해 복원하여 음성 명료도를 효과적으로 향상시킨다. 제안하는 방법에서 skip-connection에 자기지도학습 표현을 추가적으로 전달함으로써 기존의 시스템과 비교하여 깨끗한 음성 예측 성능이 향상된 마스크를 추정하도록 한다. 음성 향상의 결과를 평가하기 위한 지표로 Source-to-Distortion Ratio(SDR), Perceptual Evaluation of Speech Quality(PESQ), Short-Time Objective Intelligibility (STOI)를 사용하였으며, 본 논문에서 제안하는 방법이 SDR, PESQ, STOI 모든 지표에서 개선된 성능을 보였다.
Various studies have been conducted to apply Self-Supervised Learning (SSL) representation to speech processing. SSL is one of the unsupervised learning, which refers to a method of performing learning by generating labels on large amount of data by itself. In this progress, the model can learn the latent features of the input large amount of data. In this paper, we propose a study on speech enhancement by applying the latent speech representation extracted from SSL to a mask estimation-based U-Net model. The encoder-decoder structure of the U-Net model compresses the input signal and effectively restores it through skip-connection to improve speech clarity. At this time, the SSL representation is additionally delivered to skip-connection to estimate the mask with improved performance than the existing system. Source-to-Distortion Ratio (SDR), Perceptual Evaluation of Speech Quality (PESQ), and Short-Time Objective Intelligibility (STOI) were used as performance measure to evaluate speech enhancement, and the method proposed in this paper showed improved performance in SDR, PESQ, and STOI. Through this, we showed that the proposed method can improve speech clarity and quality.
딥 뉴럴 네트워크의 적절한 구조 및 자가-지도 학습 방법에 따른 뇌신호 데이터 표현 기술 분석 및 고찰
[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.19 No.1 2024 pp.137-142
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근, 의료 데이터 표현 분야에서 딥러닝 방법들이 사실상의 표준으로 자리잡고 있다. 하지만, 딥러닝 기술은 내재적으로 많은 양의 학습 데이터를 필요로 하므로 대규모의 데이터를 확보하기 쉽지 않은 의료 분야에서는 직접적인 적용이 어려운 실정이다. 특히 뇌신호 모달리티의 경우, 변동성이 크기 때문에 여전히 데이터 부족 문제를 가진다. 이에, 최근 연구에서는 뇌신호의 시간-공간-주파수 특징을 적절하게 추출할 수 있는 딥 뉴럴 네트워크 구조를 설계하거나, 혹은 자가-지도 학습 방법을 도입하여 뇌신호의 신경생리학적 특징을 미리 학습하도록 한다. 본 논문에서는, 최근 각광받는 기술인 뇌-컴퓨터 인터페이스 및 피험자 상태 예측 등의 관점에서 소규모데이터를 다루기 위해 적용되는 방법론에 대한 분석 및 향후 기술 방향성을 제시한다. 먼저 현재 제안되고 있는 뇌신호 표현을 위한 딥 뉴럴 네트워크 구조에 대해 분석한다. 또한 뇌신호의 특성을 잘 학습하기 위한 자가-지도 학습 방법론을 분석한다. 끝으로, 딥러닝 기반 뇌신호 분석을 위한 중요 시사점 및 방향성에 관하여 논한다.
Recently, deep learning technology has become those methods as de facto standards in the area of medical data representation. But, deep learning inherently requires a large amount of training data, which poses a challenge for its direct application in the medical field where acquiring large-scale data is not straightforward. Additionally, brain signal modalities also suffer from these problems owing to the high variability. Research has focused on designing deep neural network structures capable of effectively extracting spectro-spatio-temporal characteristics of brain signals, or employing self-supervised learning methods to pre-learn the neurophysiological features of brain signals. This paper analyzes methodologies used to handle small-scale data in emerging fields such as brain-computer interfaces and brain signal-based state prediction, presenting future directions for these technologies. At first, this paper examines deep neural network structures for representing brain signals, then analyzes self-supervised learning methodologies aimed at efficiently learning the characteristics of brain signals. Finally, the paper discusses key insights and future directions for deep learning-based brain signal analysis.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.