년 - 년
불균형 데이터 처리를 통한 침입탐지 성능향상에 관한 연구 KCI 등재
한국융합보안학회 융합보안논문지 제21권 제3호 2021.09 pp.57-66
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
침입탐지 분야에서 딥러닝과 머신러닝을 이용한 탐지성능이 검증되면서 이를 활용한 사례가 나날이 증가하고 있다. 하지만, 학습에 필요한 데이터 수집이 어렵고, 수집된 데이터의 불균형으로 인해 머신러닝 성능이 현실에 적용되는데 어려움이 있다. 본 논문에서는 이에 대한 해결책으로 불균형 데이터 처리를 위해 t-SNE 시각화를 이용한 혼합샘플링 기법을 제안한다. 이를 위해 먼저, 페이로드를 포함한 침입탐지 이벤트에 대해서 특성에 맞게 필드를 분리한다. 분리된 필드에 대해 TF-IDF 기반의 피처를 추출한다. 추출된 피처를 기반으로 혼합샘플링 기법을 적용 후 t-SNE를 이용한 데이터 시각화를 통해 불균형 데이터가 처리된 침입탐지에최적화된데이터셋을얻게된다. 공개침입탐지데이터셋CSIC2012를통해9가지샘플링기법을적용하였으며, 제안 한 샘플링 기법이 F-score, G-mean 평가 지표를 통해 탐지성능이 향상됨을 검증하였다.
As the detection performance using deep learning and machine learning of the intrusion detection field has been verified, the cases of using it are increasing day by day. However, it is difficult to collect the data required for learning, and it is difficult to apply the machine learning performance to reality due to the imbalance of the collected data. Therefore, in this paper, A mixed sampling technique using t-SNE visualization for imbalanced data processing is proposed as a solution to this problem. To do this, separate fields according to characteristics for intrusion detection events, including payload. Extracts TF-IDF-based features for separated fields. After applying the mixed sampling technique based on the extracted features, a data set optimized for intrusion detection with imbalanced data is obtained through data visualization using t-SNE. Nine sampling techniques were applied through the open intrusion detection dataset CSIC2012, and it was verified that the proposed sampling technique improves detection performance through F-score and G-mean evaluation indicators.
혼합샘플링 기법을 사용한 랜섬웨어탐지 성능향상에 관한 연구 KCI 등재
한국융합보안학회 융합보안논문지 제23권 제1호 2023.03 pp.69-77
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 아일랜드 보건당국, 美송유관 등 全세계적으로 랜섬웨어 피해가 급증하고 있으며, 사회 모든 분야에 피 해를 입히고 있다. 특히, 랜섬웨어 탐지 및 대응에 기존의 탐지방법뿐 아니라 머신러닝 등을 이용한 연구가 늘어 나고 있다. 하지만, 전통적인 머신러닝은 모델이 데이터가 많은 쪽으로 예측하는 경향이 강해 정확한 예측값을 추 출하기 어려운 문제점이 있다. 이에 다수(Majority)의 Non-Ransomware(정상코드 또는 멀웨어)와 소수의(Minorit y) Ransomware로 구성된 불균형(Imbalance) 클래스에서 샘플링 기법을 통해 불균형을 해소하고 랜섬웨어탐지 성능을 향상시키는 기법을 제안하였다. 본 실험에서는 두가지 시나리오(Binary, Multi Classification)을 사용하여 샘플링 기법이 다수 클래스의 탐지 성능을 유지하면서 소수 클래스의 탐지 성능을 개선함을 확인하였다. 특히, 제 안된 혼합샘플링 기법(SMOTE+ENN)이 10% 이상의 성능(G-mean, F1-score) 향상을 도출했다.
Recently, ransomware damage has been increasing rapidly around the world, including Irish health authorities and U.S. oil pipelines, and is causing damage to all sectors of society. In particular, research using machine learning as well as existing detection methods is increasing for ransomware detection and response. However, traditional machine learning has a problem in that it is difficult to extract accurate predictions because the model tends to predict in the direction where there is a lot of data. Accordingly, in an imbalance class consisting of a large number of non-Ransomware (normal code or malware) and a small number of Ransomware, a technique for resolving the imbalance and improving ransomware detection performance is proposed. In this experiment, we use two scenarios (Binary, Multi Classification) to confirm that the sampling technique improves the detection performance of a small number of classes while maintaining the detection performance of a large number of classes. In particular, the proposed mixed sampling technique (SMOTE+ENN) resulted in a performance(G-mean, F1-score) improvement of more than 10%.
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.15 No.3 2019 pp.682-693
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The traditional classification methods mostly assume that the data for class distribution is balanced, while imbalanced data is widely found in the real world. So it is important to solve the problem of classification with imbalanced data. In Mahalanobis-Taguchi system (MTS) algorithm, data classification model is constructed with the reference space and measurement reference scale which is come from a single normal group, and thus it is suitable to handle the imbalanced data problem. In this paper, an improved method of MTS-CBPSO is constructed by introducing the chaotic mapping and binary particle swarm optimization algorithm instead of orthogonal array and signal-to-noise ratio (SNR) to select the valid variables, in which G-means, F-measure, dimensionality reduction are regarded as the classification optimization target. This proposed method is also applied to the financial distress prediction of Chinese listed companies. Compared with the traditional MTS and the common classification methods such as SVM, C4.5, k-NN, it is showed that the MTS-CBPSO method has better result of prediction accuracy and dimensionality reduction.
Re-SSS: Rebalancing Imbalanced Data Using Safe Sample Screening
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.17 No.1 2021 pp.89-106
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Different samples can have different effects on learning support vector machine (SVM) classifiers. To rebalance an imbalanced dataset, it is reasonable to reduce non-informative samples and add informative samples for learning classifiers. Safe sample screening can identify a part of non-informative samples and retain informative samples. This study developed a resampling algorithm for Rebalancing imbalanced data using Safe Sample Screening (Re-SSS), which is composed of selecting Informative Samples (Re-SSS-IS) and rebalancing via a Weighted SMOTE (Re-SSS-WSMOTE). The Re-SSS-IS selects informative samples from the majority class, and determines a suitable regularization parameter for SVM, while the Re-SSS-WSMOTE generates informative minority samples. Both Re-SSS-IS and Re-SSS-WSMOTE are based on safe sampling screening. The experimental results show that Re-SSS can effectively improve the classification performance of imbalanced classification problems.
Improving the Error Back-Propagation Algorithm for Imbalanced Data Sets
[Kisti 연계] 한국콘텐츠학회 International journal of contents Vol.8 No.2 2012 pp.7-12
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Imbalanced data sets are difficult to be classified since most classifiers are developed based on the assumption that class distributions are well-balanced. In order to improve the error back-propagation algorithm for the classification of imbalanced data sets, a new error function is proposed. The error function controls weight-updating with regards to the classes in which the training samples are. This has the effect that samples in the minority class have a greater chance to be classified but samples in the majority class have a less chance to be classified. The proposed method is compared with the two-phase, threshold-moving, and target node methods through simulations in a mammography data set and the proposed method attains the best results.
A Statistical Perspective of Neural Networks for Imbalanced Data Problems
[Kisti 연계] 한국콘텐츠학회 International journal of contents Vol.7 No.3 2011 pp.1-5
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
It has been an interesting challenge to find a good classifier for imbalanced data, since it is pervasive but a difficult problem to solve. However, classifiers developed with the assumption of well-balanced class distributions show poor classification performance for the imbalanced data. Among many approaches to the imbalanced data problems, the algorithmic level approach is attractive because it can be applied to the other approaches such as data level or ensemble approaches. Especially, the error back-propagation algorithm using the target node method, which can change the amount of weight-updating with regards to the target node of each class, attains good performances in the imbalanced data problems. In this paper, we analyze the relationship between two optimal outputs of neural network classifier trained with the target node method. Also, the optimal relationship is compared with those of the other error function methods such as mean-squared error and the n-th order extension of cross-entropy error. The analyses are verified through simulations on a thyroid data set.
Fused Based CNN+LSTM Structure with Imbalanced Data for Fall Detection
한국AI디지털융합학회(구 한국디지털융합학회) IJICTDC Vol 8 No 1 2023.06 pp.34-48
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
Considering the aging of individuals with the increasing population, the demand for technologies that enable people to follow their daily lives unnoticed is increasing day by day. In this article, a high-performance solution to the fall and posture detection problem for CCD camera-based fall detection systems is provided. Fall detection was performed with images obtained with CCD cameras placed in different positions. Within the scope of the proposed method, two different pre-trained CNN structures were trained using two different camera images. Data fusion was applied to the high-level features obtained from these structures. Features that were fusion process applied to different classifiers were given as input and ensemble learning process was applied. Considering the performance metrics of the proposed method, it was predicted that promising results were obtained for fall detection.
CycleGAN을 이용한 편향 테이블 데이터 (Imbalanced Table Data) 오버샘플링 (Oversampling) 문제 해결 방안에 대한 연구 : 금융사기를 중심으로
한국경영정보학회 한국경영정보학회 정기 학술대회 2019년 경영정보관련 추계학술대회 2019.11 pp.436-440
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
현대 사회는 사람의 행동 하나가 데이터가 되며 이는 곧 엄청난 데이터의 흐름을 만든다. 20년 전 인터넷 속 전체 데이터의 양이 현대 사회속에서는 1초마다 저장된다. 이러한 추세는 앞으로 더욱 더 심화될 것이며 이러한 빅데이터를 활용하기에 따라서 엄청난 이점을 줄 수 있을 것으로 판단된다. 이러한 데이터의 분석을 위해서는 편향되지 않은 데이터가 필요한데 대부분의 빅데이터는 한쪽으로 편향인 불균형 상태며 이는 분석의 정확도를 떨어뜨리는 원인 중 하나이다. 또한 2종 오류의 비용이 큰 분야에서는 불균형 데이터를 사용한 분석을 믿을 수 없는 실정이기 때문에 이러한 문제점을 해결하는 것은 매우 중요하다. 정형 데이터 분야에서는 이러한 문제점을 해결하기 위해서 전통적인 통계 기법 방식의 오버샘플링이 발전해왔고 비정형 데이터에서는 딥러닝의 발전과 더불어 발전한 생성 모델이 불균형 문제의 해결책으로 떠올랐다. 본 연구에서는 비정형 데이터에서 오버샘플링을 하기 위해 자주 사용하는 생성 모델 중 CycleGAN을 정형 데이터에 맞게 변형시킬 것이다. 또한 GMM을 이용해 혼합 분포를 각각의 단일 분포로 분해하여 CycleGAN이 데이터의 특징을 더 잘 학습하게 만들 것이며 CycleGAN에 Classifier를 추가하여 좀 더 현실적인 데이터를 만드는 오버샘플링 기법을 만들고자 한다. 본 논문에서 제안하고자하는 오버샘플링 기법을 실험하기 위해 실제 금융사기에 관한 데이터를 PCA로 변조하여 개인정보를 가린 불균형 데이터를 사용할 것이다.
Structure-Preserving Data Augmentation for Imbalanced Classification of Kepler Light Curves KCI 등재
조선대학교 기초과학연구원 통합자연과학논문집(구 조선자연과학논문집) 제19권 2호 2026.06 pp.47-60
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
Detecting exoplanet transits in Kepler light curves is challenging due to severe class imbalance and the highly localized morphology of transit dips. Standard oversampling approaches such as SMOTE generate synthetic minority samples through feature-space interpolation, which can distort or blur transit structures in phase-folded representations. We introduce a structure-preserving data augmentation method tailored to folded Kepler light curves. Our approach identifies transit dip regions in positive samples and applies controlled perturbations—phase jitter, constrained depth scaling, and realistic noise injection—while maintaining transit morphology consistency. Using 2048-bin folded signals and a 1D convolutional neural network trained under KIC-grouped splits to prevent target leakage, we compare the proposed method against a weighted baseline and SMOTE. Across 10 random seeds, the proposed soft augmentation configuration achieves higher and more stable AUPRC than SMOTE and provides consistent improvement over the weighted baseline. Ablation studies indicate that phase jitter is the primary contributor to performance gains, while overly aggressive depth perturbation can degrade results. These findings highlight the importance of domain-aware, structure-preserving augmentation for robust imbalanced classification of Kepler exoplanet candidates.
SMOTE를 이용한 편중된 횡 분산계수 데이터에 대한 추정식 개발
[Kisti 연계] 한국수자원학회 한국수자원학회 논문집 Vol.54 No.12 2021 pp.1305-1316
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구에서는 과거 추적자실험결과를 이용하여 2차원 횡분산계수에 대한 새로운 추정식을 개발하고 추정식을 이용한 횡 분산계수 산정결과의 정확도를 검증했다. 다수의 추적자실험이 하폭 대 수심비가 50보다 작은 조건에서 수행되었기 때문에 기존 추적자실험결과만을 이용하여 개발한 추정식은 하폭 대 수심비가 50보다 큰 조건의 하천에 적용하는데 한계를 보인다. 따라서 특정 수리조건에 편중된 횡 분산계수 자료로부터 횡 분산계수 추정식을 개발하기 위해 SMOTE (Synthetic Minority Oversampling TEchnique)를 적용하여 기존 자료의 특성을 반영한 새로운 데이터를 생성했다. SMOTE 기법으로 하폭 대 수심비가 50보다 큰 조건에 대한 수리량과 횡 분산계수 데이터를 생성하였으며, ROC (Receiver Operating Characteristic) 곡선으로부터 생성된 데이터의 신뢰성을 검증했다. 새롭게 생성된 데이터를 포함하여 횡 분산계수 추정식을 개발했고, 추정식을 이용하여 계산한 횡 분산계수의 R<sup>2</sup>(결정계수)를 계산하여 기존 연구에서 제안한 추정식과의 정확도를 비교했다. 그 결과, 본 연구에서 개발한 추정식을 이용하여 계산한 횡 분산계수의 R<sup>2</sup>가 W/H < 50인 조건에서 0.81, 50 < W/H 인 조건에서 0.92를 나타내어 기존 추정식과 비교하여 향상된 정확도를 나타냈다.
In this study, a new empirical formula for 2D transverse dispersion coefficient was developed using the results of previous tracer test studies, and the performance of the formula was evaluated. Since many tracer test studies have been conducted under the conditions where the width-to-depth ratio is less than 50, the existing empirical formulas developed using these imbalanced tracer test results have limitations in applying to rivers with a width-to-depth ratio greater than 50. Therefore, in order to develop an empirical formula for transverse dispersion coefficient using the imbalanced tracer test data, the Synthetic Minority Oversampling TEchnique (SMOTE) was used to oversample new data representing the properties of the existing tracer test data. The hydraulic data and the transverse dispersion coefficients in conditions of width-to-depth ratio greater than 50 were oversampled using the SMOTE. The reliability of the oversampled data was evaluated using the ROC (Receiver Operating Characteristic) curve. The empirical formula of transverse dispersion coefficient was developed including the oversampled data, and the performance of the results were compared with the empirical formulas suggested in previous studies using R<sup>2</sup>. From the comparison results, the value of R<sup>2</sup> was 0.81 for the range of W/H < 50 and 0.92 for 50 < W/H, which were improved accuracy compared to the previous studies.
금융 이상 거래 탐지에서의 Semi-Hard Example Mining 기반 불균형 데이터 증강 기법 KCI 등재
한국경영정보학회 경영정보학연구 제27권 제3호 2025.08 pp.375-397
※ 기관로그인 시 무료 이용이 가능합니다.
6,000원
최근 급격히 증가하고 있는 금융 이상 거래는 막대한 경제적 손실을 일으키고 있다. 하지만 금융 이상 거래 탐지에 이용되는 데이터에서 이상 거래는 정상 거래에 비해 극히 적어 효과적인 탐지를 어렵게 하는 불균형 데이터 문제가 제기되어왔다. 본 연구는 이러한 불균형한 데이터 특성의 한계를 극복하기 위해 VAE-GAN(Variational Autoencoder-Generative Adversarial Network) 과 Semi-Hard Example Mining 기법을 결합하여, 이상 거래 데이터의 품질을 유지하면서 실제로 이상 거래이지만 정상 거래로 판단하는 거짓 음성(False Negative)을 줄이는 모델을 제안한다. 먼저, VAE-GAN을 통해 실제 거래와 유사한 소수 클래스 합성 데이터를 생성하고, Semi-Hard Example Mining으로 분류기가 헷갈리기 쉬운 사례를 집중적으로 재생성한다. 이를 신용카드 이상 거래 데이터셋에 적용한 결과, 기존 보간 기반 오버샘플링 기법(SMOTE, Borderline-SMOTE, ADASYN)과 기존 VAE-GAN 증강 대비 재현율(Recall), F2 스코어(F2 Score)가 향상됨을 확인하였다. 본 연구는 금융권 FDS(Fraud Detection System)에서 불균형 데이터 문제를 완화하고 탐지 성능을 극대화하는 데 기여할 것으로 기대한다.
The rapid rise in fraudulent financial transactions is inflicting substantial economic losses, yet effective detection remains difficult because genuine fraud represents only a tiny fraction of overall activity. To overcome this extreme class-imbalance problem, we propose a model that integrates a Variational Autoencoder Generative Adversarial Network (VAE-GAN) with Semi-Hard Example Mining (SHEM). The VAE-GAN synthesizes high-fidelity minority-class samples that closely mimic real transactions, while SHEM repeatedly targets borderline cases that the classifier is prone to misjudge, thereby reducing false negatives (fraudulent transactions incorrectly labeled as legitimate). Experiments on a benchmark credit- card-fraud dataset show that our method consistently outperforms interpolation-based oversampling techniques (SMOTE, Borderline-SMOTE, ADASYN) and a vanilla VAE-GAN baseline, achieving higher Precision, Recall, F1, and F2 scores. These results demonstrate the model’s potential to alleviate class imbalance and maximize detection performance in financial-sector fraud-detection systems(FDS).
불균형 데이터 집합에서의 의사결정나무 추론: 종합 병원의 건강 보험료 청구 심사 사례 KCI 등재
한국경영정보학회 경영정보학연구 제9권 제1호 2007.04 pp.45-65
※ 기관로그인 시 무료 이용이 가능합니다.
5,700원
다른 산업과 달리 병원/의료 산업에서는 건강 보험료 심사 평가라는 독특한 검증 과정이 필수적으로 있게 된다. 건강 보험료 심사 평가는 병원의 수익 문제 뿐 아니라 적정한 진료행위를 하는 병원이라는 이미지와도 맞물려 매우 중요한 분야이며, 특히 대형 종합병원일수록 이 부분에 많은 심사관련 인력들을 투입하여, 병원의 수익과 명예를 위해서 업무를 수행하고 있다. 본 논문은 이러한 건강보험료 청구 심사 과정에서, 사전에 수많은 진료 청구 건 중 심사 평가에서 삭감이 될 수 있는 진료 청구 건을 데이터 마이닝을 통해서 발견하여, 사전의 대비를 철저히 하고자 하는 한 국내 대형 종합병원의 사례를 소개하고자 한다. 데이터 마이닝을 적용함에 있어, 주요한 문제점 중 하나는 바로 지도학습 기법을 적용하기에 곤란한 데이터 불균형 문제가 발생하는 것이다. 이런 불균형 문제를 해소하고, 비교 조건 중에 가장 효율적인 삭감 예상 진료 건 탐지 모델을 만들어 내기 위하여, 데이터 불균형 문제의 기본 해법인 Sampling과 오분류 비용의 다양한 혼합적인 적용을 통하여, 적합한 조건을 가지는 의사결정 나무 모델을 도출하였다.
In medical industry, health insurance bill audit is unique and essential process in general hospitals. The health insurance bill audit process is very important because not only for hospital's profit but also hospital's reputation. Particularly, at the large general hospitals many related workers including analysts, nurses, and etc. have engaged in the health insurance bill audit process. This paper introduces a case of health insurance bill audit for finding reducible health insurance bill cases using decision tree induction techniques at a large general hospital in Korea. When supervised learning methods had been tried to be applied, one of major problems was data imbalance problem in the health insurance bill audit data. In other words, there were many normal(passing) cases and relatively small number of reduction cases in a bill audit dataset. To resolve the problem, in this study, well-known methods for imbalanced data sets including over sampling of rare cases, under sampling of major cases, and adjusting the misclassification cost are combined in several ways to find appropriate decision trees that satisfy required conditions in health insurance bill audit situation.
클래스 불균형 데이터의 분류 성능 향상을 위한 언어 증강과 Focal loss 를 활용한 Supervised Contrastive Learning 모델
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 2023 한국차세대컴퓨팅학회 춘계학술대회 2023.06 pp.72-75
소셜미디어의 발달로 인하여 즉각적인 소통이 활발해졌지만, 혐오표현이 유발하는 차별행위가 늘어남에 따라 혐오표현을 필터링하는 연구의 필요성이 제기되고 있다. 혐오표현은 다양한 카테고리로 구분되지만, 카테고리별로 균형 잡힌 데이터셋을 구축하기에는 어려움이 존재한다. 따라서 본 연구에서는 데이터 증강을 적용하여 혐오표현 분류 성능을 향상시킨 모델을 제시한다. Easy data augmentation techniques를 적용하여 최소 규모의 카테고리 데이터를 증강하였다. Kcbert-base 모델에 focal loss와 supervised contrastive learning을 적용하여, 동일 카테고리의 문장 유사도는 높이고, 다른 카테고리와의 문장 유사도는 낮추면서 모델을 학습시켰다. 실험 결과 증강과 focal loss를 적용하지 않은 모델에 비해 easy data augmentation techniques와 focal loss, supervised contrastive learning을 적용한 모델의 평균 정확도는 1.4%, macro f1-score는 4.4% 우수한 것을 확인하였다.
불균형 정형 데이터를 위한 SMOTE와 변형 CycleGAN 기반 하이브리드 오버샘플링 기법 KCI 등재
한국경영정보학회 경영정보학연구 제24권 제4호 2022.11 pp.97-118
※ 기관로그인 시 무료 이용이 가능합니다.
5,800원
이미지와 같은 비정형 데이터의 불균형 클래스 문제 해결에 있어 생산적 적대 신경망(generative adversarial network)에 기반한 오버샘플링 기법의 우수성이 알려짐에 따라 다양한 연구들이 이를 정형 데이터의 불균형 문제 해결에도 적용하기 시작하였다. 그러나 이러한 연구들은 데이터의 형태를 비정형 데이터 구조로 변경함으로써 정형 데이터의 특징을 정확하게 반영하지 못한다는 점이 문제로 지적되고 있다. 본 연구에서는 이를 해결하기 위해 순환 생산적 적대 신경망(cycle GAN)을 정형 데이터의 구조에 맞게 재구성하고 이를 SMOTE(synthetic minority oversampling technique) 기법과 결합한 하이브리드 오버샘플링 기법을 제안하였다. 특히 기존 연구와 달리 생산적 적대 신경망을 구성함에 있어 1차원 합성곱 신경망(1D-convolutional neural network)을 사용함으로써 기존 연구의 한계를 극복하고자 하였다. 본 연구에서 제안한 기법의 성능 비교를 위해 불균형 정형 데이터를 기반으로 오버샘플링을 진행하고 그 결과를 SMOTE, ADASYN(adaptive synthetic sampling) 등과 같은 기존 기법과 비교하였다. 비교 결과 차원이 많을수록, 불균형 정도가 심할수록 제안된 모형이 우수한 성능을 보이는 것으로 나타났다. 본 연구는 기존 연구와 달리 정형 데이터의 구조를 유지하면서 소수 클래스의 특징을 반영한 오버샘플링을 통해 분류의 성능을 향상시켰다는 점에서 의의가 있다.
As generative adversarial network (GAN) based oversampling techniques have achieved impressive results in class imbalance of unstructured dataset such as image, many studies have begun to apply it to solving the problem of imbalance in structured dataset. However, these studies have failed to reflect the characteristics of structured data due to changing the data structure into an unstructured data format. In order to overcome the limitation, this study adapted CycleGAN to reflect the characteristics of structured data, and proposed hybridization of synthetic minority oversampling technique (SMOTE) and the adapted CycleGAN. In particular, this study tried to overcome the limitations of existing studies by using a one-dimensional convolutional neural network unlike previous studies that used two-dimensional convolutional neural network. Oversampling based on the method proposed have been experimented using various datasets and compared the performance of the method with existing oversampling methods such as SMOTE and adaptive synthetic sampling (ADASYN). The results indicated the proposed hybrid oversampling method showed superior performance compared to the existing methods when data have more dimensions or higher degree of imbalance. This study implied that the classification performance of oversampling structured data can be improved using the proposed hybrid oversampling method that considers the characteristic of structured data.
TabNet 기반 생성적 적대 신경망(GAN)을 활용한 고용 빅데이터의 불균형 클래스 최적화 모델링 KCI 등재
한국기계항공기술학회(구 한국기계기술학회) 한국기계항공기술학회지(구 한국기계기술학회지) 제26권 제3호 2024.06 pp.453-461
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Handling imbalanced datasets in binary classification, especially in employment big data, is challenging. Traditional methods like oversampling and undersampling have limitations. This paper integrates TabNet and Generative Adversarial Networks (GANs) to address class imbalance. The generator creates synthetic samples for the minority class, and the discriminator, using TabNet, ensures authenticity. Evaluations on benchmark datasets show significant improvements in accuracy, precision, recall, and F1-score for the minority class, outperforming traditional methods. This integration offers a robust solution for imbalanced datasets in employment big data, leading to fairer and more effective predictive models.
Imbalanced Data Classification Based on AdaBoost-SVM
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.7 No.5 2014.10 pp.85-94
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The classification of imbalanced data is one of the most challenging problems in data mining and machine learning research. Imbalanced dataset is a form that exists in reality area, which describes truly and objectively the essential characters of something. There will appear paucity of data and flooded in the classification of imbalanced dataset. Beside problems such as loss of information and data splitting phenomenon will also appear when using the traditional machine learning methods. So how to solve the classification problem of imbalanced data will be challenging. In this paper, aiming at the above problems, a classification algorithm based on AdaBoost-SVM is proposed. In the experiments with four typical forms of imbalanced data sets in UCI were validated the effectiveness of this strategy.
Heterogeneous Ensemble of Classifiers from Under-Sampled and Over-Sampled Data for Imbalanced Data KCI 등재
국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 8 Number 1 2019.03 pp.75-81
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Data imbalance problem is common and causes serious problem in machine learning process. Sampling is one of the effective methods for solving data imbalance problem. Over-sampling increases the number of instances, so when over-sampling is applied in imbalanced data, it is applied to minority instances. Under-sampling reduces instances, which usually is performed on majority data. We apply under-sampling and over-sampling to imbalanced data and generate sampled data sets. From the generated data sets from sampling and original data set, we construct a heterogeneous ensemble of classifiers. We apply five different algorithms to the heterogeneous ensemble. Experimental results on an intrusion detection dataset as an imbalanced datasets show that our approach shows effective results.
Imbalanced Data SVM Classification Method Based on Cluster Boundary Sampling and DT-KNN Pruning
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.7 No.2 2014.04 pp.61-68
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
This paper presents a SVM classification method based on cluster boundary sampling and sample pruning. We actively explore an effective solution to solve the difficult problem of imbalanced data set classification from data re-sampling and algorithm improving. Firstly, we creatively propose the method of cluster boundary sampling, using the clustering density threshold and the boundary density threshold to determine the cluster boundaries, in order to guide the process of re-sampling more scientifically and accurately. Secondly, we put forward a new sample pruning algorithm based on dynamic threshold KNN to deal with the complexity and overlapping problem of imbalanced data set. The phenomenon of data complexity and overlapping will reduce the classification performance and generalization ability of SVM classifier. Experiments show that our method acquires obviously promotion effect in various different imbalanced data sets and it can prove the validity and st
A Study on Imbalanced Data Stream Processing Using a Mass Function SCOPUS
보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.9 No.11 2015.11 pp.91-98
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In the IOT environment, sensor data stream consists of event data from heterogeneous multi-sensors. One type of sensor may have quite a different event frequency from those other kinds of sensors, which makes most sensor data sets imbalanced. To classify an imbalanced data effectively, it is necessary to preprocess it for converting into a balanced data. This process may unify heterogeneous attributes in the imbalanced data and alleviate the difficulties for data mining on it. Mass function plays an important role in the fuzzy theory and Dempster-Shafer Theory. In this paper, using a mass function is suggested to process imbalanced data stream. A mass function is developed to compute mass values for imbalanced data sets, and an experiment is performed to investigate the validity to apply the mass function to the sensor data stream.
SVM Classification for High-dimensional Imbalanced Data based on SNR and Under-sampling SCOPUS
보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.10 No.4 2015.04 pp.105-112
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Support vector machine (SVM) is biased towards the majority class, in some case dataset is class-imbalanced and the bias is even larger for high-dimensional. In order to improve the classification accuracy of SVM on high-dimensional imbalanced data, we combine signal-noise ratio (SNR) and under-sampling technique based on K-means. In this article firstly we apply SNR into feature selection to reducing the feature amount then solve the problem of data imbalance using under-sampling technique based on K-means. To verify the feasibility of the proposed strategy, we utilize some metrics such as receiver operating characteristic curve (ROC curve) and area under the receiver operating characteristic curve (AUC value).As a result, the AUC value increased by 4%~16% before and after the process. The experimental results show that our strategy is feasible and effective exactly.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.