Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 12
No
1

4,000원

부도 데이터는 전체 데이터 중 극히 일부분을 차지하는 대표적인 불균형 데이터이다. 따라서 부도 예측 모형을 구축할 때는 반드시 데이터의 불균형 문제를 고려해야 한다. 본 연구는 이러한 데이터 불균형 문제 개선을 위한 Sampling 기법의 적용과 Classification Threshold 조정에 따른 부도 예측 모형의 성능 변화 양상을 정리하고, 시사점을 도출하는 것을 중심으로 진행하였 다. 실증 분석을 위해 A 은행의 소매 익스포져 데이터를 대상으로 부도 예측 모형을 구축하였고, 모형의 성능을 비교 분석하였다. 분석 결과 첫째, Sampling 기법을 적용할 경우 민감도 개선에 효과가 있음을 확인하였 고, Oversampling 기법에 비해 Undersampling 기법의 성능 개선 효과가 상대적으로 크게 나타났다. 둘째, Classification Threshold 조정 시 모형의 성능 변화 양상은 Sampling 기법을 적용하는 경우와 유사하 며, 정밀도 및 F1 Score 관점에서 Sampling 기법에 비해 상대적으로 큰 개선 효과를 확인하였다. 셋째, Sampling 기법 적용과 Classification Threshold 조정을 함께할 때 민감도 개선 효과가 가장 크게 나타남을 확인하였 다. 부도 예측 모형은 부도 계좌를 정상 계좌로 예측할 때 더 큰 위험을 초래하므로, 민감도가 가장 우수한 모형을 우선하여 활용하는 것이 바람직할 것이 다. 다만, 수익성을 제고할 필요가 있거나 투입 가능한 비용이 제한적인 상황이라면 F1 Score를 기준으로 모형을 선택하는 것이 적절할 것으로 보인다.

2

Texture Classification Based on Random Threshold Vector Technique SCOPUS

B.V. Ramana Reddy, M.Radhika Mani, B.Sujatha, Dr.V.Vijaya Kumar

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.5 No.1 2010.01 pp.53-62

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

A new feature set derived from the fractal geometry, called the random threshold vector (RTV) is proposed for texture analysis. The RTV is computed for different run length entropy dimensions. The run length entropy dimensions are calculated based on different thresholds. To test the rotationally invariant feature, the run length entropies are calculated in different directions. The experimental results show that the RTV contains great discriminatory information needed for a successful classification.

3

An Innovative Method for Texture Classification Based on Random Threshold and Measure of Pattern Trends

B. V. Ramana Reddy, A. Suresh, K.V.Subbaiah, Dr.B. Eswara Reddy

보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.2 No.4 2009.10 pp.7-18

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Study of different patterns on a local neighborhood of a texture plays an important role in characterization, and classification of the textures. The present paper proposes a method for measuring the occurrence factor of patterns on a randomly thresholded binary image. For this process eight simple patterns are chosen on a 3×3 neighborhood. The simple patterns are chosen in such a way that any complex pattern can be formed by grouping one or more of these simple patterns. The pattern occurrence factor of different binary images is also compared with the actual binary texture image. The experimental results on sixty four textures indicate good comparison of variation of occurrence in these patterns on different binary images of random threshold

4

A face recognition algorithm based on local binary Haar feathers which represented as Kadane optimizing multi-threshold AdaBoost was proposed according to the problems of texture shape feature representation and classification algorithm accuracy in the process of facial classifying detection and recognition, First, improve the traditional expression by using image Local binary pattern of Haar features , improve image model of texture and shape feature expression ability ; Secondly, for single threshold weak learning algorithm we can not make full use of local binary Haar feature information, resulting in a lower classification accuracy problem proposed Kadane optimizing multi-threshold AdaBoost classifier, to achieve local binary Haar feature representation of facial high accuracy recognition; Finally, through the experiments show, efficient face recognition rate can reach more than 90% by the algorithm,which is superior to the selected comparison algorithm.

5

Generalization Threshold Optimization of Fuzzy Rough Set algorithm in Healthcare Data Classification SCOPUS

Beibei Dong, Yu Liu, Benzhen Guo, Xiao Zhang

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.3 2016.03 pp.229-238

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

There is ineffective classification problem in application of K-means clustering algorithm in massive data cluster analysis. This paper presents a K-means algorithm based on generalization threshold rough set optimization weight. Firstly, utilize attribute order described method, using the average distance calculation with Laplace method to optimize the generalization threshold of fuzzy rough set , then the Euclidean distance metric is used in the calculation of the similarity of K-means algorithm, introducing the variation coefficient into the cluster analysis, clustering the Euclidean distance weighted K-means algorithm totally based on data, finally, combine the rough set algorithm based on the generalization threshold optimization and K-means clustering algorithm, applied to medical and health data classification. The K-means algorithm based on generalization threshold rough set optimization weight presented by this paper has a better effect on medical and health data classification.

6

Cone Surface Classification and Threshold Value Selection for Description of Complex Objects

Cho, Dong-Uk, Kim, Ji-Yeong, Bae, Young-Lae, Ko, Il-Seok

[Kisti 연계] 한국정보처리학회 정보처리학회논문지 B Vol.b11 No.3 2004 pp.297-302

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 컴퓨터 시각에서 가장 중요한 과제의 하나인 ,3차원 물체의 표현에 대해 원뿔형태의 기술과 표면 분류시 임계치를 자동으로 선정하는 방법에 대해 제안하고자 한다. 기존에 미분기하학에서 사용한 평균 곡률 (H)과 가우스 곡률(K)은 물체의 상당 부분을 차지하고 있는 원뿔표면에 대한 분류가 불가능하였다. 또한 평균 곡률과 가우스 곡률의 부호값에 따른 표면 분류가 실제 거리 영상에 적용시 올바로 분류가 안 되는 문제를 가지고 있었다. 이 논문에서는 기존의 이 같은 두 가지 문제를 해결하기 위해 리지와 벨리의 표면분류로부터 원뿔표면 형태(cone ridge, cone valley)를 분류해 내었다. 즉, 원뿔표면 형태의 경우 H의 값이 일정하고, 원뿔표면 형태의 경우는 H의 값이 다름을 이용하여 원뿔표면 형태를 분류하였다. 아울러 통계적인 관점에서 표면분류 임계치를 선정할 수 있는 방법을 제안하고 실험에 의해 제안한 방법의 유용성을 입증하고자 한다.

In this paper, the 3-D shape description for the objects with the cone ridge and valley surfaces, and the corresponding threshold value selection for surface classification are considered. The existing method based on the mean and Gaussian curvatures(H and K) of differential geometries cannot properly describe cone primitives, which are some of the most common objects in the real world. Also the existing method for surface classification based on the sign values of H and K has Problems in practical applications. For this, cone surface shapes are classified cone ridges and cone valleys are derived from surfaces using the fact that H values are constant case of cylinder surfaces and variable for cone surfaces, respectively. Also threshold value selection for surface classification from a statistical point of view is proposed. The effectiveness of the proposed methods are verified through experiments.

7

비대칭 오류비용을 고려한 분류기준값 최적화와 SVM에 기반한 지능형 침입탐지모형

이현욱, 안현철

[Kisti 연계] 한국지능정보시스템학회 Journal of Intelligence and Information Systems Vol.17 No.4 2011 pp.157-173

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

최근 인터넷 사용의 증가에 따라 네트워크에 연결된 시스템에 대한 악의적인 해킹과 침입이 빈번하게 발생하고 있으며, 각종 시스템을 운영하는 정부기관, 관공서, 기업 등에서는 이러한 해킹 및 침입에 의해 치명적인 타격을 입을 수 있는 상황에 놓여 있다. 이에 따라 인가되지 않았거나 비정상적인 활동들을 탐지, 식별하여 적절하게 대응하는 침입탐지 시스템에 대한 관심과 수요가 높아지고 있으며, 침입탐지 시스템의 예측성능을 개선하려는 연구 또한 활발하게 이루어지고 있다. 본 연구 역시 침입탐지 시스템의 예측성능을 개선하기 위한 새로운 지능형 침입탐지모형을 제안한다. 본 연구의 제안모형은 비교적 높은 예측력을 나타내면서 동시에 일반화 능력이 우수한 것으로 알려진 Support Vector Machine(SVM)을 기반으로, 비대칭 오류비용을 고려한 분류기준값 최적화를 함께 반영하여 침입을 효과적으로 차단할 수 있도록 설계되었다. 제안모형의 우수성을 확인하기 위해, 기존 기법인 로지스틱 회귀분석, 의사결정나무, 인공신경망과의 결과를 비교하였으며 그 결과 제안하는 SVM 모형이 다른 기법에 비해 상대적으로 우수한 성과를 보임을 확인할 수 있었다.

As the Internet use explodes recently, the malicious attacks and hacking for a system connected to network occur frequently. This means the fatal damage can be caused by these intrusions in the government agency, public office, and company operating various systems. For such reasons, there are growing interests and demand about the intrusion detection systems (IDS)-the security systems for detecting, identifying and responding to unauthorized or abnormal activities appropriately. The intrusion detection models that have been applied in conventional IDS are generally designed by modeling the experts' implicit knowledge on the network intrusions or the hackers' abnormal behaviors. These kinds of intrusion detection models perform well under the normal situations. However, they show poor performance when they meet a new or unknown pattern of the network attacks. For this reason, several recent studies try to adopt various artificial intelligence techniques, which can proactively respond to the unknown threats. Especially, artificial neural networks (ANNs) have popularly been applied in the prior studies because of its superior prediction accuracy. However, ANNs have some intrinsic limitations such as the risk of overfitting, the requirement of the large sample size, and the lack of understanding the prediction process (i.e. black box theory). As a result, the most recent studies on IDS have started to adopt support vector machine (SVM), the classification technique that is more stable and powerful compared to ANNs. SVM is known as a relatively high predictive power and generalization capability. Under this background, this study proposes a novel intelligent intrusion detection model that uses SVM as the classification model in order to improve the predictive ability of IDS. Also, our model is designed to consider the asymmetric error cost by optimizing the classification threshold. Generally, there are two common forms of errors in intrusion detection. The first error type is the False-Positive Error (FPE). In the case of FPE, the wrong judgment on it may result in the unnecessary fixation. The second error type is the False-Negative Error (FNE) that mainly misjudges the malware of the program as normal. Compared to FPE, FNE is more fatal. Thus, when considering total cost of misclassification in IDS, it is more reasonable to assign heavier weights on FNE rather than FPE. Therefore, we designed our proposed intrusion detection model to optimize the classification threshold in order to minimize the total misclassification cost. In this case, conventional SVM cannot be applied because it is designed to generate discrete output (i.e. a class). To resolve this problem, we used the revised SVM technique proposed by Platt(2000), which is able to generate the probability estimate. To validate the practical applicability of our model, we applied it to the real-world dataset for network intrusion detection. The experimental dataset was collected from the IDS sensor of an official institution in Korea from January to June 2010. We collected 15,000 log data in total, and selected 1,000 samples from them by using random sampling method. In addition, the SVM model was compared with the logistic regression (LOGIT), decision trees (DT), and ANN to confirm the superiority of the proposed model. LOGIT and DT was experimented using PASW Statistics v18.0, and ANN was experimented using Neuroshell 4.0. For SVM, LIBSVM v2.90-a freeware for training SVM classifier-was used. Empirical results showed that our proposed model based on SVM outperformed all the other comparative models in detecting network intrusions from the accuracy perspective. They also showed that our model reduced the total misclassification cost compared to the ANN-based intrusion detection model. As a result, it is expected that the intrusion detection model proposed in this paper would not only enhance the performance of IDS, but also lead to better management of FNE.

8

임계값 학습 모듈을 적용한 준지도 SAR 이미지 분류

도재준, 김선옥

[NRF 연계] 한국빅데이터학회 한국빅데이터학회지 Vol.8 No.2 2023.12 pp.177-187

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

준지도 학습(Semi-supervised learning)은 소량의 라벨이 있는 데이터와 다량의 라벨이 없는 데이터를 사용하여모델을 훈련하는 효과적인 방법이다. 그러나 많은 논문에서 준지도 학습시 하나의 고정된 임계값을 사용하여각 클래스별 서로 다른 이미지들의 특징별 차이를 고려하지 않고 임의 라벨을 만든다. 본 논문에서는 합성개구레이더(SAR) 이미지 분류 준지도 학습시 모든 클래스가 하나의 고정된 임계값을 사용하는 대신 각 클래스에 대해 서로 다른 임계값을 적용한다. 모델에 임계값 학습 모듈을 추가하여 임계값을 학습하여 클래스별로 학습되는 차이를 고려하여 클래스별로 서로 다른 임계값을 얻는다. 서로 다른 임계값을 사용한 준지도 학습 기반의 SAR 이미지 분류 방법을 적용유무를 비교하여 클래스별 임계값을 사용하는 이점에 대해 고찰하였다.

Semi-supervised learning (SSL) is an effective approach to training models using a small amount of labeled data and a larger amount of unlabeled data. However, many papers in the field use a fixed threshold when applying pseudo-labels without considering the feature-wise differences among images of different classes. In this paper, we propose a SSL method for synthetic aperture radar (SAR) image classification that applies different thresholds for each class instead of using a single fixed threshold for all classes. We propose a threshold learning module into the model, considering the differences in feature distributions among classes, to dynamically learn thresholds for each class. We compare the application of a SSL SAR image classification method using different thresholds and examined the advantages of employing class-specific thresholds.

9

선별적인 임계값 선택을 이용한 준지도 학습의 SAR 분류 기술

도재준, 유민정, 이재석, 문효이, 김선옥

[Kisti 연계] 한국군사과학기술학회 한국군사과학기술학회지 Vol.27 No.3 2024 pp.319-328

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Semi-supervised learning is a good way to train a classification model using a small number of labeled and large number of unlabeled data. We applied semi-supervised learning to a synthetic aperture radar(SAR) image classification model with a limited number of datasets that are difficult to create. To address the previous difficulties, semi-supervised learning uses a model trained with a small amount of labeled data to generate and learn pseudo labels. Besides, a lot of number of papers use a single fixed threshold to create pseudo labels. In this paper, we present a semi-supervised synthetic aperture radar(SAR) image classification method that applies different thresholds for each class instead of all classes sharing a fixed threshold to improve SAR classification performance with a small number of labeled datasets.

10

GPCR 분류에서 ART1 군집화를 위한 퍼지기반 임계값 제어 기법

조규철, 마용범, 이종식

[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.12 No.6 2007 pp.167-175

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

퍼지이론은 생명정보공학에서 지식을 표현하는데 활용되고 제어시스템 모델을 이해하는데 활용되어 왔다. 본 논문에서는 생명정보학의 응용 프로그램에서 중요한 데이터 분류에 초점을 맞추었다. 최적의 임계값 유도를 위한 GPCR 분류에서 기존의 순차기반 임계값 제어기법은 임계값 결정범위와 최적의 임계값 유도 시간의 문제점을 보였고, 이진기반 임계값 제어기법은 임계값 결정 초기에 시스템의 안정성에 대한 단점이 있었다. 이를 보완하기 위해 우리는 ART1 군집화를 위한 퍼지기반 임계값제어기법을 제안한다. 제안된 방법의 성능을 평가하기 위해 ART1 군집화를 위한 퍼지기반 임계값 제어기법을 구현하여 기존의 순차기반 임계값 제어기법과 이진기반 임계값 제어기법과의 인식률에 대한 구동시간의 변화, 임계값의 변화에 따른 시스템의 구동시간을 측정하였다. 퍼지기반 임계값제어 기법은 GPCR 데이터 분류에서 인식률과 구동시간에 대한 정보를 통해 분류 임계값을 조정하여 높은 인식률과 낮은 구동시간을 지속적으로 유도하여 안정적이고 효과적인 분류 시스템을 만들 수 있었다.

Fuzzy logic is used to represent qualitative knowledge and provides interpretability to a controlling system model in bioinformatics. This paper focuses on a bioinformatics data classification which is an important bioinformatics application. This paper reviews the two traditional controlling system models The sequence-based threshold controller have problems of optimal range decision for threshold readjustment and long processing time for optimal threshold induction. And the binary-based threshold controller does not guarantee for early system stability in the GPCR data classification for optimal threshold induction. To solve these problems, we proposes a fuzzy-based threshold controller for ART1 clustering in GPCR classification. We implement the proposed method and measure processing time by changing an induction recognition success rate and a classification threshold value. And, we compares the proposed method with the sequence-based threshold controller and the binary-based threshold controller The fuzzy-based threshold controller continuously readjusts threshold values with membership function of the previous recognition success rate. The fuzzy-based threshold controller keeps system stability and improves classification system efficiency in GPCR classification.

11

불균형 자료에서 불순도 지수를 활용한 분류 임계값 선택

장서인, 여인권

[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.34 No.5 2021 pp.711-721

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

이 논문에서는 불균형 자료에 대한 분류 분석에서 불순도지수를 이용하여 임계값을 조정하는 방법에 대해 알아본다. 이항자료에 대한 분류에서는 소수범주를 Positive, 다수범주를 Negative라고 하면, 일반적으로 사용하는 0.5 기준으로 범주를 정하면 불균형 자료에서는 특이도는 높은 반면 민감도는 상대적으로 낮게 나오는 경향이 있다. 소수범주에 속한 개체를 제대로 분류하는 것이 상대적으로 중요한 문제에서는 민감도를 높이는 것이 중요한데 이를 분류기준이 되는 임계값을 조정을 통해 높이는 방법에 대해 알아본다. 기존연구에서는 G-mean이나 F1-score와 같은 측도를 기준으로 임계값을 조정했으나 이 논문에서는 CHAID의 카이제곱통계량, CART의 지니지수, C4.5의 엔트로피를 이용하여 최적임계값을 선택하는 방법을 제안한다. 최적임계값이 여러 개 나올 수 있는 경우 해결방법을 소개하고 불균형 분류 예제로 사용되는 데이터 분석을 통해 0.5를 기준으로 ?(무엇?)을 때와 비교하여 어떤 개선이 이루어졌는지 등을 분류성능측도로 알아본다.

In this paper, we propose the method of adjusting thresholds using impurity indices in classification analysis on imbalanced data. Suppose the minority category is Positive and the majority category is Negative for the imbalanced binomial data. When categories are determined based on the commonly used 0.5 basis, the specificity tends to be high in unbalanced data while the sensitivity is relatively low. Increasing sensitivity is important when proper classification of objects in minority categories is relatively important. We explore how to increase sensitivity through adjusting thresholds. Existing studies have adjusted thresholds based on measures such as G-Mean and F1-score, but in this paper, we propose a method to select optimal thresholds using the chi-square statistic of CHAID, the Gini index of CART, and the entropy of C4.5. We also introduce how to get a possible unique value when multiple optimal thresholds are obtained. Empirical analysis shows what improvements have been made compared to the results based on 0.5 through classification performance metrics.

12

ECG 패턴 분석과 템플릿 문턱값을 통한 조기수축 부정맥분류

조익성, 조영창, 권혁숭

[Kisti 연계] 한국정보통신학회 한국정보통신학회논문지 Vol.20 No.2 2016 pp.437-444

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

일반적인 부정맥 분류 방법의 경우 심방 박동 수와 관련한 PP간격, P모양의 다양성과 같은 조건을 이용하는데, 잡음으로 인해 정확한 P파의 검출이 어렵기 때문에 잡음의 영향을 비교적 적게 받는 R파를 이용하는 것이 유리하다. 따라서 본 연구에서는 R파 중심의 ECG(electrocardiography) 패턴 분석과 템플릿 문턱치를 도입하여 조기수축 부정맥 분류 방법을 제안한다. 이를 위해 형태 연산을 통한 전 처리 과정과 차감 동작 기법을 통해 R파를 검출하였다. 이후 RR 간격의 평균 가중치와 변화율을 이용하여 먼저 조기수축 파형의 패턴을 분류하고, R파의 진폭에 대한 템플릿 문턱값을 통해 조기심실수축과 조기심방수축을 분류하는 알고리즘을 개발하였다. 제안한 방법의 우수성을 입증하기 위해 조기 심방과 심실수축이 30개 이상 포함된 MIT-BIH 6개의 레코드를 대상으로 한 R파의 평균 검출율은 99.77%의 성능을 나타내었고, 조기심실수축과 심방수축 부정맥은 각각 94.91%와 95.76%의 평균 분류율을 나타내었다.

Most methods for detecting arrhythmia require pp interval, diversity of P wave morphology, but it is difficult to detect the p wave signal because of various noise types. Therefore it is necessary to use noise-free R wave. In this paper, we propose algorithm for premature contraction arrhythmia classification through ECG pattern analysis and template threshold. For this purpose, we detected R wave through the preprocessing method using morphological filter, subtractive operation method. Also, we developed algorithm to classify premature contraction wave pattern using weighted average, premature ventricular contraction(PVC) and atrial premature contraction(APC) through template threshold for R wave amplitude. The performance of R wave detection, PVC classification is evaluated by using 6 record of MIT-BIH arrhythmia database that included over 30 PVC and APC. The achieved scores indicate the average of 99.77% in R wave detection and the rate of 94.91%, 95.76% in PVC and APC classification.

 
페이지 저장