년 - 년
Word2vec과 앙상블 분류기를 사용한 효율적 한국어 감성 분류 방안
[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.19 No.1 2018 pp.133-140
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
감성 분석에서 정확한 감성 분류는 중요한 연구 주제이다. 본 연구는 최근 많은 연구가 이루어지는 word2vec과 앙상블 방법을 이용하여 효과적으로 한국어 리뷰를 감성 분류하는 방법을 제시한다. 연구는 20 만 개의 한국 영화 리뷰 텍스트에 대해, 품사 기반 BOW 자질과 word2vec를 사용한 자질을 생성하고, 두 개의 자질 표현을 결합한 통합 자질을 생성했다. 감성 분류를 위해 Logistic Regression, Decision Tree, Naive Bayes, Support Vector Machine의 단일 분류기와 Adaptive Boost, Bagging, Gradient Boosting, Random Forest의 앙상블 분류기를 사용하였다. 연구 결과로 형용사와 부사를 포함한 BOW자질과 word2vec자질로 구성된 통합 자질 표현이 가장 높은 감성 분류 정확도를 보였다. 실증결과, 단일 분류기인 SVM이 가장 높은 성능을 나타내었지만, 앙상블 분류기는 단일 분류기와 비슷하거나 약간 낮은 성능을 보였다.
Accurate sentiment classification is an important research topic in sentiment analysis. This study suggests an efficient classification method of Korean sentiment using word2vec and ensemble methods which have been recently studied variously. For the 200,000 Korean movie review texts, we generate a POS-based BOW feature and a feature using word2vec, and integrated features of two feature representation. We used a single classifier of Logistic Regression, Decision Tree, Naive Bayes, and Support Vector Machine and an ensemble classifier of Adaptive Boost, Bagging, Gradient Boosting, and Random Forest for sentiment classification. As a result of this study, the integrated feature representation composed of BOW feature including adjective and adverb and word2vec feature showed the highest sentiment classification accuracy. Empirical results show that SVM, a single classifier, has the highest performance but ensemble classifiers show similar or slightly lower performance than the single classifier.
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.5 No.1 2012.03 pp.21-36
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Ensemble classifier is a combining approach to improve the accuracy of the simple classifiers. In this article, we introduced a new method for static handwritten signature verification based on an ensemble classifier. In our introduced method, after pre-processing stage, signature image is convolved with Gabor wavelets to compute the Gabor coefficients in different scales and directions. Three different feature sets are extracted from resulting Gabor coefficients using statistical approaches. A nearest neighbor classifier classifies each feature set by an adaptive method. The proposed ensemble classifier combines the output of the three simple classifiers, which are essentially the same. Although these simple classifiers looks the same, but the different input feature set and the adaptive thresholds related to each classifier makes them to be different with each other. Therefore, from the viewpoint of the classifiers combination, the proposed method can be considered as a feature level combination type. The proposed method was evaluated by applying on two datasets: Persian and South African signature datasets. Experimental results shown our proposed method has the lowest error rate in comparison with other methods.
A Type-2 Fuzzy Logic Ensemble SVM Classifier based on Feature Weighting
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.9 No.2 2016.02 pp.323-338
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
A New Ensemble Scheme for Predicting Human Proteins Subcellular Locations
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition vol.3 no.1 2010.03 pp.1-8
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Predicting subcellular localizations of human proteins become crucial, when new unknown proteins sequences do not have significant homology to proteins of known subcellular locations. In this paper, we present a novel approach to develop CE-Hum-PLoc system. Individual classifiers are created by selecting a fixed learning algorithm from a pool of base learners and then trained by varying feature dimensions of Amphiphilic Pseudo Amino Acid Composition. The output of combined ensemble is obtained by fusing the predictions of individual classifiers. Our approach is based on the utilization of diversity in feature and decision spaces. As a demonstration, the predictive performance was evaluated for a benchmark dataset of 12 human proteins subcellular locations. The overall accuracies reach upto 80.83% and 86.69% in jackknife and independent dataset tests, respectively. Our method has given an improved prediction as compared to existing methods for this dataset. Our CEHum-PLoc system can also be a used as a useful tool for prediction of other subcellular locations.
보안공학연구지원센터(IJUNESST) International Journal of u- and e- Service, Science and Technology vol.3 no.3 2010.09 pp.1-12
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
As the ubiquitous computing becomes popular, its applications come to real life as a form of a wide variety of ubiquitous decision support systems (UDSS). However, such ubiquity should be supported by prediction capability no matter which kind of contexts users are in. In this sense, context prediction capability, which is to predict future contexts users are going to enter sooner or later, becomes an extremely important part of ubiquitous decision support systems. This study proposes a new breed of context prediction mechanism using the Markov Blanket obtained from General Bayesian Network (GBN) as a main vehicle. To improve the prediction accuracy, ensemble of robust prediction classifiers is suggested on the basis of the GBN Markov Blanket. Three classifiers included in the ensemble mechanism are Bayesian networks, decision classifiers, and an SVM (Support Vector Machine). The proposed GBN Markov blanket-assisted ensemble classifier is applied to a real dataset of location prediction. Results were promising enough to conclude that the proposed ensemble classifier based on the GBN Markov Blanket is worthwhile for being adopted in developing a powerful context prediction purpose UDSS. Practical implications are also discussed with future research issues.
An Ensemble Classifier using Two Dimensional LDA
[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.13 No.6 2010 pp.817-824
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Linear Discriminant Analysis (LDA) has been successfully applied for dimension reduction in face recognition. However, LDA requires the transformation of a face image to a one-dimensional vector and this process can cause the correlation information among neighboring pixels to be disregarded. On the other hand, 2D-LDA uses 2D images directly without a transformation process and it has been shown to be superior to the traditional LDA. Nevertheless, there are some problems in 2D-LDA. First, it is difficult to determine the optimal number of feature vectors in a reduced dimensional space. Second, the size of rectangular windows used in 2D-LDA makes strong impacts on classification accuracies but there is no reliable way to determine an optimal window size. In this paper, we propose a new algorithm to overcome those problems in 2D-LDA. We adopt an ensemble approach which combines several classifiers obtained by utilizing various window sizes. And a practical method to determine the number of feature vectors is also presented. Experimental results demonstrate that the proposed method can overcome the difficulties with choosing an optimal window size and the number of feature vectors.
The Stock Index Prediction by the Time-Adjusted Classifier Ensemble
[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.17 No.4 2015.08 pp.1747-1758
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
There are many time-related data in the world. The properties of the attributes change and they affect the target variable differently as time changes. For example, the oil price was the major indicator for the stock price several years ago. However, the gold price is more influential than the oil price to predict the stock price nowadays. Therefore, it is reasonable to give more weights for the recent instances when we predict the target variable from the past data set. Bagging and random forests use the bootstrapped training set to construct the algorithm. Here, we introduce the modified bootstrap method, the time-adjusted bootstrap which gives the recent instance more weight to be chosen instead of the equal weight for the conventional bootstrap. We experiment with financial data set to predict the movement of KOSPI (Korea composite stock price index). We compare them for 66 cases by the combination of the different data sets or the different degree of the weights for the robust conclusion. For 62 out of 66 cases, our method improves the accuracy over the original bootstrap and the average improvement rates are 3.2% and 2.0% for bagging and random forests respectively. Consequently, bagging with the modified bootstrap shows the better performance over the original bagging and random forests with the modified bootstrap also shows the better performance over the original random forests.
A Comparative Study of Phishing Websites Classification Based on Classifier Ensemble
[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.21 No.5 2018 pp.617-625
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Phishing website has become a crucial concern in cyber security applications. It is performed by fraudulently deceiving users with the aim of obtaining their sensitive information such as bank account information, credit card, username, and password. The threat has led to huge losses to online retailers, e-business platform, financial institutions, and to name but a few. One way to build anti-phishing detection mechanism is to construct classification algorithm based on machine learning techniques. The objective of this paper is to compare different classifier ensemble approaches, i.e. random forest, rotation forest, gradient boosted machine, and extreme gradient boosting against single classifiers, i.e. decision tree, classification and regression tree, and credal decision tree in the case of website phishing. Area under ROC curve (AUC) is employed as a performance metric, whilst statistical tests are used as baseline indicator of significance evaluation among classifiers. The paper contributes the existing literature on making a benchmark of classifier ensembles for web phishing detection.
[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.27 No.3 2022 pp.33-43
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 기존의 스마트폰 카메라 센서를 사용하여 비접촉식 손바닥 기반 사용자 식별 시스템을 구축하기 위해 Attention U-Net 모델과 사전 훈련된 컨볼루션 신경망(CNN)이 있는 다채널 손바닥 이미지를 이용한 앙상블 모델을 제안한다. Attention U-Net 모델은 손바닥(손가락 포함), 손바닥(손바닥 미포함) 및 손금을 포함한 관심 영역을 추출하는 데 사용되며, 이는 앙상블 분류기로 입력되는 멀티채널 이미지를 생성하기 위해 결합 된다. 생성된 데이터는 제안된 손바닥 정보 기반 사용자 식별 시스템에 입력되며 사전 훈련된 CNN 모델 3개를 앙상블 한 분류기를 사용하여 클래스를 예측한다. 제안된 모델은 각각 98.60%, 98.61%, 98.61%, 98.61%의 분류 정확도, 정밀도, 재현율, F1-Score를 달성할 수 있음을 입증하며, 이는 저렴한 이미지 센서를 사용하고 있음에도 불구하고 제안된 모델이 효과적이라는 것을 나타낸다. 본 논문에서 제안하는 모델은 COVID-19 펜데믹 상황에서 기존 시스템에 비하여 높은 안전성과 신뢰성으로 대안이 될 수 있다.
In this paper, we propose an ensemble model facilitated by multi-channel palm images with attention U-Net models and pretrained convolutional neural networks (CNNs) for establishing a contactless palm-based user identification system using conventional inexpensive camera sensors. Attention U-Net models are used to extract the areas of interest including hands (i.e., with fingers), palms (i.e., without fingers) and palm lines, which are combined to generate three channels being ped into the ensemble classifier. Then, the proposed palm information-based user identification system predicts the class using the classifier ensemble with three outperforming pre-trained CNN models. The proposed model demonstrates that the proposed model could achieve the classification accuracy, precision, recall, F1-score of 98.60%, 98.61%, 98.61%, 98.61% respectively, which indicate that the proposed model is effective even though we are using very cheap and inexpensive image sensors. We believe that in this COVID-19 pandemic circumstances, the proposed palm-based contactless user identification system can be an alternative, with high safety and reliability, compared with currently overwhelming contact-based systems.
[NRF 연계] 서울대학교 인지과학연구소 Journal of Cognitive Science Vol.24 No.3 2023.09 pp.313-336
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Prediction of software bugs from process logs, temporal access logs, behavior analysis, etc. requires estimation of a wide variety of high-density feature sets. Extracted feature sets must be able to classify these logs into different bug categories with high accuracy, and low complexity. To perform these tasks, a wide variety of Machine Learning Models (MLMs) are proposed by researchers, and each of them varies in terms of their performance-level nuances, functional advantages, contextual limitations, and application-specific future scopes. Upon analyzing these characteristics, it was observed that existing models are highly context-specific, and cannot be applied to multidomain bug analysis datasets. Moreover, existing models do not incorporate a dynamic feature selection method, which limits their accuracy performance under multiple bug classification applications. To overcome these issues, this paper proposes design of a novel Ensemble deep learning Classifier with Feature Selection for high-efficiency Multidomain Bug Prediction under different use cases. The proposed model improves bug representation performance by combining multiple feature extraction methods, including GWO-based novel feature selection techniques. The ensemble classification model, which combines Deep Random Forest, k Nearest Neighbor, Logistic Regression, Multilayer Perceptron, Support Vector Machine, and 1D Convolutional Neural Network classifiers, achieves higher accuracy, precision, recall, and low delay compared to existing models. The model also shows faster classification speeds than existing models and can be deployed for various real-time applications. This accuracy performance was compared with various state-of-the-art models, and it was observed that the proposed model showcases 4.5% higher accuracy, 3.2% better precision, 3.9% higher recall, and 5.5% faster classification performance, which was possible due to integration of intelligent feature selection process with high efficiency classification models under multidomain scenarios.
암 분류를 위한 음의 상관관계 특징을 이용한 앙상블 분류기
[Kisti 연계] 한국정보과학회 정보과학회논문지 : 소프트웨어 및 응용 Vol.30 No.12 2003 pp.1124-1134
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근의 DNA 마이크로어레이 기술로 많은 양의 유전자 데이타를 얻을 수 있는데, 특히 암의 진단과 치료에 적용되어 암의 정확한 분류에 많은 도움을 줄 것으로 기대된다. DNA로부터 얻어지는 유전자 데이타의 양은 매우 방대하므로 이를 효과적으로 분석하는 것은 매우 중요하다. 암의 분류는 진단과 치료에 있어 매우 중요하므로 하나의 분류기에 의존한 분류 결과보다는 다수의 전문화된 분류기 결과를 결합하여 결과를 도출하는 것이 바람직하다. 일반적으로 분류기를 결합함으로써 분류 성능 및 분류 결과에 대한 신뢰도를 높일 수 있다. 앙상블 분류기의 많은 장점에도 불구하고, 오류 의존적인 분류기의 결합은 성능 향상에 한계가 있다. 본 논문에서는 암을 정확하게 분류하기 위해서 음의 상관관계를 갖는 특징으로 학습한 신경망 분류기를 결합하는 방법을 제안하고, 제안한 방법의 유용성을 체계적으로 분석하고자 한다. 세 가지 벤치마크 암 데이타에 대하여 제안한 방법을 적용하여 실험한 결과, 음의 상관관계 특징을 이용한 앙상블 분류기가 다른 분류기보다 높은 성능을 내는 것을 확인할 수 있었다.
The development of microarray technology has supplied a large volume of data to many fields. In particular, it has been applied to prediction and diagnosis of cancer, so that it expectedly helps us to exactly predict and diagnose cancer. It is essential to efficiently analyze DNA microarray data because the amount of DNA microarray data is usually very large. Since accurate classification of cancer is very important issue for treatment of cancer, it is desirable to make a decision by combining the results of various expert classifiers rather than by depending on the result of only one classifier. Generally combining classifiers gives high performance and high confidence. In spite of many advantages of ensemble classifiers, ensemble with mutually error-correlated classifiers has a limit in the performance. In this paper, we propose the ensemble of neural network classifiers learned from negatively correlated features using three benchmark datasets to precisely classify cancer, and systematically evaluate the performances of the proposed method. Experimental results show that the ensemble classifier with negatively correlated features produces the best recognition rate on the three benchmark datasets.
차량 번호판 인식을 위한 앙상블 학습기 기반의 최적 특징 선택 방법
[Kisti 연계] 대한전기학회 電氣學會論文誌 Vol.65 No.1 2016 pp.142-149
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper proposes a method to detect LP(License Plate) of vehicles in indoor and outdoor parking lots. In restricted environment, there are many conventional methods for detecting LP. But, it is difficult to detect LP in natural and complex scenes with background clutters because several patterns similar with text or LP always exist in complicated backgrounds. To verify the performance of LP text detection in natural images, we apply MB-LGP feature by combining with ensemble machine learning algorithm in purpose of selecting optimal features of small number in huge pool. The feature selection is performed by adaptive boosting algorithm that shows great performance in minimum false positive detection ratio and in computing time when combined with cascade approach. MSER is used to provide initial text regions of vehicle LP. Throughout the experiment using real images, the proposed method functions robustly extracting LP in natural scene as well as the controlled environment.
유전자 알고리즘을 이용한 림프종 암의 최적 분류기 앙상블
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2003 pp.356-358
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
DNA microarray기술의 발달로 한꺼번에 수천 개 유전자의 발현 정보를 얻는 것이 가능해졌는데, 이렇게 얻어진 데이터를 효과적으로 분류하는 시스템을 만들어놓으면 새로운 샘플이 정상상태인지, 질병을 가진 상태인지 예측할 수 있다. 분류 시스템을 위하여 여러 가지 특징선택방법들과 분류기법들을 사용할 수 있는데, 모든 상황에서 항상 뛰어난 성능을 보이는 특징선택법이나 분류기를 찾기는 힘들다. 안정되고 개선된 성능을 내기 위해서 특징-분류기의 앙상블을 이용할 수 있는데, 앙상블에 이용될 수 있는 특징선택 방법이나 분류기의 수가 많다면, 앙상블을 만들 수 있는 조합이 많아지기 때문에, 모든 조합에 대하여 앙상블 결과를 구하기는 거의 불가능하다. 이를 해결하기 위하여 본 논문에서는 유전자알고리즘을 이용하여 모든 앙상블 결과를 계산하지 않으면서 최적의 앙상블을 찾아내는 방법을 제안하였으며, 실제로 림프종 암 데이터에 적용한 결과 100%의 결합결과를 보이는 최적의 앙상블을 효과적으로 찾아내었다.
군집화와 유전 알고리즘을 이용한 거친-섬세한 분류기 앙상블 선택
[Kisti 연계] 한국정보과학회 정보과학회논문지 : 소프트웨어 및 응용 Vol.34 No.9 2007 pp.857-868
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
좋은 분류기 앙상블은 분류기간에 상호 보완성을 갖추어 높은 인식 성능을 보여야 하며, 크기가 작아 계산 효율이 좋아야 한다. 이 논문은 이러한 목적을 달성하기 위한 거친-섬세한 (coarse-to-fine)단계를 밟는 분류기 앙상블 선택 방법을 제안한다. 이 방법이 성공하기 위해서는 초기 분류기 풀 (pool)이 충분히 다양해야 한다. 이 논문에서는 여러 개의 서로 다른 분류 알고리즘과 아주 많은 수의 특징 부분집합을 결합하여 충분히 큰 분류기 풀을 생성한다. 거친 선택 단계에서는 분류기 풀의 크기를 적절하게 줄이는 것이 목적이다. 분류기 군집화 알고리즘을 사용하여 다양성을 최소로 희생하는 조건하에 분류기 풀의 크기를 줄인다. 섬세한 선택에서는 유전 알고리즘을 이용하여 최적의 앙상블을 찾는다. 또한 탐색 성능이 개선된 혼합 유전 알고리즘을 제안한다. 널리 사용되는 필기 숫자 데이타베이스를 이용하여 기존의 단일 단계 방법과 제안한 두 단계 방법의 성능을 비교한 결과 제안한 알고리즘이 우수함을 입증하였다.
The good classifier ensemble should have a high complementarity among classifiers in order to produce a high recognition rate and its size is small in order to be efficient. This paper proposes a classifier ensemble selection algorithm with coarse-to-fine stages. for the algorithm to be successful, the original classifier pool should be sufficiently diverse. This paper produces a large classifier pool by combining several different classification algorithms and lots of feature subsets. The aim of the coarse selection is to reduce the size of classifier pool with little sacrifice of recognition performance. The fine selection finds near-optimal ensemble using genetic algorithms. A hybrid genetic algorithm with improved searching capability is also proposed. The experimentation uses the worldwide handwritten numeral databases. The results showed that the proposed algorithm is superior to the conventional ones.
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2004 pp.265-267
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
경제 여건의 향상 및 생활양식의 변화로 최근 우리나라에서도 당뇨병 환자가 늘어남에 따라 당뇨병의 예측 및 치료가 중요한 관심사가 되고 있다. 본 논문은 1993년과 1995년 두 차례에 걸쳐 경기도 연천 지역 주민들의 여러 가지 신체 지수 등을 조사한 데이터를 대상으로, 1차 년도의 데이터로부터 동일한 환자가 2차 년도에 정상상태를 유지하는지 흑은 당뇨병으로 진행이 되는지를 예측하는 문제를 다룬다. 혈당량, 허리둘레 등의 수치가 당뇨병의 발병에 영향을 끼치는 것은 알려진 사실이므로, 현재의 데이터로부터 앞으로의 발병 가능성을 예측하는 것이 가능하며, 이는 환자에게 보다 정확한 정보를 알려줄 수 있으므로 의미가 있는 일이다. 예측을 위해 본 논문에서는 분류기를 사용하며, 예측율을 높이기 위해 여러 분류기를 BKS로 결합하였다. BKS (behavior knowledge space) 결합 방법은 분류기간의 독립 가정이 필요 없으며, 데이터 크기가 크고 전형적인 경우에 좋은 결과를 낼 수 있는 방법이다. BKS 결합 방법을 통해 실험을 해본 결과 단일 분류기로 실험을 한 결과보다 향상된 성능을 얻을 수 있었으며, 투표 결합 방법과 비교하여 더 좋은 성능을 보였다.
복합 분류기를 이용한 웹 문서 범주화에 관한 실험적 연구
[Kisti 연계] 한국정보관리학회 한국정보관리학회 학술대회논문집 2003 pp.73-82
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구에서는 웹 문서를 분류하기 위해 문서로부터 다양한 자질을 추출하고, 두 가지의 분류기를 통해 여러 개의 분류 예측치를 구한 다음, 그것들을 하나의 결과물로 통합하는 복합분류기를 사용하였다. 먼저 다양한 자질 집합에 대해 일반적으로 많이 사용되는 kNN(k nearest neighbor) 분류기와 나이브 베이즈(Naive Bayes) 분류기를 사용한 범주화 실험을 수행하고, 실험을 통해 나온 범주 예측치를 통합하는 복합 분류기들의 성능을 비교하였다. 또한 단일 분류기들을 통해 나온 모든 범주 예측치를 통합하는 과정을 수행하여, 단일 분류기만을 사용할 경우와 복합 분류기를 사용할 경우를 비교해 더 좋은 성능을 나타내는 분류기를 밝히고자 한다.
[NRF 연계] 한국IT정책경영학회 한국IT정책경영학회 논문지 Vol.11 No.2 2019.04 pp.1187-1193
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
IT기술의 발달로 개인화 서비스 기술이 인기를 끌게 되면서, 인간의 감정 상태를 판단하기 위한 감정 인식 기술이 연구되고 있다. 감정 인식 기술의 한 분야인 음성 감정 인식은 사용자의 음성 데이터를 분석하여 감정을 판단하는 기술이다. 감정 인식 기술은 얼굴, 표정, 뇌파와 같은 다양한 데이터를 기반으로 연구되고 있는데, 그 중에서 음성을 이용한 감정 인식 기술은 사용a자와 직접 대면하지 않고 음성만으로 감정을 인지할 수 있어 적은 공간적, 시간적 비용으로 사용자의 감정 상태를 파악할 수 있다는 장점을 가지고 있다. 본 논문에서는 음성을 이용한 감정 인식 기술의 성능을 높이기 위하여, 머신 러닝 알고리즘 중에서 분류에 높은 성능을 가진 서포트 벡터 머신, 랜덤 포레스트, xgboost 알고리즘으로 감정 모델을 학습하여 분류한 결과를 혼합하여 음성의 감정을 판단하는 방법을 제안한다. 제안한 방법은 한 가지 분류 모델을 사용하여 감정을 인식할 때 보다 최대 5.1% 높은 감정 인식률을 보였다.
As the technology of personalized customer service becomes popular with the development of IT technology, emotion recognition technology is actively being studied. Speech Emotion Recognition is a field of emotion recognition technology to judge emotion state by analyzing user's voice signal. Emotion recognition using voice has the advantage that it can be perceived only by voice without face to face with the user, and the emotion state of the user can be grasped with low spatial and temporal cost. In this paper, we propose a method to classify emotions based on speech by analyzing the extracted features from speech using the combination method with support vector machine, random forest, and xgboost algorithm has high performance for classifying. The proposed method showed a maximum 5.1% higher recognition rate than using one classification model.
유전알고리즘을 이용한 유전자발현 데이타상의 특징-분류기쌍 최적 앙상블 탐색
[Kisti 연계] 한국정보과학회 정보과학회논문지 : 소프트웨어 및 응용 Vol.31 No.4 2004 pp.525-536
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
유전발현 데이타는 생명체의 특정 조직에서 채취한 샘플을 microarray상에서 측정한 것으로, 유전자들의 발현 정도가 수치로 나타난 데이타이다. 일반적으로 정상조직과 이상조직에서 관련 유전자들의 발현정도는 차이를 보이기 때문에, 유전발현 데이타를 통하여 질병을 분류할 수 있다. 이러한 분류에 모든 유전자들이 관여하지는 않으므로 관련 유전자를 선별하는 작업인 특징선택이 필요하며, 선택된 유전자들을 적절히 분류하는 방법이 필요하다. 본 논문에서는 상관계수, 유사도, 정보이론 등에 기반을 둔 7가지 특징선택 방법과 대표적인 6가지 분류기에 대하여 특징-분류기 쌍의 최적 앙상블을 탐색하기 위한 유전자 알고리즘 기반 방법을 제안한다. 두 가지 암 관련 유전자 발현 데이타에 대하여 leave-one-out cross validation을 포함한 실험을 해본 결과, 림프종 데이타와 대장암 데이타 모두 단일 특징-분류기 쌍보다 훨씬 우수한 성능을 보이는 앙상블들을 발견할 수 있었다.
Gene expression profile is numerical data of gene expression level from organism, measured on the microarray. Generally, each specific tissue indicates different expression levels in related genes, so that we can classify disease with gene expression profile. Because all genes are not related to disease, it is needed to select related genes that is called feature selection, and it is needed to classify selected genes properly. This paper Proposes GA based method for searching optimal ensemble of feature-classifier pairs that are composed with seven feature selection methods based on correlation, similarity, and information theory, and six representative classifiers. In experimental results with leave-one-out cross validation on two gene expression Profiles related to cancers, we can find ensembles that produce much superior to all individual feature-classifier fairs for Lymphoma dataset and Colon dataset.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.