년 - 년
산업재해 예측 모델링에서 순서형 분류를 위한 순서형 의사 결정 트리 알고리즘의 응용과 한계 KCI 등재
한국기계항공기술학회(구 한국기계기술학회) 한국기계항공기술학회지(구 한국기계기술학회지) 제27권 제4호 2025.08 pp.560-565
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
This paper reviews ordinal decision tree algorithms for ordinal classification, exploring theoretical foundations, key algorithms (MDT, QMDT), specialized splitting criteria (Ordinal Gini, Weighted Information Gain), and ensemble methods. It discusses applications in healthcare and social sciences, highlighting interpretability and flexibility while acknowledging overfitting and instability. As implications for future research, this study points out advantages such as interpretability and flexibility, and limitations such as overfitting and instability.
An Effective SVM Ensemble Algorithm Based on Different Thresholds of PCA SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.2 2016.02 pp.109-118
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
This paper proposes an effective ensemble classifier, named PCAenSVM, which consists of ten weak Support Vector Machine classifiers based on different Principal Component Analysis thresholds. Those ten base Support Vector Machine classifiers are made up to fulfill classification tasks using Majority Voting strategy. Experiments are made on four UCI data sets and a data set from the Uppsala University to evaluate the performances of PCAenSVM. The results of PCAenSVM are compared with that of LibSVM and EnsembleSVM. Experimental results show that PCAenSVM has better classification accuracy than other two algorithms. Moreover, PCAenSVM has the same confidence level with the LibSVM, and its confidences of accuracy and sensitivity on those five data sets outperform that of the EnsembleSVM.
보안공학연구지원센터(IJUNESST) International Journal of u- and e- Service, Science and Technology vol.3 no.3 2010.09 pp.1-12
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
As the ubiquitous computing becomes popular, its applications come to real life as a form of a wide variety of ubiquitous decision support systems (UDSS). However, such ubiquity should be supported by prediction capability no matter which kind of contexts users are in. In this sense, context prediction capability, which is to predict future contexts users are going to enter sooner or later, becomes an extremely important part of ubiquitous decision support systems. This study proposes a new breed of context prediction mechanism using the Markov Blanket obtained from General Bayesian Network (GBN) as a main vehicle. To improve the prediction accuracy, ensemble of robust prediction classifiers is suggested on the basis of the GBN Markov Blanket. Three classifiers included in the ensemble mechanism are Bayesian networks, decision classifiers, and an SVM (Support Vector Machine). The proposed GBN Markov blanket-assisted ensemble classifier is applied to a real dataset of location prediction. Results were promising enough to conclude that the proposed ensemble classifier based on the GBN Markov Blanket is worthwhile for being adopted in developing a powerful context prediction purpose UDSS. Practical implications are also discussed with future research issues.
Ensemble Methods Applied to Classification Problem KCI 등재후보
국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.11 No.1 2019.02 pp.47-53
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The idea of ensemble learning is to train multiple models, each with the objective to predict or classify a set of results. Most of the errors from a model’s learning are from three main factors: variance, noise, and bias. By using ensemble methods, we’re able to increase the stability of the final model and reduce the errors mentioned previously. By combining many models, we’re able to reduce the variance, even when they are individually not great. In this paper we propose an ensemble model and applied it to classification problem. In iris, Pima indian diabeit and semiconductor fault detection problem, proposed model classifies well compared to traditional single classifier that is logistic regression, SVM and random forest .
Mining Educational Data to Predict Student’s academic Performance using Ensemble Methods SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.8 2016.08 pp.119-136
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Educational data mining has received considerable attention in the last few years. Many data mining techniques are proposed to extract the hidden knowledge from educational data. The extracted knowledge helps the institutions to improve their teaching methods and learning process. All these improvements lead to enhance the performance of the students and the overall educational outputs. In this paper, we propose a new student’s performance prediction model based on data mining techniques with new data attributes/features, which are called student’s behavioral features. These type of features are related to the learner’s interactivity with the e-learning management system. The performance of student’s predictive model is evaluated by set of classifiers, namely; Artificial Neural Network, Naïve Bayesian and Decision tree. In addition, we applied ensemble methods to improve the performance of these classifiers. We used Bagging, Boosting and Random Forest (RF), which are the common ensemble methods used in the literature. The obtained results reveal that there is a strong relationship between learner’s behaviors and their academic achievement. The accuracy of the proposed model using behavioral features achieved up to 22.1% improvement comparing to the results when removing such features and it achieved up to 25.8% accuracy improvement using ensemble methods. By testing the model using newcomer students, the achieved accuracy is more than 80%. This result proves the reliability of the proposed model.
보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology Vol.64 2014.03 pp.101-118
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
This paper discusses one method of clustering a high dimensional dataset using dimensionality reduction and context dependency measures (CDM). First, the dataset is partitioned into a predefined number of clusters using CDM. Then, context dependency measures are combined with several dimensionality reduction techniques and for each choice the data set is clustered again. The results are combined by the cluster ensemble approach. Finally, the Rand index is used to compute the extent to which the clustering of the original dataset (by CDM alone) is preserved by the cluster ensemble approach.
Study on the ensemble methods with kernel ridge regression
[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.23 No.2 2012 pp.375-383
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The purpose of the ensemble methods is to increase the accuracy of prediction through combining many classifiers. According to recent studies, it is proved that random forests and forward stagewise regression have good accuracies in classification problems. However they have great prediction error in separation boundary points because they used decision tree as a base learner. In this study, we use the kernel ridge regression instead of the decision trees in random forests and boosting. The usefulness of our proposed ensemble methods was shown by the simulation results of the prostate cancer and the Boston housing data.
Predicting Stock Liquidity by Using Ensemble Data Mining Methods
[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.21 No.6 2016 pp.9-19
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In finance literature, stock liquidity showing how stocks can be cashed out in the market has received rich attentions from both academicians and practitioners. The reasons are plenty. First, it is known that stock liquidity affects significantly asset pricing. Second, macroeconomic announcements influence liquidity in the stock market. Therefore, stock liquidity itself affects investors' decision and managers' decision as well. Though there exist a great deal of literature about stock liquidity in finance literature, it is quite clear that there are no studies attempting to investigate the stock liquidity issue as one of decision making problems. In finance literature, most of stock liquidity studies had dealt with limited views such as how much it influences stock price, which variables are associated with describing the stock liquidity significantly, etc. However, this paper posits that stock liquidity issue may become a serious decision-making problem, and then be handled by using data mining techniques to estimate its future extent with statistical validity. In this sense, we collected financial data set from a number of manufacturing companies listed in KRX (Korea Exchange) during the period of 2010 to 2013. The reason why we selected dataset from 2010 was to avoid the after-shocks of financial crisis that occurred in 2008. We used Fn-GuidPro system to gather total 5,700 financial data set. Stock liquidity measure was computed by the procedures proposed by Amihud (2002) which is known to show best metrics for showing relationship with daily return. We applied five data mining techniques (or classifiers) such as Bayesian network, support vector machine (SVM), decision tree, neural network, and ensemble method. Bayesian networks include GBN (General Bayesian Network), NBN (Naive BN), TAN (Tree Augmented NBN). Decision tree uses CART and C4.5. Regression result was used as a benchmarking performance. Ensemble method uses two types-integration of two classifiers, and three classifiers. Ensemble method is based on voting for the sake of integrating classifiers. Among the single classifiers, CART showed best performance with 48.2%, compared with 37.18% by regression. Among the ensemble methods, the result from integrating TAN, CART, and SVM was best with 49.25%. Through the additional analysis in individual industries, those relatively stabilized industries like electronic appliances, wholesale & retailing, woods, leather-bags-shoes showed better performance over 50%.
A Study on AI-Based Electricity Demand Forecasting - Focusing on Ensemble and Regression Methods-
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2022 pp.857-859
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 인공지능 기반의 전력 수요 데이터 예측 모델을 구축하고 이를 최종적으로 웹의 형태로 구현하는 것을 목표로 하였다. 기상청 데이터의 기후 요소를 매개변수로 삼아 전력 수요를 예측하고, 그 결과를 가시적으로 시각화하는 것까지의 전 과정을 최대한 간결하게 진행하였다. 추후 한층 더 발전된 모델을 구축할 수 있다면, 전력시장의 효율성과 경제성을 향상시켜 불필요한 에너지 낭비를 미연에 방지할 수 있을 것이라고 기대한다. 나아가 시스템 상용화를 위해 계속 연구 활동에 정진할 수 있을 것이다.
[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.22 No.11 2017 pp.111-116
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper proposes a data mining approach to predicting stock price direction. Stock market fluctuates due to many factors. Therefore, predicting stock price direction has become an important issue in the field of stock market analysis. However, in literature, there are few studies applying data mining approaches to predicting the stock price direction. To contribute to literature, this paper proposes comparing single classifiers and ensemble classifiers. Single classifiers include logistic regression, decision tree, neural network, and support vector machine. Ensemble classifiers we consider are adaboost, random forest, bagging, stacking, and vote. For the sake of experiments, we garnered dataset from Korea Stock Exchange (KRX) ranging from 2008 to 2015. Data mining experiments using WEKA revealed that random forest, one of ensemble classifiers, shows best results in terms of metrics such as AUC (area under the ROC curve) and accuracy.
데이터 전처리와 앙상블 기법을 통한 불균형 데이터의 분류모형 비교 연구
[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.27 No.3 2014 pp.357-371
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 들어 데이터 마이닝의 분류문제에 있어 목표변수의 불균형 문제가 많은 관심을 받고 있다. 이러한 문제를 해결하기 위해, 이전 연구들은 원 자료에 대하여 데이터 전처리 과정을 실시했는데, 전처리 과정에는 목표변수의 다수계급을 소수계급의 비율에 맞게 조정하는 과소표집법, 소수계급을 복원추출하여 다수계급의 비율에 맞게 조정하는 과대표집법, 소수계급에 K-최근접 이웃 방법 등을 활용하여 과대표집법을 적용 후 다수계급에는 과소표집법을 적용한 하이브리드 기법 등이 있다. 또한 앙상블 기법도 이러한 불균형 데이터의 분류 성능을 높일 수 있다고 알려져 있어, 본 논문에서는 데이터의 전처리 과정과 앙상블 기법을 함께 고려한 여러 모형들을 사용하여, 불균형 자료에 대한 이들모형의 분류성능을 비교평가한다.
There are many studies related to imbalanced data in which the class distribution is highly skewed. To address the problem of imbalanced data, previous studies deal with resampling techniques which correct the skewness of the class distribution in each sampled subset by using under-sampling, over-sampling or hybrid-sampling such as SMOTE. Ensemble methods have also alleviated the problem of class imbalanced data. In this paper, we compare around a dozen algorithms that combine the ensemble methods and resampling techniques based on simulated data sets generated by the Backbone model, which can handle the imbalance rate. The results on various real imbalanced data sets are also presented to compare the effectiveness of algorithms. As a result, we highly recommend the resampling technique combining ensemble methods for imbalanced data in which the proportion of the minority class is less than 10%. We also find that each ensemble method has a well-matched sampling technique. The algorithms which combine bagging or random forest ensembles with random undersampling tend to perform well; however, the boosting ensemble appears to perform better with over-sampling. All ensemble methods combined with SMOTE outperform in most situations.
앙상블 학습의 부스팅 방법을 이용한 악의적인 내부자 탐지 기법
[Kisti 연계] 한국정보보호학회 정보보호학회논문지 Vol.32 No.2 2022 pp.267-277
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 클라우드 및 원격 근무 환경의 비중이 증가함에 따라 다양한 정보보안 사고들이 발생하고 있다. 조직의 내부자가 원격 접속으로 기밀 자료에 접근하여 유출을 시도하는 사례가 발생하는 등 내부자 위협이 주요 이슈로 떠오르게 되었다. 이에 따라 내부자 위협을 탐지하기 위해 기계학습 기반의 방법들이 제안되고 있다. 하지만, 기존의 내부자 위협을 탐지하는 기계학습 기반의 방법들은 편향 및 분산 문제와 같이 예측 정확도와 관련된 중요한 요소를 고려하지 않았으며 이에 따라 제한된 성능을 보인다는 한계가 있다. 본 논문에서는 편향 및 분산을 고려하는 부스팅 유형의 앙상블 학습 알고리즘들을 사용하여 악의적인 내부자 탐지 성능을 확인하고 이에 대한 면밀한 분석을 수행하며, 데이터셋의 불균형까지도 고려하여 최종 결과를 판단한다. 앙상블 학습을 이용한 실험을 통해 기존의 단일 학습 모델에 기반한 방법에서 나아가, 편향-분산 트레이드오프를 함께 고려하며 유사하거나 보다 높은 정확도를 달성함을 보인다. 실험 결과에 따르면 배깅과 부스팅 방법을 사용한 앙상블 학습은 98% 이상의 정확도를 보였고, 이는 사용된 단일 학습 모델의 평균 정확도와 비교하면 악의적인 내부자 탐지 성능을 5.62% 향상시킨다.
Due to the increasing proportion of cloud and remote working environments, various information security incidents are occurring. Insider threats have emerged as a major issue, with cases in which corporate insiders attempting to leak confidential data by accessing it remotely. In response, insider threat detection approaches based on machine learning have been developed. However, existing machine learning methods used to detect insider threats do not take biases and variances into account, which leads to limited performance. In this paper, boosting-type ensemble learning algorithms are applied to verify the performance of malicious insider detection, conduct a close analysis, and even consider the imbalance in datasets to determine the final result. Through experiments, we show that using ensemble learning achieves similar or higher accuracy to other existing malicious insider detection approaches while considering bias-variance tradeoff. The experimental results show that ensemble learning using bagging and boosting methods reached an accuracy of over 98%, which improves malicious insider detection performance by 5.62% compared to the average accuracy of single learning models used.
농업기후 장기예측시스템의 보정 및 앙상블 기법 성능 평가
[Kisti 연계] 한국농림기상학회 한국농림기상학회지 Vol.27 No.3 2025 pp.162-176
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 국립농업과학원에서 운영 중인 농업기후 장기예측시스템에 적용된 통계적 보정기법(Quantile Mapping)과 앙상블 평균기법(Simple Composite Method)이 예측 성능에 미치는 영향을 정량적으로 분석하였다. 그 결과 보정기법은 기온 변수의 계통적 오차를 효과적으로 완화하는 데 기여하였으며, 특히 봄·여름철의 음의 편차와 가을철의 양의 편차를 개선하여 평균값 예측 정확도를 향상시켰다. 반면, 강수량은 여름철에 한정된 개선 효과만 관찰되었으며, 전반적인 편차 제거에는 한계가 있었다. RMSE, TCC, AC 등의 지표 분석에서는 강수 변수에서 상대적으로 더 뚜렷한 개선 효과가 나타났으나, 일부 기간에서는 오히려 성능이 저하되기도 하였다. 앙상블 평균기법은 계통적 오차 제거에는 미치지 못하였으나, 변동성 기반 예측 정확도 측면에서 뚜렷한 향상 효과를 보였다. 전 변수에서 RMSE 감소, TCC 및 AC 향상이 관찰되었으며, 특히 리드타임이 길어질수록 앙상블 효과가 더욱 두드러졌다. 강수량과 최저기온에서는 가장 큰 개선 폭이 확인되었다. 종합적으로, 두 기법은 예측 변수의 특성과 계절적 민감도에 따라 서로 다른 장단점을 지니며, 이를 상호보완적으로 병행 적용할 경우 예측 성능 향상에 효과적일 것으로 판단된다. 본 연구의 결과는 향후 농업기후 예보의 정확도 향상과 실용화 기반 확립에 기여할 수 있는 기초 자료로 활용될 수 있을 것으로 기대된다.
This study presents a quantitative evaluation of the impacts of a statistical bias correction method (Quantile Mapping) and an ensemble averaging method (Simple Composite Method) on the prediction performance of the long-range agro-climate forecasting system operated by the National Institute of Agricultural Sciences in Korea. Forecasts are evaluated against observations based on climatological means, lead time, seasonality, and initialization date. The bias correction method effectively reduces systematic bias in temperature variables, particularly mitigating negative biases in spring and summer and positive biases in autumn, thereby improving the accuracy of climatological means. However, its impact on precipitation remains limited to the summer months, showing constrained ability in addressing broader bias structures. While precipitation shows relatively notable improvements in RMSE, TCC, and AC, some deterioration in performance is also observed during extreme rainfall events. The ensemble method does not reduce systematic bias but significantly improves the accuracy of variability. Reductions in RMSE and increases in TCC and AC are observed across all variables, with more pronounced improvements at longer lead times. The largest gains appear in precipitation and minimum temperature forecasts. Overall, the two methods exhibit complementary strengths depending on variable characteristics and seasonal sensitivity. Their combined application is expected to enhance prediction skill more effectively. The findings provide valuable insights for improving the accuracy and operational applicability of long-range agro-climate forecasts.
임의변수선택 기반 앙상블 판별분석에서 변수의 상대적 중요도에 관한 연구
[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.16 No.6 2014.12 pp.2967-2978
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
이범주 판별문제에서 단일판별기보다 모형의 안정성 및 정밀도를 높이기 위한 방법론 중의 하나가 앙상블 기법이다. 앙상블 기법을 적용하는 경우에는 모형의 예측구조가 복잡하여 각 설명변수의 역할과 중요성을 확인하기 어려워 결국에는 예측결과의 해석력이 떨어진다는 단점을 가지고 있다. 본 논문에서는 임의변수선택 기반 앙상블 판별분석에서 변수의 상대적 중요도를 연구하고자 한다. 먼저 신용평점화에서 핵심기법으로 사용되는 로직스틱 회귀모형을 결합하는 임의변수선택을 이용한 앙상블 기법을 설명하고, 다음으로 각 설명변수의 역할과 상대적 중요도를 측정할 수 있는 평균 z-스코어를 이용하는 방법을 제안하여 예측결과에 대한 해석력을 살펴보고자 한다. 본 논문에서 제안한 방법론의 유한표본 성질을 규명하고 응용성을 확인하기 위하여 모의실험과 실제자료 분석을 이용한 연구를 수행하고자 한다.
A method for enhancing stability and precision of a classification method is ensemble method. Typically, ensemble method outperforms a single classifier in the binary classification. The prediction structure of the ensemble model is more complex than a single base learner so that it is difficult to identify the role and importance of the explanatory variables. Ensemble method has a problem in interpretation of the prediction result since the interpretability of the prediction result of an ensemble method may be reduced. Hence, it is hard to achieve both the performance improvement in the precision and the result understanding in the interpretation of the model because of being many variables in classification model. This paper considers a methodology to solve the interpretation problem for an ensemble method. First, we explain ensemble method using random predictor selection which combines logistic regression model, a technique being used in credit scoring. Next, as a measure for the relative importance, we adopt the mean z-score that measures role and relative importance of the explanatory variables and examine the interpretation ability of the prediction result. In order to illustrate the finite-sample performance of the considered methodology, we conduct a numerical study using both a simulated data set and real data set.
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2016 pp.401-403
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 태양광 발전사업의 투자 수요가 증가하고 있으며, 이에 따른 태양광 발전시스템 (PV시스템)의 신뢰성 및 발전 효율 향상 등을 확보할 수 있는 모니터링 시스템의 중요성이 부각되고 있다. 본 논문에서는 데이터를 앙상블 기법으로 분석하여 알려진 자동 분류 기법과 앙상블 기법을 비교해보고, 이를 바탕으로 PV시스템 고장 예측의 정확도를 향상 시키고자 한다.
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2016 pp.401-403
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 태양광 발전사업의 투자 수요가 증가하고 있으며, 이에 따른 태양광 발전시스템 (PV시스템)의 신뢰성 및 발전 효율 향상 등을 확보할 수 있는 모니터링 시스템의 중요성이 부각되고 있다. 본 논문에서는 데이터를 앙상블 기법으로 분석하여 알려진 자동 분류 기법과 앙상블 기법을 비교해보고, 이를 바탕으로 PV시스템 고장 예측의 정확도를 향상 시키고자 한다.
회계정보 기반 ESG 성과 예측에 대한 탐색적 연구 -Ensemble 기법을 중점으로-
[NRF 연계] 한국유통물류정책학회 유통물류연구 Vol.11 No.3 2024.09 pp.67-83
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 기업의 회계정보를 활용하여 ESG(환경, 사회, 지배구조) 성과를 예측하는 모델을 개발하는 것을 목표로 한다. ESG 성과는 기업의 지속 가능성과 사회적 책임을 평가하는 중요한 지표로서, 투자자와 이해관계자들에게 중요한 평가 기준이 되고 있다. 연구는 두 가지 주요 단계를 통해 진행되었다. 첫 번째 단계는데이터 전처리 단계로, ESG 성과 예측을 위해 427개의 회계 정보를 수집하고 정제하였다. 이 정보는 기업의재무 상태와 성과를 포괄적으로 반영하는 변수로 구성되었으며, 재무상태표, 손익계산서, 현금흐름표에서추출되었다. 두 번째 단계는 분석 단계로, 다양한 데이터마이닝 기법과 앙상블 기법을 적용하여 ESG 성과 예측 모델을개발하고 그 성능을 평가하였다. 데이터마이닝 기법은 대규모 데이터셋에서 패턴과 규칙을 발견하는 데 유용하며, 본 연구에서는 기본 데이터마이닝 기법과 앙상블 기법을 모두 활용하였다. 분석 결과, Random Subspace with DT(J48) 기법이 정확도(0.854), 정밀도(0.877), F-measure(0.888)에서 가장높은 값을 보이는 것으로 나타났다. 또한, 앙상블 기법이 단일 모델보다 예측 성능을 향상시키는 데 효과적임이 나타났다. Random Subspace with DT(J48) 기법을 통해 약 85.4%의 정확도로 ESG 성과를 예측할 수 있었다. 본 연구는 ESG 성과 예측 모델이 기업의 지속 가능성 평가 및 투자 결정에 유용한 정보를 제공할 수 있음을시사한다. 이 연구는 ESG 성과 예측이 기업의 장기적인 성공과 지속 가능성에 중요한 역할을 하며, 이를 예측하고 관리하는 것이 필수적임을 시사한다. 다양한 데이터마이닝 기법과 앙상블 기법을 활용한 본 연구는 ESG 성과예측의 정확도를 높이고, 실무적 적용 가능성을 증대시키는 데 기여할 것으로 기대된다.
This study aims to develop a model for predicting ESG (Environmental, Social, and Governance) performance usingcorporate accounting information. ESG performance is a critical indicator for assessing corporate sustainability andsocial responsibility, increasingly becoming a key evaluation criterion for investors and stakeholders. The study wasconducted in two primary stages. The first stage involved data preprocessing, where 427 indicators of accountinginformation were collected and refined to predict ESG performance. This data was extracted from financial statements,income statements, and cash flow statements, providing a comprehensive reflection of the companies’ financial statusand performance. The second stage focused on analysis, employing various data mining techniques and ensemble methods to developand evaluate the ESG performance prediction models. Data mining techniques are invaluable for identifying patternsand rules within large datasets. In this research, both basic data mining techniques and ensemble methods were utilized. The results indicated that the Random Subspace with DT (J48) method exhibited the highest values in terms of accuracy(0.854), precision (0.877), and F-measure (0.888). Additionally, ensemble methods demonstrated improved predictiveperformance compared to single models. This study suggests that ESG performance prediction models can provide valuable insights for evaluating corporatesustainability and making informed investment decisions. The research underscores the significant role of ESGperformance prediction in ensuring long-term corporate success and sustainability, highlighting the necessity ofpredicting and managing ESG performance. By employing various data mining techniques and ensemble methods, thisstudy contributes to enhancing the accuracy of ESG performance predictions and increasing their practical applicability.
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2023 pp.689-691
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
불균형 데이터의 분류의 성능을 향상시키기 위한 앙상블 구성 방법에 관하여 연구한다. 앙상블의 성능은 앙상블을 구성한 기계학습 모델 간의 상호 다양성에 큰 영향을 받는다. 기존 방법에서는 앙상블에 속할 모델 간의 상호 다양성을 높이기 위해 Feature Engineering 을 사용하여 다양한 모델을 만들어 사용하였다. 그럼에도 생성된 모델 가운데 유사한 모델들이 존재하며 이는 상호 다양성을 낮추고 앙상블 성능을 저하시키는 문제를 가지고 있다. 불균형 데이터의 경우에는 유사 모델 판별을 위한 기존 다양성 지표가 다수 클래스에 편향된 수치를 산출하기 때문에 적합하지 않다. 본 논문에서는 기존 다양성 지표를 개선하고 가지치기 방안을 결합하여 유사 모델을 판별하고 상호 다양성이 높은 후보 모델들을 앙상블에 포함시키는 방법을 제안한다. 실험 결과로써 제안한 방법으로 구성된 앙상블이 불균형이 심한 데이터의 분류 성능을 향상시킴을 확인하였다.
분류 앙상블 모형에서 Lasso-bagging과 WAVE-bagging 가지치기 방법의 성능비교
[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.25 No.6 2014 pp.1371-1383
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
분류 앙상블 모형이란 여러 분류기들의 예측 결과를 통합하여 더욱 정교한 예측성능을 가진 분류기를 만들기 위한 융합방법론이라 할 수 있다. 분류 앙상블을 구성하는 분류기들이 높은 예측 정확도를 가지고 있으면서 서로 상이한 모형으로 이루어져 있을 때 분류 앙상블 모형의 정확도가 높다고 알려져 있다. 하지만, 실제 분류 앙상블 모형에는 예측 정확도가 그다지 높지 않으며 서로 유사한 분류기도 포함되어 있기 마련이다. 따라서 분류 앙상블 모형을 구성하고 있는 여러 분류기들 중에서 서로 상이하면서도 정확도가 높은 것만을 선택하여 앙상블 모형을 구성해 보는 가지치기 방법을 생각할 수 있다. 본 연구에서는 Lasso 회귀분석 방법을 이용하여 분류기 중에 일부를 선택하여 모형을 만드는 방법과 가중 투표 앙상블 방법론의 하나인 WAVE-bagging을 이용하여 분류기 중 일부를 선택하는 앙상블 가지치기 방법을 비교하였다. 26개 자료에 대해 실험을 한 결과 WAVE-bagging 방법을 이용한 분류 앙상블 가지치기 방법이 Lasso-bagging을 이용한 방법보다 더 우수함을 보였다.
Classification ensemble technique is a method to combine diverse classifiers to enhance the accuracy of the classification. It is known that an ensemble method is successful when the classifiers that participate in the ensemble are accurate and diverse. However, it is common that an ensemble includes less accurate and similar classifiers as well as accurate and diverse ones. Ensemble pruning method is developed to construct an ensemble of classifiers by choosing accurate and diverse classifiers only. In this article, we proposed an ensemble pruning method called WAVE-bagging. We also compared the results of WAVE-bagging with that of the existing pruning method called Lasso-bagging. We showed that WAVE-bagging method performed better than Lasso-bagging by the extensive empirical comparison using 26 real dataset.
앙상블 기법을 이용한 잡음 환경에서의 화자인식 방법에 관한 연구
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2018 pp.457-459
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
화자인식 기술은 주어진 임의의 두 발화로부터 발화자의 일치 여부를 판단하여 등록된 화자의 목록으로부터 임의로 입력된 발화의 발화자를 식별하는 기술이다. 그러나, 배경잡음이나 반향이 존재하는 경우에는 음성신호가 왜곡되어 화자인식 성능이 저하될 수 있기 때문에 별도의 음성신호 전처리 알고리즘을 함께 사용할 수 있다. 본 논문에서는 배경잡음이 존재하는 환경에서 다수의 마이크로폰을 통해 수집한 음성신호에 대해 화자인식을 수행하는 방법으로써 parametric multi-channel Wiener filter (PMWF)를 이용한 화자일치 점수 앙상블 기법을 제안한다. 입력신호의 신호대잡음비를 기준으로 점수 결합 시 사용되는 결합계수를 정하고, Wiener filter 로 잡음을 제거하여 얻은 점수와 minimum variance distortionless response (MVDR) 빔포머를 통해 잡음을 제거하여 얻은 정수를 가중결합하는 방식으로 동일오류율을 측정한 결과, 각 전처리 알고리즘을 독립적으로 사용하여 점수를 계산한 경우보다 우수한 성능을 보임을 확인할 수 있었다.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.