년 - 년
머신러닝 기반 건물 에너지 소비량 예측 및 영향 요인 분석 KCI 등재후보
제주대학교 지능소프트웨어 교육연구소 지능정보융합과 미래교육 제4권 제19호 2025.09 pp.1-8
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 건물의 에너지 소비 효율성 증대 및 최적 관리를 목표로 머신러닝 기술을 활용한 전기 에너지 소비량 예측 모델을 개발하고, 소비에 영향을 미치는 주요 요인을 분석한다. 이를 위해 Kaggle에서 제공된 건물 에너지 소비량 데이터를 기반으로 시간적 특성을 반 영한 피처 엔지니어링과 데이터 정제 과정을 수행했다. 모델의 예측 성능을 높이기 위해 4개의 개별 머신러닝 모델(Ridge, Random Forest, XGBoost, LightGBM)과 이들의 강점을 결합한 스태킹 앙상블(Stacking Ensemble) 모델을 구축했다. 특히, LightGBM과 XGBoost 모델은 베이지안 최적화 프레임워크인 Optuna 내에서 교차 검증을 수행하는 튜닝 기법을 적용하여 잠재적 성능을 향상시켰 다. 성능 평가 결과, 개별 모델들의 예측 결과를 종합하여 최종 예측을 수행하는 스태킹 앙상블 모델이 R2 score 0.613, RMSE 5.037로 가장 우수한 성능을 보였다. 이는 안정적인 선형 모델(Ridge)과 복잡한 비선형 패턴을 학습하는 부스팅 모델(LightGBM, XGBoost)의 시너지가 예측 정확도 향상에 기여했음을 나타낸다. 또한, 설명가능 인공지능(XAI) 기법인 SHAP 분석을 통해 온도(Temperature), 냉난방 사용여부(HVACUsage), 재실 인원(Occupancy) 등이 에너지 소비량에 가장 큰 영향을 미치는 핵심 요인임을 규명하였다. 본 연구에서 개발된 모델과 분석 결과는 건물의 에너지 효율을 최적화하고 관련 정책을 수립하는데 중요한 기초 자료로 활용될 수 있다.
This study aims to develop a machine learning-based electrical energy consumption prediction model and analyze the key influencing factors to enhance the energy efficiency and optimal management of buildings. Based on building energy consumption data provided by Kaggle, advanced feature engineering and data refinement processes were performed to reflect temporal characteristics. To maximize the predictive performance of the models, five individual machine learning models (Ridge, Random Forest, Decision Tree, XGBoost, LightGBM) and a Stacking Ensemble model that combines their strengths were constructed. In particular, the LightGBM and XGBoost models were fine-tuned using a sophisticated method of performing cross-validation within the Bayesian optimization framework, Optuna, to maximize their potential performance. The performance evaluation revealed that the Stacking Ensemble model, which integrates the predictions of the individual models, demonstrated the highest performance with an R2 score of 0.613 and an RMSE of 5.037. This suggests that the synergy between a stable linear model (Ridge) and boosting models (LightGBM, XGBoost) capable of learning complex non-linear patterns contributed to the improved prediction accuracy. Furthermore, through SHAP analysis, an XAI technique, Temperature, HVACUsae, Occupancy were identified as the key factors most significantly influencing energy consumption, The model and analysis results developed in this study can serve as crucial foundational data for optimizing building energy efficiency and establishing related policies.
스태킹 앙상블 기법을 활용한 고속도로 교통정보 예측모델 개발 및 교차검증에 따른 성능 비교 KCI 등재
한국ITS학회 한국ITS학회논문지 제22권 제6호 통권110호 2023.12 pp.1-16
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
정확도가 높은 교통정보 예측은 지능형교통체계(intelligent transport systems, ITS)를 통한 교통 시설 이용자들의 혼잡 경로 회피 안내 등에서 활용되는 중요한 기능이다. 정확한 교통정보 예측을 위해 다양한 딥러닝 모델들이 발전되어 왔다. 최근에는 앙상블 기법을 활용하여 다양 한 모델들의 장단점을 결합하여 예측 정확도와 안정성을 높이고 있다. 따라서, 본 연구에서는 다양한 딥러닝 모델들을 활용하여 교통정보 예측 모델을 개발하였으며, 개발된 딥러닝 모델들 을 스태킹 앙상블(stacking ensemble)하여 성능을 개선하였다. 개별 모델들은 교통량 예측에서 10% 이내의 오차율을, 속도 예측에서 3% 이내의 오차율을 보였다. 앙상블 모델은 교차검증을 수행하지 않았을 때, 타 모델과 비교하여 더욱 높은 정확도를 보였다. 교차검증을 수행한 앙상 블 모델은 장기예측에서 타 모델보다 균일한 오차율을 보이는 것으로 나타났다.
Accurate traffic information prediction is considered to be one of the most important aspects of intelligent transport systems(ITS), as it can be used to guide users of transportation facilities to avoid congested routes. Various deep learning models have been developed for accurate traffic prediction. Recently, ensemble techniques have been utilized to combine the strengths and weaknesses of various models in various ways to improve prediction accuracy and stability. Therefore, in this study, we developed and evaluated a traffic information prediction model using various deep learning models, and evaluated the performance of the developed deep learning models as a stacking ensemble. The individual models showed error rates within 10% for traffic volume prediction and 3% for speed prediction. The ensemble model showed higher accuracy compared to other models when no cross-validation was performed, and when cross-validation was performed, it showed a uniform error rate in long-term forecasting.
A Stacking Ensemble System for Cloud Task Failure Prediction
국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.17 No.4 2025.11 pp.248-255
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Cloud computing has become a critical infrastructure for large-scale data processing and machine learning applications. However, task failures frequently occur due to distributed resources and dynamic operating environments, leading to service delays and resource waste. To address this issue, this paper proposes a task failure prediction system based on stacking ensemble learning. The proposed system is structured in three layers: log collection, prediction, and result application. Base model predictions are integrated by a metamodel to produce the final outcome. Experiments conducted with the Google Cluster Trace 2019 dataset demonstrate that the proposed system outperforms single models in terms of prediction accuracy and stability, providing a robust and scalable framework suitable for real-world cloud environments.
Stacking Ensemble Learning을 활용한 블록 탑재 시수 예측
[Kisti 연계] 대한조선학회 대한조선학회지 Vol.56 No.6 2019 pp.488-496
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The estimation of block erection work time at a dock is one of the important factors when establishing or managing the total shipbuilding schedule. In order to predict the work time, it is a natural approach that the existing block erection data would be used to solve the problem. Generally the work time per unit is the product of coefficient value, quantity, and product value. Previously, the work time per unit is determined statistically by unit load data. However, we estimate the work time per unit through work time coefficient value from series ships using machine learning. In machine learning, the outcome depends mainly on how the training data is organized. Therefore, in this study, we use 'Feature Engineering' to determine which one should be used as features, and to check their influence on the result. In order to get the coefficient value of each block, we try to solve this problem through the Ensemble learning methods which is actively used nowadays. Among the many techniques of Ensemble learning, the final model is constructed by Stacking Ensemble techniques, consisting of the existing Ensemble models (Decision Tree, Random Forest, Gradient Boost, Square Loss Gradient Boost, XG Boost), and the accuracy is maximized by selecting three candidates among all models. Finally, the results of this study are verified by the predicted total work time for one ship among the same series.
Stacking Ensemble Learning을 활용한 블록 탑재 시수 예측
[NRF 연계] 대한조선학회 대한조선학회논문집 Vol.56 No.6 2019.12 pp.488-496
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The estimation of block erection work time at a dock is one of the important factors when establishing or managing the total shipbuilding schedule. In order to predict the work time, it is a natural approach that the existing block erection data would be used to solve the problem. Generally the work time per unit is the product of coefficient value, quantity, and product value. Previously, the work time per unit is determined statistically by unit load data. However, we estimate the work time per unit through work time coefficient value from series ships using machine learning. In machine learning, the outcome depends mainly on how the training data is organized. Therefore, in this study, we use ‘Feature Engineering’ to determine which one should be used as features, and to check their influence on the result. In order to get the coefficient value of each block, we try to solve this problem through the Ensemble learning methods which is actively used nowadays. Among the many techniques of Ensemble learning, the final model is constructed by Stacking Ensemble techniques, consisting of the existing Ensemble models (Decision Tree, Random Forest, Gradient Boost, Square Loss Gradient Boost, XG Boost), and the accuracy is maximized by selecting three candidates among all models. Finally, the results of this study are verified by the predicted total work time for one ship among the same series.
Stacking Ensemble Technique for Classifying Breast Cancer
[NRF 연계] 대한의료정보학회 Healthcare Informatics Research Vol.25 No.4 2019.10 pp.283-288
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Objectives: Breast cancer is the second most common cancer among Korean women. Because breast cancer is strongly associated with negative emotional and physical changes, early detection and treatment of breast cancer are very important. As a supporting tool for classifying breast cancer, we tried to identify the best meta-learner model in a stacking ensemble when the same machine learning models for the base learner and meta-learner are used. Methods: We used machine learning models, such as the gradient boosted model, distributed random forest, generalized linear model, and deep neural network in a stacking ensemble. These models were used to construct a base learner, and each of them was used as a meta-learner again. Then, we compared the performance of machine learning models in the meta-learner to determine the best meta-learner model in the stacking ensemble. Results: Experimental results showed that using the GBM as a meta-learner led to higher accuracy than that achieved with any other model for breast cancer data and using the GLM as a meta learner led to low root-meansquared error for both sets of breast cancer data. Conclusions: We compared the performance of every meta-learner model in a stacking ensemble as a supporting tool for classifying breast cancer. The study showed that using specific models as a metalearner resulted in better performance than single classifiers, and using GBM and GLM as a meta-learner is appropriate as a supporting tool for classifying breast cancer data.
Parallel Network Model of Abnormal Respiratory Sound Classification with Stacking Ensemble
[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.26 No.11 2021 pp.21-31
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 코로나(Covid-19)의 영향으로 스마트 헬스케어 관련 산업과 비대면 방식의 원격 진단을 통한 질환 분류 예측 연구의 필요성이 증가하고 있다. 일반적으로 호흡기 질환의 진단은 비용이 많이 들고 숙련된 의료 전문가를 필요로 하여 현실적으로 조기 진단 및 모니터링에 한계가 있다. 따라서, 간단하고 편리한 청진기로부터 수집된 호흡음을 딥러닝 기반 모델을 활용하여 높은 정확도로 분류하고 조기 진단이 필요하다. 본 연구에서는 청진을 통해 수집된 폐음 데이터를 이용하여 이상 호흡음 분류모델을 제안한다. 데이터 전처리로는 대역통과필터(BandPassFilter)방법론을 적용하고 로그 멜 스펙트로그램(Log-Mel Spectrogram)과 Mel Frequency Cepstral Coefficient(MFCC)을 이용하여 폐음의 특징적인 정보를 추출하였다. 추출된 폐음의 특징에 대해서 효과적으로 분류할 수 있는 병렬 합성곱 신경망 네트워크(Parallel CNN network)모델을 제안하고 다양한 머신러닝 분류기(Classifiers)와 결합한 스태킹 앙상블(Stacking Ensemble) 방법론을 이용하여 이상 호흡음을 높은 정확도로 분류하였다. 본 논문에서 제안한 방법은 96.9%의 정확도로 이상 호흡음을 분류하였으며, 기본모델의 결과 대비 정확도가 약 6.1% 향상되었다.
As the COVID-19 pandemic rapidly changes healthcare around the globe, the need for smart healthcare that allows for remote diagnosis is increasing. The current classification of respiratory diseases cost high and requires a face-to-face visit with a skilled medical professional, thus the pandemic significantly hinders monitoring and early diagnosis. Therefore, the ability to accurately classify and diagnose respiratory sound using deep learning-based AI models is essential to modern medicine as a remote alternative to the current stethoscope. In this study, we propose a deep learning-based respiratory sound classification model using data collected from medical experts. The sound data were preprocessed with BandPassFilter, and the relevant respiratory audio features were extracted with Log-Mel Spectrogram and Mel Frequency Cepstral Coefficient (MFCC). Subsequently, a Parallel CNN network model was trained on these two inputs using stacking ensemble techniques combined with various machine learning classifiers to efficiently classify and detect abnormal respiratory sounds with high accuracy. The model proposed in this paper classified abnormal respiratory sounds with an accuracy of 96.9%, which is approximately 6.1% higher than the classification accuracy of baseline model.
Stacking Based Ensemble with Instance Selection in Random KNN
[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.22 No.4 2020.08 pp.1303-1312
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Nearest neighbor classification is a simple and quite effective method. But it has some drawbacks in that it suffers from both so called the curse of dimensionality and computational burden when there are lots of examples. A lot of features in high dimensional data degrade the performance of KNN due to the noisy and redundant variables. Hence some kinds of feature selection mechanism is essential to KNN. Random KNN simplifies the feature selection problem by utilizing an ensemble of KNNs constructed with randomly selected subsets of features. Two kinds of enhancements of the random KNN methodology are proposed in this paper. One is to add an instance selection procedure and the other is to determine the weights by training another combining model instead of simple voting in the random KNN. These two additional procedures applied to the random KNN can improve the classification performance of the original random KNN, which is verified by some data analyses with a real data set.
머신러닝 및 딥러닝 모델의 스태킹 앙상블을 이용한 단기 전력수요 예측에 관한 연구
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2021 pp.566-569
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
전력수요는 월, 요일 및 시간의 계절성(Seasonality)을 보이는 데이터이다. 각 계절성에 따라 특성이 다르기 때문에, 전력수요를 예측하기 위해서는 계절성의 특성을 고려한 다양한 모델을 선정하고, 병합하는 방법이 필요하다. 본 연구에서는 전력수요의 계절성을 고려한 다양한 예측모델을 병합하여 이용할 수 있도록 스태킹 앙상블 적용하고 실험결과를 기술한다. 또한, 162개 도시의 기상 데이터와 인구 데이터를 예측에 이용하는 방법, Regression 모델과 Time-series모델에 입력하는 특징(Feature)의 전처리 방법, 베이지안 최적화를 이용한 머신러닝 및 딥러닝 모델의 하이퍼파라메터 최적화 방법을 제시한다.
연합학습에서 Non-IID 및 불균형 데이터를 고려한 스태킹 앙상블 기반 생존분석 모델
[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.20 No.4 2025 pp.859-866
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
연합학습(Federated Learning)은 데이터를 중앙 서버로 전송하지 않고, 클라이언트에서 학습한 로컬 모델을 FedAVG(Federated Averaging) 방식으로 통합하여 글로벌 모델을 업데이트하는 분산형 학습 방식이다. FedAVG는 클라이언트 간 데이터 분포가 유사하거나 독립적인 경우에는 우수한 성능을 보이지만, 실제 의료환경과 같이 데이터가 비독립적(Non-IID)이거나 불균형한 경우에는 성능 저하 문제가 발생한다. 본 연구에서는 Non-IID 및 불균형 데이터 환경에서도 성능 저하를 최소화하기 위해, 스태킹 앙상블(Stacking Ensemble) 기법을 적용한 생존분석 기반 연합학습 모델을 제안한다. 실험 데이터는 실제 의료 환경을 반영하여 각 클라이언트의 데이터 크기와 분포를 상이하게 구성하였으며, 실험을 통해 제안한 스태킹 앙상블 기반 연합학습 방법이 기존의 FedAVG 방식보다 Non-IID 및 불균형 데이터 환경에서 더 우수한 성능을 나타냄을 확인하였다.
Federated Learning is a decentralized learning approach that updates a global model by aggregating locally trained models on each client using methods such as Federated Averaging (FedAVG), without transmitting data to a central server. While FedAVG shows strong performance when data distributions across clients are similar or independent, it suffers from performance degradation in real-world medical settings where data are often non-independent and identically distributed (Non-IID) or imbalanced. To address this issue, this study proposes a survival analysis-based federated learning model that incorporates a stacking ensemble technique to minimize performance degradation under Non-IID and imbalanced data conditions. The experimental data were designed to reflect real-world medical scenarios, with heterogeneous data sizes and distributions across clients. Experimental results demonstrate that the proposed stacking ensemble-based federated learning method outperforms the conventional FedAVG approach in Non-IID and imbalanced data environments.
한국 유역의 지역화를 통해 유출량 예측을 개선하기 위한 수문학적 후 처리된 스태킹 앙상블 모형
[Kisti 연계] 한국수자원학회 한국수자원학회 학술대회논문집 2021 p.182
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구에서는 1일부터 최대 7일까지의 시간을 두고 남한 전체의 유출량에 대한 예측 모형을 제시하고자 한다. 이를 위하여 LSM (Land Surface Model) 모형을 사용하여 유출량을 모의하였고 이 과정에서 미 계측치에 대한 유출량을 예측하기 위하여 Xgboost (Extreme Gradient Boost)를 활용하여 매개변수를 지역화하였다. 이러한 지역화 기법을 통하여 남한 전체의 유출량에 대한 그리드화 된 유출값을 얻을 수 있었다. 또한 본 연구에서는 기상 예측자료를 유출량에 대한 예측으로 변환하기 위하여 Stacking 앙상블 기반의 수문학적 후처리 기법을 사용하였다. Stacking 앙상블 기법은 Base-learner와 Meta-learner의 조합으로 이루어 지는데 본 연구에서 새롭게 사용되는 패널티 기반의 분위회귀분석 방법론은 기존의 방법론과의 비교에 있어서 유용한 것으로 파악되었다. 결과적으로 본 연구에서는 총 7일의 앞선 시간의 예측에 있어서 한반도 전체의 유출량에서 비교적 짧은 시간에 대한 예측인 1일과 2일에서의 예측은 실질적으로 사용이 가능한 것으로 파악되었다.
당뇨병 발생 예측을 위한 다층 스태킹 앙상블 모델 구축 기법
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2023 pp.426-427
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 현대인의 식습관 및 고령화로 인해 당뇨병 환자의 수가 연간 증가하고 있다. 따라서 현재는 아직 당뇨병이 발생하지 않았더라도 미래에 발생할 가능성 예측의 중요성이 커지고 있다. 기존의 당뇨병 발생 여부 진단 연구는 회귀 분석과 같은 단일 모델을 사용하여 수행된다. 그러나 당뇨병에 영향을 미치는 변수들은 복잡하게 얽혀있어 단일 모델만으로는 패턴을 충분히 학습하기 어렵다. 본 논문에서는 데이터에 적합하게 자동으로 다층 스태킹 앙상블 모델을 구성하는 알고리즘을 이용한 다층 스태킹 앙상블 모델을 제안한다. 제안하는 방법은 성능이 높은 모델들을 기준으로 층을 쌓으며 모델을 구성하며 실험 결과 다른 자동 기계학습 라이브러리와 비교해 F1 score 기준으로 최대 12.89%p의 성능 향상을 보였다.
시각적 특징과 물리적 특징에 기반한 스태킹 앙상블 모델을 이용한 과일의 자동 선별
[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.25 No.10 2022 pp.1386-1394
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
As consumption of high-quality fruits increases and sales and packaging units become smaller, the demand for automatic fruit grading systems is increasing. Compared to other crops, the quality of fruit is determined by visual characteristics such as shape, color, and scratches, rather than just physical size and weight. Accordingly, this study presents a CNN model that can effectively extract and classify the visual features of fruits and a perceptron that classifies fruits using physical features, and proposes a stacking ensemble model that can effectively combine the classification results of these two neural networks. The experiments with AI Hub public data show that the stacking ensemble model is effective for grading fruits. However, the ensemble model does not always improve the performance of classifying all the fruit grading. So, it is necessary to adapt the model according to the kind of fruit.
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2024 pp.846-849
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
다양한 분야에서 QR 코드가 급속도로 확산되면서, QR 코드를 악용하여 사용자를 악성 웹사이트로 리디렉션하는 '큐싱(Qshing)'이라는 새로운 형태의 사이버 범죄가 등장했다. 이에 본 연구에서는 일반화 성능을 향상시키기 위해 교차 검증(CV)을 활용하여 QR 코드 스캔과 관련된 악성 URL을 탐지하도록 설계된 스태킹 앙상블 모델을 제안한다. 이러한 통합은 실제 애플리케이션에서 높은 성능을 기대할 수 있도록 설계되었다. 본 연구는 이 모델이 기존의 연구보다 QR 코드 관련 사이버 위협에 대처하는 보다 효과적인 수단을 제공할 것으로 기대한다.
스태킹 앙상블 모델을 이용한 시간별 지상 오존 공간내삽 정확도 향상
[Kisti 연계] 한국지리정보학회 Journal of the Korean Association of Geographic Information Studies Vol.25 No.3 2022 pp.74-99
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
지상 오존은 차량 및 산업 현장에서 배출된 질소화합물(Nitrogen oxides; NOx)과 휘발성 유기화합물(Volatile Organic Compounds; VOCs)의 광화학 반응을 통해 생성되어 식생 및 인체에 악영향을 끼친다. 국내에서는 실시간 오존 모니터링을 수행하고 있지만 관측소 기반으로, 미관측 지역의 공간 분포 분석에 어려움이 있다. 본 연구에서는 스태킹 앙상블 기법을 활용하여 매시간 남한 지역의 지상 오존 농도를 1.5km의 공간해상도로 공간내삽하였고, 5-fold 교차검증을 수행하였다. 스태킹 앙상블의 베이스 모델로는 코크리깅(Cokriging), 다중 선형 회귀(Multi-Linear Regression; MLR), 랜덤 포레스트(Random Forest; RF), 서포트 벡터 회귀(Support Vector Regression; SVR)를 사용하였다. 각 모델의 정확도 비교 평가 결과, 스태킹 앙상블 모델이 연구 기간 내 시간별 평균 R 및 RMSE이 0.76, 0.0065ppm으로 가장 높은 성능을 보여주었다. 스태킹 앙상블 모델의 지상 오존 농도 지도는 복잡한 지형 및 도시화 변수의 특징이 잘 드러나며 더 넓은 농도 범위를 보여주었다. 개발된 모델은 매시간 공간적으로 연속적인 공간 지도를 산출할 수 있을 뿐만 아니라 8시간 평균치 산출 및 시계열 분석에 있어서도 활용 가능성이 클 것으로 기대된다.
Surface ozone is produced by photochemical reactions of nitrogen oxides(NOx) and volatile organic compounds(VOCs) emitted from vehicles and industrial sites, adversely affecting vegetation and the human body. In South Korea, ozone is monitored in real-time at stations(i.e., point measurements), but it is difficult to monitor and analyze its continuous spatial distribution. In this study, surface ozone concentrations were interpolated to have a spatial resolution of 1.5km every hour using the stacking ensemble technique, followed by a 5-fold cross-validation. Base models for the stacking ensemble were cokriging, multi-linear regression(MLR), random forest(RF), and support vector regression(SVR), while MLR was used as the meta model, having all base model results as additional input variables. The results showed that the stacking ensemble model yielded the better performance than the individual base models, resulting in an averaged R of 0.76 and RMSE of 0.0065ppm during the study period of 2020. The surface ozone concentration distribution generated by the stacking ensemble model had a wider range with a spatial pattern similar with terrain and urbanization variables, compared to those by the base models. Not only should the proposed model be capable of producing the hourly spatial distribution of ozone, but it should also be highly applicable for calculating the daily maximum 8-hour ozone concentrations.
데이터 전처리 및 물리적 변수 변환 기반의 스태킹 앙상블 기법을 활용한 풍력발전 시스템의 발전량 예측 모델 연구
[NRF 연계] 한국풍력에너지학회 풍력에너지저널 Vol.17 No.2 2026.06 pp.27-41
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This study aimed to establish an efficient and accurate power generation forecasting algorithm using machine learning to provide practical and actionable insights for power plant operators. The analysis used meteorological data and actual power generation data measured in a real-world wind farm located in South Korea. This study proposes a stacking ensemble model that combines density-based spatial clustering of applications with noise (DBSCAN)-based data cleaning and physical feature transformation to resolve the uncertainty of wind power data and maximize prediction precision. First, the DBSCAN clustering algorithm was applied to eliminate abnormal noise on the power curve. To reflect the physical laws of wind power generation, a Log V³ transformation was performed to incorporate the cubic relationship between wind speed and power output. Subsequently, a stacking ensemble architecture was designed using Random Forest, XGBoost and Linear Regression as base models and another Random Forest as the meta-model. Experimental results demonstrated that the proposed model achieved a significant reduction in error by decreasing the MAPE. This study demonstrates that the engineering design of data preprocessing plays a critical role in improving model reliability and is expected to serve as a practical forecasting framework for virtual power plants (VPP) and energy management systems (EMS). To verify the robustness of the proposed model, further investigation is required to analyze errors associated with the use of numerical weather prediction (NWP) data as input for day-ahead forecasting.
스태킹 앙상블 기반 머신러닝 모델을 활용한 파킨슨병 보행데이터 해석 연구
[Kisti 연계] 대한의용생체공학회 Journal of biomedical engineering research : the official journal of the Korean Society of Medical & Biological Engineering Vol.46 No.5 2025 pp.404-410
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This study proposes an interpretable machine learning model to assist in the early-stage classification of Parkinson's disease, aiming to support early diagnosis. A stacking ensemble approach integrating Support Vector Machine (SVM), Random Forest, and XGBoost classifiers was employed on gait data obtained from the PhysioNet Gait in Parkinson's Disease dataset, consisting of 93 patients (mean age 66.3 ± 10.9 years, Hoehn & Yahr stage 2-3) and 73 healthy controls. Feature standardization and recursive feature elimination with cross-validation (RFECV) were applied to select informative gait features. The proposed model achieved 96.6% classification accuracy and an AUC of 0.990, demonstrating strong generalization and interpretability. These findings suggest that the proposed stacking ensemble framework could serve as a foundation for future predictive models for early Parkinson's detection.
유튜브 악플 탐지를 위한 기계학습: 스태킹 앙상블 모델의 적용을 중심으로
[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.24 No.4 2022.08 pp.1583-1598
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 유튜브와 같은 개인미디어 플랫폼에서 쉽게 확산되는 악성댓글을 자동적으로 판별하는 기계학습 모델의 분류 성능을 비교하고 이를 개선하기 위한 방법으로 스태킹 앙상블 모델을 적용해 그 결과를 평가했다. 댓글 데이터는 특정 유명인이 연루된 사회적 혹은 연예계 이슈를 선정적으로 전달해 혐오 정서를 자극하는 “사이버렉카” 콘텐츠에 주목해 관련한 대표 채널의 인기 영상에 달린 댓글 59,999건을 웹 크롤링 방식으로 수집했다. 그리고 인간 코더의 라벨링 작업으로 판별된 2,851건의 악플과 데이터 균형화를 위해 무작위로 추출된 2,851건의 악플이 아닌 댓글로 구성된 5,702건의 댓글 데이터 세트를 마련했다. 분류 알고리즘으로는 logistic regression, Naïve Bayes, random forest, support vector machine 모델을 선택했다. 분석 결과, 단일 알고리즘에 기반한 모델의 성능은 평가 지표에 따라 장단점이 분명하게 드러나고 있었다. 가령, 정확도와 점수의 기준으로 가장 좋은 성능을 보여준 random forest 모델은 재현율에 있어 다른 분류 알고리즘에 비해 떨어지는 성능을 나타냈다. 반면, 스태킹 앙상블 모델은 단일 알고리즘 분류기의 장점을 취하고 약점은 보완하는 메타 분류기를 도출해 악플 분류 성능을 향상시켰다. 결국, 스태킹 앙상블 모델은 배깅이나 부스팅 앙상블 학습과는 달리 서로 다른 장점을 보이는 복수의 알고리즘을 결합하여 모델의 성능 지표에서 나타난 불균형을 해소하고 이를 통해 보다 균형있게 개선된 악플 분류 모델의 생성과 적용에 활용될 수 있음을 제시한다.
This study examined the classification performance of machine learning models that automatically detect malicious comments, which easily spread on personalized media platforms such as YouTube, and applied a stacking ensemble model as a way to improve them. Though web-crawling, 59,999 comments were collected from popular videos of representative “cyber wrecker” channels that deliver social or entertainment issues involving certain celebrities and stimualte hatred of them. In doing so, we prepared a data set of 5,702 comments, of which 2,851 were labeled as malicious by human coders and 2,851 were randomly extracted from the remaining non-malicious comments to balance the data. As classification algorithms, we selected the logistic regression, Naive Bayes, random forest, and support vector machine models. As a result, the performance of the models based on a single algorithm was clearly dependent on evaluation metrics. In particular, the random forest model, which showed the best performance on the basis of accuracy and score, fell behind other models in terms of recall. On the other hand, the stacking ensemble model improved the performance of classifying malicious and non-malicious comments by generating a meta-learning classifier that gained from single algorithm classifiers and complementing their weaknesses. Finally, we demonstrated how stacking ensemble models, as opposed to ensemble learning methods of bagging or boosting, can be used to create a better classifier of malicious comments by combining multiple algorithms with different advantages and addressing the imbalances shown in the classification evaluation metrics.
민간 이종 데이터 결합을 통한 지역사랑상품권의 매출 예측 분석: 스태킹 앙상블 기반의 추정 모델을 중심으로
[NRF 연계] 경남대학교 산업경영연구소 지역산업연구 Vol.49 No.2 2026.05 pp.31-51
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 지역사랑상품권을 발행하는 지방정부의 수가 해마다 늘고 있으며 판매 규모 또한 수십조 원에 달하는 상황에서 민간 이종 데이터의 활용을 통해 지역사랑상품권의 효과를 확인하고자 한다. 즉 정책 결과 분석에 대한 민간 이종 데이터의 효율적 활용법의 가능성을 제시하고자 스태킹 앙상블 모델을 활용하여 지역사랑상품권 가맹점의 카드사 매출 기반으로 지역사랑상품권의 매출을 추정 및 검증하고, 이 모델을 비가맹점이 가상 가맹시 기대할 수 있는 매출을 예측하여 정책적·실무적 시사점을 제시하고자 한다. 분석 결과, 첫째 스태킹 앙상블 모델을 통해 추정된 매출 분포가 가맹점의 실제 지역사랑상품권 매출과 비교하였을 때 평균 수준의 매출액뿐만 아니라 매출 분포의 전반적인 구조를 일정 수준 재현할 수 있음을 확인하였다. 이를 통해 스태킹 앙상블 모델이 비가맹점 매출 추정을 위한 기초 모델로 적용 가능한 것을 확인하였다. 둘째 업종별 지역사랑상품권의 가상 가맹 시 기대되는 매출 추정 결과에서 업종 간 차이는 있지만, 지역사랑상품권 가맹과 이에 따른 결제 수단의 추가가 소비 유입 및 매출 증진에 일정 수준 긍정적으로 작용할 가능성을 확인하였다. 본 연구는 정책 결과를 인과적으로 식별하려는 기존의 접근방식과 달리, 지역사랑상품권과 같이 접근이 제한된 정보하에서 이를 효과적으로 활용하여 매출 수준을 가늠할 수 있는 예측 기반 결과를 재시한다는 점에서 기존 연구와 차별적인 특징을 갖는다. 즉 정책 분석에 있어 민간 이종 데이터의 효율적 접근법에 대한 새로운 가능성을 제시하였다는 점에서 의의가 있으며, 실무적으로 업종별 비가맹점 점주에게 지역사랑상품권 도입에 따른 잠재적 매출 수준에 대한 정량적 수치를 제공한다는 점에서 의의가 있다.
This study evaluates the effectiveness of local currency policies by utilizing heterogeneous private-sector data, in a context where both the number of issuing local governments and transaction volumes continue to grow. Using a stacking ensemble model, the study estimates and validates local currency sales based on affiliated merchants’ credit card data and predicts sales for non-affiliated merchants, providing policy and practical implications through a prediction-based approach. The results show that the model reasonably reproduces not only average sales but also the overall distribution of actual sales among affiliated merchants, confirming its applicability for estimating non-affiliated merchant sales. Additionally, simulated results suggest that adopting local currency systems can positively influence consumer inflow and sales across industries, although the magnitude varies by sector. Unlike conventional causal inference approaches, this study proposes a prediction-based framework under limited data conditions. It contributes by demonstrating the potential of heterogeneous private data in policy analysis and by providing quantitative estimates of expected sales for non-affiliated merchants considering adoption of local currency systems.
밀리미터파 레이더 기반 비접촉식 낙상 감지를 위한 스태킹 앙상블 학습 기법 연구
[NRF 연계] 한국융합신호처리학회 융합신호처리학회 논문지 Vol.26 No.6 2025.12 pp.315-323
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
고령화 사회로의 진입과 함께 노인 낙상 사고에 대한 사회적 관심이 증가하고 있으며, 이를 조기에 감지하기 위한 다양한 기술적 접근이 연구되고 있다. 기존 영상 기반 낙상 감지 시스템은 프라이버시 침해 우려와 환경 변화에 대한 취약성이 있으며, 웨어러블 센서 기반 방식은 사용자의 착용 부담이라는 한계를 가진다. 본 논문에서는 이러한 문제를 해결하기 위해 밀리미터파(mmWave) 레이더를 활용한 비접촉식 낙상 감지 기법을 제안한다. 실내 환경에서 수집된 레이더 데이터를 기반으로 전통적인 머신러닝 모델과 딥러닝 모델을 결합한 스태킹(Stacking) 기반 앙상블 학습 구조를 설계하였다. 제안한 모델은 여러 기본 분류기의 예측 결과를 메타 분류기를 통해 통합함으로써 단일 모델 대비 향상된 낙상 감지 성능과 강인성을 확보한다. 실험 결과, 제안한 앙상블 모델은 정확도, 정밀도, 재현율 및 F1-score 측면에서 기존 단일 분류 모델보다 우수한 성능을 보였다. 본 연구는 레이더 기반 센싱과 앙상블 학습의 결합이 프라이버시 보호가 요구되는 실내 환경에서 효과적인 낙상 감지 방법이 될 수 있음을 보여준다.
With the transition to an aging society, increasing attention has been directed toward early detection of falls among elderly individuals. Vision-based fall detection systems suffer from privacy concerns and environmental sensitivity, while wearable sensor-based approaches impose a user burden. To address these limitations, this paper proposes a contactless fall detection method using millimeter-wave (mmWave) radar. Radar data collected in an indoor environment are used to design a stacking-based ensemble learning framework that combines traditional machine learning and deep learning models. The proposed method integrates predictions from multiple base classifiers through a meta-classifier to improve detection performance and robustness. Experimental results using real-world radar data demonstrate that the proposed ensemble model outperforms conventional single classifiers across accuracy, precision, recall, and F1-score. These findings indicate that radar-based sensing, combined with ensemble learning, provides a practical, privacy-preserving solution for indoor fall detection.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.