년 - 년
Random Forest 연구방법을 적용한 이륜차의 지역별 사고 심각도 비교 분석에 관한 연구 KCI 등재
한국재난정보학회 한국재난정보학회논문집 제21권 2호 통권68호 2025.06 pp.357-365
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
연구목적: 전국 전체의 이륜차 사고 건수는 점점 감소하고 있지만 사망자 수는 증가한 것으로 나타났으 며 이륜차 사고가 집중되는 시간대는 16~22시인 저녁 시간대로 나타났다. 따라서 주거지역과 다른 지 역 간의 이륜차 사고 심각도 요인의 차이를 비교할 필요가 있다. 연구방법: Random Forest를 활용해 2007년부터 2021년까지 인천광역시에서 발생한 이륜차 사고 자료와 QGIS를 사용하여 도시지역 및 비 도시지역으로 구분하고 지역에 따라 사고 심각도 분석 및 사고 심각도에 끼치는 영향요인을 분석하였 다. 연구결과: 각 지역의 상위 2개의 변수 중요도를 분석한 결과, 도시 지역 이륜차 사고 심각도에 높은 영향을 끼치는 변수는 교차로 안에서 사고가 발생한 경우, 신호위반의 경우이고, 공업지역은 차대차 정 면충돌의 경우와 피해자 차종이 화물차인 경우, 녹지지역은 신호위반의 경우, 교차로 안에서 사고가 발 생한 경우, 상업지역은 신호위반,차대차- 측면직각충돌의 경우, 비도시지역은 가해운전자가 60세 이상 인 경우, 피해자가 40~50대인 경우 사고심각도에 더 영향을 주는 것으로 나타났다. 결론: 지역의 특성에 따라 이륜차 사고의 심각도에 영향을 끼치는 요인을 집중적으로 분석하였고 본 연구로 도출된 결과를 활용하여 이륜차 사고의 심각도를 개선하기 위한 효율적이고 대응책 수립이 필요하다.
Purpose: The number of motorcycle accidents nationwide is decreasing, but the number of deaths has increased, and the time when motorcycle accidents are concentrated is between 16:00 and 22:00 in the evening. Therefore, it is necessary to compare the differences in the severity factors of motorcycle accidents between residential areas and other areas. Method: Using Random Forrest, QGIS and data on motorcycle accidents in Incheon Metropolitan City from 2007 to 2021 were used to divide them into urban and non-urban areas to derive accident severity analysis and impact factors on accident severity by region. Result: As a result of analyzing the importance of the top two variables in each region, the variables that have a high impact on the seriousness of motorcycle accidents in urban areas are traffic lights in intersections Conclusion: It is necessary to establish efficient and countermeasures to improve the severity of motorcycle accidents by analyzing in-depth factors that affect the severity of motorcycle accidents according to the characteristics of the region and utilizing the results derived from this study.
인공지능 분류모델을 활용한 농축수산물 전문 쇼핑 모바일 앱의 구매고객 예측 : 고객 행동 데이터를 기반으로 KCI 등재
한국기업경영학회 기업경영연구 제31권 제5호 통권 제117호 2024.10 pp.77-108
※ 기관로그인 시 무료 이용이 가능합니다.
7,300원
본 연구는 온라인 쇼핑 환경에서 고객의 구매 의도를 예측하는 인공지능 기반의 분류 모델을 탐색하고 평가 하는 것을 목표로 한다. 전통적인 상점 고객과 달리, 온라인 쇼핑 고객은 다양한 매체를 통해 정보를 수집하고 상품 리뷰나 구매 패턴을 참고하여 구매 결정을 내린다. 본 연구에서는 고객의 행동 데이터를 기반으로 성별, 연령, 직업과 같은 개인 기본 특성을 배제하고, 검색, 장바구니 담기, 페이지뷰 등의 고객 행동 데이터를 수집하 여 구매 가능성을 예측하였다. 이는 개인 정보 보호에 대한 우려로 인해 기본 특성 데이터의 수집이 어려운 현 실을 감안한 접근법이다. 따라서 행동 데이터만을 활용한 분석 방법의 중요성이 더욱 부각된다. 연구에서는 가 우시안 나이브 베이즈, 랜덤 포레스트, 의사 결정 트리, 로지스틱 회귀 등의 기계학습 모델을 사용하여 행동 데 이터로부터 구매 고객을 예측하였다. 각 모델의 성능을 평가하여 가장 효과적인 모델을 식별하고, 이를 통해 기 업이 구매 예측 모델을 선택할 때 참고할 수 있는 근거를 제공하고자 한다. 또한, 식별된 잠재 구매 고객에게 마케팅 자원을 집중하는 경우와 모든 고객에게 동일하게 마케팅 자원을 배분하는 경우의 구매액을 비교하여 타 겟 마케팅의 비용 효율성을 분석하였다. 본 연구는 온라인 농축수산물 리테일러들이 고객의 구매 행동을 더 정 확히 이해하고, 마케팅 자원을 효율적으로 배분하여 매출을 극대화하는 방안을 제시한다. 이를 통해 고객 행동 데이터 기반의 예측을 통한 효과적인 경영 전략 수립 방법론을 제공하고, 예측 모델을 제안하는 데 본 연구의 의의를 둔다. 고객 행동 데이터의 분석을 통해 기업은 더욱 정교한 타겟 마케팅 전략을 수립할 수 있으며, 이는 결국 매출 증대와 비용 절감으로 이어질 것이다.
This study aims to explore and evaluate AI-based classification models to predict customer purchase intentions in an online shopping environment. Unlike traditional store customers, online shoppers gather information through various media and make purchase decisions based on product reviews and purchase patterns. This study excludes personal basic characteristics such as gender, age, and occupation and instead collects and utilizes customer behavior data such as searches, cart additions, and page views to predict purchase likelihood. Given the difficulties in collecting basic characteristic data due to privacy concerns, the analysis method using only behavior data becomes even more important. This study used machine learning models such as Gaussian Naive Bayes, Random Forest, Decision Tree, and Logistic Regression to predict purchasing customers from behavior data. The performance of each model was evaluated to identify the most effective one, providing a reference for companies in selecting purchase prediction models. Additionally, we compared the purchase amounts between focusing marketing resources on identified potential buyers and distributing marketing resources equally to all customers to analyze the cost efficiency of targeted marketing. This study presents ways for online agricultural and fisheries retailers to better understand customer purchase behavior and efficiently allocate marketing resources to maximize sales. By providing a research framework and proposing prediction models, this study aims to offer methodologies for effective management strategies based on predictions from customer behavior data. Through analyzing customer behavior data, companies can develop more sophisticated targeted marketing strategies, ultimately leading to increased sales and cost savings.
랜덤 포레스트(Random Forest)의 시계열 적용에 관한 연구: 한국 물가상승률 예측 사례 분석
[NRF 연계] 한국경제학회 경제학연구 Vol.71 No.3 2023.09 pp.37-73
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본고는 한국과 미국의 물가상승률 예측에 랜덤 포레스트 모형을 적용할 때, Stationary Bootstrap이나 Moving Block Bootstrap 등 Block Bootstrap을 사용하는 것이 통상적인 독립 부트스트랩(Independent Bootstrap)을 사용하는 것에 비해 통계적으로 유의한 수준으로 예측력을 개선하지는 못한다는 것을 보인다. 그리고 FRED-MD를 참고한 총 93개의 관련 국내외 거시경제/금융 변수들을 사용하고, XGBoost, LSTM 등 다양한 머신러닝 방법을 활용하여 한국의 물가상승률을예측하고 분석한다. 2004년 9월에서 2022년 3월까지의 표본을 이용하였고, 1개월에서 12개월의 예측 대상기간(Forecast Horizon)을 고려하였다. 총 13개의 모형 중 대부분의 예측 대상기간에 있어 예측력이 우수한 모형이 존재하는 것으로나타났는데, 이는 보루타 알고리즘(Boruta Algorism)을 통해 중요한 변수로 분류된 변수들만을 랜덤 포레스트에 적용하는 모형이다. Giacomini and White(2006) 와 Hansen et al.(2009)의 검정을 통해 대부분의 예측 대상기간에서 통계적으로유의하게 예측력이 우수함을 확인하였는데, 특히 경제활동인구 및 취업자 수의증가율 등 고용시장 관련 변수, 기업경기실사지수, 주택가격 변화율 등이 물가상승률 예측에 중요한 변수로 선택되는 것으로 나타났다.
This paper first investigates whether adopting the stationary bootstrap or the moving block bootstrap, instead of the usual independent bootstrap, in the random forest method improves forecasting of stationary time series. It is shown that the block bootstrap procedures adopted in the random forest method do not make any statistically significant improvement in Korean or US inflation forecasting. Secondly, we consider inflation forecasting in Korea using 93 macroeconomic/financial variables and various machine learning methods. The samples are from September 2004 to March 2022. Comparing total 13 models, one model outperforms the rest models for most forecast horizons, which is a simplified method of the model proposed by Kim and Han (2022). The method consists of the following two steps: 1) Select important variables based on the Boruta algorithm, 2) Using only those selected variables, implement the random forest and produce a forecast. The tests by Giacomini and White (2006) and Hansen et al. (2009) show that the model provides significantly better forecasts for most forest horizons. In particular, the Boruta algorithm selected total economically active population, total employed persons, BSI, house price as important variables for Korean inflation forecasting.
Ransomware Detection using Random Forest Technique
[NRF 연계] 한국통신학회 ICT Express Vol.6 No.4 2020.12 pp.325-331
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Nowadays, the ransomware became a serious threat challenge the computing world that requires an immediate consideration to avoid financial and moral blackmail. So, there is a real need for a new method that can detect and stop this type of attack. Most of the previous detection methods followed a dynamic analysis technique which involves a complicated process. The present study proposes a novel method based on static analysis to detect ransomware. The significant characteristic of proposed method is dispensing of disassemble process by direct extraction of features from raw byte with the use of frequent pattern mining which remarkably increases the detection speed. The Gain Ratio technique was used for feature selection which exhibited that 1000 features was the optimal number for detection process. The current study involved using random forest classifier with a comprehensive analysis to the effect of both tree and seed numbers on the ransomware detection. The results showed that tree numbers of 100 with seed number of 1 achieved best results in terms of time-consuming and accuracy. The experimental evaluation revealed that the proposed method could achieve a high accuracy of 97.74% for detection ransomware.
[NRF 연계] 한국통신학회 ICT Express Vol.9 No.1 2023.02 pp.122-129
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Detection of nodes disseminating false data is a prerequisite for effective deployment of Internet of Vehicles (IoV) services. This work proposed a novel hyper-tuned ensemble Random Forest (Ens. RF) algorithm to detect false basic safety messages in IoV. Performance evaluation was done using the Vehicular Reference Misbehavior (VeReMi) dataset comprising data-centric misbehavior evaluation for vehicular networks. For validation, a comparative analysis of the performance of the proposed “Ens. RF” model, five machine learning algorithms implemented in this work, and state-of-the-art ML models from related literature was presented. The performance metrics considered are time efficiency and validation accuracy for overall misbehavior classification. Also, the results confirmed the irrelevance of data balancing in real-life scenarios. Finally, we assess the performance of our proposed system for detecting each falsification scenario using precision and recall. The result shows that the proposed algorithm outperformed others with a validation accuracy of 99.60% and a negligible 604 misclassifications out of 153,730 points.
[NRF 연계] 한국아동간호학회 Child Health Nursing Research Vol.31 No.2 2025.04 pp.85-95
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Purpose: This study aimed to identify predictive factors affecting adolescents’ subjective happiness using data from the 2023 Korea Youth Risk Behavior Survey. A random forest model was applied to determine the strongest predictive factors, and its predictive per- formance was compared with traditional regression models. Methods: Responses from a total of 44,320 students from grades 7 to 12 were analyzed. Data pre-processing involved handling missing values and selecting variables to con- struct an optimal dataset. The random forest model was employed for prediction, and SHAP (Shapley Additive Explanations) analysis was used to assess variable importance. Results: The random forest model demonstrated a stable predictive performance, with an R of .37. Mental and physical health factors were found to significantly affect subjec-2tive happiness. Adolescents’ subjective happiness was most strongly influenced by per- ceived stress, perceived health, experiences of loneliness, generalized anxiety disorder, suicidal ideation, economic status, fatigue recovery from sleep, and academic perfor- mance. Conclusion: This study highlights the utility of machine learning in identifying factors influencing adolescents’ subjective happiness, addressing limitations of traditional re- gression approaches. These findings underscore the need for multidimensional interven- tions to improve mental and physical health, reduce stress and loneliness, and provide integrated support from schools and communities to enhance adolescents’ subjective happiness.
[NRF 연계] 한국간호과학회 Asian Nursing Research Vol.19 No.4 2025.10 pp.347-353
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Purpose: The goal of this study was to explore the current status and influencing factors of nurses'intention to report adverse nursing events, using a random forest model to rank and select these factors. Methods: In August 2024, a survey was conducted among nurses at a tertiary hospital in Shanghai usinga convenience sampling method. The instruments included the General Information Questionnaire, theIntention to Report Questionnaire, the Reporting of Clinical Adverse Effects Scale (RoCAES), and theReporting Barrier Questionnaire. A random forest model and stepwise multiple linear regression wereused to analyze the factors influencing nurses' intention to report nursing adverse events. Results: The median total score for the intention to report adverse events among 823 nurses was 14(range: 0 to 15). Lasso regression identified eight key factors influencing the intention to report adverseevents: professional title, whether the nurse had reported adverse events, age, punitive culture,reporting process, reporting significance, reporting purpose, and reporting environment. Stepwiseregression analysis revealed that younger nurses, those with no prior reporting experience, and thosewho perceived reporting as important had a higher intention to report adverse events. Additionally,nurses were more likely to report adverse events when the hospital's reporting process was simpler andthe culture was less punitive. Conclusions: Nurses' intention to report adverse events in nursing is moderate. Nursing managers shouldconsider these influencing factors holistically and implement targeted measures to enhance nurses'intention to report, thereby improving nursing quality and patient safety.
Random Forest를 활용한 고속도로 교통사고 심각도 비교분석에 관한 연구 KCI 등재
한국재난정보학회 한국재난정보학회논문집 제20권 1호 통권63호 2024.03 pp.156-168
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
연구목적: 고속도로 교통사고의 추세는 증감을 반복하며 도로 종류 중 고속도로에서의 치사율은 최고치 를 나타내고 있다. 따라서 국내 실정을 반영한 개선대책 수립이 필요하다. 연구방법: Random Forest를 활 용해 2019년부터 2021년까지 전국 고속도로 노선 중 사고 다발 10개 노선에서 발생한 교통사고 자료로 사고 심각도 분석 및 사고 심각도에 미치는 영향요인을 도출하였다. 연구결과: SHAP 패키지를 활용해 상 위 10개의 변수 중요도를 분석한 결과, 고속도로 교통사고 중 사고 심각도에 높은 영향을 미치는 변수는 가해자 연령이 20세 이상 39세 미만, 시간대가 주간(06:00-18:00), 주말(토~일), 계절이 여름과 겨울, 법규 위반이 안전운전불이행, 도로 형태가 터널, 기하구조상 차로 수가 많고 제한속도가 높은 경우로 총 10개 의 독립변수에서 고속도로 교통사고 심각도와 양(+)의 상관관계를 가지는 것으로 분석되었다. 결론: 고속 도로에서의 사고 발생은 매우 다양한 요인의 복합적인 작용으로 인해 발생하므로 사고 예측에 많은 어려 움이 있지만 본 연구로 도출된 결과를 활용해 고속도로 교통사고 심각도에 영향을 주는 요인을 심층적으 로 분석해 효율적이고 합리적인 대응책 수립을 위한 노력이 필요하다.
Purpose: The trend of highway traffic accidents shows a repeating pattern of increase and decrease, with the fatality rate being highest on highways among all road types. Therefore, there is a need to establish improvement measures that reflect the situation within the country. Method: We conducted accident severity analysis using Random Forest on data from accidents occurring on 10 specific routes with high accident rates among national highways from 2019 to 2021. Factors influencing accident severity were identified. Result: The analysis, conducted using the SHAP package to determine the top 10 variable importance, revealed that among highway traffic accidents, the variables with a significant impact on accident severity are the age of the perpetrator being between 20 and less than 39 years, the time period being daytime (06:00-18:00), occurrence on weekends (Sat-Sun), seasons being summer and winter, violation of traffic regulations (failure to comply with safe driving), road type being a tunnel, geometric structure having a high number of lanes and a high speed limit. We identified a total of 10 independent variables that showed a positive correlation with highway traffic accident severity. Conclusion: As accidents on highways occur due to the complex interaction of various factors, predicting accidents poses significant challenges. However, utilizing the results obtained from this study, there is a need for in-depth analysis of the factors influencing the severity of highway traffic accidents. Efforts should be made to establish efficient and rational response measures based on the findings of this research.
팬데믹 전후 항공기 기종에 따른 기재 활용 전략의 변화 : Random Forest와 SHAP 분석을 중심으로
한국정보기술응용학회 JITAM Vol.32 No.2 2025.04 pp.83-95
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
This study investigates the changes in factors influencing aircraft utilization based on aircraft types during the pandemic period. The analysis categorizes the periods into three phases: pre-pandemic, during the pandemic, and post-pandemic, while the aircraft types are classified into two groups (C-class and E-class or higher). Utilizing SHAP and Random Forest methodologies, the study analyzes the importance of various factors for each group and examines the significance of differences in importance across the specified periods and aircraft types through group difference testing. The findings reveal that for C-class aircraft, the importance of the USD/KRW exchange rate and the lagging composite index was notably high before the pandemic. However, during the pandemic, there was a substantial surge in the significance of the JPY/KRW exchange rate, and post-pandemic, the importance of economic indicators increased. E-class or higher aircraft exhibited a similar trend, with a significant rise in the importance of the CNY/KRW exchange rate observed after the pandemic. The Kruskal-Wallis test confirmed that most average changes among the groups were not statistically significant, except for the C-class aircraft in the Random Forest analysis. This indicates that there were no substantial changes in the underlying factors between the pre-pandemic and post-pandemic periods, suggesting that the impact of the pandemic on airline fleet utilization extends beyond short-term changes. Furthermore, the findings highlight the need for airlines to enhance their sensitivity to external environments and develop utilization strategies in volatile markets. Additionally, there is a possibility that long-term dependency on network strategies may become more pronounced than reliance on external factors like the pandemic. By providing an in-depth analysis of the pandemic's impact on aircraft utilization, this study contributes to the formulation of effective utilization strategies for airlines navigating an uncertain economic environment.
GRU 분석 모델을 활용한 비트코인 시장 예측 : LSTM과 Random Forest 모델 비교 분석
한국정보기술응용학회 JITAM Vol.32 No.1 2025.02 pp.1-13
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
This study conducts a comparative analysis of various models for predicting Bitcoin prices. The models utilized in this research include the Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), and Random Forest. Their performance was evaluated using key metrics such as Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R-squared (R²). The findings reveal that the GRU model outperformed the other two models, achieving an R² value of 0.9301, which indicates its superior ability to explain the data. The LSTM model demonstrated relatively high explanatory power with an R² of 0.8252, yet it exhibited lower predictive accuracy compared to the GRU. On the other hand, the Random Forest model recorded the lowest MAE (1557.44) and RMSE (1766.95), reflecting strong short-term prediction capabilities. However, its R² of 0.6110 highlighted limitations in capturing the intricate temporal patterns of the dataset. The significance of this study lies in several aspects. First, it empirically demonstrates that the GRU model, with its more streamlined architecture, can achieve superior predictive performance compared to the LSTM and Rando Forest models. Additionally, while Random Forest showed promising results for short-term predictions, it underscored limitations in modeling the complexities of temporal data patterns. Future research should aim to incorporate broader datasets and integrate external economic variables to enhance predictive accuracy. Furthermore, the potential of ensemble approaches, combining diverse models, holds promise for further improving predictive performance in this domain.
강우·소방·사회취약성 데이터를 활용한 복합재난 위험도 예측 및 소방 자원 배치 전략 : 서초구 사례의 MRA–Random Forest 비교 분석 KCI 등재
한국재난정보학회 한국재난정보학회논문집 제21권 4호 통권70호 2025.12 pp.1169-1178
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
연구목적: 본 연구는 서울특별시 서초구를 대상으로 강우, 소방 활동, 사회 취약성 데이터를 통합하여 복합재난 위험 도를 예측하고, 이를 기반으로 소방 자원 배치 전략을 제시하는 것을 목적으로 한다. 연구자료는 2020–2024년 서초 구 시간별 강우량 자료, 화재 출동 개별 내역, 119 구조·구급활동 실적, 통계 연보에서 추출한 인구·주거·사회 취약성 지표를 사용하였다. 특히 신고일시·출동일시·도착일시를 활용하여 출동 지연, 도착시간, 총 대응 시간을 산출하고, 65세 이상 인구 비율, 장애인 등록 인구 비율, 1인 가구 비율, 반지하·지하 주택 비율을 사회 취약성 변수로 구성하였 다. 연구방법: 분석 방법으로는 먼저 화재 출동 내역을 이용해 연도·월·시간대별 출동 건수와 평균 도착시간, 평균 총 대응 시간을 산출하여 기초 통계를 제시하고, 강우량 및 사회 취약성 지표와 결합하여 복합재난 발생에 영향을 미치 는 설명변수를 구성하였다. 이후 다중회귀분석(Multiple Regression Analysis, MRA)과 Random Forest(RF)를 이용하 여 시간 단위 화재 출동 건수 및 발생 여부를 예측하고, 두 모형의 성능을 비교하였다. 2024년을 검증자료로 사용한 결과, 시간당 화재 출동 건수 예측에서 MRA 모형의 평균절대오차(MAE)는 0.24건, RF 모형의 MAE는 0.18건으로, RF가 약 25.0% 낮은 예측 오차를 보였다. 화재 발생 여부(해당 시간에 1건 이상 발생)의 분류 성능을 AUC로 비교한 결과, 로지스틱 회귀는 0.83, RF 분류기는 0.90을 기록하여 RF가 우수한 성능을 나타냈다. 연구결과: RF 변수 중요도 분석 결과, 이전 24시간 누적 강우량, 현 시각 강우량, 신고 시간의 시간대(야간 여부), 65세 이상 인구 비율, 장애인 등록 인구 비율, 반지하·지하 주택 비율, 1인 가구 비율순으로 중요도가 높게 나타났다. 이는 강우 특성뿐 아니라 고령자 및 장애인 밀집 지역, 주거 취약 지역, 1인 가구 밀집 지역이 복합재난 위험에 유의 미한 영향을 미침을 시사한다. 마지막으로, 예측된 위험도와 사회 취약성 지수를 결합하여 행정동별 우선 대응 순위를 제시하고, 이를 기반으로 소방 자원 배치 시나리오를 제안하였다. 결론: 본 연구는 기상·소방·사회통계 데이터를 결합한 복합재난 위험도 예측과 사회 취약성을 고려한 자원 배치 전략을 함께 제시함 으로써, 도시형 재난관리에 대한 데이터 기반 의사결정 지원에 기여할 수 있을 것으로 기대된다.
Purpose: This study aims to predict the risk of compound disasters in Seocho District, Seoul, by integrating rainfall, firefighting activity, and social vulnerability data, and to propose strategies for the optimal allocation of firefighting resources. Method: The research employs hourly rainfall data from 2020 to 2024, detailed fire dispatch records, 119 rescue and emergency activity logs, and demographic and social vulnerability indicators extracted from statistical yearbooks, and calculates dispatch delays, arrival times, and total response times using reported, dispatch, and arrival timestamps. Social vulnerability variables include the proportion of residents aged 65 or older, the proportion of registered persons with disabilities, the share of single-person households, and the percentage of semi-basement or basement dwellings. Result: Multiple Regression Analysis and Random Forest models are applied to predict hourly fire dispatch counts and fire occurrence, and their predictive performances are compared using 2024 data as a validation set; results show that the Random Forest model achieves a lower mean absolute error than the regression model in count prediction and a higher AUC in occurrence classification.Variable importance analysis further indicates that cumulative rainfall over the previous 24 hours, current rainfall, time of report (day/night), and the proportions of elderly residents, registered disabled residents, semi-basement or basement dwellings, and single-person households are major contributors to predicted compound disaster risk. Conclusion: Overall, the study concludes that integrating meteorological, firefighting, and socio-demographic data enables more accurate compound disaster risk prediction and supports socially responsive resource allocation strategies, thereby contributing to data-driven decision-making in urban disaster management.
Risk factors for Death in Patients with Acute Myocardial Infarction using Random Forest
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 9th International Conference on Next Generation Computing 2023 2023.12 pp.289-291
Acute myocardial infarction (AMI) is one of the leading causes of cardiovascular disease-related mortality worldwide, causing ischemic damage to the heart muscle. AMI poses a risk of sudden cardiac arrest and is associated with a high rate of recurrence, which can significantly impact daily life. Therefore, this study aims to develop a predictive model for mortality in AMI patients using machine learning. The data used in the study are from KNHDIS (2013-2022) and include demographic characteristics and disease information of discharged patients. The model was constructed using RFECV and Random Forest. Key variables influencing mortality include age, number of surgeries, and length of hospital stay. The model demonstrated high performance with an accuracy of 94% and an AUC value of 0.93.
한국ITS학회 한국ITS학회 학술대회 ITS와 함께하는 미래 스마트 시티 2022.06 pp.887-890
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
In this study, AAHT was estimated using machine learning algorithms of ensemble techniques to improve the limitations of existing studies. Among the machine learning algorithms, a random forest algorithm was used, and the traffic volume estimation accuracy was analyzed as MAPE 20.7%. It was analyzed that the higher the traffic volume level, the higher the traffic volume estimation accuracy, and the traffic volume level of 3,000 vehicles/hour or more was analyzed to be 8% or less of MAPE. It was analyzed that the accuracy secured at the current level was very high compared to the existing model.
건국대학교 KU중국연구원 Open Regional Studies Vol.5 No.5 2026.07 pp.1-24
※ 기관로그인 시 무료 이용이 가능합니다.
6,100원
This study explores the predictive factors of international students’ intention to participate in local festivals in Korea using CART and Random Forest analyses. Local festivals can provide international students with opportunities to experience local culture, interact with local communities, and adapt to Korean society. However, their participation may be constrained by information accessibility, language barriers, transportation, time, cost, and lack of companions. Using survey data collected from 509 international students in Korea, this study classified local festival participation intention into high and non-high participation intention groups based on the mean score of festival participation intention. The analysis included seven perceptual constructs—cultural interest, sense of local community belonging, expectation of social exchange, information accessibility, perceived enjoyment, perceived usefulness, and participation-facilitating conditions—together with demographic variables. The CART results showed that participation-facilitating conditions and perceived usefulness were the key variables distinguishing the high participation intention group. Random Forest analysis further confirmed the relative importance of perceived usefulness, participation-facilitating conditions, perceived enjoyment, and expectation of social exchange. The ROC-AUC values of both CART and Random Forest exceeded 0.70, indicating acceptable discriminative power. These findings suggest that international students’ intention to participate in local festivals is explained more strongly by practical usefulness and feasible participation conditions than by demographic characteristics alone. The study contributes to local festival research by extending the focus to international students and provides practical implications for local governments, festival organizers, and university international offices seeking to promote international students’ community participation.
청소년 자살생각의 예측요인 고찰: 랜덤포레스트 방법의 적용
[NRF 연계] 한국심리학회 산하 한국학교심리학회 한국심리학회지: 학교 Vol.21 No.1 2024.04 pp.1-26
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 랜덤포레스트 머신러닝 방법을 적용하여 청소년의 자살생각을 예측하는 모형을 구성하고, 모형의 예측 성과를 평가하며, 예측에 기여하는 요인의 특성을 파악하여 이를 바탕으로 청소년 자살예방 방안을 제시하는 것을 목적으로 하였다. 이를 위해 질병관리청의 2021년도 청소년 건강행태조사 자료를 분석에 사용하였으며, 구체적으로 국내 중⋅고등학교에 재학 중인 총 54,565명의 청소년의 응답자료를 분석하였다. 청소년 자살생각의 예측모형은 정신건강(예: 슬픔 및 절망감, 외로움, 스트레스 등), 신체적 건강(예: 키, 몸무게, 식이습관, 수면시간 등), 시간활용(예: 신체활동, 스마트폰 사용 시간), 건강 관련 학교 교육, 학업성적, 물리적 환경, 코로나로 인한 변화(예: 경제적 어려움, 우울감, 신체활동 등) 등의 영역에 해당하는 37개의 예측변수를 포함하였으며 랜덤포레스트 머신러닝 방법을 적용하여 추정하였다. 예측에 중요하게 역할하는 요인은 중학생 및 고등학생 모두에서 슬픔 및 절망감, 외로움, 범불안장애, 스트레스, 코로나로 인한 우울감 변화를 포함하는 정서적 특성인 것으로 나타났다. 본 연구 결과에서 확인된 청소년의 자살생각을 설명하는 정서적, 신체적, 생활 습관 관련 특성 및 학업 특성을 바탕으로 청소년 자살 예방을 위한 교육 및 상담적 개입 방안에 대해 논의하였다.
The aim of this study was to develop a predictive model using the random forest machine learning method that identifies suicidal ideation in South Korean adolescents. Additionally, this study sought to evaluate its predictive performance, identify the factors that affect such ideation, and propose measures for preventing youth suicide based on the developed model. The analysis employed responses of 54,565 adolescents attending middle and high schools in South Korea to the Korea Centers for Disease Control and Preventions’ 2021 youth health behavior survey. The adolescent suicidal ideation prediction model was estimated using 37 predictors that were categorized as being related to mental health, such as feelings of sadness and hopelessness, loneliness, and stress; physical health, such as height, weight, eating habits, and sleep duration; time utilization, such as how much physical activity they get and how much they use smartphones; health-related school education; academic performance; their physical environment; and changes due to COVID-19, such as economic difficulties, depressive symptoms, and physical activity. The results indicated that the relevant factors varied by school level, but emotional factors, such as feelings of sadness and hopelessness, loneliness, anxiety, and depressive symptoms caused by COVID-19 and generalized anxiety disorder, were the most important. This paper proposes educational and counseling intervention plans to prevent adolescent suicide based on the emotional, physical, lifestyle, and academic factors that were shown to affect adolescent suicidal ideation in this study.
랜덤 포레스트를 활용한 고등학생의 학교도서관 이용 예측 요인 탐색
[NRF 연계] 한국도서관·정보학회 한국도서관·정보학회지 Vol.57 No.1 2026.03 pp.317-336
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 랜덤 포레스트를 활용하여 고등학생의 학교도서관 이용을 예측하는 설명변수를 탐색하는 데 목적이 있다. 이를 위해 대전 지역 고등학생을 대상으로 설문조사를 실시하여 총 613명의 응답을 분석에 활용하였다. 가정 환경, 학교 환경, 학습 태도 및 습관, IT 환경, 개인 특성의 5개 영역 109개 설명변수를 설정하고, 학교도서관 이용 여부에 대한 변수의 상대적 중요도를 분석하였다. 분석 결과, 학교도서관 담당 교사 만족도, 수업 시간 도서 활용 활동, 동아리 활동 참여, 숏폼 시청 시간, 부모님 도서 추천, SNS 사용 시간, 학교도서관 관심 주제 도서 보유 등이 고등학생의 학교도서관 이용 여부를 예측하는 주요 변수로 확인되었다. 이러한 분석 결과를 근거로 학교도서관 담당 교사의 전문성 강화 및 역할 확대, 교과 연계 도서관 활용 수업 확대 및 동아리 활동과의 연계, 학생의 자율적 독서 환경 조성과 디지털 미디어의 도서관 서비스 연계, 가족 독서 문화 조성을 위한 지원 등을 제언하였다.
This study aims to explore predictor variables for high school students’ use of school libraries by employing random forest. A survey was conducted with high school students in the Daejeon region, and a total of 613 responses were used for the final analysis. A total of 109 predictor variables across five domains?home environment, school environment, learning attitudes and habits, IT environment, and personal characteristics?were established, and the relative importance of each variable in predicting school library use was analyzed. The analysis revealed that student satisfaction with school library teachers, library book utilization activities during class time, participation in club activities, short-form video viewing time, parental book recommendations, social media usage time, and availability of books on topics of interest in the school library were confirmed as key predictors of high school students’ school library use. Based on these findings, this study suggests strengthening the expertise and expanding the roles of school library teachers, expanding curriculum-linked library utilization activities and linking club activities with library use, fostering student-driven reading environments with integration of digital media into library services, and supporting the development of a family reading culture.
랜덤 포레스트 머신러닝 분석을 활용한 장애인의 정신건강 위험 요인 탐색: 임금근로 장애인을 중심으로
[NRF 연계] 한국장애인고용공단 고용개발원 장애와 고용 Vol.35 No.2 2025.05 pp.237-259
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
연구문제: 이 연구는 장애인의 정신건강 위험 요인을 랜덤 포레스트 머신러닝을 활용하여 분석하고, 주요 예측변수를 도출하는 것을 목표로 하였다. 연구방법: 한국장애인고용패널조사(2024) 8차 데이터를 활용하여 3,736명의 장애인을 대상으로 정신건강 위험군을 예측하는 53개 변수를 분석하였다. 이를 위해 랜덤 포레스트 모델의 성능을 평가하고, 주요 위험 요인을 도출하였으며, 데이터 분석은 SPSS 26.0과 통계 패키지 R 4. 4. 2 버전을 활용하였다. 연구결과: 변수 범주별로는 개인 내적 요인 4개, 장애특성 요인 2개, 경제 요인 2개, 근로 및 직무 요인 1개, 생활습관 요인 1개가 정신건강 위험 예측에 상위 10위의 높은 중요도 지수를 가진 변수들로 나타났다. 그리고 랜덤 포레스트로 구성한 예측모델의 정확도가 높아 전반적으로 신뢰할 만한 결과를 제공하였다. 논의: 이 연구는 랜덤 포레스트 머신러닝을 활용한 장애인의 정신건강 위험 예측 모델을 제시하였으며, 장애수용을 비롯한 장애인에 대한 심리·경제적 지원이 중요한 개입 요인임을 확인하였다. 이러한 점을 바탕으로 장애인 정신건강 증진에 관한 구체적인 정책적 개입과 맞춤형 지원의 필요성을 논의하였다.
Objective: This study aimed to analyze the risk factors for mental health in individuals with disabilities using the Random Forest machine learning technique and to identify key predictive variables. Methodology: The study utilized data from the 2024 Korean Panel Survey of Employment for the Disabled (KPSED), analyzing 53 variables to predict mental health risk groups among 3,736 individuals with disabilities. The performance of the Random Forest model was evaluated, and key risk factors were identified. And data analysis was conducted using SPSS 24.0 and R 4.4.2 Findings: The results indicated that four personal internal factors, two disability-related factors, two economic factors, one work-related factor, and one lifestyle factor were among the top 10 variables with the highest importance scores in predicting mental health risks. Furthermore, the prediction model constructed with the Random Forest method demonstrated high accuracy, providing reliable results. Conclusion: This study presents a predictive model for mental health risks in individuals with disabilities using machine learning, highlighting the significance of psychological and economic support alongside disability acceptance. Based on these findings, the study discusses the need for concrete policy interventions and tailored support to promote mental health in individuals with disabilities.
랜덤 포레스트 기반 분위수 회귀와 신뢰구간 보정을 통한 셰일저류층 궁극가채량 예측의 불확실성 정량화
[NRF 연계] 한국자원공학회 한국자원공학회지 Vol.62 No.4 2025.08 pp.400-413
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
이 연구에서는 Eagle Ford Shale 지역의 셰일 오일 생산정을 대상으로 랜덤 포레스트 분위수 회귀기반의 궁극가채량 예측 및 불확실성 정량화를 수행하였다. 랜덤 포레스트 분위수 회귀로부터 도출된 P90 및 P10 방식은 예측구간 포함확률(PICP)이 80%를 과도하게 초과하여 불확실성 분석에적합하지 않은 한계를 보였다. 이를 개선하기 위해 수정된 적합분위수 회귀법을 통한 신뢰구간 보정 기법을 적용하였다. 보정된 모델은 PICP 80%에 해당하는 보정된 P90 및 P10 값을 도출함으로써 예측 신뢰구간의 적정성을 확보하였다. 또한 중앙값 예측은 보정 전후 변화가 거의 없어 모델의구조적 안정성이 유지됨을 확인하였다. 제안된 방법은 셰일저류층 생산 예측의 불확실성 정량화에 대한 신뢰성과 실용성을 향상시킬 수 있는 효과적인 대안이 될 수 있다.
This study uses quantile regression forest (QRF), which is based on random forest (RF), to estimate the estimated ultimate recovery of shale oil in the Eagle Ford Shale and quantify uncertainty. Conventional prediction intervals using the P90 and P10 quantiles from the RF quantile model yielded higher prediction interval coverage probability (PICP) than the target of 80%, leading to uncertainty overestimation. A calibration method based on modified conformalized quantile regression (CQR) was used for the prediction intervals to solve this problem. By generating adjusted P90 and P10 values, the calibrated model ensured a PICP of 80% and a more appropriate uncertainty representation. The median prediction remained nearly unchanged before and after calibration, confirming the structural stability of the model. The proposed method enhances the reliability and practical applicability of uncertainty quantification in forecasting production from shale reservoirs.
취업한 발달장애인의 삶의 만족 예측요인 탐색: 랜덤 포레스트와 SHAP 기반 분석
[NRF 연계] 한국장애인고용공단 고용개발원 장애와 고용 Vol.35 No.3 2025.08 pp.327-355
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
연구 목적: 본 연구는 취업한 발달장애인의 삶의 만족에 영향을 미치는 요인을 탐색하고자 당사자 자기보고와 보호자 응답 자료를 함께 활용하여, 다양한 요인을 통합적으로 분석하였다. 연구방법:「2023년 발달장애인 일과 삶 실태조사」에서 취업자 중, ‘보통 읽기’ 집단 463명의 자료를 분석하였고, 랜덤 포레스트와 SHAP 분석을 적용하였으며, 다중회귀분석을 병행하였다. 연구결과: 20개의 주요 변수로, 당사자 응답에서 함께 일하는 사람, 자아존중감 등 9개, 보호자 응답에서 전반적 건강 상태, 당사자의 만족도 등 11개가 도출되었다. SHAP 분석은 변수의 영향 방향과 상대적 기여도를 시각화하였고, 회귀분석에서는 9개의 변수가 유의하였다. 결론 또는 시사점: 본 연구는 당사자와 보호자의 응답을 함께 분석함으로써 삶의 만족에 대한 입체적 이해를 도모하였고, 당사자 중심의 개입과 보호자 요인을 함께 고려한 정책 설계의 필요성을 제시하였다.
Purpose: This study aimed to explore the factors influencing life satisfaction among employed individuals with developmental disabilities by utilizing both self-reports from the individuals and responses from their guardians, thereby enabling a comprehensive analysis of related variables. Method: Data from 463 employed individuals in the “average reading” group were drawn from the 2023 Survey on Work and Life of Persons with Developmental Disabilities. Random forest and SHAP (Shapley Additive Explanations) analyses were conducted to examine variable importance and interaction effects, while multiple regression analysis was employed to test statistical significance. Results: A total of 20 key variables were identified?9 from the self-reports of individuals (e.g., relationships with coworkers, self-esteem) and 11 from guardians’ responses (e.g., overall health status, perceived satisfaction of the individuals). SHAP analysis visualized the relative contribution and directional influence of these variables, and 9 variables showed statistical significance in the regression analysis. Conclusion: By jointly analyzing responses from both individuals and their guardians, this study provides a multidimensional understanding of life satisfaction among people with developmental disabilities. It highlights the necessity for policy development that incorporates both individual-centered interventions and guardian-related factors.
랜덤 포레스트를 활용한 장애인 취업 예측 모형: 취업 여부 및 정규직 취업을 중심으로
[NRF 연계] 한국장애인고용공단 고용개발원 장애와 고용 Vol.29 No.3 2019.08 pp.145-165
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
연구목적: 본 연구는 최근 새롭게 주목받고 있는 머신 러닝 기법을 적용하여 장애인의 취업 여부와 정규직 취업에 영향을 미치는 변수들을 탐색적으로 규명하고자 하였다. 연구방법: 장애인고용패널 2차 웨이브의 2차 연도(2017년) 자료를 대상으로 대표적인 머신러닝 기법 중 하나인 랜덤 포레스트(Random Forest)를 통하여 179개 설명변수를 활용한 장애인 취업 예측 모형의 성능을 확인하고, 장애인의 취업에 영향을 미치는 요인들을 탐색적으로 추정해보았다. 연구결과: 분석 결과, 기존의 직장 경험 및 기대 임금 수준, 가계의 경제적 수준 등이 장애인의 취업에 중요한 요인임을 확인하였다. 그리고 장애인의 정규직 취업에는 직업 능력의 중요도가 상대적으로 높게 나타남을 확인하였다. 이와 함께 기존 선행연구들과 달리 주관적 인식 변수들이 장애인들의 취업 여부와 정규직 취업에 중요한 영향을 미치는 것 또한 확인하였다. 결론: 향후의 장애인 고용 지원 관련 연구 및 정책 입안에 있어 머신 러닝 기법의 활용 가능성을 확인하였다는 점에서 연구의 의의를 찾을 수 있다. 그리고 그 과정에서 연구자는 단순히 머신 러닝의 결과물을 수동적으로 받아들이는 객체가 아니라 연구 과정을 주도해 나가는 주체로서 기능할 수 있도록 연구자 및 정책 입안자들의 지속적인 관심과 노력이 필요함을 제안하였다.
Objectives: The purpose of this paper is to exploratively investigate predictive variables on the employment status and quality of employment of persons with disabilities. Method: This study utilized the random forest method, one of the most representative machine learning approach. Results: Our results first show that previous working experience, expect wage, and economic status of family are the substantial variables in predicting the employment status of persons with disabilities. In terms of the quality of employment, vocational capacity of the disadvantaged has a more predictive performance relatively. Moreover, this study identified subjective variables, such as quality of life and conception regarding employment, have an important predictive power on the employment status and quality of employment of persons with disabilities. Conclusion: The findings of this study, based on new approach using machine learning method, provide useful policy implications for improving the employment status and quality of employment of persons with disabilities.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.