Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 222
No
1

랜덤 포레스트(Random Forest)의 시계열 적용에 관한 연구: 한국 물가상승률 예측 사례 분석

한희준

[NRF 연계] 한국경제학회 경제학연구 Vol.71 No.3 2023.09 pp.37-73

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본고는 한국과 미국의 물가상승률 예측에 랜덤 포레스트 모형을 적용할 때, Stationary Bootstrap이나 Moving Block Bootstrap 등 Block Bootstrap을 사용하는 것이 통상적인 독립 부트스트랩(Independent Bootstrap)을 사용하는 것에 비해 통계적으로 유의한 수준으로 예측력을 개선하지는 못한다는 것을 보인다. 그리고 FRED-MD를 참고한 총 93개의 관련 국내외 거시경제/금융 변수들을 사용하고, XGBoost, LSTM 등 다양한 머신러닝 방법을 활용하여 한국의 물가상승률을예측하고 분석한다. 2004년 9월에서 2022년 3월까지의 표본을 이용하였고, 1개월에서 12개월의 예측 대상기간(Forecast Horizon)을 고려하였다. 총 13개의 모형 중 대부분의 예측 대상기간에 있어 예측력이 우수한 모형이 존재하는 것으로나타났는데, 이는 보루타 알고리즘(Boruta Algorism)을 통해 중요한 변수로 분류된 변수들만을 랜덤 포레스트에 적용하는 모형이다. Giacomini and White(2006) 와 Hansen et al.(2009)의 검정을 통해 대부분의 예측 대상기간에서 통계적으로유의하게 예측력이 우수함을 확인하였는데, 특히 경제활동인구 및 취업자 수의증가율 등 고용시장 관련 변수, 기업경기실사지수, 주택가격 변화율 등이 물가상승률 예측에 중요한 변수로 선택되는 것으로 나타났다.

This paper first investigates whether adopting the stationary bootstrap or the moving block bootstrap, instead of the usual independent bootstrap, in the random forest method improves forecasting of stationary time series. It is shown that the block bootstrap procedures adopted in the random forest method do not make any statistically significant improvement in Korean or US inflation forecasting. Secondly, we consider inflation forecasting in Korea using 93 macroeconomic/financial variables and various machine learning methods. The samples are from September 2004 to March 2022. Comparing total 13 models, one model outperforms the rest models for most forecast horizons, which is a simplified method of the model proposed by Kim and Han (2022). The method consists of the following two steps: 1) Select important variables based on the Boruta algorithm, 2) Using only those selected variables, implement the random forest and produce a forecast. The tests by Giacomini and White (2006) and Hansen et al. (2009) show that the model provides significantly better forecasts for most forest horizons. In particular, the Boruta algorithm selected total economically active population, total employed persons, BSI, house price as important variables for Korean inflation forecasting.

2

Ransomware Detection using Random Forest Technique

Ban Mohammed Khammas

[NRF 연계] 한국통신학회 ICT Express Vol.6 No.4 2020.12 pp.325-331

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Nowadays, the ransomware became a serious threat challenge the computing world that requires an immediate consideration to avoid financial and moral blackmail. So, there is a real need for a new method that can detect and stop this type of attack. Most of the previous detection methods followed a dynamic analysis technique which involves a complicated process. The present study proposes a novel method based on static analysis to detect ransomware. The significant characteristic of proposed method is dispensing of disassemble process by direct extraction of features from raw byte with the use of frequent pattern mining which remarkably increases the detection speed. The Gain Ratio technique was used for feature selection which exhibited that 1000 features was the optimal number for detection process. The current study involved using random forest classifier with a comprehensive analysis to the effect of both tree and seed numbers on the ransomware detection. The results showed that tree numbers of 100 with seed number of 1 achieved best results in terms of time-consuming and accuracy. The experimental evaluation revealed that the proposed method could achieve a high accuracy of 97.74% for detection ransomware.

3

Novel hyper-tuned ensemble Random Forest algorithm for the detection of false basic safety messages in Internet of Vehicles

Goodness Oluchi Anyanwu, Cosmas Ifeanyi Nwakanma, 이재민, Dong-Seong Kim

[NRF 연계] 한국통신학회 ICT Express Vol.9 No.1 2023.02 pp.122-129

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Detection of nodes disseminating false data is a prerequisite for effective deployment of Internet of Vehicles (IoV) services. This work proposed a novel hyper-tuned ensemble Random Forest (Ens. RF) algorithm to detect false basic safety messages in IoV. Performance evaluation was done using the Vehicular Reference Misbehavior (VeReMi) dataset comprising data-centric misbehavior evaluation for vehicular networks. For validation, a comparative analysis of the performance of the proposed “Ens. RF” model, five machine learning algorithms implemented in this work, and state-of-the-art ML models from related literature was presented. The performance metrics considered are time efficiency and validation accuracy for overall misbehavior classification. Also, the results confirmed the irrelevance of data balancing in real-life scenarios. Finally, we assess the performance of our proposed system for detecting each falsification scenario using precision and recall. The result shows that the proposed algorithm outperformed others with a validation accuracy of 99.60% and a negligible 604 misclassifications out of 153,730 points.

4

Predictive factors of adolescents’ happiness: a random forest analysis of the 2023 Korea Youth Risk Behavior Survey

김은주, 김성광, 정성혜, 류요셉

[NRF 연계] 한국아동간호학회 Child Health Nursing Research Vol.31 No.2 2025.04 pp.85-95

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Purpose: This study aimed to identify predictive factors affecting adolescents’ subjective happiness using data from the 2023 Korea Youth Risk Behavior Survey. A random forest model was applied to determine the strongest predictive factors, and its predictive per- formance was compared with traditional regression models. Methods: Responses from a total of 44,320 students from grades 7 to 12 were analyzed. Data pre-processing involved handling missing values and selecting variables to con- struct an optimal dataset. The random forest model was employed for prediction, and SHAP (Shapley Additive Explanations) analysis was used to assess variable importance. Results: The random forest model demonstrated a stable predictive performance, with an R of .37. Mental and physical health factors were found to significantly affect subjec-2tive happiness. Adolescents’ subjective happiness was most strongly influenced by per- ceived stress, perceived health, experiences of loneliness, generalized anxiety disorder, suicidal ideation, economic status, fatigue recovery from sleep, and academic perfor- mance. Conclusion: This study highlights the utility of machine learning in identifying factors influencing adolescents’ subjective happiness, addressing limitations of traditional re- gression approaches. These findings underscore the need for multidimensional interven- tions to improve mental and physical health, reduce stress and loneliness, and provide integrated support from schools and communities to enhance adolescents’ subjective happiness.

5

Factors Influencing Nurses' Intention to Report Adverse Nursing Events in a Tertiary Hospital Using a Random Forest Model

Yuxuan Chen, Wanhong Ding, Jianming Xu, Qi Zhang, Wei Qin

[NRF 연계] 한국간호과학회 Asian Nursing Research Vol.19 No.4 2025.10 pp.347-353

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Purpose: The goal of this study was to explore the current status and influencing factors of nurses'intention to report adverse nursing events, using a random forest model to rank and select these factors. Methods: In August 2024, a survey was conducted among nurses at a tertiary hospital in Shanghai usinga convenience sampling method. The instruments included the General Information Questionnaire, theIntention to Report Questionnaire, the Reporting of Clinical Adverse Effects Scale (RoCAES), and theReporting Barrier Questionnaire. A random forest model and stepwise multiple linear regression wereused to analyze the factors influencing nurses' intention to report nursing adverse events. Results: The median total score for the intention to report adverse events among 823 nurses was 14(range: 0 to 15). Lasso regression identified eight key factors influencing the intention to report adverseevents: professional title, whether the nurse had reported adverse events, age, punitive culture,reporting process, reporting significance, reporting purpose, and reporting environment. Stepwiseregression analysis revealed that younger nurses, those with no prior reporting experience, and thosewho perceived reporting as important had a higher intention to report adverse events. Additionally,nurses were more likely to report adverse events when the hospital's reporting process was simpler andthe culture was less punitive. Conclusions: Nurses' intention to report adverse events in nursing is moderate. Nursing managers shouldconsider these influencing factors holistically and implement targeted measures to enhance nurses'intention to report, thereby improving nursing quality and patient safety.

6

4,000원

연구목적: 전국 전체의 이륜차 사고 건수는 점점 감소하고 있지만 사망자 수는 증가한 것으로 나타났으 며 이륜차 사고가 집중되는 시간대는 16~22시인 저녁 시간대로 나타났다. 따라서 주거지역과 다른 지 역 간의 이륜차 사고 심각도 요인의 차이를 비교할 필요가 있다. 연구방법: Random Forest를 활용해 2007년부터 2021년까지 인천광역시에서 발생한 이륜차 사고 자료와 QGIS를 사용하여 도시지역 및 비 도시지역으로 구분하고 지역에 따라 사고 심각도 분석 및 사고 심각도에 끼치는 영향요인을 분석하였 다. 연구결과: 각 지역의 상위 2개의 변수 중요도를 분석한 결과, 도시 지역 이륜차 사고 심각도에 높은 영향을 끼치는 변수는 교차로 안에서 사고가 발생한 경우, 신호위반의 경우이고, 공업지역은 차대차 정 면충돌의 경우와 피해자 차종이 화물차인 경우, 녹지지역은 신호위반의 경우, 교차로 안에서 사고가 발 생한 경우, 상업지역은 신호위반,차대차- 측면직각충돌의 경우, 비도시지역은 가해운전자가 60세 이상 인 경우, 피해자가 40~50대인 경우 사고심각도에 더 영향을 주는 것으로 나타났다. 결론: 지역의 특성에 따라 이륜차 사고의 심각도에 영향을 끼치는 요인을 집중적으로 분석하였고 본 연구로 도출된 결과를 활용하여 이륜차 사고의 심각도를 개선하기 위한 효율적이고 대응책 수립이 필요하다.

Purpose: The number of motorcycle accidents nationwide is decreasing, but the number of deaths has increased, and the time when motorcycle accidents are concentrated is between 16:00 and 22:00 in the evening. Therefore, it is necessary to compare the differences in the severity factors of motorcycle accidents between residential areas and other areas. Method: Using Random Forrest, QGIS and data on motorcycle accidents in Incheon Metropolitan City from 2007 to 2021 were used to divide them into urban and non-urban areas to derive accident severity analysis and impact factors on accident severity by region. Result: As a result of analyzing the importance of the top two variables in each region, the variables that have a high impact on the seriousness of motorcycle accidents in urban areas are traffic lights in intersections Conclusion: It is necessary to establish efficient and countermeasures to improve the severity of motorcycle accidents by analyzing in-depth factors that affect the severity of motorcycle accidents according to the characteristics of the region and utilizing the results derived from this study.

7

Random Forest를 활용한 고속도로 교통사고 심각도 비교분석에 관한 연구 KCI 등재

이선민, 윤병조, 웃위린

한국재난정보학회 한국재난정보학회논문집 제20권 1호 통권63호 2024.03 pp.156-168

※ 기관로그인 시 무료 이용이 가능합니다.

4,500원

연구목적: 고속도로 교통사고의 추세는 증감을 반복하며 도로 종류 중 고속도로에서의 치사율은 최고치 를 나타내고 있다. 따라서 국내 실정을 반영한 개선대책 수립이 필요하다. 연구방법: Random Forest를 활 용해 2019년부터 2021년까지 전국 고속도로 노선 중 사고 다발 10개 노선에서 발생한 교통사고 자료로 사고 심각도 분석 및 사고 심각도에 미치는 영향요인을 도출하였다. 연구결과: SHAP 패키지를 활용해 상 위 10개의 변수 중요도를 분석한 결과, 고속도로 교통사고 중 사고 심각도에 높은 영향을 미치는 변수는 가해자 연령이 20세 이상 39세 미만, 시간대가 주간(06:00-18:00), 주말(토~일), 계절이 여름과 겨울, 법규 위반이 안전운전불이행, 도로 형태가 터널, 기하구조상 차로 수가 많고 제한속도가 높은 경우로 총 10개 의 독립변수에서 고속도로 교통사고 심각도와 양(+)의 상관관계를 가지는 것으로 분석되었다. 결론: 고속 도로에서의 사고 발생은 매우 다양한 요인의 복합적인 작용으로 인해 발생하므로 사고 예측에 많은 어려 움이 있지만 본 연구로 도출된 결과를 활용해 고속도로 교통사고 심각도에 영향을 주는 요인을 심층적으 로 분석해 효율적이고 합리적인 대응책 수립을 위한 노력이 필요하다.

Purpose: The trend of highway traffic accidents shows a repeating pattern of increase and decrease, with the fatality rate being highest on highways among all road types. Therefore, there is a need to establish improvement measures that reflect the situation within the country. Method: We conducted accident severity analysis using Random Forest on data from accidents occurring on 10 specific routes with high accident rates among national highways from 2019 to 2021. Factors influencing accident severity were identified. Result: The analysis, conducted using the SHAP package to determine the top 10 variable importance, revealed that among highway traffic accidents, the variables with a significant impact on accident severity are the age of the perpetrator being between 20 and less than 39 years, the time period being daytime (06:00-18:00), occurrence on weekends (Sat-Sun), seasons being summer and winter, violation of traffic regulations (failure to comply with safe driving), road type being a tunnel, geometric structure having a high number of lanes and a high speed limit. We identified a total of 10 independent variables that showed a positive correlation with highway traffic accident severity. Conclusion: As accidents on highways occur due to the complex interaction of various factors, predicting accidents poses significant challenges. However, utilizing the results obtained from this study, there is a need for in-depth analysis of the factors influencing the severity of highway traffic accidents. Efforts should be made to establish efficient and rational response measures based on the findings of this research.

8

4,500원

This study investigates the changes in factors influencing aircraft utilization based on aircraft types during the pandemic period. The analysis categorizes the periods into three phases: pre-pandemic, during the pandemic, and post-pandemic, while the aircraft types are classified into two groups (C-class and E-class or higher). Utilizing SHAP and Random Forest methodologies, the study analyzes the importance of various factors for each group and examines the significance of differences in importance across the specified periods and aircraft types through group difference testing. The findings reveal that for C-class aircraft, the importance of the USD/KRW exchange rate and the lagging composite index was notably high before the pandemic. However, during the pandemic, there was a substantial surge in the significance of the JPY/KRW exchange rate, and post-pandemic, the importance of economic indicators increased. E-class or higher aircraft exhibited a similar trend, with a significant rise in the importance of the CNY/KRW exchange rate observed after the pandemic. The Kruskal-Wallis test confirmed that most average changes among the groups were not statistically significant, except for the C-class aircraft in the Random Forest analysis. This indicates that there were no substantial changes in the underlying factors between the pre-pandemic and post-pandemic periods, suggesting that the impact of the pandemic on airline fleet utilization extends beyond short-term changes. Furthermore, the findings highlight the need for airlines to enhance their sensitivity to external environments and develop utilization strategies in volatile markets. Additionally, there is a possibility that long-term dependency on network strategies may become more pronounced than reliance on external factors like the pandemic. By providing an in-depth analysis of the pandemic's impact on aircraft utilization, this study contributes to the formulation of effective utilization strategies for airlines navigating an uncertain economic environment.

9

4,500원

This study conducts a comparative analysis of various models for predicting Bitcoin prices. The models utilized in this research include the Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), and Random Forest. Their performance was evaluated using key metrics such as Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R-squared (R²). The findings reveal that the GRU model outperformed the other two models, achieving an R² value of 0.9301, which indicates its superior ability to explain the data. The LSTM model demonstrated relatively high explanatory power with an R² of 0.8252, yet it exhibited lower predictive accuracy compared to the GRU. On the other hand, the Random Forest model recorded the lowest MAE (1557.44) and RMSE (1766.95), reflecting strong short-term prediction capabilities. However, its R² of 0.6110 highlighted limitations in capturing the intricate temporal patterns of the dataset. The significance of this study lies in several aspects. First, it empirically demonstrates that the GRU model, with its more streamlined architecture, can achieve superior predictive performance compared to the LSTM and Rando Forest models. Additionally, while Random Forest showed promising results for short-term predictions, it underscored limitations in modeling the complexities of temporal data patterns. Future research should aim to incorporate broader datasets and integrate external economic variables to enhance predictive accuracy. Furthermore, the potential of ensemble approaches, combining diverse models, holds promise for further improving predictive performance in this domain.

10

4,000원

연구목적: 본 연구는 서울특별시 서초구를 대상으로 강우, 소방 활동, 사회 취약성 데이터를 통합하여 복합재난 위험 도를 예측하고, 이를 기반으로 소방 자원 배치 전략을 제시하는 것을 목적으로 한다. 연구자료는 2020–2024년 서초 구 시간별 강우량 자료, 화재 출동 개별 내역, 119 구조·구급활동 실적, 통계 연보에서 추출한 인구·주거·사회 취약성 지표를 사용하였다. 특히 신고일시·출동일시·도착일시를 활용하여 출동 지연, 도착시간, 총 대응 시간을 산출하고, 65세 이상 인구 비율, 장애인 등록 인구 비율, 1인 가구 비율, 반지하·지하 주택 비율을 사회 취약성 변수로 구성하였 다. 연구방법: 분석 방법으로는 먼저 화재 출동 내역을 이용해 연도·월·시간대별 출동 건수와 평균 도착시간, 평균 총 대응 시간을 산출하여 기초 통계를 제시하고, 강우량 및 사회 취약성 지표와 결합하여 복합재난 발생에 영향을 미치 는 설명변수를 구성하였다. 이후 다중회귀분석(Multiple Regression Analysis, MRA)과 Random Forest(RF)를 이용하 여 시간 단위 화재 출동 건수 및 발생 여부를 예측하고, 두 모형의 성능을 비교하였다. 2024년을 검증자료로 사용한 결과, 시간당 화재 출동 건수 예측에서 MRA 모형의 평균절대오차(MAE)는 0.24건, RF 모형의 MAE는 0.18건으로, RF가 약 25.0% 낮은 예측 오차를 보였다. 화재 발생 여부(해당 시간에 1건 이상 발생)의 분류 성능을 AUC로 비교한 결과, 로지스틱 회귀는 0.83, RF 분류기는 0.90을 기록하여 RF가 우수한 성능을 나타냈다. 연구결과: RF 변수 중요도 분석 결과, 이전 24시간 누적 강우량, 현 시각 강우량, 신고 시간의 시간대(야간 여부), 65세 이상 인구 비율, 장애인 등록 인구 비율, 반지하·지하 주택 비율, 1인 가구 비율순으로 중요도가 높게 나타났다. 이는 강우 특성뿐 아니라 고령자 및 장애인 밀집 지역, 주거 취약 지역, 1인 가구 밀집 지역이 복합재난 위험에 유의 미한 영향을 미침을 시사한다. 마지막으로, 예측된 위험도와 사회 취약성 지수를 결합하여 행정동별 우선 대응 순위를 제시하고, 이를 기반으로 소방 자원 배치 시나리오를 제안하였다. 결론: 본 연구는 기상·소방·사회통계 데이터를 결합한 복합재난 위험도 예측과 사회 취약성을 고려한 자원 배치 전략을 함께 제시함 으로써, 도시형 재난관리에 대한 데이터 기반 의사결정 지원에 기여할 수 있을 것으로 기대된다.

Purpose: This study aims to predict the risk of compound disasters in Seocho District, Seoul, by integrating rainfall, firefighting activity, and social vulnerability data, and to propose strategies for the optimal allocation of firefighting resources. Method: The research employs hourly rainfall data from 2020 to 2024, detailed fire dispatch records, 119 rescue and emergency activity logs, and demographic and social vulnerability indicators extracted from statistical yearbooks, and calculates dispatch delays, arrival times, and total response times using reported, dispatch, and arrival timestamps. Social vulnerability variables include the proportion of residents aged 65 or older, the proportion of registered persons with disabilities, the share of single-person households, and the percentage of semi-basement or basement dwellings. Result: Multiple Regression Analysis and Random Forest models are applied to predict hourly fire dispatch counts and fire occurrence, and their predictive performances are compared using 2024 data as a validation set; results show that the Random Forest model achieves a lower mean absolute error than the regression model in count prediction and a higher AUC in occurrence classification.Variable importance analysis further indicates that cumulative rainfall over the previous 24 hours, current rainfall, time of report (day/night), and the proportions of elderly residents, registered disabled residents, semi-basement or basement dwellings, and single-person households are major contributors to predicted compound disaster risk. Conclusion: Overall, the study concludes that integrating meteorological, firefighting, and socio-demographic data enables more accurate compound disaster risk prediction and supports socially responsive resource allocation strategies, thereby contributing to data-driven decision-making in urban disaster management.

11

Acute myocardial infarction (AMI) is one of the leading causes of cardiovascular disease-related mortality worldwide, causing ischemic damage to the heart muscle. AMI poses a risk of sudden cardiac arrest and is associated with a high rate of recurrence, which can significantly impact daily life. Therefore, this study aims to develop a predictive model for mortality in AMI patients using machine learning. The data used in the study are from KNHDIS (2013-2022) and include demographic characteristics and disease information of discharged patients. The model was constructed using RFECV and Random Forest. Key variables influencing mortality include age, number of surgeries, and length of hospital stay. The model demonstrated high performance with an accuracy of 94% and an AUC value of 0.93.

12

4,000원

In this study, AAHT was estimated using machine learning algorithms of ensemble techniques to improve the limitations of existing studies. Among the machine learning algorithms, a random forest algorithm was used, and the traffic volume estimation accuracy was analyzed as MAPE 20.7%. It was analyzed that the higher the traffic volume level, the higher the traffic volume estimation accuracy, and the traffic volume level of 3,000 vehicles/hour or more was analyzed to be 8% or less of MAPE. It was analyzed that the accuracy secured at the current level was very high compared to the existing model.

13

청소년 자살생각의 예측요인 고찰: 랜덤포레스트 방법의 적용

이은설, 박수원, 김은하

[NRF 연계] 한국심리학회 산하 한국학교심리학회 한국심리학회지: 학교 Vol.21 No.1 2024.04 pp.1-26

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 랜덤포레스트 머신러닝 방법을 적용하여 청소년의 자살생각을 예측하는 모형을 구성하고, 모형의 예측 성과를 평가하며, 예측에 기여하는 요인의 특성을 파악하여 이를 바탕으로 청소년 자살예방 방안을 제시하는 것을 목적으로 하였다. 이를 위해 질병관리청의 2021년도 청소년 건강행태조사 자료를 분석에 사용하였으며, 구체적으로 국내 중⋅고등학교에 재학 중인 총 54,565명의 청소년의 응답자료를 분석하였다. 청소년 자살생각의 예측모형은 정신건강(예: 슬픔 및 절망감, 외로움, 스트레스 등), 신체적 건강(예: 키, 몸무게, 식이습관, 수면시간 등), 시간활용(예: 신체활동, 스마트폰 사용 시간), 건강 관련 학교 교육, 학업성적, 물리적 환경, 코로나로 인한 변화(예: 경제적 어려움, 우울감, 신체활동 등) 등의 영역에 해당하는 37개의 예측변수를 포함하였으며 랜덤포레스트 머신러닝 방법을 적용하여 추정하였다. 예측에 중요하게 역할하는 요인은 중학생 및 고등학생 모두에서 슬픔 및 절망감, 외로움, 범불안장애, 스트레스, 코로나로 인한 우울감 변화를 포함하는 정서적 특성인 것으로 나타났다. 본 연구 결과에서 확인된 청소년의 자살생각을 설명하는 정서적, 신체적, 생활 습관 관련 특성 및 학업 특성을 바탕으로 청소년 자살 예방을 위한 교육 및 상담적 개입 방안에 대해 논의하였다.

The aim of this study was to develop a predictive model using the random forest machine learning method that identifies suicidal ideation in South Korean adolescents. Additionally, this study sought to evaluate its predictive performance, identify the factors that affect such ideation, and propose measures for preventing youth suicide based on the developed model. The analysis employed responses of 54,565 adolescents attending middle and high schools in South Korea to the Korea Centers for Disease Control and Preventions’ 2021 youth health behavior survey. The adolescent suicidal ideation prediction model was estimated using 37 predictors that were categorized as being related to mental health, such as feelings of sadness and hopelessness, loneliness, and stress; physical health, such as height, weight, eating habits, and sleep duration; time utilization, such as how much physical activity they get and how much they use smartphones; health-related school education; academic performance; their physical environment; and changes due to COVID-19, such as economic difficulties, depressive symptoms, and physical activity. The results indicated that the relevant factors varied by school level, but emotional factors, such as feelings of sadness and hopelessness, loneliness, anxiety, and depressive symptoms caused by COVID-19 and generalized anxiety disorder, were the most important. This paper proposes educational and counseling intervention plans to prevent adolescent suicide based on the emotional, physical, lifestyle, and academic factors that were shown to affect adolescent suicidal ideation in this study.

14

랜덤 포레스트를 활용한 고등학생의 학교도서관 이용 예측 요인 탐색

임정훈

[NRF 연계] 한국도서관·정보학회 한국도서관·정보학회지 Vol.57 No.1 2026.03 pp.317-336

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 랜덤 포레스트를 활용하여 고등학생의 학교도서관 이용을 예측하는 설명변수를 탐색하는 데 목적이 있다. 이를 위해 대전 지역 고등학생을 대상으로 설문조사를 실시하여 총 613명의 응답을 분석에 활용하였다. 가정 환경, 학교 환경, 학습 태도 및 습관, IT 환경, 개인 특성의 5개 영역 109개 설명변수를 설정하고, 학교도서관 이용 여부에 대한 변수의 상대적 중요도를 분석하였다. 분석 결과, 학교도서관 담당 교사 만족도, 수업 시간 도서 활용 활동, 동아리 활동 참여, 숏폼 시청 시간, 부모님 도서 추천, SNS 사용 시간, 학교도서관 관심 주제 도서 보유 등이 고등학생의 학교도서관 이용 여부를 예측하는 주요 변수로 확인되었다. 이러한 분석 결과를 근거로 학교도서관 담당 교사의 전문성 강화 및 역할 확대, 교과 연계 도서관 활용 수업 확대 및 동아리 활동과의 연계, 학생의 자율적 독서 환경 조성과 디지털 미디어의 도서관 서비스 연계, 가족 독서 문화 조성을 위한 지원 등을 제언하였다.

This study aims to explore predictor variables for high school students’ use of school libraries by employing random forest. A survey was conducted with high school students in the Daejeon region, and a total of 613 responses were used for the final analysis. A total of 109 predictor variables across five domains?home environment, school environment, learning attitudes and habits, IT environment, and personal characteristics?were established, and the relative importance of each variable in predicting school library use was analyzed. The analysis revealed that student satisfaction with school library teachers, library book utilization activities during class time, participation in club activities, short-form video viewing time, parental book recommendations, social media usage time, and availability of books on topics of interest in the school library were confirmed as key predictors of high school students’ school library use. Based on these findings, this study suggests strengthening the expertise and expanding the roles of school library teachers, expanding curriculum-linked library utilization activities and linking club activities with library use, fostering student-driven reading environments with integration of digital media into library services, and supporting the development of a family reading culture.

15

랜덤 포레스트 머신러닝 분석을 활용한 장애인의 정신건강 위험 요인 탐색: 임금근로 장애인을 중심으로

박우람, 김동일

[NRF 연계] 한국장애인고용공단 고용개발원 장애와 고용 Vol.35 No.2 2025.05 pp.237-259

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

연구문제: 이 연구는 장애인의 정신건강 위험 요인을 랜덤 포레스트 머신러닝을 활용하여 분석하고, 주요 예측변수를 도출하는 것을 목표로 하였다. 연구방법: 한국장애인고용패널조사(2024) 8차 데이터를 활용하여 3,736명의 장애인을 대상으로 정신건강 위험군을 예측하는 53개 변수를 분석하였다. 이를 위해 랜덤 포레스트 모델의 성능을 평가하고, 주요 위험 요인을 도출하였으며, 데이터 분석은 SPSS 26.0과 통계 패키지 R 4. 4. 2 버전을 활용하였다. 연구결과: 변수 범주별로는 개인 내적 요인 4개, 장애특성 요인 2개, 경제 요인 2개, 근로 및 직무 요인 1개, 생활습관 요인 1개가 정신건강 위험 예측에 상위 10위의 높은 중요도 지수를 가진 변수들로 나타났다. 그리고 랜덤 포레스트로 구성한 예측모델의 정확도가 높아 전반적으로 신뢰할 만한 결과를 제공하였다. 논의: 이 연구는 랜덤 포레스트 머신러닝을 활용한 장애인의 정신건강 위험 예측 모델을 제시하였으며, 장애수용을 비롯한 장애인에 대한 심리·경제적 지원이 중요한 개입 요인임을 확인하였다. 이러한 점을 바탕으로 장애인 정신건강 증진에 관한 구체적인 정책적 개입과 맞춤형 지원의 필요성을 논의하였다.

Objective: This study aimed to analyze the risk factors for mental health in individuals with disabilities using the Random Forest machine learning technique and to identify key predictive variables. Methodology: The study utilized data from the 2024 Korean Panel Survey of Employment for the Disabled (KPSED), analyzing 53 variables to predict mental health risk groups among 3,736 individuals with disabilities. The performance of the Random Forest model was evaluated, and key risk factors were identified. And data analysis was conducted using SPSS 24.0 and R 4.4.2 Findings: The results indicated that four personal internal factors, two disability-related factors, two economic factors, one work-related factor, and one lifestyle factor were among the top 10 variables with the highest importance scores in predicting mental health risks. Furthermore, the prediction model constructed with the Random Forest method demonstrated high accuracy, providing reliable results. Conclusion: This study presents a predictive model for mental health risks in individuals with disabilities using machine learning, highlighting the significance of psychological and economic support alongside disability acceptance. Based on these findings, the study discusses the need for concrete policy interventions and tailored support to promote mental health in individuals with disabilities.

16

취업한 발달장애인의 삶의 만족 예측요인 탐색: 랜덤 포레스트와 SHAP 기반 분석

허경화

[NRF 연계] 한국장애인고용공단 고용개발원 장애와 고용 Vol.35 No.3 2025.08 pp.327-355

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

연구 목적: 본 연구는 취업한 발달장애인의 삶의 만족에 영향을 미치는 요인을 탐색하고자 당사자 자기보고와 보호자 응답 자료를 함께 활용하여, 다양한 요인을 통합적으로 분석하였다. 연구방법:「2023년 발달장애인 일과 삶 실태조사」에서 취업자 중, ‘보통 읽기’ 집단 463명의 자료를 분석하였고, 랜덤 포레스트와 SHAP 분석을 적용하였으며, 다중회귀분석을 병행하였다. 연구결과: 20개의 주요 변수로, 당사자 응답에서 함께 일하는 사람, 자아존중감 등 9개, 보호자 응답에서 전반적 건강 상태, 당사자의 만족도 등 11개가 도출되었다. SHAP 분석은 변수의 영향 방향과 상대적 기여도를 시각화하였고, 회귀분석에서는 9개의 변수가 유의하였다. 결론 또는 시사점: 본 연구는 당사자와 보호자의 응답을 함께 분석함으로써 삶의 만족에 대한 입체적 이해를 도모하였고, 당사자 중심의 개입과 보호자 요인을 함께 고려한 정책 설계의 필요성을 제시하였다.

Purpose: This study aimed to explore the factors influencing life satisfaction among employed individuals with developmental disabilities by utilizing both self-reports from the individuals and responses from their guardians, thereby enabling a comprehensive analysis of related variables. Method: Data from 463 employed individuals in the “average reading” group were drawn from the 2023 Survey on Work and Life of Persons with Developmental Disabilities. Random forest and SHAP (Shapley Additive Explanations) analyses were conducted to examine variable importance and interaction effects, while multiple regression analysis was employed to test statistical significance. Results: A total of 20 key variables were identified?9 from the self-reports of individuals (e.g., relationships with coworkers, self-esteem) and 11 from guardians’ responses (e.g., overall health status, perceived satisfaction of the individuals). SHAP analysis visualized the relative contribution and directional influence of these variables, and 9 variables showed statistical significance in the regression analysis. Conclusion: By jointly analyzing responses from both individuals and their guardians, this study provides a multidimensional understanding of life satisfaction among people with developmental disabilities. It highlights the necessity for policy development that incorporates both individual-centered interventions and guardian-related factors.

17

랜덤 포레스트 기반 분위수 회귀와 신뢰구간 보정을 통한 셰일저류층 궁극가채량 예측의 불확실성 정량화

기세일, 심재헌, 장일식

[NRF 연계] 한국자원공학회 한국자원공학회지 Vol.62 No.4 2025.08 pp.400-413

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

이 연구에서는 Eagle Ford Shale 지역의 셰일 오일 생산정을 대상으로 랜덤 포레스트 분위수 회귀기반의 궁극가채량 예측 및 불확실성 정량화를 수행하였다. 랜덤 포레스트 분위수 회귀로부터 도출된 P90 및 P10 방식은 예측구간 포함확률(PICP)이 80%를 과도하게 초과하여 불확실성 분석에적합하지 않은 한계를 보였다. 이를 개선하기 위해 수정된 적합분위수 회귀법을 통한 신뢰구간 보정 기법을 적용하였다. 보정된 모델은 PICP 80%에 해당하는 보정된 P90 및 P10 값을 도출함으로써 예측 신뢰구간의 적정성을 확보하였다. 또한 중앙값 예측은 보정 전후 변화가 거의 없어 모델의구조적 안정성이 유지됨을 확인하였다. 제안된 방법은 셰일저류층 생산 예측의 불확실성 정량화에 대한 신뢰성과 실용성을 향상시킬 수 있는 효과적인 대안이 될 수 있다.

This study uses quantile regression forest (QRF), which is based on random forest (RF), to estimate the estimated ultimate recovery of shale oil in the Eagle Ford Shale and quantify uncertainty. Conventional prediction intervals using the P90 and P10 quantiles from the RF quantile model yielded higher prediction interval coverage probability (PICP) than the target of 80%, leading to uncertainty overestimation. A calibration method based on modified conformalized quantile regression (CQR) was used for the prediction intervals to solve this problem. By generating adjusted P90 and P10 values, the calibrated model ensured a PICP of 80% and a more appropriate uncertainty representation. The median prediction remained nearly unchanged before and after calibration, confirming the structural stability of the model. The proposed method enhances the reliability and practical applicability of uncertainty quantification in forecasting production from shale reservoirs.

18

랜덤 포레스트를 활용한 장애인 취업 예측 모형: 취업 여부 및 정규직 취업을 중심으로

이기정, 김영식

[NRF 연계] 한국장애인고용공단 고용개발원 장애와 고용 Vol.29 No.3 2019.08 pp.145-165

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

연구목적: 본 연구는 최근 새롭게 주목받고 있는 머신 러닝 기법을 적용하여 장애인의 취업 여부와 정규직 취업에 영향을 미치는 변수들을 탐색적으로 규명하고자 하였다. 연구방법: 장애인고용패널 2차 웨이브의 2차 연도(2017년) 자료를 대상으로 대표적인 머신러닝 기법 중 하나인 랜덤 포레스트(Random Forest)를 통하여 179개 설명변수를 활용한 장애인 취업 예측 모형의 성능을 확인하고, 장애인의 취업에 영향을 미치는 요인들을 탐색적으로 추정해보았다. 연구결과: 분석 결과, 기존의 직장 경험 및 기대 임금 수준, 가계의 경제적 수준 등이 장애인의 취업에 중요한 요인임을 확인하였다. 그리고 장애인의 정규직 취업에는 직업 능력의 중요도가 상대적으로 높게 나타남을 확인하였다. 이와 함께 기존 선행연구들과 달리 주관적 인식 변수들이 장애인들의 취업 여부와 정규직 취업에 중요한 영향을 미치는 것 또한 확인하였다. 결론: 향후의 장애인 고용 지원 관련 연구 및 정책 입안에 있어 머신 러닝 기법의 활용 가능성을 확인하였다는 점에서 연구의 의의를 찾을 수 있다. 그리고 그 과정에서 연구자는 단순히 머신 러닝의 결과물을 수동적으로 받아들이는 객체가 아니라 연구 과정을 주도해 나가는 주체로서 기능할 수 있도록 연구자 및 정책 입안자들의 지속적인 관심과 노력이 필요함을 제안하였다.

Objectives: The purpose of this paper is to exploratively investigate predictive variables on the employment status and quality of employment of persons with disabilities. Method: This study utilized the random forest method, one of the most representative machine learning approach. Results: Our results first show that previous working experience, expect wage, and economic status of family are the substantial variables in predicting the employment status of persons with disabilities. In terms of the quality of employment, vocational capacity of the disadvantaged has a more predictive performance relatively. Moreover, this study identified subjective variables, such as quality of life and conception regarding employment, have an important predictive power on the employment status and quality of employment of persons with disabilities. Conclusion: The findings of this study, based on new approach using machine learning method, provide useful policy implications for improving the employment status and quality of employment of persons with disabilities.

19

랜덤 포레스트를 활용한 고령자의 고령친화캠퍼스 프로그램 참여의향 예측요인 탐색

남라라, 임진섭

[NRF 연계] 한국사회복지교육협의회 한국사회복지교육 Vol.73 No.2 2026.06 pp.67-94

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 지역 고령자의 고령친화캠퍼스(Age-Friendly Campus) 프로그램 참여의향에 영향을 미치는 예측요인을 탐색하고자 하였다. 이를 위해 농어촌과 중소도시에 거주하는 고령자 230명의 설문자료를 활용하였으며 인구사회학적 특성, 심리사회적 특성, 디지털활용능력 등 21개 변수를 분석에 투입하였다. 분석은 두 단계로 진행하였다. 1단계에서는 랜덤 포레스트(Random Forest)를 활용하여 변수의 상대적 중요도를 탐색하였고 2단계에서는 위계적 로지스틱 회귀분석을 통해 개별 변수의 통계적 유의성을 검토하였다. 분석 결과 5-Fold 교차검증 AUC는 .744로 나타났고 최종 모형의 분류정확도는 71.7%였다. 랜덤 포레스트 변수중요도에서는 디지털활용능력이 가장 높은 기여도를 보였으며 그 뒤를 사회활동참여, 연령, 삶의만족도, 자아통제감이 뒤를 이었다. 반면 위계적 로지스틱 회귀분석에서는 연령, 거주지역, 장애유무가 통계적으로 유의한 예측 변수로 확인되었다. 연령은 모든 모형에서 일관되게 부적 예측 변수로 나타났으며 중소도시 거주 고령자는 농어촌 거주 고령자보다 참여의향 오즈가 높았다. 디지털활용능력은 랜덤 포레스트에서는 가장 높은 변수중요도를 보였으나 로지스틱 회귀분석에서는 통계적으로 유의하지 않았다. 이는 랜덤 포레스트가 비선형적 분기 구조를 포착하는 데 강점을 가지는 반면 로지스틱 회귀분석은 다른 변수를 통제한 상태에서 개별 변수의 독립적 선형 효과를 검증한다는 점에서 비롯된 결과로 해석된다. 본 연구는 고령친화캠퍼스 프로그램 참여의향을 예측하는 데 있어 연령, 거주지역, 장애유무가 주요한 통계적 예측 변수임을 확인하였고 디지털활용능력과 사회활동참여 등은 탐색적 수준에서 중요한 예측 단서가 될 수 있음을 보여주었다. 이러한 결과는 지역대학이 고령친화캠퍼스 프로그램을 설계할 때 농어촌 고령자의 접근성, 연령대별 참여 전략, 장애 고령자의 학습 환경, 디지털 역량 지원을 함께 고려할 필요가 있음을 시사한다.

This study aimed to explore the predictors influencing older adults’ intention to participate in Age-Friendly Campus programs. To this end, survey data were collected from 230 older adults residing in rural areas and small-sized cities. A total of 21 variables, including sociodemographic characteristics, psychosocial characteristics, and digital literacy, were included in the analysis. The analysis was conducted in two stages. In the first stage, Random Forest analysis was used to examine the relative importance of the variables. In the second stage, hierarchical logistic regression analysis was conducted to test the statistical significance of individual variables. The results showed that the 5-fold cross-validation AUC was .744, and the classification accuracy of the final model was 71.7%. In the Random Forest variable importance analysis, digital literacy showed the highest contribution, followed by social participation, age, life satisfaction, and self-mastery. In contrast, the hierarchical logistic regression analysis identified age, residential area, and disability status as statistically significant predictors. Age consistently showed a negative association across all models, and older adults residing in small-sized cities showed higher odds of participation intention than those residing in rural areas. Digital literacy showed the highest variable importance in the Random Forest analysis but was not statistically significant in the logistic regression analysis. This result can be interpreted as reflecting the methodological difference between the two approaches: Random Forest is well suited to capturing nonlinear branching structures, whereas logistic regression tests the independent linear effect of each variable after controlling for other variables. This study confirmed that age, residential area, and disability status are major statistical predictors of older adults’ intention to participate in Age-Friendly Campus programs, while digital literacy and social participation may serve as important exploratory predictive cues. These findings suggest that local universities should consider rural older adults’ accessibility, age-specific participation strategies, disability- friendly learning environments, and digital literacy support when designing Age-Friendly Campus programs.

20

랜덤포레스트 머신러닝 기법을 활용한 전통적 비행이론기반 청소년 온․오프라인 비행 예측요인 연구

이택호, 김선영, 한윤선

[NRF 연계] 한국문화및사회문제심리학회 한국심리학회지: 문화 및 사회문제 Vol.28 No.4 2022.11 pp.661-690

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구에서는 청소년 비행이 지속적인 사회문제로 대두됨에 따라 청소년의 온․오프라인 비행을 예측하는 주요 요인들을 탐색하고 전통적 비행이론(사회학습이론, 일반긴장이론, 사회통제이론, 일상활동이론, 낙인이론)의 적용 가능성을 살펴보았다. 분석에 활용된 데이터는 한국아동․청소년패널조사 2010(KCYPS 2010)의 초1, 초4, 중1 패널 6차년도 데이터이다(N=4,137). 예측 모형을 구축함에 있어 전통적 통계기반의 회귀모형 대신 랜덤포레스트 머신러닝 기법을 활용함으로써 예측 성능 향상과 더불어 보다 많은 예측요인의 고려 가능성에 초점을 두었다. 랜덤포레스트 분석 결과, 청소년의 온․오프라인 비행을 설명하는 데에 전통적인 비행이론은 여전히 유효하였으며, 온라인 비행은 주로 개인적 요인(일상활동이론, 낙인이론)과, 오프라인 비행은 사회적 요인(사회학습이론, 사회통제이론)과 관련이 있는 것으로 나타났다. 또한 일반긴장이론은 온라인 비행과 오프라인 비행 모두를 예측하는 중요한 이론적 기반임을 확인할 수 있었다. 본 연구는 머신러닝 기법을 통해 청소년 비행에 영향을 주는 주요 요인을 도출하고, 기존 비행이론의 활용 가능성도 함께 고려했다는 점에서 의의가 있으며, 청소년 온․오프라인비행에 대한 예방 및 개입 방향성을 재고하는 기반을 제공할 것이라 기대된다.

Adolescent delinquency is a substantial social problem that occurs in both offline and online domains. The current study utilized random forest algorithms to identify predictors of adolescents’ online and offline delinquency. Further, we explored the applicability of classic delinquency theories (social learning, strain, social control, routine activities, and labeling theory). We used the first-grade and fourth-grade elementary school panels as well as the first-grade middle school panel (N=4,137) among the sixth wave of the nationally-representative Korean Children and Youth Panel Survey 2010 for analysis. Random forest algorithms were used instead of the conventional regression analysis to improve the predictive performance of the model and possibly consider many predictors in the model. Random forest algorithm results showed that classic delinquency theories designed to explain offline delinquency were also applicable to online delinquency. Specifically, salient predictors of online delinquency were closely related to individual factors(routine activities and labeling theory). Social factors(social control and social learning theory) were particularly important for understanding offline delinquency. General strain theory was the commonly important theoretical framework that predicted both offline and online delinquency. Findings may provide evidence for more tailored prevention and intervention strategies against offline and online adolescent delinquency.

 
1 2 3 4 5
페이지 저장