년 - 년
Random Forest를 활용한 고속도로 교통사고 심각도 비교분석에 관한 연구 KCI 등재
한국재난정보학회 한국재난정보학회논문집 제20권 1호 통권63호 2024.03 pp.156-168
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
연구목적: 고속도로 교통사고의 추세는 증감을 반복하며 도로 종류 중 고속도로에서의 치사율은 최고치 를 나타내고 있다. 따라서 국내 실정을 반영한 개선대책 수립이 필요하다. 연구방법: Random Forest를 활 용해 2019년부터 2021년까지 전국 고속도로 노선 중 사고 다발 10개 노선에서 발생한 교통사고 자료로 사고 심각도 분석 및 사고 심각도에 미치는 영향요인을 도출하였다. 연구결과: SHAP 패키지를 활용해 상 위 10개의 변수 중요도를 분석한 결과, 고속도로 교통사고 중 사고 심각도에 높은 영향을 미치는 변수는 가해자 연령이 20세 이상 39세 미만, 시간대가 주간(06:00-18:00), 주말(토~일), 계절이 여름과 겨울, 법규 위반이 안전운전불이행, 도로 형태가 터널, 기하구조상 차로 수가 많고 제한속도가 높은 경우로 총 10개 의 독립변수에서 고속도로 교통사고 심각도와 양(+)의 상관관계를 가지는 것으로 분석되었다. 결론: 고속 도로에서의 사고 발생은 매우 다양한 요인의 복합적인 작용으로 인해 발생하므로 사고 예측에 많은 어려 움이 있지만 본 연구로 도출된 결과를 활용해 고속도로 교통사고 심각도에 영향을 주는 요인을 심층적으 로 분석해 효율적이고 합리적인 대응책 수립을 위한 노력이 필요하다.
Purpose: The trend of highway traffic accidents shows a repeating pattern of increase and decrease, with the fatality rate being highest on highways among all road types. Therefore, there is a need to establish improvement measures that reflect the situation within the country. Method: We conducted accident severity analysis using Random Forest on data from accidents occurring on 10 specific routes with high accident rates among national highways from 2019 to 2021. Factors influencing accident severity were identified. Result: The analysis, conducted using the SHAP package to determine the top 10 variable importance, revealed that among highway traffic accidents, the variables with a significant impact on accident severity are the age of the perpetrator being between 20 and less than 39 years, the time period being daytime (06:00-18:00), occurrence on weekends (Sat-Sun), seasons being summer and winter, violation of traffic regulations (failure to comply with safe driving), road type being a tunnel, geometric structure having a high number of lanes and a high speed limit. We identified a total of 10 independent variables that showed a positive correlation with highway traffic accident severity. Conclusion: As accidents on highways occur due to the complex interaction of various factors, predicting accidents poses significant challenges. However, utilizing the results obtained from this study, there is a need for in-depth analysis of the factors influencing the severity of highway traffic accidents. Efforts should be made to establish efficient and rational response measures based on the findings of this research.
팬데믹 전후 항공기 기종에 따른 기재 활용 전략의 변화 : Random Forest와 SHAP 분석을 중심으로
한국정보기술응용학회 JITAM Vol.32 No.2 2025.04 pp.83-95
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
This study investigates the changes in factors influencing aircraft utilization based on aircraft types during the pandemic period. The analysis categorizes the periods into three phases: pre-pandemic, during the pandemic, and post-pandemic, while the aircraft types are classified into two groups (C-class and E-class or higher). Utilizing SHAP and Random Forest methodologies, the study analyzes the importance of various factors for each group and examines the significance of differences in importance across the specified periods and aircraft types through group difference testing. The findings reveal that for C-class aircraft, the importance of the USD/KRW exchange rate and the lagging composite index was notably high before the pandemic. However, during the pandemic, there was a substantial surge in the significance of the JPY/KRW exchange rate, and post-pandemic, the importance of economic indicators increased. E-class or higher aircraft exhibited a similar trend, with a significant rise in the importance of the CNY/KRW exchange rate observed after the pandemic. The Kruskal-Wallis test confirmed that most average changes among the groups were not statistically significant, except for the C-class aircraft in the Random Forest analysis. This indicates that there were no substantial changes in the underlying factors between the pre-pandemic and post-pandemic periods, suggesting that the impact of the pandemic on airline fleet utilization extends beyond short-term changes. Furthermore, the findings highlight the need for airlines to enhance their sensitivity to external environments and develop utilization strategies in volatile markets. Additionally, there is a possibility that long-term dependency on network strategies may become more pronounced than reliance on external factors like the pandemic. By providing an in-depth analysis of the pandemic's impact on aircraft utilization, this study contributes to the formulation of effective utilization strategies for airlines navigating an uncertain economic environment.
GRU 분석 모델을 활용한 비트코인 시장 예측 : LSTM과 Random Forest 모델 비교 분석
한국정보기술응용학회 JITAM Vol.32 No.1 2025.02 pp.1-13
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
This study conducts a comparative analysis of various models for predicting Bitcoin prices. The models utilized in this research include the Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), and Random Forest. Their performance was evaluated using key metrics such as Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R-squared (R²). The findings reveal that the GRU model outperformed the other two models, achieving an R² value of 0.9301, which indicates its superior ability to explain the data. The LSTM model demonstrated relatively high explanatory power with an R² of 0.8252, yet it exhibited lower predictive accuracy compared to the GRU. On the other hand, the Random Forest model recorded the lowest MAE (1557.44) and RMSE (1766.95), reflecting strong short-term prediction capabilities. However, its R² of 0.6110 highlighted limitations in capturing the intricate temporal patterns of the dataset. The significance of this study lies in several aspects. First, it empirically demonstrates that the GRU model, with its more streamlined architecture, can achieve superior predictive performance compared to the LSTM and Rando Forest models. Additionally, while Random Forest showed promising results for short-term predictions, it underscored limitations in modeling the complexities of temporal data patterns. Future research should aim to incorporate broader datasets and integrate external economic variables to enhance predictive accuracy. Furthermore, the potential of ensemble approaches, combining diverse models, holds promise for further improving predictive performance in this domain.
강우·소방·사회취약성 데이터를 활용한 복합재난 위험도 예측 및 소방 자원 배치 전략 : 서초구 사례의 MRA–Random Forest 비교 분석 KCI 등재
한국재난정보학회 한국재난정보학회논문집 제21권 4호 통권70호 2025.12 pp.1169-1178
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
연구목적: 본 연구는 서울특별시 서초구를 대상으로 강우, 소방 활동, 사회 취약성 데이터를 통합하여 복합재난 위험 도를 예측하고, 이를 기반으로 소방 자원 배치 전략을 제시하는 것을 목적으로 한다. 연구자료는 2020–2024년 서초 구 시간별 강우량 자료, 화재 출동 개별 내역, 119 구조·구급활동 실적, 통계 연보에서 추출한 인구·주거·사회 취약성 지표를 사용하였다. 특히 신고일시·출동일시·도착일시를 활용하여 출동 지연, 도착시간, 총 대응 시간을 산출하고, 65세 이상 인구 비율, 장애인 등록 인구 비율, 1인 가구 비율, 반지하·지하 주택 비율을 사회 취약성 변수로 구성하였 다. 연구방법: 분석 방법으로는 먼저 화재 출동 내역을 이용해 연도·월·시간대별 출동 건수와 평균 도착시간, 평균 총 대응 시간을 산출하여 기초 통계를 제시하고, 강우량 및 사회 취약성 지표와 결합하여 복합재난 발생에 영향을 미치 는 설명변수를 구성하였다. 이후 다중회귀분석(Multiple Regression Analysis, MRA)과 Random Forest(RF)를 이용하 여 시간 단위 화재 출동 건수 및 발생 여부를 예측하고, 두 모형의 성능을 비교하였다. 2024년을 검증자료로 사용한 결과, 시간당 화재 출동 건수 예측에서 MRA 모형의 평균절대오차(MAE)는 0.24건, RF 모형의 MAE는 0.18건으로, RF가 약 25.0% 낮은 예측 오차를 보였다. 화재 발생 여부(해당 시간에 1건 이상 발생)의 분류 성능을 AUC로 비교한 결과, 로지스틱 회귀는 0.83, RF 분류기는 0.90을 기록하여 RF가 우수한 성능을 나타냈다. 연구결과: RF 변수 중요도 분석 결과, 이전 24시간 누적 강우량, 현 시각 강우량, 신고 시간의 시간대(야간 여부), 65세 이상 인구 비율, 장애인 등록 인구 비율, 반지하·지하 주택 비율, 1인 가구 비율순으로 중요도가 높게 나타났다. 이는 강우 특성뿐 아니라 고령자 및 장애인 밀집 지역, 주거 취약 지역, 1인 가구 밀집 지역이 복합재난 위험에 유의 미한 영향을 미침을 시사한다. 마지막으로, 예측된 위험도와 사회 취약성 지수를 결합하여 행정동별 우선 대응 순위를 제시하고, 이를 기반으로 소방 자원 배치 시나리오를 제안하였다. 결론: 본 연구는 기상·소방·사회통계 데이터를 결합한 복합재난 위험도 예측과 사회 취약성을 고려한 자원 배치 전략을 함께 제시함 으로써, 도시형 재난관리에 대한 데이터 기반 의사결정 지원에 기여할 수 있을 것으로 기대된다.
Purpose: This study aims to predict the risk of compound disasters in Seocho District, Seoul, by integrating rainfall, firefighting activity, and social vulnerability data, and to propose strategies for the optimal allocation of firefighting resources. Method: The research employs hourly rainfall data from 2020 to 2024, detailed fire dispatch records, 119 rescue and emergency activity logs, and demographic and social vulnerability indicators extracted from statistical yearbooks, and calculates dispatch delays, arrival times, and total response times using reported, dispatch, and arrival timestamps. Social vulnerability variables include the proportion of residents aged 65 or older, the proportion of registered persons with disabilities, the share of single-person households, and the percentage of semi-basement or basement dwellings. Result: Multiple Regression Analysis and Random Forest models are applied to predict hourly fire dispatch counts and fire occurrence, and their predictive performances are compared using 2024 data as a validation set; results show that the Random Forest model achieves a lower mean absolute error than the regression model in count prediction and a higher AUC in occurrence classification.Variable importance analysis further indicates that cumulative rainfall over the previous 24 hours, current rainfall, time of report (day/night), and the proportions of elderly residents, registered disabled residents, semi-basement or basement dwellings, and single-person households are major contributors to predicted compound disaster risk. Conclusion: Overall, the study concludes that integrating meteorological, firefighting, and socio-demographic data enables more accurate compound disaster risk prediction and supports socially responsive resource allocation strategies, thereby contributing to data-driven decision-making in urban disaster management.
한국ITS학회 한국ITS학회 학술대회 ITS와 함께하는 미래 스마트 시티 2022.06 pp.887-890
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
In this study, AAHT was estimated using machine learning algorithms of ensemble techniques to improve the limitations of existing studies. Among the machine learning algorithms, a random forest algorithm was used, and the traffic volume estimation accuracy was analyzed as MAPE 20.7%. It was analyzed that the higher the traffic volume level, the higher the traffic volume estimation accuracy, and the traffic volume level of 3,000 vehicles/hour or more was analyzed to be 8% or less of MAPE. It was analyzed that the accuracy secured at the current level was very high compared to the existing model.
교통사고는 매년 수백만명의 사망자와 수천만명의 부상자를 일으키는 사회문제로 자리잡았다. 교통사고는 복합적인 원인 을 가지고 있으며 원인해결이 어려운 사회적 문제이다. 이를 과거의 사고를 기반으로 미래의 사고를 예측하는 머신러닝 모델을 제시하고자 한다. Random Forest 모델을 사용하였으며 MSE로 모델 평가를 진행하였다. 또한, 이 결과값을 통해 사회적, 경제 적인 손실을 줄이고 안전을 증진시키는 것을 목적으로 한다.
랜덤포레스트 모델을 활용한 청년층 차입자의 채무 불이행 위험 연구 KCI 등재
한국가족자원경영학회 가족자원경영과 정책(구 한국가족자원경영학회지) 제26권 3호 2022.08 pp.19-34
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
청년층 및 저소득층을 포함한 취약계층과 제2금융권을 중심으로 한 부채 불이행에 대한 우려가 증가하고 있다. 청년층의 가계부채 건전성은 최근 고용 부진, 학자금대출 부담 증가, 제2금융권에서 고금리 대출 증가 등이 복합적으로 작용하여 더욱 취약해졌다. 본 연구의 목적은 한국의 청년층 차입자를 대상으로 채무 불이행 가능성을 진단하고, 그 가능성에 영향을 주는 요인을 예측하는 것이다. 이러한 목적을 달성하기 위하여 본 연구는 2021년 「가계금융·복지조사」를 활용하고, 청년층 의 채무 불이행 가능성과 관련된 요인들을 포괄적으로 분석하기 위하여 머신러닝 알고리즘의 랜덤포레스트 방법을 적용하였 다. 청년층 차입자의 채무 불이행 위험을 예측하는 모형을 탐색한 뒤 중요도 지수를 산출하고, 중요도가 높은 설명변수들을 선별한 뒤, 주요 결정요인들의 부분 의존성 도표를 제시하고자 하였다. 최종적으로 자산대비부채비율(DTA), 의료비 비중, 가계부실위험지수(HDRI), 통신비 비중, 주거비 비중이 주요한 변인으로 나타났다.
There are growing concerns about debt insolvency among youth and low-income households. The deterioration in household debt quality among young people is due to a combination of sluggish employment, an increase in student loan burden and an increase in high-interest loans from the secondary financial sector. The purpose of this study was to explore the possibility of household debt default among young borrowers in Korea and to predict the factors affecting this possibility. This study utilized the 2021 Household Finance and Welfare Survey and used random forest algorithm to comprehensively analyze factors related to the possibility of default risk among young adults. This study presented the importance index and partial dependence charts of major determinants. This study found that the ratio of debt to assets(DTA), medical costs, household default risk index (HDRI), communication costs, and housing costs the focal independent variables.
랜덤 포레스트를 활용한 건설재해 예측모델 KCI 등재
대한건축학회지회연합회 대한건축학회연합논문집 제22권 제5호 통권 99호 2020.10 pp.81-88
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
The construction industry has seen a continuous increase in the number of victims over the last five years, despite various efforts and investments in disaster prevention. This means that we need substantial and fundamental preventive measures that are different from the ones we have now. Therefore, this study proposed a prediction model that can predict construction disasters in advance using a highly predictable random forest method. The random forest technique was used to analyze the impact factors of the construction disaster, and the model was constructed by selecting the factors of high importance from the variables. As a result of comparing the accuracy of the random forest model and the discrimination analysis model in order to verify the constructed construction disaster prediction model, it is found that the random forest model is more accurate than the discrimination analysis model. Accordingly, it is expected that disaster types that can occur by the prediction model can be accurately predicted in advance and preferentially managed to contribute to disaster prevention.
랜덤 포레스트를 활용한 저경력 교사의 이직 의도 예측 요인 분석 KCI 등재
한국교원교육학회 한국교원교육연구 제43권 제1호 2026.03 pp.1-28
※ 기관로그인 시 무료 이용이 가능합니다.
6,700원
본 연구는 랜덤 포레스트를 활용하여 저경력 교사의 이직 의도를 예측하는 요인을 규명하는 데 목적이 있다. 분석 자료는 서울교원종단연구 2020(Seoul Education Longitudinal Study of Teachers, SELST) 3차년도(2023년) 자료를 활용하였으며, 저경력 교사 634명을 연구 대상으로 선정하였다. 선행연구를 토대로 변수를 개인적 특성, 전문성〮·역량 특성, 직무 및 심리 특성, 조직 환경 특성으로 구분하여 총 77개 예측 변인을 투입하였다. 예측 변인의 중요도 분석 결과, 직무 및 심리 특성 범주에서 소진, 학교 소속감, 개인 발전 가능성이 높은 중요도를 보였고, 전문성〮·역량 특성으로는 디지털 리터러시와 교육과정 재구성 역량, 조직 환경 특성으로 조직 공정성 등이 상위 중요도를 나타냈다. 본 연구는 저경력 교사의 이직 의도를 종합적으로 예측하기 위해 다양한 요인을 제시하며, 학교 소속감 증진, 조직 공정성 인식 제고 등 이론적 및 정책적 시사점을 논의하였다.
This study aims to exploratorily identify the determinants of turnover intention among early career teachers in South Korea. Data from the 3rd wave (2023) of the Seoul Education Longitudinal Study of Teachers (SELST), focusing on a final sample of 634 early career teachers. Drawing on previous studies, 77 predictor variables were categorized into four groups: individual characteristics, professional expertise and competency characteristics, job-related and psychological characteristics, and organizational environment characteristics. Variable importance analysis indicated that factors within the job-related and psychological characteristics category were the most significant in predicting turnover intention. The results showed that, burnout, school belonging, and perceived potential for personal growth emerged as the top predictors. In the realm of professional expertise and competencies, digital literacy and curriculum reconstruction competence were identified as important predictors, while perceived organizational justice was a key factor within the organizational environment category. Based on the findings, theoretical and policy implications, including strategies to enhance school belonging, foster professional growth, and strengthen perceptions of organizational justice to reduce early career teachers’ turnover intention.
랜덤 포레스트를 활용한 도로 및 교통시설 개선방향 추정 연구 KCI 등재
한국ITS학회 한국ITS학회논문지 제20권 제6호 통권98호 2021.12 pp.37-46
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
교통사고 예방을 위해 경찰 및 지자체 등 정부기관에서는 교통시설 및 도로시설의 개선사 업을 추진하여 교통 위해 요소를 제거하고 편안한 도로 환경을 조성하는데 노력하고 있다. 이 를 위해 도로 및 교통시설을 개선 및 조정하며, 교통사고 잦은 지역의 개선사업이 대표적인 사업이다. 교통사고 잦은 지역의 개선사업은 담당자와 관계자의 주관에 따라 사업별, 지역별 편차가 발생하고 있으며, 우선순위 도출 등에 민원 및 주관성이 반영되어 사업의 효율성에 한 계가 발생하고 있다. 이를 위해 교통사고 잦은 곳 개선사업의 효과가 높은 대표사업을 대상으 로 도로여건, 교통여건, 사고여건 등을 종합적으로 고려하여 사업 대상지의 개선방향을 추정하 는 연구를 진행하였다. 연구결과 개선사업 추정 정확도가 88% 수준으로 분석되었으며, 개선방 향을 추정하는데 교통량, 사고율, 사고심각도 순으로 높은 관계가 있는 것으로 분석되었다.
Government agencies, such as police and local governments, strive to prevent traffic hazards and create a comfortable road environment by pormoting transportation and road facilities. To this end, roads and transportation facilities are enhanced and adjusted, and improvement projects in areas with frequent traffic accidents are carried out. Usually, improvement projects in areas with frequent traffic accidents vary by projects and region. Moreover, these projects are carried out under the supervision of a person in charge and related parties. Hence, civil complaints and subjectivity are reflected in deriving priorities for the improvement projects, limiting the efficiency of the project. To this end, a study was conducted to estimate the direction of improvement of the project target site. This study comprehensively considered road, traffic, and accident conditions of representative projects with high effectiveness in handling traffic accidents. The results of the study state that the accuracy of estimating the improvement project was around 88%. In addition, the study found that there was a strong relationship between traffic volume, accident rate, and accident severity in estimating the improvement direction.
PCL-R의 재범예측요인에 관한 연구 : 의사결정나무 분석과 랜덤포레스트 기법을 중심으로 KCI 등재
한국경찰연구학회 한국경찰연구 제20권 제2호 2021.06 pp.163-190
※ 기관로그인 시 무료 이용이 가능합니다.
6,700원
PCL-R은 1980년 Robert Hare 박사가 만든 Psychopath Checklist로 현재까지도 대 중적으로 쓰이면서 널리 알려진 재범위험성평가 도구 중 하나이다. 특히나 사이코패스 를 잘 구별해내는 것으로 알려진 이 검사의 경우 재범위험성평가들 중 평가 대상자의 양적, 질적 정보를 매뉴얼을 통한 면담으로 훌륭히 잘 수집해내는 것을 수많은 연구를 통해 입증하였다. 그러나 PCL-R 검사지에 속해 있는 20개의 문항들 중 어떠한 문항이 재범에 영향을 주는 가에 대한 구체적인 요인을 알아내는 것에는 한계가 있다. 본 연구 에서는 이러한 점을 알아보고자 ‘의사결정나무’와 ‘랜덤포레스트’ 라는 머신러닝 기반 분류 분석을 실시하고자 한다.
PCL-R is a Psychopath Checklist created by Dr. Robert Hare in 1980, and is one of the widely used recidivism risk assessment tools. In particular, in the case of this test, which is known to discriminate psychopaths well, it has been proven through numerous studies that it collects quantitative and qualitative information of the subject in recidivism risk assessments through manual interviews. However, there is a limit to ascertaining specific factors of which of the 20 items in the PCL-R test paper affect recidivism. To investigate this point, this study intends to conduct a machine learning-based classification analysis called ‘decision tree’ and ‘random forest’.
다문화 가정 어머니의 문화적응 유형과 분류 예측요인 탐색연구 - 잠재프로파일과 랜덤포레스트를 활용하여 - KCI 등재
인하대학교 다문화융합연구소 문화교류와 다문화교육 제10권 제2호 2021.03 pp.171-193
※ 기관로그인 시 무료 이용이 가능합니다.
6,000원
본 연구는 다문화 가정 어머니의 문화적응 유형을 살펴보고, 분류에 영향을 미치는 예측요인들을 탐색하는데 목적이 있다. 이를 위해 전국 단위 조사인 다문화청소년패널조사 2018년 8차년도 자료의 다문화 가정 어머니 1,625명을 대상으로 잠재프로파일과 랜덤포레스트를 활용하여 분석하였다. 주요 결과는 다음과 같다. 첫째, 다문화 가정 어머니의 문화적응은 동화/통합형, 분리형, 소외형 등 3개 유형으로 분류되었다. 둘째, 분류에 영향을 미치는 예측요인으로는 문화적응스트레스, 부모효능감, 자아존중감, 한국어 수준 순으로 나타났다. 이를 통해 다문화 가정 어머니가 한국사회에서 긍정적으로 적응하기 위한 제언을 하였다.
This study is to examine the acculturation profiles of mothers in multicultural families and to explore the predictors affecting on classification. For the analysis, 1,625 mothers in multicultural families were analyzed using data from the 8th waves of the Multicultural Youth Panel Survey 2018, which is a national-wide survey. A latent profile analysis and random forest were used for the analysis. The main results are as follows. First, They could be classified into three types: assimilation/integration, separation, and marginalization. Second, as predictors affecting classification, acculturative stress, parental efficacy, self-esteem, and Korean language level were in order. Based on these analysis results, suggestions were made for positive adaptation of mothers in multicultural families in Korean society.
다문화 가정 어머니의 문화적응 유형과 분류 예측요인 탐색연구 - 잠재프로파일과 랜덤포레스트를 활용하여 - KCI 등재
한국국제문화교류학회 문화교류와 다문화교육(구 문화교류연구) 제10권 제2호 2021.03 pp.171-193
※ 기관로그인 시 무료 이용이 가능합니다.
6,000원
본 연구는 다문화 가정 어머니의 문화적응 유형을 살펴보고, 분류에 영향을 미치는 예측요인들을 탐색하는데 목적이 있다. 이를 위해 전국 단위 조사인 다문화청소년패널조사 2018년 8차년도 자료의 다문화 가정 어머니 1,625명을 대상으로 잠재프로파일과 랜덤포레스트를 활용하여 분석하였다. 주요 결과는 다음과 같다. 첫째, 다문화 가정 어머니의 문화적응은 동화/통합형, 분리형, 소외형 등 3개 유형으로 분류되었다. 둘째, 분류에 영향을 미치는 예측요인으로는 문화적응스트레스, 부모효능감, 자아존중감, 한국어 수준 순으로 나타났다. 이를 통해 다문화 가정 어머니가 한국사회에서 긍정적으로 적응하기 위한 제언을 하였다.
This study is to examine the acculturation profiles of mothers in multicultural families and to explore the predictors affecting on classification. For the analysis, 1,625 mothers in multicultural families were analyzed using data from the 8th waves of the Multicultural Youth Panel Survey 2018, which is a national-wide survey. A latent profile analysis and random forest were used for the analysis. The main results are as follows. First, They could be classified into three types: assimilation/integration, separation, and marginalization. Second, as predictors affecting classification, acculturative stress, parental efficacy, self-esteem, and Korean language level were in order. Based on these analysis results, suggestions were made for positive adaptation of mothers in multicultural families in Korean society.
서울시 열섬현상 완화를 위한 바람길 및 녹지 입지 선정 제안 : 클러스터링 및 랜덤포레스트 변수 중요도를 활용한 열섬척도 제안을 중심으로
한국AI디지털융합학회(구 한국디지털융합학회) 디지털융합논문지 제6권 제1호 2019.10 pp.27-38
※ 기관로그인 시 무료 이용이 가능합니다.
4,300원
As the climate problem becomes more serious, mitigating the heat island effect is becoming an important issue. To solve the problem, it is necessary to identify the areas where the heat island effect is more severe and present appropriate measures to those areas. In this paper, we divided Seoul into 573 grids and divided into four clusters by hierarchical clustering techniques to identify the current status of the city's heat island effect. Of these cluster, we chose the one with the most urban features. We extracted the variable importance with random forest models and developed the heat island index from 1) multiplication of the variable importance and 2) the sign of the correlation coefficient between temperature and 3) the corresponding variable. Furthermore we analyzed areas where the heat island effect is severe with the index. According to the analysis, the heat island effect was particularly severe in five grids, and we suggested ways to mitigate the heat island effect for the areas concerned.
머신러닝을 활용한 도로포장 열화 예측 : 변수 선택 및 예측력 검증
한국ITS학회 한국ITS학회 학술대회 Bridging Research, Industry and Policy for Al-driven ITS 2026.04 pp.636-637
3차원 자동 측정 데이터 기반 공정 이상 탐지 예측 모델 구축
한국혁신산업학회 혁신산업기술논문지 제4권 제1호 2026.03 pp.1-9
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
제조 공정의 정밀 품질 관리에서 3차원 자동측정 데이터는 공정 상태를 정량적으로 분석할 수 있는 중요한 기반 데이터를 제공한다. 그러나 방대한 측정 데이터를 실시간으로 분석하여 공정 이상을 조기에 탐지하는 연구는 아직 제한적이다. 본 연구에서는 3차원 자동측정 데이터로부터 추출한 주요 특성을 활용하여 공정 이상을 예측하는 머신러닝 기반 모델을 구축하였다. 데이터 전처리 단계에서는 측정 데이터로부터 SPC 관리도 기반의 주요 특징을 추출하고 이상치 제거 및 정규화를 수행하였다. 모델링 단계에서는 랜덤 포레스트와 XGBoost알고리즘을 적용하여, 정상, 경고, 이상을 분류하였으며, 모델 성능은 Accuracy, Precision, Recall, F1-score로 평가하였다. 실험 결과, 랜덤 포레스트 모델과 XGBoost 모델 모두 정확도 약 99%의 높은 성능을 나타내었다. 제안한 두 모델은 3차원 자동측정 데이터를 활용한 공정 이상 예측 가능성을 보여주며, 실시간 제조지능화 시스템 구축에 효과적으로 적용될 수 있음을 나타낸다.
In precision quality control of manufacturing processes, three-dimensional automatic measurement data provide essential quantitative information for assessing process conditions. However, research on real-time analysis of large-scale measurement data for early detection of process anomalies remains limited. This paper presents a machine learning based predictive model for process anomaly detection using key features extracted from three-dimensional automatic measurement data. In the data preprocessing stage, Control chart features based on Statistical Process Control (SPC) were extracted, followed by outlier removal and normalization. In the modeling stage, Random Forest and XGBoost algorithms were applied to classify process states into normal, warning, and anomalous categories. Model performance was evaluated using accuracy, precision, recall, and F1-score metrics. According to the experimental results, both the Random Forest and XGBoost models demonstrated high predictive performance, achieving an accuracy of about 99%. The proposed model demonstrates the feasibility of real-time process anomaly prediction using three-dimensional automatic measurement data and shows strong potential for integration into intelligent manufacturing systems.
중소기업 대상 AHP와 AI 랜덤포레스트를 활용한 원격환경의 보안 전략 연구 KCI 등재
한국융합보안학회 융합보안논문지 제26권 제1호 2026.02 pp.11-17
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 분산 업무 환경에서 중소기업의 기술유출 방지를 위한 최적의 보안 전략을 제시하고자 실무자 설문 데이터를 바 탕으로 AHP(계층분석법)와 AI 랜덤 포레스트(Random Forest) 회귀 분석을 실시하였다. 실증 분석을 위해 중소기업 보안 및 전산 담당자 총 45명(정보통신, 제조업 등 10개 업종)의 응답 데이터를 활용하였다. 우선 AHP를 통해 7대 보안 대책의 투자 타당성을 분석한 결과, '문서암호화(DRM)'가 종합 가중치 0.324로 1위를 차지하여 최우선 투자 항목으로 도출되었다. 이어 AI 랜덤 포레스트를 활용하여 위협 감소 기여도를 분석한 결과, '다중인증(MFA)'(중요도 0.82)과 '시스템 모니터링 강화'(0.78)가 사고 예방의 핵심 변수임을 확인하였다. 이를 6하원칙(5W1H)에 매핑한 결과, 핵심 자산(What)은 암호화에, 실시간 운영 통제 (Who, When)는 인증과 감시에 집중하는 이원화 전략이 가장 효율적이었다. 결론적으로 중소기업은 타당성이 검증된 암호화 솔루션에 예산을 우선 배분하되, AI 기반의 지능형 인증 체계를 병행하는 맞춤형 보안 로드맵을 수립할 것을 제언한다.
This study aims to present an optimal security strategy for preventing technology leakage in Small and Medium-sized Enterprises (SMEs) within distributed work environments. To achieve this, Analytic Hierarchy Process (AHP) and AI-based Random Forest regression analysis were conducted using survey data from practitioners. For the empirical analysis, response data from 45 security and IT managers across 10 industries, including information technology and manufacturing, were utilized. First, the investment feasibility of seven major security measures was analyzed through AHP, revealing that 'Document Encryption (DRM)' ranked first with a composite weight of 0.324, identifying it as the highest priority for investment. Subsequently, an analysis of threat reduction contributions using Random Forest confirmed that 'Multi-Factor Authentication (MFA)' (importance: 0.82) and 'System Monitoring Enhancement' (importance: 0.78) are critical variables for accident prevention. Mapping these findings to the 5W1H (Six Ws) framework showed that a dual-track strategy— concentrating on encryption for core assets (What) and focusing on authentication and surveillance for real-time operational control (Who, When) —is most efficient. In conclusion, this study suggests that SMEs should prioritize budget allocation for validated encryption solutions while simultaneously establishing a customized security roadmap that incorporates AI-based intelligent authentication systems.
머신러닝 기반 MOOC 중도 탈락 예측모델 개발 KCI 등재
한국정보교육학회 정보교육학회논문지 제29권 제6호 2025.12 pp.947-955
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
MOOC의 높은 중도 탈락 문제가 대두됨에 따라서 본 연구는 이를 예측하는 머신러닝 기반의 시간 분할 검증을 적용한 예측모델을 개발하고자 하였다. 이를 위해 edX의 학습 데이터를 사용하여 시간 분할 설계(temporal split design)를 통해 2023년 학습 데이터(N=1,601)로 모델을 훈련하고, 2024년 학습 데이터(N=1,827)로 검증하였다. 모델 평가는 로지스틱 회귀 분석과 랜덤 포레스트 기법을 사용하였다. 탐색적 분석 결과, 학습 중후반기의 퀴즈 수행도 가 가장 중요한 예측 변수로 나타났으며, 데이터 분포가 다른 환경에서도 중도 이탈의 초기 징후를 효과적으로 탐 지하는 것으로 확인되었다. 각 변수의 기여도를 시각화하기 위해 SHAP 분석을 수행한 결과, 퀴즈 응답의 지속성 이 중도 이탈 예측에 가장 중요한 요소로 확인되었다. 이 결과는 검증된 예측 모델을 활용한 조기 경보 시스템이 실제 교육 현장에서 효과적으로 구현될 수 있음을 보여준다.
With the increasing concern over high dropout rates in massive open online courses (MOOCs), the present study endeavored to develop a machine learning–based time-series prediction model to forecast learner attrition. The model was trained on the 2023 dataset (N=1,601) and validated on the 2024 dataset (N=1,827) through a temporal split design, with the utilization of learning log data from edX. The performance of the model was evaluated using logistic regression and random forest techniques. Exploratory analyses revealed that quiz performance during the mid-to-late stages of the course was the most significant predictor of dropout, and the model effectively detected early signs of attrition even in environments with different data distributions. To visualize the contribution of each variable, SHAP analysis was conducted, confirming that the consistency of quiz completion was the most influential factor in predicting dropout. These results demonstrate the efficacy of implementing an early warning system based on a validated prediction model in real educational settings.
4,000원
인터넷 서비스의 확산과 함께 정교해지는 피싱(Phishing) 공격은 개인 정보 탈취 및 금융 피해를 유발하는 심각한 보안 위 협으로 대두되고 있다. 기존의 피싱 탐지 체계는 주로 구글 세이프 브라우징(Google Safe Browsing)이나 피쉬탱크 (PhishTank)와 같은 블랙리스트(Blacklist) 방식에 의존해 왔다. 이 방식은 알려진 위협에 대해서는 신속하고 정확한 차단이 가능하나, 제로데이(Zero-day) 공격을 탐지하지 못하는 치명적인 한계를 가진다. 본 연구에서는 이러한 한계를 극복하기 위해 URL의 어휘적 특징을 기반으로 하는 다양한 인공지능 모델의 탐지 성능을 비교 분석하였다. 실험 대상 모델로는 전통적인 휴 리스틱 알고리즘과 머신러닝 모델인 로지스틱 회귀(Logistic Regression), 서포트 벡터 머신(SVM), 랜덤 포레스트(Random Forest), 그리고 딥러닝 모델인 CNN(1D)과 LSTM을 선정하였다. 실험 결과, 휴리스틱 방식은 44.5%의 저조한 정확도를 보인 반면, SVM(RBF 커널) 모델은 97.0%의 정확도와 0.970의 F1-Score를 기록하며 가장 우수한 성능을 나타냈다. 특히 딥러닝 모 델인 CNN(94.5%)과 LSTM(76.1%) 대비 SVM은 0.165초라는 빠른 추론 속도를 보여 실시간 탐지 환경에서 성능과 효율성의 최적 균형을 갖춘 모델임을 입증하였다.
As internet services proliferate, phishing attacks are becoming increasingly sophisticated and are emerging as a serious security threat that causes the theft of personal information and financial damage. Existing phishing detection systems have primarily relied on blacklist methods such as Google Safe Browsing or PhishTank. While this approach enables the rapid and accurate blocking of known threats, it has a critical limitation in its inability to detect zero-day attacks. To overcome these limitations, this study comparatively analyzed the detection performance of various artificial intelligence models based on the lexical features of URLs. The models selected for the experiment included traditional heuristic algorithms, machine learning models such as Logistic Regression, Support Vector Machine(SVM), and Random Forest, as well as deep learning models like CNN(1D) and LSTM. The experimental results showed that while the heuristic method yielded a poor accuracy of 44.5%, the SVM(RBF kernel) model demonstrated the superior performance, recording an accuracy of 97.0% and an F1-Score of 0.970. In particular, compared to the deep learning models CNN(94.5%) and LSTM(76.4%), SVM demonstrated a fast inference speed of 0.165 seconds, proving it to be the model with the optimal balance between performance and efficiency in a real-time detection environment.
치매 조기 진단을 위한 음성 기반 AI 성능 비교 연구 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제9권 11호 2025.11 pp.2843-2852
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 치매 조기 진단을 위해 어르신의 음성 데이터를 Mel-Spectrogram으로 변환하여 CNN, ViT(Vision Transformer), Random Forest 등 다양한 인공지능 모델의 분류 성능을 비교하였다. 데이터는 영어 권 공개 데이터셋(ADReSS-2020 등)에서 수집하였으며, Mel-Spectrogram 이미지로 변환 후 모델에 입력하였다. 모든 모델의 정확도는 약 61~62%, F1-score는 0.60 내외로, 전반적으로 유사한 성능을 보였다. ViT 모델이 AUC 0.72로 가장 높은 분류력을 보였으나, 모든 모델에서 치매군과 정상군의 오진률이 임상적으로 유의미한 수준으로 나타났다. 연구 결과, 단일 음향 특성만으로는 치매의 조기 진단에 한계가 있음을 확인하였다. 따라서 데이터 규모 확대, 다양한 특성(언어·음향·인지정보) 융합, 딥러닝 구조의 고도화가 필요함을 시사하였다. 향후 연구에서는 한국 어 등 다양한 언어와 환경에서의 데이터수집, 실시간 진단 및 임상 적용성 검증이 추가적으로 요구된다.
This study compares the classification performance of various artificial intelligence models—CNN, ViT (Vision Transformer), and Random Forest—for early dementia diagnosis by converting elderly speech data into Mel-Spectrograms. The data were collected from publicly available English-language datasets (such as ADReSS-2020) and transformed into Mel-Spectrogram images for model input. All models achieved an accuracy of approximately 61–62%, with F1-scores around 0.60, indicating generally similar performance. The ViT model demonstrated the highest discriminative power with an AUC of 0.72, but the misclassification rates for both dementia and normal groups were clinically significant across all models. The findings confirm that relying solely on acoustic features has limitations for early dementia diagnosis. Therefore, expanding dataset size, integrating multiple features (linguistic, acoustic, cognitive), and advancing deep learning architectures are necessary. Future studies should focus on collecting data in various languages and environments, as well as verifying real-time diagnostic and clinical applicability.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.