Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 263
No
1

4,300원

본 연구는 사회심리적 영역에 해당하는 대학만족도, 구직효능감, 진로발달, 자기효능감, 스트레스 등을 대 상으로 학업중단 예측요인을 밝히고자 하였다. 분석에는 R(4.2.2)프로그램을 활용하여 랜덤포레스트 분석을 실시하 였다. 오차행렬과 ROC의 AUC를 통해 예측모형 성능을 확인하였으며 MDA, SHAP값을 도출하여 주요 예측 요 인을 분석하였다. 분석 결과 ‘미래 유망 직업에 대한 고민 여부’, ‘진로 선택 시 본인 의견 존중 정도’, ‘취업에 필 요한 정보 활용 정도’, ‘자신감’, ‘해야 할 일에 대한 실행력’, ‘적성을 고려한 취업 방법 인지 정도’, ‘장래희망 성취 를 위해 지금 할 일 인지 정도’, ‘미래에 대한 불확실함 및 불안에 의한 스트레스’, ‘대학만족도’ 등이 주요 예측 요 인으로 나타났다. 이와 같은 결과를 통해 학생들에게 필요한 역량을 ‘자기이해능력’, ‘의사결정 및 실천능력’, ‘진로 및 직업정보 활용능력’, ‘불안관리’로 제안하고 이를 지원하기 위한 대학 정책과 후속 연구를 제언하였다.

This study aimed to identify predictors of dropping out of school, targeting college satisfaction, job-seeking efficacy, career development, self-efficacy, and stress, which are social psychological domains. For the analysis, random forest analysis was performed using the R program. The performance of the prediction model was confirmed through the error matrix and AUC of the ROC, and the MDA and SHAP values were derived to analyze the main predictive factors. The results of the analysis showed ‘worries about promising future jobs’, ‘the extent to which one’s own opinion is respected when choosing a career’, ‘the extent to which information necessary for employment is utilized’, ‘self-confidence’, ‘execution ability for things to be done’, ‘the extent to which one is aware of employment methods that take aptitude into account’, ‘the extent to which one is aware of what to do now to achieve future aspirations’, ‘stress due to uncertainty and anxiety about the future’, and ‘university satisfaction’. Based on these results, the competencies required for students were suggested as ‘self-understanding ability’, ‘decision-making and practical ability’, ‘the ability to utilize career and job information’, and ‘anxiety management’, and university policies and follow-up research were suggested to support them.

2

랜덤포레스트를 이용한 낙동강 본류의 남조류 발생 영향인자 분석 KCI 등재

정우석, 김성은, 김영도

한국습지학회 한국습지학회지 제23권 제1호 2021.02 pp.27-34

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 연구에서는 랜덤포레스트를 이용하여 8개 보 지점별 남조류 발생 주요 영향인자를 도출하고, 조류경보제 기반의 범주형 예측모델을 개발하였다. 8개 보 지점의 랜덤포레스트의 변수 중요도를 살펴본 결과, 상류의 보 지점들은 남조 류 발생에 있어 보 운영에 따른 영향을 직접적으로 받는 것으로 나타났다. 이는 효율적인 보 운영을 통한 남조류 관리 가 가능할 수 있다는 것을 의미한다. 중류 구간은 DO와 E.C가 주요 영향인자로 도출되었다. 공간적으로 구미와 김천 에 대규모 산업공단들이 밀집되어 있으며, 환경기초시설의 배출량이 큰 영향을 끼치는 구간이다. 따라서 폭염 및 가 뭄 시기에 중류 유역에서 배출되는 환경기초시설의 방류는 본류의 E.C를 증가하게 하고 남조류 발생을 촉진 시키는 것으로 나타났다. 중·하류에 위치한 보 지점들은 폭염 및 가뭄의 영향을 가장 많이 받는 지역으로 여름철 가뭄에 따른 남조류 대발생에 대비하여 선제적인 관리가 필요한 지점으로 나타났다. 본 연구를 통해 지점별 남조류 발생 영향인자 를 도출하였으며, 맞춤형 조류관리를 위한 정책적 의사결정의 기초자료로 제공될 수 있을 것으로 판단된다.

In this study, the main influencing factors of the occurrence of cyanobacteria at each of the eight Multifunctional weirs were derived using a random forest, and a categorical prediction model based on a Algal bloom warning system was developed. As a result of examining the importance of variables in the random forest, it was found that the upstream points were directly affected by weir operation during the occurrence of cyanobacteria. This means that cyanobacteria can be managed through efficient security management. DO and E.C were indicated as major influencers in midstream. The midstream section is a section where large-scale industrial complexes such as Gumi and Gimcheon are concentrated as well as the emissions of basic environmental facilities have a great influence. During the period of heatwave and drought, E.C increases along with the discharge of environmental facilities discharged from the basin, which promotes the outbreak of cyanobacteria. Those monitoring sites located in the middle and lower streams are areas that are most affected by heat waves and droughts, and therefore require preemptive management in preparation for the outbreak of cyanobacteria caused by drought in summer. Through this study, the characteristics of cyanobacteria at each point were analyzed. It can provide basic data for policy decision-making for customized cyanobacteria management.

3

4,200원

본 연구의 핵심 질문은 미래 시점의 질병 발생 확률을 예측하는 미래 예측(forward prediction)에 비해 사건 기준 예후 감지(event-anchored pre-diagnostic prediction)가 성능 우위를 보이는가이다. 이 질문에 대한 연구자들의 대답은 긍정적이다. 이는 전자가 본질적으로 높은 불확실성을 수반하는 반면 후자는 진단 이전에 이미 존재하는 잠재적 생리 신호를 탐지하는 접근법이기 때문이다. 이와 같이 두 접근법은 동일한 입력 변수를 사용하지만 목표 변수의 정의와 시간 구조가 상이하며, 이러한 차이가 예측 성능에 미치는 영향을 분석하는 것이 본 연구의 핵심 목적이다. 본 연구에서는 환자 10,000명과 비환자 10,000명으로 구성된 총 20,000개의 ECG 기반 데이터를 구축하였으며, 각 데이터는 진단 시점을 기준으로 한 시간 정보를 포함하도록 설계하였다. QTc 간격, 심박수, 신호 통계 특성 등 ECG 기반 변수들을 추출하여 동일한 입력 변수 집합을 구성하고, 이를 기반으로 3년, 5년, 7년, 10년, 15년의 다양한 시간 구간에 대한 미래 예측 모델과 진단 시점을 기준으로 정렬된 예후 감지모델을 구축하였다. 모델 성능은 계층적 5겹 교차검증과 ROC 곡선 아래 면적(AUC)을 통해 평가하였다. 분석 결과, 모든 시간 구간에서 예후감지 모델이 미래 예측 모델보다 일관되게 높은 성능을 보였으며, 이러한 성능 차이는 시간적으로 가까운 구간에서 더욱 크게 나타났다. 또한 모델 유형, 변수 구성, 랜덤 시드 변화에도 불구하고 결과는 강건하게 유지되었다. 추가 분석을 통해 ECG 신호는 진단 수년 이전부터 질병과 관련된 잠재적 정보를 포함하고 있음을 확인하였다. 본 연구는 이러한 성능 차이가 특정 모델에 의존하는 것이 아니라 예측 문제의 구조적 차이에 기인함을 보여준다. 이는 의료 인공지능에서 문제 정의의 중요성을 강조하며, 장기 질병 예측에서 시간 정렬 방식이 예측 성능에 미치는 영향을 이해하는 데 중요한 시사점을 제공한다.

The guiding research question of this study is whether event-anchored pre-diagnostic prediction outperforms forward prediction, which estimates the probability of future disease occurrence. The answer provided by this study is affirmative. This is because forward prediction inherently involves a high level of uncertainty, whereas pre-diagnostic prediction focuses on detecting latent physiological signals that already exist prior to clinical diagnosis. Although both approaches utilize the same set of input features, they differ fundamentally in the definition of the target variable and the temporal structure of the data. Investigating how these differences affect predictive performance constitutes the primary objective of this study. To this end, we constructed a large-scale ECG-based dataset consisting of 20,000 samples, including 10,000 patients and 10,000 non-patients. Each record was designed to include temporal information relative to the diagnosis event. ECG-derived features, such as corrected QT interval (QTc), heart rate, and signal statistical characteristics, were extracted to form a consistent set of input variables. Based on these features, forward prediction models were developed across multiple temporal horizons (3, 5, 7, 10, and 15 years), while pre-diagnostic models were aligned relative to the diagnosis event. Model performance was evaluated using stratified 5-fold cross-validation and the area under the receiver operating characteristic curve (AUC). The results consistently demonstrate that pre-diagnostic models outperform forward prediction models across all temporal horizons, with the performance gap being more pronounced at shorter temporal distances. Furthermore, the findings remain robust across different model types, feature configurations, and random seeds. Additional analyses reveal that ECG signals contain detectable latent disease-related information several years prior to diagnosis. These findings suggest that the observed performance differences are not driven by model choice, but rather by fundamental differences in problem structure. This study highlights the importance of problem formulation in medical artificial intelligence and provides important insights into how temporal alignment influences predictive performance in long-term disease risk modeling.

4

랜덤 포레스트를 활용한 청년층의 정규직 첫 직장 영향요인 탐색 KCI 등재

유승완, 이찬

한국취업진로학회 취업진로연구 제15권 4호 통권52호 2025.12 pp.167-193

※ 기관로그인 시 무료 이용이 가능합니다.

6,600원

본 연구는 청년층의 첫 직장에서 정규직으로 진입하는 데 영향을 미치는 요인을 규명하고자, 청년패 널조사(YP2021) 1·2차 자료를 활용하여 랜덤 포레스트(Random Forest)와 로지스틱 회귀분석을 적용하 였다. 기존 연구들이 취업 여부나 객관적 스펙 중심으로 설명에 그친 데 비해, 본 연구는 고용의 질인 고용형태(정규직/비정규직)에 초점을 맞추어 분석함으로써 새로운 시각을 제시한다. 랜덤 포레스트 분석 결과, 정규직 진입에 가장 중요한 변수는 지역 요인(경기, 울산)이었으며, 자아존 중감·구직효능감 등 심리적 요인, 구직활동(지원 횟수), 스펙 중요성 인식(직무경험·영어회화 능력 등)이 높은 예측력을 보였다. 특히 객관적 스펙 보유 여부보다 개인이 특정 스펙을 얼마나 중요하게 인식하는 지가 정규직 취업 가능성을 더 잘 설명하는 것으로 나타났다. 부분의존성 도표 분석에서는 지역, 구직 기간, 심리 변수 등에서 비선형적 영향이 확인되었다. 로지스틱 회귀분석 결과, 여러 설명변수 중 통계적으로 유의한 변수는 경기·울산 지역 거주뿐이었다. 이는 지역 간 고용 접근성과 산업구조의 차이가 청년층의 정규직 진입 가능성에 구조적으로 작용함을 시사한다. 반면 심리적 요인과 스펙 인식 요인은 예측력은 높았으나 통계적 유의미한 영향은 확인하지 못했다. 본 연구는 첫 직장의 고용형태에 영향을 미치는 요인을 개인특성, 스펙 준비, 스펙 인식, 직업가치관, 심리적 요인을 포함한 다면적 접근으로 분석하여 객관적 스펙 중심 설명의 한계를 보완하였다. 또한 청 년층이 특정 스펙의 중요성을 과대평가하는 경향과 실제 취업효과 간의 괴리를 확인함으로써, 스펙 중 심 준비에서 직무적합성 기반의 전략적 경력설계 체계가 필요함을 제안한다. 나아가 지역 간 노동시장 접근성 불균형을 완화할 수 있는 지역 기반 청년 일자리 정책, 심리역량 기반 진로지원 프로그램의 강 화, 스펙 해석 능력을 높이는 커리어 교육의 중요성을 정책적 시사점으로 제시한다.

This study investigates factors influencing young adults’ entry into regular employment at their first job using Youth Panel Survey (YP2021) data and applying Random Forest (RF) and logistic regression analyses. Unlike prior studies centered on simple employment status or objective credentials, this study focuses on employment quality and incorporates individual, skill-related, psychological, and regional variables. The RF results indicate that regional factors—especially residence in Gyeonggi—are the strongest predictors of regular employment, followed by psychological variables (self-esteem, job-search efficacy), job-search intensity, and perceived importance of specific skills. Perceived skill importance showed higher predictive power than actual credential possession. Partial dependence plots reveal nonlinear effects across key variables. In contrast, the logistic regression results show statistical significance only for regional variables such as Gyeonggi and Ulsan, suggesting that regional industrial structures and labor market accessibility play decisive roles in shaping regular employment outcomes. Psychological and skill-perception variables exhibited predictive value but did not reach statistical significance. These findings highlight a mismatch between young adults’ perceptions of skill importance and the actual determinants of regular employment. The study underscores the need for regionally tailored youth employment policies, psychological capital–based career support, and career education that helps young job seekers interpret and align their competencies with job requirements.

5

9,000원

본 연구는 머신러닝 기법을 활용하여 ESG(Environmental, Social, and Governance) 펀드와 일반 펀드의 주식 포트폴리오 편입 결정 요인을 분석하였다. 편입 결정 요인으로는 9개의 기업 특성 변수와 4개의 ESG 평가 등급을 사용하였으며, Random forest, XGBoost(Extreme Gradient Boosting), LightGBM(Light Gradient Boosting Machine) 알고리즘을 적용하여 주식의 펀드 편입 여부를 예측하는 모델을 구축하였다. 2022년 12월말 기준 한국의 공모 액티브 주식형 펀드 데이터 를 대상으로 한 분석 결과, ESG 펀드와 일반 펀드 모두에서 ‘기업의 규모’가 펀드 편입 결정에 가장 큰 영향을 미치는 요인으로 나타났으며, ESG 평가 등급은 전반적으로 다른 기업 특성 변수에 비해 상대적 중요도가 낮은 것으로 확인되었다. 다만, ESG 펀드의 경우 일반 펀드 대비 환경 등급과 종합 ESG 평가 등급이 펀드 편입 결정에 미치는 영향력이 상대적으로 더 크게 나타났다. 본 연구는 머신러닝 기법을 활용하여 기존 선형 모형의 한계를 보완하고, 펀드 편입 결정 요인을 보다 정교하게 규명하였다는 점에서 의의가 있다. 또한 본 연구는 머신러닝 기법을 통해 ESG 펀드가 실제로 ESG 요소를 어느 정도 반영하는지 검증함으로써 펀드 운용의 투명성을 높이고, ESG 펀드 공시체계 개선을 위한 정책적 시사점을 제공한다는 점에서도 의미가 크다. 아울러 본 연구에서 제시한 분석 틀은 향후 펀드 평가와 투자전략 수립에 활용할 수 있는 실무적·학문적 기반을 제공할 것으로 기대된다.

This study analyzes the determinants of stock selection decisions for ESG funds and conventional equity funds using machine learning techniques. Specifically, the analysis employs nine firm-specific variables and four ESG rating categories (Environmental, Social, Governance, and Composite ESG) to predict stock inclusion in equity fund portfolios. The analysis employs decision tree-based machine learning algorithms --- Random forest, XGBoost, and LightGBM. Based on data from Korean public active equity funds as of December 2022, we find that firm size is the most significant factor in fund inclusion decisions for both ESG and conventional funds. Although ESG ratings exhibit relatively lower importance compared to other firm-specific factors, the Environmental and Composite ESG ratings show a significantly higher impact within ESG funds than in conventional funds. The study underscores the value of machine learning techniques in uncovering the determinants of fund inclusion decisions and offers practical and policy implications for ESG investment.

6

4,600원

본 연구는 의료 데이터 분석에서 단순모형과 복잡모형의 방법론적 우수성을 비교하였다. 두 모형의 우수성 비교를 위한 실증분석을 위해 피 마 인디언 당뇨(Pima Indians Diabetes) 자료를 사용해 0값을 결측으로 처리하고 중앙값 대치·표준화를 거친 뒤 70:30 분할 및 동일 전처리 파 이프라인으로 로지스틱 회귀(단순), 랜덤포레스트, 다층퍼셉트론(MLP)을 학습·평가하였다. 두 모형의 성능평가를 위해서는 임계값 의존 지표 (정확도, 민감도, 특이도, 정밀도(PPV), 음성예측도, 조화평균(F1), 균형정확도, 매튜스 상관계수(MCC))와 임계값 무관·보정 지표(ROC-AUC, PR-AUC, Brier(↓), 보정 절편/기울기, Hosmer-Lemeshow(HL), 임상 순이득(DCA))를 적용하였다. 실증분석 결과, 로지스틱 회귀는 정확도 0.7662, 민감도 0.7160, F1 0.6824, 균형정확도 0.7547, MCC 0.4995로 전반적으로 가장 견조했으며 ROC-AUC 역시 0.8365로 최고였다. 반면 랜덤포레스트는 특이도 0.8733과 PPV 0.6885로 확증 성능이 우수했고, 보정 품질에서도 Brier 0.1594, 보정 절편 0.1356, 기울기 0.9914, HL p=0.7606으로 이상적 수준에 근접했다. 다층퍼셉트론은 동일 조건에서 상대적으로 열세를 보였다. 한편 DCA 차원에서는 전 구간 공통으로 모형 간 격차는 크지 않되, 임계값 선택에 따라 승자가 바뀌었다(0.10: 랜덤포레스트, 0.20~0.30: 로지스틱 회귀). 따라서 실제 적용 시 목표 pt 와 FN/FP 비용구조를 먼저 정하고, 그 범위에서 순이득이 가장 큰 모형-컷오프 조합을 선택하는 것이 합리적이다. 종합하자면, 특정 모형의 방 법론적 우수성은 절대적이지 않고 상대적이며, 분석목적에 따라 선택이 필요하다. 구체적으로 선별(미탐 최소화)이 목표일 때는 로지스틱 회 귀가 적합하며, 확증(위양성 최소화)이나 확률 기반 의사결정에는 보정이 우수한 랜덤포레스트가 적합하다. 따라서 분석모형의 최적 선택은 유병률·FN/FP 비용·임계확률을 반영한 DCA와 함께 이뤄져야 한다.

This study compared the methodological superiority of simple models and complex models in medical data analysis. For an empirical analysis to compare the superiority of the two model classes, we used the Pima Indians Diabetes dataset, treated zeros as missing, performed median imputation and standardization, split the data 70:30, and trained/evaluated logistic regression (simple), random forest, and multilayer perceptron (MLP) under an identical preprocessing pipeline. For performance evaluation of the two model classes, we applied threshold-dependent metrics (accuracy, sensitivity, precision, specificity, F1, balanced accuracy, MCC) and threshold-independent and calibration metrics (ROC-AUC, PR-AUC, Brier, calibration intercept/slope, Hosmer–Lemeshow, clinical net benefit (DCA)). In the empirical results, logistic regression was overall the most robust, with accuracy 0.7662, sensitivity 0.7160, F1 0.6824, balanced accuracy 0.7547, and MCC 0.4995, and also achieved the highest ROC-AUC of 0.8365. By contrast, random forest showed superior rule-in performance with specificity 0.8733 and PPV 0.6885, and was close to an ideal level in calibration quality, with Brier 0.1594, calibration intercept 0.1356, slope 0.9914, and HL p=0.7606. The MLP was relatively inferior under the same conditions. Meanwhile, in the DCA dimension, the gaps between models were not large across the range, but the winner changed with the choice of threshold (0.10: random forest; 0.20–0.30: logistic regression). Therefore, in practical application, it is reasonable to first specify the target pt and the FN/FP cost structure, and then choose the model–cutoff combination that yields the largest net benefit within that range. In sum, the methodological superiority of a given model is not absolute but relative, and selection should depend on the analytic objective. Specifically, when the goal is screening (minimizing missed positives), logistic regression is appropriate, whereas for rule-in (minimizing false positives) or probability-based decision making, random forest with superior calibration is suitable. The optimal choice of analysis model should be made together with DCA that reflects prevalence, FN/FP costs, and the threshold probability.

7

랜덤포레스트를 이용한 기상 환경에 따른 이상기온 분류 KCI 등재후보

김윤수, 송광윤, 장인홍

조선대학교 기초과학연구원 통합자연과학논문집(구 조선자연과학논문집) 제17권 1호 2024.03 pp.1-12

※ 기관로그인 시 무료 이용이 가능합니다.

4,300원

Many abnormal climate events are occurring around the world. The cause of abnormal climate is related to temperature. Factors that affect temperature include excessive emissions of carbon and greenhouse gases from a global perspective, and air circulation from a local perspective. Due to the air circulation, many abnormal climate phenomena such as abnormally high temperature and abnormally low temperature are occurring in certain areas, which can cause very serious human damage. Therefore, the problem of abnormal temperature should not be approached only as a case of climate change, but should be studied as a new category of climate crisis. In this study, we proposed a model for the classification of abnormal temperature using random forests based on various meteorological data such as longitudinal observations, yellow dust, ultraviolet radiation from 2018 to 2022 for each region in Korea. Here, the meteorological data had an imbalance problem, so the imbalance problem was solved by oversampling. As a result, we found that the variables affecting abnormal temperature are different in different regions. In particular, the central and southern regions are influenced by high pressure (Mainland China, Siberian high pressure, and North Pacific high pressure) due to their regional characteristics, so pressure-related variables had a significant impact on the classification of abnormal temperature. This suggests that a regional approach can be taken to predict abnormal temperatures from the surrounding meteorological environment. In addition, in the event of an abnormal temperature, it seems that it is possible to take preventive measures in advance according to regional characteristics.

8

4,000원

본 연구에서는 서울 소재 한 전문대학 학생들을 대상으로 하여 최소한의 인구통계학적 변수와 1학년 1학기 성적을 활용하여 학생들의 최종 학적 상태를 예측하고자 하였다. XGBoost와 LightGBM 모델을 사용한 결과, 이러한 변수들이 학생 들의 제적 여부 예측에 유의미한 것을 발견하였다. 이는 학업 시작 초기의 성적이 학업 중단의 중요한 지표가 될 수 있음을 시 사한다. 또한, 전문대학의 학제가 최종 학적에 영향을 미칠 가능성을 확인하였으며, 이는 학업 기간이 학생들의 학업 중단 결 정에 중요한 요소임을 나타낸다. 전문대학에서 조기 학업 중단 의도를 파악하는 데 있어 심리적, 사회적, 경제적 요인에 의존 하지 않고 학업 성취도만을 기준으로 모델링을 시도하였다. 이는 향후 학업 중단에 대한 조기 경보 시스템 구축에 도움이 될 것으로 기대된다.

This study utilized minimum number of demographic variables and first-semester GPA of students to predict the final academic status of students at a vocational college in Seoul. The results from XGBoost and LightGBM models revealed that these variables significantly impacted the prediction of students' dismissal. This suggests that early academic performance could be an important indicator of potential academic dropout. Additionally, the possibility that academic years required to award an associate degree at the vocational college could influence the final academic status was confirmed, indicating that the duration of study is a crucial factor in students' decisions to discontinue their studies. The study attempted to model without relying on psychological, social, or economic factors, focusing solely on academic achievement. This is expected to aid in the development of an early warning system for preventing academic dropout in the future.

9

교통사고 정보를 이용한 과실비율 산정 모델 개발 KCI 등재

한음, 박기옥, 강희진, 이요셉, 윤일수

한국ITS학회 한국ITS학회논문지 제21권 제6호 통권104호 2022.12 pp.36-56

※ 기관로그인 시 무료 이용이 가능합니다.

5,700원

국내에서 발생하는 교통사고는 손해보험협회에서 작성한 「자동차사고 과실비율 인정기준」 에 따라 과실비율을 산정하며, 이를 통해 보험사의 합의나 판결이 내려진다. 하지만, 과실비율 산정에 있어 분쟁이 빈번하게 일어나고 있다. 따라서, 교통사고 발생 시 경찰공무원에 의해 작 성되는 교통사고 정보를 이용하여 「자동차사고 과실비율 인정기준」상의 교통사고 유형을 신 속하게 확인할 수 있다면, 보다 효과적인 대응이 가능할 것으로 사료된다. 이에 본 연구에서는 경찰에 의해 작성된 교통사고 정보를 학습시켜 「자동차사고 과실비율 인정기준」에서 제시하 는 교통사고 유형으로 분류하는 모델을 개발하고자 한다. 특히, 데이터마이닝을 통해 경찰청 교통사고 데이터에서 「자동차사고 과실비율 인정기준」의 교통사고 유형으로 분류하는 데 필 요한 핵심어들을 추출하였다. 그리고, 키워드를 의사결정나무 및 랜덤 포레스트 모델을 통해 학습시켜 교통사고 유형을 도출하는 모델을 개발하였다.

Traffic accidents occur in Korea are calculated with the 「Automobile Accident Negligence Ratio Certification Standard」 prepared by the ‘General Insurance Association of Korea’ and the insurance company's agreement or judgment is made. However, disputes are frequently occurring in calculating the negligence ratio. Therefore, it is thought that a more effective response would be possible if accident type according to the standard could be quickly identified using traffic accident information prepared by police. Therefore, this study aims to develop a model that learns the accident information prepared by the police and classifies it to match the accident type in the standard. In particular, through data mining, keywords necessary to classify the accident types of the standard were extracted from the accident data of the police. Then, models were developed to derive the types of accidents by learning the extracted keywords through decision trees and random forest models.

10

4,000원

본 연구는 자동차 구매 계약에 대하여 해약 예측 성능이 우수한 최적의 머신러닝 모델을 구현하는 것으로 국내 수입 자동차 딜러사에서 판매, 고객, 재고 관리를 위해 사용하고 있는 영업지원시스템(SFA)에 축적되어 있는 계약, 해약, 판매 정보를 머신러닝 모델에 적용하여 해약 가능성을 예측하였다. 영업지원시스템(SFA)의 2015년부터 2020년까지의 계약, 해약, 판매 데이터 27,208건을 추출하여 분석 프로그램 파이썬 주피터 노트북으로 데이터 전처리 및 검증 후에 로지스틱 회귀, 서포트 벡터 머신, 램덤 포레스트, 가우시안 NB, 인공신경망 모델에 적용 및 학습하고 신규 데이터를 이용하여 최종 결과를 예측하였다. 학습 데이터 셋은 1개의 인덱스와 13개의 독립변수, 1개의 목표변수로 각각의 머신러닝 모델 성능을 분석하였으며 성능 평가 척도인 정확도, 정밀도, 재현율, F1-Score, AUC 값으로 최적의 머신러닝 모델을 구현하였다. 본 연구 결과, 차량 구매 계약 정보를 머신 러닝 모델에 적용하여 해약 예측의 가능성을 확인하였으며 머신러닝 모델 중에 서포트 벡터 머신과 딥 러닝 기반 인공 신경망이 예측 성능이 우수한 것으로 확인되었다. 학문적 의의로는 수입차 영업지원시스템(SFA)의 축적된 데이터를 머신러닝 기술과 통합하여 일반화된 예측 모델 개발 및 활용 가능성을 제시하였으며, 계약 고객에 대한 해약 예측을 통하여 잠재적인 해약 고객을 집중적으로 관리함으로써 판매 실패를 미연에 방지하고 매출 향상에 기여할 수 있는 실무적인 시사점을 제시하였다.

11

안드로이드 퍼미션(permission) 정보를 특징정보로 하여 기계학습 기반으로 안드로이드 멀웨어(malware)를 탐지하는 연구들이 수행되어 왔다. 그러나, 대부분의 기존 연구들에서는 기본 퍼미션(built-in permissions)들만을 추출하여 활용하였으며, 최근 일부 연구들에서 커스텀 퍼미션(custom permission)들을 도입하였다. 본 논문에서는 실험을 통해 안드로이드 멀웨어 탐지에서 커스텀 퍼미션들의 영향을 분석한다. 랜덤 포레스트(Random Forest) 분류기를 적용한 실험 결과, 커스텀 퍼미션들이 멀웨어 탐지 정확도를 1.53% 향상시켜 주었다.

12

고교생의 교사 인식 관련 예측 요인 탐색 KCI 등재

김영식, 문찬주, 박환보

한국교원교육학회 한국교원교육연구 제37권 제1호 2020.03 pp.465-494

※ 기관로그인 시 무료 이용이 가능합니다.

7,000원

본 연구는 한국교육고용패널Ⅱ 1차년도(2016년) 자료에 대하여 랜덤 포레스트 기법을 활용 하여, 고교생의 담임 교사 및 학교 교사들에 대한 인식을 예측하는 요인을 탐색하였다. 또한 모형의 예측성과를 검토함으로써 도출된 랜덤 포레스트의 추정 모형이 적합한 모형임을 확인 하고자 하였다. 이와 함께 고교생의 담임 및 학교 교사 인식에 대한 예측력이 상대적으로 높 은 주요 변수들을 활용하여 중다회귀분석 모형과 학교 고정효과 모형, 학교 확률효과 모형 분석을 실시함으로써 세 가지 모형 중 담임 교사 및 학교 교사에 대한 인식을 예측하는데 보 다 효과적인 모형을 확인하고자 하였다. 분석 결과, 부모의 양육 태도, 학교 내 진로교육·활동에 대한 만족도, 수업태도 및 자존감 등과 같은 학생 개인의 정의적 특성, 담임 선생님이 학생을 대하는 방식 등이 담임 교사 및 학교 교사들에 대한 고교생의 긍정적인 인식에 영향을 미치는 것으로 나타났다. 그리고 담임 교사 인식을 예측하는데 있어서는 확률효과모형이, 학교 교사 인식을 예측하는데 있어서는 고정효과모형이 적절함을 통계적으로 확인하였다. 본 연구는 이상과 같은 분석 결과에 기반하여 교사들에 대한 긍정적인 인식을 제고하기 위 한 정책적인 시사점을 제공하였으며, 보다 엄밀한 관련 요인 탐색을 위한 학술적 제언 또한 함께 제공하였다.

This study explored factors that predict the perception of high school students on homeroom teachers and school teachers by utilizing the random forest technique for the first year of the Korean Education Employment Panel II (2016) and examined whether the estimated model of the random forest was appropriate. In addition, this study tried to identify more effective models among the three models of the multiple regression model, the fixed-effect model, and the random effect model in predicting the perception of the students. The analysis found that the following variables affected the positive perception of students on homeroom teachers and school teachers: parenting attitude, satisfaction with career education and activities in schools, the affective characteristics of students such as class attitude and self-esteem, and the way homeroom teachers treat their students. It also statistically confirmed that the random effect model was appropriate for predicting the perception of students on homeroom teachers and the fixed effect model was appropriate for predicting the perception of students on school teachers. Based on the results, this study provided policy implications for enhancing students’ positive perception on teachers and also provided academic suggestions for a more rigorous exploration of relevant factors.

13

한국 중서부 논 습지에 서식하는 도롱뇽(Hynobius leechii)의 번식지 특성

도민석, 서재화, 손석준, 최그린, 유나경, 정지화, 구교성, 이상철, 남형규

한국양서ㆍ파충류학회 한국양서ㆍ파충류학회지 Volume 11 Number 1 2020.03 pp.21-29

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

논 습지는 자연 습지를 대체할 수 있는 야생동물들의 주요 서식공간으로 세계적으로 논 습지의 보전 및 관리 를 위한 다양한 생태적 연구들이 수행되고 있다. 양서류 는 논 습지를 번식공간으로 이용하는 대표적인 분류군 중 하나로 서식지 파괴와 수질 오염으로 개체수가 급감 하고 있다. 본 연구에서는 한국 중서부 논 습지에 서식 하는 도롱뇽의 번식지 환경특성을 파악하고, 한국 중서 부 논 습지가 번식지로 적합한지 확인하였다. 이를 위해 2016년부터 2017년까지 총 40개의 논을 2개월간(3월-4 월) 도룡뇽의 번식 유무와 경관, 대기, 물리, 수 환경요 인을 파악하였다. 그 결과, 중서부 논 습지에서 도롱뇽 의 번식이 확인된 지점은 총 40개 중 8개로, 도롱뇽의 번식지는 비번식지 보다 경관적으로 논의 면적이 좁고 산림의 면적이 넓으며, 해안과의 거리가 멀고, 고도가 높은 지역으로 나타났고, 수 환경적으로 EC와 TDS, NaCl이 낮은 지역으로 확인되었다. 도롱뇽의 번식지 선 택에 가장 큰 영향을 끼친 주요 환경요인은 모두 경관 요인으로 파악되었다. 종합적으로 한국 중서부의 산림 주변에 위치한 논 습지는 도롱뇽들의 번식지로써 적합 한 환경을 갖추고 있었으며, 수 환경 또한 양서류의 서 식 권장 기준 범위에 적합했다. 하지만, 도롱뇽이 번식 하지 않는 해안가 주변의 평지에 위치한 논 습지에서는 EC, TDS가 매우 높게 확인되어 양서류의 서식에 부정 적인 영향을 끼칠 것으로 예상된다.

Rice fields are major habitats of wildlife and they can substitute natural wetlands. Therefore, various ecological studies have been conducted for the conservation and management of rice fields worldwide. Amphibians are one of the representative taxonomic groups using rice fields as their breeding sites and their populations are rapidly decreasing due to habitat destruction and water pollution. This study examined the environmental characteristics of Wonsan salamander(Hynobius leechii)’s breeding sites in rice fields in the mid-western region of South Korea and evaluated if they were suitable breeding sites for Wonsan salamanders. The results of this study showed that Wonsan salamander used eight sites out of the surveyed 40 rice field sites in the mid-western region of South Korea for their breeding. To achieve study objectives, this study evaluated the breeding of Wonsan salamanders, landscape factors, atmospheric factors, physical factors, and aquatic factors from March to April in 2016 and 2017. The Wonsan salamander’s breeding sites had smaller rice field areas and larger forest areas, were farther from the coast, and were located at higher altitude compare to Wonsan salamander’s nonbreeding sites. These breeding sites also had low EC, TDS, and NaCl. It was found that all environmental factors affecting the breeding site selection greatly were landscape factors. In conclusion, rice fields near forests in the mid-western region of South Korea had an environment suitable for Wonsan salamander breeding and their aquatic environment also met the recommended conditions for amphibian species’ breeding. However, Wonsan salamanders didn’t breed in rice fields on plains located near the coast and the EC and TDS of them were high, which could adversely affect the inhabitation of amphibian species.

14

4,000원

대화시스템은 인간과 컴퓨터의 상호작용에 새로운 패러다임이 되고 있다. 자연어로써 상호작용함으로써 인간 은 보다 자연스럽고 편리하게 각종 서비스를 누릴 수 있게 되었다. 대화시스템의 구조는 일반적으로 음성 인식, 자연 어 이해, 문맥 파악 등의 여러 모듈의 파이프라인으로 이뤄지는데, 본 연구에서는 자연어 이해 모듈의 도메인 분류 문 제를 풀기 위해 convolutional neural network, random forest 등의 기계학습 모델을 비교하였다. 사람이 직접 태 깅한 총 7개 서비스 도메인 데이터에 대하여 각 문장의 도메인을 분류하는 실험을 수행하였고 random forest 모델 이 F1 score 0.97 이상으로 가장 높은 성능을 달성한 것을 보였다. 향후 다른 기계학습 모델들을 추가 실험함으로써 도메인 분류 성능 개선을 지속할 계획이다.

Dialog system is becoming a new dominant interaction way between human and computer. It allows people to be provided with various services through natural language. The dialog system has a common structure of a pipeline consisting of several modules (e.g., speech recognition, natural language understanding, and dialog management). In this paper, we tackle a task of domain classification for the natural language understanding module by employing machine learning models such as convolutional neural network and random forest. For our dataset of seven service domains, we showed that the random forest model achieved the best performance (F1 score 0.97). As a future work, we will keep finding a better approach for domain classification by investigating other machine learning models.

15

알츠하이머 병 (Alzheimer's disease, AD) 에서 해마는 신경 세포 및 신경 섬유 얽힘 스레드 침착으로 인해 영향 을 받는 첫 번째 구조물 중 하나이며 결국 신경 세포의 손실을 초래합니다. 또한 최근에 많은 자기 공명 영상 연구에 서 AD 환자는 건강한 피검자에 비해 해마의 체적이 작다는 의견을 제시했다. 해마 자기 공명 영상 체적은 AD를 위한 잠재적인 바이오 마커이지만 작은 구조로 인해 수동 분할 및 낮은 시각화의 한계가 에 의해 어려움이 있다. 이 러한 문제를 해결하기 위해 대부분의 이미지 분석가는 뇌 구조의 자동 세분화를 위해 Freesurfer와 FSL에 익숙했 습니다. 본 연구에서는 다른 그룹의 AD 환자의 조기 진단을 위해 Freesurfer (v.6.0) 자동화 도구 상자를 사용하여 해마 영역을 추출하는 방법을 이용한다. 38 명의 환자가 AD에 속하며, 46 명의 환자가 MCI에 속한다 (안정-MCI, 18 개월 후에 AD로 전환되지 않음) 및 36 명의 환자가 MCIc에 속한다 (전환-MCI, 18 개월 이내에 AD로 전환), 나머지 38 명의 노인은 정상 대상 (NC)에 속하며 이러한 모든 환자들의 정보는 알츠하이머병 Neuroimaging Initiative database (ADNI)에서 다운로드 되었습니다. AD, MCIs, MCIc 및 노인 대조군 (NC) 환자를 구별하 기 위해 우리는 구입한 해마 체적을 사용했습니다. 여기서 차원적 감소를 위해 우리는 다양한 학습 기반의 ISOMAP 기법을 사용했다. 이 기법은 많은 기능들 중 중요한 특징만을 선택하고 나중에 랜덤 포레스트와 softmax 분류기를 사용하여 이진 분류 문제들을 분류했다.

In Alzheimer’s disease (AD), the hippocampus is among the first structures, which is affected, due to the neuropil and neurofibrillary tangle thread deposition, which eventually results in a neuronal loss. Moreover, recently a large number of magnetic resonance imagining studies have stated that AD patients has get smaller hippocampus volume as compared to healthy subjects. Hippocampal magnetic resonance imaging volumetric is a potential biomarker for an AD but is hindered by the limitation of manual segmentation and low visualization because of its small structure. To solve that problems, most image analyst used to Freesurfer and FSL for automatic segmentation of brain structures. In this paper, we apply to extract hippocampus region using Freesurfer (v.6.0) automated toolbox for early diagnosis of AD subjects with different groups. a total 158 baseline subjects were used for the experiment, from which 38 patients belong to AD, 46 patients belong to MCIs (stable-MCI, not converted to AD after 18th month of periods) and 36 patients belong to MCIc (converted-MCI, converted to AD within 18th month of periods), and remaining 38 patients belong to elderly normal subjects (NC), and all these subjects were downloaded from Alzheimer’s disease Neuroimaging Initiative database (ADNI). We used the gained hippocampal volumes to discriminate between AD, MCIs, MCIc, and elderly controls (NC) patients. Here, for dimensionality reduction purpose we have used manifold learning based ‘isometric feature mapping’ ISOMAP technique, which only selects the important feature from a bunch of features and later random forest and softmax classifier were used to classify the binary classification problems.

16

랜덤 포레스트(Random Forest)의 시계열 적용에 관한 연구: 한국 물가상승률 예측 사례 분석

한희준

[NRF 연계] 한국경제학회 경제학연구 Vol.71 No.3 2023.09 pp.37-73

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본고는 한국과 미국의 물가상승률 예측에 랜덤 포레스트 모형을 적용할 때, Stationary Bootstrap이나 Moving Block Bootstrap 등 Block Bootstrap을 사용하는 것이 통상적인 독립 부트스트랩(Independent Bootstrap)을 사용하는 것에 비해 통계적으로 유의한 수준으로 예측력을 개선하지는 못한다는 것을 보인다. 그리고 FRED-MD를 참고한 총 93개의 관련 국내외 거시경제/금융 변수들을 사용하고, XGBoost, LSTM 등 다양한 머신러닝 방법을 활용하여 한국의 물가상승률을예측하고 분석한다. 2004년 9월에서 2022년 3월까지의 표본을 이용하였고, 1개월에서 12개월의 예측 대상기간(Forecast Horizon)을 고려하였다. 총 13개의 모형 중 대부분의 예측 대상기간에 있어 예측력이 우수한 모형이 존재하는 것으로나타났는데, 이는 보루타 알고리즘(Boruta Algorism)을 통해 중요한 변수로 분류된 변수들만을 랜덤 포레스트에 적용하는 모형이다. Giacomini and White(2006) 와 Hansen et al.(2009)의 검정을 통해 대부분의 예측 대상기간에서 통계적으로유의하게 예측력이 우수함을 확인하였는데, 특히 경제활동인구 및 취업자 수의증가율 등 고용시장 관련 변수, 기업경기실사지수, 주택가격 변화율 등이 물가상승률 예측에 중요한 변수로 선택되는 것으로 나타났다.

This paper first investigates whether adopting the stationary bootstrap or the moving block bootstrap, instead of the usual independent bootstrap, in the random forest method improves forecasting of stationary time series. It is shown that the block bootstrap procedures adopted in the random forest method do not make any statistically significant improvement in Korean or US inflation forecasting. Secondly, we consider inflation forecasting in Korea using 93 macroeconomic/financial variables and various machine learning methods. The samples are from September 2004 to March 2022. Comparing total 13 models, one model outperforms the rest models for most forecast horizons, which is a simplified method of the model proposed by Kim and Han (2022). The method consists of the following two steps: 1) Select important variables based on the Boruta algorithm, 2) Using only those selected variables, implement the random forest and produce a forecast. The tests by Giacomini and White (2006) and Hansen et al. (2009) show that the model provides significantly better forecasts for most forest horizons. In particular, the Boruta algorithm selected total economically active population, total employed persons, BSI, house price as important variables for Korean inflation forecasting.

17

SWAT 및 random forest를 이용한 기후변화에 따른 한강유역의 수생태계 건강성 지수 영향 평가

우소영, 정충길, 김진욱, 김성준

[Kisti 연계] 한국수자원학회 한국수자원학회 논문집 Vol.51 No.10 2018 pp.863-874

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구에서는 SWAT 모형과 random forest를 이용하여 미래 기후변화에 따른 한강유역($34,148km^2$)의 수생태계 건강성을 평가하였다. 국립환경과학원에서 8년간(2008~2015년) 봄철(4~6월)에 모니터링한 부착돌말류 지수(TDI), 저서형 대형무척추동물지수(BMI), 어류평가지수(FAI)는 0~100점, A~E등급으로 평가되며, 이를 본 연구에서 사용하였다. 수생태 건강성에 영향을 미치는 변수로는 수질(T-N, $NH_4$, $NO_3$, T-P, $PO_4$)과 수온을 선정하였으며, 수질 오염도가 낮은 경우에는 수생태계 건강성 점수가 광범위하게 분포되지만 수질 오염도가 높은 경우 수생태계 건강성 점수가 낮아지는 역상관관계를 확인하였다. 기계학습의 분류 분석 기법 중 하나인 random forest 모델을 이용한 세 개의 수생태 건강성 지수 등급분류 결과 정밀도, 재현율, f1-score 모두 0.81 이상의 예측 정확도를 나타내었다. 기상청의 HadGEM3-RA RCP 4.5와 8.5 시나리오를 적용한 미래 SWAT 수문, 수질 결과 기저유출의 증가로 인해 질소 계열 수질 농도는 기준년도 대비 최대 43.2% 증가하였고, 지표유출 감소로 인해 인 계열수질 오염도는 최대 18.9% 감소하는 것으로 분석되었다. 미래 FAI, BMI의 등급은 개선되는 경향을 보이지만 TDI는 등급이 악화되는 것으로 나타났다. 이를 통해 TDI는 질소 계열 수질에 민감하고 FAI, BMI는 인 계열 수질에 더 민감하다고 판단하였다.

The purpose of this study is to evaluate the future climate change impact on stream aquatic ecology health of Han River watershed ($34,148km^2$) using SWAT (Soil and Water Assessment Tool) and random forest. The 8 years (2008~2015) spring (April to June) Aquatic ecology Health Indices (AHI) such as Trophic Diatom Index (TDI), Benthic Macroinvertebrate Index (BMI) and Fish Assessment Index (FAI) scored (0~100) and graded (A~E) by NIER (National Institute of Environmental Research) were used. The 8 years NIER indices with the water quality (T-N, $NH_4$, $NO_3$, T-P, $PO_4$) showed that the deviation of AHI score is large when the concentration of water quality is low, and AHI score had negative correlation when the concentration is high. By using random forest, one of the Machine Learning techniques for classification analysis, the classification results for the 3 indices grade showed that all of precision, recall, and f1-score were above 0.81. The future SWAT hydrology and water quality results under HadGEM3-RA RCP 4.5 and 8.5 scenarios of Korea Meteorological Administration (KMA) showed that the future nitrogen-related water quality in watershed average increased up to 43.2% by the baseflow increase effect and the phosphorus-related water quality decreased up to 18.9% by the surface runoff decrease effect. The future FAI and BMI showed a little better Index grade while the future TDI showed a little worse index grade. We can infer that the future TDI is more sensitive to nitrogen-related water quality and the future FAI and BMI are responded to phosphorus-related water quality.

18

영산강 유역에서 Sentinel-1 SAR와 Random Forest 방법론을 활용한 고해상도 토양수분 산정

박기진, 박종민

[Kisti 연계] 한국수자원학회 한국수자원학회 논문집 Vol.57 No.12 2024 pp.1085-1098

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

토양수분은 수문순환, 그리고 지표와 대기 사이의 상호작용을 이해하는 데 중요한 요소이다. 이에 따라 국·내외에서는 토양수분 모니터링과 관련된 연구가 활발히 이루어지고 있다. 본 연구에서는 유럽우주국(European Space Agency)의 Sentinel-1 Synthetic Aperture Radar (SAR) 영상과 Random Forest (RF) 방법론을 활용하여 영산강 유역의 고해상도(10m) 토양수분 추정 모델을 개발하였다. 2015년 5월부터 2023년 8월까지의 SAR 영상과 4개의 지상 관측소에서 수집한 토양수분 자료를 훈련 및 검증 자료로 구분하여 모델을 개발하고 추정값에 대한 통계적 검증을 수행하였다. 훈련 과정에 활용하지 않은 검증 자료와 RF 기반 모델을 통해 산정한 토양수분 추정값 사이의 통계적 분석을 수행한 결과, 상관계수(correlation coefficient; R)는 0.75, 일치도(Index of Agreement)는 0.83으로 유의미한 통계치를 도출하였다. 추가적인 검증을 위해 RF 기반 모델을 활용하여 유역 평균 토양수분을 산정하고 European Centre for Medium-Range Weather Forecasts Reanalysis v5-Land 강수 및 Soil Moisture Active Passive (SMAP)/Sentinel-1 토양수분과 비교하였다. 검증 결과, 강수 사상에 따라 유역 평균 토양수분 추정값이 증가하는 경향(R = 0.30)을 나타내었다. 또한, SMAP/Sentinel-1을 활용하여 산정한 유역 평균 토양수분과 RF 기반 모델을 통해 산정한 유역 평균 토양수분을 각각 지점 관측자료와 비교하였을 때, RF 기반 모델(R = 0.37)이 SMAP/Sentinel-1 (R = 0.28)보다 높은 정확도를 나타내었다. 계절적으로는 가을에 가장 높은 유역 평균 토양수분(32.60%)을 나타내었고, 겨울에 가장 낮은 값(30.66%)을 나타내었다. 이러한 분석 과정을 통해, 본 연구에서 개발한 RF 기반 모델이 강수 사상과 계절 변동을 모의할 수 있는 것으로 판단하였다.

Soil moisture (SM) is a key variables in understanding the hydrological cycle and interactions between the land surface and the atmosphere. Consequently, research on SM monitoring has been actively conducted both domestically and internationally. In this study, a high-resolution (10m) SM estimation model for the Yeongsan River watershed was developed using Sentinel-1 Synthetic Aperture Radar (SAR) imagery from the European Space Agency and the Random Forest (RF) method. The model was trained and validated with SAR imagery from May 2015 to August 2023 and SM data collected from four ground observation sites. Statistical verification was conducted for the estimated values. Statistical analysis of the Sentinel-1 and RF-based SM model estimates against independent observed data yielded a correlation coefficient (R) of 0.75 and an index of agreement of 0.83, indicating significant statistical performance. For further validation, the high-resolution SM estimates were used to calculate watershed-averaged SM and compared with precipitation data from the European Centre for Medium-Range Weather Forecasts Reanalysis v5-Land and Soil Moisture Active Passive (SMAP)/Sentinel-1 SM data. The validation results showed an increase in watershed-averaged SM estimates in response to precipitation events (R = 0.30). Additionally, when comparing the watershed-averaged SM estimates derived from the model and those from SMAP/Sentinel-1 with site observed data, the model (R = 0.37) demonstrated higher accuracy than SMAP/Sentinel-1 (R = 0.28). Seasonally, the highest watershed-averaged SM (32.60%) was observed in autumn, while the lowest (30.66%) was observed in winter. This analysis suggests that the model developed in this study is capable of simulating precipitation events and seasonal variations.

19

Ransomware Detection using Random Forest Technique

Ban Mohammed Khammas

[NRF 연계] 한국통신학회 ICT Express Vol.6 No.4 2020.12 pp.325-331

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Nowadays, the ransomware became a serious threat challenge the computing world that requires an immediate consideration to avoid financial and moral blackmail. So, there is a real need for a new method that can detect and stop this type of attack. Most of the previous detection methods followed a dynamic analysis technique which involves a complicated process. The present study proposes a novel method based on static analysis to detect ransomware. The significant characteristic of proposed method is dispensing of disassemble process by direct extraction of features from raw byte with the use of frequent pattern mining which remarkably increases the detection speed. The Gain Ratio technique was used for feature selection which exhibited that 1000 features was the optimal number for detection process. The current study involved using random forest classifier with a comprehensive analysis to the effect of both tree and seed numbers on the ransomware detection. The results showed that tree numbers of 100 with seed number of 1 achieved best results in terms of time-consuming and accuracy. The experimental evaluation revealed that the proposed method could achieve a high accuracy of 97.74% for detection ransomware.

20

Novel hyper-tuned ensemble Random Forest algorithm for the detection of false basic safety messages in Internet of Vehicles

Goodness Oluchi Anyanwu, Cosmas Ifeanyi Nwakanma, 이재민, Dong-Seong Kim

[NRF 연계] 한국통신학회 ICT Express Vol.9 No.1 2023.02 pp.122-129

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Detection of nodes disseminating false data is a prerequisite for effective deployment of Internet of Vehicles (IoV) services. This work proposed a novel hyper-tuned ensemble Random Forest (Ens. RF) algorithm to detect false basic safety messages in IoV. Performance evaluation was done using the Vehicular Reference Misbehavior (VeReMi) dataset comprising data-centric misbehavior evaluation for vehicular networks. For validation, a comparative analysis of the performance of the proposed “Ens. RF” model, five machine learning algorithms implemented in this work, and state-of-the-art ML models from related literature was presented. The performance metrics considered are time efficiency and validation accuracy for overall misbehavior classification. Also, the results confirmed the irrelevance of data balancing in real-life scenarios. Finally, we assess the performance of our proposed system for detecting each falsification scenario using precision and recall. The result shows that the proposed algorithm outperformed others with a validation accuracy of 99.60% and a negligible 604 misclassifications out of 153,730 points.

 
1 2 3 4 5
페이지 저장