Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 40
No
1

소규모 데이터의 예측 성능 개선방안 연구 KCI 등재후보

이승용

한국디지털정책학회 디지털정책학회지 제5권 제2호 2026.06 pp.137-145

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 연구의 목적은 소규모 구조화 데이터 환경에서 예측 정확도 향상을 위한 머신러닝 기반 방법론을 비교 분석하고 모델 복잡도에 따른 일반화 성능과 과적합 특성을 검토하는 데 있다. 시계열 신용 카드 이용 데이터를 활용하여 규제 기반 선형모형과 다양한 앙상블 모델 및 Stacking Ensemble의 예측 성능을 비교한 결과, 규제 기반 선형모형이 가장 안정적인 성능을 보인 반면 비선형고복잡도 모델은 과적합으로 인해 테스트 성능이 저하되었고, 모델 복잡도가 증가할수록 일반화 성능이 감소하는 경향이 확인되었다. 또한 Stacking Ensemble은 성능을 일부 개선하였으나 선형모형의 안정성을 크게 초과하지는 못하였으며, 통계적 검정 결과 모델 간성능 차이는 유의하게 나타나 모델 선택이 예측 성능에 중요한 영향을 미침이 확인되었다. 본 연구는 소규모 데이터 환경에서 복잡한 모델보다 규제 기반 단순모형이 더 안정적인 예측 성능을 제공함을 실증적으로 보여주며 데이터 특성에 맞는 모델 선택의 중요성을 확인하였다. 본 연구는 소규모 데이터를 보유한 기업 등에 필요한 예측 방법론을 적절히 선택하는데 도움을 줄 수 있을 것이다. 향후 연구에서는 중소기업 등에 적합한 머신러닝 기반의 예측 정확성 향상방안을 연구해 보겠다.

The purpose of this study is to compare and analyze machine learning-based methods for improving prediction accuracy in a small-scale structured data environment and to examine generalization performance and overfitting according to model complexity. Using time-series credit card usage data, it is compared regularization-based linear models, various ensemble models, and a Stacking Ensemble. The results show that regularization-based linear models achieved the most stable performance, while nonlinear high-complexity models experienced overfitting and lower test performance, with generalization performance decreasing as complexity increased. The Stacking Ensemble improved performance to some extent but did not surpass the stability of linear models. Statistical tests confirmed significant differences among models, indicating that model selection has a meaningful impact on predictive accuracy. The findings suggest that simpler regularized models are more reliable in small-data environments and highlight the importance of selecting models appropriate to data characteristics. This study can help to appropriately select the prediction methodology required for companies with small data. In future research, we will study machine learning-based prediction accuracy improvement measures suitable for SMEs.

2

Analysis of machine learning and deep learning prediction models for sepsis and neonatal sepsis: A systematic review

A. Safiya Parvin, B. Saleena

[NRF 연계] 한국통신학회 ICT Express Vol.9 No.6 2023.12 pp.1215-1225

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Sepsis and Neonatal sepsis are major challenges in global healthcare because they cause life-threatening organ dysfunction in intensive care adult and pediatric patients due to downregulated host response to a particular infection. Early clinical identification of sepsis is difficult, and failure to provide prompt treatment can often lead to crucial stages and increase the rates of fatality. Thus an intense study is needed to determine and categorize sepsis in its initial stage. The complexity of varying clinical statistics makes it difficult to attain a precise definition in pediatrics. The advanced Machine Learning (ML) and Deep Learning (DL) technologies in the implementation of protocols show promising real-time models for predicting sepsis at the primary stage and thereby reducing the mortality rate. This review article contemplates the complete list of procedures through which sepsis and neonatal sepsis are speculated by ML and DL and concentrates specifically on data available in the adult emergency care unit as well as the neonatal intensive care unit. The survey process was carried out by searching terms related to ML and DL merged with topics concerning sepsis and neonatal sepsis. The literature analysis was carried out from Scopus, Web of Science, and PubMed databases for the period from 2015 to 2022. The assessment of the risk of bias was carried out for the eleven selected papers using the Prediction Model Risk of Bias Assessment Tool (PROBAST). The eleven papers were selected from different medical care units based on the performance measure AUROC, which ranges from 0.68 to 0.95. Five papers involving ML/DL models reduce the bias and lessen risk occurrence. Five papers generate an increase in bias but can be applied to new data. One paper works with above twenty-five features has high-risk probability but predicts patients within 5?6 h in the future. This survey portrays the role of prediction models that supports the researchers and clinicians for better decision-making and antibiotic administration at an earlier stage.

3

4,000원

본 논문은 수중 IoT 네트워크에서 센서의 전력 소비를 줄이고 네트워크의 처리량을 향상하는 수중 링크 적응 방법을 제안한다. 링크 적응 방법의 하나인 AMC(Adaptive Modulation and Coding) 기술은 SNR(Signal Noise Rate)과 BER(Bit Error Rate)의 강한 상관관계를 이용하지만, 수중에 바로 적용하는 것은 어렵다. 따라서 수중 환경에 적합한 머신러닝 기반의 AMC 기술을 제안한다. 제안하는 MCS(Modulation Coding and Scheme) 예측 모델은 수중 채널 환경에서 목표 BER 값을 달성하기 위한 통신 방법을 예측한다. 예측된 통신 방법을 실제 수중 무선 통신에서 적용하는 것은 현실적으로 어렵기 때문에 본 논문에서는 높은 정확도의 BER 예측 모델을 사용 해 MCS 예측 모델의 성능을 확인한다. 결과적으로 제안하는 AMC 기술은 통신 성공 확률을 올림으로써 머신러닝 의 적용 가능성을 확인시켰다.

This paper proposes a link adaptation method for Underwater Internet of Things (IoT), which reduces power consumption of sensor nodes and improves the throughput of network in underwater IoT network. Adaptive Modulation and Coding (AMC) technique is one of link adaptation methods. AMC uses the strong correlation between Signal Noise Rate (SNR) and Bit Error Rate (BER), but it is difficult to apply in underwater IoT as it is. Therefore, we propose the machine learning based AMC technique for underwater environments. The proposed Modulation Coding and Scheme (MCS) prediction model predicts transmission method to achieve target BER value in underwater channel environment. It is realistically difficult to apply the predicted transmission method in real underwater communication in reality. Thus, this paper uses the high accuracy BER prediction model to measure the performance of MCS prediction model. Consequently, the proposed AMC technique confirmed the applicability of machine learning by increase the probability of communication success.

4

5,100원

본 연구는 웨어러블 기기에서 수집된 라이프로그 데이터를 활용하여 고령화 사회에서 증가하고 있는 치매를 조기에 진단하여 관리할 수 있는 예측 모델을 개발하고, 이를 기반으로 한 상업적 활용 전략을 제안하는 것을 목표로 하였다. 이 연구는 전문의의 병리진단을 기반으로 한 60~80대 174명의 대상자로부터 수집된 12,184개의 라이프로그 정보(수면 및 활동 정보)와 치매 진단 데이터를 활용하였다. 연구 과정에서 수면과 활동 데이터를 포함하는 다차원적인 데이터셋을 표준화 하였고 다양한 머신러닝 알고리즘으로 분석하였으며, 가장 높은 ROC-AUC점수를 보여준 랜덤 포레스트 모델이 가장 우수한 성능을 보였다. 또한 ablation test를 통해 수면과 관련된 변수들과 활동과 관련 변수들의 제외가 모델 예측력에 미치는 영향을 평가하였고, 이러한 변수들이 모델의 예측력에 유의미한 영향력을 가지고 있음을 확인하였다. 마지막으로, 개발된 모델의 상업적 활용 전략의 가능성을 탐구함으로써, 치매 예방 시스템의 상업적 확산을 위한 새로운 방향을 제안하였다.

This study aimed to propose early diagnosis and management of dementia, which is increasing in aging societies, and suggest commercial utilization strategies by leveraging digital healthcare technologies, particularly lifelog data collected from wearable devices. By introducing new approaches to dementia prevention and management, this study sought to contribute to the field of dementia prediction and prevention. The research utilized 12,184 pieces of lifelog information (sleep and activity data) and dementia diagnosis data collected from 174 individuals aged between 60 and 80, based on medical pathological diagnoses. During the research process, a multidimensional dataset including sleep and activity data was standardized, and various machine learning algorithms were analyzed, with the random forest model showing the highest ROC-AUC score, indicating superior performance. Furthermore, an ablation test was conducted to evaluate the impact of excluding variables related to sleep and activity on the model's predictive power, confirming that regular sleep and activity have a significant influence on dementia prevention. Lastly, by exploring the potential for commercial utilization strategies of the developed model, the study proposed new directions for the commercial spread of dementia prevention systems.

5

기계학습을 활용한 장애인 취업 예측모형 탐색 및 분석

정대영, 염동문

[NRF 연계] 한국장애인고용공단 고용개발원 장애와 고용 Vol.34 No.3 2024.08 pp.23-42

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

연구문제 또는 목적: 본 연구는 기계학습 알고리즘을 활용하여 장애인의 취업 여부를 예측하는 최적의 모델을 도출하고, 도출된 모델을 기반으로 장애인의 취업 여부 예측에 기여하는 변인들을 분석하고자 하였다. 연구방법: 장애인고용패널조사 2차 웨이브 7차년도 자료에서 3,328명을 분석대상으로 장애인의 취업 여부에 영향을 미치는 64개의 설명변수를 선정하였다. 기계학습은 Orange 3.36.2를 사용하여 9가지 알고리즘(로지스틱회귀분석, 의사결정나무, Random Forest, Gradient Boosting, AdaBoost, SVM, 인공신경망, Naive Bayes, kNN)을 적용하여 최적의 예측 모델을 결정하고, 예측에 기여하는 설명변수의 중요도를 분석하였다. 연구결과: 장애인의 취업 여부에 대한 최적의 예측 모델은 Gradient Boosting 이었으며, 중요도가 가장 높은 설명변수는 근로 가능정도로 나타났으며, 직업적 능력 요인과 가구 요인, 취업활동 요인, 취업제도 및 환경 요인이 중요한 것으로 분석되었다. 결론 또는 시사점: 장애인의 취업을 지원하고, 효과적인 취업지원 프로그램과 관련된 정책을 새롭게 개발하는 데 필요한 기초 자료로서의 역할을 할 수 있을 것으로 기대된다.

Purpose or Objectives : This study aims to derive the optimal model for predicting employment status of people with disabilities using machine learning algorithms and to analyze the contributing variables based on the derived model. Method: Sixty-four explanatory variables influencing the employment status of people with disabilities were selected. Machine learning was conducted using Orange 3.36.2, applying nine algorithms to determine the optimal prediction model and analyze the importance of explanatory variables contributing to the prediction. Results: The optimal prediction model for the employment status of people with disabilities was found to be Gradient Boosting. The most significant explanatory variable was the degree of work ability, with vocational ability factors, household factors, job-seeking activity factors, and employment system and environment factors also being important. Conclusion or Implications: The findings are expected to serve as foundational data necessary for supporting the employment of people with disabilities and developing effective employment support programs and related policies.

7

Could Machine Learning-Based Gestational Diabetes Mellitus Prediction Models Replace Traditional Screening Test?

황종윤

[NRF 연계] 한국모자보건학회 한국모자보건학회지 Vol.28 No.4 2024.10 pp.153-155

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

8

Recent advances in applications of machine learning in cervical cancer research: a focus on prediction models

Syed S Abrar, Seoparjoo Azmel Mohd Isa, Suhaily Mohd Hairon, Mohd Pazudin Ismail, Mohd Nasrullah Bin Nik Ab Kadir

[NRF 연계] 대한산부인과학회 Obstetrics & Gynecology Science Vol.68 No.4 2025.07 pp.247-259

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Artificial intelligence (AI) and machine learning (ML) are transforming cervical cancer research and offering advancements in diagnosis, prognosis, screening, and treatment. This review explores ML applications with particular emphasis on prediction models. A comprehensive literature search identified studies using ML for survival prediction, risk assessment, and treatment optimization. ML-driven prognostic models integrate clinical, histopathological, and genomic data to improve survival prediction and patient stratification. Screening methods, including deep-learning-based cytology analysis and human papillomavirus detection, enhance accuracy and efficiency. ML-driven imaging techniques facilitate early and precise cancer diagnosis, while risk prediction models assess susceptibility based on demographic and genetic factors. AI also optimizes treatment planning by predicting therapeutic responses and guiding personalized interventions. Despite significant progress, challenges remain regarding data availability, model interpretability, and clinical implementation. Standardized datasets, external validation, and cross-disciplinary collaborations are crucial for implementing ML innovations in clinical settings. Subsequent investigations should prioritize joint initiatives among data scientists, healthcare providers, and health authorities to translate AI innovations into real-world applications and to enhance the impact of ML on cervical cancer care. By synthesizing recent developments, this review highlights the potential of ML to improve clinical outcomes and shaping the future of cervical cancer management.

9

Metabolic Syndrome Prediction Using Machine Learning Models with Genetic and Clinical Information from a Nonobese Healthy Population

Choe, Eun Kyung, Rhee, Hwanseok, Lee, Seungjae, Shin, Eunsoon, Oh, Seung-Won, Lee, Jong-Eun, Choi, Seung Ho

[Kisti 연계] 한국유전체학회 Genomics & informatics Vol.16 No.4 2018 pp.31-32

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The prevalence of metabolic syndrome (MS) in the nonobese population is not low. However, the identification and risk mitigation of MS are not easy in this population. We aimed to develop an MS prediction model using genetic and clinical factors of nonobese Koreans through machine learning methods. A prediction model for MS was designed for a nonobese population using clinical and genetic polymorphism information with five machine learning algorithms, including naïve Bayes classification (NB). The analysis was performed in two stages (training and test sets). Model A was designed with only clinical information (age, sex, body mass index, smoking status, alcohol consumption status, and exercise status), and for model B, genetic information (for 10 polymorphisms) was added to model A. Of the 7,502 nonobese participants, 647 (8.6%) had MS. In the test set analysis, for the maximum sensitivity criterion, NB showed the highest sensitivity: 0.38 for model A and 0.42 for model B. The specificity of NB was 0.79 for model A and 0.80 for model B. In a comparison of the performances of models A and B by NB, model B (area under the receiver operating characteristic curve [AUC] = 0.69, clinical and genetic information input) showed better performance than model A (AUC = 0.65, clinical information only input). We designed a prediction model for MS in a nonobese population using clinical and genetic information. With this model, we might convince nonobese MS individuals to undergo health checks and adopt behaviors associated with a preventive lifestyle.

10

Do Machine Learning Methods Outperform Traditional Statistical Models in Crime Prediction? A Comparison Between Logistic Regression and Neural Networks

나종민, 오경석, 송주영, 박형아

[NRF 연계] 서울대학교 행정대학원 Journal of Policy Studies Vol.36 No.1 2021.03 pp.1-13

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Although machine learning (ML) methods have recently gained popularity in both academia and industry as alternative risk assessment tools for efficient decision-making, inconsistent patterns are observed in the existing literature regarding their competitiveness and utility in predicting various outcomes. Drawing on a sample of the general youth population in the U.S., we compared the predictive accuracy of logistic regression (LR) and neural networks (NNs), which are the most widely applied approaches in conventional statistics and contemporary ML methods, respectively, by adopting many theoretically relevant predictors of the future arrest outcome. Even after fully implementing rigorous ML protocols for model tuning and up-sampling and down-sampling procedures recommended in recent literature to optimize learning algorithms, NNs did not yield substantially improved performance over LR if we still rely on a conventional dataset with relatively small sample sizes and a limited number of predictors. Nonetheless, we encourage more rigorous, comprehensive, and diverse evaluation research for a complete understanding of the ML potential in predictive capacity and the contingencies in which modern ML methods can perform better than conventional parametric statistical models.

11

Comparison of Models for the Prediction of Medical Costs of Spinal Fusion in Taiwan Diagnosis-Related Groups by Machine Learning Algorithms

Ching-Yen Kuo, Liang-Chin Yu, Hou-Chaung Chen, Chien-Lung Chan

[NRF 연계] 대한의료정보학회 Healthcare Informatics Research Vol.24 No.1 2018.01 pp.29-37

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Objectives: The aims of this study were to compare the performance of machine learning methods for the prediction of themedical costs associated with spinal fusion in terms of profit or loss in Taiwan Diagnosis-Related Groups (Tw-DRGs) andto apply these methods to explore the important factors associated with the medical costs of spinal fusion. Methods: A dataset was obtained from a regional hospital in Taoyuan city in Taiwan, which contained data from 2010 to 2013 on patients ofTw-DRG49702 (posterior and other spinal fusion without complications or comorbidities). Naive-Bayesian, support vectormachines, logistic regression, C4.5 decision tree, and random forest methods were employed for prediction using WEKA3.8.1. Results: Five hundred thirty-two cases were categorized as belonging to the Tw-DRG49702 group. The mean medicalcost was US $4,549.7, and the mean age of the patients was 62.4 years. The mean length of stay was 9.3 days. The lengthof stay was an important variable in terms of determining medical costs for patients undergoing spinal fusion. The randomforest method had the best predictive performance in comparison to the other methods, achieving an accuracy of 84.30%,a sensitivity of 71.4%, a specificity of 92.2%, and an AUC of 0.904. Conclusions: Our study demonstrated that the randomforest model can be employed to predict the medical costs of Tw-DRG49702, and could inform hospital strategy in terms ofincreasing the financial management efficiency of this operation.

12

Two Machine Learning Models for Mobile Phone Battery Discharge Rate Prediction Based on Usage Patterns

Chantrapornchai, Chantana, Nusawat, Paingruthai

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.12 No.3 2016 pp.436-454

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This research presents the battery discharge rate models for the energy consumption of mobile phone batteries based on machine learning by taking into account three usage patterns of the phone: the standby state, video playing, and web browsing. We present the experimental design methodology for collecting data, preprocessing, model construction, and parameter selections. The data is collected based on the HTC One X hardware platform. We considered various setting factors, such as Bluetooth, brightness, 3G, GPS, Wi-Fi, and Sync. The battery levels for each possible state vector were measured, and then we constructed the battery prediction model using different regression functions based on the collected data. The accuracy of the constructed models using the multi-layer perceptron (MLP) and the support vector machine (SVM) were compared using varying kernel functions. Various parameters for MLP and SVM were considered. The measurement of prediction efficiency was done by the mean absolute error (MAE) and the root mean squared error (RMSE). The experiments showed that the MLP with linear regression performs well overall, while the SVM with the polynomial kernel function based on the linear regression gives a low MAE and RMSE. As a result, we were able to demonstrate how to apply the derived model to predict the remaining battery charge.

13

Comparison of intracranial pressure prediction in hydrocephalus patients among linear, non-linear, and machine learning regression models in Thailand

Trakulpanitkit Avika, Tunthanathip Thara

[NRF 연계] 대한중환자의학회 Acute and Critical Care Vol.38 No.3 2023.08 pp.362-370

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Background: Hydrocephalus (HCP) is one of the most significant concerns in neurosurgical patients because it can cause increased intracranial pressure (ICP), resulting in mortality and morbidity. To date, machine learning (ML) has been helpful in predicting continuous outcomes. The primary objective of the present study was to identify the factors correlated with ICP, while the secondary objective was to compare the predictive performances among linear, non-linear, and ML regression models for ICP prediction. Methods: A total of 412 patients with various types of HCP who had undergone ventriculostomy was retrospectively included in the present study, and intraoperative ICP was recorded following ventricular catheter insertion. Several clinical factors and imaging parameters were analyzed for the relationship with ICP by linear correlation. The predictive performance of ICP was compared among linear, non-linear, and ML regression models. Results: Optic nerve sheath diameter (ONSD) had a moderately positive correlation with ICP (r=0.530, P<0.001), while several ventricular indexes were not statistically significant in correlation with ICP. For prediction of ICP, random forest (RF) and extreme gradient boosting (XGBoost) algorithms had low mean absolute error and root mean square error values and high R2 values compared to linear and non-linear regression when the predictive model included ONSD and ventricular indexes. Conclusions: The XGBoost and RF algorithms are advantageous for predicting preoperative ICP and establishing prognoses for HCP patients. Furthermore, ML-based prediction could be used as a non-invasive method.

14

의사 결정과 빅데이터-기계학습 예측 모형: 좋은 예측에는 인과적 지식이 반드시 필요할까?

허원기

[NRF 연계] 한국과학철학회 과학철학 Vol.27 No.2 2024.07 pp.27-62

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

실천적 의사 결정의 맥락에서 예측은 합리적 선택을 뒷받침하는 정보를 제공한다. 빅데이터 기반의 기계학습 모형(이하 ML 예측 모형)은 다양한 분야의 예측 과제에서 획기적인 진전을 이루었다고 평가받는다. 그러나 이러한 ML 예측 모형을 사용하는 것에 대한 비판도 상당하다. 특히 ML 예측 모형이 인과관계가 아닌 상관관계에 의존한다는 점은 ML 예측 모형의 예측이 진정으로 좋은 예측인지 질문하게 만든다. 본 논문에서는 예측 문제를 두 종류로 구분하여 이 질문에 답할 것이다. 의사 결정에 필요한 예측은 ‘사전 개입을 위한 예측’과 ‘대응 개입을 위한 예측’이라는 둘로 나눌 수 있는데, 이 구분되는 두 예측은 의사 결정에 다른 방식으로 기여한다. 전자의 경우에는 좋은 예측을 위해서 인과 지식이 강하게 요구되지만, 후자의 경우에는 상관 지식만으로도 좋은 예측에 충분할 수 있다. 따라서, 상관관계에 기초하는 ML 예측 모형을 사용하는 것은 대응 개입의 맥락에서 충분히 좋은 예측 활동을 구성할 수 있다.

In the context of practical decision-making, predictions provide information that supports rational choices. Big data-based machine learning models (hereafter ML prediction models) have been acclaimed for making breakthrough advancements in various predictive tasks. However, considerable criticism exists regarding using these ML prediction models. Specifically, the fact that ML prediction models depend on correlation rather than causation raises the question of whether their predictions can be genuinely good. This paper will address this question by distinguishing prediction problems into two types. Predictions for decision-making can be divided into ‘predictions for prior intervention’ and ‘predictions for reactive intervention’. These two distinct types of predictions contribute to decision-making in different ways. For prior intervention, causal knowledge is strongly required for good predictions, whereas for reactive intervention, correlational knowledge alone can suffice for good predictions. Therefore, using ML prediction models based on correlations can constitute good prediction activities in the context of reactive interventions.

15

데이터 전처리를 고려한 하수처리장 머신러닝 모델 개발

심규대, 김효상, 박찬수, 김동균, 김신걸

[Kisti 연계] 한국수자원학회 한국수자원학회 학술대회논문집 2023 p.495

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 하수처리장 운영시스템 자료를 활용하여, 머신러닝 기반의 예측 모델을 개발하고, 모델 정확도 향상에 대하여 검토하였다. 하수처리장에 설치된 각종 센서를 통해 실시간으로 자료가 모니터링되고 있으며, 수집된 자료는 운영시스템에 저장된다. 하수처리장 시스템은 설정된 값과 센서의 측정값을 비교해 이상치가 발생하면 운영자가 즉각적으로 조치하여 문제를 해결하고 있으나, 비정상적인 상황 발생시 이를 대처할 시간이 부족하여 적절한 조치가 이루어지지 못하는 경우가 발생 되고 있다. 따라서, 이러한 문제점을 해결하기 위해 A 하수처리장 운영자료를 활용하여 결과 예측이 신속하고 신뢰도 높은 머신러닝 기반의 예측 모델을 개발하고자 하였다. 모델의 예측 정확도 및 신뢰성을 향상하기 위하여 결과에 영향을 미치는 주요 영향 인자를 분석하고, 이를 기반으로 모델의 추가 분석 및 개선을 수행하여 모델의 예측력을 평가하였다. 금회 연구는 데이터 전처리를 과정을 통한 인사이트를 도출하고 이를 활용하여 하수처리장 운영자료 예측 정확도를 높일 수 있었으며, 이 결과를 바탕으로 다른 하수처리장의 모델 개발시에도 유용하게 활용이 가능할 것으로 검토되었다.

16

머신러닝 기반 대학생 중도 탈락 예측 모델의 성능 비교

정석봉, 김두연

[Kisti 연계] 한국시뮬레이션학회 한국시뮬레이션학회논문지 Vol.32 No.4 2023 pp.19-26

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

전국 대학생의 중도 탈락 비율의 증가는 학생 개인 뿐만 아니라 대학과 사회에 심각한 부정적 영향을 끼친다. 본 연구에서는 중도 탈락이 예상되는 학생을 사전에 식별하기 위하여, 각 대학의 학사관리 시스템에서 손쉽게 얻을 수 있는 학적 데이터를 기반으로 머신러닝 분야의 결정트리, 랜덤 포레스트, 로지스틱 회귀 및 딥러닝 기반의 중도 탈락 예측 모델을 구축하고, 그 성능을 비교·분석하였다. 분석 결과 로지스틱 회귀 기반 예측 모델의 재현율이 가장 높았으나 f-1 및 auc 값이 낮은 한계를 보였고, 랜덤 포레스트 기반의 예측 모델의 경우 재현율을 제외한 다른 모든 지표에서 가장 우수한 성능을 보였다. 또한 예측 기간에 따른 예측 모델의 성능을 확인하기 위하여 예측 기간을 단기(1개 학기 이내), 중기(2개 학기 이내) 및 장기(3개 학기 이내)로 나누어 분석해 본 결과, 장기 예측 시 가장 높은 예측력을 보였다. 본 연구를 통해 각 대학은 중도 탈락이 예상되는 학생들을 조기에 식별하고, 이들에 대한 집중 관리를 통해 중도 탈락 비율을 줄이며 나아가 대학 재정 안정화에 기여할 수 있을 것으로 기대된다.

The increase in the dropout rate of college students nationwide has a serious negative impact on universities and society as well as individual students. In order to proactive identify students at risk of dropout, this study built a decision tree, random forest, logistic regression, and deep learning-based dropout prediction model using academic data that can be easily obtained from each university's academic management system. Their performances were subsequently analyzed and compared. The analysis revealed that while the logistic regression-based prediction model exhibited the highest recall rate, its f-1 value and ROC-AUC (Receiver Operating Characteristic - Area Under the Curve) value were comparatively lower. On the other hand, the random forest-based prediction model demonstrated superior performance across all other metrics except recall value. In addition, in order to assess model performance over distinct prediction periods, we divided these periods into short-term (within one semester), medium-term (within two semesters), and long-term (within three semesters). The results underscored that the long-term prediction yielded the highest predictive efficacy. Through this study, each university is expected to be able to identify students who are expected to be dropped out early, reduce the dropout rate through intensive management, and further contribute to the stabilization of university finances.

17

머신러닝 기반의 하수처리장 예측 모델 평가 및 개발

심규대, 김효상, 장근수, 김동균, 김영모

[Kisti 연계] 한국수자원학회 한국수자원학회 학술대회논문집 2023 p.499

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

최근 컴퓨터 성능 향상과 새로운 머신러닝 알고리즘 개발됨에 따라, 각 분야별 연구자들이 이를 활용한 연구를 다양하게 수행하고 있으며, 하수처리시설의 경우에는 막대한 양의 운영자료가 축척됨에 따라 머신러닝을 활용한 다양한 연구가 가속화 되고 있다. 기존 하수처리장의 물리학적 모델은 적용된 영향 인자에 여러 가지 가정이 고려되어 모델 정확도가 부정확해지는 경향이 있었으며, 이러한 문제점을 보완하기 위해 하수처리장의 수집된 운영자료 및 머신러닝 기반의 예측 모델을 활용하여 예측 모델 정확도를 향상하는 선행 연구들이 진행되고 있다. A 하수처리장의 부지 내에 설치된 센서를 통하여 운영자료가 중앙제어실 서버에 실시간으로 저장되는 자료를 활용하여 NN (Neural Network), SVM (Support Vector Machine), RF (Random Forest) 등과 같은 다양한 머신러닝 모델을 적용하였고, 하수처리장 운영자료를 적용할 경우 어느 모델이 가장 높은 성능이 나타나는지 인사이트를 도출하고자 하였다. 금회 연구는 A 하수처리장을 대상으로 여러 머신러닝 기반 예측 모델을 개발하고, 각 모델의 예측정확도를 서로 평가함으로써, 머신러닝 모델 최적화를 수행할 수 있었다. 이번 연구에서 도출된 결과를 활용하여 하수처리장 예측 모델 최적화를 진행할 경우, 향후 비교적 짧은 시간에 하수처리장 머신러닝 기반 예측 모델 개발이 가능하다는 점에 의의가 있다.

18

머신러닝과 딥러닝을 이용한 부동산 지수 예측 모델 비교

박수민, 이연재, 박주현, 박주아, 임진섭, 김현희

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2021 pp.1156-1159

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

수도권을 중심으로 한 부동산 가격 상승이 지속적으로 진행되고 있다. 한국은행에서는 기준금리 인상으로 과열된 부동산 시장의 안정을 바라고 있다. 하지만 기준금리 인상이 부동산 시장에 미치는 영향이 크지 않다고 보는 시각도 많다. 이에 본 논문에서는 머신러닝과 딥러닝을 이용하여 서울 지역의 부동산 매매지수를 예측하고 기준금리를 추가 변수로 이용하여 결과를 비교하였다. 실험 결과 선형적으로 증가 중인 시장 특성상 전통적 모델인 선형회귀가 우수한 성능을 보였으며, 기준 금리를 변수로 추가한 경우 예측력이 근소하게 증가하였으나 그 영향은 크지 않음을 볼 수 있었다.

19

온라인 뉴스와 거시경제 지표, 금융 지표, 기술적 지표, 관심도 지표를 이용한 코스닥 상장 기업의 기계학습 기반 주가 변동 예측

김화련, 홍승혜, 홍헬렌

[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.24 No.3 2021 pp.448-459

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, we propose a method of predicting the next-day stock price fluctuations of 10 KOSDAQ-listed companies in 5G, autonomous driving, and electricity sectors by training SVM, XGBoost, and LightGBM models from macroeconomic·financial market indicators, technical indicators, social interest indicators, and daily positive indices extracted from online news. In the three experiments to find out the usefulness of social interest indicators and daily positive indices, the average accuracy improved when each indicator and index was added to the models. In addition, when feature selection was performed to analyze the superiority of the extracted features, the average importance ranking of the social interest indicator and daily positive index was 5.45 and 1.08, respectively, it showed higher importance than the macroeconomic financial market indicators and technical indicators. With the results of these experiments, we confirmed the effectiveness of the social interest indicators as alternative data and the daily positive index for predicting stock price fluctuation.

20

머신러닝 기반의 온실 VPD 예측 모델 비교

장경민, 이명배, 임종현, 오한별, 신창선, 박장우

[Kisti 연계] 한국정보처리학회 정보처리학회논문지/소프트웨어 및 데이터 공학 Vol.12 No.3 2023 pp.125-132

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구에서는 식물의 영양분 흡수에 따른 식물 성장뿐만 아니라 기공 기능 및 광합성에도 영향을 끼치는 온실의 수증기압차(VPD, Vapor Pressure Deficit)예측을 위한 머신러닝 모델들의 성능을 비교해보았다. VPD 예측을 위해 온실 내·외부 환경요소 및 시계열 데이터의 시간적 요소들과의 상관관계를 확인하고 상관관계가 높은 요소들이 VPD에 어떤 영향을 미치는지 확인하였다. 예측 모델의 성능을 분석하기 전 분석 시계열 데이터의 양(1일, 3일, 7일), 간격(20분, 1시간)이 예측 성능에 미치는 영향을 확인하여 데이터의 양과 간격을 조절하였다. 마지막으로 4개의 머신러닝 예측 모델(XGB Regressor, LGBM Regressor, Random Forest Regressor 등)을 적용하여 모델별 예측 성능을 비교했다. 모델의 예측 결과로 20분 간격의 1일의 데이터를 사용했을 때 LGBM에서 MAE는 0.008, RMSE는 0.011의 가장 높은 예측 성능을 보였다. 또한 20분 후 VPD 예측에 가장 큰 영향을 미치는 요소는 환경적 요인보다는 과거 20분 전의 VPD(VPD_y__71)임을 확인하였다. 본 연구의 결과를 활용하여 VPD 예측을 통해 작물의 생산성을 높이고, 온실의 결로, 병 발생 예방 등이 가능하다. 향후 온실의 환경 데이터 예측뿐만 아니라 더 나아가 생산량 예측, 스마트팜 제어 모델 등 다양한 분야에 활용할 수 있을 것이다.

In this study, we compared the performance of machine learning models for predicting Vapor Pressure Deficits (VPD) in greenhouses that affect pore function and photosynthesis as well as plant growth due to nutrient absorption of plants. For VPD prediction, the correlation between the environmental elements in and outside the greenhouse and the temporal elements of the time series data was confirmed, and how the highly correlated elements affect VPD was confirmed. Before analyzing the performance of the prediction model, the amount and interval of analysis time series data (1 day, 3 days, 7 days) and interval (20 minutes, 1 hour) were checked to adjust the amount and interval of data. Finally, four machine learning prediction models (XGB Regressor, LGBM Regressor, Random Forest Regressor, etc.) were applied to compare the prediction performance by model. As a result of the prediction of the model, when data of 1 day at 20 minute intervals were used, the highest prediction performance was 0.008 for MAE and 0.011 for RMSE in LGBM. In addition, it was confirmed that the factor that most influences VPD prediction after 20 minutes was VPD (VPD_y__71) from the past 20 minutes rather than environmental factors. Using the results of this study, it is possible to increase crop productivity through VPD prediction, condensation of greenhouses, and prevention of disease occurrence. In the future, it can be used not only in predicting environmental data of greenhouses, but also in various fields such as production prediction and smart farm control models.

 
1 2
페이지 저장