Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 169
No
1

기계학습 응용 및 학습 알고리즘 성능 개선방안 사례연구 KCI 등재

이호현, 정승현, 최은정

한국디지털정책학회 디지털융복합연구 제14권 제2호 2016.02 pp.245-258

※ 기관로그인 시 무료 이용이 가능합니다.

4,600원

본 논문에서는 기계학습과 관련된 다양한 사례들에 대한 연구를 바탕으로 기계학습 응용 및 학습 알고리즘의 성능 개선 방안을 제시한다. 이를 위해 기계학습 기법을 적용하여 결과를 얻어낸 문헌을 자료로 수집하고 학문 분야로 나누어 각 분야에서 적합한 기계학습 기법을 선택 및 추천하였다. 공학에서는 SVM, 의학에서는 의사결정나무, 그 외 분야에서는 SVM이 빈번한 이용 사례와 분류/예측의 측면에서 그 효용성을 보였다. 기계학습의 적용 사례 분석을 통해 응용 방안의 일반적 특성화를 꾀할 수 있었다. 적용 단계는 크게 3단계로 이루어진다. 첫째, 데이터 수집, 둘째, 알고리즘을 통한 데이터 학습, 셋째, 알고리즘에 대한 유의미성 테스트 이며, 각 단계에서의 알고리즘의 결합을 통해 성능을 향상시킨다. 성능 개선 및 향상의 방법은 다중 기계학습 구조 모델링과 +α 기계학습 구조 모델링 등으로 분류한다.

This paper aims to present the way to bring about significant results through performance improvement of learning algorithm in the research applying to machine learning. Research papers showing the results from machine learning methods were collected as data for this case study. In addition, suitable machine learning methods for each field were selected and suggested in this paper. As a result, SVM for engineering, decision-making tree algorithm for medical science, and SVM for other fields showed their efficiency in terms of their frequent use cases and classification/prediction. By analyzing cases of machine learning application, general characterization of application plans is drawn. Machine learning application has three steps: (1) data collection; (2) data learning through algorithm; and (3) significance test on algorithm. Performance is improved in each step by combining algorithm. Ways of performance improvement are classified as multiple machine learning structure modeling, +α machine learning structure modeling, and so forth.

2

Supervised Machine Learning for Frailty Classification using Physical Performance Measures in Older Adults

Si-hyun Kim

[NRF 연계] KEMA학회 Journal of Musculoskeletal Science and Technology Vol.9 No.1 2025.06 pp.36-43

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Background Frailty is an important condition to detect in its early stages to prevent progression to more severe stages in older adults. Age-related declines in physical performance are strongly associated with frailty. Purpose This study aims to develop a frailty classification model by comparing the performance of machine learning models based on physical performance measures in community-dwelling older adults. Study design A cross-sectional study Methods Physical performance data were collected from older adults aged ≥65 years. Frailty classification models were developed using logistic regression, support vector machine (SVM), K-nearest neighbors (KNN), decision tree, and random forest. Clinical features including short physical performance battery, single-leg stance, SARC-F, body mass index, and mini-mental state examination (MMSE) were used as input variables for model development. The performance of each model was evaluated using accuracy, sensitivity, specificity, precision, F1-score, and area under the receiver operating characteristic curve (AUC). Permutation feature importance was employed to identify key predictors of frailty. Results The KNN model demonstrated the highest classification performance, achieving an accuracy of 0.93, an F1-score of 0.95, and an AUC of 0.86, indicating its suitability for frailty assessment. The logistic regression model achieved an accuracy of 0.86, an F1-score of 0.89, and an AUC of 0.98. The random forest model showed similar results, with an accuracy of 0.86, an F1-score of 0.88, and an AUC of 0.96. The SVM model recorded an accuracy of 0.79, an F1-score of 0.84, and an AUC of 0.80. The decision tree model showed the lowest performance, with an accuracy of 0.71, an F1-score of 0.78, and an AUC of 0.64. Feature importance analysis revealed that MMSE and SARC-F were the most influential predictors in the KNN model. Conclusions This study demonstrates that KNN is well-suited for identifying subtle variations in physical function that contribute to frailty. The results highlight its potential for clinical implementation in automated frailty screening. Feature importance analysis provides insight into key predictors, supporting personalized assessment strategies. However, due to the small sample size, further research is needed to assess the generalizability of frailty classification models in larger populations.

3

Identifying machine learning techniques for classification of target advertising

Jin-A Choi, Kiho Lim

[NRF 연계] 한국통신학회 ICT Express Vol.6 No.3 2020.09 pp.175-180

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

There have been numerous applications of artificial intelligence (AI) technologies to online advertising, especially to optimize the reach of target audiences. Previous studies show that improved computational power significantly advances granular audience targeting capabilities. This study investigates and classifies various machine learning techniques that are used to enhance targeted online advertising. Twenty-three machine learning-based online targeted advertising strategies are identified and classified largely into two categories, user-centric and content-centric approaches. The paper also identifies an underexamined area, algorithm-based detection of click frauds, to illustrate how machine learning approaches can be integrated to preserve the viability of online advertising.

4

Evaluation of the Machine Learning Classifier in Wafer Defects Classification

Jessnor Arif Mat Jizat, Anwar P.P. Abdul Majeed, Ahmad Fakhri Ab. Nasir, Zahari Taha, Edmund Yuen

[NRF 연계] 한국통신학회 ICT Express Vol.7 No.4 2021.12 pp.535-539

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, an evaluation of machine learning classifiers to be applied in wafer defect detection is described. The objective is to establish the best machine learning classifier for Wafer Defect Detection application. k-Nearest Neighbours (k-NN), Logistic Regression, Stochastic Gradient Descent, and Support Vector Machine were evaluated with 3 defects categories and one non-defect category. The key metrics for the evaluation are classification accuracy, classification precision and classification recall. 855 images were used to train, test and validate the classifier. Each image went through the embedding process by InceptionV3 algorithms before the evaluated classifier classifies the images.

5

A comparative study of classification and prediction of Cardio-Vascular Diseases (CVD) using Machine Learning and Deep Learning techniques

M. Swathy, K. Saruladha

[NRF 연계] 한국통신학회 ICT Express Vol.8 No.1 2022.03 pp.109-116

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Cardio-Vascular Diseases (CVD) is found to be rampant in the populace leading to fatal death. The statistics of a recent survey reports that the mortality rate is expanding due to obesity, cholesterol, high blood pressure and usage of tobacco among the people. The severity of the disease is piling up due to the above factors. Studying about the variations of these factors and their impact on CVD is the demand of the hour. This necessitates the usage of modern techniques to identify the disease at its outset and to aid a markdown in the mortality rate. Artificial Intelligence and Data Mining domains have a research scope with their enormous techniques that would aassist in the prediction of the CVD priory and identify their behavioural patterns in the large volume of data. The results of these predictions will help the clinicians in decision making and early diagnosis, which would reduce the risk of patients becoming fatal. This paper compares and reports the various Classification, Data Mining, Machine Learning, Deep Learning models that are used for prediction of the Cardio-Vascular diseases. The survey is organized as threefold: Classification and Data Mining Techniques for CVD, Machine Learning Models for CVD and Deep Learning Models for CVD prediction. The performance metrics used for reporting the accuracy, the dataset used for prediction and classification, and the tools used for each category of these techniques are also compiled and reported in this survey.

6

Novel Method of Classification in Knee Osteoarthritis: Machine Learning Application Versus Logistic Regression Model

Jung Ho Yang, Jae Hyeon Park, Seong-Ho Jang, Jaesung Cho

[NRF 연계] 대한재활의학회 Annals of Rehabilitation Medicine Vol.44 No.6 2020.12 pp.415-427

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Objective To present new classification methods of knee osteoarthritis (KOA) using machine learning and compare its performance with conventional statistical methods as classification techniques using machine learning have recently been developed. Methods A total of 84 KOA patients and 97 normal participants were recruited. KOA patients were clustered into three groups according to the Kellgren-Lawrence (K-L) grading system. All subjects completed gait trials under the same experimental conditions. Machine learning-based classification using the support vector machine (SVM) classifier was performed to classify KOA patients and the severity of KOA. Logistic regression analysis was also performed to compare the results in classifying KOA patients with machine learning method. Results In the classification between KOA patients and normal subjects, the accuracy of classification was higher in machine learning method than in logistic regression analysis. In the classification of KOA severity, accuracy was enhanced through the feature selection process in the machine learning method. The most significant gait feature for classification was flexion and extension of the knee in the swing phase in the machine learning method. Conclusion The machine learning method is thought to be a new approach to complement conventional logistic regression analysis in the classification of KOA patients. It can be clinically used for diagnosis and gait correction of KOA patients.

7

Classification of self-care patterns in Korean adults with prediabetes using unsupervised machine learning: a secondary data analysis

조미경, 허명륜

[NRF 연계] 한국기초간호학회 Journal of korean biological nursing science Vol.27 No.4 2025.11 pp.586-597

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

PurposeThis study aimed to classify self-care patterns among Korean adults with prediabetes using an unsupervised machine learning approach. The classification was grounded in Orem’s Self-Care Theory, focusing on self-care demands, self-care agencies, and self-care behaviors. MethodsA secondary data analysis was conducted using the 2023 Korea National Health and Nutrition Examination Survey. Variables were selected and categorized according to the theoretical components of Orem’s model. Principal component analysis was applied for dimensionality reduction, followed by K-means clustering to identify distinct self-care pattern groups. All variables were standardized using min-max normalization. Group differences were examined using analysis of variance and the chi-square test. ResultsThree self-care pattern groups were identified: the high self-care performance group, the latent self-care risk group, and the self-care vulnerable group. These groups exhibited distinct profiles across self-care demands, agencies, and behaviors. Significant intergroup differences were also observed in education level, income, health literacy, fasting blood glucose, and hemoglobin A1c levels. ConclusionSelf-care patterns among adults with prediabetes can be effectively classified through unsupervised learning techniques. The findings highlight the importance of developing tailored nursing interventions that consider multidimensional self-care profiles. This study underscores the applicability of Orem’s Self-Care Theory and demonstrates the potential of machine learning in identifying at-risk subgroups for early intervention.

8

4,000원

Automated Search Engine Optimization (SEO) is crucial for streamlining processes, ensuring consistency, and adapting to changes, thereby enhancing a website's overall success and visibility in the competitive online landscape. This research introduces a dataset and a baseline method for classifying website SEO ranks into three categories. Using 26 keywords, data was collected from 780 web pages across various Google rankings, and 36 ranking factors were employed to predict their rank. Key considerations for webpage preparation include anchor text, backlinks, Ref Domain, unique visits, and text length. The Random Forest model exhibited superior performance, achieving an average accuracy of 72% in predicting actual search rankings. The significance of this automated approach lies in identifying web pages requiring SEO improvements, leading to enhanced search engine rankings.

9

Diabetes classification model using machine learning

Doo-Bin Kim, Han-Joon Yang, Joo-Wan Hong

대한디지털의료영상학회 대한디지털의료영상학회논문지 Volume 26 Number 1 2024.04 pp.1-6

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

당뇨병은 혈당 수치가 오랜 시간 지속되는 대사질환이다. 당뇨병은 자체 질환에 대한 위험성과 더불어 심혈관과 뇌혈 관, 망막병증, 족부병변 등 질환 다양한 합병증을 유발하는 질병이다. 이렇게 발생하는 합병증으로 인해 사망률을 높이는 원인이 되고 있다. 이에 따라 당뇨병 조기발견과 합병증에 대한 치료가 수반되어야하며, 이를 위해 빠른 진단이 필요하 다. 이에 본 연구에서는 당뇨병 환자 데이터를 이용하여 인공지능 기법 중 로지스틱 회귀분석 머신러닝 알고리즘을 이용한 당뇨병 분류에 대한 머신러닝 모델을 구축하고자 한다. 머신러닝 모델은 scikit-learn을 이용하여 학습용 데이터 80%, 검증용 데이터 20%로 구분하고 교차검증을 5로 나누어 모델을 구성하였다. 로지스틱 회귀분석을 이용한 머신러닝 결과 정확도는 평균 92.24%, AUC 0.95, Recall 0.9294, Precision 0.927, F1 score 0.9248, Kappa 0.6983 이었으 며, 허리-엉덩이 둘레비율이 당뇨병 분류에 가장 큰 영향을 주는 변수였다. 간단한 모델을 통해 높은 성능의 모델을 구축할 수 있었으며, 개인 생활습관 등에 관한 변수를 추가한다면, 더욱 우수한 성능의 머신러닝 모델을 구축할 수 있을 것으로 사료된다.

Diabetes is a metabolic disease in which blood sugar levels last for a long time. Diabetes is a disease that causes various complications of diseases such as cardiovascular and cerebrovascular, retinopathy, and diabetes foot ulcer in addition to the risk of its own disease. Complications that occur in this way are responsible for increasing the mortality rate. Accordingly, early detection of diabetes and treatment for complications should be accompanied, and for this, rapid diagnosis is required. In this study intends to build a machine learning model for diabetes classification using logistic regression machine learning algorithm among artificial intelligence techniques using diabetes patient data. Machine learning model was divided into 80% of training data and 20% of verification data using scikit-learn, and the cross-validation was divided into 5 to construct a model. Accuracy of machine learning results using logistic regression analysis was 92.24% on average, AUC 0.95, Recall 0.9294, Precision 0.927, F1 score 0.9248, and Kappa 0.6983, and the waist-hip circumference ratio was the most influential variable for diabetes classification. A high-performance model could be built through a simple model, and if variables related to personal lifestyle are added, it is believed that a better performance machine learning model can be built.

10

4,000원

비만은 세계보건기구에서 질병으로 규정하고 있으며, 신체 내외부적인 영향을 나타낸다. 본 연구는 머신러닝을 이용 하여 생활패턴에 따른 비만도를 예측하고자 한다. 머신러닝에 이용한 데이터는 오픈데이터를 사용하였으며, 머신러닝 모델은 구글 코랩, 파이썬을 이용하고 모델 구성은 Light Gradient Boosting Machine, Extreme Gradient Boosting, Decision Tree Classifier, K Neighbors Classifier, Naive bayes 총 5개의 모델로 구성하였다. 각 모델 성능 평가 지표는 정확도, area under curve, 재현율, 정밀도, F1-score로 평가하였다. 해당 데이터를 활용한 5개 모델 중 Light Gradient Boosting Machine이 모든 지표에서 성능이 가장 우수 했으며, 지표에 대한 결과는 정확도 0.9601, area under curve 0.9981, 재현율 0.9601, 정밀도 0.9611, F1 score 0.9601이었다. 본 연구를 통해 기본적인 생활 패턴에 따른 비만도 예측을 통해 비만에 대한 사전 예방이 가능할 것으로 사료된다.

Obesity is defined as a disease by the World Health Organization and indicates internal and external influences on the body. This study aims to predict obesity according to lifestyle patterns using machine learning. The data used for machine learning used open data, and the machine learning model used Google Colab and Python. The model configuration consisted of a total of five models: Light Gradient Boosting Machine, Extreme Gradient Boosting, Decision Tree Classifier, K Neighbors Classifier, and Naive Bayes. The performance evaluation indices for each model were accuracy, area under curve, recall, precision, and F1-score. Among the five models using the data, Light Gradient Boosting Machine showed the best performance in all indices, and the results for the indices were accuracy 0.9601, area under curve 0.9981, recall 0.9601, precision 0.9611, and F1-score 0.9601. Through this study, it is believed that it will be possible to prevent obesity in advance by predicting obesity according to basic lifestyle patterns.

11

A Study on the Drug Classification Using Machine Learning Techniques KCI 등재후보

Anmol Kumar Singh, Ayush Kumar, Adya Singh, Akashika Anshum, Pradeep Kumar Mallick

중소기업융합학회 산업과 과학 제3권 제2호 2024.06 pp.8-16

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문에서는 인구통계학적, 생리학적 특성을 기반으로 환자에게 가장 적합한 약물을 예측하는 것을 목표로 하는 약물 분류 시스템을 제시한다. 데이터 세트에는 적절한 약물을 결정하기 위한 목적으로 연령, 성별, 혈압(BP), 콜레스테롤 수치, 나트륨 대 칼륨 비율(Na_to_K)과 같은 속성들이 포함된다. 본 연구에 사용된 모델은 KNN(K-Nearest Neighbors), 로지스틱 회귀 분석 및 Random Forest이다. 하이퍼파라미터를 최적화하기 위해 5겹 교차 검증을 갖춘 GridSearchCV를 활용하였으며, 각 모델은 데이터 세트에서 훈련 및 테스트 되었다. 초매개변수 조정 유무에 관계없이 각 모델의 성능은 정확도, 혼동 행렬, 분류 보고서와 같은 지표를 사용하여 평가되었다. GridSearchCV를 적용하지 않은 모델의 정확도는 0.7, 0.875, 0.975인 반면, GridSearchCV를 적용한 모델의 정확도는 0.75, 1.0, 0.975로 나타났다. GridSearchCV는 로지스틱 회귀 분석을 세 가지 모델 중 약물 분류에 가장 효과적인 모델로 식별했으며, K-Nearest Neighbors가 그 뒤를 이었고 Na_to_K 비율은 결과를 예측하는 데 중요한 특징인 것으로 밝혀졌다.

This paper shows the system of drug classification, the goal of this is to foretell the apt drug for the patients based on their demographic and physiological traits. The dataset consists of various attributes like Age, Sex, BP (Blood Pressure), Cholesterol Level, and Na_to_K (Sodium to Potassium ratio), with the objective to determine the kind of drug being given. The models used in this paper are K-Nearest Neighbors (KNN), Logistic Regression and Random Forest. Further to fine-tune hyper parameters using 5-fold cross-validation, GridSearchCV was used and each model was trained and tested on the dataset. To assess the performance of each model both with and without hyper parameter tuning evaluation metrics like accuracy, confusion matrices, and classification reports were used and the accuracy of the models without GridSearchCV was 0.7, 0.875, 0.975 and with GridSearchCV was 0.75, 1.0, 0.975. According to GridSearchCV Logistic Regression is the most suitable model for drug classification among the three-model used followed by the K-Nearest Neighbors. Also, Na_to_K is an essential feature in predicting the outcome.

12

5,100원

본 연구는 2023년 충청남도 홍성군 산불피해지 대상으로 Maximum Likelihood Classification(MLC), Random Forest(RF), Support Vector Machine(SVM), K-Nearest Neighbors(KNN) 이상 4가지 머신러닝 알고리 즘 모델의 산불피해 강도 분류 성능을 파악하기 위해 실시되었다. 검증데이터와 비교한 결과, 모델 총 성능은 Random Forest(RF)의 Overall Accuacy(OA)와 Cohen’s Kappa 계수가 97.2%와 0.95로 가장 높은 성능을 보였다. RF는 여러 개의 의사결정 나무들을 결합하여 최종 예측을 하므로 데이터의 불균형을 해소할 수 있다. 또한, 산불피해 강도에 따른 성능은 모든 머신러닝 알고리즘 모델에서 열해 피해지가 수관화 및 지표화 피해지에 비해 낮은 탐지 성능을 나타냈다. 이는 수관화ㆍ열해 피해지의 분광 반사율 패턴이 유사하고, 열해·지표화 피해지의 경우 복잡한 수관층 구조로 인해 탐지 성능이 낮은 것으로 판단된다. 이와 같은 머신러닝 알고리즘 모델을 통한 산불피해 강도별 면적을 산출은 신속한 복구계획 수립뿐만 아니라 산불로 인한 온실가스 배출량 산정에 기초자료를 제공할 수 있을 것이다.

This study aims to evaluate the wildfire severity classification performance of four machine learning algorithm models, Maximum Likelihood Classification(MLC), Random Forest(RF), Support Vector Machine(SVM) and K-Nearest Neighbors(KNN), in the wildfire areas of Hongseong-gun, Chungcheongnam-do, in 2023. Compared with the verification dataset, the overall performance of the models indicated that the Random Forest(RF) showed the highest accuracy. The highest performance of the Random Forest(RF) model can be explained by its ability to combine multiple decision trees for a final prediction. Furthermore, the performance based on wildfire severity showed that all machine learning algorithm models had lower detection performance in heat damage areas compared with crown and surface fire areas. Crown and heat damage areas have similar spectral reflectance patterns, and heat damage and surface fire areas are considered to have shown lower performance due to the complex canopy structure. Machine learning based burn severity estimation provide critical baselines for rapid recovery planning and greenhouse gas emissions calculations.

13

This study compares and analyzes the performance of LIME-based machine learning methods (Gaussian Naive Bayes (GNB), Highly-Efficient Logistic Regression (LR), Linear Support Vector Machine (SVM), and Triple-layer Neural Network (TNN)) using three medical datasets. High-dimensional data increases the likelihood of overfitting in learning algorithms due to the curse of dimensionality. To address this, LIME is utilized to compute the importance of key features contributing to the model's predictions. Based on this, features are selected. The LIME technique generates multiple samples by perturbing the data in the local region. Subsequently, a simple linear model is used to evaluate the impact of each feature on the predictions. Features with high importance derived from this process are selected for model retraining. As a result, it was confirmed that learning time could be reduced while maintaining or even improving performance with a smaller number of features. Consequently, by selecting necessary features, the curse of dimensionality issue is alleviated, and accuracy can be maintained or improved using fewer features in the Hepatitis C Prediction Dataset, Breast Cancer Wisconsin (Prognostic) Dataset, and Glioma Grading Clinical and Mutation Features Dataset.

14

In heterogeneous landscapes, high-resolution land cover classification is vital for planning, ecological monitoring, and green infrastructure management. This study evaluates the performance of five machine learning algorithms: Decision Tree (DT), Naïve Bayes, Support Vector Machine (SVM), Random Tree (RT), and K-Nearest Neighbors (KNN), using UAV multispectral imagery and object-based image analysis (OBIA). Five scenarios were designed to compare algorithm accuracy. Additionally, a sixth scenario applied the best-performing algorithm to a feature subset selected through Recursive Feature Elimination (RFE), to examine the effect of feature optimization. RT achieved the highest overall classification accuracy (76.75%) and Kappa coefficient (0.7066), while SVM showed limited performance in complex environments. Height features contributed most to accuracy improvements, followed by spectral and geometric features. In class-specific analysis, the Naïve Bayes algorithm yielded the highest Producer’s Accuracy (90.86%) for forest-type land cover but had a lower User’s Accuracy (70.41%), indicating overclassification. In contrast, RT showed more balanced performance (PA = 87.09%, UA = 85.71%), suggesting greater reliability. The results demonstrate the benefits of integrating algorithm selection with feature optimization to improve classification accuracy in complex settings. This approach provides methodological insights for fine-scale mapping of vegetated areas and supports future applications in landscape monitoring and urban green space assessment.

15

4,000원

Accuracy improvement of classification model becomes main research objective in various fields. Selecting important features and removing outliers of a dataset are two effective solutions for improving model accuracy. Information Gain is one of the feature selection methods that can be considered as a solution for selecting important features of a dataset. Information Gain selects the variable that maximizes the information gain, which in turn minimizes the entropy and best splits the dataset into groups for effective classification. Aside of selecting important feature, removing outlier is also necessary for improving accuracy of the classification model. Density-Based Spatial Clustering of Applications with Noise (DBSCAN) is one of the powerful outlier removal methods which can identify with significant accuracy the clusters of random shape and size in large databases corrupted with noise. Therefore, in this study, we propose the accuracy improvement of heart disease classification model using Information Gain and DBSCAN applied to various machine learning algorithms. One publicly available heart disease dataset (Cleveland) is utilized in this study to build the classification model. The results showed that after implementing Information Gain, the accuracy of the model applied to Gaussian Naïve Bayes, Logistic Regression, Multi-Layer Perceptron, Support Vector Machine, Decision Tree, Random Forest, and Extreme Gradient Boosting algorithms increases as much as 1.31% in average. The accuracy also increases when DBSCAN is applied to the model after utilizing Information Gain, with the number of improvements is around 0.62%.

16

머신러닝 분류기법을 활용한 신생 유튜버의 생존 및 수익창출에 관한 연구 KCI 등재

김호익, 김한민

한국경영정보학회 경영정보학연구 제25권 제2호 2023.05 pp.57-76

※ 기관로그인 시 무료 이용이 가능합니다.

5,500원

본 연구는 목적은 디지털 플랫폼인 YouTube에서 최근 채널을 만든 크리에이터와 유튜버의 성공 여부를 분류 분석을 통해 알아보고자 함이다. 이를 위하여 과학기술 카테고리의 유튜버 채널 실제 정보들을 바탕으로 평균 동영상 업로드 횟수, 평균 영상 길이, 선택 가능한 다국어 자막 개수, 운영 중인 다른 소셜 네트워크 채널의 정보를 식별하였다. 식별한 정보와 머신러닝 기법을 활용하여 초기 유튜버들의 성공 여부인 수익창출 여부를 분류 분석하였으며, 분석결과, 인공 신경망 알고리즘이 초기 유튜버의 성공 또는 실패를 예측하는 데 가장 정확한 결과를 제공하고 있음을 발견했다. 또한, 제시된 다섯 가지 요인은 분석결과 향상에 기여하는 것으로 나타났다. 본 연구는 유튜브를 시작하고자 하는 신규 개인 창업가, 현재 유튜브를 운영하고 있는 인플루언서, 이러한 디지털 플랫폼을 활용하고자 하는 기업들에게 디지털 플랫폼의 다양한 접근 방식과 활용 방향에 대해 제언한다.

This study classifies the success of creators and YouTubers who have created channels on YouTube recently, which is the most influential digital platform. Based on the actual information disclosure of YouTubers who are in the field of science and technology category, video upload cycle, video length, number of selectable multilingual subtitles, and information from other social network channels that are being operated, the success of YouTubers using machine learning was classified and analyzed, which is the closest to the YouTube revenue structure. Our findings showed that neural network algorithm provided the best performance to predict the success or failure of YouTubers. In addition, our five factors contributed to improve the performance of the classification. This study has implications in suggesting various approaches to new individual entrepreneurs who want to start YouTube, influencers who are currently operating YouTube, and companies who want to utilize these digital platforms. We discuss the future direction of utilizing digital platforms.

17

4,000원

대화시스템은 인간과 컴퓨터의 상호작용에 새로운 패러다임이 되고 있다. 자연어로써 상호작용함으로써 인간 은 보다 자연스럽고 편리하게 각종 서비스를 누릴 수 있게 되었다. 대화시스템의 구조는 일반적으로 음성 인식, 자연 어 이해, 문맥 파악 등의 여러 모듈의 파이프라인으로 이뤄지는데, 본 연구에서는 자연어 이해 모듈의 도메인 분류 문 제를 풀기 위해 convolutional neural network, random forest 등의 기계학습 모델을 비교하였다. 사람이 직접 태 깅한 총 7개 서비스 도메인 데이터에 대하여 각 문장의 도메인을 분류하는 실험을 수행하였고 random forest 모델 이 F1 score 0.97 이상으로 가장 높은 성능을 달성한 것을 보였다. 향후 다른 기계학습 모델들을 추가 실험함으로써 도메인 분류 성능 개선을 지속할 계획이다.

Dialog system is becoming a new dominant interaction way between human and computer. It allows people to be provided with various services through natural language. The dialog system has a common structure of a pipeline consisting of several modules (e.g., speech recognition, natural language understanding, and dialog management). In this paper, we tackle a task of domain classification for the natural language understanding module by employing machine learning models such as convolutional neural network and random forest. For our dataset of seven service domains, we showed that the random forest model achieved the best performance (F1 score 0.97). As a future work, we will keep finding a better approach for domain classification by investigating other machine learning models.

18

감정인식은 인간과 컴퓨터 상호작용 분야에서 다양한 생체신호를 활용하여 지속적으로 연구되고 있다. 특히, 뇌파 (Electroencephalogram, EEG)는 음성과 표정보다 객관적이고, 접근이 용이하기 때문에, 감정연구에서 활발하게 활용되고 있다. 본 논문에서는 총 12명의 한국인을 대상으로 6개의 감정범주인 화남(Anger), 흥분(Excitement), 두려움(Fear), 행복(Happiness), 슬픔(Sadness), 그리고 중립(Neutral)로 구성된 4분간의 감정영상을 시청하 는 상황에서 뇌파 데이터를 수집하였다. 수집된 데이터는 전처리 과정을 거친 후, 파워스펙트럼(power spectrum) 값을 추출하였다. 이를 활용하여 Support Vector Machine(SVM)과 Long Short Term Memory(LSTM) 모델을 통하여 감정분류를 진행하였다. 실험방법은 개인의 차이를 고려하여 실험참가자 독립적(Subject-independent)방 식으로 진행되었다. 그 결과, LSTM이 SVM 보다 6개의 감정에 대하여 12.91% 향상된 80.46%의 분류정확도를 얻었다. 또한, LSTM의 레이어(Layer)수를 변경하여 분류를 진행한 결과로, 레이어(Layer)수가 2개일 때, 82.17%로 1.71% 향상되었다. 결론적으로, 본 논문에서는 6개의 감정범주에 대한 감정분류로 82.17%의 높은 정 확도를 얻었으며, LSTM은 뇌파와 같은 시간 의존형 데이터에 적합한 알고리즘임을 확인한 데에 의의가 있다.

EEG-based emotion recognition is a challenging task in human-computer interaction(HCI). Electroencephalogram(EEG) has significant advantages in emotion recognition since it directly reflects brain activity. This study recorded EEG signals from 12 participants watching Korean movie clips for 4 minutes, defined as six emotions: anger, excitement, fear, happiness, sadness, and neutral state. We also extract spectral power values as the feature using Fast Fourier Transform(FFT). Next, we have performed three classification models, Support Vector Machine(SVM), Long Short Term Memory(LSTM) with subject-dependent experiments. As a result, the best accuracy of 81.46% has been achieved in the LSTM model for classifying six emotional states, while SVM accuracy is 68.64%. In addition, promising results have been obtained for the LSTM model of 82.89% with two layers when applied for different layers. In conclusion, we present that the LSTM model specializes in the time-sequence data processing.

20

척추 압박 골절은 골다공증의 가장 흔한 합병증으로 기능적 제한과 요통을 유발하며, 생존율이 감소하는 등 심각한 건강 문제로 대두되고 있습니다. 척추 압박 골절은 정확한 진단과 적절한 치료 및 관리가 매우 중요하나, 현재의 진단 방법은 전문의의 판단에 의존하기 때문에 관찰자 간 변동성으로 인해 진단 불일치가 발생할 수 있습니다. 이에 비해 라디오믹스 특징은 정량적인 측정 지표로써, 전문의 간의 주관적 판단 차이를 줄이고 진단의 일관성을 높일 수 있습니다. 이번 연구에서는 방사선 영상 특징을 활용하여 약 470개의 개별 척추에 대해 척추 압박 골절의 양성 및 악성을 분류하는 머신 러닝 모델을 학습했습니다. 임상의가 판단한 개별 척추에 대해 추출한 방사선 영상 특징을 사용하여 다양한 머신 러닝 모델을 훈련한 결과, 정확도 0.72, 재현률 0.78, 정밀도 0.72, 그리고 F1 점수 0.75의 성능을 얻을 수 있었습니다. 이를 기반으로 진단 보조 측면에서 임상 현장에 적용한다면 척추 압박 골절 진단의 정확성과 일관성을 향상시킬 수 있으며, 이는 불필요한 추가 검사나 치료를 줄이고, 환자에게 적절한 치료 계획을 조기에 수립할 수 있게 하여 치료 효과를 높일 수 있습니다.

 
1 2 3 4 5
페이지 저장