This study aims to identify an AI classification model that is optimal for the classification of linguistic data. For this purpose, three commonly used classification models (XGBoost classifier, Random Forest Classifier, and SVM classifier) are compared in terms of their performance. Specifically, the three models are trained to classify the input data into essays and dialogues based on the syntactic complexity-related characteristics that distinguish between essays and dialogues. To determine if a model performing well on balanced data also performs well on imbalanced data, the three models’ performances are measured under two conditions: when the training dataset is balanced and when it is imbalanced. The performances of the trained models on the first test dataset are evaluated using accuracy, F1-score, normalized confusion matrix, and the area under the receiver operating characteristic curve. The performances on the second test dataset are assessed in terms of accuracy, confusion matrix, precision, and recall. The results demonstrate that the Random Forest Classifier has the best performance among the three models regardless of the balance of training data.
목차
1. Introduction 2. Background 2.1 XGBoost Classifier 2.2 Random Forest Classifier 2.3 Support Vector Machine Classifier 3. Method 3.1 Data 3.2 Procedure 4. Results 4.1 Results from XGBoost Classifier 4.2 Results from Random Forest Classifier 4.3 Results from Support Vector Machine Classifier 5. Discussion and Conclusion References [Abstract]
국제언어인문학회 [INTERNATIONAL ASSOCIATION FOR HUMANISTIC STUDIES IN LANGUAGE]
설립연도
2000
분야
인문학>언어학
소개
국제언어인문학회는 '언어를 통한 인문학 연구'의 필요성에 동감하는 여러 전공분야 학자들의 뜻을 담고 있습니다. 언어에 초점을 맞추는 것은, 다양한 전공분야의 참여에서 생겨날 수 있는 '이질적 집합'의 상황을 극복하기 위한 장치입니다. 현재로서는 작은 불씨를 지핀 것에 불과합니다. 그러나 이렇게 일구어진 불꽃이 새로운 학풍의 바람결에 커다란 섬광으로 빛나게 될 날이 올 것을 우리는 확신합니다. 우리의 학회와 학술지는 인문학 불변의 가치와 시대적 사명을 인식하는 국내외의 학자들을 향해 활짝 개방되어 있습니다. 특정 전공의 범위를 넘어서서 철학, 문학, 언어학, 종교, 역사, 문화, 예술 등의 시각에서 언어의 본질을 토론할 기회가 될 것입니다.