Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 9
No
1

Finding an Optimal Classification Model for Analyzing Linguistic Data KCI 등재

Wonbin Kim

국제언어인문학회 인문언어 제26권 2호 2024.12 pp.205-236

※ 기관로그인 시 무료 이용이 가능합니다.

7,300원

This study aims to identify an AI classification model that is optimal for the classification of linguistic data. For this purpose, three commonly used classification models (XGBoost classifier, Random Forest Classifier, and SVM classifier) are compared in terms of their performance. Specifically, the three models are trained to classify the input data into essays and dialogues based on the syntactic complexity-related characteristics that distinguish between essays and dialogues. To determine if a model performing well on balanced data also performs well on imbalanced data, the three models’ performances are measured under two conditions: when the training dataset is balanced and when it is imbalanced. The performances of the trained models on the first test dataset are evaluated using accuracy, F1-score, normalized confusion matrix, and the area under the receiver operating characteristic curve. The performances on the second test dataset are assessed in terms of accuracy, confusion matrix, precision, and recall. The results demonstrate that the Random Forest Classifier has the best performance among the three models regardless of the balance of training data.

2

감정 분석은 디지털 텍스트를 분석하여 메시지의 감정적 어조가 긍정적인지, 부정적인지 또는 중립적인지를 확인하 는 프로세스이다. 오늘날 회사는 이메일, 고객 지원 채팅 트랜스크립트, 소셜 미디어 댓글 및 리뷰와 같은 대량의 텍스트 데이터를 스캔하여 주제에 대한 글쓴이의 태도를 자동으로 확인할 수 있다. 기업은 감정 분석의 인사이트를 활용하여 고객 서비스를 개선하고 브랜드 평판을 높인다. 그러나 언어의 복잡성과 반어적 표현 등의 사용으로 긍정, 부정 판단이 잘못 분류되는 사례가 있다. 본 논문은 랜덤 포레스트 분류기를 기반으로 이러한 문제 해결방법을 제안 하여 네이버 뉴스기사 5개월 분량으로 학습하고 교차검증을 실시하였으며, ROC 곡선에 기반한 AUC 값 중 우수한 데이터 세트로 학습하고 검증하였다.

Sentiment analysis is the process of analyzing digital text to determine whether the emotional tone of the message is positive, negative, or neutral. Today, companies can scan large amounts of text data such as emails, customer support chat transcripts, and social media comments and reviews to automatically determine the writer's attitude on a topic. Companies use insights from sentiment analysis to improve customer service and enhance brand reputation. However, there are cases where positive and negative judgments are misclassified due to the complexity of language and the use of ironic expressions. This paper proposed a method to solve this problem based on a random forest classifier, learned from 5 months of Naver news articles, performed cross-validation, and learned and verified with an excellent data set of AUC values based on the ROC curve.

3

랜덤포레스트 기반 지하철 통행 경로 추정

윤현수, 이은학, 김동규

한국ITS학회 한국ITS학회 학술대회 그린뉴딜 정책에 따른 ITS의 확대 추진 및 고도화 2020.11 pp.151-156

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

4

Random Forest Classifier-based Ship Type Prediction with Limited Ship Information of AIS and V-Pass

Jeon, Ho-Kun, Han, Jae Rim

[Kisti 연계] 대한원격탐사학회 대한원격탐사학회지 Vol.38 No.4 2022 pp.435-446

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Identifying ship types is an important process to prevent illegal activities on territorial waters and assess marine traffic of Vessel Traffic Services Officer (VTSO). However, the Terrestrial Automatic Identification System (T-AIS) collected at the ground station has over 50% of vessels that do not contain the ship type information. Therefore, this study proposes a method of identifying ship types through the Random Forest Classifier (RFC) from dynamic and static data of AIS and V-Pass for one year and the Ulsan waters. With the hypothesis that six features, the speed, course, length, breadth, time, and location, enable to estimate of the ship type, four classification models were generated depending on length or breadth information since 81.9% of ships fully contain the two information. The accuracy were average 96.4% and 77.4% in the presence and absence of size information. The result shows that the proposed method is adaptable to identifying ship types.

5

Performance of Random Forest Classifier for Flood Mapping Using Sentinel-1 SAR Images

Chu, Yongjae, Lee, Hoonyol

[Kisti 연계] 대한원격탐사학회 대한원격탐사학회지 Vol.38 No.4 2022 pp.375-386

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The city of Khartoum, the capital of Sudan, was heavily damaged by the flood of the Nile in 2020. Classification using satellite images can define the damaged area and help emergency response. As Synthetic Aperture Radar (SAR) uses microwave that can penetrate cloud, it is suitable to use in the flood study. In this study, Random Forest classifier, one of the supervised classification algorithms, was applied to the flood event in Khartoum with various sizes of the training dataset and number of images using Sentinel-1 SAR. To create a training dataset, we used unsupervised classification and visual inspection. Firstly, Random Forest was performed by reducing the size of each class of the training dataset, but no notable difference was found. Next, we performed Random Forest with various number of images. Accuracy became better as the number of images in creased, but converged to a maximum value when the dataset covers the duration from flood to the completion of drainage.

6

An Assessment of a Random Forest Classifier for a Crop Classification Using Airborne Hyperspectral Imagery

Jeon, Woohyun, Kim, Yongil

[Kisti 연계] 대한원격탐사학회 대한원격탐사학회지 Vol.34 No.1 2018 pp.141-150

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Crop type classification is essential for supporting agricultural decisions and resource monitoring. Remote sensing techniques, especially using hyperspectral imagery, have been effective in agricultural applications. Hyperspectral imagery acquires contiguous and narrow spectral bands in a wide range. However, large dimensionality results in unreliable estimates of classifiers and high computational burdens. Therefore, reducing the dimensionality of hyperspectral imagery is necessary. In this study, the Random Forest (RF) classifier was utilized for dimensionality reduction as well as classification purpose. RF is an ensemble-learning algorithm created based on the Classification and Regression Tree (CART), which has gained attention due to its high classification accuracy and fast processing speed. The RF performance for crop classification with airborne hyperspectral imagery was assessed. The study area was the cultivated area in Chogye-myeon, Habcheon-gun, Gyeongsangnam-do, South Korea, where the main crops are garlic, onion, and wheat. Parameter optimization was conducted to maximize the classification accuracy. Then, the dimensionality reduction was conducted based on RF variable importance. The result shows that using the selected bands presents an excellent classification accuracy without using whole datasets. Moreover, a majority of selected bands are concentrated on visible (VIS) region, especially region related to chlorophyll content. Therefore, it can be inferred that the phenological status after the mature stage influences red-edge spectral reflectance.

7

A Cross-cultural Analysis of Monologues Using Random Forest Classifier

김원빈

[NRF 연계] 한국영어학학회 영어학연구 Vol.31 No.1 2025.02 pp.97-119

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The objective of this study is to examine how English monologues produced by Asian learners of English vary across different countries. For the analysis of the differences across countries, this study investigated whether the monologues of learners from each country exhibit characteristics of essays or dialogues regarding syntactic complexity. To determine whether the monologues are closer to essays or dialogues, a supervised machine learning technique, specifically the Random Forest Classifier, was employed and the essay and dialogue data from the International Corpus Network of Asian Learners of English were used to train the model. The Random Forest Classifier was trained to classify monologues into essays and dialogues on the basis of syntactic complexity indices. After the trained model sorted the monologues from each country into essays and dialogues, the countries were grouped based on the counts of essays and dialogues. For further analysis, Euclidean distance was utilized to identify the country whose monologues’ syntactic complexity most closely resembles that observed in Korea’s monologues, as well as the country whose monologues’ syntactic complexity most closely matches that observed in English-speaking countries’ monologues. The results demonstrated that (i) Hong Kong, the Philippines, Singapore, and the set of English-speaking countries were assigned to a group with a higher count of essays, while China, Indonesia, Japan, Korea, Pakistan, Taiwan, and Thailand were assigned to a group with a higher count of dialogues; (ii) Thailand’s monologues are most similar to those of Korea; (iii) Singapore’s monologues are most similar to those of English-speaking countries.

8

Bag-of-Feature 특징과 랜덤 포리스트를 이용한 의료영상 검색 기법

손정은, 곽준영, 고병철, 남재열

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2012 pp.601-603

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 의료영상의 특성을 반영하여 영상의 그래디언트 방향 값을 특징으로 하는 Oriented Center Symmetric Local Binary Patterns (OCS-LBP) 특징을 개발하고 추출된 특징 값에 대해 차원을 줄이고 의미 있는 특징 단위로 재 생성하기 위해 Bag-of-Feature (BoF)를 적용하였다. 검색을 위해서는 기존의 영상 검색 방법과는 다르게, 학습 영상을 이용하여 랜덤 포리스트 (Random Forest)를 사전에 학습시켜 데이터베이스 영상을 N 개의 클래스로 자동 분류 시키고, 질의로 입력된 영상을 같은 방법으로 랜덤 포리스트에 적용하여 상위 확률 값을 갖는 2 개의 클래스에서만 K-nearest neighbor 방법으로 유사 영상을 검색결과로 제시하는 새로운 영상검색 방법을 제시하였다. 실험결과에서 본 논문의 우수성을 증명하기 위해 일반적인 유사성 측정 방법과 랜덤 포리스트를 이용한 방법의 검색 성능 및 시간을 비교하였고, 검색 성능과 시간 면에서 상대적으로 매우 우수한 성능을 보여줌을 증명하였다.

9

Bag-of-Feature 특징과 랜덤 포리스트를 이용한 의료영상 검색 기법

손정은, 곽준영, 고병철, 남재열

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2012 pp.601-603

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 의료영상의 특성을 반영하여 영상의 그래디언트 방향 값을 특징으로 하는 Oriented Center Symmetric Local Binary Patterns (OCS-LBP) 특징을 개발하고 추출된 특징 값에 대해 차원을 줄이고 의미 있는 특징 단위로 재 생성하기 위해 Bag-of-Feature (BoF)를 적용하였다. 검색을 위해서는 기존의 영상 검색 방법과는 다르게, 학습 영상을 이용하여 랜덤 포리스트 (Random Forest)를 사전에 학습시켜 데이터베이스 영상을 N 개의 클래스로 자동 분류 시키고, 질의로 입력된 영상을 같은 방법으로 랜덤 포리스트에 적용하여 상위 확률 값을 갖는 2 개의 클래스에서만 K-nearest neighbor 방법으로 유사 영상을 검색결과로 제시하는 새로운 영상검색 방법을 제시하였다. 실험결과에서 본 논문의 우수성을 증명하기 위해 일반적인 유사성 측정 방법과 랜덤 포리스트를 이용한 방법의 검색 성능 및 시간을 비교하였고, 검색 성능과 시간 면에서 상대적으로 매우 우수한 성능을 보여줌을 증명하였다.

 
페이지 저장