Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 101
No
1

4,200원

본 논문에서는 여러 연구기관에서 논의하고 있는 데이터 기반 평가 방법론 중 토픽모델링 기법을 이용하여 계량 적인 값을 도출하고 그 과정에서 실제 전문가들이 수행하는 국가연구개발사업과제와 이를 법률과 정책실무에서 다루는 국회 상임위원회 간의 정책적 인식 차이가 있는지 ICT 분야를 중심으로 파악해 보고자 한다. 먼저 HAN 모델로 사업과제 데이터를 학습하여 ICT 문서를 분류하는 모델을 만들고, 해당 모델을 통해 분류된 ICT 문서를 대상으로 LDA 토픽모델링 분석을 수행하여 국가연구개발사업과제 데이터와 국회 상임위원회 회의록에서 도출 된 토픽과 분포를 비교한다. 구체적으로 총 26개의 토픽이 도출되었으며, 각 토픽이 포함하는 단어와 문서 분포 비율을 살펴봤을 때, 국가사업과제는 상대적으로 전문적인 주제의 문서가 많았으며, 국회 상임위원회는 상대적으로 사회적이고 대중적인 문제를 다루는 것으로 나타나 인식에 다소 차이가 있는 것으로 보였다. 인식의 차이를 수치적으로 확인할 수 있는 만큼, 향후 정책이나 과제 평가에 사용할 수 있는 지표에 대한 기초연구로 활용 가능할 것이다.

In this paper, numerical values are derived using topic modeling among data-based evaluation methodologies discussed by various research institutes. In addition, we will focus on the ICT field to see if there is a difference in policy perception between the national R&D project and standing committee. First, we create model for classifying ICT documents by learning R&D project data using HAN model. And we perform LDA topic modeling analysis on ICT documents classified by applying the model, compare the distribution with the topics derived from the R&D project data and proceedings of standing committees. Specifically, a total of 26 topics were derived. Also, R&D project data had professionally topics, and the standing committee-discuss relatively social and popular issues. As the difference in perception can be numerically confirmed, it can be used as a basic study on indicators that can be used for future policy or project evaluation.

2

Document Classification Model Using Web Documents for Balancing Training Corpus Size per Category

Park, So-Young, Chang, Juno, Kihl, Taesuk

[Kisti 연계] 한국정보통신학회 Journal of information and communication convergence engineering Vol.11 No.4 2013 pp.268-273

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, we propose a document classification model using Web documents as a part of the training corpus in order to resolve the imbalance of the training corpus size per category. For the purpose of retrieving the Web documents closely related to each category, the proposed document classification model calculates the matching score between word features and each category, and generates a Web search query by combining the higher-ranked word features and the category title. Then, the proposed document classification model sends each combined query to the open application programming interface of the Web search engine, and receives the snippet results retrieved from the Web search engine. Finally, the proposed document classification model adds these snippet results as Web documents to the training corpus. Experimental results show that the method that considers the balance of the training corpus size per category exhibits better performance in some categories with small training sets.

3

Category Factor Based Feature Selection for Document Classification

Kang Yun-Hee

[Kisti 연계] 한국콘텐츠학회 International journal of contents Vol.1 No.2 2005 pp.26-30

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

According to the fast growth of information on the Internet, it is becoming increasingly difficult to find and organize useful information. To reduce information overload, it needs to exploit automatic text classification for handling enormous documents. Support Vector Machine (SVM) is a model that is calculated as a weighted sum of kernel function outputs. This paper describes a document classifier for web documents in the fields of Information Technology and uses SVM to learn a model, which is constructed from the training sets and its representative terms. The basic idea is to exploit the representative terms meaning distribution in coherent thematic texts of each category by simple statistics methods. Vector-space model is applied to represent documents in the categories by using feature selection scheme based on TFiDF. We apply a category factor which represents effects in category of any term to the feature selection. Experiments show the results of categorization and the correlation of vector length.

4

Word-Level Embedding to Improve Performance of Representative Spatio-temporal Document Classification

Byoungwook Kim, Hong-Jun Jang

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.19 No.6 2023 pp.830-841

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Tokenization is the process of segmenting the input text into smaller units of text, and it is a preprocessing task that is mainly performed to improve the efficiency of the machine learning process. Various tokenization methods have been proposed for application in the field of natural language processing, but studies have primarily focused on efficiently segmenting text. Few studies have been conducted on the Korean language to explore what tokenization methods are suitable for document classification task. In this paper, an exploratory study was performed to find the most suitable tokenization method to improve the performance of a representative spatio-temporal document classifier in Korean. For the experiment, a convolutional neural network model was used, and for the final performance comparison, tasks were selected for document classification where performance largely depends on the tokenization method. As a tokenization method for comparative experiments, commonly used Jamo, Character, and Word units were adopted. As a result of the experiment, it was confirmed that the tokenization of word units showed excellent performance in the case of representative spatio-temporal document classification task where the semantic embedding ability of the token itself is important.

5

Inferring CEFR and Its Companion Volume Reading Comprehension Indices Based on Japanese Document Classification Method with Binary Classification KCI 등재

宮崎佳典, Vuong Hong Duc, 谷誠司, 安志英, 元裕璟

한국일본학회 일본학보 제125권 2020.11 pp.153-175

※ 기관로그인 시 무료 이용이 가능합니다.

6,000원

최근 학습중인 언어를 사용하여 구체적으로 무엇을 할 수 있는지를 나타내는 범용 체계에 큰 관심이 모아지고 있다. 그 중에서도 2001년에 유럽위원회가 발표한 Common European Framework of Reference for Languages (CEFR)는 언어능력의 국제표 준으로 세계적으로 평가가 높다. 2017년에는 그것을 보완하는 CEFR Companion Volume이 공개되어, PreA1 레벨이 추가되는 등 더욱 더 레벨이 세분화되었다. CEFR 를 사용한 연구와 실천 예는 영어를 비롯한 많은 언어에서 이루어지고 있는 반면, 일 본어 교육을 염두에 둔 CEFR연구는 수적으로도 여전히 적으며, 일본어 CEFR준수 텍 스트 코퍼스도 현재까지의 연구 결과 존재하지 않는다. 본 연구에서는, 코퍼스를 작성 할 때 발생하는, 예문에 CEFR의 독해력을 반영하는 Can-Do Statements (CDS)를 부여 하는 노력을 경감하기 위해 자동분류 실장(実装)에 대해 지속적으로 연구하고 있다. 분류 방법에는 Support Vector Machine과 랜덤 포레스트에 의한 지도 학습을 적용하 고, 기계 학습을 위한 예문의 특징량으로써 문서 유형, 전문성, 문장 길이, 한자 비율 4 개를 사용한다. Pre-A1 레벨은 종전 레벨 군과 난이도와 구성언어요소에 큰 차이가 있기 때문에, 모든 레벨의 CDS를 한 번에 분류하는 과거 방식에 비해 2 단계에 따른 CDS 분류에 따른 정확도 향상을 목표로 하였다. 또한, Web 어플리케이션의 개발을 실시하여, 자동 분류 알고리즘을 내부 구현함으로써, 주어진 예문에 대하여 이에 대응 하는CDS를 자동 부여하는 기능과, 특정 CDS를 선택함으로써 이에 해당하는 예문 리 스트를 그 확실성 순으로 제공하는 기능을 제공하고 있다.

In recent years, a lot of attention is being paid to general-purpose frameworks that show what can be concretely done in using a target language for learners. In 2017, the Common European Framework of Reference for Languages (CEFR) Companion Volume was released. This volume complements the CEFR initially published in 2001, which is widely considered as an international standard for language ability, and introduces a Pre-A1 level. Conversely, there are few studies on CEFR for Japanese language education, and from the past studies, it was noted that there are no Japanese CEFR compliant text corpora. Thus, the present study aims to classify example sentences according to their corresponding Can-Do Statements (CDSs) to reduce efforts in creating a corpus. Support Vector Machine and Random Forest were applied to the classification approach where document types, specialty, sentence length, and kanji ratio have been given as the features of example sentences. The Pre-A1 level has a great difference in difficulty level and constituent language elements from the previous level groups. Therefore, our study seeks to improve the accuracy through binary classification combined with incorporation of the past method of classifying all levels of CDSs at once. Moreover, we also developed a web application that would help attach CDSs efficiently to example sentences and provide example sentence collections corresponding to specific CDSs.

6

CRM ; Web document classification for Mass-Customized online service in e-CRM based on Fuzzy Logic

Iraj Mahdavi, Nam Jae Cho, Hyun Soo Han, Babak Shirazi

한국경영정보학회 한국경영정보학회 정기 학술대회 2005년 추계학술대회 2005.11 pp.147-149

※ 기관로그인 시 무료 이용이 가능합니다.

3,000원

7

사례기반 추론을 이용한 한글 문서분류 시스템 KCI 등재

이재식, 이종운

한국경영정보학회 Asia Pacific Journal of Information Systems 제12권 제2호 2002.06 pp.179-195

※ 기관로그인 시 무료 이용이 가능합니다.

5,100원

9

건설 리스크 도출을 위한 SVM 기반의 건설프로젝트 문서 분류 모델 개발

강동욱, 조민건, 차기춘, 박승희

[Kisti 연계] 대한토목학회 대한토목학회논문집 Vol.43 No.6 2023 pp.841-849

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

건설프로젝트는 공기 지연, 건설 재해 등 다양한 요인으로 인한 리스크가 존재한다. 이러한 건설 리스크를 기반으로 건설프로젝트의 공사 기간의 산정 방법은 주로 감독자 경험에 의존한 주관적 판단으로 이루어지고 있다. 또한, 공기 지연과 건설 재해로 지연된 건설프로젝트 일정을 맞추기 위한 무리한 단축 시공은 부실시공 등의 부정적인 결과를 초래하며, 지연된 일정으로 인한 사회 기반 시설물 부재로 경제적 손실이 발생한다. 이러한 건설프로젝트의 리스크 해결을 위한 데이터 기반의 과학적 접근과 통계적 분석이 필요한 실정이다. 실제 건설프로젝트에서 수집되는 데이터는 비정형 텍스트 형태로 저장되어 있어 데이터를 기반으로 한 리스크를 적용하기 위해서는 데이터 전처리에 많은 인력과 비용을 수반하기 때문에 텍스트 마이닝을 활용한 데이터 분류 모델을 통한 기초자료를 요구한다. 따라서, 본 연구에서는 건설프로젝트 문서를 수집하여 텍스트 마이닝을 활용하여 SVM(Support Vector Machine) 기반의 데이터 분류 모델을 통해 리스크 관리를 위한 문서 기초자료 생성 분류 모델을 개발하였다. 향후 연구 결과를 통해 정량적인 분석을 통해서 건설프로젝트 공정관리 등에 있어 효율적이고 객관적인 기초자료로 활용되어 리스크 관리가 가능해질 것으로 기대된다.

Construction projects have risks due to various factors such as construction delays and construction accidents. Based on these construction risks, the method of calculating the construction period of the construction project is mainly made by subjective judgment that relies on supervisor experience. In addition, unreasonable shortening construction to meet construction project schedules delayed by construction delays and construction disasters causes negative consequences such as poor construction, and economic losses are caused by the absence of infrastructure due to delayed schedules. Data-based scientific approaches and statistical analysis are needed to solve the risks of such construction projects. Data collected in actual construction projects is stored in unstructured text, so to apply data-based risks, data pre-processing involves a lot of manpower and cost, so basic data through a data classification model using text mining is required. Therefore, in this study, a document-based data generation classification model for risk management was developed through a data classification model based on SVM (Support Vector Machine) by collecting construction project documents and utilizing text mining. Through quantitative analysis through future research results, it is expected that risk management will be possible by being used as efficient and objective basic data for construction project process management.

10

토픽모델링과 딥 러닝을 활용한 생의학 문헌 자동 분류 기법 연구

육지희, 송민

[Kisti 연계] 한국정보관리학회 정보관리학회지 Vol.35 No.2 2018 pp.63-88

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 LDA 토픽 모델과 딥 러닝을 적용한 단어 임베딩 기반의 Doc2Vec 기법을 활용하여 자질을 선정하고 자질집합의 크기와 종류 및 분류 알고리즘에 따른 분류 성능의 차이를 평가하였다. 또한 자질집합의 적절한 크기를 확인하고 문헌의 위치에 따라 종류를 다르게 구성하여 분류에 이용할 때 높은 성능을 나타내는 자질집합이 무엇인지 확인하였다. 마지막으로 딥 러닝을 활용한 실험에서는 학습 횟수와 문맥 추론 정보의 유무에 따른 분류 성능을 비교하였다. 실험문헌집단은 PMC에서 제공하는 생의학 학술문헌을 수집하고 질병 범주 체계에 따라 구분하여 Disease-35083을 구축하였다. 연구를 통하여 가장 높은 성능을 나타낸 자질집합의 종류와 크기를 확인하고 학습 시간에 효율성을 나타냄으로써 자질로의 확장 가능성을 가지는 자질집합을 제시하였다. 또한 딥 러닝과 기존 방법 간의 차이점을 비교하고 분류 환경에 따라 적합한 방법을 제안하였다.

This research evaluated differences of classification performance for feature selection methods using LDA topic model and Doc2Vec which is based on word embedding using deep learning, feature corpus sizes and classification algorithms. In addition to find the feature corpus with high performance of classification, an experiment was conducted using feature corpus was composed differently according to the location of the document and by adjusting the size of the feature corpus. Conclusionally, in the experiments using deep learning evaluate training frequency and specifically considered information for context inference. This study constructed biomedical document dataset, Disease-35083 which consisted biomedical scholarly documents provided by PMC and categorized by the disease category. Throughout the study this research verifies which type and size of feature corpus produces the highest performance and, also suggests some feature corpus which carry an extensibility to specific feature by displaying efficiency during the training time. Additionally, this research compares the differences between deep learning and existing method and suggests an appropriate method by classification environment.

11

CEFRに対応した日本語例文自動分類システムのBERT適用による精度改善の試み KCI 등재

宮崎佳典, Cao Hoai Giang, 谷誠司, 安志英, 元裕璟

한국일본학회 일본학보 제137권 2023.11 pp.43-62

※ 기관로그인 시 무료 이용이 가능합니다.

5,500원

최근 CanDo에 의한 언어 능력 척도의 하나로 CEFR(유럽 언어 공통 기준)가 관심을 받고 있으며, 외국어 교육에 분야에서도 전세계적으로 도입되고 있는 것은 잘 알려진 사 실이다. 한편, 일본어 교육을 위한 CEFR의 연구 사례는 많지 않으며, 일본어 CEFR 준거 텍스트 코퍼스에 관한 연구도 거의 이루어지고 있지 않다. 본 연구에서는 코퍼스를 작성 할 때, 예문에 CEFR의 독해력을 나타내는 CDS(능력 기술문)에 관한 정보를 부여하는 것 이 상정되었을 때 그 노력을 경감하기 위한 자동 분류의 구현을 지속적으로 연구하고 있 다. 분류를 위한 특징량으로 현재 문서 타입, 전문성, 문장, 한자율을 채용하고 있으며, 그 속의 문서 타입이나 전문성 분류를 위해 fastText가 이용되고 있다. 이에 대해 본 논문 에서는 지금까지 많은 자연언어 처리 태스크에서 유효하게 동작하는 것이 확인되어 정평 이 있는 BERT 알고리즘(Bidirectional Encoder Representations from Transformers)을 새롭게 적용하여 시도한 결과, 예측 정밀도를 향상시키는 것에 성공하였으며, 예측한 CDS의 수 를 보다 적절하게 억제할 수 있었다. 향후 예측할 수 있는 수를 줄임으로써 적합률을 높 이고, 보다 효과적인 특징량의 작성을 검토하여 CDS 정보가 더 많은 데이터를 수집하고 자 한다. 또한 이번 연구 대상은 CEFR가 초기에 설정한 6단계의 언어 능력 레벨(초급 레 벨의 A1부터 A2, B1, B2, C1와 최상급 레벨의 C2까지) 중, 모국어화자라도 난이도가 높 은 C1과 C2를 뺀 A1, B2을 더하여 2017년 CEFR을 보완한 것으로 공개된 CEFR Companion Volume에서 새롭게 추가된 PreA1 레벨도 포함하였다. 또한 기술 항목은 Reading, Writing, Speaking, Listening, Interaction 중 Reading에 초점을 맞추고, 그에 대응하 는 CDS는 34으로 세었다. 즉, 본 연구는 입력 예문을 34개의 라벨로 분류하는 것을 중심 으로 하고 있다는 것을 의미한다.

The Common European Framework of Reference for Languages (CEFR) is an example of a Can-Do language proficiency scale that has attracted large attention in recent years and has been introduced in foreign language education worldwide. On the other hand, there are a few examples of CEFR research related to Japanese language education. As far as the present authors investigated, there is no Japanese CEFR-compliant text corpus. In the current research, to create a corpus, we focused on the implementation of automatic classification in order to reduce the effort of adding Can-Do Statement (CDS), thereby enhancing the reading comprehension of CEFR in the example sentences. Document type, specialty, sentence length, and Kanji rate are commonly used as the parameters for classification; however, the current version of the aforementioned implementation uses fastText to identify document type and specialty. The present study attempted to apply the Bidirectional Encoder Representations from Transformers (BERT) algorithm, which has been confirmed to work effectively in many natural language processing tasks. The findings of the research showed that prediction accuracy was improved and it was possible to suppress the number of CDSs for more appropriate prediction. Prospects include improving the precision by further narrowing down the number of predictions, creating more compelling features, and collecting more data using CDS information. The targets in this study were CEFR proficiency levels A1, A2, B1, B2, and PreA1 for Reading skill items (34 CDSs correspond to those levels).

12

지능적인 웹문서 분류를 위한 구조 및 프로세스 설계 연구 KCI 등재후보

장영철

한국디지털정책학회 디지털융복합연구 제6권 제4호 2008.12 pp.177-183

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

This paper aims to offer a solution based on intelligent document classification to create a user-centric information retrieval system allowing user-centric linguistic expression. So, structures expressing user intention and fine document classifying process using EBL, similarity, knowledge base, user intention, are proposed. To overcome the problem requiring huge and exact semantic information, a hybrid process is designed integrating keyword, thesaurus, probability and user intention information. User intention tree hierarchy is build and a method of extracting group intention between key words and user intentions is proposed. These structures and processes are implemented in HDCI(Hybrid Document Classification with Intention) system. HDCI consists of analyzing user intention and classifying web documents stages. Classifying stage is composed of knowledge base process, similarity process and hybrid coordinating process. With the help of user intention related structures and hybrid coordinating process, HDCI can efficiently categorize web documents in according to user's complex linguistic expression with small priori information.

13

4,000원

스캔된 인보이스에 특화된 서류 관리 자동화 시스템 구축에있어서 추출된 금전적 데이터의 정확도에대한 엄격한 요구는 인보이스 테이블을 위한 발생적 모델 설계에서 자체 인증 절차를 포함하는 것을 필요로 한다. 가격 = 단가 x 구매수량과 같은 내부적 관계식을 활용한 단순한 인증 절차를 사용하는 것이 전형적 방법론이다. 본 논문에서는 영상내 테이블 헤더 부분의 탐색과 탐색된 헤더의 컬럼 구분자를 활용하는 개선된 자동 인증 절차를 갖춘 인보이스내 정보 추출 모델을 제안한다.

Development of automated document management system specified for scanned invoice images suffers from rigorous accuracy requirements for extraction of monetary data, which necessiate automatic validation on the extracted values for a generative invoice table model. Use of certain internal constraints such as “amount = unit price times quantity” is typical implementation. In this paper, we propose a noble invoice information extraction model with improved auto-validation method by utilizing table header detection and column classification.

14

문서분류의 이론과 변천에 관한 연구 - 조선조이후 현행 ‘정부공문서분류’까지 -

최정태, 이주연

한국기록관리학회 한국기록관리학회지 제3권 제2호 2003.12 pp.1-32

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

15

At present, information retrieval systems are simply expressed with a combination of keyword search according to the direct keyword matching method to get the information that users need. Because of this, documents retrieval systems serve too many documents due to term ambiguity. This makes the user need extra time and effort to get closer the document. To overcome these problems, this paper proposes the information retrieval system based on the content that connects documents according to the degree of semantic link that expresses a fuzzy value by fuzzy function. This paper also proposes an algorithm that produces a hierarchical structure using the degree of concept and content among documents. As a result, we are able to select and to provide user-interested documents.

16

Document Classification Using N-gram and Word Semantic Similarity

Mei-ying Ren, Sinjae Kang

보안공학연구지원센터(IJUNESST) International Journal of u- and e- Service, Science and Technology Vol.8 No.8 2015.08 pp.111-118

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

This paper mainly conducted two series of experiments. One is investigation between language dependent and language independent features. Bi-grams in Korean experiments and uni-grams in Chinese contributed most as basic features. And another one is utilization of Korean WordNet to improve the performance of Korean document classification. Korean WordNet is a Korean Lexical Semantic Network. Language independent features seem can lead better performance and stable. The performance of Korean text classification was improved by using Korean WordNet.

17

WordNet-based Hybrid VSM for Document Classification SCOPUS

Luda Wang, Peng Zhang, Shouping Gao

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.1 2016.01 pp.185-200

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Many text classifications depend on statistical term measures or synsets to implement document representation. Such document representations ignore the lexical semantic contents or relations of terms, leading to losing the distilled mutual information. This work proposed a synthetic document representation method, WordNet-based hybrid VSM, to solve the problem. This method constructed a data structure of semantic-element information to characterize lexical semantic contents, and support disambiguation of word stems. As a template, lexical semantic vector consisting of lexical semantic contents was built in the lexical semantic space of corpus, and lexical semantic relations are marked on the vector. Then, it connects with special term vector to form the eigenvector in hybrid VSM. Applying algorithm NWKNN, on text corpus Reuter-21578 and its adjusted version, the experiments show that the eigenvector performs F1 measure better than document representations based on TF-IDF.

18

A Study on Naïve Bayes Classifier Based Document Classification Scheme with an Apriori Feature Extraction SCOPUS

Jong-Yeol Yoo, Min-Ho Lee, Grace Aloyce, Dong-Min Yang

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.7 2016.07 pp.207-214

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

A document classifier is an essential tool for classifying the various types of documents being generated in the Big Data era. In recent years, the wide variety of information services available for use with smartphones and portable mobile devices (tablets) have provided a technique that efficiently classifies the quality of sorted data. A common type of document classification scheme is the naïve Bayes classifier. The Naïve Bayes scheme is based on performance classification, which varies widely depending on the method of extraction used in the document. In this paper, we propose a system model that offers feature extraction methods which combine frequency with associated words. This model is then applied to the Naïve Bayes classifier to precisely classify documents. This method is proposed as an alternative to using traditional classification techniques. In addition, experiments will be evaluated by the existing document classification techniques and the proposed techniques.

19

The Impact of Feature Reduction Techniques on Arabic Document Classification SCOPUS

Abdullah Ayedh, Guanzheng Tan, Hamdi Rajeh

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.6 2016.06 pp.67-80

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Feature reduction are common techniques that used to improve the efficiency and accuracy of the document classification systems. The problems associated with these techniques are the highly dimensionality of the feature space and The difficulty of selecting the important features for understanding the document in question. The document usually consists of several parts and the important features that more closely associated with the topic of the document are appearing in the first parts or repeated in several parts of the document. Therefore, the position of the first appearance of a word and the compactness of the word considered as factors that determine the important features using the information within a document. This study, explored the impact of combining three feature weighting methods that depend on inverse document frequency (IDF), namely, Term frequency (TFiDF), the position of the first appearance of a word (FAiDF), and the compactness of the word (CPiDF) on the classification accuracy. In addition, we have investigated different feature selection techniques, namely, Information gain (IG), Goh and Low (NGL) coefficients, Chi-square Testing (CHI), and Galavotti-Sebastiani-Simi Coefficient (GSS) in order to improve the performance for Arabic document classification system. Experimental analysis on Arabic datasets reveals that the proposed methods have a significant impact on the classification accuracy, and in most cases the FAiDF feature weighting performed better than CPiDF and TFiDF. The results also clearly showed the superiority of the GSS over the other feature selection techniques and achieved 98.39% micro-F1 value when using a combination of TFiDF, FAiDF, and CPiDF as feature weighting method.

20

Is Naïve Bayes a Good Classifier for Document Classification? SCOPUS

S.L. Ting, W.H. Ip, Albert H.C. Tsang

보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.5 No.3 2011.07 pp.37-46

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Document classification is a growing interest in the research of text mining. Correctly identifying the documents into particular category is still presenting challenge because of large and vast amount of features in the dataset. In regards to the existing classifying approaches, Naïve Bayes is potentially good at serving as a document classification model due to its simplicity. The aim of this paper is to highlight the performance of employing Naïve Bayes in document classification. Results show that Naïve Bayes is the best classifiers against several common classifiers (such as decision tree, neural network, and support vector machines) in term of accuracy and computational efficiency.

 
1 2 3 4 5
페이지 저장