Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 98
No
1

지능적인 웹문서 분류를 위한 구조 및 프로세스 설계 연구 KCI 등재후보

장영철

한국디지털정책학회 디지털융복합연구 제6권 제4호 2008.12 pp.177-183

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

This paper aims to offer a solution based on intelligent document classification to create a user-centric information retrieval system allowing user-centric linguistic expression. So, structures expressing user intention and fine document classifying process using EBL, similarity, knowledge base, user intention, are proposed. To overcome the problem requiring huge and exact semantic information, a hybrid process is designed integrating keyword, thesaurus, probability and user intention information. User intention tree hierarchy is build and a method of extracting group intention between key words and user intentions is proposed. These structures and processes are implemented in HDCI(Hybrid Document Classification with Intention) system. HDCI consists of analyzing user intention and classifying web documents stages. Classifying stage is composed of knowledge base process, similarity process and hybrid coordinating process. With the help of user intention related structures and hybrid coordinating process, HDCI can efficiently categorize web documents in according to user's complex linguistic expression with small priori information.

2

자연어 처리를 이용한 문서 분류 분야에서도 전통적인 방법에서 벗어나 단어 임베딩을 활용한 합성곱 신경망과 순환 신경망 등 심층 신경망을 이용한 다양한 연구가 진행되고 있다. 본 논문에서는 Word2Vec과 두 개의 계층으로 구성 된 양방향 장단기 기억 네트워크를 이용한 특허 문서의 IPC(International Patents Classification) 자동 분류 모델을 제안한다. IPC는 세계지식재산권기구에서 제정한 국제적으로 통일된 특허 분류 기준이며, 각 국가의 공인된 기관에서 수작업으로 분류하고 있다. IPC 자동 분류를 위하여 입력 시퀀스에 Word2Vec을 이용한 단어 임베딩가 중치를 사용한다. 그리고 가중치가 부여된 시퀀스를 두 개의 계층을 갖는 깊은 구조의 양방향 장단기 기억 네트워크 신경망에 입력하여 IPC를 분류한다. 실험 결과 특허 문서의 분류 정확도가 합성곱 신경망 보다는 약 7% 향상되었 으며, 순환 신경망을 단일로 이용하는 것 보다는 약 5% 향상된 것을 확인할 수 있었다. 또한 전통적인 방법인 나이 브 베이시안, 로지스틱 분류 및 서포트 벡터 머신보다는 5~12% 이상 우수한 성능을 나타내었다.

There are various studies using Deep Neural Network such as CNN(Convolutional Neural Network) and RNN(Recurrent Neural Network) that utilize word embedding in document classification using natural language processing out of traditional methods. In this paper, we propose the IPC(International Patents Classification) automatic classification model of patent documents using two layers BLSTM (Bidirectional Long Short Term memory) network. The IPC is an internationally uniform standard for patent classification established by the World Intellectual Property Organization and is categorized by hand in authorized agencies in each country. For the IPC automatic classification, we use word embedding weight with Word2Vec in the input sequences. And they are classified by entering a weighted sequences into a deep neural network with two layers BLSTM. The experimental results showed that the accuracy of classification is improved by about 7% than that of CNN, and about 5% than that of single layer LSTM that is a field of RNN. Also it showed more than 5~12% higher performance than traditional methods such as Naive Bayes, Logistic and Support Vector Machine classification.

3

Inferring CEFR and Its Companion Volume Reading Comprehension Indices Based on Japanese Document Classification Method with Binary Classification KCI 등재

宮崎佳典, Vuong Hong Duc, 谷誠司, 安志英, 元裕璟

한국일본학회 일본학보 제125권 2020.11 pp.153-175

※ 기관로그인 시 무료 이용이 가능합니다.

6,000원

최근 학습중인 언어를 사용하여 구체적으로 무엇을 할 수 있는지를 나타내는 범용 체계에 큰 관심이 모아지고 있다. 그 중에서도 2001년에 유럽위원회가 발표한 Common European Framework of Reference for Languages (CEFR)는 언어능력의 국제표 준으로 세계적으로 평가가 높다. 2017년에는 그것을 보완하는 CEFR Companion Volume이 공개되어, PreA1 레벨이 추가되는 등 더욱 더 레벨이 세분화되었다. CEFR 를 사용한 연구와 실천 예는 영어를 비롯한 많은 언어에서 이루어지고 있는 반면, 일 본어 교육을 염두에 둔 CEFR연구는 수적으로도 여전히 적으며, 일본어 CEFR준수 텍 스트 코퍼스도 현재까지의 연구 결과 존재하지 않는다. 본 연구에서는, 코퍼스를 작성 할 때 발생하는, 예문에 CEFR의 독해력을 반영하는 Can-Do Statements (CDS)를 부여 하는 노력을 경감하기 위해 자동분류 실장(実装)에 대해 지속적으로 연구하고 있다. 분류 방법에는 Support Vector Machine과 랜덤 포레스트에 의한 지도 학습을 적용하 고, 기계 학습을 위한 예문의 특징량으로써 문서 유형, 전문성, 문장 길이, 한자 비율 4 개를 사용한다. Pre-A1 레벨은 종전 레벨 군과 난이도와 구성언어요소에 큰 차이가 있기 때문에, 모든 레벨의 CDS를 한 번에 분류하는 과거 방식에 비해 2 단계에 따른 CDS 분류에 따른 정확도 향상을 목표로 하였다. 또한, Web 어플리케이션의 개발을 실시하여, 자동 분류 알고리즘을 내부 구현함으로써, 주어진 예문에 대하여 이에 대응 하는CDS를 자동 부여하는 기능과, 특정 CDS를 선택함으로써 이에 해당하는 예문 리 스트를 그 확실성 순으로 제공하는 기능을 제공하고 있다.

In recent years, a lot of attention is being paid to general-purpose frameworks that show what can be concretely done in using a target language for learners. In 2017, the Common European Framework of Reference for Languages (CEFR) Companion Volume was released. This volume complements the CEFR initially published in 2001, which is widely considered as an international standard for language ability, and introduces a Pre-A1 level. Conversely, there are few studies on CEFR for Japanese language education, and from the past studies, it was noted that there are no Japanese CEFR compliant text corpora. Thus, the present study aims to classify example sentences according to their corresponding Can-Do Statements (CDSs) to reduce efforts in creating a corpus. Support Vector Machine and Random Forest were applied to the classification approach where document types, specialty, sentence length, and kanji ratio have been given as the features of example sentences. The Pre-A1 level has a great difference in difficulty level and constituent language elements from the previous level groups. Therefore, our study seeks to improve the accuracy through binary classification combined with incorporation of the past method of classifying all levels of CDSs at once. Moreover, we also developed a web application that would help attach CDSs efficiently to example sentences and provide example sentence collections corresponding to specific CDSs.

4

CRM ; Web document classification for Mass-Customized online service in e-CRM based on Fuzzy Logic

Iraj Mahdavi, Nam Jae Cho, Hyun Soo Han, Babak Shirazi

한국경영정보학회 한국경영정보학회 정기 학술대회 2005년 추계학술대회 2005.11 pp.147-149

※ 기관로그인 시 무료 이용이 가능합니다.

3,000원

5

사례기반 추론을 이용한 한글 문서분류 시스템 KCI 등재

이재식, 이종운

한국경영정보학회 Asia Pacific Journal of Information Systems 제12권 제2호 2002.06 pp.179-195

※ 기관로그인 시 무료 이용이 가능합니다.

5,100원

7

CEFRに対応した日本語例文自動分類システムのBERT適用による精度改善の試み KCI 등재

宮崎佳典, Cao Hoai Giang, 谷誠司, 安志英, 元裕璟

한국일본학회 일본학보 제137권 2023.11 pp.43-62

※ 기관로그인 시 무료 이용이 가능합니다.

5,500원

최근 CanDo에 의한 언어 능력 척도의 하나로 CEFR(유럽 언어 공통 기준)가 관심을 받고 있으며, 외국어 교육에 분야에서도 전세계적으로 도입되고 있는 것은 잘 알려진 사 실이다. 한편, 일본어 교육을 위한 CEFR의 연구 사례는 많지 않으며, 일본어 CEFR 준거 텍스트 코퍼스에 관한 연구도 거의 이루어지고 있지 않다. 본 연구에서는 코퍼스를 작성 할 때, 예문에 CEFR의 독해력을 나타내는 CDS(능력 기술문)에 관한 정보를 부여하는 것 이 상정되었을 때 그 노력을 경감하기 위한 자동 분류의 구현을 지속적으로 연구하고 있 다. 분류를 위한 특징량으로 현재 문서 타입, 전문성, 문장, 한자율을 채용하고 있으며, 그 속의 문서 타입이나 전문성 분류를 위해 fastText가 이용되고 있다. 이에 대해 본 논문 에서는 지금까지 많은 자연언어 처리 태스크에서 유효하게 동작하는 것이 확인되어 정평 이 있는 BERT 알고리즘(Bidirectional Encoder Representations from Transformers)을 새롭게 적용하여 시도한 결과, 예측 정밀도를 향상시키는 것에 성공하였으며, 예측한 CDS의 수 를 보다 적절하게 억제할 수 있었다. 향후 예측할 수 있는 수를 줄임으로써 적합률을 높 이고, 보다 효과적인 특징량의 작성을 검토하여 CDS 정보가 더 많은 데이터를 수집하고 자 한다. 또한 이번 연구 대상은 CEFR가 초기에 설정한 6단계의 언어 능력 레벨(초급 레 벨의 A1부터 A2, B1, B2, C1와 최상급 레벨의 C2까지) 중, 모국어화자라도 난이도가 높 은 C1과 C2를 뺀 A1, B2을 더하여 2017년 CEFR을 보완한 것으로 공개된 CEFR Companion Volume에서 새롭게 추가된 PreA1 레벨도 포함하였다. 또한 기술 항목은 Reading, Writing, Speaking, Listening, Interaction 중 Reading에 초점을 맞추고, 그에 대응하 는 CDS는 34으로 세었다. 즉, 본 연구는 입력 예문을 34개의 라벨로 분류하는 것을 중심 으로 하고 있다는 것을 의미한다.

The Common European Framework of Reference for Languages (CEFR) is an example of a Can-Do language proficiency scale that has attracted large attention in recent years and has been introduced in foreign language education worldwide. On the other hand, there are a few examples of CEFR research related to Japanese language education. As far as the present authors investigated, there is no Japanese CEFR-compliant text corpus. In the current research, to create a corpus, we focused on the implementation of automatic classification in order to reduce the effort of adding Can-Do Statement (CDS), thereby enhancing the reading comprehension of CEFR in the example sentences. Document type, specialty, sentence length, and Kanji rate are commonly used as the parameters for classification; however, the current version of the aforementioned implementation uses fastText to identify document type and specialty. The present study attempted to apply the Bidirectional Encoder Representations from Transformers (BERT) algorithm, which has been confirmed to work effectively in many natural language processing tasks. The findings of the research showed that prediction accuracy was improved and it was possible to suppress the number of CDSs for more appropriate prediction. Prospects include improving the precision by further narrowing down the number of predictions, creating more compelling features, and collecting more data using CDS information. The targets in this study were CEFR proficiency levels A1, A2, B1, B2, and PreA1 for Reading skill items (34 CDSs correspond to those levels).

8

4,000원

스캔된 인보이스에 특화된 서류 관리 자동화 시스템 구축에있어서 추출된 금전적 데이터의 정확도에대한 엄격한 요구는 인보이스 테이블을 위한 발생적 모델 설계에서 자체 인증 절차를 포함하는 것을 필요로 한다. 가격 = 단가 x 구매수량과 같은 내부적 관계식을 활용한 단순한 인증 절차를 사용하는 것이 전형적 방법론이다. 본 논문에서는 영상내 테이블 헤더 부분의 탐색과 탐색된 헤더의 컬럼 구분자를 활용하는 개선된 자동 인증 절차를 갖춘 인보이스내 정보 추출 모델을 제안한다.

Development of automated document management system specified for scanned invoice images suffers from rigorous accuracy requirements for extraction of monetary data, which necessiate automatic validation on the extracted values for a generative invoice table model. Use of certain internal constraints such as “amount = unit price times quantity” is typical implementation. In this paper, we propose a noble invoice information extraction model with improved auto-validation method by utilizing table header detection and column classification.

9

Classification of Arabic Documents by a Model of Fuzzy Proximity with a Radial Basis Function

Taher ZAKI, Driss MAMMASS, Abdellatif ENNAJI, F. NOUBOUD

보안공학연구지원센터(IJFGCN) International Journal of Future Generation Communication and Networking vol.3 no.4 2010.12 pp.31-42

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

In this paper we propose a model of classification based on the principle of the fuzzy proximity of the terms within the documents. Given the heterogeneous nature of the Arabic documents in our possession, we have studied for this purpose the research model based on the semantic proximity of terms and inspired from the classic Boolean model. Our approach is based on the assumption that more the occurrences of terms in query are close with good connectivity in the extracted semantic graph from the set of document , more this document is relevant to this query. We propose a measure that provides a contextual and semantic search. We used not only a semantic graph to highlight the semantic connections between terms, but also an auxiliary dictionary to increase the connectivity of the graph and therefore the discrimination of documents relevant to the query.

11

At present, information retrieval systems are simply expressed with a combination of keyword search according to the direct keyword matching method to get the information that users need. Because of this, documents retrieval systems serve too many documents due to term ambiguity. This makes the user need extra time and effort to get closer the document. To overcome these problems, this paper proposes the information retrieval system based on the content that connects documents according to the degree of semantic link that expresses a fuzzy value by fuzzy function. This paper also proposes an algorithm that produces a hierarchical structure using the degree of concept and content among documents. As a result, we are able to select and to provide user-interested documents.

12

Document Classification Using N-gram and Word Semantic Similarity

Mei-ying Ren, Sinjae Kang

보안공학연구지원센터(IJUNESST) International Journal of u- and e- Service, Science and Technology Vol.8 No.8 2015.08 pp.111-118

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

This paper mainly conducted two series of experiments. One is investigation between language dependent and language independent features. Bi-grams in Korean experiments and uni-grams in Chinese contributed most as basic features. And another one is utilization of Korean WordNet to improve the performance of Korean document classification. Korean WordNet is a Korean Lexical Semantic Network. Language independent features seem can lead better performance and stable. The performance of Korean text classification was improved by using Korean WordNet.

13

WordNet-based Hybrid VSM for Document Classification SCOPUS

Luda Wang, Peng Zhang, Shouping Gao

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.1 2016.01 pp.185-200

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Many text classifications depend on statistical term measures or synsets to implement document representation. Such document representations ignore the lexical semantic contents or relations of terms, leading to losing the distilled mutual information. This work proposed a synthetic document representation method, WordNet-based hybrid VSM, to solve the problem. This method constructed a data structure of semantic-element information to characterize lexical semantic contents, and support disambiguation of word stems. As a template, lexical semantic vector consisting of lexical semantic contents was built in the lexical semantic space of corpus, and lexical semantic relations are marked on the vector. Then, it connects with special term vector to form the eigenvector in hybrid VSM. Applying algorithm NWKNN, on text corpus Reuter-21578 and its adjusted version, the experiments show that the eigenvector performs F1 measure better than document representations based on TF-IDF.

14

A Study on Naïve Bayes Classifier Based Document Classification Scheme with an Apriori Feature Extraction SCOPUS

Jong-Yeol Yoo, Min-Ho Lee, Grace Aloyce, Dong-Min Yang

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.7 2016.07 pp.207-214

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

A document classifier is an essential tool for classifying the various types of documents being generated in the Big Data era. In recent years, the wide variety of information services available for use with smartphones and portable mobile devices (tablets) have provided a technique that efficiently classifies the quality of sorted data. A common type of document classification scheme is the naïve Bayes classifier. The Naïve Bayes scheme is based on performance classification, which varies widely depending on the method of extraction used in the document. In this paper, we propose a system model that offers feature extraction methods which combine frequency with associated words. This model is then applied to the Naïve Bayes classifier to precisely classify documents. This method is proposed as an alternative to using traditional classification techniques. In addition, experiments will be evaluated by the existing document classification techniques and the proposed techniques.

15

The Impact of Feature Reduction Techniques on Arabic Document Classification SCOPUS

Abdullah Ayedh, Guanzheng Tan, Hamdi Rajeh

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.6 2016.06 pp.67-80

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Feature reduction are common techniques that used to improve the efficiency and accuracy of the document classification systems. The problems associated with these techniques are the highly dimensionality of the feature space and The difficulty of selecting the important features for understanding the document in question. The document usually consists of several parts and the important features that more closely associated with the topic of the document are appearing in the first parts or repeated in several parts of the document. Therefore, the position of the first appearance of a word and the compactness of the word considered as factors that determine the important features using the information within a document. This study, explored the impact of combining three feature weighting methods that depend on inverse document frequency (IDF), namely, Term frequency (TFiDF), the position of the first appearance of a word (FAiDF), and the compactness of the word (CPiDF) on the classification accuracy. In addition, we have investigated different feature selection techniques, namely, Information gain (IG), Goh and Low (NGL) coefficients, Chi-square Testing (CHI), and Galavotti-Sebastiani-Simi Coefficient (GSS) in order to improve the performance for Arabic document classification system. Experimental analysis on Arabic datasets reveals that the proposed methods have a significant impact on the classification accuracy, and in most cases the FAiDF feature weighting performed better than CPiDF and TFiDF. The results also clearly showed the superiority of the GSS over the other feature selection techniques and achieved 98.39% micro-F1 value when using a combination of TFiDF, FAiDF, and CPiDF as feature weighting method.

16

Is Naïve Bayes a Good Classifier for Document Classification? SCOPUS

S.L. Ting, W.H. Ip, Albert H.C. Tsang

보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.5 No.3 2011.07 pp.37-46

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Document classification is a growing interest in the research of text mining. Correctly identifying the documents into particular category is still presenting challenge because of large and vast amount of features in the dataset. In regards to the existing classifying approaches, Naïve Bayes is potentially good at serving as a document classification model due to its simplicity. The aim of this paper is to highlight the performance of employing Naïve Bayes in document classification. Results show that Naïve Bayes is the best classifiers against several common classifiers (such as decision tree, neural network, and support vector machines) in term of accuracy and computational efficiency.

17

Feature based Star Rating of Reviews: A Knowledge-Based Approach for Document Sentiment Classification

Shaishav Agrawal, Tanveer J. Siddiqui

보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.5 No.4 2012.10 pp.95-110

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

This paper presents a novel knowledge-based approach for star rating of reviews. It uses SentiWordNet and linguistic heuristics to determine sentiment orientation of sentences, which is used to assign a positive, negative and objective score to document to achieve 5-star rating of movie reviews. A method for generating ratings based on individual features is also presented. The experimental results on sentiment scale dataset demonstrate the effectiveness of our approach.

18

점증적으로 변화하는 데이터를 이용한 지능화된 웹 정보 분류 방법에 관한 연구

박길철, 박성식, 김양석, 강병호

보안공학연구지원센터(JSE) 보안공학연구논문지 Vol.2 No.3 2005.11 pp.186-192

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

19

특징선택과 특징가중의 융합을 통한 웹문서분류 성능의 개선 KCI 등재

이아람, 김한준, 현만

국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제13권 제4호 2013.08 pp.141-148

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

기계학습을 이용한 자동분류시스템은 학습과정을 통해 분류모델을 구축하고 이를 기반으로 미분류 데이터를 특정 카테고리로 분류한다. 기계학습 기반 자동분류 시스템의 성능은 분류모델의 구성 인자인 특징의 품질에 크게 의 존한다. 문서 데이터의 경우 특징 집합을 생성하기 위해 문서내의 출현단어와 문서의 구조적 정보를 활용한다. 특히 웹문서로부터 특징을 추출하기 위해 단어뿐만 아니라 태그, 하이퍼링크 정보를 분석할 수 있다. 최근 웹문서의 분류 기법에 대한 연구는 기계학습 알고리즘보다 특징 생성 및 가공 기술에 초점을 맞추고 있다. 이에 본 논문은 웹문서의 분류모델을 개선하기 위해 단어, 태그, 하이퍼링크 정보로부터 고품질의 특징을 선별 추출하여 가중치를 자동으로 부 여하는 기법을 제안한다. Web-KB 문서집합을 이용한 다양한 실험을 통해 제안 기법의 우수성을 보인다.

Automated classification systems which utilize machine learning develops classification models through learning process, and then classify unknown data into predefined set of categories according to the model. The performance of machine learning-based classification systems relies greatly upon the quality of features composing classification models. For textual data, we can use their word terms and structure information in order to generate the set of features. Particularly, in order to extract feature from Web documents, we need to analyze tag and hyperlink information. Recent studies on Web document classification focus on feature engineering technology other than machine learning algorithms themselves. Thus this paper proposes a novel method of incorporating feature selection and weighting which can improves classification models effectively. Through extensive experiments using Web-KB document collections, the proposed method outperforms conventional ones.

20

단어패턴 빈도를 이용한 단문 오피니언 문서 분류기법의 실험적 평가 KCI 등재

장재영, 김일민

국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제12권 제5호 2012.10 pp.243-253

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

데이터 마이닝의 문서분류 기술에서 발전된 오피니언 마이닝은 이제 국외뿐만 아니라 국내 산업에서 중요한 관심분야로 자리잡아가고 있다. 오피니언 마이닝의 핵심은 문서에서 감정 단어를 추출하여 긍정/부정 여부를 얼마나 정확하게 판별하느냐를 평가하는 것이다. 국내에서도 이에 관련된 많은 연구가 이루어 졌으나 아직 실용적으로 적용 할 만큼의 분류 정확도를 보이지 않고 있다. 한국어의 경우 비문법적 표현, 감정단어의 다양성 등으로 인해 문서의 극성을 판별하기가 쉽지 않기 때문이다. 본 논문에서는 문법적 요소를 최대한 배제하고 단어패턴의 빈도만을 고려한 새로운 오피니언 문서 분류기법을 제안한다. 제안된 방법에서는 문서를 단어들의 리스트로 추상화한 후, 패턴들의 빈 도를 이용하여 기계학습 알고리즘을 적용한다. 이후에 적절한 스코어 함수를 적용하여 문서의 극성을 판별한다. 또한 제안된 기법의 정확도를 평가하기 위해서 실험결과를 제시한다.

An opinion mining technique which was developed from document classification in area of data mining now becomes a common interest in domestic as well as international industries. The core of opinion mining is to decide precisely whether an opinion document is a positive or negative one. Although many related approaches have been previously proposed, a classification accuracy was not satisfiable enough to applying them in practical applications. A opinion documents written in Korean are not easy to determine a polarity automatically because they often include various and ungrammatical words in expressing subjective opinions. Proposed in this paper is a new approach of classification of opinion documents, which considers only a frequency of word patterns and excludes the grammatical factors as much as possible. In proposed method, we express a document into a bag of words and then apply a learning algorithm using a frequency of word patterns, and finally decide the polarity of the document using a score function. Additionally, we also present the experiment results for evaluating the accuracy of the proposed method.

 
1 2 3 4 5
페이지 저장