Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 2
No
1

Keybert와 Bertopic을 활용한 텍스트마이닝 연구동향 분석 KCI 등재

길완제, 장환영, 신인수

한국EA학회 정보화연구 제21권 2호 2024.06 pp.159-169

※ 기관로그인 시 무료 이용이 가능합니다.

4,200원

대량의 텍스트 데이터를 분석하고 활용하는 텍스트 마이닝 기법은 공학 분야뿐만 아니라 사회 과학과 교육 등 거의 모든 학문 분야에서 널리 사용되고 있다. 특히 최근 대규모 언어 모델의 급속한 발전은 기존 텍스트 마이닝 기법의 한계를 보완하는 혁신적인 방법들을 도입하는 데 기여하고 있다. 본 연구의 목적은 국내 학술 및 학위 논문을 수집하여 최신 텍스트 마이닝 기법을 활용해 분석하는 것 이다. 이를 위해 학술연구정보서비스(RISS) 데이터베이스에서 ‘텍스트 마이닝’을 키워드로 논문을 수 집하였고, 수집된 논문들에 대해 키워드 분석과 토픽 모델링을 수행하였다. 키워드 분석에서는 TFIDF를 활용한 빈도 기반 분석과 BERT 기반의 KeyBERT를 활용한 분석을 비교하였다. 또한, 토픽 모델링 분석에서는 기존 통계 기반의 LDA 기법과 최신 언어 모델인 BERT 기반의 토픽 모델링 기 법인 BERTopic을 비교하였다. 그 결과, BERT 기반의 토픽 분석이 응집도(Coherence Score) 점수 에서 보다 우수한 성능을 나타냈다. 특히, Bertopic에서 한국어 임베딩 모델과 Keybert 기반의 토픽 추출이 다국어 모델과 문장 기반의 추출보다 더 높은 응집도 점수를 기록하였다. 본 연구는 이러한 결 과를 통해 한국어 텍스트 마이닝에서 최신 기법들의 적용과 활용 가능성을 제시하고자 한다.

Text mining techniques for analyzing and utilizing large-scale text data are widely used not only in engineering but also in social sciences, education, and almost all academic fields. This study aims to collect and analyze domestic academic and thesis papers on text mining using the latest text mining techniques. For this purpose, papers were collected from the Research Information Sharing Service (RISS) database using the keyword ‘text mining’, and keyword analysis and topic modeling were conducted on the collected papers. In keyword analysis, frequency-based analysis using TF-IDF and analysis using BERT-based KeyBERT were compared. Additionally, in topic modeling analysis, traditional statistical-based LDA techniques were compared with BERTopic, a topic modeling technique based on the latest BERT language model. The results showed that BERT-based topic analysis demonstrated superior performance in terms of coherence score. Particularly, the topic extraction based on Korean embedding models and Keybert recorded higher coherence scores compared to those based on multilingual models and sentence-based extraction. Through these findings, this study aims to present the applicability and potential of the latest techniques in Korean text mining.

2

Optimizing Information Retrieval in Dark Web Academic Literature : A Study Using KeyBERT for Keyword Extraction and Clustering

Yosua Setyawan Soekamto, Leonard Christopher Limanjaya, Yoshua Kaleb Purwanto, Bongjun Choi, Seung-Keun Song, Dae-Ki Kang

국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.16 No.4 2024.12 pp.203-208

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The exponential increase in publications and the interconnected nature of sub-domains make traditional methods of information extraction and organization inadequate. This inefficiency can impede scientific progress and innovation. To address these challenges, this research leverages the ability of Bidirectional Encoder Representations from Transformers for keyword extraction (KeyBERT) and integrates with K-Means clustering to organize topics from large datasets effectively. Analyzing a dataset of 47,627 articles from SCOPUS in the domains of Reinforcement Learning and Computer Vision. An ablation study demonstrates the generalizability of the approach across these fields, with the optimal number of clusters determined to be three using the Elbow Method. The results demonstrate that KeyBERT is effective in extracting and organizing topics within these domains, with a particular focus on applications such as medical imaging, autonomous driving, and real-time detection systems. This methodology offers a scalable solution for organizing vast academic datasets, enabling researchers to extract meaningful insights efficiently and apply this approach to other domains.

 
페이지 저장