Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 16
No
1

A graph mining-based approach to analyze the dynamics of the Twitter community of COVID-19 misinformation disseminators

Hussna Asma Ul, Islam Risul, Alam Md Golam Rabiul, UDDINJIA, 아시라프임란, 사마드 엠디 압두스

[NRF 연계] 한국통신학회 ICT Express Vol.10 No.6 2024.12 pp.1280-1287

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This study explores the global problem of misinformation dissemination on social media, particularly Twitter, due to the COVID-19 pandemic. It identifies prominent disseminators, investigates the spread of false information and the ecosystem of disinformation spreaders, and assesses their online personalities. We track the interaction among fake news spreaders using the User?User Interaction Graph. The study reveals a rapidly growing population of disseminators, including professional spreaders, with over 3% dominating the others. The collaboration among fake news spreaders is high, highlighting the need for further research using publicly available online data to understand the community spreading malicious misinformation about COVID-19.

2

4,000원

1990년대 초반 데이터마이닝 방법론이 등장한 이후 현재까지 정확성과 신속성을 향상시키기 위한 기법들이 꾸준히 연구되고 있다. 본 연구에서는 데이터마이닝 방법론 중 하나인 연관규칙 마이닝 과정에서 작업 수행 시간을 최소화하기 위하여 그래프 구조를 사용하였다. 본 연구에서 제시한 알고리즘의 특징은 다음과 같다. 첫째, 데이터베이스 스캐닝을 하여 연관규칙을 찾는데 필요한 모든 정보를 그래프에 저장하였다. 둘째, 트랜잭션 그룹핑을 통해 동일한 레코드 값에 대한 접근 횟수를 최소화하였다. 셋째, 동일한 데이터에서 파라메터의 값만 변경하여 새로 연관규칙을 찾으려 하는 경우, 기 생성된 그래프의 데이터를 재사용함으로써 연관규칙을 찾는 시간을 최소화하였다. 넷째, 빈발항목집합을 찾는 과정에서 후보항목집합의 생성을 배제하였다. 마지막으로, 알고리즘의 성능을 측정하기 위하여 전통적인 연관규칙분석 방법인 Apriori Algorithm과 메모리 기반의 데이터 클러스터링 방법을 사용한 Cluster-based Association Rule Algorithm을 구현하여 비교한 결과, 본 연구에서 제시한 알고리즘이 기존 연구보다 효율성이 높게 나타남을 확인하였다.

3

그래프마이닝을 활용한 빈발 패턴 탐색에 관한 연구

홍준석

한국정보기술응용학회 JITAM Vol.26 No.1 2019.02 pp.65-75

※ 기관로그인 시 무료 이용이 가능합니다.

4,200원

As the use of semantic web based on XML increases in the field of data management, a lot of studies to extract useful information from the data stored in ontology have been tried based on association rule mining. Ontology data is advantageous in that data can be freely expressed because it has a flexible and scalable structure unlike a conventional database having a predefined structure. On the contrary, it is difficult to find frequent patterns in a uniformized analysis method. The goal of this study is to provide a basis for extracting useful knowledge from ontology by searching for frequently occurring subgraph patterns by applying transaction-based graph mining techniques to ontology schema graph data and instance graph data constituting ontology. In order to overcome the structural limitations of the existing ontology mining, the frequent pattern search methodology in this study uses the methodology used in graph mining to apply the frequent pattern in the graph data structure to the ontology by applying iterative node chunking method. Our suggested methodology will play an important role in knowledge extraction.

4

특허 문서로부터 키워드 추출을 위한 위한 텍스트 마이닝 기반 그래프 모델 KCI 등재

이순근, 임영문, 엄완섭

대한안전경영과학회 대한안전경영과학회지 제17권 제4호 2015.12 pp.335-342

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

The increasing interests on patents have led many individuals and companies to apply for many patents in various areas. Applied patents are stored in the forms of electronic documents. The search and categorization for these documents are issues of major fields in data mining. Especially, the keyword extraction by which we retrieve the representative keywords is important. Most of techniques for it is based on vector space model. But this model is simply based on frequency of terms in documents, gives them weights based on their frequency and selects the keywords according to the order of weights. However, this model has the limit that it cannot reflect the relations between keywords. This paper proposes the advanced way to extract the more representative keywords by overcoming this limit. In this way, the proposed model firstly prepares the candidate set using the vector model, then makes the graph which represents the relation in the pair of candidate keywords in the set and selects the keywords based on this relationship graph.

5

Mining Strongly Correlated Sub-graph Patterns by Considering Weight and Support Constraints SCOPUS

Gangin Lee, Unil Yun

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.8 No1 2013.01 pp.197-206

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Frequent graph mining is one of famous data mining fields that receive the most attention, and its importance has been raised continually as recent databases in the real world become more complicated. Weighted frequent graph mining is an approach for applying importance of objects in the real world to the graph mining, and numerous studies related to this have been conducted so far. However, all of the results obtained from this approach do not become actually useful information, and a significant portion of them may be meaningless ones even though they are weighted frequent sub-graph patterns. To overcome this problem, in this paper, we propose a novel method which can consider whether any sub-graph pattern has close correlation among elements in the pattern, called MSCG (Mining Strongly Correlated sub-Graph). In experimental results, we demonstrate that our MSCG outperforms a state-of-the-art method with respect to runtime and memory usage.

6

Research Trends on Graph-Based Text Mining

Jae-Young Chang, Il-Min Kim

보안공학연구지원센터(IJSH) International Journal of Smart Home Vol.8 No.4 2014.07 pp.37-46

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Since text mining has been assumed to apply for unformatted text (document), it is necessary to represent text with simplified models. One of the most commonly used models is the vector space model, in which text is represented as a bag of words. Recently, many researches tried to apply a graph-based text model for representing semantic relationships between words. In this paper, we surveyed research trends of graph-based text representation models for text mining. We summarized the models, their features and forecasted further researches.

7

Research Trends on Graph-Based Text Mining SCOPUS

Jae-Young Chang, Il-Min Kim

보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.8 No.4 2014.04 pp.147-156

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Since text mining has been assumed to apply for unformatted text (document), it is necessary to represent text with simplified models. One of the most commonly used models is the vector space model, in which text is represented as a bag of words. Recently, many researches tried to apply a graph-based text model for representing semantic relationships between words. In this paper, we surveyed research trends of graph-based text representation models for text mining. We summarized the models, their features and forecasted further researches.

8

텍스트 마이닝은 비정형화된 텍스트를 분석하여 그 안에 내재된 패턴, 추세, 분포 등의 고급정보들을 추출하 는 분야이다. 텍스트 마이닝은 기본적으로 비정형 데이터를 가정하므로 텍스트를 단순화된 모델로 표현하는 것이 필 요하다. 현재까지 가장 많이 사용되고 있는 모델은 텍스트를 단순한 단어들의 집합으로 표현한 벡터공간 모델이다. 그러나 최근 들어 단어들의 의미적 관계까지 표현하기 위해 그래프를 이용한 텍스트 표현 모델을 많이 사용하고 있 다. 본 논문에서는 텍스트 마이닝을 위한 기존의 연구 중에서 그래프에 기반한 텍스트 표현 모델의 방법들과 그들의 특징들을 기술한다. 또한 그래프 기반 텍스트 마이닝의 향후 발전방향에 대해서도 논한다.

Text Mining is a research area of retrieving high quality hidden information such as patterns, trends, or distributions through analyzing unformatted text. Basically, since text mining assumes an unstructured text, it needs to be represented as a simple text model for analyzing it. So far, most frequently used model is VSM(Vector Space Model), in which a text is represented as a bag of words. However, recently much researches tried to apply a graph-based text model for representing semantic relationships between words. In this paper, we survey research trends of graph-based text representation models for text mining. Additionally, we also discuss about future models of graph-based text mining.

9

그래프마이닝을 활용한 빈발 패턴 탐색에 관한 연구

홍준석

[Kisti 연계] 한국데이타베이스학회 Journal of information technology applications & management Vol.26 No.1 2019 pp.65-75

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

As the use of semantic web based on XML increases in the field of data management, a lot of studies to extract useful information from the data stored in ontology have been tried based on association rule mining. Ontology data is advantageous in that data can be freely expressed because it has a flexible and scalable structure unlike a conventional database having a predefined structure. On the contrary, it is difficult to find frequent patterns in a uniformized analysis method. The goal of this study is to provide a basis for extracting useful knowledge from ontology by searching for frequently occurring subgraph patterns by applying transaction-based graph mining techniques to ontology schema graph data and instance graph data constituting ontology. In order to overcome the structural limitations of the existing ontology mining, the frequent pattern search methodology in this study uses the methodology used in graph mining to apply the frequent pattern in the graph data structure to the ontology by applying iterative node chunking method. Our suggested methodology will play an important role in knowledge extraction.

10

그래프 마이닝에서 그래프 동형판단연산의 향상기법

노영상, 윤은일, 김명준

[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.14 No.10 2009 pp.251-258

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

그래프마이닝에서 그래프패턴의 동형판단문제는 지수함수적 계산시간을 요구하기 때문에 그래프마이닝의 전체수행시간에서 동형판단 연산이 차지하는 비율이 매우 높다. 그러므로 그래프마이닝 알고리즘은 그래프동형판단을 최대한 효율적으로 할 필요가 있다. 본 논문은 그래프마이닝에서 빠른 수행시간을 보이는 gaston 알고리즘의 동형판단효율성을 증가시켜 수행시간을 평가해 보았으며, 제시한 방법으로 인해 더욱 향상된 성능을 보인다.

Data mining is a method that extract useful knowledges from huge size of data. Recently, a focussing research part of data mining is to find interesting patterns in graph databases. More efficient methods have been proposed in graph mining. However, graph analysis methods are in NP-hard problem. Graph pattern mining based on pattern growth method is to find complete set of patterns satisfying certain property through extending graph pattern edge by edge with avoiding generation of duplicated patterns. This paper suggests an efficient approach of reducing computing time of pattern growth method through pattern growth's property that similar patterns cause similar tasks. we suggest pruning methods which reduce search space. Based on extensive performance study, we discuss the results and the future works.

11

텍스트 마이닝 기반의 그래프 모델을 이용한 미발견 공공 지식 추론

허고은, 송민

[Kisti 연계] 한국정보관리학회 정보관리학회지 Vol.31 No.1 2014 pp.231-250

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

정보통신기술의 발달로 학술 정보의 양이 기하급수적으로 증가하였고 방대한 양의 텍스트 데이터를 처리하기 위한 자동화된 텍스트 처리의 필요성이 대두되었다. 생의학 문헌에서 생물학적 의미와 치료 효과 등에 대한 정보를 발견해내는 바이오 텍스트 마이닝은 문헌 내의 각 개념들 간의 유의미한 연관성을 발견하여 의학 영역에서 상당한 시간과 비용을 줄여준다. 문헌 기반 발견 연구로 새로운 생의학적 가설들이 발견되었지만 기존의 연구들은 반자동화된 기법으로 전문가의 개입이 필수적이며 원인과 결과의 한가지의 관계만을 밝히는 제한점이 있다. 따라서 본 연구에서는 중간 개념인 B를 다수준으로 확장하여 다양한 관계성을 동시출현 개체와 동사 추출을 통해 확인한다. 그래프 기반의 경로 추론을 통해 각 노드 사이의 관계성을 체계적으로 분석하여 규명할 수 있었으며 새로운 방법론적 시도를 통해 기존에 밝혀지지 않았던 새로운 가설 제시의 가능성을 기대할 수 있다.

Due to the recent development of Information and Communication Technologies (ICT), the amount of research publications has increased exponentially. In response to this rapid growth, the demand of automated text processing methods has risen to deal with massive amount of text data. Biomedical text mining discovering hidden biological meanings and treatments from biomedical literatures becomes a pivotal methodology and it helps medical disciplines reduce the time and cost. Many researchers have conducted literature-based discovery studies to generate new hypotheses. However, existing approaches either require intensive manual process of during the procedures or a semi-automatic procedure to find and select biomedical entities. In addition, they had limitations of showing one dimension that is, the cause-and-effect relationship between two concepts. Thus;this study proposed a novel approach to discover various relationships among source and target concepts and their intermediate concepts by expanding intermediate concepts to multi-levels. This study provided distinct perspectives for literature-based discovery by not only discovering the meaningful relationship among concepts in biomedical literature through graph-based path interference but also being able to generate feasible new hypotheses.

12

온라인 리뷰 텍스트마이닝과 GAT 분석을 활용한 애슬레저 의류 소비자의 만족 및 재구매 요인 연구

이홍주, 김가원, 정수인

[NRF 연계] 한국고객만족경영학회 고객만족경영연구 Vol.27 No.4 2025.12 pp.155-170

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

온라인 시장의 확산과 함께 소비자들의 온라인 리뷰(review)는 중요한 소비 정보로 활용되고 있다. 이에, 본 연구는 이러한 소비환경 변화를 고려하여, 소비자의 실제 경험에 기반한 만족 요인과 재구매 요인을 연구하였다. 이를 위해 국내 대표 온라인 플랫폼인 N사의 쇼핑몰에서 애슬레저 의류 리뷰 12,442건을 수집하고, 텍스트마이닝과 GAT(Graph Attention Network)기반의 분석을 수행하였다. 특히 설문조사 방식이 가지는 사전 정의 변수의 한계를 보완하기 위해, 리뷰 문장을 SBERT 임베딩과 코사인 유사도 기반 K-NN 그래프로 구조화하고, GAT 모델을 통해 문장 간 의미적 관계와 주요 요인의 기여도를 정량적으로 도출하였다. 연구 결과, 재구매 의도는 ‘사이즈 조정(한 치수, 작게)’과 ‘가격 대비 가치(이 가격에)’가 핵심 요인으로 나타났으며, 실제 리뷰에서도 초기 착용 경험을 바탕으로 치수를 조정해 재구매하는 패턴이 확인되었다. 특히, 만족 요인의 경우 ‘편안함(편하고)’이 가장 높은 기여도를 보였고, ‘사이즈’, ‘핏도’, ‘착용감’ 등 신체 적합성이 중요한 요인으로 도출되었다. 또한 기존 애슬레저 연구에서 다뤄지지 않았던 ‘빠른 배송’ 요인이 상위에 포함되면서, 온라인 쇼핑 환경에서 배송 경험이 소비자 만족의 핵심 요소로 부상하고 있음을 확인하였다. 아울러 ‘가성비’는 만족과 재구매 두 영역에서 공통적으로 중요한 역할을 하는 요인으로 나타났다. 본 연구에서는 온라인 리뷰에 기반한 딥러닝 분석을 통해 실제 소비 경험에서 비롯된 핵심 요인을 체계적으로 규명함으로써, 설문 중심 연구가 포착하기 어려웠던 새로운 재구매 요인을 분석할 수 있었다. 아울러, 기업 관점에서 사이즈 옵션 세분화, 기능성 소재 개선, 가격–품질 균형 전략, 배송 품질 관리 등을 통해 소비자 만족과 재구매를 높일 수 있는 전략적 시사점을 도출하였다.

With the expansion of online markets, consumer-generated online reviews have become an important source of consumption-related information. Considering this shift in the consumption environment, the present study investigates satisfaction factors and repurchase drivers based on consumers’ actual product experiences. To this end, 12,442 athleisure wear reviews were collected from N Company’s major online shopping platform in Korea, and text mining combined with a Graph Attention Network (GAT) model was employed. To overcome the limitations of survey-based research?which relies on predefined variables?this study transformed review sentences into a K-NN graph using SBERT embeddings and cosine similarity, and then applied a GAT model to quantify semantic relationships between sentences and identify the relative contribution of key factors. The results show that the major drivers of repurchase intention were ‘size adjustment(one size up/down)’ and ‘value for money(for this price).’ Review content further confirmed that many consumers repurchased items after adjusting the size based on their initial wearing experience. For satisfaction factors, ‘comfort (comfortable)’ exhibited the highest contribution, followed by ‘size,’ ‘fit,’ and ‘wearability,’ indicating that body-fit?related attributes are critical determinants of satisfaction. Notably, ‘fast delivery,’ a factor rarely addressed in prior athleisure research, emerged as a top contributor, highlighting the increasing importance of delivery experience in online shopping contexts. Additionally, ‘value for money’ played a consistently important role in both satisfaction and repurchase domains. By applying deep learning analysis to large-scale online reviews, this study systematically identified key drivers grounded in actual consumption experiences, thereby uncovering repurchase determinants that are difficult to capture through traditional survey methods. From a managerial perspective, the findings provide actionable insights for improving consumer satisfaction and repurchase rates through size-option diversification, functional material enhancement, balanced price?quality positioning, and strengthened delivery service quality.

13

토픽 모델링을 이용한 댓글 그래프 기반 소셜 마이닝 기법

이상연, 이건명

[Kisti 연계] 한국지능시스템학회 Journal of Korean Institute of Intelligent Systems Vol.24 No.6 2014 pp.640-645

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

인터넷 상에서 많은 사람들은 사용자 간의 의사소통과 정보 공유, 사회적 관계를 생성하기 위한 방법으로 소셜 네트워크 서비스를 이용한다. 그 중 대표적인 트위터는 하루에 수백만 건의 소셜 데이터가 발생하기 때문에 수집되고 있는 데이터의 양이 엄청나다. 이 방대한 양의 데이터로부터 의미 있는 정보를 추출하는 소셜 마이닝이 집중적으로 연구되고 있다. 트위터는 일반적으로 유용한 정보 혹은 공유하고자 하는 내용을 팔로잉-팔로워 관계를 이용해 쉽게 전달하고 리트윗할 수 있다. 소셜 미디어에서 트윗 데이터에 대한 토픽 모델링은 이슈를 추적하기 위한 좋은 도구이다. 짧은 텍스트 기반인 트윗 데이터의 제한점을 극복하기 위해, 사용자를 노드로 사용자간 댓글과 리트윗 메시지의 여부를 간선으로 하는 그래프 구조를 갖는 댓글 그래프의 개념을 소개한다. 토픽 모델링의 대표적인 방법인 LDA 토픽 모델이 짧은 텍스트 데이터에 대해 비효율적인 것을 보완하기 위한 방법으로, 이 논문에서는 짧은 문서의 수를 줄이고 마이닝 결과의 질을 향상시키기 위한 댓글 그래프를 사용하는 토픽 모델링 방법을 소개한다. 제안한 모델은 토픽 모델링 방법으로 LDA 모델을 사용하였으며, 7일간 수집한 트윗 데이터에 대한 실험 결과를 보인다.

Many people use social network services as to communicate, to share an information and to build social relationships between others on the Internet. Twitter is such a representative service, where millions of tweets are posted a day and a huge amount of data collection has been being accumulated. Social mining that extracts the meaningful information from the massive data has been intensively studied. Typically, Twitter easily can deliver and retweet the contents using the following-follower relationships. Topic modeling in tweet data is a good tool for issue tracking in social media. To overcome the restrictions of short contents in tweets, we introduce a notion of reply graph which is constructed as a graph structure of which nodes correspond to users and of which edges correspond to existence of reply and retweet messages between the users. The LDA topic model, which is a typical method of topic modeling, is ineffective for short textual data. This paper introduces a topic modeling method that uses reply graph to reduce the number of short documents and to improve the quality of mining results. The proposed model uses the LDA model as the topic modeling framework for tweet issue tracking. Some experimental results of the proposed method are presented for a collection of Twitter data of 7 days.

14

그래프를 이용한 빈발 서비스 탐사

황정희

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.19 No.3 2018 pp.471-477

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

시간의 변화에 따라 사용자의 관심도는 변화한다. 이 논문에서는 유비쿼터스 환경에서 연령, 시기, 계절 등에 따라 변화하는 사용자의 서비스 관심도를 고려하기 위하여 서비스에 대한 관심도를 동적 가중치로 부여하여 사용자에게 적합한 서비스를 추천하기 위한 방법을 제안한다. 사용자에게 제공한 서비스 이력 데이터를 기준으로 시기나 연령에 따른 일반적인 서비스 규칙을 저장하고, 실시간으로 변화하는 서비스의 관심도를 고려한 최신의 서비스 규칙을 지속적으로 추가하여 사용자의 관심 변화를 반영하는 서비스를 제공하기 위한 방법이다. 이를 위해 사용자에게 제공하는 일련의 서비스는 트랜잭션으로 고려하고 서비스는 항목으로 고려하여 서비스의 연관관계를 그래프로 표현하고, 이를 기반으로 빈발 서비스 항목을 발견한다. 발견된 빈발 서비스 항목은 사용자에게 유용한 최신의 정보 서비스를 의미한다.

As time changes, users change their interest. In this paper, we propose a method to provide suitable service for users by dynamically weighting service interests in the context of age, timing, and seasonal changes in ubiquitous environment. Based on the service history data presented to users according to the age or season, we also offer useful services by continuously adding the most recent service rules to reflect the changing of service interest. To do this, a set of services is considered as a transaction and each service is considered as an item in a transaction. And also we represent the association of services in a graph and extract frequent service items that refer to the latest information services for users.

15

웹 마이닝을 위한 웹 문서 하이퍼링크와 웹 접근로그를 통합한 방향그래프

박철현, 이성대, 곽용원, 전성환, 박휴찬

[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2005 pp.16-18

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

웹은 사용자가 원하는 정보를 쉽고 정확하게 검색할 수 있도록 웹 문서를 자료구조화하여 보다 신뢰성 있는 패턴을 추출하고 사용자의 특성과 행동 패턴을 적용하여 개인화 하여야한다. 본 논문에서는 개인화하기 위한 전처리 과정으로서 웹 문서를 구조화 하는 방법을 제안한다. 제안 방법은 기본적으로 웹 문서 태그의 하이퍼링크를 깊이 우선 탐색 알고리즘을 사용하여 방향그래프를 만드는 것이다. 이때 웹 문서 태그 탐색 시 플래시, 스크립트 등의 찾기 힘든 하이퍼링크를 찾는 문제와 '뒤로' 버튼 사용 시 웹 접근로그에 기록되지 않는 문제점을 보완한다. 이를 위해 클릭 스트림을 스택에 저장하여 이미 만들어진 방향그래프와 비교하여 새롭게 찾은 정점과 간선을 추가함으로써 보다 신뢰성높은 방향그래프를 만든다.

16

텍스트 마이닝을 활용한 감정 비율 단어 그래프

김장민, 이연동, 조영석

[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.25 No.5 2023.10 pp.1749-1757

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

SNS, 논문, 설문조사 주관식 문항 답변과 같은 자연어로 이루어진 비정형 데이터는 텍스트 마이닝을 이용하여 분석 결과를 비교하거나 시각화하는 경우가 일반적이다. 년, 분기, 월, 요일과 같은 시간을 나타내는 임의의 구간을 설정하여 텍스트 데이터를 분석할 경우 전체 구간 중 어떤 구간에 데이터가 가장 많고 적은지, 전체 구간 중 구간별로 많이 사용된 감정 단어가 무엇인지, 특정 구간에 있는 텍스트 데이터가 상대적으로 얼마큼 많이 긍정보다 부정적으로 작성되었는지 판단해야 할 경우가 있다. 본 연구에서는 2019년부터 2022년까지 “지방대”와 관련된 뉴스 기사를 수집하기 위해 네이버에서 “지방대”라고 검색한 뒤 네이버 뉴스라고 표시된 기사만을 수집하여 위의 세 가지 정보를 한 번에 전달할 수 있는 감정 비율 단어 그래프를 제안한다. 감정 비율 단어 그래프는 텍스트 데이터를 년, 분기, 월, 요일과 같은 시간을 나타내는 임의의 구간 기준으로 나눈 뒤 감성 사전에 있는 감정 점수를 텍스트 데이터에 부여하여 만들어진 그래프이다. 감정 비율 단어 그래프를 시각화할 때 파이계수도 같이 활용하여 단어를 표시한다면 특정 구간에서 감정 단어와 관련성이 가장 큰 단어가 무엇인지에 대한 정보를 추가로 전달할 수 있다.

Unstructured data consisting of natural language such as SNS, papers, and questionnaire subjective question answers are generally compared or visualized using text mining. When analyzing text data by setting a random interval representing a time such as year, quarter, month, and day, it may be necessary to determine which interval has the most data, which sentiment words are used a lot for each interval, and how much text data in a particular interval is written negatively than positive. As a way to solve this problem, this study proposes an sentiment ratio word graph that can deliver the above three information at once. An sentiment ratio word graph is a graph created by dividing text data by a random interval standard representing time such as year, quarter, month, and day of the week and then assigning the sentiment score in the sentiment dictionary to the text data. When visualizing an sentiment ratio word graph, if you also use the pie coefficient to display words, you can further convey information about which words are most relevant to the sentiment word in a particular interval.

 
페이지 저장