Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 96
No
1

5,500원

In this paper, as one of the attempts to effectively utilize the vast amount of text data, I have introduced a text mining technique called Topic Model into the field of Japanese studies. Concretely, the texts of the reading comprehension parts of the previous format JLPT for the past 20 years were collected, and Topic Model analysis was carried out. The following points were made clear by such a study. First of all, it was confirmed from actual data that the subjects of the previous format JLPT tried to avoid topic-specific biases when selecting and producing the texts for the questions. Next, the text can be statistically classified into four main topics: “Private relationships such as family and work,” “Communications related to schedules,” “Public relations related to the country and society,” and “Economic activity.” The techniques and results of topic model analysis in this paper were empirical analyzes of actual existing questions. It is considered significant in that it can be applied to all fields of Japanese studies that are needed. Of course, the discussion in this paper is limited to the texts of the previous format JLPT, not the new format JLPT, and the amount of data is relatively small, although it covers all the data for the past 20 years. In addition, a comparative analysis with other texts was not possible. Therefore, it seems that there is still room for improvement in this paper, but I would like to address this as a future issue.

2

4,000원

본 연구는 노인일자리사업의 사회적 논의구조를 분석하기 위해 대표적인 대중매체인 신문기사에서 다루어지는 노인일자리 관련 주요 토픽들과 시계열적 특성을 분석하였다. 이를 위해 뉴스 통합 데이터베이스인 빅카인즈에 수록된 11개 중앙지와 8개 경제지의 노인일자리사업 관련 기사 1107개에 대해 잠재디리클레할당 방법을 이용한 토픽분석을 실시해 언론 기사에 내재된 노인일자리사업의 잠재토픽을 추출하였다. 분석결과 노인일자리사업에 대한 일반적 정보전 달, 지자체 사업 홍보, 노후생활, 고용효과, 시장연계 등 5개의 잠재토픽이 추출되었는데 2015년까지 대부분의 언론 기사가 일반적 정보전달과 지자체 사업홍보에 국한되어 있어 노인일자리사업의 정체성에 대한 사회적 논의가 형성되지 못하였음을 알 수 있었던 반면 2015년 이후부터 노인일자리사업의 소득, 안전 등 노후생활 효과 관련 주제가 다루어지 는 비중이 증가했으며 특히 문재인 정부 출범이후 고용효과와 관련된 기사가 압도적인 비중을 차지하게 되었음을 발견 할 수 있었다. 본 연구는 이러한 결과에 근거해 향후 노인일자리사업의 질적측면 및 고용효과 측면을 증진시킬 수 있는 방안에 대한 고민의 필요성과 고용프레임 이외의 대안적 프레임 제시의 필요성을 제안하였다.

This study aims to find the structure of social disussion on government ‘Senior job program’ by analyzing 1107 newspaper articles on ‘senior job program’ from 11 major newspaper articles and 8 financial newspapers. Topic modeling via latent dirichlet allocation model was employed for analysis and as result, 5 latent topics were extracted as follows : general information, local government project propaganda, senior life related issues, employment creation effect and market relations. Until 2015, most of the articles focused on the first two topics, indicating not much discourse was formed concerning the characteristics of the program. However, after 2015, the third topic started to increase and after the launch of Moon Jae In government, there has been a drastic increase in the employment creation related topic indicating that current social discourse mirrored by the media is definitely focused on employment creation aspect of senior job program. Based on the result, this study suggests the necessity to increase the quality and also enhance employment aspects of Senior job program.

3

4,000원

As block chain is considered as a new technology for transforming business and government, many attention have been paid to the advantages that block chains offer. However, there is still a lack of research on technological characteristics and direction block chain as a foundation for innovative digital businesses. Therefore, this research analyzes the existing literature on this issue based on a text mining semi-automated approach in order to identify the main trends. This research helps predict potential changes in block chain technology.

4

Discovering Community Interests Approach to Topic Model with Time Factor and Clustering Methods

Ho, Thanh, Thanh, Tran Duy

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.17 No.1 2021 pp.163-177

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Many methods of discovering social networking communities or clustering of features are based on the network structure or the content network. This paper proposes a community discovery method based on topic models using a time factor and an unsupervised clustering method. Online community discovery enables organizations and businesses to thoroughly understand the trend in users' interests in their products and services. In addition, an insight into customer experience on social networks is a tremendous competitive advantage in this era of ecommerce and Internet development. The objective of this work is to find clusters (communities) such that each cluster's nodes contain topics and individuals having similarities in the attribute space. In terms of social media analytics, the method seeks communities whose members have similar features. The method is experimented with and evaluated using a Vietnamese corpus of comments and messages collected on social networks and ecommerce sites in various sectors from 2016 to 2019. The experimental results demonstrate the effectiveness of the proposed method over other methods.

5

5,100원

이 연구는 10·29 이태원 참사와 관련된 트윗을 대량으로 수집하고 정제한 후 DMR 토픽모델링으로 분석하였다. 모델 검증을 통해 가장 적절한 토픽의 수로 15개를 설정하였다. 총 15개 토픽을 클러스 터링하여 정보 반응, 감정 반응, 정부 비판 반응. 정치적 대립 반응, 2차 피해 반응, 공감 반응 및 기타 등 7개 그룹으로 묶을 수 있었다. 토픽의 분포를 시계열적으로도 살펴보았는데, ‘대책 미비’, ‘책임’, ‘애도와 추모’ 토픽의 비중은 시간이 지날수록 감소한 반면, ‘막말 논란’, ‘분노와 비난’은 시간이 지날수록 증가하였다. 이와 함께 전체 기간을 3개의 시기로 나누어 토픽의 차이를 살펴본 결과 시기별로 토픽 비중의 차이가 나타났는데 ‘애도기’와 비교해 볼 때 ‘수사기’와 ‘국정 조사기’에 서는 ‘수사와 구속’, ‘진상 규명’ 토픽의 비중이 증가하는 것을 볼 수 있었다. 이러한 토픽의 구조와 토픽의 시간별 변화에 따른 분석 결과에 기반하여 재난의 정치화와 공동체의 회복에 관점에서 연구 결과의 함의를 논의하였다.

This research involved the extensive collection and refinement of tweets related to the 10·29 Itaewon Tragedy, followed by analysis using DMR topic modeling. Through model validation, the optimal number of topics was determined to be 15, clustered into seven groups: informational, emotional, government criticism, political confrontation, secondary harm, empathetic responses, and miscellaneous. Examining the temporal distribution of topics revealed a decrease in the weight of topics such as 'Lack of government preparedness,' 'Accountability for the disaster,' and 'Condolences and memorials for victims' over time, while topics like 'Controversy over disrespect for victims' families' and 'Anger and Blame' increased. Further, dividing the overall period into three phases demonstrated variations in topic weights. Compared to the 'Mourning' phase, the 'Investigation' and 'Parliamentary Inquiry' phases indicated an increase in the weight of topics related to 'Investigation and detainment of those responsible' and 'Demand for the truth about the disaster.' The research concludes by discussing the implications for the politicization of disasters and community recovery.

6

The Topic Detection and Tracking with Topic Sensitive Language Model

Ruifang He, Bing Qin, Ting Liu, Sheng Li

한국어정보학회 한국어정보학 제8권 1호 2006.06 pp.64-68

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

In this paper, we explore the language model with topic sensitive features for the topic detection and tracking, formulate the relationship among the Chinese internet new words, language model with topic sensitive feature and the scheduling logic and the interval temporal reasoning and the key techniques. we use the Chinese internet new words to strengthen the detection and tracking of the topic and try to employ the scheduling logic and interval temporal reasoning to educe the reciprocal influences of events. At last we summarize the potential issues and the future work.

7

4,000원

8

LDA와 BERTopic을 이용한 토픽모델링의 증강과 확장 기법 연구

김선욱, 양기덕

[Kisti 연계] 한국정보관리학회 정보관리학회지 Vol.39 No.3 2022 pp.99-132

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구의 목적은 LDA 토픽모델링 결과와 BERTopic 토픽모델링 결과를 합성하는 방법론인 Augmented and Extended Topics(AET)를 제안하고, 이를 사용해 문헌정보학 분야의 연구주제를 분석하는 데 있다. AET의 실제 적용결과를 확인하기 위해 2001년 1월부터 2021년 10월까지의 Web of Science 내 문헌정보학 학술지 85종에 게재된 학술논문 서지 데이터 55,442건을 분석하였다. AET는 서로 다른 토픽모델링 결과의 관계를 WORD2VEC 기반 코사인 유사도 매트릭스로 구축하고, 매트릭스 내 의미적 관계가 유효한 범위 내에서 매트릭스 재정렬 및 분할 과정을 반복해 증강토픽(Augmented Topics, 이하 AT)을 추출한 뒤, 나머지 영역에서 코사인 유사도 평균값 순위와 BERTopic 토픽 규모 순위에 대한 조화평균을 통해 확장토픽(Extended Topics, 이하 ET)을 결정한다. 최적 표준으로 도출된 LDA 토픽모델링 결과와 AET 결과를 비교한 결과, AT는 LDA 토픽모델링 토픽을 한층 더 구체화하고 세분화하였으며 ET는 유효한 토픽을 발견하였다. AT(Augmented Topics)의 성능은 LDA 이상이었으며 ET(Extended Topics)는 일부 경우를 제외하고 대부분 LDA와 유사한 수준의 성능을 나타내었다.

The purpose of this study is to propose AET (Augmented and Extended Topics), a novel method of synthesizing both LDA and BERTopic results, and to analyze the recently published LIS articles as an experimental approach. To achieve the purpose of this study, 55,442 abstracts from 85 LIS journals within the WoS database, which spans from January 2001 to October 2021, were analyzed. AET first constructs a WORD2VEC-based cosine similarity matrix between LDA and BERTopic results, extracts AT (Augmented Topics) by repeating the matrix reordering and segmentation procedures as long as their semantic relations are still valid, and finally determines ET (Extended Topics) by removing any LDA related residual subtopics from the matrix and ordering the rest of them by F<sub>1</sub> (BERTopic topic size rank, Inverse cosine similarity rank). AET, by comparing with the baseline LDA result, shows that AT has effectively concretized the original LDA topic model and ET has discovered new meaningful topics that LDA didn't. When it comes to the qualitative performance evaluation, AT performs better than LDA while ET shows similar performances except in a few cases.

9

WV-BTM: SNS 단문의 주제 분석을 위한 토픽 모델 정확도 개선 기법

송애린, 박영호

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.19 No.1 2018 pp.51-58

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

SNS의 사용자와 데이터량이 폭발적으로 증가함에 따라, SNS 빅 데이터를 기반으로 한 연구들이 활발히 진행되고 있다. 특히 소셜 마이닝 분야에서는 비 분류된 대용량 SNS 텍스트 데이터로부터 각 텍스트 별 유사성을 파악하고, 그로부터 트렌드를 추출하기 위해 대표적인 토픽 모델 기법인 LDA를 사용한다. 그러나 LDA는 단문 데이터에 대하여 비 빈발 단어 출현으로 인한 의미 희박성(semantic sparsity)으로 인해 양질의 주제 추론이 어렵다는 한계를 가진다. BTM 연구는 이와 같은 LDA의 한계점을 두 단어의 조합을 통해 개선하였으나, BTM 또한 조합된 단어 중 높은 빈도수의 단어에 더 큰 영향을 받아 각 주제와의 연관성을 고려한 가중치 계산이 불가능하다는 한계점을 지닌다. 본 논문은 단어 간의 의미적 연관성을 반영함으로써 기존 연구 BTM의 정확도를 개선하는 방안을 모색한다.

As the amount of users and data of NS explosively increased, research based on SNS Big data became active. In social mining, Latent Dirichlet Allocation(LDA), which is a typical topic model technique, is used to identify the similarity of each text from non-classified large-volume SNS text big data and to extract trends therefrom. However, LDA has the limitation that it is difficult to deduce a high-level topic due to the semantic sparsity of non-frequent word occurrence in the short sentence data. The BTM study improved the limitations of this LDA through a combination of two words. However, BTM also has a limitation that it is impossible to calculate the weight considering the relation with each subject because it is influenced more by the high frequency word among the combined words. In this paper, we propose a technique to improve the accuracy of existing BTM by reflecting semantic relation between words.

10

소셜 빅데이터로 알아본 코로나19와 가족생활: 토픽모델 접근

박선영, 이재림

[Kisti 연계] 한국콘텐츠학회 한국콘텐츠학회논문지 Vol.21 No.3 2021 pp.282-300

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구의 목적은 코로나19 확산으로 가족생활에서 급격한 변화가 일어난 1차 확산기에 블로그와 온라인 카페에 게시된 소셜 빅데이터를 분석하여 키워드를 파악하고, 게시글에 잠재된 주요 토픽을 발견하는 것이다. 강화된 사회적 거리두기가 처음 시행되었던 2020년 2월 23일부터 4월 19일까지 네이버와 다음의 블로그 및 카페에 게시된 글 중 '코로나'와 '가족' 또는 '코로나'와 '가정'이 함께 언급된 문서 총 351,734건을 분석하였다. 수집된 데이터는 전처리를 거쳐 텍스트 마이닝 기법으로 분석하였다. TF-IDF 가중치 값을 바탕으로 상위 100개 단어를 살펴보았으며, 잠재디리클레할당 방식의 토픽모델 분석을 통해 총 22개 토픽을 도출하고 토픽명을 부여하였다. 연구결과, 코로나19가 가족의 일상생활에 미친 전방위적 영향이 나타났으며, 특히 식생활, 주거생활, 여가생활, 종교생활, 자녀돌봄, 자녀교육, 가족관계, 가족의례 등에서 변화가 두드러졌다. 더불어, 가족 관련 국내 문헌에서는 잘 논의되지 않던 건강공동체로서의 가족을 시사하는 토픽도 등장하였다.

The purpose of this study was to explore what social media posts tell us about family life during the COVID-19 pandemic by examining the keywords and topics underlying posts on blogs and online forums. Our criteria for web crawling were (a) blog and forum posts on Naver and Daum, the top portal sites in Korea, (b) posts between February 23 and April 19, 2020, the period of the first heightened social distancing orders, and (c) inclusion of "COVID" and "family" or "COVID" and "home." We analyzed 351,734 posts using TF-IDF values and topic modeling based on latent Dirichlet allocation. We identified and named 22 topics including COVID-19 prevention, family infection, family health, dietary life and changes, religious life, stuck at home, postponed school year, family events, travel and vacations, concerns about family and friends, anxiety and stress, disaster and damage, COVID-19 warning text messages, family support policies, Shin-cheon-ji and Daegu. The results show that COVID-19 impacted various domains of family life including health, food, housing, religion, child care, education, rituals, and leisure as well as relationships and emotions.

11

BERTopic을 활용한 불면증 소셜 데이터 토픽 모델링 및 불면증 경향 문헌 딥러닝 자동분류 모델 구축

고영수, 이수빈, 차민정, 김성덕, 이주희, 한지영, 송민

[Kisti 연계] 한국정보관리학회 정보관리학회지 Vol.39 No.2 2022 pp.111-129

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

불면증은 최근 5년 새 환자가 20% 이상 증가하고 있는 현대 사회의 만성적인 질병이다. 수면이 부족할 경우 나타나는 개인 및 사회적 문제가 심각하고 불면증의 유발 요인이 복합적으로 작용하고 있어서 진단 및 치료가 중요한 질환이다. 본 연구는 자유롭게 의견을 표출하는 소셜 미디어 'Reddit'의 불면증 커뮤니티인 'insomnia'를 대상으로 5,699개의 데이터를 수집하였고 이를 국제수면장애분류 ICSD-3 기준과 정신의학과 전문의의 자문을 받은 가이드라인을 바탕으로 불면증 경향 문헌과 비경향 문헌으로 태깅하여 불면증 말뭉치를 구축하였다. 구축된 불면증 말뭉치를 학습데이터로 하여 5개의 딥러닝 언어모델(BERT, RoBERTa, ALBERT, ELECTRA, XLNet)을 훈련시켰고 성능 평가 결과 RoBERTa가 정확도, 정밀도, 재현율, F1점수에서 가장 높은 성능을 보였다. 불면증 소셜 데이터를 심층적으로 분석하기 위해 기존에 많이 사용되었던 LDA의 약점을 보완하며 새롭게 등장한 BERTopic 방법을 사용하여 토픽 모델링을 진행하였다. 계층적 클러스터링 분석 결과 8개의 주제군('부정적 감정', '조언 및 도움과 감사', '불면증 관련 질병', '수면제', '운동 및 식습관', '신체적 특징', '활동적 특징', '환경적 특징')을 확인할 수 있었다. 이용자들은 불면증 커뮤니티에서 부정 감정을 표현하고 도움과 조언을 구하는 모습을 보였다. 또한, 불면증과 관련된 질병들을 언급하고 수면제 사용에 대한 담론을 나누며 운동 및 식습관에 관한 관심을 표현하고 있었다. 발견된 불면증 관련 특징으로는 호흡, 임신, 심장 등의 신체적 특징과 좀비, 수면 경련, 그로기상태 등의 활동적 특징, 햇빛, 담요, 온도, 낮잠 등의 환경적 특징이 확인되었다.

Insomnia is a chronic disease in modern society, with the number of new patients increasing by more than 20% in the last 5 years. Insomnia is a serious disease that requires diagnosis and treatment because the individual and social problems that occur when there is a lack of sleep are serious and the triggers of insomnia are complex. This study collected 5,699 data from 'insomnia', a community on 'Reddit', a social media that freely expresses opinions. Based on the International Classification of Sleep Disorders ICSD-3 standard and the guidelines with the help of experts, the insomnia corpus was constructed by tagging them as insomnia tendency documents and non-insomnia tendency documents. Five deep learning language models (BERT, RoBERTa, ALBERT, ELECTRA, XLNet) were trained using the constructed insomnia corpus as training data. As a result of performance evaluation, RoBERTa showed the highest performance with an accuracy of 81.33%. In order to in-depth analysis of insomnia social data, topic modeling was performed using the newly emerged BERTopic method by supplementing the weaknesses of LDA, which is widely used in the past. As a result of the analysis, 8 subject groups ('Negative emotions', 'Advice and help and gratitude', 'Insomnia-related diseases', 'Sleeping pills', 'Exercise and eating habits', 'Physical characteristics', 'Activity characteristics', 'Environmental characteristics') could be confirmed. Users expressed negative emotions and sought help and advice from the Reddit insomnia community. In addition, they mentioned diseases related to insomnia, shared discourse on the use of sleeping pills, and expressed interest in exercise and eating habits. As insomnia-related characteristics, we found physical characteristics such as breathing, pregnancy, and heart, active characteristics such as zombies, hypnic jerk, and groggy, and environmental characteristics such as sunlight, blankets, temperature, and naps.

12

토픽 모델링 기반 비대면 강의평 분석 및 딥러닝 분류 모델 개발

한지영, 허고은

[Kisti 연계] 한국문헌정보학회 한국문헌정보학회지 Vol.55 No.4 2021 pp.267-291

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

2020년 신종 코로나바이러스 감염증(코로나19)으로 인한 전 세계적인 팬데믹으로 교육 현장에도 큰 변화가 있었다. 대학에서는 보조 교육 수단으로 생각했던 원격수업을 전면 도입하였고 비대면 수업이 일상화되어 교수자와 학생들은 새로운 교육환경에 적응하기 위해 큰 노력을 기울이고 있다. 이러한 변화 속에서 비대면 강의의 질적 향상을 위하여 강의 만족도 영향요인에 관한 연구가 필요하다. 본 연구는 코로나 전과 후로 변화된 대학 강의 만족도 영향요인을 파악하기 위해 빅데이터를 활용한 새로운 방법론을 제시하고자 한다. 토픽 모델링을 활용하여 코로나 전과 후의 강의평을 분석하고 이를 통해 강의 만족도 영향요인을 파악하여 대학교육이 나아가야 할 방향성을 제언하였다. 또한, 딥러닝 언어 모델인 KoBERT를 기반으로 0.84의 F1-score를 보이는 토픽 분류 모델을 구축함으로써 강의의 만족, 불만족 요인을 다각도로 파악할 수 있으며 이를 통해 강의 만족도의 지속적인 질적 향상에 기여할 수 있다.

Due to the global pandemic caused by COVID-19 in 2020, there have been major changes in the education sites. Universities have fully introduced remote learning, which was considered as an auxiliary education, and non-face-to-face classes have become commonplace, and professors and students are making great efforts to adapt to the new educational environment. In order to improve the quality of non-face-to-face lectures amid these changes, it is necessary to study the factors affecting lecture satisfaction. Therefore, This paper presents a new methodology using big data to identify the factors affecting university lecture satisfaction changed before and after COVID-19. We use Topic Modeling method to analyze lecture reviews before and after COVID-19, and identify factors affecting lecture satisfaction. Through this, we suggest the direction for university education to move forward. In addition, we can identify the factors of satisfaction and dissatisfaction of lectures from multiangle by establishing a topic classification model with an F1-score of 0.84 based on KoBERT, a deep learning language model, and further contribute to continuous qualitative improvement of lecture satisfaction.

13

도로 위의 군비경쟁 : LDA 토픽모델을 활용한 SUV의 인기 요인 탐구 KCI 등재

전승봉, 고태경

한국디지털정책학회 디지털융복합연구 제18권 제10호 2020.10 pp.239-252

※ 기관로그인 시 무료 이용이 가능합니다.

4,600원

본 연구자들은 텍스트 마이닝을 활용하여 SUV 선호 증가의 요인을 탐색하고자 한다. 온라인 자동차 커뮤니티인 보배드림에서 2005년부터 2019년까지 작성된 SUV 관련 게시글 32,679개를 수집한 후, LDA 토픽모델링 기법을 적용 하였다. 분석 결과, SUV 담화에서 주요한 토픽으로 등장한 ‘안전’이 범죄로부터 개인의 위험에 주목한 기존 연구와 달리 교통사고 및 고속주행 상황에서의 안전을 의미하는 것으로 드러났다. 한국 사회의 SUV 소비는 개인이 운전하면서 느끼 는 불안과 위험에 대한 대비 수단을 의미한다고 볼 수 있다는 것이다. 또한, 이와 같은 위험 인식 저변에는 불평등 증대 로 인해 감소하는 타인에 대한 신뢰가 작동한다고 할 수 있다.

By using text mining, we explore the factors responsible for an increase in SUV preference. We collected 32,679 posts related to SUVs from “Bobaedream,” the largest online automobile community in South Korea, and applied the LDA topic model. While previous studies have explained the SUV boom as an individual’s risk aversion strategy from crime, the result shows that the topic of ‘Safety’ appears to be an important factor in the SUV discourse in the context of a car accident and high-speed driving situation. To conclude, the consumption of SUVs in Korean society serves as a mean to prevent anxiety and danger to individuals when driving. We insist that decreasing social trust, caused by an increase in inequality, underlies the perception of risk on the road.

14

4,000원

택배 업계의 서비스 품질 향상과 고객 만족에 영향을 미치는 요인들을 다양한 차원에서 분석한 연구가 많이 이루어져 왔다. 이러한 연구의 대부분은 한정된 응답자를 대상으로 설문조사나 인터뷰 등의 질적 방법이 사용 되어 자료의 형식과 내용, 응답자의 범위가 제한적이라는 한계가 있다. 이에 본 연구에서는 장기간에 축적된 불특 정 다수의 소비자 상담 사례와 업계 종사자의 불만이 반영된 기사문을 대상으로 삼아 택배 서비스에서 소비자와 공급자가 지적하는 불만에 관한 주요 토픽을 탐색하고 분석하여 선행 연구의 미비점을 보완하고자 하였다. 또한, 이러한 토픽이 시점에 따라 어떻게 변화하는지 흐름을 분석함으로써 최근에 제기된 새로운 토픽을 발굴하고 시사 점을 제시하고자 하였다. 이 결과 지연/분실/오배송 토픽과 택배 산업 경쟁 심화 토픽이 중심을 이루는 것으로 나타났다. 토픽 트렌드 분석 결과 최근 국제 택배 상담 내용이 다소 늘었고, 아파트 택배 배송과 관련된 갈등이 많이 다루어져 이를 정부 정책에 반영하거나 연구 주제로 다루어 볼 수 있을 것이다. 연구 결과 나타난 토픽은 선행 연구에서 다루어진 내용이 주를 이루지만 내밀한 상담사례와 학술 문헌 등 다른 자료와 분석 방법을 추가하면 더 새롭고 가치있는 토픽을 도출할 수 있을 것이라고 기대한다.

Many studies have been conducted to analyze factors that affect customer satisfaction, and service quality improvement in the parcel delivery industry. Most of these studies have a limited number of respondents using methods such as surveys and interviews. Therefore, this study aims to supplement the shortcomings of previous studies, by searching and analyzing the common major topics related to the complaints pointed out by consumers and suppliers in the parcel delivery service with cases of consumer counseling, and articles that reflect the complaints of workers in the industry. In addition, by analyzing the trend of these topics, we attempted to discover new topics and suggest implications. In conclusion, topics such as delay/lost/wrong deliveries as well as the fierce competition in the parcel delivery industry, turned out to be central aspects. As a result of the topic trend analysis, talks with international couriers have recently increased, and many conflicts related to apartment parcel delivery have been dealt with. The topics presented in this study are mainly focused on the contents of previous studies, but we expect that new and valuable topics can be derived by adding other data and analysis methods, such as internal counseling and academic literature.

15

LDA 토픽 모델을 활용한 포스트 Covid-19 시대의 소상공인 지원정책 분석 KCI 등재

서경도, 최정일, 최판암, 정재림

대한산업경영학회 산업융합연구(구 대한산업경영학회지) 제22권 제6호 2024.06 pp.51-59

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문은 COVID-19와 같은 팬데믹 상황에서 소상공인에게 실질적으로 도움이 되는 정부 정책을 제언하는데 목적 이 있다. 이를 위해 ‘COVID-19 소상공인 지원’, ‘COVID-19 감염병 대응체계에 따른 소상공인 영향’, ‘COVID-19 소상공 인 경제정책’ 키워드를 중심으로 뉴스 기사를 크롤링하여 텍스트 마이닝 분석의 키워드 빈도분석과 워드클라우드 분석을 수 행하였고, LDA 토픽 모델링 분석을 통해 주요 이슈를 파악하였다. LDA 토픽 모델링을 수행한 결과 소상공인 지원 정책은 정 부의 현금성 지원과 금융지원으로 토픽 레이블을 구성하였고, COVID-19 감염병 대응체계에 따른 소상공인 영향은 정부 주 도의 방역체계와 개인 주도의 방역체계로 토픽 레이블을 구성하였으며, COVID-19 경제정책은 경제위기와 자생력을 갖추기 위한 소상공인 정책으로 토픽 레이블을 구성하였다. 구성한 토픽레이블을 중심으로 향후 팬데믹 상황에서 소상공인 피해 감면정책과 소상공인이 시장경쟁력 제고 정책에 대해 파악할 수 있는 기초자료를 제공하고자 하였다.

The purpose of the paper is to suggest government policies that are practically helpful to small business owners in pandemic situations such as COVID-19. To this end, keyword frequency analysis and word cloud analysis of text mining analysis were performed by crawling news articles centered on the keywords "COVID-19 Support for Small Businesses", "The Impact of Small Businesses by Response System to COVID-19 Infectious Diseases", and "COVID-19 Small Business Economic Policy", and major issues were identified through LDA topic modeling analysis. As a result of conducting LDA topic modeling, the support policy for small business owners formed a topic label with government cash and financial support, and the impact of small business owners according to the COVID-19 infectious disease response system formed a topic label with a government-led quarantine system and an individual-led quarantine system, and the COVID-19 economic policy formed a topic label with a policy for small business owners to acquire economic crisis and self-sustainability. Focusing on the organized topic label, it was intended to provide basic data for small business owners to understand the damage reduction policy for small business owners and the policy for enhancing market competitiveness in the future pandemic situation.

16

다이나믹 토픽 모델을 활용한 D(Data)ㆍN(Network)ㆍA(A.I) 중심의 연구동향 분석 KCI 등재

우창우, 이종연

한국융합학회 한국융합학회논문지 제11권 제9호 2020.09 pp.21-29

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

최근 디지털 사회의 도래로 다양한 데이터가 폭발적으로 증가하고, 그중 문헌 내 주제어를 도출하는 토픽 모델링 에 관한 연구가 활발히 진행되고 있다. 본 논문의 연구목표는 토픽 모델링 방법 중 하나인 DTM(Dynamic Topic Model) 모델을 적용해 D.N.A.(Data, Network, A.I) 분야에 대한 연구동향을 탐색하는데 있다. 실험 데이터는 최근 6년간(2015∼2020) ICT(Information and Communication Technology) 분야 중 기술대분류가 SWㆍAI에 해당하 는 연구과제 1,519개 사업에 대해 DTM 모델을 적용하였다. 실험결과로, D.N.A. 분야의 기술 키워드 Big data, Cloud, Artificial Intelligence와 확장된 의미의 기술 키워드 Unstructured, Edge Computing, Learning, Recognition 등 이 매년 연구에 표출되었으며, 해당 키워드 들이 특정 연구과제에 종속되지 않고 다른 연구과제에서도 포괄적으로 연구 되고 있음을 확인하였다. 끝으로 본 논문의 연구결과는 향후 D.N.A. 분야에 대한 정책기획ㆍ과제기획 등 연구개발 기획 과정과 기업의 기술 확보전략ㆍ마케팅 전략 등 다양한 곳에 활용될 수 있을 것으로 기대한다.

The Topic Modeling research, the methodology for deduction keyword within literature, has become active with the explosion of data from digital society transition. The research objective is to investigate research trends in D.N.A.(Data, Network, Artificial Intelligence) field using DTM(Dynamic Topic Model). DTM model was applied to the 1,519 of research projects with SWㆍA.I technology classifications among ICT(Information and Communication Technology) field projects between 6 years(2015∼2020). As a result, technology keyword for D.N.A. field; Big data, Cloud, Artificial Intelligence, extended keyword; Unstructured, Edge Computing, Learning, Recognition was appeared every year, and accordingly that the above technology is being researched inclusively from other projects can be inferred. Finally, it is expected that the result from this paper become useful for future policyㆍR&D planning and corporation’s technologyㆍmarketing strategy.

17

비정형 보안 인텔리전스 보고서 기반 토픽 자동 추출 모델 KCI 등재

허윤아, 이찬희, 김경민, 임희석

한국융합학회 한국융합학회논문지 제10권 제6호 2019.06 pp.33-39

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

지능형 사이버 공격 기법이 다양화됨에 따라 보안 침해 사건, 글로벌 범죄 등의 사건 발생이 증가하고 있다. 지능형 공격을 예측하고 대응하기 위해서는 공격 기법의 특성, 수법, 유형을 파악해야 한다. 이를 위해 수많은 보안 기업 회사에서는 다양한 공격 기법을 빠르게 파악하고 더 큰 피해를 막기 위해 보안 인텔리전스 보고서를 배포한다. 하지만 각 기업에서 배포하는 보고서에 대한 형식이 맞춰져 있지 않으며, 대량의 비정형 보안 인텔리전스 보고서가 배포되고 있다. 본 논문은 비정형한 보안 인텔리전스 보고서에 대한 문제점을 고려하여 정형화된 데이터로 추출하는 방안을 제안 한다. 또한, 대량의 보안 인텔리전스 보고서를 파악하기 위해 소요되는 시간을 줄이고자 대량의 보고서를 주제별로 분류 할 수 있는 보안 인텔리전스 보고서 토픽 자동 추출 모델을 제안한다.

As cyber attack methods are becoming more intelligent, incidents such as security breaches and international crimes are increasing. In order to predict and respond to these cyber attacks, the characteristics, methods, and types of attack techniques should be identified. To this end, many security companies are publishing security intelligence reports to quickly identify various attack patterns and prevent further damage. However, the reports that each company distributes are not structured, yet, the number of published intelligence reports are ever-increasing. In this paper, we propose a method to extract structured data from unstructured security intelligence reports. We also propose an automatic intelligence report analysis system that divides a large volume of reports into sub-groups based on their topics, making the report analysis process more effective and efficient.

18

4,000원

본 논문의 연구목표는 LDA(Latent Dirichlet Allocation) 모델을 적용하여 국가연구개발사업을 통해 수행되고 있는 ICT(Information and Communication Technology) 분야의 연구과제에 대한 주요 연구 토픽과 동향을 탐색 하는데 있다. 연구방법에는 NTIS(National Science and Technology Information Service)로부터 최근 5년간 국가 연구개발사업의 전체 연구과제 정보를 다운로드받고 이를 정보통신기획평가원(IITP)의 EZone 시스템과 매칭하여 ICT 분야 연구과제 5,200건을 확보하고, 토픽모델링 기법중 하나인 LDA 모델을 적용하여 연구토픽과 연구동향을 조사하였 다. 실험결과로, ICT분야 연구과제에 대한 연구토픽은 인공지능, 빅데이터, 사물인터넷(Internet of Things)과 같은 지능정보기술로 확인되었고 연구동향에는 초실감미디어에 관한 연구가 활발히 진행되고 있음을 확인하였다. 끝으로 본 논문에서 진행된 국가연구개발사업에 대한 토픽모델링 결과는 향후 ICT분야 연구개발 계획 및 전략수립, 정책, 과제기 획 등 중요한 정보로 활용될 수 있을 것이다.

The research objectives investigates main research topics and trends in the information and communication technology(ICT) field, Korea using LDA(Latent Dirichlet Allocation), one of the topic modeling techniques. The experimental dataset of ICT research and development(R&D) project of 5,200 was acquired through matching with the EZone system of IITP after downloading R&D project dataset from NTIS(National Science and Technology Information Service) during recent five years. Consequently, our finding was that the majority research topics were found as intelligent information technologies such as AI, big data, and IoT, and the main research trends was hyper realistic media. Finally, it is expected that the research results of topic modeling on the national R&D foundation dataset become the powerful information about establishment of planning and strategy of future’s research and development in the ICT field.

19

Employing Latent Dirichlet Allocation Model for Topic Extraction of Chinese Text SCOPUS

Qihua Liu

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.7 2016.07 pp.51-66

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The hidden topic model of Chinese text, which possesses complicated semantics, is urgently needed, since China has occupied an increasingly significant role during the booming development of globalization over recent years. This paper details and elaborates the basic process of extracting latent Chinese topics by demonstrating a Chinese topic extraction schema based on Latent Dirichlet Allocation (LDA) model. Furthermore, the application was practiced in CCL, an authoritative Chinese corpus, to extract topics for its nine categories. With rigorous empirical analysis, extracting the LDA results has a considerably higher average precision rate as opposed to other three comparable Chinese topic extraction techniques; however the average recall rate is worse than KNN and almost the same with the PLSI model. Moreover, the recall rate and precision rate of LDA-CH is worse than LDA-EH. Therefore, the LDA model should be improved to adapt to the distinctive feature of Chinese words with the purpose of making it better for Chinese topic extraction.

20

Service Clustering by Leveraging a Context-Sensitive Approach SCOPUS

Lantian Guo, Tao Yang, Huixiang Zhang, Dejun Mu, Zhe Li, Yang Li

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.11 No.12 2016.12 pp.433-442

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Service technology has gained increasing popularity in recent communication software applied in many domains. With a growing number of services that share same or similar functionalities, clustering services help improve both service composition and mashup creation. To achieve service clustering, utilizing probabilistic topic model to extract and characterize the service description documents as corresponding topics is an available scheme. However, unlike short text in social networks, the descriptions of published services possess higher dimensionality and sparse functional information. With traditional LDA (Latent Dirichlet Allocation) model to implement topic extraction makes topics unclear. To address that challenge, we conduct a context sensitive approach to generate context sensitive vector for merging the words with similar context before loading to LDA model, referred to as CV-LDA (Context Vector LDA). Through F1-Measure of clustering and topic perplexity analysis in the real-world dataset, it is shown that the proposed approach outperforms traditional LDA model in service clustering.

 
1 2 3 4 5
페이지 저장