년 - 년
파이썬 기반 규칙 중심 텍스트 분석을 활용한 한·중 한국어교육 학술논문의 구조·주제·정서 경향 탐색 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제10권 3호 2026.03 pp.825-833
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 파이썬 기반 규칙 중심 텍스트 분석을 활용하여 한·중 한국어교육 학술논문의 구조 요소, 주제 분포 및 정서 지표의 분포 경향을 탐색적으로 분석하였다. 분석 대상은 CNKI와 KCI에 등재된 한국어교육 관련 학술논문 각 150편으로, PDF 문헌을 텍스트 데이터로 변환한 후 동일한 전처리 및 분석 절차를 적용하였다. 구조 분석에서는 자동 판별 모델을 사용하지 않고, 핵심 어휘의 출현 비율과 위치 분포를 규칙 기반으로 계량하였다. 주제 분석은 군집 안정도가 높은 문헌을 중심으로 키워드 분포와 주제 집중 경향을 살펴보았다. 정서 분석은 사전 및 규칙 기반 방식을 활용하여 논문별 정서 강도 지표를 산출하고 범주별 분포를 제시하였다. 분석 결과, 한·중 논문은 전반적인 구조 구성 비율과 위치 분포에서 유사한 경향을 보였으나, 일부 구조 범주에서 어휘 밀도와 주제 집중 양상에 차이가 관찰되었다. 또한 정서 지표는 양국 논문 모두 중립 범주에 집중되는 경향을 나타냈다. 본 연구는 대규모 학술 텍스트를 대상으로 한 규칙 기반 자동 분석의 적용 가능성을 제시하고, 한국어교육 연구에서 활용 가능한 방법론적 시사점을 제공한다.
This study explores structural elements, topic distribution, and sentiment indicators in Korean language education research articles published in China and Korea using Python-based rule-oriented text analysis. A total of 300 articles indexed in CNKI and KCI were analyzed after converting PDF documents into text data and applying a unified preprocessing procedure. Structural analysis was conducted through rule-based quantification of keyword frequency and positional distribution rather than automated classification models. Topic distribution was examined by focusing on documents with high cluster coherence to identify keyword concentration and thematic tendencies. Sentiment analysis employed lexicon- and rule-based methods to calculate sentiment intensity indicators for each article. The results show that articles from both countries share similar overall structural proportions and positional patterns, while differences were observed in lexical density and topic concentration within certain structural categories. Sentiment indicators in both datasets were largely concentrated in the neutral range. This study demonstrates the applicability of rule-based automated text analysis for large-scale academic corpora and offers methodological implications for future research in Korean language education.
Too Much Information – Trying to Help or Deceive? An Analysis of Yelp Reviews KCI 등재 SCOPUS
한국경영정보학회 Asia Pacific Journal of Information Systems 제33권 제2호 2023.06 pp.261-281
※ 기관로그인 시 무료 이용이 가능합니다.
5,700원
The proliferation of online customer reviews has completely changed how consumers purchase. Consumers now heavily depend on authentic experiences shared by previous customers. However, deceptive reviews that aim to manipulate customer decision-making to promote or defame a product or service pose a risk to businesses and buyers. The studies investigating consumer perception of deceptive reviews found that one of the important cues is based on review content. This study aims to investigate the impact of the information amount of review on the review truthfulness. This study adopted the Information Manipulation Theory (IMT) as an overarching theory, which asserts that the violations of one or more of the Gricean maxim are deceptive behaviors. It is regarded as a quantity violation if the required information amount is not delivered or more information is delivered; that is an attempt at deception. A topic modeling algorithm is implemented to reveal the distribution of each topic embedded in a text. This study measures information amount as topic diversity based on the results of topic modeling, and topic diversity shows how heterogeneous a text review is. Two datasets of restaurant reviews on Yelp.com, which have Filtered (deceptive) and Unfiltered (genuine) reviews, were used to test the hypotheses. Reviews that contain more diverse topics tend to be truthful. However, excessive topic diversity produces an inverted U-shaped relationship with truthfulness. Moreover, we find an interaction effect between topic diversity and reviews’ ratings. This result suggests that the impact of topic diversity is strengthened when deceptive reviews have lower ratings. This study contributes to the existing literature on IMT by building the connection between topic diversity in a review and its truthfulness. In addition, the empirical results show that topic diversity is a reliable measure for gauging information amount of reviews.
Impact of Topic Distribution on Review Sentiment: A Comparative Study between South Korea and the U.S. KCI 등재 SCOPUS
한국경영정보학회 Asia Pacific Journal of Information Systems 제32권 제3호 2022.09 pp.514-536
※ 기관로그인 시 무료 이용이 가능합니다.
6,000원
Online reviews offer valuable information to businesses by reflecting consumer experiences about their products and services. Two crucial aspects of online reviews are the topics consumers choose to address, and the sentiments expressed in their reviews. Building upon previous literature that shows online reviews are context- dependent, we employ the Expectation-Confirmation Theory (ECT) to examine the impact of topic distribution on review sentiment in South Korea and the U.S. during pre- and post-pandemic periods. After applying a topic modeling to Airbnb app review data, we measure the contribution of each topic on review sentiment using SHAP values. Our results indicate variations in topic distribution trends between 2018 and 2021. In addition, the order and magnitude of topics’ impact on review sentiment change between pre- and post-pandemic periods for both countries. This study can help businesses understand how topics and sentiments associated with their products and services changed after the pandemic and thus identify areas of improvement.
빅데이터 토픽모델링과 감성분석을 활용한 물공급과정에서의 수질사고 기사 분석
[Kisti 연계] 한국수자원학회 한국수자원학회 논문집 Vol.55 No.suppl1 2022 pp.1235-1249
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구에서는 웹 크롤링 방법을 이용한 자료수집, 텍스트 마이닝을 활용한 데이터 분석과 같은 빅데이터 분석기법을 이용하여 국내 상수도 수질사고에 대한 전개양상 분석을 수행하였다. 상수도 시스템의 수질사고 빅데이터 뉴스의 추출을 위한 웹크롤링 기법을 적용하고 정확한 수질사고 뉴스를 획득하고자 알고리즘을 절차화하여 제시하였다. 또한 대규모 수질사고의 경우 사고발생에 따른 사고인지, 사고확산, 사고대응, 사고해결 등과 같은 전개양상이 나타나므로, 각 단계에 따른 적절한 뉴스기사를 추출하고, 이에 따른 정보분석을 실시하였다. 즉, 각 단계 별 주요 키워드, 감성분석을 통한 수질사고 전개양상분석을 사례기반으로 상세히 실시하고 그 의미를 분석, 도출하였다. 제안된 방법론을 2020년 발생한 인천광역시 유충사고기간에 적용하여 분석하였다. 그 결과, 수질사고와 같은 소비자에게 직접적인 영향을 미치는 정보의 공개가 제한된 상황에서 사고발생시 장기간의 피해 지속성이 있는 수질사고에 대한 뉴스 기사 언론보도의 논조 및 소비자의 긍부정도가 시간에 따라 명확히 변화됨을 확인할 수 있었다. 이것은 공급자 입장에서의 수질사고의 전개양상은 시설물의 빠른 복구도 매우 중요하지만 소비자의 긍정도를 높이기 위한 소비자 중심의 정책마련의 필요성을 제시하고 있다.
This study applied the web crawling technique for extracting big data news on water quality accidents in the water supply system and presented the algorithm in a procedural way to obtain accurate water quality accident news. In addition, in the case of a large-scale water quality accident, development patterns such as accident recognition, accident spread, accident response, and accident resolution appear according to the occurrence of an accident. That is, the analysis of the development of water quality accidents through key keywords and sentiment analysis for each stage was carried out in detail based on case studies, and the meanings were analyzed and derived. The proposed methodology was applied to the larval accident period of Incheon Metropolitan City in 2020 and analyzed. As a result, in a situation where the disclosure of information that directly affects consumers, such as water quality accidents, is restricted, the tone of news articles and media reports about water quality accidents with long-term damage in the event of an accident and the degree of consumer pride clearly change over time. could check This suggests the need to prepare consumer-centered policies to increase consumer positivity, although rapid restoration of facilities is very important for the development of water quality accidents from the supplier's point of view.
Detecting Poetic Metaphors by LDA-based Topic Distribution
[NRF 연계] 중앙대학교 인문콘텐츠연구소 인공지능인문학연구 Vol.5 2020.04 pp.77-93
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
It is difficult to automatically extract a metaphor from Chinese poetry. In Chinese poetry, a metaphor appears when a word has a different, implicit connotation from its original, explicit significance. The meaning of a word in a non-literary text is its original, explicit sense. Thereby, we assume the metaphorical word, which has different nuances in a poem and non-literary texts (which form a semantically inconsistent pair). Depending on the text, a word is semantically inconsistent. For example, a “moon” is a satellite of the Earth in a non-literary setting, while in the poem “Quiet Night Thoughts,” the term “moon” means homesickness. Hence, the “moon” is an SIP in “Quiet Night Thoughts” and non-literary texts. This paper aims to detect SIPs in Chinese poems and non-literary texts. In particular, we discern SIP based on latent Dirichlet allocation (LDA) topic modeling. Subsequently, the proposed method has been evaluated by discovering SIP in Chinese poetry and non-literary texts.
[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.23 No.4 2020 pp.595-602
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
We propose a generative probabilistic model with Dirichlet prior distribution for topic modeling and text similarity analysis. It assigns a topic and calculates text correlation between documents within a corpus. It also provides posterior probabilities that are assigned to each topic of a document based on the prior distribution in the corpus. We then present a Gibbs sampling algorithm for inference about the posterior distribution and compute text correlation among 50 abstracts from the papers published by IEEE. We also conduct a supervised learning to set a benchmark that justifies the performance of the LDA (Latent Dirichlet Allocation). The experiments show that the accuracy for topic assignment to a certain document is 76% for LDA. The results for supervised learning show the accuracy of 61%, the precision of 93% and the f1-score of 96%. A discussion for experimental results indicates a thorough justification based on probabilities, distributions, evaluation metrics and correlation coefficients with respect to topic assignment.
[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.23 No.7 2020 pp.883-890
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this paper, we propose a variational expectation-maximization algorithm that computes posterior probabilities from Latent Dirichlet Allocation (LDA) model. The algorithm approximates the intractable posterior distribution of a document term matrix generated from a corpus made up by 50 papers. It approximates the posterior by searching the local optima using lower bound of the true posterior distribution. Moreover, it maximizes the lower bound of the log-likelihood of the true posterior by minimizing the relative entropy of the prior and the posterior distribution known as KL-Divergence. The experimental results indicate that documents clustered to image classification and segmentation are correlated at 0.79 while those clustered to object detection and image segmentation are highly correlated at 0.96. The proposed variational inference algorithm performs efficiently and faster than Gibbs sampling at a computational time of 0.029s.
영어 교과서, EBS 교재, 대학수학능력시험의 읽기 지문에 대한 코퍼스 기반 소재별 어휘 사용 양상 분석
[NRF 연계] 학습자중심교과교육학회 학습자중심교과교육연구 Vol.19 No.4 2019.02 pp.711-729
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 고등학교 영어 교과서, EBS 교재, 대학수학능력시험 읽기 지문의 소재와 소재별 어휘 사용 양상을 비교·분석하고 소재와 소재별 어휘 차이가 학습자에게 학습 부담을 주는지 살펴보았다. 이를 위하여 2009 개정 교육과정에 따른 영어 교과서, EBS 교재, 수능의 읽기 지문에 대해 2017년 1월부터 8월까지 코퍼스로 구축한 후, 영어과 교육과정 권장 소재 목록과 비교하고, 소재를 분류하여 읽기 지문별 소재 분포를 확인하고, 교육과정, BNC-COCA-25000, GSL-AWL 어휘 목록과 비교하여 소재별 어휘 사용 양상을 분석하였다. 그 연구 결과는 다음과 같다. 첫째, 교과서는 고른 소재 분포를 보였으며, EBS 교재와 수능은 1번, 18번, 19번 소재에 편중된 분포를 보였다. 둘째, 교과서와 달리 EBS 교재와 수능에서는 가장 많이 등장하는 소재의 교육과정 권장 어휘 반영 비율이 가장 낮았다. 셋째, EBS 교재와 수능에 등장한 1번, 18번, 19번 소재의 경우 BNC-COCA-25000 어휘 밴드 중 가장 높은 수준의 어휘를 요구하였다. 넷째, 교과서와 달리 EBS 교재와 수능은 학술적인 성격의 소재가 상당수 있었다. 교과서가 교육과정을 충실히 반영하여 소재와 어휘 면에서 통제가 이루어지고 있는 반면 EBS 교재와 수능은 편중된 소재 분포를 보였으며, 자주 등장하는 소재의 어휘 학습 부담이 커서 교과서 읽기 지문의 소재별 어휘 수준의 현실적 조정이 필요하다는 교육적 시사점을 얻었다.
The purpose of this study is to analyze the topic distribution and vocabulary level in high school English textbooks, EBS materials, and CCASTs, and to examine the differences in topic distribution and vocabulary among the different materials. A corpus was constructed with sixteen English I and II textbooks, three EBS materials published in 2016, and the CSATs from 2013 to 2016 respectively. Then these corpora were analyzed to find out how topics are distributed based on the 19 topic categories in the 2009 National Curriculum. In addition, they were analyzed with the corpus-based word lists: the 2009 National Curriculum Word List, BNC-COCA-25000, and GSL-AWL to examine the vocabulary level of the reading materials. The results of the study are as follows: First, the textbooks consist of various types of topics based on the curriculum, but in the EBS materials and the CSATs, three out of the 19 topics were highly dominant, which were about personal life, general education and academic education. Second, the textbooks contain more words from the 2009 National Curriculum Word List than the EBS materials and CSATs do. Third, the result of the comparison with the BNC-COCA-25000 revealed that the vocabulary levels of the EBS materials and CASTs are much higher than those of the textbooks. Fourth, the result of comparison with the GSL-AWL showed that the percentage of the AWL in the reading passages of the textbooks did not exceed 10%. However, the percentage of AWL in the EBS materials and CSATs was relatively higher than the textbooks. The suggestions based on the above results are as follows: First, it is essential to include the word list of the CASTs in developing textbooks. Second, guidelines for topic distribution must be provided considering the learners’ grades and purposes of materials. Finally, it is needed to analyze topic distribution based on question types and various factors within topic distribution.
토픽모델링을 이용한 지역 간 의료인력 불균형 원인분석: 수도권-비수도권 의료인력 종사자 직무 인식 비교분석
[NRF 연계] 한국보건사회연구원 보건사회연구 Vol.45 No.4 2025.12 pp.467-491
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
지역 간 의료인력 불균형은 단순한 인력 분포 문제를 넘어, 의료 접근성 저하와 건강 형평성 악화로 이어질 수 있는 구조적 문제로 대두되고 있다. 특히 수도권과 비수도권 간 병원 근무 환경 및 직무 만족도에 대한 체계적인 비교 연구는 아직 미흡한 실정이다. 본 연구는 수도권과 비수도권 간 의료인력 분포 불균형의 원인을 의료 종사자의 근무 인식을 통해 규명하고자 하였다. 이를 위해 기업 리뷰 플랫폼 ‘잡플래닛’에 게시된 수도권 5개 및 비수도권 18개 상급종합병원 종사자의 리뷰 4,537건을 수집하였으며, ‘장점’, ‘단점’, ‘경영진에게 바라는 점’ 항목을 중심으로 Python 기반의 LDA 토픽모델링을 실시하였다. 분석 결과, 수도권 병원 종사자는 높은 보상 체계, 다양한 임상 경험, 체계적인 수련 환경 등을 장점으로 언급한 반면, 비수도권 병원 종사자는 과중한 업무, 수직적인 조직문화, 낮은 보상 수준 등을 주요 기피 요인으로 지적하였다. 이는 의료인의 근무지 선택이 단순한 개인 선호를 넘어 병원 구조와 제도, 근무환경의 질적 차이에 기반함을 시사한다. 본 연구는 의료인의 실제 경험이 반영된 정성적 데이터를 분석함으로써, 수련 체계 개편, 지역 근무 유인을 위한 인센티브 확대 등 정책적 개선 방향을 제안하고, 의료인력의 지역 간 분포 불균형 해소를 위한 실증적 함의를 제공한다.
Regional disparities in the medical workforce constitute a structural issue that can undermine healthcare accessibility and equity. However, comparative evidence on work environments between metropolitan and non-metropolitan hospitals remains limited. This study examines the causes of workforce imbalance by analyzing employee perceptions. A total of 4,537 reviews from JobPlanet, covering five metropolitan and 18 non-metropolitan tertiary hospitals, were collected and analyzed using Python-based Latent Dirichlet Allocation (LDA) on the categories of strengths, weaknesses, and suggestions. The results showed that metropolitan hospitals were associated with favorable perceptions, particularly regarding compensation, clinical exposure, and training systems. In contrast, non-metropolitan hospitals were linked to negative perceptions, including heavy workloads, hierarchical culture, and lower pay. These findings suggest that practice location decisions among healthcare personnel are shaped by institutional and environmental factors rather than individual preference alone. By utilizing experience-based qualitative data, this study provides empirical insights into structural drivers of workforce maldistribution and highlights the need for training reforms and targeted regional retention incentives.
데이터 분산 서비스를 활용한 실시간 시험자료 토픽 설계
[Kisti 연계] 한국정보통신학회 한국정보통신학회논문지 Vol.21 No.7 2017 pp.1447-1454
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
실시간 시험자료 토픽은 시험을 수행하는 네트워크에 연결되어 있는 여러 계측 장비로부터 실시간으로 데이터를 수신하여 분석 처리하고 계측장비로 데이터를 제공하거나 가시화 할 수 있는 일종의 패킷 형태를 의미한다. 기존 UDP 통신프로토콜을 활용한 구조에서는 모든 계측장비들이 전송하는 데이터를 하나의 패킷으로 통합 설계하여 계측장비들의 필요유무와 상관없이 송 수신 하는 한계점이 존재하였다. 하지만 DDS(Data Distribution Service)를 활용하여 제안하는 시스템의 토픽 설계는 다음과 같은 장점들이 있다. 각 시스템에서 사용하는 플랫폼에 유연하게 공통된 API를 사용하여 개발이 가능하며 향후 장비 업그레이드 시 필요 토픽의 추가 선언 등 최소 작업만 필요하고 전체 시스템을 재설계하지 않아도 된다. 또한 시스템 간 연계를 위한 계측장비 및 시스템이 추가로 도입 시에도 공통 메시지 포맷을 적용하여 개발하기 때문에 기존 장비의 수정이 불필요하여 시스템의 확장이 용이하다. 추가 장비의 도입은 토픽의 QoS(Quality of Service) 튜닝을 통하여 우선 적용할 수 있기 때문에 통신의 성능을 조정 및 유지할 수 있다. 본 논문에서는 이종 시스템간의 플랫폼과 통신 프로토콜을 통합 설계한 DDS 미들웨어를 활용하여 새로운 센서 및 계측장비 도입 시 기존 시스템 구성장비들의 수정과 시스템의 별도 통신 커넥션, 신규 장비의 도입 및 업그레이드에 따른 시스템 S/W 재설계를 지양하는 토픽의 설계를 통해 보다 효율적인 자료 전달체계를 제안하고자 한다.
The realtime test data topic means that process for the data efficiently from many kinds of measurement device at the test range. There are many measurement devices in test range. The test range require accurate observation and determine on test object. In this realtime test data slaving framework system, the system can produce variety of test informations and all these data also must be transmitted to test information management or display system in realtime. Using RTI DDS(Data Distribution Service) middle ware Ver 5.2, we can product the efficiency of system usability and QoS(Quality of Service) requirements. So the application user enables to concentrate on applications, not middle ware. As the reason, Complex function is provided by the DDS, not the application such as Visualization Software. In this paper, I suggest the realtime test data topic on slaving framework of realtime test data based on DDS at the test range system.
텍스트마이닝을 활용한 식품유통 플랫폼에 대한소비자 인식 분석 - 토픽모델링 기법을 중심으로 -
[NRF 연계] 한국외식경영학회 외식경영연구 Vol.24 2021.11 pp.71-100
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
COVID-19로 인해 비대면 문화가 확산되면서 소비자들의 식품유통 플랫폼 이용률이 증가하고 있다. 본 연구는 식품유통 플랫폼의 최근 2년간 소비자 리뷰를 토픽모델링 기법으로 분석하였다. 이를 통해 식품유통 플랫폼의 대표 3사에 대한 소비자 리뷰의 차이와 코로나 발생으로 식품유통 플랫폼에 대한 소비자 인지하는 키워드에 차이가 존재하는지 확인하였다. 이를 위해 2019년, 2020년 총 2개 연도의 ‘마켓컬리’, ‘쿠팡’, ‘SSG’ 어플리케이션 리뷰 총 17,954건의 데이터를 App Store와 Google Store에서 크롤링하여 수집하였다. 연구결과, 2019년에 비해 2020년의 리뷰의 개수 및 평점 평균이 유의미하게 증가하였으며 각 플랫폼별 증가·감소한 단어와 소비자들이 인지하고 있는 키워드의 차이를 확인하였다. 코로나 발생 이후 온라인 식품유통 플랫폼의 전반적인 이용률과 만족도는 크게 상승하였고, 코로나로 인한 환경 변화로 인해 플랫폼별 서비스 차이를 소비자들이 인지하고 있음을 확인하였다. 본 연구는 식품유통 플랫폼 이해관계자들에게 소비자의 니즈 및 불만 사항에 대한 유용한 정보를 제공하였다. 사회적 이슈에 대한 소비자들의 관심과 행동의 영향력을 파악하여 소비자들의 니즈와 사회 참여 욕구의 발견으로 플랫폼 마케팅을 위한 기초자료로 활용될 수 있다.
As the non-face-to-face culture spreads due to COVID-19, consumers' use of food distribution platforms is increasing. Based on consumer reviews, this study analyzed the difference in consumer perception of the three representative food distribution platforms (Market Kurly, Coupang, and SSG) and the changes in consumer perception of food distribution platforms before and after COVID-19 using topic modeling. The results confirm that the overall utilization and satisfaction of the food distribution platform increased after the outbreak of COVID-19 in 2020, and changes in word frequencies indicate consumers’ awareness of the difference in services by platform and the impact of COVID-19. This study provides useful information on consumer needs for food distribution platform stakeholders, and can be used as basic data for platform marketing since it informs of consumers' social participation needs by capturing the influence of consumers' interest in social issues and consequent behavior on the use of food distribution platforms.
토픽 모델링을 이용한 트위터 데이터의 공간 분포 패턴 분석
[Kisti 연계] 한국지역지리학회 한국지역지리학회지 Vol.23 No.2 2017 pp.376-387
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 트위터를 대상으로 트윗 공간 데이터에서 지리적 의미를 탐색하기 위한 방법을 모색하였다. 트윗 공간 데이터의 구축 과정 및 지리적 분석의 프레임워크를 정립하고 지리적 연구 방법론을 제안하였다. 이를 위해 본 연구는 제주도의 GPS 좌표 참조 트윗(geotweet)을 대상으로 트윗의 내용적 특성과 트윗 발생 위치의 공간 분포 특성을 확인하였다. 제주도 좌표 참조 트윗에서는 지명 또는 장소명이 많이 출현하였는데, 이는 자신의 위치를 알리고자하는 의도로 파악하였다. 트윗의 공간 분포는 제주공항을 중심으로 한 일부 관광지 주변으로 핫스팟이 확인되었고, 이는 제주도 유동인구 핫스팟과 유사한 패턴을 보였다. 주제 중심의 트윗 분석을 위해 본 연구에서는 토픽 모델링 알고리즘을 이용하여 분석하였다. 분석 결과, 주제의 지리적 위치와 트윗의 내용은 서로 관련이 있음을 알 수 있었다. 마지막으로 본 연구는 토픽 모델링 분석을 통해 방대한 트윗 데이터의 내용에 상응하는 지역 분포 특성을 직관적으로 확인하는데 유용하게 활용될 수 있다는 것을 확인하였다.
This paper attempts to analyze the geographical characters of Twitter data and presents analysis potentials for social network analysis in geography. First, this paper suggests a methodology for a topic modeling-based approach in order to identify the geographical characteristics of tweets, including an analysis flow of Twitter data sets, tweet data collection and conversion, textural pre-processing and structural analysis, topic discovery, and interpretation of tweets' topics. GPS coordinates referencing tweets(geotweets) were extracted among sampled Twitter data sets because it contains the tweet place where it was created. This paper identifies a correlated relationship between some specific topics and local places in Jeju. This correlation is closely associated with some place names and local sites in Jeju Island. We assume it is the intention of tweeters to record their tweet places and to share and retweet with other tweeters in some cases. A surface density map shows the hotspots of tweets, detecting around some specific places and sites such as Jeju airport, sightseeing sites, and local places in Jeju Island. The hotspots show similar patterns of the floating population of Jeju, especially the thirty-year age group. In addition, a topic modeling algorithm is applied for the geographical topic discovery and comparison of the spatial patterns of tweets. Finally, this empirical analysis presents that Twitter data, as social network data, provide geographical significance, with topic modeling approach being useful in analyzing the textural features reflecting the geographical characteristics in large data sets of tweets.
어휘의 동시 발생 빈도와 분포를 이용한 다중 주제 회의록 요약
[Kisti 연계] 한국컴퓨터정보학회 한국컴퓨터정보학회 학술대회논문집 2015 pp.13-16
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 어휘의 동시 발생 (co-occurrence) 빈도와 분포를 이용한 회의록 요약방법을 제안한다. 회의록은 일반 문서와 달리 문서에 여러 세부적인 주제들이 나타나며, 잘못된 형식의 문장, 불필요한 잡담들을 포함하고 있기 때문에 이러한 특징들이 문서요약 과정에서 고려되어야 한다. 기존의 일반적인 문서요약 방법은 하나의 주제를 기반으로 문서 전체에서 가장 중요한 문장으로 요약하기 때문에 다중 주제 회의록 요약에는 적합하지 않다. 제안한 방법은 먼저 어휘의 동시 발생 (co-occurrence) 빈도를 이용하여 회의록 분할 (segmentation) 과정을 수행한다. 다음으로 주제의 구분에 따라 분할된 각 영역 (block)의 중요 단어 집합 생성, 중요 문장 추출 과정을 통해 회의록의 중요 문장들을 선별한다. 마지막으로 추출된 중요 문장들의 위치, 종속 관계를 고려하여 최종적으로 회의록을 요약한다. AMI meeting corpus를 대상으로 실험한 결과, 제안한 방법이 baseline 요약 방법들보다 요약 비율에 따른 평가 및 요약문의 세부 주제별 평가에서 우수한 요약 성능을 보임을 확인하였다.
구문 분석 말뭉치를 이용한 ‘은/는’의 분포 연구: 주제부 설정을 위한 시론
[NRF 연계] 서강대학교 언어정보연구소 언어와 정보 사회 Vol.54 2025.03 pp.1-28
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper aims to examine the distribution of the josa eun/neun and discuss the necessity of positing the topic part. When examining the distance between the governor and dependent in the morphologically and syntactically analyzed corpus, unlike other josa, phrases that included the josa eun/neun that appeared at the beginning of a sentence often did not depend on the closest predicate. In order to reflect this phenomenon when analyzing Korean sentence structure, it is useful to view Korean sentences as having an inherent position for topic. Positing a topic part in the sentence structure of Korean has explanatory power in the following points. First, whatever sentence constituents, such as the subject, the object, and the adverb, are topicalized in the same way. Second, the realization of the topic is consistent in embedding sentences and conjunctive sentences. Third, whether eun/neun expresses a topic or a contrast is determined by the sentence structure.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.