년 - 년
토픽모델을 활용한 택배 서비스 소비자와 종사자의 불만 사항 분석 KCI 등재
중소기업융합학회 융합정보논문지(구 중소기업융합학회논문지) 제10권 제2호 2020.02 pp.39-48
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
택배 업계의 서비스 품질 향상과 고객 만족에 영향을 미치는 요인들을 다양한 차원에서 분석한 연구가 많이 이루어져 왔다. 이러한 연구의 대부분은 한정된 응답자를 대상으로 설문조사나 인터뷰 등의 질적 방법이 사용 되어 자료의 형식과 내용, 응답자의 범위가 제한적이라는 한계가 있다. 이에 본 연구에서는 장기간에 축적된 불특 정 다수의 소비자 상담 사례와 업계 종사자의 불만이 반영된 기사문을 대상으로 삼아 택배 서비스에서 소비자와 공급자가 지적하는 불만에 관한 주요 토픽을 탐색하고 분석하여 선행 연구의 미비점을 보완하고자 하였다. 또한, 이러한 토픽이 시점에 따라 어떻게 변화하는지 흐름을 분석함으로써 최근에 제기된 새로운 토픽을 발굴하고 시사 점을 제시하고자 하였다. 이 결과 지연/분실/오배송 토픽과 택배 산업 경쟁 심화 토픽이 중심을 이루는 것으로 나타났다. 토픽 트렌드 분석 결과 최근 국제 택배 상담 내용이 다소 늘었고, 아파트 택배 배송과 관련된 갈등이 많이 다루어져 이를 정부 정책에 반영하거나 연구 주제로 다루어 볼 수 있을 것이다. 연구 결과 나타난 토픽은 선행 연구에서 다루어진 내용이 주를 이루지만 내밀한 상담사례와 학술 문헌 등 다른 자료와 분석 방법을 추가하면 더 새롭고 가치있는 토픽을 도출할 수 있을 것이라고 기대한다.
Many studies have been conducted to analyze factors that affect customer satisfaction, and service quality improvement in the parcel delivery industry. Most of these studies have a limited number of respondents using methods such as surveys and interviews. Therefore, this study aims to supplement the shortcomings of previous studies, by searching and analyzing the common major topics related to the complaints pointed out by consumers and suppliers in the parcel delivery service with cases of consumer counseling, and articles that reflect the complaints of workers in the industry. In addition, by analyzing the trend of these topics, we attempted to discover new topics and suggest implications. In conclusion, topics such as delay/lost/wrong deliveries as well as the fierce competition in the parcel delivery industry, turned out to be central aspects. As a result of the topic trend analysis, talks with international couriers have recently increased, and many conflicts related to apartment parcel delivery have been dealt with. The topics presented in this study are mainly focused on the contents of previous studies, but we expect that new and valuable topics can be derived by adding other data and analysis methods, such as internal counseling and academic literature.
국내 간호전문직관 연구 주제 동향 : 텍스트네트워크분석과 토픽모델링의 융합 KCI 등재
한국융합학회 한국융합학회논문지 제12권 제9호 2021.09 pp.295-305
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
본 연구의 목적은 국내에서 발표된 간호전문직관 연구 주제 동향을 양적 내용분석을 통해 탐색하는 것이다. 연구방법은 학술논문수집, 단어 정제 및 추출, 자료 분석의 절차를 수행하였다. 351편의 논문을 수집하여 영문초록에서 단어를 추출하여 텍스트네트워크를 개발하였고, 네트워크분석과 토픽모델링을 융합하여 자료를 분석하였다. 연구결과 핵심 주제는 간호사, 간호전문직관, 간호학생, 간호, 전문직자아개념, 보건의료인, 만족, 임상역량, 자기효능감 등이었 다. 토픽모델링을 통해 간호사 전문직관, 간호학생 전문직관, 간호전문직 정체성, 간호역량의 토픽그룹을 파악하였다. 시간이 흘러도 핵심 주제는 변화가 없었지만, 1990년대 역할갈등, 윤리적 가치, 2000년대 셀프리더십, 사회화, 2010년 대 임상실습스트레스, 지지체계와 같은 주제들이 부상하였다. 결론적으로 본 연구를 통해 국내에서 임상간호사와 간호 학생의 간호전문직관과 이에 영향을 미치는 요소들에 대한 연구가 활발하게 발표되고 있었으나, 간호전문직관 형성 및 향상에 효과적인 다차원적인 중재 전략을 모색한 연구는 부족하였음을 알 수 있었다.
The purpose of this study is to explore the trend of nursing professional research topics published domestically through quantitative content analysis. The research method performed procedures for collecting academic papers, refining and extracting words, and data analysis. A text network was developed by collecting 351 papers and extracting words from the abstract, and network analysis and topic modeling were performed. The core-topics were nurses, nursing professionalism, nursing students, nursing care, professional self-concept, health care professionals, satisfaction, clinical competence, and self-efficacy. Through topic modeling, topic groups of nurse’s professionalism, nursing students’ professionalism, nursing professional identity, and nursing competency were identified. Over time, core-topics remained unchanged, but topics such as role conflict and ethical values in the 1990s, self-leadership and socialization in the 2000s, and clinical practice stress and support systems in the 2010s have emerged. In conclusion, it is necessary to facilitate multidimensional interventional research to improve nursing professionalism of clinical nurses and nursing students.
Effect of Business Relatedness on Mergers and Acquisitions: Using 10-k Annual Reports
한국경영정보학회 한국경영정보학회 정기 학술대회 Digital Inclusion in Post Pandemic Era 2021.06 pp.92-96
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Business relatedness or similarity is a widely used measure in merger and acquisition studies. The effectiveness of the relatedness measure is central for managerial judgement and performance analysis of diversification decisions. Traditionally the Standard Industrial Classification (SIC Code) is used to represent business similarity, however the search for a more effective metric have led to the creation of alternative industry classifications and business relatedness measures. This paper uses text mining and topic modelling techniques to create new text based similarity metrics. The model is created from texts retrieved from 10-k annual reports submitted to the Securities and Exchange Commission. From these reports two sections of relevance are selected and used as sources for possible business capabilities that form our similarity measure. The new metrics are based on the topic distributions obtained from our Topic Model and on a dictionary of terms that had an impact on the topic model creation. These measures are then compared against the traditionally used SIC Code and a simple cosine similarity comparison of the 10-k text of merging firms.
위기관리 이론과 실천 한국위기관리논집 제19권 제12호 2023.12 pp.59-75
※ 기관로그인 시 무료 이용이 가능합니다.
5,100원
이 연구는 10·29 이태원 참사와 관련된 트윗을 대량으로 수집하고 정제한 후 DMR 토픽모델링으로 분석하였다. 모델 검증을 통해 가장 적절한 토픽의 수로 15개를 설정하였다. 총 15개 토픽을 클러스 터링하여 정보 반응, 감정 반응, 정부 비판 반응. 정치적 대립 반응, 2차 피해 반응, 공감 반응 및 기타 등 7개 그룹으로 묶을 수 있었다. 토픽의 분포를 시계열적으로도 살펴보았는데, ‘대책 미비’, ‘책임’, ‘애도와 추모’ 토픽의 비중은 시간이 지날수록 감소한 반면, ‘막말 논란’, ‘분노와 비난’은 시간이 지날수록 증가하였다. 이와 함께 전체 기간을 3개의 시기로 나누어 토픽의 차이를 살펴본 결과 시기별로 토픽 비중의 차이가 나타났는데 ‘애도기’와 비교해 볼 때 ‘수사기’와 ‘국정 조사기’에 서는 ‘수사와 구속’, ‘진상 규명’ 토픽의 비중이 증가하는 것을 볼 수 있었다. 이러한 토픽의 구조와 토픽의 시간별 변화에 따른 분석 결과에 기반하여 재난의 정치화와 공동체의 회복에 관점에서 연구 결과의 함의를 논의하였다.
This research involved the extensive collection and refinement of tweets related to the 10·29 Itaewon Tragedy, followed by analysis using DMR topic modeling. Through model validation, the optimal number of topics was determined to be 15, clustered into seven groups: informational, emotional, government criticism, political confrontation, secondary harm, empathetic responses, and miscellaneous. Examining the temporal distribution of topics revealed a decrease in the weight of topics such as 'Lack of government preparedness,' 'Accountability for the disaster,' and 'Condolences and memorials for victims' over time, while topics like 'Controversy over disrespect for victims' families' and 'Anger and Blame' increased. Further, dividing the overall period into three phases demonstrated variations in topic weights. Compared to the 'Mourning' phase, the 'Investigation' and 'Parliamentary Inquiry' phases indicated an increase in the weight of topics related to 'Investigation and detainment of those responsible' and 'Demand for the truth about the disaster.' The research concludes by discussing the implications for the politicization of disasters and community recovery.
The Topic Detection and Tracking with Topic Sensitive Language Model
한국어정보학회 한국어정보학 제8권 1호 2006.06 pp.64-68
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
In this paper, we explore the language model with topic sensitive features for the topic detection and tracking, formulate the relationship among the Chinese internet new words, language model with topic sensitive feature and the scheduling logic and the interval temporal reasoning and the key techniques. we use the Chinese internet new words to strengthen the detection and tracking of the topic and try to employ the scheduling logic and interval temporal reasoning to educe the reciprocal influences of events. At last we summarize the potential issues and the future work.
THE TOPIC DETECTION AND TRACKING WITH TOPIC SENSITIVE LANGUAGE MODEL
한국어정보학회 한국어정보학회 국제학술대회 종전과 광복 60주년 및 안의사 서거 95주년 기념 2005.08 pp.324-327
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
トピックモデルを用いた日本語テキスト マイニングの研究 -旧JLPTの読解の既出問題に対する分析を中心に- KCI 등재
한국일본언어문화학회 일본언어문화 제61집 2022.12 pp.27-46
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
In this paper, as one of the attempts to effectively utilize the vast amount of text data, I have introduced a text mining technique called Topic Model into the field of Japanese studies. Concretely, the texts of the reading comprehension parts of the previous format JLPT for the past 20 years were collected, and Topic Model analysis was carried out. The following points were made clear by such a study. First of all, it was confirmed from actual data that the subjects of the previous format JLPT tried to avoid topic-specific biases when selecting and producing the texts for the questions. Next, the text can be statistically classified into four main topics: “Private relationships such as family and work,” “Communications related to schedules,” “Public relations related to the country and society,” and “Economic activity.” The techniques and results of topic model analysis in this paper were empirical analyzes of actual existing questions. It is considered significant in that it can be applied to all fields of Japanese studies that are needed. Of course, the discussion in this paper is limited to the texts of the previous format JLPT, not the new format JLPT, and the amount of data is relatively small, although it covers all the data for the past 20 years. In addition, a comparative analysis with other texts was not possible. Therefore, it seems that there is still room for improvement in this paper, but I would like to address this as a future issue.
도로 위의 군비경쟁 : LDA 토픽모델을 활용한 SUV의 인기 요인 탐구 KCI 등재
한국디지털정책학회 디지털융복합연구 제18권 제10호 2020.10 pp.239-252
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
본 연구자들은 텍스트 마이닝을 활용하여 SUV 선호 증가의 요인을 탐색하고자 한다. 온라인 자동차 커뮤니티인 보배드림에서 2005년부터 2019년까지 작성된 SUV 관련 게시글 32,679개를 수집한 후, LDA 토픽모델링 기법을 적용 하였다. 분석 결과, SUV 담화에서 주요한 토픽으로 등장한 ‘안전’이 범죄로부터 개인의 위험에 주목한 기존 연구와 달리 교통사고 및 고속주행 상황에서의 안전을 의미하는 것으로 드러났다. 한국 사회의 SUV 소비는 개인이 운전하면서 느끼 는 불안과 위험에 대한 대비 수단을 의미한다고 볼 수 있다는 것이다. 또한, 이와 같은 위험 인식 저변에는 불평등 증대 로 인해 감소하는 타인에 대한 신뢰가 작동한다고 할 수 있다.
By using text mining, we explore the factors responsible for an increase in SUV preference. We collected 32,679 posts related to SUVs from “Bobaedream,” the largest online automobile community in South Korea, and applied the LDA topic model. While previous studies have explained the SUV boom as an individual’s risk aversion strategy from crime, the result shows that the topic of ‘Safety’ appears to be an important factor in the SUV discourse in the context of a car accident and high-speed driving situation. To conclude, the consumption of SUVs in Korean society serves as a mean to prevent anxiety and danger to individuals when driving. We insist that decreasing social trust, caused by an increase in inequality, underlies the perception of risk on the road.
LDA 토픽 모델을 활용한 포스트 Covid-19 시대의 소상공인 지원정책 분석 KCI 등재
대한산업경영학회 산업융합연구(구 대한산업경영학회지) 제22권 제6호 2024.06 pp.51-59
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 COVID-19와 같은 팬데믹 상황에서 소상공인에게 실질적으로 도움이 되는 정부 정책을 제언하는데 목적 이 있다. 이를 위해 ‘COVID-19 소상공인 지원’, ‘COVID-19 감염병 대응체계에 따른 소상공인 영향’, ‘COVID-19 소상공 인 경제정책’ 키워드를 중심으로 뉴스 기사를 크롤링하여 텍스트 마이닝 분석의 키워드 빈도분석과 워드클라우드 분석을 수 행하였고, LDA 토픽 모델링 분석을 통해 주요 이슈를 파악하였다. LDA 토픽 모델링을 수행한 결과 소상공인 지원 정책은 정 부의 현금성 지원과 금융지원으로 토픽 레이블을 구성하였고, COVID-19 감염병 대응체계에 따른 소상공인 영향은 정부 주 도의 방역체계와 개인 주도의 방역체계로 토픽 레이블을 구성하였으며, COVID-19 경제정책은 경제위기와 자생력을 갖추기 위한 소상공인 정책으로 토픽 레이블을 구성하였다. 구성한 토픽레이블을 중심으로 향후 팬데믹 상황에서 소상공인 피해 감면정책과 소상공인이 시장경쟁력 제고 정책에 대해 파악할 수 있는 기초자료를 제공하고자 하였다.
The purpose of the paper is to suggest government policies that are practically helpful to small business owners in pandemic situations such as COVID-19. To this end, keyword frequency analysis and word cloud analysis of text mining analysis were performed by crawling news articles centered on the keywords "COVID-19 Support for Small Businesses", "The Impact of Small Businesses by Response System to COVID-19 Infectious Diseases", and "COVID-19 Small Business Economic Policy", and major issues were identified through LDA topic modeling analysis. As a result of conducting LDA topic modeling, the support policy for small business owners formed a topic label with government cash and financial support, and the impact of small business owners according to the COVID-19 infectious disease response system formed a topic label with a government-led quarantine system and an individual-led quarantine system, and the COVID-19 economic policy formed a topic label with a policy for small business owners to acquire economic crisis and self-sustainability. Focusing on the organized topic label, it was intended to provide basic data for small business owners to understand the damage reduction policy for small business owners and the policy for enhancing market competitiveness in the future pandemic situation.
다이나믹 토픽 모델을 활용한 D(Data)ㆍN(Network)ㆍA(A.I) 중심의 연구동향 분석 KCI 등재
한국융합학회 한국융합학회논문지 제11권 제9호 2020.09 pp.21-29
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 디지털 사회의 도래로 다양한 데이터가 폭발적으로 증가하고, 그중 문헌 내 주제어를 도출하는 토픽 모델링 에 관한 연구가 활발히 진행되고 있다. 본 논문의 연구목표는 토픽 모델링 방법 중 하나인 DTM(Dynamic Topic Model) 모델을 적용해 D.N.A.(Data, Network, A.I) 분야에 대한 연구동향을 탐색하는데 있다. 실험 데이터는 최근 6년간(2015∼2020) ICT(Information and Communication Technology) 분야 중 기술대분류가 SWㆍAI에 해당하 는 연구과제 1,519개 사업에 대해 DTM 모델을 적용하였다. 실험결과로, D.N.A. 분야의 기술 키워드 Big data, Cloud, Artificial Intelligence와 확장된 의미의 기술 키워드 Unstructured, Edge Computing, Learning, Recognition 등 이 매년 연구에 표출되었으며, 해당 키워드 들이 특정 연구과제에 종속되지 않고 다른 연구과제에서도 포괄적으로 연구 되고 있음을 확인하였다. 끝으로 본 논문의 연구결과는 향후 D.N.A. 분야에 대한 정책기획ㆍ과제기획 등 연구개발 기획 과정과 기업의 기술 확보전략ㆍ마케팅 전략 등 다양한 곳에 활용될 수 있을 것으로 기대한다.
The Topic Modeling research, the methodology for deduction keyword within literature, has become active with the explosion of data from digital society transition. The research objective is to investigate research trends in D.N.A.(Data, Network, Artificial Intelligence) field using DTM(Dynamic Topic Model). DTM model was applied to the 1,519 of research projects with SWㆍA.I technology classifications among ICT(Information and Communication Technology) field projects between 6 years(2015∼2020). As a result, technology keyword for D.N.A. field; Big data, Cloud, Artificial Intelligence, extended keyword; Unstructured, Edge Computing, Learning, Recognition was appeared every year, and accordingly that the above technology is being researched inclusively from other projects can be inferred. Finally, it is expected that the result from this paper become useful for future policyㆍR&D planning and corporation’s technologyㆍmarketing strategy.
비정형 보안 인텔리전스 보고서 기반 토픽 자동 추출 모델 KCI 등재
한국융합학회 한국융합학회논문지 제10권 제6호 2019.06 pp.33-39
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
지능형 사이버 공격 기법이 다양화됨에 따라 보안 침해 사건, 글로벌 범죄 등의 사건 발생이 증가하고 있다. 지능형 공격을 예측하고 대응하기 위해서는 공격 기법의 특성, 수법, 유형을 파악해야 한다. 이를 위해 수많은 보안 기업 회사에서는 다양한 공격 기법을 빠르게 파악하고 더 큰 피해를 막기 위해 보안 인텔리전스 보고서를 배포한다. 하지만 각 기업에서 배포하는 보고서에 대한 형식이 맞춰져 있지 않으며, 대량의 비정형 보안 인텔리전스 보고서가 배포되고 있다. 본 논문은 비정형한 보안 인텔리전스 보고서에 대한 문제점을 고려하여 정형화된 데이터로 추출하는 방안을 제안 한다. 또한, 대량의 보안 인텔리전스 보고서를 파악하기 위해 소요되는 시간을 줄이고자 대량의 보고서를 주제별로 분류 할 수 있는 보안 인텔리전스 보고서 토픽 자동 추출 모델을 제안한다.
As cyber attack methods are becoming more intelligent, incidents such as security breaches and international crimes are increasing. In order to predict and respond to these cyber attacks, the characteristics, methods, and types of attack techniques should be identified. To this end, many security companies are publishing security intelligence reports to quickly identify various attack patterns and prevent further damage. However, the reports that each company distributes are not structured, yet, the number of published intelligence reports are ever-increasing. In this paper, we propose a method to extract structured data from unstructured security intelligence reports. We also propose an automatic intelligence report analysis system that divides a large volume of reports into sub-groups based on their topics, making the report analysis process more effective and efficient.
LDA 토픽모델링을 통한 ICT분야 국가연구개발사업의 주요 연구토픽 및 동향 탐색 KCI 등재
한국융합학회 한국융합학회논문지 제11권 제7호 2020.07 pp.9-18
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문의 연구목표는 LDA(Latent Dirichlet Allocation) 모델을 적용하여 국가연구개발사업을 통해 수행되고 있는 ICT(Information and Communication Technology) 분야의 연구과제에 대한 주요 연구 토픽과 동향을 탐색 하는데 있다. 연구방법에는 NTIS(National Science and Technology Information Service)로부터 최근 5년간 국가 연구개발사업의 전체 연구과제 정보를 다운로드받고 이를 정보통신기획평가원(IITP)의 EZone 시스템과 매칭하여 ICT 분야 연구과제 5,200건을 확보하고, 토픽모델링 기법중 하나인 LDA 모델을 적용하여 연구토픽과 연구동향을 조사하였 다. 실험결과로, ICT분야 연구과제에 대한 연구토픽은 인공지능, 빅데이터, 사물인터넷(Internet of Things)과 같은 지능정보기술로 확인되었고 연구동향에는 초실감미디어에 관한 연구가 활발히 진행되고 있음을 확인하였다. 끝으로 본 논문에서 진행된 국가연구개발사업에 대한 토픽모델링 결과는 향후 ICT분야 연구개발 계획 및 전략수립, 정책, 과제기 획 등 중요한 정보로 활용될 수 있을 것이다.
The research objectives investigates main research topics and trends in the information and communication technology(ICT) field, Korea using LDA(Latent Dirichlet Allocation), one of the topic modeling techniques. The experimental dataset of ICT research and development(R&D) project of 5,200 was acquired through matching with the EZone system of IITP after downloading R&D project dataset from NTIS(National Science and Technology Information Service) during recent five years. Consequently, our finding was that the majority research topics were found as intelligent information technologies such as AI, big data, and IoT, and the main research trends was hyper realistic media. Finally, it is expected that the research results of topic modeling on the national R&D foundation dataset become the powerful information about establishment of planning and strategy of future’s research and development in the ICT field.
Identifying Negative Sentiment with Sentiment Based LDA and Support Vector Machine Classification SCOPUS
보안공학연구지원센터(IJCA) International Journal of Control and Automation Vol.9 No.9 2016.09 pp.331-342
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
With the increasing development of online collaborative platforms, there emerge massive subjective texts. However, due to the massive negative news about eroticism, violence, extremity and corruption, as well as the influences of agitators and provocateurs, it is quite likely that Internet users can be turned from conscious individuals into unconscious groups, which contributes to the accumulation of public negative sentiment. In this work, we focus on the identification of sentiment and especially negative sentiment. Specifically, we introduce sentiment layer to the basic LDA topic model to map the texts into a lower dimensional space of topics and sentiment. Besides, we also consider the sentiment dictionary based sentiment feature word extraction method. By feeding the feature words into Support Vector Machine (SVM) classifier, we get the sentiment tendency of texts. Our experiments prove the efficiency of proposed method.
On-Line labeled Topic Model Based on Global and Local Topic
보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.8 No.12 2015.12 pp.269-282
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
A large number of electronic documents are labeled using human-interpretable annotations. High-efficiency text mining on such data set requires generative model that can flexibly comprehend the significant of observed labels while simultaneously uncovering topics within unlabeled documents. This paper presents a novel and generalized on-line labeled topic model based on global and local topic (GL-OLT) tracking the time evolution of topics in a sequentially organized multi-labeled corpus. GL-OLT topic model has an incrementally update principle based on time slices by an on-line fashion, and each label has not only a set of local topics, but also has several global topics. Empirical results are presented to demonstrate significant improvements accuracy of label predictive, and lower perplexity and high performance of our proposed model when compared with other models.
Time Label Topic Model SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.1 2015.02 pp.213-226
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Most of the models not aware of these dependencies on document time stamps. Not modeling time can confound co-occurrence patters and results in exchangeability of topic problem, which is important factor to deal with when finding dynamic topic discovery. This limitation has thus motivated work on developing a generalized framework for incorporating time information into topic models. Consequently, a topic model named Topics over Time (TOT) is proposed, which introduces a time node in topic model to handle the exchangeability of topics problem. However it lacks the capability to accommodate data type of side information. In this paper, we present a generative time LDA-style topic model with a variety of side information named Time Label Topic(TLT), which can find not only how the latent low-dimensional structure of document-response pairs changes over time, but also overcome the exchangeability of topics problem. Empirical results demonstrate significant improvements accuracy of time stamp and response variable prediction, and lower perplexity of our proposed model and dominance over other models.
Character Type Classification via Probabilistic Topic Model
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.5 No.2 2012.06 pp.123-140
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In this paper, we propose a method for character type classification based on a probabilistic topic model. The topic model is originally developed for topic discovery in text analysis using bag-of-words representation. Recent studies have shown the model is also useful for image analysis. We adopt the probabilistic topic model for character type classification. In our method, character type classification is carried out by classifying image patches based on their topic proportions. Since the performance of the method depends on a visual vocabulary generated by image feature extraction, we compare several feature extraction and description methods, and examine the relations to classification performance. In addition, by extending the method, we propose a coarse-to-fine approach to achieve stable character type classification for a small image patch. For that purpose, firstly, we partition an image into several patches which contain enough information to estimate the model parameters via EM algorithm. Then, each patch is subdivided into smaller patches. Estimation on the small patch is carried out by MAP-technique with a prior reflecting topic proportion of its parent patch. Through the experiments, we show accurate character type classification is made possible by the probabilistic topic model.
Image Holistic Scene Understanding Based on Global Contextual Features and Bayesian Topic Model
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.9 No.2 2016.02 pp.247-270
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Image holistic scene understanding based on global contextual features and Bayesian topic model is proposed. The model integrates three basic subtasks: the scene classification, image annotation and semantic segmentation. The model takes full advantage of global feature information in two aspects. On the one side, the performance of image scene classification and image annotation are boosted by incorporating image global contextual features; On the other side, the performance of image semantic segmentation is also boosted by new superpixel region segmentation method and new superpixel regions and patch feature representation. 1) For image scene classification and image annotation: (1) We improve the feature engineering methods by using the PHOW proposed by Vedaldi [1]; (2) Furthermore, global contextual features are learned by semantic features. 2) For semantic segmentation: (1) We improve the super-pixel segmentation method by using UCM in the literature [2]; (2)We proposed new feature representation for super-pixel region and patches by incorporating DSIFT, texton filter banks, RGB color, HOG, LBP and location features. The experiments testify that model performance has raised on all three sub-tasks.
Employing Latent Dirichlet Allocation Model for Topic Extraction of Chinese Text SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.7 2016.07 pp.51-66
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The hidden topic model of Chinese text, which possesses complicated semantics, is urgently needed, since China has occupied an increasingly significant role during the booming development of globalization over recent years. This paper details and elaborates the basic process of extracting latent Chinese topics by demonstrating a Chinese topic extraction schema based on Latent Dirichlet Allocation (LDA) model. Furthermore, the application was practiced in CCL, an authoritative Chinese corpus, to extract topics for its nine categories. With rigorous empirical analysis, extracting the LDA results has a considerably higher average precision rate as opposed to other three comparable Chinese topic extraction techniques; however the average recall rate is worse than KNN and almost the same with the PLSI model. Moreover, the recall rate and precision rate of LDA-CH is worse than LDA-EH. Therefore, the LDA model should be improved to adapt to the distinctive feature of Chinese words with the purpose of making it better for Chinese topic extraction.
토픽모델을 활용한 명문대 재학생의 학벌에 관한 인식 분석 KCI 등재
국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.10 No.3 2024.06 pp.503-512
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
이 연구는 한국 사회에서 명문대로 분류되는 한 대학의 학생이 작성한 학벌에 대한 글쓰기 과제를 분석하여 이들이 가진 학벌에 대한 인식을 확인하고 내재한 의미를 분류한 연구이다. 분석에서 활용한 방법은 토픽 모델 중 잠재 디리클레 할당 방 법으로 총 172편의 문서를 분석한 후 각 토픽에서 빈출한 키워드가 자주 등장하는 문서를 중심으로 학생의 인식을 탐색하 였다. 분석 결과 도출한 토픽은 학벌의 순기능(토픽 1), 양날의 검(토픽 2), 권력공동체(토픽 3), 승리의 징표(토픽 4), 학벌 의 역기능(토픽 5)의 다섯 가지이다. 각 토픽에서 가장 빈번하게 제시되는 단어를 정리하면 다음과 같다. 토픽 1에서는 ‘개 인’, ‘지위’, ‘수단’이, 토픽 2는 ‘정의(定義)’, ‘학교’, ‘의미’가, 토픽 3은 ‘사람’, ‘출신’, ‘권력‘이, 토픽 4는 ‘대학(교)’, ‘능력’, ‘노 력’이, 토픽 5는 ‘학력’, ‘우리나라’, ‘출신’이었다. 이상의 분석을 통해 우리는 명문대 학생이 학벌을 논할 때 계급과 학벌 공동체, 사회와의 관련성을 통하여 계급재생산을 고려하지만 인종 및 민족와 같이 학벌에 영향을 미치는 기타 요인에 대하여는 크게 관심을 두지 않고 있음을 확인하였다. 앞으로의 관련 강의에서 보다 다양한 요인과 학벌의 관련성을 다룰 필요가 있다.
This study examines the essays of academic background, written by students from a university, which is classified into prestigious universities in Korean society. By Latent Dirichlet Allocation, 172 essays were analyzed to explore the students’ perspectives of the academic fractionalism. The analysis identified five topics such as, functional aspects (Topic 1), double-edged nature (Topic 2), power communities (Topic 3), symbols of victory (Topic 4), and dysfunctional aspects (Topic 5). The most frequently appearing keywords are 'individual,' 'status,' and 'means' in Topic 1, 'definition,' 'school,' and 'meaning' in Topic 2, 'people,' 'origin,' and 'power' in Topic 3, 'university,' 'ability,' and 'effort' in Topic 4, and 'academic achievement,' 'South Korea,' and 'origin' in Topic 5. By exploring the topics, we found that students regarded class reproduction by education as important social issues and they showed little interest in other factors influencing academic fractionalism, such as race or ethnicity. these findings suggest that professars, who teach the impact of education on academic fractionalism, deal with the influence of diverse factors on academic fractionalism.
한국의 군사혁신(RMA) 담론 연구: 구조토픽 모델(Structural Topic Model)을 중심으로
[NRF 연계] 국방대학교 국가안전보장문제연구소 국방연구 Vol.65 No.4 2022.12 pp.119-146
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구의 목적은 군사혁신(RMA: Revolution in Military Affairs)에 관한 양적 담론 분석의 한 종류인 구조토픽 모델(STM: Structural Topic Model)을 통해 한국 군사혁신의 독특성을 탐색하고 국방 분야 발전에 주는 함의를 도출하는 것이다. 사회적 환경이 구성원 간 상호작용의 패턴에 영향을 주며 상호작용의 패턴에는 구성원의 관념이 내재해 있다는 이론을 적용하여 한국의 특수한 안보환경이 한국 군사혁신 담론에 영향을 주었을 것이라는 가설을 설정했다. 가설의 검정을 위해 2000년부터 최근까지 출판된 군사혁신 관련 연구 논문의 텍스트 데이터를 수집한 후, 양적 담론 분석의 한 종류인 STM을 활용하여 군사혁신 담론에 내재한 토픽을 추출했다. STM 분석 결과는 본 연구에서 제시하고 있는 가설을 강력히 지지하고 있다. 한국의 군사혁신 담론에는 군사혁신의 개념과 관련된 토픽이 더 많이 출현하는 것으로 나타났으며 전략적 수준의 미국 군사혁신과 연관된 토픽 및 4차 산업혁명과 관련된 중국 군사혁신에 대한 토픽이 덜 출현했다. 본 연구는 군사혁신 관련 전문가 집단 양성의 강화, 기술 중심적 접근의 회피, 전략적 및 국제정치적 맥락을 고려한 군사혁신을 추구해야 한다는 함의를 제시한다.
This study explores a structure of discourse regrading the Korean RMA(Revolution in Military Affairs) with STM(Structural Topic Model). Analyzing the discourse about the RMA has a great implication in terms of military and international politics, as well. To this end, we suggest a hypothesis that there exits an unique aspect in the Korean RMA comparing to others, due to a distinctive security environment. For testing the theoretical proposition, we collected text data from 220 academic articles about the RMA. The STM test results strongly support our hypothesis that the RMA discourse structure of South Korea is significantly different from it of all around the world. We found out that a topic regarding concept of RMA turns out to more likely appear in the Korean RMA discourse, while the topics about the U.S. and Chinese RMA on a strategic and international level less likely emerge in it. Our test results provide several important policy directions for improving the South Korean Security.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.