년 - 년
워드 임베딩은 자연어처리 분야에서 중요한 기술로, word2vec이 대표적인 알고리즘으로 알려져 있다. word2vec 을 비롯한 사전기반의 워드 임베딩 알고리즘들은 단어의 형태소특징을 사용하지 않는 방식, 즉 단어를 하나의 개체 로 사용하기 때문에 학습에 사용된 단어에 대해서만 단어 벡터를 만들 수 있는 한계를 가지고 있다. FastText는 이 문제를 해결하기 위해 제안된 알고리즘으로, 부속단어들의 조합으로 워드 임베딩을 하며, 이에 따라 학습에 사용된 적이 없는 단어에 대해서도 단어 벡터를 만들 수 있다. FastText는 형태소적 특징을 사용하기 때문에, word2vec 방식에 비하여 구문적 부분에서는 강점이 있고, 의미적 부분에서는 약점이 있다. 이 논문에서는 부속단어의 역문서 빈도를 이용하여 FastText를 개선하는 방법을 제시하며, FastText가 가지고 있는 의미적 부분에서의 약점을 극복 하고자 한다. 실험결과는 구문적 부분에서의 손실이 거의 없이 의미적부분에서 개선이 있었음을 보여준다. 또한 이 방법은 부속단어를 이용한 워드 임베딩에 모두 적용할 수 있다. 중의어를 구별하여 워드 임베딩하기 위해 고안된 확 률적 FastText에도 역문서 빈도를 적용 실험하고, 결과를 통해 성능이 향상되었음을 확인하고자 한다.
Word Embedding is important in natural language processing, and word2vec is known as a representative algorithm. Word2vec and many other dictionary based word-imbedding algorithms have limitations in creating word vectors only for words used in learning, because they does not use the words’ morphological feature. FastText is a proposed algorithm to solve this problem, word embedding in a combination of sub-words, thus creating a word vector for words that have never been used in learning. Because FastText uses morphological features, it has strengths in syntactic and weekness in semantic compared to word2vec. In this paper, the method of improving FastText is presented by using the inverse document frequency of the subword, and was intended to overcome the weakness in the semantic part of FastText. The results of the experiment show that there has been improvement in semantic tests with little loss in syntactic tests. this method can be applied to any word embedding algorithms using subwords. The probabilistic FastText designed to distinguish multi-sense words and was also tested with the inverse document frequency, and the results confirmed that the performance is improved.
기업의 조직문화가 재무성과에 미치는 영향에 대한 연구 : 텍스트 분석과 패널 데이터 방법을 이용하여 KCI 등재
한국경영정보학회 경영정보학연구 제26권 제1호 2024.02 pp.269-288
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
본 연구의 주요 목표는 기업의 조직문화가 재무성과에 어떤 영향을 주는지를 실증적으로 탐색하는데 있다. 이를 위해 우리나라의 대표적인 온라인 구인․구직 플랫폼인 잡플래닛(JobPlanet)으로부터 KOSPI 200에 포함된 58개의 기업을 선정하였고, 그 기업들의 조직문화를 파악하기 위하여 2014년부터 2022년, 9년 동안 해당 기업의 전․현직 구성원들이 잡플래닛에서 작성한 81,067개의 리뷰 데이터를 수집하여 분석에 이용하였다. 리뷰 데이터로부터 해당 기업의 조직문화를 정의하기 위하여 본 연구에서는 대표적인 텍스트 분석 기법인 Word2Vec와 FastText 분석 방법을 이용하여 Guiso et al.(2015)가 정의한 5가지의 조직문화 가치(Innovation, Integrity, Quality, Respect, and Teamwork)와 연관된 키워드들을 수정․보완․확장함으로써 새로운 조직문화 사전(Culture Dictionary)을 구축하였다. 이 사전을 기반으로 각 기업의 리뷰 데이터마다 어떤 조직문화 가치와 연관된 키워드가 많이 등장하였는지를 탐색하여 봄으로써 기업에서 어떤 문화가치가 상대적으로 강하게 나타나는지를 탐색하여 보았다. 한 걸음 더 나아가서 어떤 문화가치가 재무성과에 통계적으로 유의한 영향을 미치는지도 탐색하여 보았다. 연구 결과, 혁신과 창의성이 강조되는 혁신문화(Innovation)와 고객과 시장을 중시하는 조직문화(Quality)가 기업의 미래가치와 성장성을 나타내는 지표인 Tobin’s Q에 긍정적인 영향을 미치는 것을 확인할 수 있었고, 기업의 수익성을 나타내는 지표인 ROA에는 5가지의 조직문화 비율 변수 중 고객과 시장을 중시하는 조직문화(Quality)만이 통계적으로 유의미한 영향을 미치는 것을 확인할 수 있었다. 본 연구는 기존의 설문과 사례분석 기반의 조직문화 관련 논문과는 달리 대규모의 텍스트 데이터를 분석하여 조직문화를 탐색하고자 한 점에서 차별성을 찾을 수 있을 것이다.
The main objective of this study is to empirically explore how the organizational culture influences financial performance of companies. To achieve this, 58 companies included in the KOSPI 200 were selected from an online job platform in South Korea, JobPlanet. In order to understand the organizational culture of these companies, data was collected and analyzed from 81,067 reviews written by current and former members of these companies on JobPlanet over a period of 9 years from 2014 to 2022. To define the organizational culture of each company based on the review data, this study utilized well-known text analysis techniques, namely Word2Vec and FastText analysis methods. By modifying, supplementing, and extending the keywords associated with the five organizational culture values (Innovation, Integrity, Quality, Respect, and Teamwork) defined by Guiso et al. (2015), this study created a new Culture Dictionary. By using this dictionary, this study explored which cultural values-related keywords appear most often in the review data of each company, revealing the relative strength of specific cultural values within companies. Going a step further, the study also investigated which cultural values statistically impact financial performance. The results indicated that the organizational culture focusing on innovation and creativity (Innovation) and on customers and the market (Quality) positively influenced Tobin’s Q, an indicator of a company's future value and growth. For the indicator of profitability, ROA, only the organizational culture emphasizing customers and the market (Quality) showed statistically significant impact. This study distinguishes itself from traditional surveys and case analysis-based research on organizational culture by analyzing large-scale text data to explore organizational culture.
워드 임베딩을 이용한 질의 기반 한국어 문서 요약 분석 및 비교 KCI 등재
국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제19권 제6호 2019.12 pp.161-167
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
현재 ICT 기반의 웹 서비스 발달과 빠른 최신 기술의 보급으로 인하여 생성되는 정보의 양이 기하급수적으로 증가하고 있다. 이와 더불어 사용자들은 자신이 원하는 정보를 얻기 위해서는 많은 시간과 노력을 필요로 한다. 문서요약 기법은 사용자에게 주어진 문서의 문장과 핵심 단어들을 분석하여 효과적으로 요약문을 생성해주는 기술이다. 특히 한국 어로 이루어진 문서는 언어의 특성상 기존 언어 분석 기법들을 적용하기 어렵다는 문제점이 있다. 따라서 한국어의 특성 을 고려한 문서요약기법에 대한 연구가 필수적이다. 본 논문은 워드 임베딩 기법인 Word2Vec과 FastText를 활용하여 질의 기반의 한국어 문서요약 기법을 제안하고 그 결과를 비교 분석한다.
Recently, the amount of created information has been rising rapidly by dissemination of state of the art and developing of the various web service based on ICT. In additionally, the user has to need a lot of times and effort to find the necessary information which is the user want to know it in the mount of information. Document summarization is the technique that making and providing the summary of given document efficiently by analyzing and extracting the key sentences and words. However, it is hard to apply the previous of word embedding technique to the document which is composed by korean language for analyzing contents in the document due to the character of language. In this paper, we propose the new query-focused korean document summarization by exploiting word embedding technique such as Word2Vec and FastText, and then compare the both result of performance.
FastText 알고리즘을 이용한 사용자 지정 키워드 기반 동영상 요약 시스템
[Kisti 연계] 한국컴퓨터정보학회 한국컴퓨터정보학회 학술대회논문집 2023 pp.693-694
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 FastText 알고리즘을 기반으로 한 사용자 지정 키워드 기반 동영상 요약 시스템을 제안한다. 사용자가 키워드를 입력하면 시스템은 해당 키워드와 관련된 단어들을 FastText를 통해 추출하며, 이를 STT (Speech-to-Text)로 변환된 동영상에서 타임 스탬프 기반으로 인식한다. 인식된 키워드와 관련된 내용은 클립 형식으로 요약되어 사용자에게 제공된다. 본 연구의 목적은 숏폼 콘텐츠 환경에서 효과적인 콘텐츠 추출 및 제공을 통해 사용자 경험과 정보 제공의 효율성을 향상시키기 위함이다. 제안된 시스템은 사용자 지정 키워드에 맞춰 다양한 동영상 플랫폼에서 효율적인 영상 요약을 제공함으로써 온라인 동영상 환경에서 큰 혁신을 이끌어낼 것으로 기대된다.
fastText와 OpenCV를 이용하여 크리에이터 맞춤 영상자막 수정 방법 연구
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2019 pp.566-567
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
영상으로 콘텐츠를 개발하는 사람들을 '크리에이터'라 칭한다. 이들이 사람들에게 재미를 주고 이목을 끌기 위해 표준어 이외에 다양한 유행어와 신조어들을 만들어내며 이들을 영상뿐만 아니라 자막으로 활용하게 된다. 이러한 자막이 있는 영상 제작시 대본을 제작하는데 있어 자유도가 높은 크리에이터들의 특징상 맞춤법 오류 및 오타의 문제가 생긴다. 하지만 영상제작 도구에는 맞춤법 검사 기능이 없어 검사를 미리 하기에는 어려운 점이 있다. 우리는 이 문제점을 해결하기 위해 영상을 완성 하고 최종 검토를 할 때 맞춤법 검사를 하기 쉽도록 프로그램을 개발한다. OpenCV를 통해 영상의 자막을 글자로써 인식을 하고, fastText 모델을 통해 인식된 글자가 맞춤법에 맞는지 크리에이터에게 제안해주는 맞춤형 프로그램을 개발하고자 한다.
한국어 텍스트의 토큰화 방법에 관한 언어학적 연구 - fastText 단어 임베딩을 이용하여 -
[NRF 연계] 연세대학교 언어정보연구원 언어사실과 관점 Vol.55 2022.02 pp.309-354
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
.
Predicting the gender of Korean personal names using fastText
[NRF 연계] 한국음운론학회 음성음운형태론연구 Vol.27 No.3 2021.12 pp.483-500
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Male and female names tend to have distinct phonotactic characteristics in many languages. This paper explores the use of fastText, a neural-network text-classifier using sub-word information, in predicting the gender of Korean personal names, and compares the results with the results from a maximum-entropy model of phonotactics (Hayes and Wilson 2008). In this study, fastText is trained with training data consisting of 6400 Korean personal names, labeled with male and female. The model is tested with testing data of 35 Korean names. The fastText results positively correlated with Korean speakers’ ratings on the gender of the names. It outperformed the maximum-entropy model in terms of correlation with human ratings and accuracy of the labels. Yet, while the maximum-entropy model has OT-style constraints allowing generative linguists to interpret the results, fastText does not offer such interpretability. An error analysis is presented for the names where the models made incorrect predictions, using OT-style constraints.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.