년 - 년
An Experimental Comparison of the Usability of Rule-based and Natural Language Processing-based Chatbots KCI 등재 SCOPUS
한국경영정보학회 Asia Pacific Journal of Information Systems 제30권 제4호 2020.12 pp.832-846
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
Service organizations increasingly adopt data-based intelligent engines called chatbots in support of the interaction between customers and the companies. Two different types of chatbots have been suggested and introduced by companies leading the adoption of this emerging technology: rule-based chatbots and natural language processing-based chatbots. While the differences between these two types of technologies look relatively clear, the organizational and practical impacts of the differences have not been systematically explored. This study performed an experiment to compare the use of the two different types of chatbots used in practice by two comparable organizations. These two types of actual chatbots were used by Korean on-line shopping malls with similar business models (mobile shopping), length of history, size and reputation. The comparison was made based on such dimensions as usability, searchability, reliability and attractiveness. Contraty to conventional expectation that the superiority in technology will produce superior usability, the results show mixed superiority. The discussion on the reasons is presented.
한국피부과학연구원 아시안뷰티화장품학술지 제22권 제4호 통권 제82호 2024.12 pp.553-581
※ 기관로그인 시 무료 이용이 가능합니다.
6,900원
목적: 본 연구는 “학술연구정보서비스”에서 “아시안뷰티화장품학술지”와 국내 학술지 논문 제목에서 “화장품”을 키워드로 사용 하는 학술지의 연구 논문 초록을 크롤링하여, 이들 두 집단은 과연 차별화된 정체성으로 포지셔닝되어 있으며, 어떠한 유사도와 차 이점을 가지고 있을까? 에 대한 의문으로 시작하였다. 방법: 본 연구를 수행하기 위해 Python version 3.10.9 프로그램을 이용하여 학술연구정보서비스에서 아시안뷰티화장품학술지 433편의 연구 논문과 ‘화장품’이라는 키워드로 5,232편의 연구 논문을 크롤링하 여 전처리 과정을 거쳐서 군집 분석, 공저자 네트워크 분석, 토픽 모델링 분석, 유사도 분석을 시행하였다. 결과: 두 학술지 집단은 모두 공통적으로 화장품 산업의 핵심 요소인 피부, 소비자 행동, 브랜드 전략 등 다양한 주제를 다루고 있는 것으로 나타났으나, 세 부적으로는 연구 초점이 다르게 나타났으며, 정체성이 차별화되어 있는 것으로 나타났다. 결론: 공저자 네트워크 분석에서 나타난 각 중심성 분석 결과는 연구의 시너지 효과 극대화를 위한 각 공저자의 역할과 영향력을 알 수 있었으며, 아시안뷰티화장품학술지는 비교 대상 군과는 다른 고유한 차별화된 정체성을 가지고 있는 것으로 나타났다.
Purpose: The aim of this study was to determine whether the “Asian Journal of Beauty and Cosmetology” and domestic journals that use “cosmetics” as a keyword in the abstract of their paper, are positioned with differentiated identities and to identify the similarities and differences between them. All articles were procured from the “Research Information Sharing Service.” Methods: Python version 3.10.9 was used in this study to identify 433 research papers from the “Asian Journal of Beauty and Cosmetology” and 5,232 research papers with the keyword “cosmetic” from the “Research Information Sharing Service.” After preprocessing the articles, a co-author network analysis was performed, followed by cluster analysis and topic modeling analysis. Four types of similarity analyses were performed based on the results obtained. Results: The two journal groups were found to commonly cover a variety of topics such as skin, consumer behavior, and brand strategy, that are central to the cosmetics industry. However, the research topics had different central focuses, indicating distinct identities. Conclusion: The centrality analysis results from the co-authorship network analysis revealed the roles and influence of each co-author in maximizing the synergy of research. The identity of the Asian Journal of Beauty and Cosmetology was found be unique and differentiated, compared with the comparison group.
目的: 这篇研究始于一个关于以下问题:在“学术研究信息服务”中《亚洲美容学术杂志》与国内学术期刊的论文题 目中以“化妆品”为关键词的期刊的论文摘要进行网络爬行,对这两个群体是否真正具有差异化的本质定位,他 们有哪些相同点和不同点? 方法: 本研究使用Python 3.10.9版本识别出来自《亚洲美容与美容杂志》的433篇研 究论文以及来自“学术研究信息服务”的5,232篇关键词为“化妆品”的研究论文。对文章进行预处理后,进行合著 者网络分析,然后进行聚类分析和主题建模分析。根据获得的结果进行了四种类型的相似性分析。结果: 发现这 两个期刊组普遍涵盖了化妆品行业核心的各种主题,例如皮肤、消费者行为和品牌战略。然而,研究主题的中 心点不同,表现出不同的本性。结论: 合着网络分析的中心性分析结果揭示了每位合著者在最大化研究协同作用 中的作用和影响。与对照组相比,《亚洲美容与美容杂志》的本质是独特且有区别的。
4,000원
트위터 사용자는 팔로우, 리트윗 등을 사용하여 자신이 관심 있어 하는 트윗을 찾는다. 하지만 사용자가 3억여 명에 달하는 트위터에서 사용자가 관심 있는 트윗을 찾기는 힘든 일이다. 이를 해결하기 위해 본 논문에서는 사용자 맞춤형 트윗 추천 시스템을 개발하였다. 우선, 사용자에게 추천할 수 있을 만한 가치가 있는 트윗을 수집하기 위해 현재 트랜드를 수집하 고, 트랜드에 대해 이야기하는 인기 있는 트윗들을 수집한다. 이후 사용자를 분석하고 맞춤형 트윗을 추천하기 위해 사용자 의 트윗과 수집한 트윗을 범주화한다. 최종적으로 웹서비스를 이용하여 사용자에게 본인과 카테고리가 일치하는 트윗과 관 심사가 일치하는 사용자를 추천해준다. 결과적으로 67.2%로 적절한 트윗을 추천하였다.
Twitter users use ‘Following’, ‘Retweet’ and so on to find tweets that they are interested in. However, it is difficult for users to find tweets that are of interest to them on Twitter, which has more than 300 million users. In this paper, we developed a customized tweet recommendation system to resolve it. First, we gather current trends to collect tweets that are worth recommending to users and popular tweets that talk about trends. Later, to analyze users and recommend customized tweets, the users’ tweets and the collected tweets are categorized. Finally, using Web service, we recommend tweets that match with user categorization and users whose interests match. Consequentially, we recommended 67.2% of proper tweet.
자연어처리(NLP) 기반 텍스트 마이닝을 활용한 스타트업 경영성과 연구동향 분석 : RISS DB 중심으로 KCI 등재
대한산업경영학회 산업융합연구(구 대한산업경영학회지) 제23권 제4호 2025.04 pp.1-16
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
본 연구는 자연어 처리(NLP) 기반 텍스트 마이닝 기법을 활용하여 스타트업 경영성과에 대한 국내연구 동향을 체계적 으로 분석하는 것을 목적으로 한다. 한국교육학술정보원(RISS)에서 ‘스타트업 경영성과’, ‘스타트업 경영’을 키워드로, 전 기간의 초록을 웹 크롤링하여 수집하였으며 최종 323편의 문헌이 분석에 활용되었다. 분석과정에서 형태소분석, 단어빈도분석(TF), 역 문서 빈도분석(TF-IDF), N-gram 분석, 네트워크 분석, CONCOR 기법이 적용되었다. 분석결과, ‘창업’, ‘규제’, ‘영향’, ‘혁신’, ‘기술’과 같은 키워드가 핵심 주제로 도출되었으며, 자금 조달, 기술 혁신, 정책 지원, 기업가 정신이 주 연구 분야로 확인되었다. 키워드 간 구조적 연관성을 파악하였으며 7개의 군집을 도출하였다. 본 연구는 데이터 기반 분석을 통해 기존 문헌과의 차별성을 제시하였다. 향후 연구에서는 시계열 분석, 해외연구 동향분석, 머신러닝 기법을 활용한 연구가 필요할 것이다.
This study aims to systematically analyze domestic research trends on startup business performance using text mining techniques based on natural language processing (NLP). Abstracts were collected through web crawling from the RISS (Korea Education and Research Information Sharing Service) database using the keywords “startup business performance” and “startup management” across all available years, resulting in a final dataset of 323 academic papers. During the analysis process, various methods were applied, including morphological analysis, term frequency (TF), term frequency–inverse document frequency (TF-IDF), N-gram analysis, network analysis, and CONCOR clustering. The analysis revealed that core keywords such as “entrepreneurship,” “regulation”, “impact”, “innovation”, and “technology” emerged as central themes. Major research areas included funding, technological innovation, policy support, and entrepreneurship. Furthermore, the structural relationships among keywords were identified, and seven clusters were derived. This study contributes to the literature by offering a data-driven perspective that distinguishes itself from prior qualitative research. Future studies should consider time-series analysis, comparative analysis of international research trends, and the application of machine learning techniques.
논문 및 특허 데이터를 활용한 자연어 처리 기반 ICT 동향 분석 연구 KCI 등재
한국EA학회 정보화연구 제17권 2호 2020.06 pp.127-140
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
많은 기업들은 지속적으로 발전하는 정보통신기술(ICT)에 발맞추고자 이를 파악하는데 높은 관 심을 갖고 있다. 기존의 ICT 동향 분석에 관한 연구들은 비교적 적은 양의 데이터를 사용하거나 특정 종류의 데이터 혹은 특정 분야만을 중점적으로 분석하였다. 하지만, 정확한 기술 동향 파악을 위해서 는 한정된 분야가 아닌 전반적인 분야에 대한 분석이 필요하다. 따라서 본 논문에서는 전문가가 선정 한 주요 ICT 용어를 기반으로 논문 및 특허 데이터를 활용한 ICT 동향 분석 연구를 진행한다. 구체 적으로, ICT 용어가 포함된 문서의 빈도수, 이를 기반으로 한 단어의 위상, 그리고 단어 간 네트워크 분석을 수행한다. 2018년 및 2019년 발행 문서에 대해 각각 분석을 진행한 후, 결과를 비교하여 앞 으로의 ICT 동향을 제시한다.
Companies are increasingly interested in the identification of the trends in information and communication technology (ICT) as ICTs are rapidly developing. Previous research on the analysis of ICT trends utilized relatively small sized data or focused only on a specific data or field. In order to identify ICT trends clearly, it is necessary to analyze data covering the overall fields. To this end, in this paper, trend analysis is conducted using manuscript and patent data based on the ICT related keywords selected by experts. In detail, keyword frequency analysis, positioning of keywords based on their frequencies, and keyword network analysis are performed. Documents published in 2018 and 2019 are analyzed separately and compared to identify the future technology trends.
인공지능 기반 자연어처리를 적용한 욕창간호기록 분석 KCI 등재
한국융합학회 한국융합학회논문지 제12권 제10호 2021.10 pp.365-372
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구의 목적은 자연어처리에 의해 생성된 욕창간호진술문의 특성을 파악하고, 욕창 단계판별 예측정확 도를 평가하기 위함이다. 욕창관련 간호기록은 서술통계를 이용하여 분석하였고, 워드클라우드 생성기를 활용하여 욕창예방 간호기록에서 단어의 특성을 파악하였다. 딥러닝을 이용하여 욕창단계판별 정확도(accuracy ratio) 를 구하였다. 연구결과, 욕창의 단계에 대한 기록 중 2단계와 심부조직손상의심단계가 각각 23.1% 와 23.0 % 로 가장 많았고, 빈도수가 높은 핵심단어는 홍반, 수포, 가피, 부위, 크기 등으로 나타났다. 예측의 정확도가 높은 단계 는 0단계, 심부조직손상의심단계, 2단계 순으로 나타났다. 따라서, 이를 활용하여 임상적 의사결정지지 시스템으 로 개발된다면, 임상간호사의 욕창관리역량 향상 전략 개발에 기초가 될 수 있을 것이다.
The purpose of this study was to examine the statements characteristics of the pressure ulcer nursing record by natural langage processing and assess the prediction accuracy for each pressure ulcer stage. Nursing records related to pressure ulcer were analyzed using descriptive statistics, and word cloud generators (http://wordcloud.kr) were used to examine the characteristics of words in the pressure ulcer prevention nursing records. The accuracy ratio for the pressure ulcer stage was calculated using deep learning. As a result of the study, the second stage and the deep tissue injury suspected were 23.1% and 23.0%, respectively, and the most frequent key words were erythema, blisters, bark, area, and size. The stages with high prediction accuracy were in the order of stage 0, deep tissue injury suspected, and stage 2. These results suggest that it can be developed as a clinical decision support system available to practice for nurses at the pressure ulcer prevention care.
자연어처리 모델을 이용한 이커머스 데이터 기반 감성 분석 모델 구축 KCI 등재
한국융합학회 한국융합학회논문지 제11권 제11호 2020.11 pp.33-39
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
자연어 처리 분야에서 번역, 형태소 태깅, 질의응답, 감성 분석등 다양한 영역의 연구가 활발히 진행되고 있다. 감성 분석 분야는 Pretrained Model을 전이 학습하여 단일 도메인 영어 데이터셋에 대해 높은 분류 정확도를 보여주 고 있다. 본 연구에서는 다양한 도메인 속성을 가지고 있는 이커머스 한글 상품평 데이터를 이용하고 단어 빈도 기반의 BOW(Bag Of Word), LSTM[1], Attention, CNN[2], ELMo[3], KoBERT[4] 모델을 구현하여 분류 성능을 비교하였 다. 같은 단어를 동일하게 임베딩하는 모델에 비해 문맥에 따라 다르게 임베딩하는 전이학습 모델이 높은 정확도를 낸다 는 것을 확인하였고, 17개 카테고리 별, 모델 성능 결과를 분석하여 실제 이커머스 산업에서 적용할 수 있는 감성 분석 모델 구성을 제안한다. 그리고 모델별 용량에 따른 추론 속도를 비교하여 실시간 서비스가 가능할 수 있는 모델 연구 방향을 제시한다.
In the field of Natural Language Processing, Various research such as Translation, POS Tagging, Q&A, and Sentiment Analysis are globally being carried out. Sentiment Analysis shows high classification performance for English single-domain datasets by pretrained sentence embedding models. In this thesis, the classification performance is compared by Korean E-commerce online dataset with various domain attributes and 6 Neural-Net models are built as BOW (Bag Of Word), LSTM[1], Attention, CNN[2], ELMo[3], and BERT(KoBERT)[4]. It has been confirmed that the performance of pretrained sentence embedding models are higher than word embedding models. In addition, practical Neural-Net model composition is proposed after comparing classification performance on dataset with 17 categories. Furthermore, the way of compressing sentence embedding model is mentioned as future work, considering inference time against model capacity on real-time service.
자연어 처리 모델을 활용한 블록 코드 생성 및 추천 모델 개발 KCI 등재
한국정보교육학회 정보교육학회논문지 제26권 제3호 2022.06 pp.197-207
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
본 논문에서는 코딩 학습 중 학습자의 인지 부하 감소를 목적으로 자연어 처리 모델을 이용하여 전이학습 및 미세조정을 통해 블록 프로그래밍 환경에서 이미 이루어진 학습자의 블록을 학습하여 학습자에게 다음 단계에서 선택가능한 블록을 생성하고 추천해주는 머신러닝 기반 블록 코드 생성 및 추천 모델을 개발하였다. 모델 개발을 위해 훈련용 데이터셋은 블록 프로그래밍 언어인 ‘엔트리’ 사이트의 인기 프로젝트 50개의 블록 코드를 전 처리하여 제작하였으며, 훈련 데이터셋과 검증 데이터셋 및 테스트 데이터셋으로 나누어 LSTM, Seq2Seq, GPT-2 모델을 기반으로 블록 코드를 생성하는 모델을 개발하였다. 개발된 모델의 성능 평가 결과, GPT-2가 LSTM과 Seq2Seq 모델보다 문장의 유사도를 측정하는 BLEU와 ROUGE 지표에서 더 높은 성능을 보였다. GPT-2 모델을 통해 실제 생성된 데이터를 확인한 결과 블록의 개수가 1개 또는 17개인 경우를 제외하면 BLEU와 ROUGE 점수에서 비교적 유사한 성능을 내는 것을 알 수 있었다.
In this paper, we develop a machine learning based block code generation and recommendation model for the purpose of reducing cognitive load of learners during coding education that learns the learners block that has been made in the block programming environment using natural processing model and fine-tuning and then generates and recommends the selectable blocks for the next step. To develop the model, the training dataset was produced by pre-processing 50 block codes that were on the popular block programming language web site ‘Entry’. Also, after dividing the pre-processed blocks into training dataset, verification dataset and test dataset, we developed a model that generates block codes based on LSTM, Seq2Seq, and GPT-2 model. In the results of the performance evaluation of the developed model, GPT-2 showed a higher performance than the LSTM and Seq2Seq model in the BLEU and ROUGE scores which measure sentence similarity. The data results generated through the GPT-2 model, show that the performance was relatively similar in the BLEU and ROUGE scores except for the case where the number of blocks was 1 or 17.
보안공학연구지원센터(IJSH) International Journal of Smart Home Vol.4 No.2 2010.04 pp.1-10
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The one of the upcoming research stream of Computer Science and Engineering is Natural Language Processing (NLP), which is widely being used in design of Interfaces for Human-Computer Interaction. One of the basic applications of NLP, is design of “"Smart Homes”", in which; based on user input certain actions can be performed either locally inside the house or globally outside the house. The smart homes designed are either based on remote input provided to the system from one place in the house or limited to only certain type of actions which the system can handle. The current state of art of the Smart Home design methodologies, does not includes the design of customized systems capable of handling inputs from different gender i.e., in different pitches, similarly the methodologies does not provides facilities to handle input in different language. The systems existing are not capable of understanding the context of situation and determining actions based on context. The paper describes a Smart Home application which can be used by elderly people living alone in the home to serve their basic needs, which specifically includes security issues. The system is initially based on generic architecture, and can be further customized to user needs. The main component of the generic architecture is ability to fix a certain language or set of languages in which the input will be provided to the system. This enables the user to interact with the system in multiple languages, thus the specific instruction set is not limited to one language. The interface designed is based on speech input, can handle multiple languages and is not gender specific. The system designed is also “"Context Specific”", to understand the context of current state in which input is given to the system and perform the necessary action. The context understanding feature makes the system more specific to understand the urgency of action to be performed based on input. The system can deployed for a building incorporating a wireless sensor network, and provide a high quality, efficient contextsensitive data transmission facility.
국제문화기술진흥원 International Journal of Advanced Culture Technology(IJACT) Volume 12 Number 4 2024.12 pp.408-417
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Since the emergence of ChatGPT, transformer-based language models have become highly popular. This study utilizes a transformer-based approach to measure Korean sentence similarity for software testing. By doing so, we propose test cases using metamorphic relationships. The performance of the transformer models is then compared using similarity measures. First, we create a test set by transforming sentences from the Defense Daily according to specific rules. We then input these transformed sentences into the RoBERTa, Electra, and T5 models. We check whether the similarity measure between the original sentence and its variant satisfies the metamorphic relationship. The performance of each model is then compared using a similarity measure. In our experiments, the RoBERTa model satisfied metamorphic relations in MR5 and MR6, which involved transforming nouns and verbs into synonyms, and in MR7, which involved altering sentence order, with accuracy rates of 80%, 85%, and 88%, respectively. All tests passed except MR7 (77%) for the Electra model and MR1 (73%) for the T5 model. Finally, we compared the performance of each model. In the comparison, the Electra model outperformed the T5 model (99.96%) and the RoBERTa model (99.62%) with an accuracy of 99.97%.
Generating ER Diagrams from Requirement Specifications Based On Natural Language Processing SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.2 2015.04 pp.61-70
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
An Entity Relationship (ER) data model is a high level conceptual model that describes information as entities, attributes, and relationships. Entity relationship modeling designed to facilitate database design. The abstract nature of Entity Relationship diagrams can be discouraging task to both designers and student alike. This paper deals with the problem of extracting ER elements from natural language specifications using Natural Language Processing (NLP). The approach provides the opportunity of using natural language documents as a source of knowledge for generating ER data model. The structural approach is used to parse specification syntactically based a predefined set of on heuristics rules. Extracted words with its Part Of Speech (POS) mapped into entities, attributes and relationships, which are the basic elements of ER diagrams.
보안공학연구지원센터(IJSH) International Journal of Smart Home Vol.10 No.3 2016.03 pp.181-190
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Unicode font is used in coding system that assign a unique code to every symbol of scripts irrespective of their platform, and language. The Greek Unicoder receives 16-bit hexadecimal code of alphabet. The device has been designed to convert Greek language into different languages that our people could understand. This Unicode reader code has been implemented on 28nm FPGA platform called Kintex-7 FPGA. In this paper we are using frequency scaling technique and Design goal. In this paper power analysis is our main concern and we have studied about the power analysis at different frequencies keeping the temperature constant at 25 degree Celsius and maintaining the constant air flow.
보안공학연구지원센터(IJUNESST) International Journal of u- and e- Service, Science and Technology Vol.9 No.8 2016.08 pp.75-86
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In this paper we have aimed to design an energy efficient and thermally aware Latin Unicode Reader. Our design is based on 28nm FPGA (Kintex-7) and 40nm FPGA (Artix-7). In order to test the portability of our design, we are operating our design with respective frequency of different mobile architecture. For thermal analysis of our energy efficient design, we have taken temperatures of four different regions from reference. Latin Unicode reader takes 16-bit hexadecimal code of alphabet and clock input. At the end we can conclude that the maximum power consumption is at 2.2GHz and minimum power consumption is at 1.2GHz. When we talk in terms of temperature we can see that maximum power is consumed at 329.85K and minimum power is consumed at 294.15K. And also the power dissipation is less in the case of 40nm (Artix-6) and is more in the case of 28nm (Kintex-7). Changing the parameter (Temperature) doesn’t affect the clock power in both cases (Gated and Non-gated).
트랜스포머 기반 효율적인 자연어 처리 방안 연구 KCI 등재
국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제23권 제4호 2023.08 pp.115-119
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
현재의 인공지능에서 사용되는 자연어 처리 모델은 거대하여 실시간으로 데이터를 처리하고 분석하는 것은 여러 가지 어려움들을 야기하고 있다. 이런 어려움을 해결하기 위한 방법으로 메모리를 적게 사용해 처리의 효율성을 개선하 는 방법을 제안하고 제안된 모델의 성능을 확인하였다. 본 논문에서 제안한 모델의 성능평가를 위해 적용한 기법은 BERT[1] 모델의 어텐션 헤드 개수와 임베딩 크기를 작게 조절해 큰 말뭉치를 나눠서 분할 처리 후 출력값의 평균을 통해 결과를 산출하였다. 이 과정에서 입력 데이터의 다양성을 주기위해 매 에폭마다 임의의 오프셋을 문장에 부여하였 다. 그리고 모델을 분류가 가능하도록 미세 조정하였다. 말뭉치를 분할 처리한 모델은 그렇지 않은 모델 대비 정확도가 12% 정도 낮았으나, 모델의 파라미터 개수는 56% 정도 절감되는 것을 확인하였다.
The natural language processing models used in current artificial intelligence are huge, causing various difficulties in processing and analyzing data in real time. In order to solve these difficulties, we proposed a method to improve the efficiency of processing by using less memory and checked the performance of the proposed model. The technique applied in this paper to evaluate the performance of the proposed model is to divide the large corpus by adjusting the number of attention heads and embedding size of the BERT[1] model to be small, and the results are calculated by averaging the output values of each forward. In this process, a random offset was assigned to the sentences at every epoch to provide diversity in the input data. The model was then fine-tuned for classification. We found that the split processing model was about 12% less accurate than the unsplit model, but the number of parameters in the model was reduced by 56%.
[Kisti 연계] 한국정보관리학회 정보관리학회지 Vol.23 No.2 2006 pp.167-183
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 이메일에 나타난 감성정보 메타데이터 추출에 있어 자연언어처리에 기반한 방식을 적용하였다. 투자분석가와 고객 사이에 주고받은 이메일을 통하여 개인화 정보를 추출하였다. 개인화란 이용자에게 개인적으로 의미 있는 방식으로 콘텐츠를 제공함으로써 온라인 상에서 관계를 생성하고, 성장시키고, 지속시키는 것을 의미한다. 전자상거래나 온라인 상의 비즈니스 경우, 본 연구는 대량의 정보에서 개인에게 의미 있는 정보를 선별하여 개인화 서비스에 활용할 수 있도록, 이메일이나 토론게시판 게시물, 채팅기록 등의 텍스트를 자연언어처리 기법에 의하여 자동적으로 메타데이터를 추출할 수 있는 시스템을 구현하였다. 구현된 시스템은 온라인 비즈니스와 같이 커뮤니케이션이 중요하고, 상호 교환되는 메시지의 의도나 상대방의 감정을 파악하는 것이 중요한 경우에 그러한 감성정보 관련 메타데이터를 자동으로 추출하는 시도를 했다는 점에서 연구의 가치를 찾을 수 있다.
This paper describes a metadata extraction technique based on natural language processing (NLP) which extracts personalized information from email communications between financial analysts and their clients. Personalized means connecting users with content in a personally meaningful way to create, grow, and retain online relationships. Personalization often results in the creation of user profiles that store individuals' preferences regarding goods or services offered by various e-commerce merchants. We developed an automatic metadata extraction system designed to process textual data such as emails, discussion group postings, or chat group transcriptions. The focus of this paper is the recognition of emotional contents such as mood and urgency, which are embedded in the business communications, as metadata.
발달장애인을 위한 자연어처리 인공지능 --보조공학적 필요성과 가능성-
[NRF 연계] 성신여자대학교 인문과학연구소 人文科學硏究 Vol.51 2025.02 pp.285-320
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 자연어처리 기반 인공지능이 발달장애인을 위한 보조공학으로 기능할 능성을 제시하며, 향후 연구 및 기술 개발의 방향을 모색한다. 인공지능 기술이 발전을 거듭하면서, 다양한 장애 보조공학으로 활용되고 있다. 여러 인공지능 기술 가운데 자연어처리는 특히 발달장애인의 학습과 의사소통을 지원하는 데 유용할 수 있다. 본 연구는 자연어처리 기반 보조공학이 발달장애인에게 제공할 수 있는 세 가지 주요 이점을 논의한다. 첫째, 텍스트 분석의 방법을 통하여 저렴하면서도 신속한 진단/평가 도구로 활용될 수 있다. 둘째, 발달장애인의 일상적인 생활과 학습 과정에서 지속적인 지원 도구로 기능한다. 셋째, 언어적 패턴 분석을 통해 연구자의 분석 정확도를 향상하고, 데이터 기반의 객관적인 연구 환경을 조성할 수 있다. 넷째, 특수교육 현장에서 교사의 보조자로 기능하며, 발달장애 학생들의 사회적 상호작용을 촉진할 수 있다. 종합하자면, 자연어처리 환경은 멀티모달 인공지능에 대비하여 발달장애인의 조기 진단이나 의사소통 지원에 더 효과적으로 대응할 수 있는 수단이 된다. 그러나 이러한 기술을 효과적으로 구현하기 위해서는 몇 가지 선결 과제가 있다. 첫째, 학제 간 협력을 통해 발달장애인 연구와 자연어처리 연구를 유기적으로 결합해야 한다. 둘째, 한국어에 특화된 자연어처리 모델을 개발하여 영어 중심의 연구 환경을 극복할 필요가 있다. 셋째, 개인정보 보호와 모델 편향성 제거 등 윤리적 고려가 필수적이다.
This study explores the potential of Natural Language Processing-based Artificial Intelligence as an assistive technology for individuals with (neuro)developmental disorders. Additionally, it examines the necessary directions for future technological advancements. As AI continues to evolve, various assistive tools have been developed to support individuals with disabilities. Among these, NLP holds particular promise for aiding those with developmental disorders. This study highlights three key benefits of NLP-based assistive technology. First, text-based analysis methods can serve as a cost-effective diagnostic tool. Second, analyzing linguistic patterns enables a more objective and data-driven research environment. Third, NLP can assist special education teachers and facilitate social interaction among students with developmental disorders. However, several prerequisites must be addressed to ensure the effective implementation of this technology. First, interdisciplinary collaboration is essential to integrate research on developmental disorders with NLP-based AI. Second, developing Korean-specific NLP models is crucial, as current English-based research cannot be straightforwardly applied to Korean individuals. Third, ethical considerations, such as data privacy protection and bias mitigation, must be thoroughly examined and discussed.
언어학적 특징을 고려한 자연어 처리 기반 한국어 요약 시스템
[NRF 연계] 한국지식정보기술학회 (사)한국지식정보기술학회논문지 Vol.18 No.2 2023.04 pp.389-398
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
인터넷을 통해 대규모 데이터가 유통되면서 인터넷 이용자들은 자신에게 필요한 데이터를 찾기 어려워졌다. 또한, 데이터의 형태가 텍스트인 경우 텍스트를 압축 및 요약하는 작업이 요구되고 있는 실정이다. 본 논문에서는 한국어 텍스트에 대해서 학습한 Transformer Encoder-Decoder 기반의 KoBART(Korean Bidirectional and Auto-Regressive Transformers) 모델을 활용하여 한국어로 구성된 텍스트를 압축 및 요약하기에 효율적인 시스템을 구축하였다. 본 시스템은 추출요약을 수행하는 전처리기와 생성요약을 수행하는 KoBART 모델로 구성하였다. 전처리기는 한국어의 언어학적 특징을 고려하여 특정 문구가 출현하였을 경우 해당 문장을 중심으로 추출요약을 수행하고 KoBART 모델은 전처리기가 처리하지 않은 텍스트들에 대해서 생성요약을 수행한다. 제안하는 시스템은 한국어로 구성된 텍스트를 압축 및 요약하기 위하여 한국어의 언어학적 특징을 고려한 전처리기와 사전학습 언어모델인 KoBART 모델을 활용하였으며 일반적인 추출요약 모델과 생성요약 모델에 비해 우수한 성능을 보였다. 이는 특정 국가의 언어가 가지는 특징을 보다 상세하게 분석하고 활용하는 방안이 우수한 성능의 사전학습 언어모델만을 활용하는 것보다 좋은 결과를 기대할 수 있다는 점을 시사한다. 본 논문이 언어학적 특징과 우수한 성능의 사전학습 언어모델의 시너지를 전파하는데 선도하는 연구가 될 수 있을 것으로 기대된다.
As large-scale data is distributed through the Internet, it has become difficult for Internet users to find the data they need. In addition, when the form of data is text, it is required to compress and summarize the text. In this paper, we construct an efficient system for compressing and summarizing text composed of Korean using the Transformer Encoder-Decoder-based KoBART (Korean Bidirectional and Auto-Regulatory Transformers) model. This system consisted of a preprocessor that performs an extraction summary and a KoBART model that performs a generation summary. The preprocessor performs an extraction summary based on the sentence when a specific phrase appears considering the linguistic characteristics of the Korean language, and the KoBART model performs a generation summary on texts not processed by the preprocessor. The proposed system used the preprocessor considering the linguistic features of Korean and the KoBART model, which is a pre-learning language model, to compress and summarize text composed of Korean, and showed superior performance compared to the general extraction summary model and generation summary model. This suggests that a method of analyzing and utilizing the characteristics of a specific country's language in more detail can expect better results than using only a pre-learning language model with excellent performance. It is expected that this paper will be a leading study in spreading the synergy of the pre-learning language model with linguistic features and excellent performance.
운전면허 학과시험의 의미 편향성과 평가 공정성에 대한 자연어처리 기반 분석
[NRF 연계] 한국경찰학회 한국경찰학회보 Vol.27 No.2 2025.04 pp.67-90
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
운전면허 학과시험은 교통법규에 대한 이해와 안전 운전 습관을 평가하기 위한 중요한 절차이다. 그러나 지속적인 제도 개선에도 불구하고높은 평균 합격률이 유지되고 있으며, 이에 따라 응시자가 법규에 대한정확한 이해 없이 직관에 의존해 정답을 추정할 수 있는 시험 구조인지에대한 검토가 필요하다. 본 연구에서는 1,000개의 현행 운전면허 학과시험 문제를 대상으로 안전지향적 키워드의 영향력을 분석하고, 시험의 공정성과 실질적인 법규 평가 기능을 점검하였다. 이를 위해 TF-IDF(Term Frequency-Inverse Document Frequency) 기반의 키워드 도출, SBERT(Sentence-BERT)를 활용한 선택지 의미 분석, 랜덤 시뮬레이션기반 정답 예측 실험을 수행하였다. 분석 결과, 일부 문항에서는 안전지향적 키워드를 포함한 선택지가 실제 정답과 높은 연관성을 보였으며, 특히문장형과 사진형 문제에서 의미 유사도 기반 전략이 무작위 전략보다 높은정답률을 나타냈다. 그러나 전체 시험 수준에서는 이러한 전략의 성능차이가 통계적으로 유의하지 않았으며, 교통법규에 대한 학습 없이 합격하기는 어려운 구조임이 확인되었다. 이는 운전면허 학과시험이 일정 수준의 공정성을 유지하고 있음을 보여주나, 일부 문제 유형에서는 선택지표현이 편향적으로 구성될 여지가 있음을 시사한다. 본 연구는 운전면허학과시험의 평가 방식이 실질적인 법규 이해를 보다 정밀하게 측정하도록 개선될 필요성을 제기하며, 향후 관련 연구 및 제도 개선 논의에 기초자료로 활용될 수 있을 것이다.
The written driver's license test in South Korea serves as a critical procedure for evaluating one’s understanding of traffic laws and safe driving habits. Despite ongoing reforms, the test continues to exhibit high average pass rates, raising concerns that test-takers may rely more on moral intuition than on actual knowledge of traffic regulations. This study examines the potential influence of safety-oriented keywords on answer selection by analyzing 1,000 official test items from the current written exam. Using natural language processing techniques, including Term Frequency-Inverse Document Frequency (TF-IDF) for keyword extraction and Sentence-BERT (SBERT) for semantic similarity analysis, the study simulates answer prediction through a semantic similarity-based strategy and compares it to a random selection strategy. The results indicate that in certain question types?especially sentence-based questions? options containing safety-oriented keywords tend to correlate more strongly with correct answers. However, this pattern was not statistically significant across the entire test set, suggesting that the current test structure still requires a solid understanding of traffic laws to pass. While the test appears to maintain a certain level of fairness, the presence of linguistic biases in some question types implies that further refinement in test design may be warranted. This study contributes to ongoing discussions on improving the written driver's license exam to better assess actual knowledge and understanding of traffic regulations.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.