Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 16
No
1

대화시스템에서 자연어 생성은 대화관리 단계에서 결정한 시스템 발화의 의미표현을 사람이 이해할 수 있는 자연어 로 생성하는 것이다. 기존의 자연어 생성 연구는 의미표현에 대하여 매우 제한된 종류의 발화만을 생성하거나 문법 적으로 불완전한 발화를 생성한다는 문제점이 있다. 그래서 본 논문에서는 문제점들을 동시에 처리하기 위하여 Long Short Term Memory 기반의 언어모델을 이용한 한국어 자연어 생성 모델을 제안한다. 특히 우리는 시스템 발화의 다양성과 문법적 정확성을 높이기 위하여 빔서치 디코딩을 적용한다. 실험은 어절, 형태소, 음절단위에 따라 개별적으로 진행하였으며, 생성한 문장들은 정량적, 정성적 평가를 모두 진행하였다. 그 결과 형태소 단위로 학습한 제안모델에 빔서치 디코딩을 적용한 방법은 가장 좋은 성능을 보였다. 실제로 해당 생성 문장은 정량평가 결과에서 BLEU 지표는 0.86, Slot Error Rate 지표는 0.03을 기록하였으며 정성평가 역시 문법적으로 정확하고 문맥적 으로 충분히 자연스러운 결과임을 확인하였다.

Natural language generation in the dialogue system is a task that transforms the semantic frame of the system utterance determined in the dialogue management phase into a natural language that can be understood by humans. Existing studies have still faced some obstacles in that only very limited types of utterances or grammatically incomplete ones are generated from the semantic frames. In order to address these issues simultaneously, we propose a Korean natural language generation model using a long short term memory based language model. In particular, we exploit the beam search decoding method to obtain system utterances with diverse structures and grammatical correctness. The experiments were conducted individually with respect to the word, morpheme, and syllable units, and the generated utterances were evaluated in both quantitative and qualitative ways. As a result, the morpheme-based model with the beam search decoding has achieved the most robust result of all. In fact, in the quantitative evaluation result of the generated sentence, the BLEU-4 score was 0.86 and the SER was 0.03, and the qualitative evaluation was also confirmed to be grammatically correct and contextually natural.

2

자연어 처리 모델을 활용한 블록 코드 생성 및 추천 모델 개발 KCI 등재

전인성, 송기상

한국정보교육학회 정보교육학회논문지 제26권 제3호 2022.06 pp.197-207

※ 기관로그인 시 무료 이용이 가능합니다.

4,200원

본 논문에서는 코딩 학습 중 학습자의 인지 부하 감소를 목적으로 자연어 처리 모델을 이용하여 전이학습 및 미세조정을 통해 블록 프로그래밍 환경에서 이미 이루어진 학습자의 블록을 학습하여 학습자에게 다음 단계에서 선택가능한 블록을 생성하고 추천해주는 머신러닝 기반 블록 코드 생성 및 추천 모델을 개발하였다. 모델 개발을 위해 훈련용 데이터셋은 블록 프로그래밍 언어인 ‘엔트리’ 사이트의 인기 프로젝트 50개의 블록 코드를 전 처리하여 제작하였으며, 훈련 데이터셋과 검증 데이터셋 및 테스트 데이터셋으로 나누어 LSTM, Seq2Seq, GPT-2 모델을 기반으로 블록 코드를 생성하는 모델을 개발하였다. 개발된 모델의 성능 평가 결과, GPT-2가 LSTM과 Seq2Seq 모델보다 문장의 유사도를 측정하는 BLEU와 ROUGE 지표에서 더 높은 성능을 보였다. GPT-2 모델을 통해 실제 생성된 데이터를 확인한 결과 블록의 개수가 1개 또는 17개인 경우를 제외하면 BLEU와 ROUGE 점수에서 비교적 유사한 성능을 내는 것을 알 수 있었다.

In this paper, we develop a machine learning based block code generation and recommendation model for the purpose of reducing cognitive load of learners during coding education that learns the learners block that has been made in the block programming environment using natural processing model and fine-tuning and then gen­erates and recommends the selectable blocks for the next step. To develop the model, the training dataset was produced by pre-processing 50 block codes that were on the popular block programming language web site ‘Entry’. Also, after dividing the pre-processed blocks into training dataset, verification dataset and test dataset, we developed a model that generates block codes based on LSTM, Seq2Seq, and GPT-2 model. In the results of the performance evaluation of the developed model, GPT-2 showed a higher performance than the LSTM and Seq2Seq model in the BLEU and ROUGE scores which measure sentence similarity. The data results generated through the GPT-2 model, show that the performance was relatively similar in the BLEU and ROUGE scores ex­cept for the case where the number of blocks was 1 or 17.

3

생성형 AI 기반 감정형 일기 앱 서비스 설계 연구 KCI 등재후보

김유성, 김규석, 안영휘, 김도연

중소기업융합학회 산업과 과학 제4권 제6호 2025.11 pp.43-49

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

정신건강과 정서적 웰빙에 대한 사회적 관심이 높아지고 있는 현대 시대에서 감정을 표현 및 기록할 수 있는 디지털 도구에 대한 수요도 동시에 증가하고 있는 상황에서 감정일기 서비스가 정서 인식과 자기 성찰에 긍정적인 영향을 미친다는 점을 기존 연구들에서 반복적으로 제시해 왔다. 대표적으로 생성형 AI 기반 감정저널링 시스템인 MindScape 연구에서 대학생을 대상 으로 한 실험에서 부정 정서가 11% 감소, 긍정 정서는 7% 증가하는 효과를 확인했고 임상 환경에서 우울장애 환자를 대상으로 한 MindfulDiary 연구에서는 사용자의 감정 기록 빈도 증가, 감정 맥락 파악 용이성, 치료자와의 공감적 소통 향상 등 정성적·정 량적 효과가 실증되었으며 Resonance 시스템의 무작위 통제 실험(RCT)에서도 사용자에게서 PHQ-8 점수의 유의미한 감소가 관찰되어 생성형 AI 기반 저널링 도구의 정서 개입 효과가 재확인되었다. 이러한 사례는 생성형 AI가 단순한 텍스트 생성 기술을 넘어 정서 지원 도구로서 실제적 효용을 갖고 있음을 보여준다. 이러한 사회적, 학술적 수요에 착안해서 사용자의 감정을 기반으 로 맞춤형 일기 콘텐츠를 생성하는 생성형 AI 기반 감정형 일기 앱 서비스의 설계와 구현을 본 연구에서 제안한다. 감정 분류 체계와 사용자 인터페이스, 생성형 인공지능을 결합하고 감정 입력에 따라 개인화된 일기문을 생성하도록 설계 및 구현했다. 향후 감정 기반 콘텐츠 생성 서비스 개발 관련 분야의 다양한 연구로 확장될 것이다.

In a contemporary era where social interest in mental health and emotional well-being is growing, demand for digital tools that enable people to express and record emotions is also increasing. Previous studies have repeatedly shown that emotional diary services have a positive impact on emotional awareness and self-reflection. For example, in a study on MindScape, a generative AI-based emotion journaling system, an experiment on college students confirmed an 11% decrease in negative emotions and a 7% increase in positive emotions. In a study on MindfulDiary on patients with depressive disorders in a clinical setting, qualitative and quantitative effects were demonstrated, such as an increase in the frequency of users recording their emotions, an ease of understanding the context of their emotions, and an improvement in empathic communication with therapists. These examples demonstrate that generative AI goes beyond simple text generation and has practical utility as an emotional support tool. Inspired by these social and academic needs, this study proposes the design and implementation of a generative AI-based emotional diary app service that generates personalized diary content based on the user's emotions. We designed and implemented a system that combines an emotion classification system, a user interface, and generative AI to generate personalized diary entries based on emotional input. This will be expanded into various research areas related to the development of emotion-based content generation services.

4

스포츠 경기 이벤트 데이터를 활용한 실시간 자동 중계문 생성 시스템 개발 연구 KCI 등재

염두승, 박병권

한국골프학회 골프연구 제19권 특별호(제1호) 2025.05 pp.117-127

※ 기관로그인 시 무료 이용이 가능합니다.

4,200원

[목적] 본 연구는 축구 경기 이벤트 데이터를 기반으로 실시간 자동 중계문 생성 시스템을 구현하고, 템플릿 기반 방식과 KoAlpaca 기반 딥러닝 방식의 성능을 비교·분석하고자 하였다. [방법] StatsBomb Open Data에서 수집한 실제 축구 경기 데이터를 활용하였으며, 슈팅, 패스, 파울 등 주요 이벤트 1057건 중 약 120건을 정제하여 중계문 생성 시스템에 입력하였다. 템플릿 방식은 사전 정의된 규칙 기반 문장을, 딥러닝 방식은 KoAlpaca-Polyglot-12.8B 언어모델을 활용하여 중계문을 자동 생성하였고, 생성 결과는 BLEU 및 ROUGE 지표를 통해 정량적으로 평가하였다. [결과] 두 방식 모두 실시간 중계문 생성이 가능하였으나, KoAlpaca 기반 방식이 문장 표현의 다양성과 감정 전달 측면에서 더 높은 성능을 나타냈으며, BLEU-4 및 ROUGE-L 지표에서 템플릿 대비 향상된 결과를 보였다. [결론] 본 연구는 자동화된 스포츠 중계문 생성 시스템의 기초 구조를 제시하였으며, 향후 다양한 종목 및 언어 확장, 사용자 맞춤형 중계 서비스 개발로의 응용 가능성을 제시하였다.

[Purpose] This study aims to develop a real-time automatic sports commentary generation system using football match event data, comparing the performance of a template-based method and a deep learning-based KoAlpaca model. [Method] A total of approximately 120 key events, including shots, passes, and fouls, were extracted from StatsBomb Open Data for the 2018 FIFA World Cup match between Korea and Germany. Commentary was generated using rule-based sentence templates and the KoAlpaca- Polyglot-12.8B language model. The generated texts were evaluated quantitatively using BLEU and ROUGE metrics. [Result] Both approaches were able to generate commentary in real time. However, the KoAlpaca-based method demonstrated greater variation and contextual relevance in expression, achieving higher scores on BLEU-4 and ROUGE-L. [Conclusion] This study proposes a foundational framework for AI-based automatic sports commentary systems and suggests directions for future application, including expansion to various sports and languages and user-personalized commentary services.

5

A story generation task is to develop a system that can continuously generate natural, consistent, and coherent stories for consecutive scenes. Recently transformer-based language models have shown considerable results at the sentence-level generation for learning human-writing ability. However, it is very crucial to understand the way of developing the story using a combination of various contents. Recent works have mainly focused on human-guided AI story generation methods in which humans as guidance determine the next storyline, and the system creates a story that reflects the storyline well. This study focuses on the way of replacing the human role with the AI-based model. Based on this, this study deals with the methodology for creating a long story spanning multiple scenes rather than creating a story at the level of one scene. In this regard, we propose a novel AI-guided story generation framework with automatic storyline generator. It is a pipeline structure consisting of two modules such as a storyline generator and a story generator, which enables the continuous creation of coherent stories. Particularly, we transform the storyline generation problem into a multiple-choice QA problem to predict the next storyline. This study shows the possibility of generating continuous stories for multiple scenes without any human intervention.

6

A Survey of Automatic Code Generation from Natural Language

Shin, Jiho, Nam, Jaechang

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.17 No.3 2021 pp.537-555

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Many researchers have carried out studies related to programming languages since the beginning of computer science. Besides programming with traditional programming languages (i.e., procedural, object-oriented, functional programming language, etc.), a new paradigm of programming is being carried out. It is programming with natural language. By programming with natural language, we expect that it will free our expressiveness in contrast to programming languages which have strong constraints in syntax. This paper surveys the approaches that generate source code automatically from a natural language description. We also categorize the approaches by their forms of input and output. Finally, we analyze the current trend of approaches and suggest the future direction of this research domain to improve automatic code generation with natural language. From the analysis, we state that researchers should work on customizing language models in the domain of source code and explore better representations of source code such as embedding techniques and pre-trained models which have been proved to work well on natural language processing tasks.

7

영한 기계번역의 자연어 생성 연구

홍성룡

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.6 No.1 2005 pp.89-94

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

기계번역에서 자연어 생성의 목적은 입력언어의 어구 분석을 이용하여 그 문장의 의미를 변환해주는 목적 언어를 생성하는 것이다. 그것은 언어적 구조 낱말 전사. 대화체 언어, 어휘적 정보 등을 포함해야 한다. 본 연구에서는 대화체 자동 기계번역 시스템 구현계획의 일부인 음성, 음운 분야에서 담당하게 될 음성인식과 음성합성 알고리듬을 확립하기 위한 한국어 특질에 대한 기초조사를 하고자 한다. 또한 기계번역의 단계를 분석하여 형태소 분석 단계와 구문 분석 단계, 의미 분석 단계로 구분한다. 형태소 분석은 입력 문장을 받아 분리된 형태소를 사전 내에서 검색하여·품사 정보를 얻고 이웃하는 단어와의 접속 관계가 문법적으로 올바르게 되었는지를 점검한다. 본 연구의 결과가 대화체 기계번역 시스템 구현계획의 종합적 입장에서는 단순한 기초조사일 수 있지만, 한국어의 교육 및 기계번역 이해의 측면에서는 그 자체로 가치를 지닌다고 할 수 있겠다. 따라서 교육적 측면에서의 직접적 활용을 여러 측면에서 고려할 수 있을 것이다.

In machine translation the goal of natural language generation is to produce an target sentence transmitting the meaning of source sentence by using an parsing tree of source sentence and target expressions. It provides generator with linguistic structures, word mapping, part-of-speech, lexical information. The purpose of this study is to research the Korean Characteristics which could be used for the establishment of an algorism in speech recognition and composite sound. This is a part of realization for the plan of automatic machine translation. The stage of MT is divided into the level of morphemic, semantic analysis and syntactic construction.

8

생성형 AI를 활용한 자연어 요구사항 분석 기반의 반자동 코드 생성

진예진, 김장환, 김영철

[Kisti 연계] 한국정보처리학회 정보처리학회논문지 Vol.14 No.10 2025 pp.804-812

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

최근 생성형 AI 도구는 소프트웨어 개발 현장에서 많이 사용되고 있다. 그러나, 아직은 AI 기반 생성 결과에 대한 신뢰성을 보장할 수 없다. 또한, 현재도 AI 기반 코드 생성은 요구사항의 명확한 반영과 설계에 대한 추적이 부족하다. 이를 해결하기 위해, 소프트웨어 프로세스인 자연어 요구사항으로부터 UML 설계 및 스켈레톤 코드 생성까지 생성형 AI를 부분적으로 적용하고, 메타모델 기반 모델 변환하는 메커니즘을 제안한다. 이를 위해 변환 규칙 프로그래밍을 Metamodel Transformation Engine을 통해 자동 변환을 수행한다. 이는 자연어 요구사항의 정합성, 설계의 추적성, 그리고 코드의 품질을 확보할 수 있는 자동화된 개발 프로세스를 제공하고자 한다. 이를 통해 생성형 AI 코드의 신뢰성을 보완하고, 고품질의 소프트웨어를 개발하고자 한다. 향후 연구로, 이러한 메커니즘 통해 인간 중심의 코드와 AI가 부분적으로 접목된 코드를 비교하고자 한다. 그 결과로 어떠한 코드가 더 좋은 품질을 갖는 지 확인하는 것이 목표이다.

Recently, generative AI tools have been increasingly utilized in the field of software development. However, the reliability of AI-generated outputs still cannot be guaranteed. In particular, current AI-based code generation often lacks a clear reflection of requirements and sufficient traceability of the design process. To address these limitations, we propose a mechanism that partially applies generative AI and applies metamodel-driven model transformation across the software process, from natural language requirements to skeleton code generation via UML design. We define transformation rules that are automatically applied on a Metamodel Transformation Engine. This approach may provide an automated development process that ensures the consistency of natural language requirements, the traceability of design artifacts, and the quality of the generated code. In future research, we aim to compare human-written code with AI-assisted code through the proposed mechanism. The objective is to verify which approach ensures higher software quality and, by doing so, enhances the reliability of AI-generated code while supporting the development of high-quality software.

9

자연어 처리 모델을 활용한 퍼징 시드 생성 기법

김동영, 전상훈, 류민수, 김휘강

[Kisti 연계] 한국정보보호학회 정보보호학회논문지 Vol.32 No.2 2022 pp.417-437

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Fuzzing에서 seed corpus의 품질은 취약점을 보다 빠르게 찾기 위해서 중요한 요소 중 하나라고 할 수 있다. 이에 dynamic taint analysis와 symbolic execution 기법 등을 활용하여 효율적인 seed corpus를 생성하는 연구들이 진행되어왔으나, 높은 전문 지식이 요구되고, 낮은 coverage로 인해 광범위한 활용에 제약이 있었다. 이에 본 논문에서는 자연어 처리 모델인 Sequence-to-Sequence 모델을 기반으로 seed corpus를 생성하는 DDRFuzz 시스템을 제안한다. 본 논문에서 제안하는 시스템은 멀티미디어 파일을 입력값으로 하는 5개의 오픈소스 프로젝트를 대상으로 관련 연구들과 비교하여 효과를 검증하였다. 실험 결과, DDRFuzz가 coverage와 crash count 측면에서 가장 뛰어난 성능을 나타냄을 확인할 수 있었고, 또한 신규 취약점을 포함하여 총 3개의 취약점을 탐지하였다.

The quality of the fuzzing seed file is one of the important factors to discover vulnerabilities faster. Although the prior seed generation paradigm, using dynamic taint analysis and symbolic execution techniques, enhanced fuzzing efficiency, the yare not extensively applied owing to their high complexity and need for expertise. This study proposed the DDRFuzz system, which creates seed files based on sequence-to-sequence models. We evaluated DDRFuzz on five open-source applications that used multimedia input files. Following experimental results, DDRFuzz showed the best performance compared with the state-of-the-art studies in terms of fuzzing efficiency.

10

Test Cases Generation and Transformer Language Models Performance Comparison for Korean Natural Language Processing-based AI SW KCI 등재

In-Kyoun Lee, Min-Seong Jin, Gwang-Woon Lee, Gun-Woo Park

국제문화기술진흥원 International Journal of Advanced Culture Technology(IJACT) Volume 12 Number 4 2024.12 pp.408-417

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Since the emergence of ChatGPT, transformer-based language models have become highly popular. This study utilizes a transformer-based approach to measure Korean sentence similarity for software testing. By doing so, we propose test cases using metamorphic relationships. The performance of the transformer models is then compared using similarity measures. First, we create a test set by transforming sentences from the Defense Daily according to specific rules. We then input these transformed sentences into the RoBERTa, Electra, and T5 models. We check whether the similarity measure between the original sentence and its variant satisfies the metamorphic relationship. The performance of each model is then compared using a similarity measure. In our experiments, the RoBERTa model satisfied metamorphic relations in MR5 and MR6, which involved transforming nouns and verbs into synonyms, and in MR7, which involved altering sentence order, with accuracy rates of 80%, 85%, and 88%, respectively. All tests passed except MR7 (77%) for the Electra model and MR1 (73%) for the T5 model. Finally, we compared the performance of each model. In the comparison, the Electra model outperformed the T5 model (99.96%) and the RoBERTa model (99.62%) with an accuracy of 99.97%.

11

Toon Image Generation of Main Characters in a Comic from Object Diagram via Natural Language Based Requirement Specifications

Janghwan Kim, Jihoon Kong, Hee-Do Heo, Sam-Hyun Chun, R. Young Chul Kim

국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 13 Number 1 2024.03 pp.85-91

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Currently, generative artificial intelligence is a hot topic around the world. Generative artificial intelligence creates various images, art, video clips, advertisements, etc. The problem is that it is very difficult to verify the internal work of artificial intelligence. As a requirements engineer, I attempt to create a toon image by applying linguistic mechanisms to the current issue. This is combined with the UML object model through the semantic role analysis technique of linguists Chomsky and Fillmore. Then, the derived properties are linked to the toon creation template. This is to ensure productivity based on reusability rather than creativity in toon engineering. In the future, we plan to increase toon image productivity by incorporating software development processes and reusability.

12

음성인식과 자연어 처리 딥러닝을 통한 전자의무기록 자동 생성 시스템 KCI 등재

손현곤, 류기환

국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.9 No.3 2023.05 pp.731-736

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

최근 의료 현장은 전자의무기록, 전자건강기록 등의 의료 기록을 전산화하여 저장하고 관리하는 시스템이 의무적으로 적 용되거나 전체 의료 현장에 보급되어 환자 개개인의 과거 의료 기록을 추가적인 의료 행위에 활용하고 있다. 그러나 일반적인 의료 문진 및 상담 간 발생하는 의료진과 환자 간의 대화는 별도로 기록되거나 저장되지 않고 있어 추가적인 환자의 주요 정보 는 효율적으로 활용되지 못하고 있다. 이에 따라, 의료 문진 현장에서 발생하는 의료진과 환자와의 대화를 저장하고 이를 텍스 트 데이터로 변환하여 주요한 문진 내용만 자동으로 추출, 요약하여 정보화하는 음성인식과 자연어 처리 딥러닝을 통한 의료 상담 요약문을 자동으로 생성하는 전자의무기록 시스템을 제안한다. 본 시스템은 의료 종사자와 환자의 의료 상담 내용의 인식 과정을 거쳐서 텍스트 정보를 획득한다. 이렇게 획득된 텍스트를 복수의 문장으로 구분하고, 생성된 문장에 포함된 복수 키워드 의 중요도를 산출한다. 산출된 중요도를 기반으로 복수의 문장에 순위를 매기고, 순위를 기반으로 문장들을 요약하여 최종 전자 의무기록 데이터를 생성한다. 제안하는 시스템 성능은 정량적 분석을 통하여 우수함을 확인한다.

Recently, the medical field has been applying mandatory Electronic Medical Records (EMRs) and Electronic Health Records (EHRs) systems that computerize and manage medical records, and distributing them throughout the entire medical industry to utilize patients' past medical records for additional medical procedures. However, the conversations between medical professionals and patients that occur during general medical consultations and counseling sessions are not separately recorded or stored, so additional important patient information cannot be efficiently utilized. Therefore, we propose an electronic medical record system that uses speech recognition and natural language processing deep learning to store conversations between medical professionals and patients in text form, automatically extracts and summarizes important medical consultation information, and generates electronic medical records. The system acquires text information through the recognition process of medical professionals and patients' medical consultation content. The acquired text is then divided into multiple sentences, and the importance of multiple keywords included in the generated sentences is calculated. Based on the calculated importance, the system ranks multiple sentences and summarizes them to create the final electronic medical record data. The proposed system's performance is verified to be excellent through quantitative analysis.

13

자연어 생성기반 뉴스 보도 패턴 일반화 및 뉴스 구성에 따른 분류 가능성 소규모 LSTM 생성 데이터를 통한 내용 및 표현 형식 기반 뉴스 유형화 원리 고찰

윤호영, 안도현

[NRF 연계] 한국언론학회 커뮤니케이션 이론 Vol.19 No.1 2023.03 pp.84-123

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문은 인공지능 자연어 생성 모델을 통해 보도된 기사들의 일반화된 보도 내용데이터를 만들고, 이를 활용하여 이후 지도학습기반 뉴스 클러스터링 방식을 제안하는 연구이다. 보다 구체적으로는 뉴스를 수집하고 문장 기반 패턴 분석 방법을 활용하여, 이미 작성된 기사에서 쓰이는 단어와 주요 문장의 패턴이 추론된 자연어 생성기사 문장을 만들어낸다. 생성된 문장은 보도된 기사의 기본적인 보도 내용 및 관행을 보여주는 보도된 내용들의 일반화된 특질을 보여주는 것으로 본다. 그 다음 생성된 문장과 수집된 데이터 문장간의 내용 특질 유사성을 레벤슈타인 거리와 ROUGE 지표로 비교하여 컴퓨터가 만들어낸 문장과 실제 기사 문장 간의 내용과 표현상의괴리를 측정함으로써, 보도된 뉴스를 빠르게 유형화하는 방법을 제안한다. 본 글에서는 이러한 방법이 적용되는 과정을 소규모 데이터로 감염병 백신 보도를 주제로시연하고, 해당 방법이 가지는 의의와 향후 연구 가능성을 논의한다.

This paper proposes using natural language generation model with LSTM neural network for clustering news reporting pattern. More specifically, it suggests to collect news articles and to utilize a sentence-based natural language generation that can infer the general patterns of news reporting and then to compare the similarity of content features between the generated sentences and the collected data sentences. Levenstein distance and ROUGE-L metric are used for the comparison. These two metrics are to measure the content and expression similarity etween the computer-generated sentences and the actual article sentences. In doing so, we propose a rapid news clustering method that can be used for supervised learning in the later analysis stage. In this article, we demonstrate the application of these methods using small-scale data on infectious disease vaccine coverage and discuss the potential benefits of this method.

14

단어 간 관계 패턴 학습을 통한 하이퍼네트워크 기반 자연 언어 문장 생성

석호식, 작가멧, 장병탁

[Kisti 연계] 한국정보과학회 정보과학회논문지 : 소프트웨어 및 응용 Vol.37 No.3 2010 pp.205-213

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 단어간 관계 패턴을 학습한 후 이에 기반하여 자연 언어 문장을 생성하는 방법을 소개한다. 기존의 문장 생성 방법론에서는 내재된 문법 규칙의 존재를 가정하거나 템플릿을 사용하고 있으나, 본 논문에서 소개하는 방법론에서는 태깅 등의 부가 정보 없이 단어의 동시 등장 빈도만을 활용하여 단어간 관계 패턴을 학습한다. 단어간 관계 패턴은 하이퍼네트워크 방법론에 기반하여 학습되었다. 학습이 진행됨에 따라 하이퍼네트워크의 복잡도가 높아지며, 학습 모델에 축적되는 언어 관계 패턴의 수가 증가한다. 학습된 모텔의 유효성은 학습 패턴에 기반한 자연 언어 문장 생성을 통해 확인하였다. 실험 결과 학습이 진행됨에 따라 문법적으로 성립하는 문장의 비율이 향상하였다. 파서를 이용하여 생성된 문장을 구성하는 문법 규칙을 분석한 후 문법 규칙의 분포를 학습에 사용한 코퍼스의 문법 규칙 분포와 비교한 결과 학습에 사용된 코퍼스의 문법적 특성을 학습할 수 있는 잠재력을 갖고 있음을 확인하였다.

We introduce a natural language sentence generation (NLG) method based on learning of word-association patterns. Existing NLG methods assume the inherent grammar rules or use template based method. Contrary to the existing NLG methods, the presented method learns the words-association patterns using only the co-occurrence of words without additional information such as tagging. We employ the hypernetwork method to analyze and represent the words-association patterns. As training going on, the model complexity is increased. After completing each training phase, natural language sentences are generated using the learned hyperedges. The number of grammatically plausible sentences increases after each training phase. We confirm that the proposed method has a potential for learning grammatical properties of training corpuses by comparing the diversity of grammatical rules of training corpuses and the generated sentences.

15

뉴 미디어 시대의 자연어 데이터베이스 구축 - 영상KWIC 자동생성 기술을 중심으로 -

손영석

[NRF 연계] 한국일본어학회 일본어학연구 Vol.73 2022.09 pp.55-71

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Collection of examples is an important part that determines the basis of language research on actually used language data. The conventional collection method of usage examples can be primarily divided into two cases: (i) A case where a researcher personally collects examples from media, such as newspapers or novels, and (ii) a case where a researcher uses a public corpus constructed based on these media. However, the distinction between such cases became diluted with the emergence of new media. This is because of the shift that occurred from a print-based to a digital-based era, allowing anyone, including individual researchers, to collect examples and build a corpus easily. The software and online homepage introduced in this study are typical examples of easy establishment of a “multimedia corpus equipped with ‘keyword in context’ (KWIC) video automatic generation technique” based on user videos. Further studies need to be conducted to examine the characteristics of language research in an era where the boundary between collection of examples and the corpus construction is blurred.

16

M세대의 인성발달 과제: 자연어 사용을 중심으로

조정호

[NRF 연계] 한국인격교육학회 인격교육 Vol.10 No.1 2016.04 pp.65-78

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

오늘날 모바일 기기는 인간 삶의 중요한 부분이 되었다. 자연어를 사용하는 대면적 의사소통을 대체하는 모바일 커뮤니케이션이 급속히 증가하였다. 이것은 대부분의 사람들이 M세대(Generation Mobile, Gen M)로 살아가고 있음을 나타낸다. 그래서 이 연구에서는 M세대의 인성발달 과제를 자연어 사용 측면에서 탐색하였다. 본 연구의 결과, 인성은 인간관계를 전제로 하며, 인지적 영역에서는 주관성으로부터 객관성으로 발달해나가고, 정서 영역에서는 이기심으로부터 이타심으로 발달해나가며, 의지 영역에서는 타율성으로부터 자율성으로 발달해나간다. 이러한 인성은 일상의 인간관계에서 일어나는 현상이고, 일상에서의 인간관계는 자연어를 사용하여 이루어진다. 이를 고려하면 M세대는 대면적 인간관계 경험이 줄어드는 가운데 인성발달의 지체에 빠질 우려가 있으며, 모바일에 중독된 아동과 청소년들은 심각한 정서발달의 문제를 일으키고 있다. 이들은 모바일 기기가 대면적 의사소통을 대체하는 가운데 타인과 자연어를 사용하면서 예절, 인내, 감정 조절과 같은 정서적 인성을 올바르게 발달시키는 상호작용의 기회를 상실하였다. 따라서 M세대의 건전한 인성발달을 위하여 자연어를 사용하는 대면적 상호작용의 교육기회를 늘리고 관련 교육 프로그램을 강화할 필요가 있다.

Mobile devices these days have become an important part of our daily lives. Mobile use has increased rapidly in our daily lives over face-to-face communication with natural language use. This means that most people live as part of Generation Mobile (Gen M). This paper aims to search for the tasks for moral character development (MCD) of Gen M with a focus on the use of natural language. According to the previous study, in the tropism model, moral characteristics of humans develop from subjectivity toward objectivity in cognition, from egoism (self-love) toward altruism (other-love) in affect, and from heteronomy toward autonomy in volition. Moral character, therefore, is closely related to human relations in real life, and human relations are directly related to the use of natural language. Considering this, Gen M may have exhibit a certain amount of retardation in the affective domain of MCD such as emotional development, because mobile-phone addiction among children and youths raises serious problems of moral character. They have less opportunity for face-to-face interaction to develop facets of human relationships including manners, perseverance, and emotional control with the use of natural language as mobile communication reduces face-to-face conversations. Mobile devices may lead to us becoming emotionless robots ourselves. Consequently, it is necessary to increase various educational opportunities and programs of face-to-face interaction with use of natural language in order to develop a sound moral character for Gen M.

 
페이지 저장