Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 5
No
1

통계 언어모델 기반 객관식 빈칸 채우기 문제 생성 KCI 등재

박영기

한국정보교육학회 정보교육학회논문지 제20권 제2호 2016.04 pp.197-206

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

빈칸 채우기 문제는 학생들이 학습 내용을 제대로 이해했는지 확인하기 위해 널리 사용되어 왔다. 이런 유형의 문제를 컴퓨터 알고리즘에 의해 자동으로 생성하는 많은 방법들이 제안되어 왔지만, 대부분 어떤 부분을 빈칸으로 만들면 좋을지에 대해 집중했기 때문에 적절한 보기를 자동으로 생성하는 연구는 미흡했다. 본 논문에서는 빈칸이 주어졌다고 가정하고, 이에 어울리는 보기를 자동 생성하는 알고리즘을 제안한다. 본 알고리즘은 통계 언어 모델에 기반하여 보기를 생성하기 때문에, 사람이 생성하는 경우보다 출제자에 편향되지 않은 보기를 제공할 수 있다. 또, 확률값에 기반하여 난이도를 자동으로 조절하는 것이 가능하기 때문에, 직접 사람이 문제를 만드는 것에 비해 상당한 비용 절감 효과가 있다. TEPS 문법, 어휘 시험에 대해 적용하여 실험한 결과, 사람과 유사한 결과를 생성함을 확인하였다. 향후 스마트 교육 분야에서 높은 활용도를 보일 것으로 기대한다.

A fill-in-the-blank with choices are widely used in classrooms in order to check whether students’ understand what is being taught. Although there have been proposed many algorithms for generating this type of questions, most of them focus on preparing sentences with blanks rather than generating multiple choices. In this paper, we propose a novel algorithm for generating multiple choices, given a sentence with a blank. Because the algorithm is based on a statistical language model, we can generate relatively unbiased result and adjust the level of difficulty with ease. The experimental results show that our approach automatically produces similar multiple-choices to those of the exam writers.

2

본 고에서는 MAP(maximum a priori) 추정에 기반을 둔 한국어 언어모델 적응을 제안한다. 먼저 언어모델을 위 한 기본 단위로 통계적 특징을 이용하는 WPM(word-piece model)을 제안한다. 이를 이용한 언어모델 적응 방법 으로 MAP 적응 알고리즘을 제안하였고 언어모델을 적응하지 않을 경우 및 전통적인 동적 주변 적응 방식과 비교하 였다. 성능 실험을 위해서 먼저 9천만 문장을 사용하여 베이스라인 언어 모델을 구했고 동일한 도메인에서 1천만 문장으로 시험한 결과 복잡도가 393.6 ppl(perplexity)을 구할 수 있었다. 베이스라인 언어 모델을 사용하여 SMS 분야로 시험한 결과가 적응 전 673.1 ppl에서 동적 주변적응을 하였을 경우에는 338.2 ppl, MAP 적응 알 고리즘을 사용한 경우에는 282.8 ppl이 되었다. 또한 동영상 강의 문장을 사용할 경우에도 적응 전에는 1340 ppl 을 보였으나 MAP 알고리즘에 의한 언어 적응 후에는 219.7ppl로 나왔다. 결론적으로 한국어에서 WPM을 기본 단위로 사용하고 MAP 언어모델 적응을 한 경우에는 베이스라인 언어모델의 복잡도보다 SMS, 동영상 각각의 도메 인에서 28.2%, 44.2% 감소되었다.

In this paper, we propose a Korean language model adaptation based on maximum a priori (MAP) estimation. The word-piece model (WPM) based on the statistical characteristic is proposed to use as basic units for language model. And we have compared our proposed MAP adaption algorithm with dynamic marginal adaptation algorithm for our language model adaption as well as language model without adaptation. For this purpose, we have built a baseline language model using 90 million sentences, which yields the perplexity (ppl) of 393.6 when experimental 10 million sentences are used as test sentences in the same domain. In the domain of short message service (SMS), we get the ppl of 673.1 when the language adaptation is not applied. However we can get the ppl of 282.8 after MAP adaption algorithm, the ppl of 338.2 after dynamic marginal adaption algorithm, respectively. And in the domain of video lecture, we get the same trend of performance, in which the ppl of 1340 before language adaptation reduces to the ppl of 219.7 after MAP language adaptation. In conclusion, MAP language adaptation algorithm yields ppl reduction of 28.2 % in the domain of SMS, 44.2 % in the domain of video lecture, respectively.

3

The Postprocessing of Optical Character Recognition Based on Statistical Noisy Channel and Language Model

Chang, JasonJ.S., Chen, Shun-Der

[Kisti 연계] 한국언어정보학회 한국언어정보학회 학술대회논문집 1995 pp.127-132

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

4

문맥의존 철자오류 후보 생성을 위한 통계적 언어모형 개선

이정훈, 김민호, 권혁철

[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.20 No.2 2017 pp.371-381

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The performance of the statistical context-sensitive spelling error correction depends on the quality and quantity of the data for statistical language model. In general, the size and quality of data in a statistical language model are proportional. However, as the amount of data increases, the processing speed becomes slower and storage space also takes up a lot. We suggest the improved statistical language model to solve this problem. And we propose an effective spelling error candidate generation method based on a new statistical language model. The proposed statistical model and the correction method based on it improve the performance of the spelling error correction and processing speed.

5

심층신경망으로 가는 통계 여행, 세 번째 여행: 언어모형과 트랜스포머

김유진, 황인준, 장기석, 이윤동

[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.37 No.5 2024 pp.567-582

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

지난 10년의 기간 심층신경망의 비약적 발전은 언어모형의 개발과 그 발전을 함께 해 왔다. 언어모형은 초기 RNN을 이용한 encoder-decoder 모형의 형태로 개발되었으나, 2015년 attention이 등장하고, 2017년 transformer가 등장하여 혁명적 기술로 성장하였다. 본 연구에서는 언어모형의 발전과정을 간략하게 살펴보고, 트랜스포머의 작동원리와 기술적 요소에 대하여 구체적으로 살펴본다. 동시에 언어모형, 트랜스포머와 관련되는 통계모형과, 방법론에 대하여 함께 검토한다.

Over the past decade, the remarkable advancements in deep neural networks have paralleled the development and evolution of language models. Initially, language models were developed in the form of Encoder-Decoder models using early RNNs. However, with the introduction of Attention in 2015 and the emergence of the Transformer in 2017, the field saw revolutionary growth. This study briefly reviews the development process of language models and examines in detail the working mechanism and technical elements of the Transformer. Additionally, it explores statistical models and methodologies related to language models and the Transformer.

 
페이지 저장