Earticle

현재 위치 Home

Research Article

Automatic Detection of Korean Learner Errors Based on K-ELECTRA
K-ELECTRA 기반 한국어 학습자 오류 자동 검출 연구

첫 페이지 보기
  • 발행기관
    한국인공지능교육학회 바로가기
  • 간행물
    인공지능연구 논문지 KCI 등재후보 바로가기
  • 통권
    Vol.7 No.2 (2026.06)바로가기
  • 페이지
    pp.72-81
  • 저자
    GilJa So
  • 언어
    영어(ENG)
  • URL
    https://www.earticle.net/Article/A488488

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

원문정보

초록

영어
Demand for error-detection components that automatically identify errors in learners' writing is increasing in Korean-as-a-foreign-language learning tools. General-purpose large language models (LLMs), however, suffer from hallucinations and inaccuracies on Korean grammatical knowledge, motivating learner-specialized approaches that fine-tune Korean-only pretrained models on learner corpora. This study empirically examines whether combining the National Institute of Korean Language (NIKL) Korean Learner Corpus with K-ELECTRA, a Korean-only pretrained model, can achieve detection performance suitable for deployment in learning tools. The two-stage detection target of this study is, first, to determine whether a sentence entered by a learner contains an error and, second, to identify the grammatical category (lexical, function-word, or orthographic) in which the error occurs. Building on the observation that the ErrorArea annotation scheme of the NIKL Learner Corpus directly corresponds to this target, the ErrorArea axis was adopted as the common basis for both detection and evaluation. A preprocessing pipeline resolving the structural mismatch between the NIKL XML and K-ELECTRA subword tokenization was designed alongside a BIO sequence labeling procedure; an ErrorArea head was placed on top of the K-ELECTRA encoder as the primary classification head; and a Focal Loss with γ=2.0 was adopted to handle class imbalance. On 27,764 evaluation sentences the proposed model attained sentence-level F0.5 of 0.8452 (precision 0.8541, recall 0.8115), satisfying the precision-first requirement of learning tools, where false alarms are considered more harmful than misses. The results show that learning-tool-grade detection precision can be achieved using only the combination of a Korean-only pretrained model and a learner corpus.
한국어
한국어를 외국어로 배우는 학습자를 지원하는 학습 도구의 핵심 컴포넌트로서, 학습자의 작문에 포함된 오류를 자동으로 검출하는 검사기의 수요가 증가하고 있다. 그러나 범용 거대 언어 모델(LLM)은 한국어 문법 지식에서 환각과 부정확성을 보이는 한계가 있어, 한국어 전용 사전학습 모델을 학습자 말뭉치로 미세 조정하는 학습자 특화 접근의 필요성이 부각된다. 본 연구는 국립국어원 한국어 학습자 말뭉치(NIKL)와 한국어 전용 사전학습 모델 K-ELECTRA를 결합하여, 학습 도구 적용 가능 수준의 한국어 학습자 오류 검출이 가능한지를 실증적으로 검토하였다. 본 연구의 2단계 검출 목표는 1차로 학습자가 입력한 문장의 오류 존재를 검출하고, 2차로 오류가 발생한 문법 범주(어휘·기능어·표기)를 식별하는 것이다. NIKL 학습자 말뭉치의 오류 부위(ErrorArea) 주석 체계가 이 목표와 직접 대응한다는 점에 착안하여, 오류 부위 축을 검출과 평가의 공통 기준으로 채택하였다. NIKL XML과 K-ELECTRA 서브워드 토큰화 사이의 구조적 불일치를 해소하는 전처리 파이프라인과 BIO 시퀀스 레이블링 절차를 설계하고, K-ELECTRA 인코더 위에 오류 부위 헤드를 주 분류 헤드로 둔 검출기를 구성하였으며, 클래스 불균형에 대응하기 위해 γ를 2.0으로 설정한 Focal Loss를 채택하였다. 평가 집합 27,764문장에 대한 실험 결과 문장 수준 F0.5 0.8452(정밀도 0.8541, 재현율 0.8115)를 달성하여, 거짓 경보가 누락보다 부정적이라는 학습 도구의 정밀도 우선 요건을 충족하였다. 본 연구의 결과는 한국어 전용 사전학습 모델과 학습자 말뭉치의 결합만으로도 학습 도구 수준의 검출 정밀도를 달성할 수 있음을 보인다.

목차

ABSTRACT
요약
I. Introduction
II. Related Work
1. Generative AI for Korean Learner Support
2. Korean Grammatical Error Correction
3. Studies Using the Korean Learner Corpus
III. Implementation of the Korean Error Detector
1. Data
2. Preprocessing
3. Training the Error Detector
IV. Experiments and Results
1. Experimental Environment and Configuration
2. Evaluation Methodology
3. Sentence-Level Error Detection Performance
4. Detection Performance by Erro Category
5. Error Analysis
6. Discussion and Directions for Improvement
V. Conclusion
Data and Code Availability
Acknowledgement
References

저자

  • GilJa So [ 소길자 | Department of Computer Engineering, Youngsan University, Republic of Korea ] Corresponding Author

참고문헌

자료제공 : 네이버학술정보

간행물 정보

발행기관

  • 발행기관명
    한국인공지능교육학회 [Korean Association of Artificial Intelligence Education]
  • 설립연도
    2019
  • 분야
    사회과학>교육학
  • 소개
    인공지능 기반의 융합 사회의 도래로 사회 전반에서 인공지능의 소양과 역량에 대한 요구가 증가하고 있습니다. 알파고 이후 인공지능은 우리 생활의 일부가 되고 있고 인공지능 기술이 융합 산업의 핵심으로 대두되었습니다. 인공지능기술이 다른 분야를 만났을 때 창출되는 가치는 자동차, 반도체, 스마트폰의 부가가치를 모두 합친 것보다 초월하고 있고 인공지능 역량을 가진 인재는 세상의 변화를 주도하는 막강한 영향력을 갖게 되었습니다. 이러한 인재의 양성은 혁신 기업의 존망을 좌우하게 되었고 국가의 경쟁력으로 이어지고 있습니다. 이것이 인공지능교육의 필요성이며 이를 이끌 단체로서 인공지능교육학회가 있습니다. 한국인공지능교육학회는 인공지능 기술과 융합적 역량을 가진 인재를 양성하고 미래 사회에서 인공지능이 인간을 위한 기술로 전개될 수 있도록 교육의 기반을 마련하고자 합니다. 학회에서는 인공지능에 관한 산학연 연계의 학문을 발전시키고 국가 발전에 기여하는 인재를 양성하는 등 다양한 방면에서 인공지능교육의 발전을 위해 노력하겠습니다. 또한 글로벌 인공지능과 융합 기술 분야에서 우리나라가 선도할 수 있도록 다양한 연구와 학술활동 그리고 국내외 공유의 장을 만들어 가도록 하겠습니다 .

간행물

  • 간행물명
    인공지능연구 논문지 [Journal of The Korean Association of Artificial Intelligence Education]
  • 간기
    연3회
  • pISSN
    2733-404X
  • 수록기간
    2020~2026
  • 등재여부
    KCI 등재후보
  • 십진분류
    KDC 000 DDC 006

이 권호 내 다른 논문 / 인공지능연구 논문지 Vol.7 No.2

    피인용수 : 0(자료제공 : 네이버학술정보)

    함께 이용한 논문 이 논문을 다운로드한 분들이 이용한 다른 논문입니다.

      페이지 저장