생성형 AI를 이용한 통번역대학원생의 한영 영한 번역 평가 - 인간 평가자와 생성형 AI의 번역 평가 비교
Translation Quality Assessment Using LLMs : Comparing Human and LLM Ratings of Graduate Students’ Korean–English and English–Korean Translations.
This study examines the potential of large language models (LLMs) as translation assessment tools in graduate-level translator education by comparing human scores with those from ChatGPT-4o and Gemini 3.5. Utilizing a shared rubric, the study analyzes English-Korean and Korean-English translations produced by ten graduate students. Statistical analysis revealed strong, significant correlations between human and LLM mean scores in both directions (r = .785 for English-Korean; r = .863 for Korean-English). However, analysis by rubric category revealed significant differences in the terminology category, whereas content, syntax, and vocabulary & register showed none. Qualitative analysis suggests that LLMs tended to penalize proper-name, title, and spelling errors more strictly, while human assessors relied more on holistic judgment considering the broader translation and assessment context. The results suggest that LLMs can support preliminary scoring, formative feedback, and learner self-review, provided that explicit rubrics and carefully designed prompts are used with human oversight.
목차
1. 서론 2. 선행연구 검토 3. 연구 방법 3.1. 데이터 수집 3.2. 데이터 분석 3.2.1. 통계 분석 3.2.2. 정성적 분석 4. 분석 결과 4.1. 평가자간 신뢰도 (inter-rater reliability) 4.2. 항목별 비교 및 이상치(outlier) 분석 4.2.1. 세부 항목별 평가 비교 4.2.2. 개별 번역물 분석 5. 결론 참고문헌