영어 쓰기 수행평가에 적용한 AI 채점 : 엄격성 설정에 따른 평가 영역별 교사 채점과의 합치도
A Comparative Study of Human and AI Scoring in EFL Writing Assessments : Focusing on Agreement across Evaluation Domains.
This study investigates the consistency and agreement between scores generated by human raters and an artificial intelligence (AI)-based automated scoring system in the context of high school English writing performance assessments. To evaluate the pedagogical validity of AI-embedded scoring tools, 140 student essays were independently assessed by three raters and a commercial AI system using an analytical rubric. Statistical analysis utilizing Intraclass Correlation Coefficients was conducted to measure inter-rater reliability, the internal consistency of the AI system, and the degree of human-AI agreement. The results indicate that while human raters maintained moderate-to-high agreement and the AI system demonstrated high stability across iterated measures, the agreement between the two varied significantly across construct domains. Notably, a consistently low correlation was identified in the domain of Expression of Opinion. The findings suggest that while AI systems reliably process structured linguistic features such as lexis, grammar and textual coherence, they may operationalize different construct interpretations than human raters when evaluating context-sensitive competencies such as discursive meaning and rhetorical strategies. The paper concludes by underscoring the need for collaborative human-AI assessment models and further validation of construct validity in automated evaluation.
목차
Abstract 1. 서론 2. 연구배경 2.1. 영어 쓰기 수행평가 2.2. 인공지능 기반 영작문 평가 관련 연구 3. 연구방법 3.1. 연구설계 3.2. 연구대상 및 채점자 구성 3.3. 평가도구 및 루브릭 3.4. 채점 절차 및 자료 분석 4. 연구 결과 4.1. 인간 채점 및 AI 채점 조건별 기술통계 4.2. 인간 채점과 AI 채점 간 합치도 분석 5. 논의 및 결론 참고문헌