Earticle

다운로드

TOEFL11 코퍼스에서 주제 프롬프터의 자동 분류
Automatic classification of prompts in the TOEFL11 corpus

  • 간행물
    인문언어 KCI 등재 바로가기
  • 권호(발행년)
    제20권 1호 (2018.06) 바로가기
  • 페이지
    pp.157-177
  • 저자
    윤태진
  • 언어
    한국어(KOR)
  • URL
    https://www.earticle.net/Article/A334650

원문정보

초록

영어
The aim of the paper is to develop classification models of the prompts of the TOEFL essays in the TOEFL11 corpus. The corpus is a collection of TOEFL essays written in response to one of 8 prompts of various topics and by test-takers of different proficiency levels who are from 11 different countries. The number of essays is 11,000 for each language (that is, 121,000 in total). The paper aims at developing prompt classification models using an automatic method of Support Vector Machine (SVM), to which a number of different features are fed: The input features to the model include high frequency words which are observed in the raw essay texts, and high frequency nouns which are extracted from a POS-tagged essay texts. High frequency nouns among three different proficiency levels are also used as input features. The results indicated that even though high frequency words taken from raw textual materials performed quite well with an accuracy of 90.4%, the words tagged as nouns did even better with an accuracy of 97.3%. The inspecting of high frequency nouns revealed that the words were independently distributed among prompts with nearly no overlapping across different prompts. The classification test of essay samples of different proficiency levels confirmed that the accuracy rate of automatically classifying prompts by observing the frequency occurrence of nouns in texts increased in general as the proficiency levels of the essay samples increase. The paper serves as a foundation for further details studies on topic modeling used by learners of English.

목차

1. 서론
 2. 연구방법
  2.1 Corpus
  2.2 데이터의 선처리 (Preprocessing)
 3. 결과 분석 및 논의
 4. 결론
 인용문헌
 [Abstract]

저자

  • 윤태진 [ Tae-Jin Yoon | 성신여자대학교 ]

참고문헌

자료제공 : 네이버학술정보

    간행물 정보

    • 간행물
      인문언어 [LINGUA HUMANITATIS]
    • 간기
      반년간
    • pISSN
      1598-2130
    • 수록기간
      2000~2025
    • 등재여부
      KCI 등재
    • 십진분류
      KDC 705 DDC 405