감성 변화와 강조 부사 편향을 활용한 한국어 비꼼 탐지 모델과 XAI 기반 단축학습 분석
A Korean Sarcasm Detection Model Using Sentiment Shift and Emphatic Adverb Bias with XAI-Based Shortcut Learning Analysis
This study aims to detect sarcasm that implicitly conveys negativity or mockery within Korean dialogue contexts. Sentiment shift between context and response, along with the biased occurrence of emphatic adverbs, are defined as key linguistic features and empirically validated through analysis of the KoCoSa dataset. Based on these features, SKoELECTRA is proposed by incorporating a three-dimensional sentiment feature vector and emphatic adverb special token marking into KoELECTRA. Through seven comparative experiments, the model applying solely the sentiment feature achieves a Balanced Accuracy of 76.5%, outperforming the baseline model KLUE-RoBERTa-large (74.9%), which has approximately three times more parameters. This result demonstrates that injecting domain-specific linguistic features into a lightweight model can surpass a larger model, suggesting that precise feature design is a more critical factor than model complexity. Furthermore, XAI analysis combining attention weight visualization and counterfactual masking experiments provides quantitative evidence suggesting a shortcut learning tendency arising from the model's over-reliance on emphatic adverbs, and discusses this as a structural limitation of the proposed approach.
한국어
본 연구는 한국어 대화 맥락에서 부정이나 조롱의 의미를 내포하는 비꼼 문장을 탐지하는 것을 목표로 한다. 문맥과 응답 간의 감성 변화와 강조 부사의 편향적 출현을 핵심 언어적 특성으로 정의하고 KoCoSa 데이터셋 분석을 통해 이를 실증한다. 이를 기반으로 KoELECTRA에 3차원 감성 피처 벡터와 강조 부사 특수 토큰 마킹을 결합한 SKoELECTRA를 제안한다. 7개의 비교 실험 결과 감성 피처만을 단독 적용한 모델이 Balanced Accuracy 76.5%를 기록하며 파라미터 수 약 3배 규모의 기준선 모델 KLUE-RoBERTa-large(74.9%)를 상회하는 최고 성능을 달성하였다. 이는 경량 모델에도메인 특화 피처를 주입하는 것만으로 대형 모델을 능가할 수 있음을 보여주며 모델 복잡도보다 핵심 피처의 정교한 설계가 더 중요한 요인임을 시사한다. 나아가 어텐션 가중치 시각화와 반사실 마스킹 실험을 결합한 XAI 분석을 통해 모델이 강조 부사에 과의존하는 단축학습 경향이 있음을 시사하는 정량적 근거를 제시하고 이를 모델의 구조적 한계로 논의한다.
목차
요약 Abstract 1. 서론 2. 관련 연구 2.1. 혐오 표현 탐지 2.2. 영어권 비꼼 탐지 연구 2.3. 대화 맥락을 활용한 감성·비꼼 탐지 3. 데이터셋 및 데이터 전처리 3.1. 사용한 데이터셋 3.2. 언어적 특징 분석 3.3. 전처리 파이프라인 4. 제안 모델:SkoELECTRA 4.1. 텍스트 인코딩 4.2. 감성 피처 4.3. 특징 융합 및 선형 분류 4.4. 하이퍼파라미터 5. 실험 설계 5.1. 비교 실험 설계 5.2. 평가 지표 6. 실험 결과 및 분석 6.1. 전체 성능 비교 6.2. 오류 분석 7. XAI 분석 7.1. 분석 목적 7.2 어텐션 가중치 시각화 7.3 반사실 마스킹 실험 8. 결론 REFERENCES
한국EA학회는 전사적 관점의 아키텍처 개념 및 원칙을 국내 민간기업 및 정부기관에 적용 확산시키고, EA 및 관련 분야의 연구, 전문인력의 양성 및 정책적 건의 등을 통해 기업 및 정부기관의 경쟁력 및 생산성을 향상시키고, 우리나라 지식 기반 산업 등의 고도화를 도모하는 것을 목적으로 합니다.