LLM-as-a-Judge 기반 지각 지표의 활용에 대한 탐색적 분석 : 반려동물 CQA 플랫폼에서의 답변 채택 여부를 중심으로
An Exploratory Analysis of LLM-as-a-Judge Based Perceptual Metrics : Focusing on Answer Adoption in a Pet CQA Platform
Answer quality evaluation in online community question answering (CQA) platforms has traditionally relied on indicators evaluating answer quality. In particular, both text-based indicators, which capture the objective quality of text, and perceptual indicators, which reflect perceived subjective quality, have been utilized. However, perceptual indicators are costly to collect and prone to quality issues. This study examines the applicability of using LLM-as-a-Judge–based perceptual indicators as an alternative in a Pet CQA platform. Machine learning results show that models incorporating both text-based and LLM-as-a-Judge–based perceptual indicators outperform the baseline in classifying answer acceptance. SHAP, LIME, and counterfactual analyses reveal that perceptual indicators play a central role, demonstrating that understandability and question relevance are key factors influencing answer acceptance, while text-based indicators have limited impact. Comparative analyses further show that accepted answers score higher across all text-based and perceptual indicators. This study contributes to the literature by confirming the applicability of LLM-as-a-Judge–based perceptual indicators in Pet CQA and identifying understandability and question relevance as key determinants of answer acceptance.
한국어
온라인 커뮤니티 질의응답 플랫폼 (Community Question Answering; 이하 CQA)에서 답변의 채택 가능성을 이해하기 위하여, 기존에는 답변의 품질을 평가한 지표를 활용하였다. 특히 텍스트의 특성으로 답변의 객관적 품질을 계산한 텍스트 지표와 지각된 주관적 품질인 지각 지표를 함께 활용해 왔으나, 지각 지표는 수집 비용이 높고 품질 저하 문제가 있다. 본 연구는 대안으로 반려동물 CQA에서 LLM-as-a-Judge 기반 지각 지표의 활용 가능성을 검토하였다. 첫째, 기계학습 분석 결과, 텍스트 지표와 LLM-as-a-Judge 기반 지각 지표를 활용한 모든 모델이 baseline에 비해 답변 채택 여부 분류 성능이 향상되었다. SHAP, LIME, 반사실적 설명 등 분석 결과, 이해가능성과 질문과의 관련성 등의 지각 지표가 분류에 핵심적으로 기여했으며, 텍스트 지표의 영향은 제한적이었다. 둘째, 채택 및 미채택 답변의 비교분석 결과, 채택 답변이 미채택 답변에 비해 모든 지각 지표의 값이 높은 것으로 나타났다. 이 연구는 반려동물 CQA 맥락에서 LLM-as-a-Judge 기반 지각 지표의 활용성과, 이해가능성과 질문과의 관련성이 채택 여부에 중요함을 밝혔다는 점에서 의의가 있다.
목차
요약 Abstract 1. 서론 2. 연구 방법 2.1 분석 데이터 수집 2.2 분석 데이터 전처리 2.3 변수 정의 및 측정 2.4 실험 설계 및 분석 방법 3. 분석 결과 3.1 지각 지표와 인간 평가자와의 상관분석 결과 3.2 기계학습 분석 결과 3.3 채택 및 미채택 답변의 품질 비교 분석 결과 4. 논의 5. 결론 Acknowledgements 참고문헌