We combine Gaussian Mixture Model-based clustering with SHAP (SHapley Additive exPlanations), an explainable AI technique, to segment loan customers based on credit risk characteristics and identify the key drivers of each cluster. Using unsecured loan data from a Korean financial institution, we classify customers into eight distinct clusters that differ in asset levels, loan structures, number of financial institutions used, and residential stability. We use SHAP analysis to find the variables that most strongly influence cluster assignment and to clearly explain how customer traits relate to credit risk. Some clusters show elevated risk due to complex debt structures or heavy reliance on non-bank lenders. Others demonstrate stable repayment behavior despite having limited credit histories or high liquidity. These patterns challenge conventional risk assessment frameworks that rely on simple thresholds or static indicators. By integrating probabilistic clustering with explainable machine learning, our approach reveals the nuanced and heterogeneous nature of borrower risk profiles. This framework supports more transparent and precise credit risk analysis and offers practical insights for lenders seeking to understand diverse customer segments.
한국어
본 연구는 설명 가능한 인공지능(Explainable AI) 기법인 SHAP(SHapley Additive exPlanations)과 GMM 기반 군집화(Gaussian Mixture Model-based Clustering)를 결합하여, 대출 고객의 리스크 특성을 중심으로 세분화하고 각 군집의 형성 요인을 정량적으로 분석하였다. 국내 신용대출 고객 데이터를 기반으로 GMM을 적용한 결과, 고객은 상이한 금융 행태와 리스크 특성을 지닌 여덟 개의 군집으로 분류되었으며, 자 산 규모, 대출 구조, 거래 기관 수, 주소 안정성 등에서 차이를 보였다. SHAP 분석을 통해 각 군집 형성에 영향을 미친 핵심 변수들을 도출하고, 변수별 기여도를 정량적으로 분석함으로써 군집 간 리스크 특성을 설명하였다. 분석 결과, 일부 군집은 복잡한 대출 구조나 비은행 금융기관에 대한 높은 의존도로 인해 신용 위험이 집중된 반면, 일부 군집은 짧 은 금융 이력이나 높은 유동성을 보유하고 있음에도 안정적인 상환 성향을 나타내는 등 전통적인 리스크 평가 기준과는 다른 양상을 보였다. 이는 금융 리스크의 해석에 있어 단일 지표나 정형화된 기준에 의존하기보다는, 고객군의 맥락적 특성과 다차원적 변수 간 관계를 종합적으로 고려해야 함을 시사한다.
목차
요약 I. 서론 II. 이론적 배경 및 선행연구 1. 군집 분석을 활용한 고객 세분화 2. 설명 가능한 인공지능(XAI)과 SHAP 기법 3. 선행연구 III. 연구 방법론 1. 가우시안 혼합 모형(GMM)을 활용한 군집화 2. 군집화 평가 방법 및 군집 수 결정 3. SHAP 기반 군집 해석 IV. 분석 및 결과 1. 분석 데이터 2. 군집분석과 XAI 3. 결과분석 V. 결론 참고문헌 Abstract