Earticle

현재 위치 Home

Human-Machine Interaction Technology (HIT)

Enhancing Gemstone Identification Generalization Performance : A Hybrid Alignment Model (HAM) for MSDA

첫 페이지 보기
  • 발행기관
    국제인공지능학회(구 한국인터넷방송통신학회) 바로가기
  • 간행물
    The International Journal of Advanced Smart Convergence 바로가기
  • 통권
    Volume 14 Number 4 (2025.12)바로가기
  • 페이지
    pp.42-69
  • 저자
    Choolha Hwang, Dongha Shim
  • 언어
    영어(ENG)
  • URL
    https://www.earticle.net/Article/A481176

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

원문정보

초록

영어
This paper presents a Hybrid Alignment Model (HAM) based on a multi-source domain adaptation (MSDA) framework that designates directly captured, unprocessed images (Set C) as a fixed reference domain. Domain bias remains a major obstacle to generalization, particularly for models trained on visually enhanced web or academic images (Sets A and B), where Top-1 accuracy drops by up to 27.2 percentage points (pp) when transferred to real-world data (A→A = 0.988 vs. A→C = 0.716). We aim to overcome this domain bias, which severely degrades out-of-domain generalization in computer-vision–based gemstone identification. HAM builds on the 2,048-dimensional embedding of a pre-trained ResNet-50 backbone and integrates three complementary alignment strategies—maximum mean discrepancy (MMD, global), adversarial alignment (Adv., local), and contrastive alignment (Conf., class-boundary)—to mitigate domain shift at multiple levels. Decision-boundary stability is further enhanced through three auxiliary strategies—pseudo-labeling (PL), adaptive template matching (ATM), and dynamic loss weighting (DLW). Empirically, HAM achieves substantial gains over the baseline: Top-1 accuracy improves from 0.716 to 0.935 (+21.9 pp) for A→C and from 0.775 to 0.956 (+18.1 pp) for B→C, effectively mitigating domain bias and enabling robust cross-domain generalization. Explainable AI (XAI) analyses using Grad-CAM, LIME, SHAP, and Integrated Gradients (IG) confirm that the model consistently focuses on expert gemstone cues—light dispersion, surface texture, and facet boundaries— supporting interpretability and trust. To the best of our knowledge, this study presents the first empirical integration of all three core alignment mechanisms (MMD, Adv., Conf.) within a unified MSDA framework, offering an explainable and practically reliable approach for real-world gemstone identification.

목차

Abstract
1. Introduction
2. Related Work
2.1 Dataset Bias and Domain Generalization
2.2 Limitations of Prior Work and Differentiation of This Study
3. Proposed Methodology
3.1 Overview and Data Configuration
3.2 Proposed Framework and Training Setup
3.3 Experiment Configuration and Evaluation
3.4 Set C Dataset Refinement and Curation Strategy
3.5 HAM-Based Experiment Scenario Definition and Configuration
4. Results
4.1 Domain Bias Diagnostic Experiments (40 Scenarios)
4.2 Analysis of Key Results from 40 Scenarios
4.3 Proposed model validation experiments (HAM-based 24 scenarios)
4.4 Comparison and Interpretation of HAM-Based Auxiliary Strategies
4.5 Performance Visualization and Explainability (XAI) Analysis
4.6 Key Findings and Implications
5. Discussion
5.1 Main Contributions
5.2 Limitations and Future Work
6. Conclusion
Reference

저자

  • Choolha Hwang [ Ph.D. Candidate, Department of Nano IT Design Fusion, Seoul National University of Science and Technology, Korea ]
  • Dongha Shim [ Professor, Department of Manufacturing Systems and Design Engineering, Seoul National University of Science and Technology, Korea ] Corresponding Author

참고문헌

자료제공 : 네이버학술정보

간행물 정보

발행기관

  • 발행기관명
    국제인공지능학회(구 한국인터넷방송통신학회) [The International Association for Artificial Intelligence]
  • 설립연도
    2000
  • 분야
    공학>전자/정보통신공학
  • 소개
    인터넷방송, 인터넷 TV , 방송 통신 네트워크 및 관련 분야에 대한 국내는 물론 국제적인 학술, 기술의 진흥발전에 공헌하고 지식 정보화 사회에 기여하고자 한다.

간행물

  • 간행물명
    The International Journal of Advanced Smart Convergence
  • 간기
    계간
  • pISSN
    2288-2847
  • eISSN
    2288-2855
  • 수록기간
    2012~2025
  • 십진분류
    KDC 326 DDC 380

이 권호 내 다른 논문 / The International Journal of Advanced Smart Convergence Volume 14 Number 4

    피인용수 : 0(자료제공 : 네이버학술정보)

    함께 이용한 논문 이 논문을 다운로드한 분들이 이용한 다른 논문입니다.

      페이지 저장