This paper presents a Hybrid Alignment Model (HAM) based on a multi-source domain adaptation (MSDA) framework that designates directly captured, unprocessed images (Set C) as a fixed reference domain. Domain bias remains a major obstacle to generalization, particularly for models trained on visually enhanced web or academic images (Sets A and B), where Top-1 accuracy drops by up to 27.2 percentage points (pp) when transferred to real-world data (A→A = 0.988 vs. A→C = 0.716). We aim to overcome this domain bias, which severely degrades out-of-domain generalization in computer-vision–based gemstone identification. HAM builds on the 2,048-dimensional embedding of a pre-trained ResNet-50 backbone and integrates three complementary alignment strategies—maximum mean discrepancy (MMD, global), adversarial alignment (Adv., local), and contrastive alignment (Conf., class-boundary)—to mitigate domain shift at multiple levels. Decision-boundary stability is further enhanced through three auxiliary strategies—pseudo-labeling (PL), adaptive template matching (ATM), and dynamic loss weighting (DLW). Empirically, HAM achieves substantial gains over the baseline: Top-1 accuracy improves from 0.716 to 0.935 (+21.9 pp) for A→C and from 0.775 to 0.956 (+18.1 pp) for B→C, effectively mitigating domain bias and enabling robust cross-domain generalization. Explainable AI (XAI) analyses using Grad-CAM, LIME, SHAP, and Integrated Gradients (IG) confirm that the model consistently focuses on expert gemstone cues—light dispersion, surface texture, and facet boundaries— supporting interpretability and trust. To the best of our knowledge, this study presents the first empirical integration of all three core alignment mechanisms (MMD, Adv., Conf.) within a unified MSDA framework, offering an explainable and practically reliable approach for real-world gemstone identification.
목차
Abstract 1. Introduction 2. Related Work 2.1 Dataset Bias and Domain Generalization 2.2 Limitations of Prior Work and Differentiation of This Study 3. Proposed Methodology 3.1 Overview and Data Configuration 3.2 Proposed Framework and Training Setup 3.3 Experiment Configuration and Evaluation 3.4 Set C Dataset Refinement and Curation Strategy 3.5 HAM-Based Experiment Scenario Definition and Configuration 4. Results 4.1 Domain Bias Diagnostic Experiments (40 Scenarios) 4.2 Analysis of Key Results from 40 Scenarios 4.3 Proposed model validation experiments (HAM-based 24 scenarios) 4.4 Comparison and Interpretation of HAM-Based Auxiliary Strategies 4.5 Performance Visualization and Explainability (XAI) Analysis 4.6 Key Findings and Implications 5. Discussion 5.1 Main Contributions 5.2 Limitations and Future Work 6. Conclusion Reference
Choolha Hwang [ Ph.D. Candidate, Department of Nano IT Design Fusion, Seoul National University of Science and Technology, Korea ]
Dongha Shim [ Professor, Department of Manufacturing Systems and Design Engineering, Seoul National University of Science and Technology, Korea ]
Corresponding Author