Earticle

현재 위치 Home

문화 융합(CC)

다양한 비전 인코더를 활용한 이미지 캡셔닝 성능 비교 연구
A Comparative Study on the Performance of Image Captioning with Various Vision Encoders

첫 페이지 보기
  • 발행기관
    국제문화기술진흥원 바로가기
  • 간행물
    The Journal of the Convergence on Culture Technology (JCCT) KCI 등재 바로가기
  • 통권
    Vol.11 No.6 (2025.11)바로가기
  • 페이지
    pp.665-671
  • 저자
    남기훈
  • 언어
    한국어(KOR)
  • URL
    https://www.earticle.net/Article/A486581

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

원문정보

초록

영어
This study compared and analyzed the effects of various encoder structures on image captioning performance. Image captioning is a technology that describes input images using natural language sentences, offering diverse application possibilities such as assisting the visually impaired, image search, and autonomous driving. In this study, models were implemented by applying ResNet50, VGG16, Vision Transformer, Swin Transformer, Swin Transformer V2, and FasterViT-4 encoders based on a common decoder architecture. The Flickr8k dataset was used for training, and quantitative evaluation metrics such as BLEU, METEOR, ROUGE, CIDEr, and SPICE were applied alongside qualitative assessments considering contextual appropriateness. As a result of the experiment, the Transformer-based encoder generated more natural and rich captions due to its ability to understand global context, while the CNN-based encoder tended to produce fragmented descriptions focused on local features. FasterViT-4 demonstrated competitive performance by balancing computational efficiency and contextual expressiveness. In the contextual relevance evaluation, Vision Transformer, FasterViT, and Swin-based models demonstrated superior results, confirming that the scale of pre-training data and structural features influence performance. This study provides practical insights for encoder selection and hybrid architecture design when developing image captioning models.
한국어
본 연구는 다양한 인코더 구조가 이미지 캡셔닝 성능에 미치는 영향을 비교·분석하였다. 이미지 캡셔닝은 입력 이미지를 자연어 문장으로 기술하는 기술로, 시각 장애인 보조, 이미지 검색, 자율주행 등 다양한 응용 가능성을 가진다. 본 연구에서는 공통된 디코더 구조를 기반으로 ResNet50, VGG16, Vision Transformer, Swin Transformer, Swin Transformer V2 및 FasterViT-4 인코더를 적용하여 모델을 구현하였다. 학습에는 Flickr8k 데이터셋을 사용하였으며, BLEU, METEOR, ROUGE, CIDEr, SPICE 등 정량적 평가 지표와 문맥 적합성을 고려한 정성적 평가를 수행하였다. 실험 결과, Transformer 기반 인코더는 전역적 문맥 이해 능력으로 더 자연스럽고 풍부한 캡션을 생성하였으며, CNN 기반 인코더는 지역적 특징 중심으로 단편적 묘사가 많았다. FasterViT-4는 연산 효율성과 문맥 표현력에서 균형을 보여 경쟁력 있는 성능을 나타냈다. 문맥 적합성 평가에서는 Vision Transformer와 FasterViT, Swin 계열이 우수한 결과를 보였으며, 사전 학습 데이터 규모와 구조적 특징이 성능에 영향을 미침을 확인하였다. 본 연구는 이미지 캡셔닝 모델 설계 시 인코더 선택 및 하이브리드 구조 설계 방향에 실질적 시사점을 제공한다.

목차

요약
Abstract
Ⅰ. 서론
Ⅱ. 관련 연구
Ⅲ. 연구 방법
Ⅳ. 실험 및 결과
Ⅴ. 결론
References

저자

  • 남기훈 [ Ki Hun Nam | 정회원, 서경대학교 미래융합학부2 부교수 ] 제1저자

참고문헌

자료제공 : 네이버학술정보

간행물 정보

발행기관

  • 발행기관명
    국제문화기술진흥원 [The International Promotion Agency of Culture Technology]
  • 설립연도
    2009
  • 분야
    공학>공학일반
  • 소개
    본 진흥원은 문화기술(Culture Technology) 관련 산·학·연·관으로 구성된 비영리 단체이다. 문화기술(CT)은 정보통신기술(ICT), 문화적 사고 기반의 예술, 인문학, 디자인, 사회과학기술이 접목된 신융합기술(New Convergence Technology, NCT)로 정의한다. 인간의 삶의 질을 향상시키고, 진보된 방향으로 변화시키고, 문화기술 관련 분야의 학술 및 기술의 발전과 진흥에 공헌하기 위하여, 제3조의 필요한 사업을 행함을 그 목적으로 한다.

간행물

  • 간행물명
    The Journal of the Convergence on Culture Technology (JCCT) [문화기술의 융합]
  • 간기
    격월간
  • pISSN
    2384-0358
  • eISSN
    2384-0366
  • 수록기간
    2015~2026
  • 등재여부
    KCI 등재
  • 십진분류
    KDC 600 DDC 700

이 권호 내 다른 논문 / The Journal of the Convergence on Culture Technology (JCCT) Vol.11 No.6

    피인용수 : 0(자료제공 : 네이버학술정보)

    함께 이용한 논문 이 논문을 다운로드한 분들이 이용한 다른 논문입니다.

      페이지 저장