Earticle

현재 위치 Home

A Probabilistic Visual Question Answering Model Based VQA
VQA 기반의 확률론적 시각적 질문 답변 모델

첫 페이지 보기
  • 발행기관
    한국컴퓨터게임학회 바로가기
  • 간행물
    컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) KCI 등재 바로가기
  • 통권
    제35권 제3호 (2022.09)바로가기
  • 페이지
    pp.65-73
  • 저자
    Manva Trivedi, Sabah Mohammed
  • 언어
    영어(ENG)
  • URL
    https://www.earticle.net/Article/A418293

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

원문정보

초록

영어
Visual data is present everywhere and natural language is a way of communication understandable to humans. Visual Question Answering (VQA) is a system which takes image as an input and a question about the image and generates a natural language answer using complex reasoning. Thus, a VQA needs detailed understanding of the image and complex reason to predict the answer. Given its multimodal structure and possible real-world implementations, VQA is a challenge of critical importance for artificial intelligence. The architectures and hyperparameters used in deep neural networks for VQA have a big impact on their results. This project introduces a pretrained model (VGGNet) to extract image features and Word2Vec to embed the words and LSTM to get word features from the question and after combining the results will predict the answer having highest probability.
한국어
시각적 데이터는 도처에 존재하며 자연어는 인간이 이해할 수 있는 의사소통 수단이다. VQA(Visual Question Answering)는 이미지를 이미지에 대한 입력과 질문으로 취하고 복잡한 추론을 사용하여 자연어 답변을 생성 하는 시스템이다. 따라서, VQA는 답을 예측하기 위해 이미지에 대한 자세한 이해와 복잡한 이유가 필요하 다. 멀티모달 구조와 가능한 실제 구현을 고려할 때, VQA는 인공지능에게 매우 중요한 과제이다. VQA를 위한 심층 신경망에 사용되는 아키텍처와 하이퍼 파라미터는 결과에 큰 영향을 미친다. 이 프로젝트는 이미 지 특징을 추출하기 위해 사전 훈련된 모델(VGGNet)과 단어를 내장하기 위해 Word2Vec를 도입하고 질문에 서 단어 특징을 얻기 위해 LSTM을 도입하고 결과를 결합한 후 가장 높은 확률을 가진 답을 예측한다.

목차

ABSTRACT
1. Introduction
2. Literature Review
3. Research Methodology
3.1 Visual Question Answering Model Prototype
3.2 The Proposed Model
3.3 Image pre-processing for feature engineering
3.4 Understanding type of Questions
3.5 Process Diagram – Proposed Model
3.6 How does the probabilistic model work?
4. Analysis and Evaluation
4.1 Hyperparameter Tuning
4.2 Model Layers
4.3 Results
5. Conclusion
References
<국문초록>
<결론 및 향후 연구>

저자

  • Manva Trivedi [ Graduate Student, Computer Science Department, Lakehead University, Canada ] Corresponding Author
  • Sabah Mohammed [ Professor, Computer Science Department, Lakehead University, Canada ]

참고문헌

자료제공 : 네이버학술정보

간행물 정보

발행기관

  • 발행기관명
    한국컴퓨터게임학회 [Korean Society for Computer Game]
  • 설립연도
    2002
  • 분야
    공학>컴퓨터학
  • 소개
    1. 게임산업을 활성화 하고, 2. 게임기술과 기술 인력을 양산할 수 있도록 교육기관의 교과과정을 개발하고, 3. 관련기술에 대한 연구발표회, 강연회, 강습회 등을 개최하며, 4. 학회지, 논문지 및 관련 문헌을 발간하고, 5. 게임 기술 개발을 위한 국제화, 표준화 등을 지원하고, 6. 산.학.연.관이 협동할 수 있는 국제적 학술교류 및 협력을 지원하고, 7. 회원 상호간의 공동 이익과 친목을 증진시킨다.

간행물

  • 간행물명
    컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) [Journal of Computer Games and Contents]
  • 간기
    월간
  • pISSN
    3091-7409
  • eISSN
    3092-3638
  • 수록기간
    2002~2026
  • 등재여부
    KCI 등재
  • 십진분류
    KDC 691 DDC 793

이 권호 내 다른 논문 / 컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) 제35권 제3호

    피인용수 : 0(자료제공 : 네이버학술정보)

    함께 이용한 논문 이 논문을 다운로드한 분들이 이용한 다른 논문입니다.

      페이지 저장