Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 7
No
1

4,000원

본 연구는 기업의 AI Transformation(AX)을 지원하기 위해, Vision Language Model(VLM) 기반 지능 형 문서처리 플랫폼을 설계하고, Qwen2.5VL-7B를 활용한 영수증 처리 프로토타입을 구현하였다. 제안된 플랫폼 은 3-Tier 마이크로서비스 아키텍처를 기반으로, 프롬프트 관리 체계와 기능별 모듈화를 통해 유연하고 확장 가능 한 구조를 구현하였다. 실험 결과, 평균 91.7%의 정보 추출 정확도를 달성하였으며, 사전 템플릿 없이 다양한 문서 형식에 대응 가능한 처리 유연성을 바탕으로 실무 적용 가능성을 입증하였다. 본 연구는 OCR 중심 기술의 한계를 보완하는 프롬프트 기반 VLM 아키텍처를 실증적으로 제시하고, 금융·물류·의료 등 산업 전반에서 적용 가능한 문 서 자동화 기반을 제공하였다는 점에서 학문적·실무적 의의를 갖는다.

This study supports corporate AI Transformation (AX) by designing a document processing platform based on a Vision Language Model (VLM) and implementing a prototype using Qwen2.5VL-7B. The platform employs a three-tier microservice architecture with prompt management and modular components to ensure flexibility and scalability. Experiments showed an average information extraction accuracy of 91.7%, and the system demonstrated practical applicability by handling diverse document formats without predefined templates. This research provides an empirical implementation of a prompt-based VLM architecture that overcomes limitations of OCR technologies, offering academic and practical value as a foundation for document automation across sectors such as finance, logistics, and healthcare.

2

Out-of-Stock (OOS) detection is a critical task in retail management, yet traditional automated methods suffer from significant limitations. Existing approaches, whether based on classical image processing or closed-set deep learning models, lack scalability and robustness, often requiring extensive retraining for every new shelf layout or product type. This paper proposes a novel and flexible OOS detection pipeline that leverages the power of Open- Vocabulary Object Detection (OVD). Instead of attempting to directly detect "empty space," our method first identifies all present products using an OVD model guided by flexible text prompts. An inverse occupancy mask is then generated to identify potential OOS regions, which are subsequently refined through a robust multi-stage post-processing filter. Experiments on our custom dataset of 149 diverse retail shelf images demonstrate the superiority of this approach. Our method achieves an OOS detection Accuracy of 87.9%, vastly outperforming the baseline approach (70.1% Accuracy). Furthermore, by applying per-shelf optimized prompts and parameters, our model's Accuracy increases to 96.84%, highlighting its high adaptability and effectiveness for realworld retail environments.

3

4,000원

Modern enterprises maintain extensive repositories of business documents within their intranet systems, creating a critical need for automated processing capabilities of image-based documents to enhance operational efficiency. Unlike standardized forms, most business documents are semi-structured, with layouts and field positions varying widely across organizations and document types. This complexity has generated substantial demand for advanced information extraction and organization technologies, capable of handling irregular structures and diverse schemas. However, conventional Optical Character Recognition (OCR) approaches, which prioritize textual recognition, encounter significant limitations when processing complex forms due to their reliance on location-based extraction. Similarly, Key Information Extraction (KIE) techniques often require domain-specific pre-training, resulting in considerable learning and adapting costs for novel document formats. To address these challenges, this study proposes an innovative process for effectively extracting and organizing key elements from semi-structured documents by employing Visual Language Models (VLMs) that process documents as image inputs and concurrently analyze visual and linguistic information. The proposed framework determines superior extraction accuracy, economic efficiency, and even user satisfaction by exploiting both semantic textual content and spatial positioning as visual cues. Experimental results demonstrate that the VLM-based framework outperforms existing OCR and KIE solutions across multiple evaluation dimensions, while the integration of human-in-the-loop verification processes establishes a practical framework for semi-structured document automation (e.g., commercial invoice) with immediate applicability in fast-changing enterprise environments.

4

VLM을 이용한 해양로봇 자율주행 시스템 최적화 연구

김민정, 김태연, 김병모, 최원석

[Kisti 연계] 한국로봇학회 로봇학회논문지 Vol.21 No.2 2026 pp.140-147

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In the recent developments, contrary to previous question-answering systems, VLMs (Vision-Language Models) are capable to analyze the semantic meaning with visual interpretation simutanously. Such semantic capabilities of the VLM has been exploited in the autonomous control systems. Requirements of techniques for applications are challenging due to high versatility of the scenes at marine environments. Compared to static scenarios, in ocean, the surface changes continuously due to waves effected by currents and winds. Therefore, the difficulty of localizations and eliminations of noise from sensory data causes wrong decision makings for autonomous systems. Controls based on those inaccurate estimates of the position and status of the maritine robots converge to incorrect navigations. The commonly practiced procedures of the autonomous sytem includes visual input, semantic analysis using VLM, decision inferences, and motor command excutions. In this paper, to overcome challenges of the marine environments, improvement of autonomous control performance by integrating perception and decision making procedures are developed. Therefore, the time delay and information loss have been decreased. The new structure has shown enhancement on reactivity and reliability of the decision makings in maritime navigations. In simulation, a comparative analysis was carried out based on ROS (Robot Operating System) and Gazebo between a non-integrated model (with separate image interpretation and generation of the commands) and the integrated one. The integrated system also had a faster response time and higher decision reliability across trials. However, as the number of commands increased, consistency decreased, resulting in a slight increase in the total driving distance. Nevertheless, the performance was smoother and more robust than that of the conventional system overall. Important items for future work are, for example, better stability behavior with frequent command updates and reduced computational demands that guarantee that VLM-driven systems can be used for reliable real-time navigation in challenging maritime environments.

5

작물 수확 자동화를 위한 시각 언어 모델 기반의 환경적응형 과수 검출 기술

남창우, 송지민, 진용식, 이상준

[Kisti 연계] 대한임베디드공학회 대한임베디드공학회논문지 Vol.19 No.2 2024 pp.73-81

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Recently, mobile manipulators have been utilized in agriculture industry for weed removal and harvest automation. This paper proposes a domain adaptive fruit detection method for harvest automation, by utilizing OWL-ViT model which is an open-vocabulary object detection model. The vision-language model can detect objects based on text prompt, and therefore, it can be extended to detect objects of undefined categories. In the development of deep learning models for real-world problems, constructing a large-scale labeled dataset is a time-consuming task and heavily relies on human effort. To reduce the labor-intensive workload, we utilized a large-scale public dataset as a source domain data and employed a domain adaptation method. Adversarial learning was conducted between a domain discriminator and feature extractor to reduce the gap between the distribution of feature vectors from the source domain and our target domain data. We collected a target domain dataset in a real-like environment and conducted experiments to demonstrate the effectiveness of the proposed method. In experiments, the domain adaptation method improved the AP50 metric from 38.88% to 78.59% for detecting objects within the range of 2m, and we achieved 81.7% of manipulation success rate.

6

비전-언어 모델 기반 이미지 독립적 시각 의미 인증(VSA) 프레임워크의 설계 및 보안성 실증

정필성

[Kisti 연계] 한국정보보호학회 정보보호학회논문지 Vol.36 No.3 2026 pp.935-948

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

기존 이미지 기반 그래픽 패스워드는 특정 이미지에 종속되어 이미지 유출 및 숄더서핑(shoulder surfing) 공격에 취약하다. 본 논문은 이미지 자체가 아닌 이미지 속 '시각적 의미(visual semantic)'를 비밀키로 활용하는 이미지 독립적 시각 의미 인증(Visual Semantic Authentication, VSA) 프레임워크를 제안한다. Grounding DINO와 CLIP을 결합하여 객체, 속성, 수량, 2×2 4분면 공간 정보를 추출하고, 유연한 범위 논리(FlexibleRange Logic)와 해시 기반 비밀 바인딩(hash-based secret binding)으로 보안성과 가용성을 동시에 확보한다. COCO 2017 벤치마크 1,009장에 대해 5단계 난이도 16개 정책을 평가한 결과, 규칙 복잡도에 따라 FAR이 51.73%에서 0%까지 단조 감소하는 보안 기울기(security gradient)를 확인하였다. 또한 벤치마크에서 정책별 양성 이미지를 자동 식별하여 FRR을 평가하고, 정책 복잡도와 이미지 가용성 간의 보안-사용성 트레이드오프를 정량적으로 제시하였다.

Conventional graphical passwords are tied to specific images, making them vulnerable to image leakage and shoulder-surfing attacks. This paper proposes an Image-Agnostic Visual Semantic Authentication (VSA) framework that uses visual semantics within images as secret keys. The framework combines Grounding DINO and CLIP to extract objects, attributes, quantities, and 2×2 quadrant spatial information, while Flexible Range Logic and hash-based secret binding ensure both usability and server-opaque protection. Experiments on 1,009 COCO 2017 benchmark images across 16 security policies in five difficulty levels demonstrate a monotonic FAR gradient from 51.73% to 0% with increasing rule complexity. FRR is also evaluated by automatically identifying positive images from the benchmark, and the trade-off between policy complexity and image availability is quantitatively characterized.

7

LVLN: 시각-언어 이동을 위한 랜드마크 기반의 심층 신경망 모델

황지수, 김인철

[Kisti 연계] 한국정보처리학회 정보처리학회논문지/소프트웨어 및 데이터 공학 Vol.8 No.9 2019 pp.379-390

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 시각-언어 이동 문제를 위한 새로운 심층 신경망 모델인 LVLN을 제안한다. LVLN 모델에서는 자연어 지시의 언어적 특징과 입력 영상 전체의 시각적 특징들 외에, 자연어 지시에서 언급하는 주요 장소와 랜드마크 물체들을 입력 영상에서 탐지해내고 이 정보들을 추가적으로 이용한다. 또한 이 모델은 자연어 지시 내 각 개체와 영상 내 각 관심 영역, 그리고 영상에서 탐지된 개별 물체 및 장소 간의 서로 연관성을 높일 수 있도록 맥락 정보 기반의 주의 집중 메커니즘을 이용한다. 그뿐만 아니라, LVLN 모델은 에이전트의 목표 도달 성공율을 향상시키기 위해, 목표를 향한 실질적인 접근을 점검할 수 있는 진척 점검기 모듈도 포함하고 있다. Matterport3D 시뮬레이터와 Room-to-Room (R2R) 벤치마크 데이터 집합을 이용한 다양한 실험들을 통해, 본 논문에서 제안하는 LVLN 모델의 높은 성능을 확인할 수 있었다.

In this paper, we propose a novel deep neural network model for Vision-and-Language Navigation (VLN) named LVLN (Landmark-based VLN). In addition to both visual features extracted from input images and linguistic features extracted from the natural language instructions, this model makes use of information about places and landmark objects detected from images. The model also applies a context-based attention mechanism in order to associate each entity mentioned in the instruction, the corresponding region of interest (ROI) in the image, and the corresponding place and landmark object detected from the image with each other. Moreover, in order to improve the success rate of arriving the target goal, the model adopts a progress monitor module for checking substantial approach to the target goal. Conducting experiments with the Matterport3D simulator and the Room-to-Room (R2R) benchmark dataset, we demonstrate high performance of the proposed model.

 
페이지 저장