년 - 년
멀티모달 정보와 계층 구조를 반영한 장면 인지 기반의 영상 요약
한국경영정보학회 한국경영정보학회 정기 학술대회 Generative AI and the Next Computing Revolution : From Automation to Creative Disruption 2025.05 pp.234-237
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
2024년 기준 전 세계 인터넷 사용자 중 92%가 매월 온라인 비디오를 시청하며, 주간 평균 시청 시간은 17시간에 달한다. 이처럼 영상 콘텐츠의 양이 기하급수적으로 증가하면서, 정보 과잉 속에서 핵심 내용을 빠르게 파악하기 어려워지고 있다. 이에 따라 영상 요약 기술의 필요성이 더욱 강조되고 있다. 기존 영상 요약 기법은 주로 프레임 단위 중요도 예측에 집중하지만, 영상의 시간적 구조나 의미 있는 사건을 충분히 반영하지 못하는 한계가 있다. 실제 영상은 프레임(frame), 샷(shot), 장면(scene), 시퀀스(sequence)로 이어지는 복합적 서사 구조를 가지므로, 시간적 흐름이나 사적 구조를 가지며, 이 구조를 고려한 요약은 중복을 줄이고 맥락을 보존하는데 중요하다. 본 연구는 샷과 장면 경계, 멀티모달 이벤트를 탐지하고 어텐션 기반으로 서사 흐름을 반영하는 장면 인지 기반 요약 프레임워크를 제안한다.
다중 뷰 비디오로부터 두드러진 정보 추출은 인터뷰, 인트라 뷰간 상관관계와 계산 비용 때문에 매우 어려운 영역입 니다. 매우 높은 계산 복잡성을 지닌 멀티 뷰 비디오에서 키프레임을 추출하기 위해 개발된 몇 가지 기술이 있습니 다. 이 논문에서, 우리는 내부에 존재하는 엔트로피와 복잡한 정보를 사용하여 멀티 뷰 비디오의 키프레임 추출 접 근 방식을 제시합니다. 첫 번째 단계에서는 프레임 사이의 SSIM값을 기반으로 각 보기에서 전체 비디오의 대표 샷 을 추출합니다. 두 번째 단계에서는 서로 다른 보기의 모든 샷 프레임에 대한 엔트로피와 복잡성 점수가 계산됩니 다. 마지막으로 엔트로피와 복잡성 점수가 가장 높은 프레임은 키 프레임으로 간주됩니다. 제안된 시스템은 사용 가 능한 Office벤치마크 데이터 세에서 주관적으로 평가되며, 정확성과 시간 복잡성의 측면에서 결과는 편리합니다.
Salient information extraction from multi-view videos is a very challenging area because of interview, intra-view correlations, and computational complexity. There are several techniques developed for keyframes extraction from multi-view videos with very high computational complexities. In this paper, we present a keyframes extraction approach from multi-view videos using entropy and complexity information present inside frame. In first step, we extract representative shots of the whole video from each view based on structural similarity index measurement (SSIM) difference value between frames. In second step, entropy and complexity scores for all frames of shots in different views are computed. Finally, the frames with highest entropy and complexity scores are considered as keyframes. The proposed system is subjectively evaluated on available office benchmark dataset and the results are convenient in terms of accuracy and time complexity.
Viewer's Affective Feedback for Video Summarization
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.11 No.1 2015 pp.76-94
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
For different reasons, many viewers like to watch a summary of films without having to waste their time. Traditionally, video film was analyzed manually to provide a summary of it, but this costs an important amount of work time. Therefore, it has become urgent to propose a tool for the automatic video summarization job. The automatic video summarization aims at extracting all of the important moments in which viewers might be interested. All summarization criteria can differ from one video to another. This paper presents how the emotional dimensions issued from real viewers can be used as an important input for computing which part is the most interesting in the total time of a film. Our results, which are based on lab experiments that were carried out, are significant and promising.
[Kisti 연계] 한국문헌정보학회 한국문헌정보학회지 Vol.56 No.1 2022 pp.95-117
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 시선 및 뇌파 정보를 이용하여 오디오-비주얼(audio-visual, AV) 시맨틱스 기반의 동영상 요약 방법들을 개발하고 평가해 보았다. 이를 위해서 27명의 대학생들을 대상으로 시선추적과 뇌파 실험을 수행하였다. 평가 결과, 뇌파와 동공크기 데이터를 함께 사용한 방법의 평균 재현율(0.73)이 뇌파 또는 동공크기 데이터만을 사용한 방법의 평균 재현율(뇌파: 0.50, 동공크기: 0.68)보다 높게 나타났다. 또한 AV 시맨틱스 기반의 개인화된 동영상 요약의 평균 재현율(0.57)이 AV 시맨틱스 기반의 일반적인 동영상 요약의 평균 재현율(0.69)보다 낮게 나타난 원인들을 분석하였다. 끝으로, AV 시맨틱스 기반 동영상 요약 방법과 텍스트 시맨틱스 기반 동영상 요약 방법 간의 차이 및 특성도 비교분석해 보았다.
This study developed and evaluated audio-visual (AV) semantics-based video summarization methods using eye tracking and electroencephalography (EEG) data. For this study, twenty-seven university students participated in eye tracking and EEG experiments. The evaluation results showed that the average recall rate (0.73) of using both EEG and pupil diameter data for the construction of a video summary was higher than that (0.50) of using EEG data or that (0.68) of using pupil diameter data. In addition, this study reported that the reasons why the average recall (0.57) of the AV semantics-based personalized video summaries was lower than that (0.69) of the AV semantics-based generic video summaries. The differences and characteristics between the AV semantics-based video summarization methods and the text semantics-based video summarization methods were compared and analyzed.
[Kisti 연계] 한국문헌정보학회 한국문헌정보학회지 Vol.52 No.4 2018 pp.91-110
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 비디오 스킴의 자동 생성을 위한 비디오 요약 알고리즘을 제안하고 이를 평가하였다. 제안된 알고리즘은 ERP(Event Related Potentials) 기반의 주제 적합성 모형, MMR(Maximal Marginal Relevance) 기법 및 판별분석기법을 사용하여 구현하였다. 제안한 ERP/MMR 기반 알고리즘을 이용하여 구성한 비디오 스킴의 품질과 유용성을 내재적 및 외재적 평가를 통해서 검증하였다. 내재적 및 외재적 평가에서 ERP/MMR 방법들의 평가 점수들은 각각 경쟁 기준으로 사용한 SBD(Shot Boundary Detection) 방법의 평가 점수 보다 유의미한 차이를 보이며 높게 나왔다. 그러나 이 두 평가에서 ERP/MMR(${\lambda}=0.6$) 방법의 평가 점수와 ERP/MMR(${\lambda}=1.0$) 방법의 평가 점수 간에 통계적으로 유의미한 차이는 없는 것으로 나타났다.
We proposed a video summarization algorithm based on an ERP (Event Related Potentials)-based topic relevance model, a MMR (Maximal Marginal Relevance), and discriminant analysis to generate a semantically meaningful video skim. We then conducted implicit and explicit evaluations to evaluate our proposed ERP/MMR-based method. The results showed that in the implicit and explicit evaluations, the average scores of the ERP / MMR methods were statistically higher than the average score of the SBD (Shot Boundary Detection) method used as a competitive baseline, respectively. However, there was no statistically significant difference between the average score of ERP/MMR (${\lambda}=0.6$) method and that of ERP/MMR (${\lambda}=1.0$) method in both assessments.
Video summarization with adjustable ratio of highlights and story
국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.17 No.4 2025.11 pp.53-66
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Most current video summarization techniques focus on converting long videos into short summary clips by detecting semantically important parts of the original video. However, many summarization applications, such as personal summaries of weddings, travel videos, movie trailers, or short summaries for search, require summaries that include the entire content of the original video, not just the important parts. This paper proposes a technique for generating summary videos by combining highlight scenes representing important parts of the video with story-conveying scenes capturing the overall flow of the video. To achieve this objective, we introduce the concept of 'diversity contribution,' which numerically quantifies how much each individual scene contributes to creating a diverse scene composition in the summary video. The higher the proportion of scenes with high diversity contribution values, the more comprehensively the summary video incorporates the entire content of the original video. In this study, we developed an algorithm to create a summary video by adjusting the diversity contribution and importance of scenes in proportion, and verified its operation by implementing an actual summary application system. Furthermore, we confirmed through experiments that the accuracy of the summary technique presented in this paper was over 90%.
Event Detection Based Approach for Soccer Video Summarization Using Machine learning SCOPUS
보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.7 No2 2012.05 pp.63-80
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Many soccer fans prefer to watch a summary of football games as watching a whole soccer match needs a lot of time. Traditionally, soccer videos were analyzed manually, however this costs valuable time. Therefore, it is necessary to have a tool for doing the video anal- ysis and summarization job automatically. Automatic soccer video summarization is about extracting important events from soccer matches in order to produce general summaries for the most important moments in which soccer viewers may be interested. This paper presents a machine learning (ML) based event detection and summarization system for em- phasizing important events during soccer matches. The proposed system rstly segments the whole video stream into small video shots, then it classies the resulted shots into dierent shot-type classes. Afterwards, the system applies two machine learning algorithms, namely; support vector machine (SVM) and articial neural network (ANN), for emphasizing impor- tant segments with logo appearance with addition to detecting the caption region providing information about the score of the game. Subsequently, the system detects vertical goal posts and goal net. Finally, the most important events during the match are highlighted in the resulted soccer video summary. Experiments on real soccer videos demonstrate encourag- ing results. The proposed approach greatly reduces workload and enhances the accuracy of summarizing soccer video matches with reference to both recall and precision performance measurement criteria.
보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.9 No.11 2014.11 pp.397-408
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Massive Open Online Courses (MOOC) platforms provide a rich environment for knowledge creation through its massiveness and inherited collaborative tools. However, it also restricts spontaneous knowledge sharing by the existing LMS barriers between the main multimedia content and the collaborative tools. None the less, the collaboration still massive due to the number of participants. The separation of the multimedia content and the discussion tools is the first focus point of this paper. Moreover, this article is presenting a new added value to the MOOC architecture so to link the learner’s discussions and its summary with the multimedia contents. The added-value component involves a summarization algorithm that summarizes the shared collaborative textual discussion collected from the various learners viewing relevant MOOC multimedia/video contents. The affectivity of the summarization component was tested using the popular ROUGE software package from University of Southern California. The new MOOC architecture represents an enhanced learning environment that enables learners to share the multimedia information along with its annotated collaborative information with the power of summarizing the final outcome of the presented annotations relevant to a specific shared multimedia content.
[Kisti 연계] 한국컴퓨터정보학회 한국컴퓨터정보학회 학술대회논문집 2023 pp.701-702
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
비대면 교육이 증가함에 따라 강의, 특강과 같은 정보성 동영상의 수가 급격히 많아지고 있다. 이러한 정보성 동영상을 보아야 하는 학습자들은 자원과 시간을 효율적으로 활용할 수 있는 동영상 이해 및 학습 시스템이 필요하다. 본 논문에서는 GPT-3 모델과 KoNLPy 사용하여 동영상 요약을 수행하고 키워드 기반 해당 영상 프레임으로 바로 갈 수 있는 시스템의 개발내용에 대해 기술한다. 이를 통해 동영상 콘텐츠를 효과적으로 활용하여 학습자들의 학습 효율성을 향상시킬 수 있을 것으로 기대한다.
동영상 실시간 시청시 유발전위(ERP) N400 속성을 이용한 주제무관 쇼트 선별 자동영상요약 연구
[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.20 No.8 2017 pp.1258-1270
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
'Semantic gap' has been a year-old problem in automatic video summarization, which refers to the gap between semantics implied in video summarization algorithms and what people actually infer from watching videos. Using the external EEG bio-feedback obtained from video watchers as a solution of this semantic gap problem has several another issues: First, how to define and measure noises against ERP waveforms as signals. Second, whether individual differences among subjects in terms of noise and SNR for conventional ERP studies using still images captured from videos are the same with those differently conceptualized and measured from videos. Third, whether individual differences of subjects by noise and SNR levels help to detect topic-irrelevant shots as signals which are not matched with subject's own semantic topical expectations (mis-match negativity at around 400m after stimulus on-sets). The result of repeated measures ANOVA test clearly shows a 2-way interaction effect between topic-relevance and noise level, implying that subjects of low noise level for video watching session are sensitive to topic-irrelevant visual shots, while showing another 3-way interaction among topic-relevance, noise and SNR levels, implying that subjects of high noise level are sensitive to topic-irrelevant visual shots only if they are of low SNR level.
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2023 pp.694-695
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 ChatGPT 를 각 분야에 활용하는 연구가 활발하게 이루어지고 있다. ChatGPT 는 최신 자연어 처리 모델로, 텍스트를 통해 입출력을 진행한다. 본 논문에서는 이러한 ChatGPT 를 활용하여 영상을 효과적으로 요약할 수 있는 새로운 접근 방식을 제시한다. STT 기술을 사용하여 영상의 자막에 대한 텍스트 파일을 추출하고 이를 ChatGPT 로 요약한다. 최종적으로 기존 텍스트와의 유사도 분석을 통해 유사도가 높은 부분을 선택하여 영상을 편집하고 요약한다.
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2006 pp.44-48
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
디지털 카메라 기술의 발전과 보급으로 공공건물의 보안 카메라부터 개인 휴대 단말기의 카메라까지 동영상 데이터를 수집할 수 있는 수단이 크게 늘었으며, 그 활용 또한 매우 일반화되었다. 동영상 데이터는 문서나 음성 등의 다른 데이터보다 훨씬 구체적이고 사실적인 정보를 포함하므로 과거의 기억을 정리하고 복원하기 위한 유용한 방법이 될 수 있다. 동영상 데이터의 증가와 함께 동영상 요약에 대한 연구가 최근에 활발히 진행되고 있는데, 이들 연구의 대부분은 하나의 동영상을 요약하고 분석하기 위한 것이다. 본 논문에서는 사무실에 여러 대의 카메라를 설치하여 데이터를 저장하며, 이렇게 수집된 동영상 데이터를 효과적으로 요약하고 검색하는 시스템을 구축한다. 동일한 이벤트를 여러 방향에서 바라보고, 그 상황을 가장 잘 설명한 카메라를 선택 할 수 있다는 점에서 멀티 카메라의 사용은 장점을 갖는다. 사전에 정의된 이벤트에 따라 전문가가 어노테이션을 부여하도록 하였으며, 전문가가 설정한 유틸리티에 따라 카메라 선택 및 요약이 이루어진다. 다양한 옵션에 따라 요약된 결과로 사용자 평가를 수행하였다.
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2003 pp.199-201
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
MPEG-21 환경에서의 DI(Digital Item)은 MPEG-21 프레임워크 내에서 표준화된 표현 형식, 식별 체계, 서술 형식을 따르는 구조화된 디지털 객체이며, 유통, 처리의 최소 단위이다 따라서. 이러한 DI가 MPEG-21 멀티미디어 프레임워크 환경에서 사용자 터미널에 전달되었을 때 어떻게 처리되어야 될 것인지를 규정하는 것은 매우 중요한 과제이며. 이와 관련한 기술이 DIP(Digital Item Processing)이다. 본 논문에서는 DIP의 한 응용 예로서 멀티미디어 콘텐츠를 계층적으로 기술하는 Video Summary의 응용 방안에 대한 연구 결과를 제시하고자 한다.
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2002 pp.7-10
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 MPEG-7의 활동도 기술자를 이용한 비디오 기술을 제안한다. 제안한 방법은 압축상태의 비디오 자료에서 직접 움직임 벡터들을 추출, 각 프레임들의 활동도의 강도를 계산하고 프레임의 흐름에 따라 계산된 활동도의 변화량에 대해 퓨리에 변환을 적용하여 얻어진 주파수 성분을 분석하여 활동도의 시간적 분포도를 계산한다. 계산된 강도 및 분포도는 MPEG-7의 표준에 따르기 위해 양자화하여 비디오 요약에 이용한다.
실시간 상황 인식을 위한 센서 운용 모드 기반 항공 영상 요약 기법
[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.20 No.6 2015 pp.87-97
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
항공 영상 요약은 무인항공기를 통해 획득된 전체 영상의 내용을 제한된 시간 내에 효과적으로 브라우징 함으로써 감시 정찰 지역에 대한 상황 인식을 가능하게 하는 기술이다. 항공 영상의 정확한 요약을 수행하기 위해 본 논문에서는 센서 운용 모드를 집중감시, 전역감시 그리고 구역감시모드로 구분하고 해당 센서 운용 모드의 특성을 고려하여 항공 영상 요약을 수행한다. 특히 집중감시 모드에서의 영상 요약은 화면 내 움직임이 있는 관심 객체의 지속적인 추적을 기반으로 수행되며 이를 위해 본 논문에서는 지역 움직임 벡터(partitioning motion vector)와 해당 벡터가 발생한 영역에서의 시공간적 중요도 지도(spatiotemporal saliency map)를 활용한 움직임 반응 추적 기법을 제안한다. 제안하는 알고리즘의 효율성과 적합성을 확인하기 위해 실 항공 영상을 대상으로 실험을 수행하였다. 도출된 실험 결과를 통해 제안하는 방법은 전체 항공 영상에서의 영상 요약을 위해 센서 운용 모드에 따라 정확한 대표 프레임을 검출하였으며 이에 따라 대용량의 무인항공기 획득 영상이 효과적으로 요약될 수 있음을 확인하였다.
An Aerial video summarization is not only the key to effective browsing video within a limited time, but also an embedded cue to efficiently congregative situation awareness acquired by unmanned aerial vehicle. Different with previous works, we utilize sensor operation mode of unmanned aerial vehicle, which is global, local, and focused surveillance mode in order for accurately summarizing the aerial video considering flight and surveillance/reconnaissance environments. In focused mode, we propose the moving-react tracking method which utilizes the partitioning motion vector and spatiotemporal saliency map to detect and track the interest moving object continuously. In our simulation result, the key frames are correctly detected for aerial video summarization according to the sensor operation mode of aerial vehicle and finally, we verify the efficiency of video summarization using the proposed mothed.
[Kisti 연계] 한국멀티미디어학회 한국멀티미디어학회 학술대회논문집 2002 pp.77-80
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
비디오 데이터에서 캡션은 비디오의 중요한 부분과 내용을 나타내는 가장 보편적인 방법이다. 본 논문에서는 축구 비디오에서 캡션이 갖는 특징을 분석하고 캡션에 의한 키 프레임을 추출하도록 하며, 비디오 요약 생성 규칙에 따라 요약된 비디오를 생성하도록 한다. 키 프레임 추출은 이벤트 발생에 따른 캡션의 등장과 캡션 내용의 변화를 추출하는 것으로 탬플리트 매칭과 지역적 차영상을 통하여 추출하며 샷의 재설정 통하여 중요한 이벤트를 포함한 요약된 비디오를 생성하도록 한다.
[Kisti 연계] 한국멀티미디어학회 한국멀티미디어학회 학술대회논문집 2001 pp.245-248
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
비디오 데이터에서 캡션은 비디오의 중요한 부분과 내용을 나타내는 가장 보편적이 방법이다. 본 논문에서는 축구 비디오에서 캡션이 갖는 특징을 분석하고 캡션에 의한 키 프레임을 추출하도록 하며, 비디오 요약 생성 규칙에 따라 요약된 비디오를 생성하도록 한다. 키 프레임 추출은 이벤트 발생에 따른 캡션의 등장과 캡션 내용의 변화를 추출하는 것으로 탬플리트 매칭과 지역적 차영상을 통하여 추출하며 샷의 재설정 통하여 중요한 이벤트를 포함한 요약된 비디오를 생성하도록 한다.
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2006 pp.544-548
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 광범위한 지역을 감시하기 위해 설치된 여러 대의 카메라로부터 획득된 비디오에 대해 물체를 기반으로 한 비디오 요약 시스템을 제안한다. 제안된 시스템은 시야가 겹쳐지지 않은 다수의 CCTV 카메라를 통해서 촬영한 비디오들을 30분 단위로 나누어 비디오 데이터베이스를 구축하고 시간별, 카메라별 비디오 검색이 가능하다. 비디오에서 물체기반 키프레임을 추출하여 카메라별, 사람별로 비디오를 요약할 수 있도록 하였다. 또한 임계치에 따라 키프레임 검색정도를 조절함으로써 비디오 요약정도를 조절할 수 있다. 이렇게 검색된 키프레임에 대한 카메라별, 시간별 통계를 통해서 감시지역의 물체기반 이벤트를 간단히 확인해 볼 수 있다.
[Kisti 연계] 한국신호처리시스템학회 한국신호처리.시스템학회 논문지 Vol.6 No.4 2005 pp.163-168
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 인물 기반의 비디오 요약 방법으로써 비디오 내 음성정보를 이용하여 화자 인식 기법을 통한 등장인물 중심의 요약 기법을 제안한다. 먼저, 얼굴 영역을 포함하는 장면을 중심으로 비디오로부터 배우의 대사에 해당하는 음성 정보를 분리하고, 화자 인식 기법을 수행하여 등장인물 별로 분류하였다. 화자인식 기법은 각 화자별로 MFCC(Mel Frequency Cepstrum Coefficient) 값을 추출하고 GMM(Gaussian Mixture Model)을 이용하여 분류한다. 본 논문에서는 4명의 등장인물에 대해 GMM을 학습시키고 4명 중 1명을 검출하는 실험을 통해 학습된 GMM 분류기가 실험 비디오에 대해 0.138 정도의 오분류율을 보임을 확인하였다.
In this paper, we propose a character-based summarization algorithm using speaker identification method from the dialog in video. First, we extract the dialog of shots containing characters' face and then, classify the scene according to actor/actress by performing speaker identification. The classifier is based on the GMM(Gaussian Mixture Model) using the 24 values of MFCC(Mel Frequency Cepstrum Coefficient). GMM is trained to recognize one actor/actress among four who are all trained by GMM. Our experiment result shows that GMM classifier obtains the error rate of 0.138 from our video data.
[NRF 연계] 한국융합신호처리학회 융합신호처리학회 논문지 Vol.6 No.4 2005.10 pp.163-168
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 인물 기반의 비디오 요약 방법으로써 비디오 내 음성정보를 이용하여 화자 인식 기법을 통한 등장인물 중심의 요약 기법을 제안한다. 먼저, 얼굴 영역을 포함하는 장면을 중심으로 비디오로부터 배우의 대사에 해당하는 음성 정보를 분리하고, 화자 인식 기법을 수행하여 등장인물 별로 분류하였다. 화자인식 기법은 각 화자별로 MFCC(Mel Frequency Cepstrum Coefficient) 값을 추출하고 GMM(Gaussian Mixture Model)을 이용하여 분류한다. 본 논문에서는 4명의 등장인물에 대해 GMM을 학습시키고 4명 중 1명을 검출하는 실험을 통해 학습된 GMM 분류기가 실험 비디오에 대해 0.138 정도의 오분류율을 보임을 확인하였다.
In this paper, we propose a character-based summarization algorithm using speaker identification method from the dialog in video. First, we extract the dialog of shots containing characters' face and then, classify the scene according to actor/actress by performing speaker identification. The classifier is based on the GMM(Gaussian Mixture Model) using the 24 values of MFCC(Mel Frequency Cepstrum Coefficient). GMM is trained to recognize one actor/actress among four who are all trained by GMM. Our experiment result shows that GMM classifier obtains the error rate of 0.138 from our video data.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.