Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 23
No
1

최근 딥러닝 기술이 발전함에 따라 많은 컴퓨터 비전 연구 분야에서 주목할 만한 성과들이 지속적으로 나오고 있다. 단일 이미지를 기반으로 사람의 2차원 및 3차원 포즈를 추정하는 연구에서도 비약적인 성능향상을 보여주고 있으 며, 많은 연구자들이 문제의 범위를 확장하며 활발한 연구 활동을 진행하고 있다. 사람의 포즈 추정은 다양한 응용 분야가 존재하고, 특히 이미지나 비디오 분석에서 사람의 포즈는 행동 및 상태, 의도 파악을 위한 핵심 요소가 되기 때문에 상당히 중요한 연구 분야이다. 이러한 배경에 따라 본 논문은 단일 이미지를 기반으로 한 사람의 포즈 추정 기술에 대한 연구 동향을 살펴보고자 한다. 강인하고 정확한 문제 해결을 위해 다양한 연구 활동 결과가 존재한다는 점에서 본 논문에서는 사람의 포즈 추정 연구를 2차원 및 3차원 포즈 추정에 대해서 나누어 살펴보고자 한다. 끝으 로 연구에 필요한 데이터 세트 및 사람의 포즈 추정 기술을 적용하는 다양한 연구 사례를 살펴볼 것이다.

With the recent development of deep learning technology, remarkable achievements have been made in many research areas of computer vision. Deep learning has also made dramatic improvement in two-dimensional or three-dimensional human pose estimation based on a single image, and many researchers have been expanding the scope of this problem. The human pose estimation is one of the most important research fields because there are various applications, especially it is a key factor in understanding the behavior, state, and intention of people in image or video analysis. Based on this background, this paper surveys research trends in estimating human poses based on a single image. Because there are various research results for robust and accurate human pose estimation, this paper introduces them in two separated subsections: 2D human pose estimation and 3D human pose estimation. Moreover, this paper summarizes famous data sets used in this field and introduces various studies which utilize human poses to solve their own problem.

2

카메라를 사용하여 사람의 자세를 추정하는 연구는 비전 연구에서 오래된 주제 중 하나이다. 최근에 도 비전에 기반하여 사람의 자세를 추정하기 위한 연구들이 진행되고 있지만 단일 카메라만으로 정확 한 3D 자세를 추정하기 위해 중요한 깊이 값과 실시간 성능을 위한 추론 속도를 함께 추구하기에는 한계가 있었다. 이러한 문제들을 해결하기 위하여 본 논문에서 제안하는 방법을 활용하여 깊이 값 추 론을 통한 정확성과 실시간으로 사용할 수 있는 추론 속도를 가지는 방법을 제안한다.

3

Prior-free 3D human pose estimation in a video using limb-vectors

Memon Anam, Arain Qasim, Pirzada Nasrullah, Shaikh Akram, Sulaiman Adel, Al Reshan Mana Saleh, Alshahrani Hani, Shaikh Asadullah

[NRF 연계] 한국통신학회 ICT Express Vol.10 No.6 2024.12 pp.1266-1272

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Estimating accurate 3D human poses from a monocular video is fundamental to various computer vision tasks. Existing methods exploit 2D-to-3D pose lifting, multiview images, and depth sensors to model spatio-temporal dependencies. However, depth ambiguities, occlusions, and larger temporal receptive fields pose challenges to these approaches. To address this, we propose a novel prior-free DCNN-based 3D human pose estimation method for monocular image sequences using limb vectors. Our method comprises two subnetworks: a limb direction estimator and a limb length estimator. The limb direction estimator utilizes a fully convolutional network to model limb direction vectors across a temporal window. We show that network complexity can be significantly reduced by utilizing dilated convolutional operations and a relatively smaller receptive field while maintaining estimation accuracy. Moreover, the limb length estimator captures stable limb length estimations from a reliable frame set. Our model has shown superior performance compared to existing methods on the Human3.6M and MPI-INF-3DHP datasets.

4

Often most of the modern human people are suffering from a long time of working or studying on the stationary pose. Subsequently, the health of our life is highly threatened to be exacerbated by chronic orthopedic diseases. In order to solve this social problem, we suggest pose detection that can have the people who have deleterious postures be notified. By using nowadays advanced computer vision techniques, in this paper we suggest the posture recognition module to enhance our quality of life. While most posture recognition recognizes only one person's posture, we made our pipeline to perform posture recognition for multiple people through images obtained through a single camera. One of the big problems in measuring people's postures is that it is necessary to distinguish the various body structures and postures of people. For this, posture images and labeling of various people are required. We created pose images of people of various body types through images of a small number of people through skeleton-based coordinates augmentation. We made a posture classifier using various models and observed the improvement of augmentation performance for each model. Through this, we found that the postures of various people can be measured using a relatively small data set. In particular, for deep-learning models that require a lot of data, generalization performance was greatly improved.

5

4,000원

6

인공지능 인체 자세추정 기반 3차원 다중객체추적에 관한 연구 KCI 등재

남성현, 김영원

국제차세대융합기술학회 차세대융합기술학회논문지 제7권 8호 2023.08 pp.1182-1189

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문에서는 별도의 고사양 카메라 또는 센서를 사용하지 않고, 단일 RGB 카메라와 경량화된 자세추 정 모델만으로 3차원 다중 객체 추정 방법에 대하여 제안하였다. 이를 위하여 자세추정 모델로부터 추정한 2D 관 절 키포인트 좌표값으로 검출 박스를 추정하였으며, 검출 박스의 jitter 오류를 보완하고자 OneEuroFilter를 사용 하였다. 검출된 박스의 중심점을 사용하여 객체의 좌우 이동을 추적하였으며, 박스의 크기와 높이 그리고 비율을 사용한 깊이 추정을 통하여 객체 추적 알고리즘을 개발하였다. 또한, 다중 객체 추적을 위하여 객체가 검출된 프 레임을 ID 값으로 하여 전후 프레임간 검출된 관절 키포인트의 유사도를 검사하여 ID 매칭을 수행하였다.

This paper proposed a 3D MOT method with only a single RGB camera and a lightweight pose estimation model without using a separate high-end camera or sensor. To this end, the detection box was estimated from the 2D joint keypoint coordinate estimated from the pose estimation model and OneEuroFilter was used to compensate for the jitter error in the detection box. The center point of the detected box was used to track the left and right movement of the object, and an object tracking algorithm was developed through depth estimation using the size, height, and ratio of the box. In addition, for multiple object tracking, ID matching was performed by examining the similarity of joint key points detected between current and previous frames using the frame in which the object was detected as an ID value.

7

인공지능을 활용한 머신 비전 기술의 발달로 수많은 포즈 인식 솔루션이 출시되어 이들을 활용하여 인체 관절 좌표를 쉽게 추출할 수 있다. 이 추출된 여러 관절 좌표의 고유값을 이용하여 인체 움직임을 인식하는 방법을 제시한다. 2D 카메라 입력 영상의 각 인체의 3D좌표를 획득하고 각 좌표의 움직임의 총량을 계산하기 위해 관절 좌표의 고유값을 사용하여 동작을 인식한다. 이를 통해 대표적인 몇 개의 관절 좌표를 활용하여 각도, 방향등으로 동작을 인식하는 방법에 비해 효율적이고 노이즈등에 강하며 전체 움직임을 종합적으로 판단할 수 있다.

8

현재 평균 수명 증가와 더불어 고령자의 일상 생활 중 안전 사고가 빈번히 발생하고 있으며, 이 중 낙상 사고가 큰 비중을 차지하고 있다. 피해를 최소화하기 위해서는 낙상 사고가 발생 시 신속하게 응급상황 대응기관으로 사고 발생 여부 및 위치를 알리는 기술이 필요하다. 본 연구에서는 가속도 센서가 탑재된 스마트 밴드와 영상 기반 자세 정보를 동시에 사용하여 낙상 여부를 판단하는 기법을 제안한다. 특히 복합 센서의 취득 정보 처리 지연시간을 일정 수준으로 관리하여, 응급 상황시에도 이를 활용할 수 있도록 하는 것을 목표로 한다. 이를 확인하기 위해 처리 프로세스의 지연시간을 분석하는 실험을 진행하였다.

9

Pose estimation which is regard as the cross-technology of computer vision and pattern recognition, and an important prerequisite for human behavior understanding. Human pose estimation which use the probability theory, machine learning, pattern recognition, graph theory and other theories to get the position, the deflection angle of the various parts of the body. Then make the detection and estimation parameters for the human body pose. When the image has interference in the background, color and scale changed, human pose complex, self-occlusion and interpersonal interaction occlusion may make the precision and accuracy of pose estimation face great challenge. Thus, according to the above problems, this paper use the advanced model of the human body as contour model to descript the complex pose, in order to make the model more accurate and suitable for various human pose, we pre-clustering the human body pose of the training samples before we trained the model and in order to ensure the accuracy of the pose we use robust segmentation of multi-view with a novel shape prior. The experiment shows that the algorithm performs better than the classic algorithm on the public datasets.

10

Lightening of Human Pose Estimation Algorithm Using MobileViT and Transfer Learning

Kunwoo Kim, Jonghyun Hong, Jonghyuk Park

[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.28 No.9 2023 pp.17-25

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 매개변수가 더 적고, 빠르게 추정 가능한 MobileViT 기반 모델을 통해 사람 자세 추정 과업을 수행할 수 있는 모델을 제안한다. 기반 모델은 합성곱 신경망의 특징과 Vision Transformer의 특징이 결합한 구조를 통해 경량화된 성능을 입증한다. 본 연구에서 주요 매커니즘이 되는 Transformer는 그 기반의 모델들이 컴퓨터 비전 분야에서도 합성곱 신경망 기반의 모델들 대비 더 나은 성능을 보이며, 영향력이 커지게 되었다. 이는 사람 자세 추정 과업에서도 동일한 상황이며, Vision Transformer기반의 ViTPose가 COCO, OCHuman, MPII 등 사람 자세 추정 벤치마크에서 모두 최고 성능을 지키고 있는 것이 그 적절한 예시이다. 하지만 Vision Transformer는 매개변수의 수가 많고 상대적으로 많은 연산량을 요구하는 무거운 모델 구조를 가지고 있기 때문에, 학습에 있어 사용자에게 많은 비용을 야기시킨다. 이에 기반 모델은 Vision Transformer가 많은 계산량을 요구하는 부족한 Inductive Bias 계산 문제를 합성곱 신경망 구조를 통한 Local Representation으로 극복하였다. 최종적으로, 제안 모델은 MS COCO 사람 자세 추정 벤치마크에서 제공하는 Validation Set으로 ViTPose 대비 각각 5분의 1과 9분의 1만큼의 3.28GFLOPs, 972만 매개변수를 나타내었고, 69.4 Mean Average Precision을 달성하여 상대적으로 우수한 성능을 보였다.

In this paper, we propose a model that can perform human pose estimation through a MobileViT-based model with fewer parameters and faster estimation. The based model demonstrates lightweight performance through a structure that combines features of convolutional neural networks with features of Vision Transformer. Transformer, which is a major mechanism in this study, has become more influential as its based models perform better than convolutional neural network-based models in the field of computer vision. Similarly, in the field of human pose estimation, Vision Transformer-based ViTPose maintains the best performance in all human pose estimation benchmarks such as COCO, OCHuman, and MPII. However, because Vision Transformer has a heavy model structure with a large number of parameters and requires a relatively large amount of computation, it costs users a lot to train the model. Accordingly, the based model overcame the insufficient Inductive Bias calculation problem, which requires a large amount of computation by Vision Transformer, with Local Representation through a convolutional neural network structure. Finally, the proposed model obtained a mean average precision of 0.694 on the MS COCO benchmark with 3.28 GFLOPs and 9.72 million parameters, which are 1/5 and 1/9 the number compared to ViTPose, respectively.

11

A Kidnapping Detection Using Human Pose Estimation in Intelligent Video Surveillance Systems

Park, Ju Hyun, Song, KwangHo, Kim, Yoo-Sung

[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.23 No.8 2018 pp.9-16

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, a kidnapping detection scheme in which human pose estimation is used to classify accurately between kidnapping cases and normal ones is proposed. To estimate human poses from input video, human's 10 joint information is extracted by OpenPose library. In addition to the features which are used in the previous study to represent the size change rates and the regularities of human activities, the human pose estimation features which are computed from the location of detected human's joints are used as the features to distinguish kidnapping situations from the normal accompanying ones. A frame-based kidnapping detection scheme is generated according to the selection of J48 decision tree model from the comparison of several representative classification models. When a video has more frames of kidnapping situation than the threshold ratio after two people meet in the video, the proposed scheme detects and notifies the occurrence of kidnapping event. To check the feasibility of the proposed scheme, the detection accuracy of our newly proposed scheme is compared with that of the previous scheme. According to the experiment results, the proposed scheme could detect kidnapping situations more 4.73% correctly than the previous scheme.

12

Fast Random-Forest-Based Human Pose Estimation Using a Multi-scale and Cascade Approach

Chang, Ju Yong, Nam, Seung Woo

[Kisti 연계] 한국전자통신연구원 ETRI journal Vol.35 No.6 2013 pp.949-959

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Since the recent launch of Microsoft Xbox Kinect, research on 3D human pose estimation has attracted a lot of attention in the computer vision community. Kinect shows impressive estimation accuracy and real-time performance on massive graphics processing unit hardware. In this paper, we focus on further reducing the computation complexity of the existing state-of-the-art method to make the real-time 3D human pose estimation functionality applicable to devices with lower computing power. As a result, we propose two simple approaches to speed up the random-forest-based human pose estimation method. In the original algorithm, the random forest classifier is applied to all pixels of the segmented human depth image. We first use a multi-scale approach to reduce the number of such calculations. Second, the complexity of the random forest classification itself is decreased by the proposed cascade approach. Experiment results for real data show that our method is effective and works in real time (30 fps) without any parallelization efforts.

13

A Spatial-Temporal Three-Dimensional Human Pose Reconstruction Framework

Nguyen, Xuan Thanh, Ngo, Thi Duyen, Le, Thanh Ha

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.15 No.2 2019 pp.399-409

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Three-dimensional (3D) human pose reconstruction from single-view image is a difficult and challenging topic. Existing approaches mostly process frame-by-frame independently while inter-frames are highly correlated in a sequence. In contrast, we introduce a novel spatial-temporal 3D human pose reconstruction framework that leverages both intra and inter-frame relationships in consecutive 2D pose sequences. Orthogonal matching pursuit (OMP) algorithm, pre-trained pose-angle limits and temporal models have been implemented. Several quantitative comparisons between our proposed framework and recent works have been studied on CMU motion capture dataset and Vietnamese traditional dance sequences. Our framework outperforms others by 10% lower of Euclidean reconstruction error and more robust against Gaussian noise. Additionally, it is also important to mention that our reconstructed 3D pose sequences are more natural and smoother than others.

14

An Accurate Forward Head Posture Detection using Human Pose and Skeletal Data Learning

Jong-Hyun Kim

[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.28 No.8 2023 pp.87-93

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 사용자의 골격 자세를 분석하여 네트워크 학습 기반으로 거북목 자세를 정확하고 효율적으로 판별하는 시스템을 제안한다. 거북목 증후군이란 목이 구부정하게 앞으로 나오는 자세를 오래 유지함으로써 목의 자세가 바뀌고 뒷목, 어깨, 허리 등에 통증이 생기는 증상을 말하며, 수술이나 약물치료보다 평소의 자세 습관이 효과적이라고 알려져 있다. 기존의 방법들은 웹캠을 이용한 합성곱 신경망을 이용하였고, 이러한 접근법은 영상의 명도와 조명, 피부 색 등에 영향을 받기 때문에 특정 인물에 대해서만 수행되는 문제가 있다. 본 논문에서는 이 문제를 완화하고 자 영상으로부터 골격을 추출하고, 정면보다는 측면에 해당하는 데이터를 학습하여 이전 기법보다 효율적이고 정확하게 거북목 자세를 찾아낸다. 결과적으로 이전 기법에 비해 다양한 실험 장면에서 정확도가 되었음을 보여준다.

In this paper, we propose a system that accurately and efficiently determines forward head posture based on network learning by analyzing the user's skeletal posture. Forward head posture syndrome is a condition in which the forward head posture is changed by keeping the neck in a bent forward position for a long time, causing pain in the back, shoulders, and lower back, and it is known that daily posture habits are more effective than surgery or drug treatment. Existing methods use convolutional neural networks using webcams, and these approaches are affected by the brightness, lighting, skin color, etc. of the image, so there is a problem that they are only performed for a specific person. To alleviate this problem, this paper extracts the skeleton from the image and learns the data corresponding to the side rather than the frontal view to find the forward head posture more efficiently and accurately than the previous method. The results show that the accuracy is improved in various experimental scenes compared to the previous method.

15

Robust 2D human upper-body pose estimation with fully convolutional network

Lee, Seunghee, Koo, Jungmo, Kim, Jinki, Myung, Hyun

[Kisti 연계] 테크노프레스 Advances in robotics research Vol.2 No.2 2018 pp.129-140

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

With the increasing demand for the development of human pose estimation, such as human-computer interaction and human activity recognition, there have been numerous approaches to detect the 2D poses of people in images more efficiently. Despite many years of human pose estimation research, the estimation of human poses with images remains difficult to produce satisfactory results. In this study, we propose a robust 2D human body pose estimation method using an RGB camera sensor. Our pose estimation method is efficient and cost-effective since the use of RGB camera sensor is economically beneficial compared to more commonly used high-priced sensors. For the estimation of upper-body joint positions, semantic segmentation with a fully convolutional network was exploited. From acquired RGB images, joint heatmaps accurately estimate the coordinates of the location of each joint. The network architecture was designed to learn and detect the locations of joints via the sequential prediction processing method. Our proposed method was tested and validated for efficient estimation of the human upper-body pose. The obtained results reveal the potential of a simple RGB camera sensor for human pose estimation applications.

16

트랜스포머 기반의 다중 시점 3차원 인체자세추정

최승욱, 이진영, 김계영

[Kisti 연계] 한국스마트미디어학회 스마트미디어저널 Vol.12 No.11 2023 pp.48-56

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

3차원 인체자세추정은 스포츠, 동작인식, 영상매체의 특수효과 등의 분야에서 널리 활용되고 있는 기술이다. 이를 위한 여러 방법들 중 다중 시점 3차원 인체자세추정은 현실의 복잡한 환경에서도 정밀한 추정을 하기 위해 필수적인 방법이다. 하지만 기존 다중 시점 3차원 인체자세추정 모델들은 3차원 특징 맵을 사용함에 따라 시간 복잡도가 높은 단점이 있다. 본 논문은 계산 복잡도가 적은 트랜스포머 기반 기존 단안 시점 다중 프레임 모델을 다중 시점에 대한 3차원 인체자세추정으로 확장하는 방법을 제안한다. 다중 시점으로 확장하기 위하여 먼저 2차원 인체자세 검출자 CPN(Cascaded Pyramid Network)을 활용하여 획득한 4개 시점의 17가지 관절에 대한 2차원 관절좌표를 연결한 8차원 관절좌표를 생성한다. 그 다음 이들을 패치 임베딩 한 뒤 17×32 데이터로 변환하여 트랜스포머 모델에 입력한다. 마지막으로, 인체자세를 출력하는 MLP(Multi-Layer Perceptron) 블록을 매 반복 마다 사용한다. 이를 통해 4개 시점에 대한 3차원 인체자세추정을 동시에 수정한다. 입력 프레임 길이 27을 사용한 Zheng[5]의 방법과 비교했을 때 제안한 방법의 모델 매개변수의 수는 48.9%, MPJPE(Mean Per Joint Position Error)는 20.6mm(43.8%) 감소했으며, 학습 횟수 당 평균 학습 소요 시간은 20배 이상 빠르다.

The technology of Three-dimensional human posture estimation is used in sports, motion recognition, and special effects of video media. Among various methods for this, multi-view 3D human pose estimation is essential for precise estimation even in complex real-world environments. But Existing models for multi-view 3D human posture estimation have the disadvantage of high order of time complexity as they use 3D feature maps. This paper proposes a method to extend an existing monocular viewpoint multi-frame model based on Transformer with lower time complexity to 3D human posture estimation for multi-viewpoints. To expand to multi-viewpoints our proposed method first generates an 8-dimensional joint coordinate that connects 2-dimensional joint coordinates for 17 joints at 4-vieiwpoints acquired using the 2-dimensional human posture detector, CPN(Cascaded Pyramid Network). This paper then converts them into 17×32 data with patch embedding, and enters the data into a transformer model, finally. Consequently, the MLP(Multi-Layer Perceptron) block that outputs the 3D-human posture simultaneously updates the 3D human posture estimation for 4-viewpoints at every iteration. Compared to Zheng[5]'s method the number of model parameters of the proposed method was 48.9%, MPJPE(Mean Per Joint Position Error) was reduced by 20.6 mm (43.8%) and the average learning time per epoch was more than 20 times faster.

17

RGB 이미지 기반 인간 동작 추정을 통한 투구 동작 분석

우영주, 주지용, 김영관, 정희용

[Kisti 연계] 한국스마트미디어학회 스마트미디어저널 Vol.13 No.4 2024 pp.16-22

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

투구는 야구의 시작이라 할 만큼 야구에서 주요한 부분을 차지한다. 투구 동작의 정확한 분석은 경기력 향상과 부상 예방 측면에서 매우 중요하다. 올바른 투구 동작을 분석할 때, 현재 주로 사용되는 모션캡처는 환경적으로 치명적인 단점들이 몇 가지 존재한다. 본 논문에서 우리는 이러한 단점들이 존재하는 모션캡처를 대체하기 위하여 RGB 기반의 Human Pose Estimation(HPE) 모델을 활용한 투구 동작의 분석을 제안하며 이에 대한 신뢰도를 검증하기 위해 모션캡처 데이터와 HPE 데이터의 관절 좌표를 Dynamic Time Warping(DTW) 알고리즘의 비교를 통해 두 데이터의 유사도를 검증하였다.

Pitching is a major part of baseball, so much so that it can be said to be the beginning of baseball. Analysis of accurate pitching motions is very important in terms of performance improvement and injury prevention. When analyzing the correct pitching motion, the currently used motion capture method has several critical environmental drawbacks. In this paper, we propose analysis of pitching motion using the RGB-based Human Pose Estimation (HPE) model to replace motion capture, which has these shortcomings, and use motion capture data and HPE data to verify its reliability. The similarity of the two data was verified by comparing joint coordinates using the Dynamic Time Warping (DTW) algorithm.

18

가려진 사람의 자세추정을 위한 의미론적 폐색현상 증강기법

배현재, 김진평, 이지형

[Kisti 연계] 한국정보처리학회 정보처리학회논문지/소프트웨어 및 데이터 공학 Vol.11 No.12 2022 pp.517-524

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

사람의 자세추정(Human pose estimation)은 사람의 관절 키포인트를 추출하여 자세를 추정하는 방법이다. 폐색현상(Occlusion)이 발생하면, 사람의 관절이 가려지므로 관절 키포인트 추출 성능이 낮아진다. 폐색현상은 총 3가지로 행동할 때 스스로 가려짐, 다른 사물에 의해 가려짐과 배경에 의해 가려짐으로 크게 나뉜다. 본 논문에서는 폐색현상 증강기법을 활용하여 효과적인 자세추정방법을 제안한다. 자세추정방법이 지속적으로 연구되어왔지만, 자세추정방법의 가려짐 현상에 관한 연구는 상대적으로 부족한 상태이다. 이를 해결하기 위해 저자는 사람의 관절을 타겟팅하여 의도적으로 가리는 데이터 증강기법을 제안한다. 본 논문에서의 실험 결과는 의도적으로 폐색현상 증강기법을 활용하면 폐색현상에 강인하며 성능이 올라간 것을 보여준다.

Human pose estimation is a method of estimating a posture by extracting a human joint key point. When occlusion occurs, the joint key point extraction performance is lowered because the human joint is covered. The occlusion phenomenon is largely divided into three types of actions: self-contained, covered by other objects, and covered by background. In this paper, we propose an effective posture estimation method using a masking phenomenon enhancement technique. Although the posture estimation method has been continuously studied, research on the occlusion phenomenon of the posture estimation method is relatively insufficient. To solve this problem, the author proposes a data augmentation technique that intentionally masks human joints. The experimental results in this paper show that the intentional use of the blocking phenomenon enhancement technique is strong against the blocking phenomenon and the performance is increased.

19

엣지 디바이스와 카메라 센서 퓨전을 활용한 사람 자세 데이터 자동 수집 시스템

김영근, 김승현, 김정곤, 김원중

[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.19 No.1 2024 pp.189-196

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

지능형 선별 관제 시스템의 잦은 오탐지로 인해 관제 요원들의 업무 능률 및 시장 신뢰도 저하 문제가 꾸준히 보고되고 있다. 오탐지 문제 개선을 위해 새 AI 모델을 개발하거나 교체하는 것은 기회비용이 크므로, 훈련 데이터 세트 품질을 향상하여 문제를 개선하는 것이 현실적이다. 그러나 소규모 조직은 데이터 세트 수집 및 정제 역량이 부족한 실정이다. 이에 본 논문에서는 사람 자세 추정 모델을 중심으로 엣지 디바이스와 카메라 센서 퓨전을 활용한 사람 자세 데이터 자동 수집 시스템을 제안한다. 이 시스템은 네트워크 말단에서 현장 데이터를 직접 수집하고 레이블링하는 과정을 실시간으로 처리하도록 만들어, 중앙으로 집중되는 연산 부하를 분산시킨다. 또한 현장 데이터를 직접 레이블링하므로 새로운 훈련 데이터 구축에 도움을 준다.

Frequent false positives alarm from the Intelligent Selective Control System have raised significant concerns. These persistent issues have led to declines in operational efficiency and market credibility among agents. Developing a new model or replacing the existing one to mitigate false positives alarm entails substantial opportunity costs; hence, improving the quality of the training dataset is pragmatic. However, smaller organizations face challenges with inadequate capabilities in dataset collection and refinement. This paper proposes an automatic human pose data collection system centered around a human pose estimation model, utilizing camera-based sensor fusion techniques and edge devices. The system facilitates the direct collection and real-time processing of field data at the network periphery, distributing the computational load that typically centralizes. Additionally, by directly labeling field data, it aids in constructing new training datasets.

20

다시점 준지도 학습 기반 3차원 휴먼 자세 추정

김도엽, 장주용

[Kisti 연계] 한국방송공학회 방송공학회논문지 Vol.27 No.2 2022 pp.174-184

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

3차원 휴먼 자세 추정 모델은 다시점 모델과 단시점 모델로 분류될 수 있다. 일반적으로 다시점 모델은 단시점 모델에 비하여 뛰어난 자세 추정 성능을 보인다. 단시점 모델의 경우 3차원 자세 추정 성능의 향상은 많은 양의 학습 데이터를 필요로 한다. 하지만 3차원 자세에 대한 참값을 획득하는 것은 쉬운 일이 아니다. 이러한 문제를 다루기 위해, 우리는 다시점 모델로부터 다시점 휴먼 자세 데이터에 대한 의사 참값을 생성하고, 이를 단시점 모델의 학습에 활용하는 방법을 제안한다. 또한, 우리는 각각의 다시점 영상으로부터 추정된 자세의 일관성을 고려하는 다시점 일관성 손실함수를 제안하여, 이것이 단시점 모델의 효과적인 학습에 도움을 준다는 것을 보인다. Human3.6M과 MPI-INF-3DHP 데이터셋을 사용한 실험은 제안하는 방법이 3차원 휴먼 자세 추정을 위한 단시점 모델의 학습에 효과적임을 보여준다.

3D human pose estimation models can be classified into a multi-view model and a single-view model. In general, the multi-view model shows superior pose estimation performance compared to the single-view model. In the case of the single-view model, the improvement of the 3D pose estimation performance requires a large amount of training data. However, it is not easy to obtain annotations for training 3D pose estimation models. To address this problem, we propose a method to generate pseudo ground-truths of multi-view human pose data from a multi-view model and exploit the resultant pseudo ground-truths to train a single-view model. In addition, we propose a multi-view consistency loss function that considers the consistency of poses estimated from multi-view images, showing that the proposed loss helps the effective training of single-view models. Experiments using Human3.6M and MPI-INF-3DHP datasets show that the proposed method is effective for training single-view 3D human pose estimation models.

 
1 2
페이지 저장