년 - 년
OpenVINO-based Mixed- Precision Quantization for Accelerating RTMPose Inference
국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.17 No.4 2025.11 pp.328-337
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Post-Training Quantization (PTQ), a method for model quantization without the need for retraining, is under active investigation for the deployment of real-time human pose estimation techniques on resourceconstrained edge devices. However, conventional PTQ techniques exhibit a critical limitation when applied to the SimCC output format utilized by the RTMPose model, leading to severe performance degradation due to the high sensitivity of its coordinate representation. This study proposes an optimization framework to overcome this challenge by systematically identifying optimal quantization targets through a Singular Value Decomposition (SVD)-based sensitivity analysis and by correcting the distorted output distribution via realtime post-processing. Experimental results demonstrate that the proposed framework achieves a 1.56-fold model compression while successfully retaining 88.8% of the accuracy of the original FP32 model. This research is anticipated to significantly enhance the practical deployment of models with sensitive output architectures, such as RTMPose, thereby contributing to a wide range of applications in real-time human pose estimation.
양자화 기반의 모델 압축을 이용한 ONNX 경량화 KCI 등재
국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제21권 제1호 2021.02 pp.93-98
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
딥 러닝의 발전으로 다양한 AI 기반의 응용이 많아지고, 그 모델의 규모도 매우 커지고 있다. 그러나 임베디드 기기와 같이 자원이 제한적인 환경에서는 모델의 적용이 어렵거나 전력 부족 등의 문제가 존재한다. 이를 해결하기 위해 서 클라우드 기술 또는 오프로딩 기술을 활용하거나, 모델의 매개변수 개수를 줄이거나 계산을 최적화하는 등의 경량화 방법이 제안되었다. 본 논문에서는 다양한 프레임워크들의 상호 교환 포맷으로 사용되고 있는 ONNX(개방형 신경망 교환 포맷) 포맷에 딥러닝 경량화 방법 중 학습된 모델의 양자화를 적용한다. 경량화 전 모델과의 신경망 구조와 추론 성능을 비교하고, 양자화를 위한 다양한 모듈 방식를 분석한다. 실험을 통해 ONNX의 양자화 결과, 정확도는 차이가 거의 없으며 기존 모델보다 매개변수 크기가 압축되었으며 추론 시간 또한 전보다 최적화되었음을 알 수 있었다.
Due to the development of deep learning and AI, the scale of the model has grown, and it has been integrated into other fields to blend into our lives. However, in environments with limited resources such as embedded devices, it is exist difficult to apply the model and problems such as power shortages. To solve this, lightweight methods such as clouding or offloading technologies, reducing the number of parameters in the model, or optimising calculations are proposed. In this paper, quantization of learned models is applied to ONNX models used in various framework interchange formats, neural network structure and inference performance are compared with existing models, and various module methods for quantization are analyzed. Experiments show that the size of weight parameter is compressed and the inference time is more optimized than before compared to the original model.
[Kisti 연계] 대한전자공학회 대한전자공학회 학술대회논문집 2002 pp.1240-1243
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this paper, we propose an adaptive digital image watermarking algorithm using successive subband quantization (SSQ) and perceptual model based on wavelet domain. The watermark is embedded into the perceptually significant coefficients (PSCs) of image. The PSCs in the baseband are selected according to the amplitude of the coefficients and the high frequency subbands are selected by SSQ. To embed the watermark, we use perceptual model. The perceptual model is based on the computation of the noise visibility function (NVF) and embed at the texture and edge region stronger embedded watermarks.
[Kisti 연계] 한국로봇학회 로봇학회논문지 Vol.5 No.1 2010 pp.70-76
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper addresses a design issue of "model complexity and performance trade-off" in the application of bandwidth extension (BWE) methods to the intra-frame predictivevector quantization problem of wideband speech. It discusses model-based linear and non-linear prediction methods and presents a comparative study of them in terms of prediction gain. Through experimentation, the general trend of saturation in performance (with the increase in model complexity) is observed. However, specifically, it is also observed that there is no significant difference between HMM and GMM-based BWE functions.
[Kisti 연계] 제어로봇시스템학회 제어로봇시스템학회 학술대회논문집 1991 pp.131-135
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
A quantization error model is suggested for analog to frequency(A/F) converter in strapdown inertial navigation system(SDINS),which is characterized by some white noise exciting the state variables. Also, effects on the performance of SDINS by analog to digital(A/D) converter and A/F converter are analyzed and compared via covariance simulation. As a result, A/F converter turns out to be superior to the A/D converter with respect to the induced navigation error and the difficulty in circuit realization. The quantization error model developed in this paper appears to be useful for optimal filter design.
멀티웨이브릿 변환 영역 기반의 연속 부대역 양자화 및 지각 모델을 이용한 적응 워터마킹
[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.6 No.7 2003 pp.1149-1158
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 멀티웨이브릿 변환영역에서 연속부대역 양자화 및 지각 모델을 이용한 내용기반 적응적 워터마킹 기법을 제안한다. 제안한 방법의 워터마크는 멀티웨이브릿을 통해 분해된 계수들 중 지각적 중요계수(perceptually significant coefficients, PSCs)에 삽입된다. 고주파 부대역에서의 PSC는 연속부대역양자화(successive subband quantization, SSQ)에 의해 결정된다. 문턱값은 각 부대역내의 최대계수의 절반에서 결정된다. 지각모델은 워터마크 삽입을 위한 국부적 영상 특성을 가지는 NVF (noise visibility function)에 기반한 통계적 방법을 적용한다. 이 모델은 워터마크가 노이즈특성을 가지므로 정상상태 일반화 가우스모델을 사용한다. 또한 워터마크는 각 부대역 영역의 분산과 형상계수 (shape parameter)에 의해 추정함으로써 평탄영역과 에지나 텍스쳐 영역에 따라 내용 기반 적응적 척도를 얻는다. 제안한 멀티웨이브릿 변환 기반에서의 워터마크 삽입 방법에 대한 실험 결과 우수한 강인성과 비가시성을 확인하였다.
Content adaptive watermark embedding algorithm using a stochastic image model in the multiwavelet transform is proposed in this paper. A watermark is embedded into the perceptually significant coefficients (PSCs) of each subband using multiwavelet transform. The PSCs in high frequency subband are selected by SSQ, that is, by setting the thresholds as the one half of the largest coefficient in each subband. The perceptual model is applied with a stochastic approach based on noise visibility function (NVF) that has local image properties for watermark embedding. This model uses stationary Generalized Gaussian model characteristic because watermark has noise properties. The watermark estimation use shape parameter and variance of subband region. it is derive content adaptive criteria according to edge and texture, and flat region. The experiment results of the proposed watermark embedding method based on multiwavelet transform techniques were found to be excellent invisibility and robustness.
멀티웨이브릿 변환 기반에서 연속 부대역 양자화 및 지각 모델을 이용한 적응 워터마킹 기술
[Kisti 연계] 대한전자공학회 대한전자공학회 학술대회논문집 2002 pp.121-124
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper presents an adaptive digital image watermarking scheme that uses successive subband quantization (SSQ) and perceptual modeling. Our approach performs a multiwavelet transform to determine the local image properties optimal and the watermark embedding location. The multiwavelet used in this paper is the DGHM multiwavelet with approximation order 2 to reduce artifacts in the reconstructed image. A watermark is embedded into the perceptually significant coefficients (PSC) of the image in each subband. The PSCs in high frequency subbands are selected by setting the thresholds to one half of the largest coefficient in each subband. After the PSCs in each subband are selected, a perceptual model is combined with a stochastic approach based on the noise visibility function to produce the final watermark.
양자화 입력을 고려한 연속시간 T-S 퍼지 시스템을 위한 이벤트 트리거 모델예측제어
[Kisti 연계] 대한전기학회 電氣學會論文誌 Vol.66 No.9 2017 pp.1364-1372
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this paper, a problem of event-triggered model predictive control is investigated for continuous-time Takagi-Sugeno (T-S) fuzzy systems with input quantization. To efficiently utilize network resources, event-trigger is employed, which transmits limited signals satisfying the condition that the measurement of errors is over the ratio of a certain level. Considering sampling and quantization, continuous Takagi-Sugeno (T-S) fuzzy systems are regarded as a sector bounded continuous-time T-S fuzzy systems with input delay. Then, a model predictive controller (MPC) based on parallel distributed compensation (PDC) is designed to optimally stabilize the closed loop systems. The proposed MPC optimize the objective function over infinite horizon, which can be easily calculated and implemented solving linear matrix inequalities (LMIs) for every event-triggered time. The validity and effectiveness are shown that the event triggered MPC can stabilize well the systems with even smaller average sampling rate and limited actuator signal guaranteeing optimal performances through the numerical example.
벡터 양자화 변분 오토인코더 기반의 폴리 음향 생성 모델을 위한 잔여 벡터 양자화 적용 연구
[Kisti 연계] 한국음향학회 한국음향학회지 Vol.43 No.2 2024 pp.243-252
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근에 연구되기 시작한 폴리(Foley) 음향 생성 모델 중 벡터 양자화 변분 오토인코더(Vector Quantized-Variational AutoEncoder, VQ-VAE) 구조와 Pixelsnail 등 생성모델을 활용한 생성 기법은 중요한 연구대상 중 하나이다. 한편, 딥러닝 기반의 음향 신호의 압축/복원 분야에서는 기존의 VQ-VAE 구조에 비해 잔여 벡터 양자화 기술이 더 적합한 것으로 보고되고 있으며, 따라서 본 논문에서는 폴리 음향 생성 분야에서도 잔여 벡터 양자화 기술이 효과적으로 적용될 수 있을지 연구하고자 한다. 이를 위하여 본 논문에서는 기존의 VQ-VAE 기반의 폴리 음향 생성 모델에 잔여 벡터 양자화 기술을 적용하되, Pixelsnail 등 기존의 다른 모델과 호환이 가능하고 연산 자원의 소모를 늘리지 않는 모델을 고안하여 그 효과를 확인하고자 하였다. 효과를 검증하기 위하여 DCASE2023 Task7의 데이터를 활용하여 실험을 진행하였으며, 그 결과 평균적으로 0.3 가량의 Fréchet audio distance 의 향상을 보이는 것을 확인하였다. 다만 그 성능 향상의 정도가 제한적이었으며, 이는 연산 자원의 소모를 유지하기 위하여 시간-주파수축의 분해능이 저하된 영향으로 판단된다.
Among the Foley sound generation models that have recently begun to be studied, a sound generation technique using the Vector Quantized-Variational AutoEncoder (VQ-VAE) structure and generation model such as Pixelsnail are one of the important research subjects. On the other hand, in the field of deep learning-based acoustic signal compression, residual vector quantization technology is reported to be more suitable than the conventional VQ-VAE structure. Therefore, in this paper, we aim to study whether residual vector quantization technology can be effectively applied to the Foley sound generation. In order to tackle the problem, this paper applies the residual vector quantization technique to the conventional VQ-VAE-based Foley sound generation model, and in particular, derives a model that is compatible with the existing models such as Pixelsnail and does not increase computational resource consumption. In order to evaluate the model, an experiment was conducted using DCASE2023 Task7 data. The results show that the proposed model enhances about 0.3 of the Fréchet audio distance. Unfortunately, the performance enhancement was limited, which is believed to be due to the decrease in the resolution of time-frequency domains in order to do not increase consumption of the computational resources.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.