년 - 년
커널밀도추정(Kernel Density Estimation)과 합성곱 신경망(Convolution Neural Network)을 이용한 아파트 가격 예측 KCI 등재
한국상업경영학회(구 한국상업교육학회) 상업경영연구(구 상업교육연구) 제37권 제6호 2023.12 pp.33-53
※ 기관로그인 시 무료 이용이 가능합니다.
5,700원
주택가격 예측은 개인, 기관, 정부에 이르기까지 다양한 이해관계자들이 관심을 갖는 주요한 주 제이다. 지금까지 통계적 모델을 통해 예측을 해왔지만, 최근 딥러닝 및 머신러닝을 이용하여 예측 하는 연구가 많다. 특히 이미지 정보를 활용한 합성곱신경망(Convolutional Neural Networks)을 적 용한 연구도 진행되고 있다. 그러나 아직 국내에서는 이미지 정보를 이용한 연구는 살펴볼 수 없었 다. 본 연구는 커널밀도추정(Kernel Density Estimation)을 활용하여 인공위성 이미지와 상가 및 편 의시설 밀도 정보를 해당 구역에 표시하고 이를 입력변수로 활용하여 서울시 아파트 가격을 예측하 고자 한다. 본 논문에서는 2021년부터 2022년까지의 주택 거래 기록과 지역소득정보, 인구 밀도, 연령대별 인구 데이터를 수집하였고, 네이버 클라우드 플랫폼의 Static Map API를 통해 인공위성사진을 구역 별로 나누어서 수집했다. 특히 공공데이터포털과 서울 열린데이터 광장에서 수집한 상가, 병원, 공 원, 지하철역, 학교 등 정보는 커널밀도추정을 활용해 밀도 정보로 변환하였다. 이렇게 다양한 형태 로 수집된 데이터는 전처리 과정을 거쳐 면적당 단가로 계산된 아파트 실거래가를 예측하였다. 예 측모델로서는 회귀 모델, 다층 인공신경망 모델, 그리고 합성곱신경망 모델을 이용하였고 이들의 예 측력을 비교하였다. 연구결과는 다음과 같다. 첫째, 회귀 모델 중에서는 인구통계 및 주택관련 변수 에 커널밀도추정으로 구해진 상가 및 편의시설 밀도 정보를 추가한 모델이 예측력이 높았다. 둘째, 다층인공신경망(Multilayer Artificial Neural Network) 모델은 회귀 모델에 비해 상대적으로 높은 성능을 보였다. 셋째, 합성곱신경망과 다층인공신경망을 융합한 모델은 인공위성 이미지 피처, 인구 통계 및 주택 관련 변수, 그리고 커널밀도추정 특성을 모두 사용한 경우 기존 다층 인공신경망과 회귀모델과 비교시 가장 우수한 성능을 보였다. 이 결과는 인공위성 이미지가 주택 가격 예측에 유 용한 정보를 제공하는 것을 시사하며 향후 주택가격예측에 인공위성 이미지, 혹은 커널밀도추정으 로 표시되는 밀도 정보를 합성곱신경망을 이용하는 경우 예측력이 높아졌기에, 향후 이를 개선하는 연구를 기대한다.
House price forecasting is a major topic of interest to a wide range of stakeholders, from individuals to institutions to governments. Until now, predictions have been made using statistical models, but recently, there have been many studies using deep learning and machine learning. In particular, researchers are applying CNN (Convolutional Neural Networks) using image information. However, there have been no studies using image information in Korea. This paper aims to predict apartment prices in Seoul by using Kernel Density Estimation (KDE) to display satellite images and density information of shopping malls and amenities in the area and use them as input variables. In this paper, we collected housing transaction records, local income information, population density, and population data by age group from 2021 to 2022, and collected satellite images divided into zones through Naver Cloud Platform’s Static Map API. In particular, information on shopping malls, hospitals, parks, subway stations, and schools collected from the Open government data portal(www.data.go.kr) and Seoul open data square(data.seoul.go.kr) was converted into density information using KDE. The data collected in various forms was preprocessed to predict the actual transaction price of apartments calculated as a unit price per area. Regression model, multi-layer artificial neural network model, and CNN model were used as prediction models, and their predictive power was compared. The results are as follows. First, among the regression models, the model that adds shopping centre and amenity density information obtained from KDE to demographic and housing-related variables has a high predictive power. Second, the MLP model has a relatively high performance compared to the regression model. Third, the fusion of CNN and MLP performed the best when satellite image features, demographic and housing-related variables, and KDE characteristics were all used, compared to traditional multilayer neural networks and regression models. These results suggest that satellite imagery provides useful information for house price prediction, and we look forward to future research to improve the predictive power of satellite imagery and density information represented by KDE using CNN for house price prediction.
평생직업교육학원 시설의 분포 특성과 입지 요인 분석 KCI 등재
한국지역개발학회 한국지역개발학회지 32권 3호 통권 112집 2020.09 pp.123-146
※ 기관로그인 시 무료 이용이 가능합니다.
6,100원
본 연구는 비형식 형태의 평생교육기관 중 가장 큰 규모를 차지하고 노동시장 환경 변화에 더 신속하게 대응하며 급증세를 보이는 평생직업교육학원을 중심으로 종류별 시설의 분포 현황과 입지적 특성 및 공간 분포에 영향을 미치는 사회경제적 요인을 추정하는 데 목적을 두고 실증분석을 진행하였다. 분석 결과, 모든 학원 종류별 시설의 공간입지는 거주 및 활동 인구가 밀집된 곳에 집중적으로 분포하며, 특히 직업기술 분야 시설의 입지는 20‧30대 밀레니얼 세대 거주자들이 많은 곳이면서 동시에 산업‧경제 활동을 수행하는 사업체의 분포도가 높은 곳, 특히 전문직과 사무직의 화이트칼라 계층의 취업 분포가 높은 곳과 유의한 관계를 맺었다. 평생직업교육학원의 경우 일부 특정 지역을 중심으로 시설 입지가 집중되는 양상을 보이므로 공공의 지역 차원에서 다양한 평생직업교육 참여 기회를 제공하고, 특히 취업과 연계되어 직업능력향상을 도모하는 프로그램을 생산해 공급 보완하며 학원 시설의 지역 간 차이에서 발생하는 수요자의 지리적 접근성의 격차를 완화하는 지역 밀착형의 평생교육체제가 구축되어야 할 것이다.
In the new normal era, as the importance of life-long vocational competency development activities increased, the demand for life-long vocational education of adults, especially for vocational education and training, increased remarkably. Therefore, this study conducts an empirical analysis with the aim of estimating the socio-economic factors affecting the locational distribution and spatial characteristics by type of non-formal life-long vocational education and training facilities. At the results of spatial empirical analysis, due to the nature of market-oriented private institute services, the location of all life-long vocational education and training facilities are concentrated at major hotspots in Seoul city. In particular, occupational skill of life-long vocational education and training facilities are concentrated in major business-commercial-residence mixed areas have a significant relationship with highly educated professionals. The Gangnam area is analyzed as a place where private tutoring facilities are clustered for all age group participants. At the public level, for reducing the gap in geographical accessibility of participants in regional distribution differences in the life-long vocational education and training facilities, it is necessary to operate and provide various programs that reflect the age, education level, and job distribution of participants. It will be possible to establish a life-long vocational education system in closely related the region.
규제정책이 기업성과에 미치는 동태적 영향 분석 : 건설업을 중심으로 KCI 등재
한국무역통상학회 무역통상학회지 제24권 제2호 2024.04 pp.151-169
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
This paper analyzes how corporate financial performance is affected by changes in government policy and regulation, focusing on the construction industry. We analyze the performance of construction firms listed on the KOSPI and KOSDAQ markets from 2001 to 2022 from both financial and non-financial perspectives. In particular, construction firms listed on the KOSPI and KOSDAQ, which are closely related to the domestic real estate market, can experience significant changes in profitability depending on government policies and regulations. We use kernel density estimation to convert discrete financial indicators into continuous probability distributions to examine changes over time. To understand this more precisely, we conduct a dynamic panel analysis to examine the impact of government policy and regulatory variables on the profitability of construction firms. The empirical results show that government regulations affect the profitability of construction firms. The result of robustness analysis shows that the impacts of government policies and regulations on between KOSPI firms and KODAQ are somewhat different.
A Small Sample Text Classification Algorithm based on Kernel Density Estimation
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.8 No.6 2015.06 pp.235-244
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.6 No.1 1995 pp.85-96
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The purpose of this paper is to propose the kernel density estimator and the jackknife kernel density estimator in the presence of k's unidentified outliers, and to compare the small sample performances of the proposed estimators in a sense of mean integrated square error(MISE).
Kernel Density Estimation in the L$^{\infty}$ Norm under Dependence
[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.27 No.2 1998 pp.153-163
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
We investigate density estimation problem in the L$^{\infty}$ norm and show that the iii optimal minimax rates are achieved for smooth classes of weakly dependent stationary sequences. Our results are then applied to give uniform convergence rates for various problems including the Gibbs sampler.
Kernel Density Estimation for Polarization Measure
[NRF 연계] 경북대학교 사회과학기초자료연구소 연구방법논총 Vol.6 No.2 2021.07 pp.29-59
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this study, I propose a new estimator for the DER index that is a general measure of polarization introduced by Duclos et al.(2004). The existing estimators for the DER index are based on the empirical distribution function and a Rosenblatt-Parzen density estimator. The empirical distribution function, however, suffers from lack of smoothness. Hence, I suggest new estimators for the DER index using a new class of nonparametric kernel density estimators provided by Mynbaev and Martins-Filho(2010). I show that my estimators for polarization measure are consistent and establish asymptotic normality with a rate of convergence for the distribution of . A small Monte Carlo study reveals that my estimator performs well relative to the existing estimator for the DER index in terms of bias and mean squared error.
Adaptive Kernel Density Estimation
[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.2 No.1 1995 pp.99-111
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
It is shown that the adaptive kernel methods can potentially produce superior density estimates to the fixed one. In using the adaptive estimates, problems pertain to the initial choice of the estimate can be solved by iteration. Also, simultaneous recommended for variety of distributions. Some data-based method for the choice of the parameters are suggested based on simulation study.
Transformation in Kernel Density Estimation
[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.3 No.1 1992 pp.17-24
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The problem of estimating symmetric probability density with high kurtosis is considered. Such densities are often estimated poorly by a global bandwidth kernel estimation since good estimation of the peak of the distribution leads to unsatisfactory estimation of the tails and vice versa. In this paper, we propose a transformation technique before using a global bandwidth kernel estimator. Performance of density estimator based on proposed transformation is investigated through simulation study. It is observed that our method offers a substantial improvement for the densities with high kurtosis. However, its performance is a little worse than that of ordinary kernel estimator in the situation where the kurtosis is not high.
Distributed Channel Allocation Using Kernel Density Estimation in Cognitive Radio Networks
[Kisti 연계] 한국전자통신연구원 ETRI journal Vol.34 No.5 2012 pp.771-774
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Typical channel allocation algorithms for secondary users do not include processes to reduce the frequency of switching from one channel to another caused by random interruptions by primary users, which results in high packet drops and delays. In this letter, with the purpose of decreasing the number of switches made between channels, we propose a nonparametric channel allocation algorithm that uses robust kernel density estimation to effectively schedule idle channel resources. Experiment and simulation results demonstrate that the proposed algorithm outperforms both random and parametric channel allocation algorithms in terms of throughput and packet drops.
On Bias Reduction in Kernel Density Estimation
[Kisti 연계] 한국통계학회 한국통계학회 학술대회논문집 2000 pp.65-73
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Kernel estimator is very popular in nonparametric density estimation. In this paper we propose an estimator which reduces the bias to the fourth power of the bandwidth, while the variance of the estimator increases only by at most moderate constant factor. The estimator is fully nonparametric in the sense of convex combination of three kernel estimators, and has good numerical properties.
A Berry-Esseen Type Bound in Kernel Density Estimation for a Random Left-Truncation Model
[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.21 No.2 2014 pp.115-124
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this paper we derive a Berry-Esseen type bound for the kernel density estimator of a random left truncated model, in which each datum (Y) is randomly left truncated and is sampled if $Y{\geq}T$, where T is the truncation random variable with an unknown distribution. This unknown distribution is estimated with the Lynden-Bell estimator. In particular the normal approximation rate, by choice of the bandwidth, is shown to be close to $n^{-1/6}$ modulo logarithmic term. We have also investigated this normal approximation rate via a simulation study.
On the Plug-in Bandwidth Selectors in Kernel Density Estimation
[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.18 No.2 1989 pp.107-117
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
A stronger result than that of Park and Marron (1994) is proved here on the asymptotic distribution of the plug-in bandwidth selector. The new result is that the plug-in bandwidth selector may have the rate of convergence ($n^{-4/13}$ with less smoothness conditions on the unknown density functions than as described in Park and Marron's paper. Together with this, a class of various plug-in bandwidth selectors are considered and their asymptotic distributions are given. Finally, some ideas of possible improvements on those plug-in bandwidth selectors are provided.
ECG Denoising by Modeling Wavelet Sub-Band Coefficients using Kernel Density Estimation
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.8 No.4 2012 pp.669-684
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Discrete wavelet transforms are extensively preferred in biomedical signal processing for denoising, feature extraction, and compression. This paper presents a new denoising method based on the modeling of discrete wavelet coefficients of ECG in selected sub-bands with Kernel density estimation. The modeling provides a statistical distribution of information and noise. A Gaussian kernel with bounded support is used for modeling sub-band coefficients and thresholds and is estimated by placing a sliding window on a normalized cumulative density function. We evaluated this approach on offline noisy ECG records from the Cardiovascular Research Centre of the University of Glasgow and on records from the MIT-BIH Arrythmia database. Results show that our proposed technique has a more reliable physical basis and provides improvement in the Signal-to-Noise Ratio (SNR) and Percentage RMS Difference (PRD). The morphological information of ECG signals is found to be unaffected after employing denoising. This is quantified by calculating the mean square error between the feature vectors of original and denoised signal. MSE values are less than 0.05 for most of the cases.
[Kisti 연계] 한국수학교육학회 한국수학교육학회지시리즈E:수학교육논문집 Vol.16 2003 pp.139-147
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
While the statistical concept 'order statistics' has a great number of applications in our society ranging from industry to military analysis, it is not necessarily an easy concept to understand for many people. Adding some interesting simulation activities of this concept to the probability or statistics curriculum, however, can enhance the learning curve greatly. A hands-on and a graphic calculator based activities of a missile simulation were introduced by Kim(2003) in the context of order statistics. This article revisits the two activities in his paper and point out a dilemma that occurs from the violation of an assumption on two deviation parameters associated with the missile simulation. A third activity is introduced to resolve the dilemma in the terms of a kernel density estimation which is a nonparametric approach.
[Kisti 연계] 한국전산구조공학회 전산구조공학 Vol.32 No.3 2019 pp.173-181
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
제한된 실험 데이터로부터 확률분포함수를 추정하기 위해서 KDE가 많이 사용되고 있다. KDE에 의한 분포함수는 대역폭 선택법에 따라서 실험 데이터에 대해 평활하거나 과대적합된 커널 추정치를 생성한다. 본 연구에서는 Silverman's rule of thumb, rule using adaptive estimate, oversmoothing rule을 사용해서 각 방법에 따른 정확성과 보수적인 성향을 비교하였다. 비교를 위해서 단봉분포와 다봉분포를 가지는 실제 모델을 가정하고 통계적 시뮬레이션을 수행한 다음 다양한 데이터의 개수에 따른 추정된 분포함수의 정확도와 보수성을 비교하였다. 또한, 간단한 신뢰성 예제를 통해 대역폭 선택법에 따른 KDE의 추정된 분포가 신뢰성 해석 결과에 어떻게 영향을 미치는지 확인하였다.
To estimate probabilistic distribution function from experimental data, kernel density estimation(KDE) is mostly used in cases when data is insufficient. The estimated distribution using KDE depends on bandwidth selectors that smoothen or overfit a kernel estimator to experimental data. In this study, various bandwidth selectors such as the Silverman's rule of thumb, rule using adaptive estimates, and oversmoothing rule, were compared for accuracy and conservativeness. For this, statistical simulations were carried out using assumed true models including unimodal and multimodal distributions, and, accuracies and conservativeness of estimating distribution functions were compared according to various data. In addition, it was verified how the estimated distributions using KDE with different bandwidth selectors affect reliability analysis results through simple reliability examples.
움직임 인식응용을 위한 커널 밀도 추정 기반 학습용 데이터 증폭 기법
[Kisti 연계] 한국산업정보학회 한국산업정보학회논문지 Vol.27 No.4 2022 pp.19-27
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
머신러닝(ML, Machine Learning)기반 응용에서의 인식성능은 적용된 모델의 종류와 크기, 학습환경 및 학습에 사용되는 데이터 등 다양한 요인에 따라 결정된다. 특히 학습에 사용되는 데이터가 충분치 않을 경우 인식성능이 저하되거나 과적합(Overfitting)등의 문제가 발생하기도 한다. 이미지 인식을 주요 대상으로 하는 기존 연구들은 학습을 위한 데이터셋이 풍부하고 검증된 데이터셋을 사용하여 학습 및 인식성능을 평가할 수 있다. 하지만 사용된 센서, 인식의 대상, 인식 상황이 다른 특정 응용들의 경우 데이터셋을 직접 구축해야 한다. 이런 경우, ML모델의 성능은 데이터의 양과 품질에 따라 달라진다. 본 논문에서는 이용 가능한 학습용 데이터가 충분치 않은 움직임 인식응용에 효율적으로 사용될 수 있는 비모수 추정 방식의 일종인 커널 밀도 추정 알고리즘을 사용하여 학습용 데이터를 증폭한 후, 사용된 커널의 종류에 따라, 원본 데이터의 수 및 증폭 비율에 따라 증폭된 데이터가 원본 데이터의 특징을 잘 반영하는지 인식 정확도 변화를 토대로 비교 분석한다. 실험결과, 본 연구에서 사용한 움직임 인식응용에서는 좁은 대역폭을 가진 Tophat 커널로 증폭된 데이터셋에서 최대 14.31%의 인식 정확도 향상을 확인하였다.
In general, the performance of ML(Machine Learning) application is determined by various factors such as the type of ML model, the size of model (number of parameters), hyperparameters setting during the training, and training data. In particular, the recognition accuracy of ML may be deteriorated or experienced overfitting problem if the amount of dada used for training is insufficient. Existing studies focusing on image recognition have widely used open datasets for training and evaluating the proposed ML models. However, for specific applications where the sensor used, the target of recognition, and the recognition situation are different, it is necessary to build the dataset manually. In this case, the performance of ML largely depends on the quantity and quality of the data. In this paper, training data used for motion recognition application is augmented using the kernel density estimation algorithm which is a type of non-parametric estimation method. We then compare and analyze the recognition accuracy of a ML application by varying the number of original data, kernel types and augmentation rate used for data augmentation. Finally experimental results show that the recognition accuracy is improved by up to 14.31% when using the narrow bandwidth Tophat kernel.
가우시안 커널 밀도 추정 함수를 이용한 오토인코더 기반 차량용 침입 탐지 시스템
[Kisti 연계] 한국전기전자학회 Journal of IKEEE Vol.28 No.1 2024 pp.6-13
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 비지도학습 모델인 오토인코더와 가우시안 커널 밀도 추정 함수를 이용하여 차량용 CAN 네트워크에서 비정상적인 데이터를 탐지하는 방안을 제안한다. 제안하는 오토인코더 모델은 정상 데이터에서 CAN 프레임의 ID만으로 학습시킨다. 이후 가우시안 커널 밀도 추정 함수를 이용하여 구한 최적의 프레임 개수와 손실 임계값을 가지는 모델을 사용하여 비정상 데이터를 효과적으로 탐지한다. DoS 공격, Gear 스푸핑 공격, RPM 스푸핑 공격, Fuzzy 공격 등 4가지 공격 데이터로 오토인코더 기반 IDS를 검증하였으며 성능을 평가하였다. 기존 비지도학습 기반 모델들과 비교했을 때 우수한 성능을 나타냈으며 모든 평가 지표에서 99% 이상의 성능을 나타냈다.
This paper proposes an approach to detect abnormal data in automotive controller area network (CAN) using an unsupervised learning model, i.e. autoencoder and Gaussian kernel density estimation function. The proposed autoencoder model is trained with only message ID of CAN data frames. Afterwards, by employing the Gaussian kernel density estimation function, it effectively detects abnormal data based on the trained model characterized by the optimally determined number of frames and a loss threshold. It was verified and evaluated using four types of attack data, i.e. DoS attacks, gear spoofing attacks, RPM spoofing attacks, and fuzzy attacks. Compared with conventional unsupervised learning-based models, it has achieved over 99% detection performance across all evaluation metrics.
커널 밀도 추정을 이용한 Fuzzy C-means의 초기 원형 설정
[Kisti 연계] 한국컴퓨터정보학회 한국컴퓨터정보학회 학술대회논문집 2011 pp.85-88
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Fuzzy C-Means (FCM) 알고리듬은 가장 널리 사용되는 군집화 알고리듬 중 하나로 다양한 응용 분야에서 사용되고 있다. 하지만 FCM은 여러 가지 문제점을 가지고 있으며 초기 원형 설정이 그 중 하나이다. FCM은 국부 최적해에 수렴하므로 초기 원형 설정에 따라 클러스터링 결과가 달라진다. 이 논문에서는 이러한 FCM의 초기 원형 설정 문제를 개선하기 위하여 커널밀도 추정 (kernel density estimation) 기법을 활용하는 방법을 제안한다. 제안한 방법에서는 먼저 커널 밀도 추정을 수행한 후 밀도가 높은 지역에 클러스터의 초기 원형을 설정하고 원형이 설정된 영역의 밀도를 감소시키는 과정을 반복함으로써 효율적으로 초기 원형을 설정할 수 있다. 제안된 방법이 일반적으로 사용되는 무작위 초기화 방법에 비해 효율적이라는 사실은 실험결과를 통해 확인할 수 있다.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.