년 - 년
시간적 요소를 활용한 교통량 이상치 및 결측치 보정 모델 KCI 등재
한국ITS학회 한국ITS학회논문지 제24권 제3호 통권119호 2025.06 pp.37-52
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
본 연구는 월(Month), 시간(Hour), 휴일(Day off) 여부를 종합한 복합 시간적 요소(Temporal Factors)를 활용하여, 실시간 교통 데이터의 이상치와 결측치를 정밀하게 처리하는 보정 모델을 제안한다. 모델은 이 시간적 요소로 데이터를 그룹화 후 그룹 내 Z-score로 이상치를 탐지하며, 결측치는 시간적 요소 그룹 내 평균 기반 단계적 보간 방식을 결합한 파이프라인을 구성한다. 모델의 성능을 검증하기 위해 인천시 1,569개 도로의 교통량 데이터를 기반으로, 실무에서 널 리 쓰이는 기법들과 비교 평가를 수행했다. 그 결과, 복원 및 예측 정확도 실험 모두에서 제안 모델이 다른 기법 조합들보다 통계적으로 유의미하게 우수한 성능을 보이는 것을 확인했다. 이는 계절성, 일별 주기, 휴일 등 복합적 시간 요소를 반영하는 것이 예측 정확도 향상에 매우 효과적임을 입증하며, 실시간 데이터 전처리를 위한 본 모델의 높은 실용적 가치를 시사한다.
This study proposes an integrated correction method that effectively handles outliers and missing values in real-time traffic data, using data from 1,569 roads in Incheon between 2022 and 2024. The proposed method first removes outliers empirically, then constructs an integrated pipeline by combining "hourly Z-score" with "hourly average imputation." To validate this approach, we assembled 35 models by combining seven outlier-detection techniques and five missing-value imputation methods, including those commonly used in practice. We then conducted experiments involving artificially generated outliers and missing values, as well as performance comparisons using an LSTM prediction model. The results demonstrate that the proposed method outperforms all other combinations in both verification tests. This suggests that a simple, statistically based preprocessing strategy incorporating hourly characteristics is highly effective for improving urban traffic flow forecasts and has significant potential for real-time environments.
한국ITS학회 한국ITS학회 학술대회 2004년 한국ITS학회 정기총회 및 추계학술대회 2004.11 pp.270-275
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
누락교통량자료 보정방법에서 강우의 영향 고려 KCI 등재
한국ITS학회 한국ITS학회논문지 제14권 제2호 통권58호 2015.04 pp.1-13
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
교통량자료는 매우 다양한 분야에서 사용되는 기초자료이다. 교통량자료는 도로교통량조사를 통하여 수집되며, 도로교통량조사 중 기계식 장비를 사용하여 365일 24시간 지속적으로 수집되는 자료를 상시교통량자료라고 한다. 상시교통량자료는 장비의 오작동 및 여러 원인으로 교통량자료누락이 발생하는 경우가 있다. 누락된 교통량자료는 여러 누락보정방법을 적용하여 보정을 수행하고 있다. 하지만, 기존의 누락보정방법론들은 기상에 대한 영향을 전혀 고려하지 않은 실정이다. 따라서 본 연구에서는 기상 중 강우의 영향을 고려한 누락교통량자료 보정방법에 대한 연구를 수행하였다. 이를 위해 우선 일반국도에서 수집한 교통량자료와 기상청의 기상자료의 매칭을 수행하였으며, 이후 일반국도의 특성별로 군집분석 수행 및 분석대상지점 선정을 진행하였다. 세 가지 보정 기법들(평균대체법/자기회귀모형/EM 기법)을 사용하여 전체 자료에서 누락보정을 수행하는 것과 강우일의 자료만을 가지고 누락보정을 수행하여 보정값의 정확도를 평가하였다. 분석 결과 모든 보정방법 및 분석지점에서 과거 강우일의 교통량자료만을 가지고 보정한 경우가 더 정확한 보정값을 산출하는 것으로 분석되었다.
Traffic volume data is basic information that is used in a wide variety of fields. Existing missing traffic volume data imputation method did not take the effect on the rainfall. This research analyzed considering of the rainfall effect in missing traffic volume data imputation method. In order to consider the effect of rainfall, established the following assumption. When missing of traffic volume data generated in rainy days it would be more accurate to use only the traffic volume data of the past rainy days. To confirm this assumption, compared for accuracy of imputed results at three kinds of imputation method(Unconditional Mean, Auto Regression, Expectation-Maximization Algorithm). The analysis results, the case on consideration of the rainfall effect was more low error occurred.
하천 수온 측정 자료의 결측 유형별 보간 기법 적용 특성 비교 KCI 등재
한국습지학회 한국습지학회지 제27권 제4호 2025.12 pp.348-356
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
다양한 수질 측정자료가 오염원 추적 및 수질 환경 평가 등에 널리 활용되고 있으며, 이에 따라 현장 모니터링을 통한 수질자료 취득을 위한 노력이 계속되고 있다. 하지만 현장 수질 모니터링의 특성상 센서 오류, 고장 및 강우에 따른 재해 등 다양한 원인에 따른 결측이 발생할 수 있으며 수질 모니터링 결과의 신뢰도 확보를 위한 결측 관리의 중요성이 커지고 있다. 본 연구에서는 현장의 수질 특성을 확인할 수 있는 수질 환경 변수 중 하나인 수온에 대하여 4가지 유형의 결측을 생성하고, 2개의 통계기반 보간 기법인 선형 보간 (Linear)과 다항 보간 (Polynomial) 그리고 2개의 머신러닝 기반 모형인 K-Nearest Neighbors (KNN) 및 autoencoder (AE)를 적용한 총 4개의 보간 모형을 적용하여 성능을 비교하였다. 4개의 결측 유형은 단기 결측 (Case 1), 장기결측 (Case 2), 첨두 구간 전후의 급격한 수질변화 구간의 결측 (Case 3), 첨두 및 저점을 포함한 장기간의 수질변화 구간의 결측 (Case 4)으로 구분되었다. 분석결과 단기 결측이 발생되는 Case 1 및 3에서는 Linear 모형이 RSR 0.26 및 0.76으로 가장 우수한 보간 성능을 보였으며, 장기간의 결측을 포함하는 Case 2이 경우 AE가 RSR 0.63으로 가장 우수한 성능을 보이는 것을 확인하였다. Case 4는 KNN (k=3)의 RSR 이 0.66으로 가장 우수한 성능을 보였으며, AE의 RSR이 0.68로 KNN에 비해 다소 낮은 성능을 보였지만 그 차이는 크지 않았다. 본 연구를 통해 결측 유형에 따라 보간 모형의 성능에 차이가 있음을 확인할 수 있었다.
Various water quality measurements are used to track pollution sources and assess water environments. As such, efforts to collect water quality data through field monitoring continue to expand. However, due to the nature of field monitoring, missing values are often observed as a result of sensor errors, equipment failures, and external factors such as rainfall or disasters. This highlights the growing importance of managing missing data to ensure the reliability of water quality monitoring results. This study generated four types of missing patterns for water temperature, a key indicator of field water quality conditions. Then, four imputation methods were applied. The methods included two traditional statistical approaches (linear interpolation and polynomial interpolation) and two machine learning models (K-nearest neighbors (KNN) and autoencoder (AE)). The four missing data scenarios were defined as follows: short-term missing (Case 1), long-term missing (Case 2), missing around peak values with rapid water quality change (Case 3), and extended missing periods including both peaks and troughs (Case 4). The results showed that the linear model achieved the best performance for Cases 1 and 3, with RSR values of 0.26 and 0.76, respectively. For Case 2, AE achieved the highest performance with an RSR of 0.63. In Case 4, KNN (k=3) showed the best result with an RSR of 0.66, followed closely by AE with an RSR of 0.68. These findings indicate that imputation performance varies depending on the missing data pattern.
공분산 구조모형의 적합도 평가에 있어서 결측치 처리 방법 비교: 완전정보최대우도, 다중대체, 베이지안 접근법을 중심으로
[NRF 연계] 한국심리학회 한국심리학회지: 일반 Vol.33 No.2 2014.06 pp.507-533
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
통계 모형을 이용한 데이터 분석에서 적절한 결측치 처리가 최종적인 통계적 결론 도출에 결정적인 영향을 미칠 수 있다는 것은 이미 잘 알려져 있다. 따라서 현재까지 다양한 결측치 처리 기법이 제안되어왔고, 그 중 최근 활발히 연구되고 있는 기법으로서 완전정보최대우도(full information maximum likelihood), 다중대체(multiple imputation), 그리고 베이지안(Bayesian) 접근법이 있다. 이 중 완전정보최대우도법은 사회과학 통계에서 가장 많이 사용되는 방법으로서 많은 심리학 연구자들에게도 이미 알려져 있는 방법이나, 다른 두 방법, 즉 다중대체 및 베이지안 접근법은 아직은 많은 심리학 연구자들에게 생소한 방법이다. 따라서 본 논문에서는 이 세 가지 방법을 공분산 구조모형의 맥락에서 소개하고 주요 특징을 비교함으로써 심리학 연구자들의 결측치 처리에 관한 이해를 돕고자 하였다. 공분산구조모형의 적용 과정은 다른 통계모형과 달리 자료와 모형간의 적합도 지수를 계산하고 이를 바탕으로 모형의 적절성을 판단하는 과정을 포함한다. 최대우도법에서는 카이자승 통계량을 기준으로 다양한 적합도 지수가 파생되어 제안되었으며, 다중대체법은 D2와 D3 통계량이, 그리고 베이지안 접근법에서는 사후예측모형검증(posterior predictive model checking) 기법이 사용된다. 최대우도법의 카이자승 통계량은 다양한 맥락에서 독립적으로 연구된 결과가 있으나, 다중대체법 및 베이지안 접근법에서 제안된 방법은 공분산 구조모형의 맥락에서 그 성능이 평가된 적이 없으며 서로 비교된 적도 없다. 따라서, 본 논문에서는 동일한 구조방정식모형의 자료-모형간 적합성 판단에 있어서 세 가지 다른 결측치 처리 방법이 어떤 영향을 미치는지 비교 분석하였다. 구체적으로, 본 논문에서는 모의실험(simulation) 기법을 사용하여, 모집단에서 데이터가 종단적 측정 불변성을 지지하지 않는 경우(longitudinal measurement non-invariance)를 가정하고, 옳은 모형을 사용한 경우의 제 1 종 오류율과 부분측정불변성(partial measurement invariance)을 가정한 모형을 적용한 경우의 검정력을 최대우도, 다중대체, 그리고 베이지안 접근법 별로 추정/비교하였다. 본 연구의 결과는 결측치가 존재하는 데이터를 공분산 구조모형을 이용하여 분석하고자 하는 연구자들에게 결측치 처리 기법의 특성을 이해하고 최종 모형의 적합성을 판단하는데 적절한 지침을 제공할 것으로 기대된다.
In practical applications of any statistical modeling, including structural equation modeling(SEM), virtually every data set contains missing values. It is a well known fact that improper handling of missing data can exert harmful impact on subsequent statistical inferences in a variety of ways to varying degrees. In the context of SEM, the full information maximum likelihood(FIML) has been arguably the most popular method for addressing missing data. Despite of being yet less widely known to majority of applied researchers as flexible alternatives to FIML, multiple imputation (MI) procedures and Bayesian approaches have recently begun to emerge as viable solutions among many applied researchers. An important objective of this article is to introduce these methods to applied researchers in an accessible manner using SEM as the context. Structural equation modeling actually involves the process of proposing, estimating, and evaluating the researcher’s hypothesis that is believed to be underlying and purported in generating the observed data. Therefore, it is essential to evaluate the overall goodness-of-fit of the posited model in any given application. FIML, MI and Bayesian approaches, respectively, yield the chi-square, , , and the posterior predictive modeling checking (PPMC) p-value as statistical tools for the assessment of data-model fit. Another important objective of this article is to study performance of these model evaluation tools in the context of SEM. Further, relative performance of these data-model fit assessment tools is to be evaluated with respect to their Type I error rates and power. The performance of these assessment tools, except the chi-square statistics, has never been evaluated nor been compared within the context of SEM. The initial results provided in the present article is believed to not only enhance the knowledge base regarding the characteristics of these assessment tools under missing data, but also provide an initial guideline for the proper use of these assessment tools in the real-world data analysis especially in the application of SEM with missing data.
적응형 k-NN 기법을 이용한 UTIS 속도정보 결측값 보정처리에 관한 연구 KCI 등재
한국ITS학회 한국ITS학회논문지 제13권 제3호 통권53호 2014.06 pp.66-77
※ 기관로그인 시 무료 이용이 가능합니다.
4,300원
UTIS(Urban Traffic Information System)는 프로브차량을 활용하여 도시지역의 구간통행시간 정보를 직접 수집하는 방식으로 타 검지체계에 비해 상대적으로 정확한 링크 속도정보를 산출할 수 있다. 하지만, 현재 UTIS에서는 프로브차 량(Probe Vehicle) 및 노변기지국(RSE)의 부족, 시스템 오류 등 다양한 요인에 의해 링크 속도정보의 수집이 누락되는 결측 구간이 발생되고 있다. 본 연구에서는 보다 정확한 여행시간 정보를 제공하기 위한 방안으로 k-NN 알고리즘을 기 반으로 결측속도 정보를 효율적으로 보정할 수 있는 새로운 보정모형을 제안하였다. 제안 모형은 각 후보개체(이력 시 계열 데이터)의 분포 특성에 따라 최근접이웃 개수를 탄력적으로 조정하는 적응형 k-NN 모형이다. 모형 평가 결과, 제 안 모형이 결측정보를 효과적으로 보정‧처리할 수 있는 동시에 ARIMA 등 타 모형에 비해 보정 오차를 크게 감소시킬 수 있는 것으로 분석되었다. 본 연구에서 제안된 결측 보정 모형은 UTIS 중앙교통정보센터에 직접 적용하여 교통정보 서비스 품질을 향상시키데 활용될 계획이다.
UTIS(Urban Traffic Information System) directly collects link travel time in urban area by using probe vehicles. Therefore it can estimate more accurate link travel speed compared to other traffic detection systems. However, UTIS includes some missing data caused by the lack of probe vehicles and RSEs on road network, system failures, and other factors. In this study, we suggest a new model, based on k-NN algorithm, for imputing missing data to provide more accurate travel time information. New imputation model is an adaptive k-NN which can flexibly adjust the number of nearest neighbors(NN) depending on the distribution of candidate objects. The evaluation result indicates that the new model successfully imputed missing speed data and significantly reduced the imputation error as compared with other models(ARIMA and etc). We have a plan to use the new imputation model improving traffic information service by applying UTIS Central Traffic Information Center.
4,000원
현재 여러 지자체에서 혼잡한 도시교통의 이동성 및 안전성을 향상시키기 위해 첨단교통관리체계(ITS)를 구축 운영중인데 이러한 시스템에서 수집하는 교통상황에 대한 실시간 자료가 노면상황, 악천후, 통신 및 장비자체의 결함 등으로 인해 수많은 자료가 결측된다. 이러한 결측 자료로 인해 통행시간 예측 및 각종 연구가 불가능한 경우가 발생하며 또한 도로의 계획과 기하구조 설계시 기본 자료가 되는 AADT 및 DHV 등의 교통 파라메터들이 과소 또는 과대 추정될 수 있어서 심각한 손해를 끼칠수 있다. 따라서 본 연구에서는 부득이하게 누락되는 교통량 자료에 대해 전 후기간 평균, 회귀 모형, EM, 시계열 모형들을 활용한 대체기법들을 살펴보았고, 그 결과 시계열 모형을 이용한 대체의 경우 MAPE, 불균등계수, RMSE 가 각각 5.0, 0.030, 110으로 가장 좋은 결과를 보였고 나머지 대체기법들은 평가지표에 따라 조금씩 다른 결과를 보였으나 대체로 만족할 만한 수준의 결과를 낳았다
There are many cities installing ITS(Intelligent Transportation Systems) and running TMC(Trafnc Management Center) to improve mobility and safety of roadway transportation by providing roadway information to drivers. There are many devices in ITS which collect real-time traffic data. We can obtain many valuable traffic data from the devices. But it's impossible to avoid missing traffic data for many reasons such as roadway condition, adversary weather, communication shutdown and problems of the devices itself. We couldn't do any secondary process such as travel time forecasting and other transportation related research due to the missing data. If we use the traffic data to produce AADT and DHV, essential data in roadway planning and design, We might get skewed data that could make big loss. Therefore, He study have explored some imputation techniques such as heuristic methods, regression model, EM algorithm and time-series analysis for the missing traffic volume data using some evaluating indices such as MAPE, RMSE, and Inequality coefficient. We could get the best result from time-series model generating 5.0, 0.03 and 110 as MAPE, Inequality coefficient and RMSE, respectively. Other techniques produce a little different results, but the results were very encouraging.
한국ITS학회 한국ITS학회 학술대회 2003년 한국ITS학회 정기총회 및 추계학술대회 2003.11 pp.221-225
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Missing Data Imputation Based on Grey System Theory
보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.7 No.2 2014.03 pp.347-356
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
This paper proposed a new weighted KNN data filling algorithm based on grey correlation analysis (GBWKNN) by researching the nearest neighbor of missing data filling method. It is aimed at that missing data is not sensitive to noise data and combined with grey system theory and the advantage of the K nearest neighbor algorithm. The experimental results on six UCI data sets showed that its filling accuracy is better than the traditional method of K nearest neighbor and filling algorithm presented by Huang and Lee.
보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.9 No.2 2015.02 pp.211-218
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
As our goal, we are interested in estimating the degree of software reliability based on software development project data. It is widely-known that several software development attributes which are measured can be used to evaluate and predict software reliability/quality via multi-variable analyses. In this article, we focus on the data treatment method which is needed prior to the software reliability assessment, since the software development data sets often include missing data. This paper discusses the method of data preparation against missing data and their effectiveness by using the Random Forest as a multi-variable analysis.
결측값 대체를 위한 데이터 재현 기법 비교 KCI 등재
국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.10 No.1 2024.01 pp.603-608
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
무응답 및 결측값은 표본 탈락, 설문조사에 대한 답변 회피 등으로 발생하며 정보의 손실 및 편향된 추론의 가 능성이 있는 문제가 발생하게 되며, 이 경우 결측값을 적절한 값으로 바꾸는 대체가 필요하게 된다. 본 논문에서는 결측값에 대한 대체 방법으로 제안되었던 평균 대체, 다중회귀 대체, 랜덤 포레스트 대체, K-최근접 이웃 대체, 그리 고 딥러닝을 기본으로 한 오토인코더 대체와 잡음제거 오토인코더 대체 방법을 비교한다. 결측값을 대체하는 이러한 방법들에 대해 설명하고, 연속형의 모의실험 데이터와 실제 데이터에 접목시켜 각 방법들을 비교하였다. 비교 결과 대부분의 경우에서 다중 대체 방법인 랜덤 포레스트 대체 방법과 잡음제거 오토인코더 대체 방법의 성능이 좋았음을 확인하였다.
Nonresponse and missing values are caused by sample dropouts and avoidance of answers to surveys. In this case, problems with the possibility of information loss and biased reasoning arise, and a replacement of missing values with appropriate values is required. In this paper, as an alternative to missing values imputation, we compare several replacement methods, which use mean, linear regression, random forest, K-nearest neighbor, autoencoder and denoising autoencoder based on deep learning. These methods of imputing missing values are explained, and each method is compared by using continuous simulation data and real data. The comparison results confirm that in most cases, the performance of the random forest imputation method and the denoising autoencoder imputation method are better than the others.
Multiple Imputation for Missing Data in the KLoSA Study
[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.9 No.5 2007.10 pp.2085-2095
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Most survey data include missing values due to nonresponse. Especially, sensitive questions such as income or assets tend to show higher percentage of missing values. When missing values occur, complete-case analysis may lead to biased estimates in parameters. Korean Longitudinal Study of Aging(KLoSA) is a longitudinal study to evaluate aging trends in the Korean population and apply the results to the social welfare and labor policy. In 2006, KLoSA collected baseline data. We conduct multiple imputation based on hotdeck to handle missing values in the KLoSA baseline data. In this study, we explain the imputation method for filling in missing values and discuss the results of imputation.
[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.26 No.4 2019 pp.411-430
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Tumor development is driven by complex combinations of biological elements. Recent advances suggest that molecularly distinct subtypes of breast cancers may respond differently to pathway-targeted therapies. Thus, it is important to dissect pathway disturbances by integrating multiple molecular profiles, such as genetic, genomic and epigenomic data. However, missing data are often present in the -omic profiles of interest. Motivated by genomic data integration and imputation, we present a new statistical framework for pathway significance analysis. Specifically, we develop a new strategy for imputation of missing data in large-scale genomic studies, which adapts low-rank, structured matrix completion. Our iterative strategy enables us to impute missing data in complex configurations across multiple data platforms. In turn, we perform large-scale pathway analysis integrating gene expression, copy number, and methylation data. The advantages of the proposed statistical framework are demonstrated through simulations and real applications to breast cancer subtypes. We demonstrate superior power to identify pathway disturbances, compared with other imputation strategies. We also identify differential pathway activity across different breast tumor subtypes.
Parametric fractional imputation for nonignorable missing data
[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.41 No.3 2012 pp.291-303
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Parameter estimation with missing data is a frequently encountered problem in statistics. Imputation is often used to facilitate the parameter estimation by simply applying the complete-sample estimators to the imputed dataset. In this article, we consider the problem of parameter estimation with nonignorable missing data using the approach of parametric fractional imputation proposed by Kim (2011). Using the fractional weights, the E-step of the EM algorithm can be approximated by the weighted mean of the imputed data likelihood where the fractional weights are computed from the current value of the parameter estimates. Calibration fractional imputation is also considered as a way for improving the Monte Carlo approximation in the fractional imputation. Variance estimation is also discussed. Results from two simulation studies are presented to compare the proposed method with the existing methods. A real data example from the Korea Labor and Income Panel Survey (KLIPS) is also presented.
A Study on Imputation for Missing Data using the Kriging
[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.17 No.6 2015.12 pp.2857-2866
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Geostatistics is a branch of statistical theory concerned with problems of spatial serial data, interpolation and mapping of distributed data, and related problems. The missing data are produced in many fields concerned with geostatistical data. Generally, in case of continuous data, we use the average as a method for imputation of missing data. However, using an average is no longer a good method to impute for the geostatistical missing data as there exist a spatial autocorrelation among them. In this study, we suggest the ‘kriging’ that is spatial statistics method as a method to impute for those spatial missing data, and then compare it with general average method in aspect to imputation superiority. To show the usefulness of the proposed method, we perform a small simulation study based on the PRESS statistic and show an empirical example with rainfall data from the Philippines.
Application of SOLAS to the Multiple Imputation for Missing Data
[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.14 No.3 2003 pp.579-590
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
When we analyze incomplete data, i.e., data with missing values, we need treatment for the missing values. A common way to deal with this problem is to delete the cases with missing values. Various other methods have been developed. Among them are EM algorithm and regression algorithm which can estimate missing values and impute the missing elements with the estimated values. In this paper, we introduce multiple imputation software SOLAS which generates multiple data sets and imputes with them.
Application of NORM to the Multiple Imputation for Multivariate Missing Data
[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.13 No.2 2002 pp.105-113
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The statistical analysis of incomplete data sometimes requires handling of incomplete observations. Towards this end, each case with some missing values generally should be deleted, namely, resulting in only use of non-missing cases. EM algorithm(Dempster et al., 1977) which involves prediction and estimation steps is a general method among others. In this article, we use the free software NORM developed for multiple imputation, which uses DA(Data Augmentation) algorithm in its imputation, and evaluate its efficiency through a numerical example.
Application of Multiple Imputation Method in Analyzing Data with Missing Continuous Covariates
[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.21 No.4 2008 pp.659-664
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Missing continuous covariates are pervasive in the use of generalized linear models for medical data. Multiple imputation is the most common and easy-to-do method of dealing with missing covariate data. However, there are always serious warnings in using this method. There should be concern to make imputed values more proper. In this paper, proper imputation from posterior predictive distribution is developed for implementing with arbitrary priors. We use empirical distribution of the posterior for approximating the posterior predictive distribution, to sample from it. This method is preferable in comparison with a presented imputation method of us which uses a full model to impute missing values using available software. The proposed methods are implemented on glucocorticoid data.
Bias Reduction by Imputation for Linear Panel Data Models with Nonrandom Missing
[NRF 연계] 한국계량경제학회 계량경제학보 Vol.29 No.1 2018.03 pp.1-25
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
When no variables are observed for endogenous non-respondents of panel data, bias correction is available only for a limited class of instrumental variable estimators, which require strong conditions for consistency and often suffer from substantial efficiency loss. In this paper we examine a convenient alternative method of imputing the missing explanatory variables and then using standard bias-correction procedures for sample selection. Various bias-corrected estimators are derived and their performances are compared by Monte Carlo experiments. Results verify efficiency loss by the instrumental variable estimators and suggest that the imputation method is practically useful if it is applied to first-difference regression.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.