Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 112
No
1

본 논문에서는 데이터마이닝에서 분류·예측 모형의 성과를 높이기 위하여 기존의 변수들의 상호작용효과를 고려한 새로운 변수를 추가는 방법에 대하여 기술한다. 이를 위하여 연관분석을 통하여 종속변수의 값을 성공으로 만드는 데에 유사한 행위를 하는 변수들을 골라내고, 그 변수들의 값을 곱으로 하는 새로운 변수를 추가하여 모형을 구축해 보았다. 주식시장의 이상매매데이터에 적용한 결과, 새로운 모형은 기존의 모형과 비교하여 나은 성과를 나타내었다.

This paper introduces a noble method of variable selection utilizing interaction effects between variables to enhance performances of classification/prediction models in data mining. The proposed method utilizes the association of data mining and finds variables which affect the target variable in some way. Those new variables are further screened, multplied to form interaction effect, and finally input to the model. As a result, the model added with the new variables performed better than compared models.

2

4,800원

본 연구에서는 공공기관에서 제공하고 있는 정량적 자료를 이용하여 자연재해로 인한 피해액을 추 정하기 위한 회귀분석 모형을 제안한다. 이를 위해, 공공 부문에서 225개 기초지방자치단체에 대해 2001년부터 2014년까지의 자연재해 세부지표 값과 피해액 자료를 수집하여 회귀분석을 수행한다. 본 연구에서는 회귀모형이 사용하는 독립변수의 개수를 줄이기 위해 변수선정 방법론을 활용한다. 이를 통해, 기존 자연재해 위험지표에서 사용하던 32개 세부지표 중 주요 변수만을 선정하여 14개 독립변수로 구성된 자연재해 피해액 추정 모형을 정의하였다. 제안된 모형의 유효성을 검증하기 위해, 모형을 이용하여 추정된 피해액 예측치, 자연재해 위험지표 평가결과, 국민안전처 지역안전도 평가결과, 실제 피해액을 비교⋅분석하였다. 분석 결과, 본 연구에서 도출된 피해액 예측치가 기존 자연재해 위험지표와 지역안전도에 비해 자연재해 피해액과 더 높은 상관관계를 갖는 것을 확인할 수 있었다. 본 연구에서 제안하고 있는 회귀분석 기반 자연재해 피해액 추정 모형을 활용하면 향후 자연재해 발생 양상의 변화로 인한 자연재해 피해액의 변화를 미리 예측할 수 있으며, 제안된 회귀 분석 모형의 회귀계수를 가중치로 활용하면 기존 자연재해 위험지표의 성능 및 정확도를 향상시킬 수 있을 것이다.

In this study, we propose a regression model for estimating damage amounts of natural disasters based on the quantitative data provided by public institutions. Based on quantitative data for 225 primary local governments during the years of 2001 to 2014, we define a regression model to have 14 independent variables by applying variable selection methods to the variables from 32 sub-indexes used in Natural Disaster Risk Index(NDRI). For verifying the proposed regression model, we compare the estimates from the model with NDRI, Regional Safety Grades(RSG), and actual amounts of damages from natural disasters. From the analysis results, we can find the regression model give better estimates for the amounts of damages than NDRI and RSG. If we utilize the proposed model, we can estimate the future changes in the amounts of damages using the changes in natural disasters and improve the performance of NDRI using the standard regression coefficients from the regression analysis.

3

서울 아파트매매가격지수 예측을 위한 베이지안 변수선택 기법

이창훈, 강규호, 안지희

[NRF 연계] 한국경제학회 경제학연구 Vol.68 No.1 2020.03 pp.153-190

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

정확한 장단기 주택가격 예측은 효율적인 부동산 정책수립과 운용, 부동산 투자및 은행의 담보대출 위험관리 등에 필수적임에도 불구하고 잠재적인 예측변수의수가 너무 많아 예측을 실시하는 데 기술적 어려움이 존재한다. 본 연구는 이러한예측변수와 모형 불확실성 문제를 고려함으로써 보다 정확한 장단기 서울 아파트매매가격지수 예측분포를 도출하고자 기존의 베이지안 변수선택 기법을 보완 및확장한 새로운 변수선택 알고리즘을 제안한다. 본 연구에서 제시된 변수선택기법은 각 예측변수의 선택여부 뿐만 아니라 종속변수에 미치는 영향의 방향까지 식별한다는 점에서 기존 기법과 차별화된다. 표본 외 예측결과, 중장기 예측에서 본연구의 베이지안 변수선택 기법을 적용한 모형이 다른 모든 경쟁모형들보다 예측력에서 우월한 것으로 나타났다. 반면 서울 아파트매매가격지수의 높은 지속성으로 인해 단기 예측에서는 p차 자기회귀모형의 예측력이 가장 우수했다.

Accurate house price forecasts are essential for efficient policy-making, investment, and risk management of mortgage loan. Nevertheless, there are few empirical studies on the Korean house price prediction. This seems to be because of the large number of variables. In this study, we provide a new Bayesian variable selection method for the Seoul apartment price index forecasting, considering the uncertainty of the variables. To do this, we extend the standard Bayesian variable selection by using a more flexible and interpretable spike-and-slab prior. This method consists of two stages: variable selection and predictive density simulation. According to our out-of-sample forecasting experiment, the proposed model outperforms the standard variable selection method and first-order auto-regressive (AR(p) model in the medium- and long-term horizons. Meanwhile, the AR(p) is found to be the best for the short-term forecasting due to the high persistence of the price index growth.

4

범죄 및 피해자 특성과 범죄피해 내용의 관계 탐색: 랜덤포레스트 알고리즘에 기초한 변인선택

한유화, 이우열

[NRF 연계] 한국법심리학회 한국심리학회지:법 Vol.13 No.2 2022.07 pp.121-145

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 범죄 및 피해자 특성과 범죄피해 내용의 관련성을 확인하기 위하여 2010년부터 2018년까지 격년으로 수집된 전국범죄피해조사 자료에 랜덤포레스트 알고리즘을 적용하였다. 전체 자료 중 범죄피해경험이 있는 사례 및 관심 변인을 선별하여 분석자료를 구성하였으며, 총 3080건 자료의 성별, 연령(생애주기단계), 범죄유형, 가해자 면식여부, 반복피해 여부, 심리적 피해내용(우울함, 고립감, 극심한 두려움, 신체증상, 대인관계 문제, 사람을 피해 이사, 자살 충동, 자살 시도) 및 범죄피해 후 감정변화(자기보호 자신감, 자존감, 타인에 대한 신뢰감, 사법기관에 대한 신뢰감 및 사법제도와 법에 대한 존중감의 변화)를 나타내는 변인들이 분석자료에 포함되었다. 전통적 통계기법을 적용하기 어려운 자료의 특성을 고려하여, 본 연구는 범죄피해 내용(심리적 피해내용과 감정변화)을 이용하여 범죄 및 피해자 특성을 예측하기 위한 랜덤포레스트 알고리즘을 다섯 번 실행하고, VSURF 함수를 이용하여 범죄 및 피해자 특성을 잘 예측하는 범죄피해 내용 변인들을 선택하였다. 분석 결과, 범죄유형과 우울함, 극심한 두려움 및 신체증상의 관련성, 가해자 면식여부와 신체증상 및 대인관계 문제의 관련성, 반복피해 여부와 사법제도와 법에 대한 존중감 변화의 관련성이 확인되었다. 성별과 생애주기단계(청소년/성인/노인)는 각각 극심한 두려움과 자기보호 자신감 변화와 관련이 있는 것으로 확인되었으나 의미를 부여하기 위해서는 추가적 경험자료가 필요할 것으로 판단되었다. 본 연구의 결과는 범죄피해평가제도의 실효성을 높이기 위해 전문가 교육과정에 범죄 및 피해자 특성과 범죄피해 내용에 관한 지식과 사례교육의 제공 및 면담전략과 법률지식에 관한 교육 강화가 필요함을 시사한다.

The current study applied the random forest algorithm to Korean crime victim survey data collected biennially between 2010 and 2018 to explore the relationship between crime/victim characteristics and the victim’s criminal damages. A total of 3,080 cases including gender, age (life cycle stage), type of crime, perpetrator acquisition, repeated victimization, psychological damage (depression, isolation, extreme fear, somatic symptoms, interpersonal problems, moving out to avoid people, suicidal impulses, suicide attempts), and emotional changes after victimization (changes in self-protection confidence, self-esteem, confidence in others, confidence in legal institutions, and respect for Korean legal system/law) were analyzed. Considering the features of data that are difficult to apply traditional statistical techniques, this study implemented random forest algorithms to predict crime and victim characteristics using the victim’s criminal damages (psychological damage and emotional change) and selected good predictors using VSURF function in VSURF package for R. As a result of the analysis, it was confirmed that the relationship between the type of crime and depression, extreme fear, somatic symptoms, and interpersonal problems, between perpetrator acquisition and somatic symptoms and interpersonal problems, and between repeated victimization and changes in respect for Korean legal system/law. Gender and life cycle stage (youth/adult/elderly) were found to be related to extreme fear and changes in self-protection confidence, respectively. However, more empirical evidence should be aggregated to explain the results as meaningful. The results of this study suggest that it is necessary to enhance the experts’ knowledge and educate them on cases about the relationship between crime/victim characteristics and criminal damage. Strengthening their interview strategy and knowledge about law/rules were also needed to increase the effectiveness of the Korean victim assessment system.

5

4,600원

본 연구는 기업 규모별 수출액에 영향을 미치는 무역원활화와 교역국의 소득 수준의 하위 변수를 대상으로 변수선택 분석을 진행하였으며, 단계적 회귀와 Lasso 회귀분석을 각기 적용 하여 그 결과를 비교하였다. 종속변수는 대기업과 중견기업, 중소기업별 수출액, 독립변수는 OECD의 무역원활화지표(TFIs)의 11개 하위 차원과 교역국의 소득 수준을 저소득, 중저소득, 고중소득, 고소득 국가로 더미 변수화하였다. 주요 결과로 단계적 회귀와 Lasso 회귀분석에서 공통으로 선택된 변수는 중소기업 수출액의 경우, 정보 가용성, 무역공동체 참여, 내부협력, 중견기업 수출액은 정보 가용성, 무역공동체 참여, 사전심사, 자동화, 거버넌스와 공정성, 대기업 수출액은 정보 가용성, 내부협력으로 나타났다. 이 외에도 교역국의 소득 수준은 단계적 회귀분석에서 중소기업의 수출액에 대해서만 고중소득, 중저소득, 저소득 국가가 선택되었고 Lasso 회귀분석의 결과가 단계적 회귀분석이 비해 최적 변수의 수가 감소하여 상대적으로 간결한 모형을 제안하는 것으로 나타났다.

This study conducted an analysis to identify the vriables among the sub-dimensions of the OECD’s Trade Facilitation Indcx and the inocme level of trading countries, categorized as low, lower-middle, upper-middle, and high income, that affect the export amounts differentiated by firm size: large, medium sized, and small-to-medium sized enterprises. To achieve this, both Stepwise Regression and Lasso Regression analyses were applied, with the export amount serving as the dependent variable, and the results from both methods were compared. The primary findings revealed that the variables commonly selected by both Stepwise Regression and Lasso Regression analyses include information availability, participation in the trading community, and internal cooperation for the exports of small and medium sized enterprises. For large corporations, the variables determining the export amount were identified as information availability and internal cooperation. Additionally, the Stepwise Regression analysis on the income level of trading countries identified that exports of small and medium sized enterprises are significantly associated with upper-middle, lower-middle, and low income countries. Contrastingly, the Lasso Regression analysis indicated a reduction in the number of optimal variables compared to the Stepwise Regression analysis, suggesting a more concise model. This comparative analysis underscores the importance of specific trade facilitation factors and the economic status of trading countries in influencing export amounts, with a noted efficiency in model complexity reduction through Lasso Regression.

6

5,700원

본 연구에서는 우리나라 대기업과 중견·중소기업별 수출액에 대한 무역원활화 측정지표의 변수선택 차이를 살펴보기 위하여 측정지표로는 세계은행(World Bank)의 물류성과지표 (Logistics Performance Index, LPI)를 사용하였다. WTO는 무역원활화를 수출입 과정의 단순화를 지향하는 것이라 설명하고 있으며 여전히 국제 무역 및 운송 등의 과정에서 보이지 않는 많은 장벽이 존재하고 이는 기업에게 비용 등의 부담과 불확실성을 야기한다. 이에 기존 무역원활화 관련 선행연구가 국가 수준(country level)의 시사점을 제시한 것과 비교하여 기업 수준(firm level) 관점으로 접근하여 실증분석을 진행하였으며 추정량의 문제를 해결하기 위한 능형회귀분석 및 변수선택 기법에 따른 선택된 변수들의 기업크기별 차이를 살펴보기 위해 lasso 회귀분석을 적용하였다. 그 결과, 물류성과지표의 6가지 하위 변수 중, 대기업은 거의 유의하게 선택된 변수가 없는 것에 비해 중견·중소기업은 국제 운송, 물류 경쟁력 및 품질, 위탁화물 추적이 유의한 변수로 나타났다. 이러한 결과는 중견·중소기업은 물류 서비스 등의 소프트웨어적 속성을 지닌 변수 에서 발생하는 불확실성에 대해 민감하다는 점과 교역국의 상이한 물류 환경이 유발하는 기업 경영활동의 불확실성을 감소시킬 수 있도록 정책 방안이 유관기관에서도 수립되어야 한다는 시사점을 제시한다. 또한 세계적인 무역장벽 완화와 환경 개선이 이루어지고 있는 만큼, 향후 대기업은 물론 중견・중소기업 또한 해외시장 진출 기회 증가로 이어질 것이라 예상하는 바이다.

In this study, the Logistics Performance Index (LPI) of the World Bank was used to measure the difference in the variable selection of trade facilitation measures for exports by the leading Korean conglomerates and other small to medium enterprises. The WTO explains that trade facilitation aims to simplify the import and export process, and there still are many barriers not seen in the course of international trade and transportation which cause the enterprises burden and uncertainty of cost. The previous studies related to existing trade facilitation suggested the implications of country level. On the other hand, in this study, the empirical analysis was conducted by approaching from firm level perspective. In this analysis, lasso regression analysis was applied to investigate the difference between the selected variables according to the firm size by ridge regression analysis and to explore the variable selection technique to solve the problem of the estimator. As a result, among the six sub-parameters of LPI, leading conglomerates had little or no selected variables, while for small to medium enterprises, international shipments, logistics competitiveness and quality, and their tracking and tracing were significant variables. These results show that small to medium enterprises are sensitive to the uncertainty arising from variables with software attributes such as logistics services and that policy measures are needed to reduce the uncertainty of the enterprise management activities caused by the different logistics environment of their trading partners. In addition, as global trade barriers are eased and the trade environment is improving, we expect that leading conglomerates, as well as small to medium enterprises, will increase their opportunities to enter overseas markets.

7

4,000원

9

정보 소득율 기반의 변수 선택을 통한 영화 관객 수 예측

박현목, 최상현

한국정보기술응용학회 JITAM Vol.26 No.3 2019.06 pp.19-27

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

In this study, we propose a methodology for predicting the movie audience based on movie information that can be easily acquired before opening and effectively distinguishing qualitative variables. In addition, we constructed a model to estimate the number of movie audiences at the time of data acquisition through the configured variables. Another purpose of this study is to provide a criterion for categorizing success of movies with qualitative characteristics. As an evaluation criterion, we used information gain ratio which is the node selection criterion of C4.5 algorithm. Through the procedure we have selected 416 movie data features. As a result of the multiple linear regression model, the performance of the regression model using the variables selection method based on the information gain ratio was excellent.

10

10,000원

본 연구에서는 기술적 특성 및 비정상적 덤핑입찰 등 건설공사 입찰담합 손해액 산정을 위한 계량모형의 추정 과정에서 고려해야 할 다양한 요인들에 대해 분석했다. 주요 분석결과에 의하면, 건설공사 입찰담합의 손해액 산정은 계량모형의 설정뿐만 아니라 모형 추정에 사용되는 표본 자료의 선택에 따라 달라질 수 있다는 것을 확인하였다. 특히 모형 접근법 및 변수 선정에 따른 평균 손해액 추정치의 차이는 작게 도출된 반면, 개별 입찰 건들의 손해액 산정 결과는 계량모형과 표본 자료의 선택에 따라 크게 달라짐을 보였다. 이는 체계적인 손해액 추정을 위해서는 입찰 자료의 특성, 계량 모형의 장단점 등을 고려하여 각 담합 건에 적합한 손해액 산정법을 모색할 필요가 있다는 것을 시사한다.

This paper investigates various factors that should be considered in setting up and estimating econometric models for assessing overcharges of bid-rigging cases in construction industry. The main finding shows that the average of estimated overcharges seems robust to model approaches and variable selections, whereas the estimated damages for each individual bid-rigging cases differ substantially depending upon the choice of econometric models and sample data. This observation indicates that we should search for the most appropriate approach for assessing collusive overcharges for each individual contracts by employing a systematic approach to model specification and variable selection on a contract-by-contract basis. These results suggest important implications on rational model specification and systematic variable selection for assessing overcharges in construction industry bid-rigging cases.

11

Category Variable Selection Method for Efficient Clustering

Jun Heo, Chae Yun Kim, Yong-Gyu Jung

국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 2 Number 2 2013.11 pp.40-42

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Recent medical industry is an aging society and the application of national health insurance, with state-of-the-art research and development, including the pharmaceutical market is greatly increased. The nation's health care industry through new support expansion and improve the quality of life for the research and development will be needed. In addition, systemic administration of basic medical supplies , or drugs are needed , the drug at the same time managing how systematic analysis of pharmaceutical ingredients , based on data through the purchase of new medicines and pharmaceutical ingredients automatically classified by analyzing the statistics of drug purchases and the future a system that can predict a patient is needed. In this study, the drugs to the patient according to the component analysis and predictions for future research techniques, k-means clustering and k-NN (Nearest Neighbor) Comparative studies through experiments using the techniques employ a more efficient method to study how to proceed . In this study, the effects of the drugs according to the respective components in time according to the number of pieces in accordance with the patient by analyzing the statistics by predicting future patient better medical industry can be built.

12

Simultaneous outlier detection and variable selection via difference-based regression model and stochastic search variable selection

Park, Jong Suk, Park, Chun Gun, Lee, Kyeong Eun

[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.26 No.2 2019 pp.149-161

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this article, we suggest the following approaches to simultaneous variable selection and outlier detection. First, we determine possible candidates for outliers using properties of an intercept estimator in a difference-based regression model, and the information of outliers is reflected in the multiple regression model adding mean shift parameters. Second, we select the best model from the model including the outlier candidates as predictors using stochastic search variable selection. Finally, we evaluate our method using simulations and real data analysis to yield promising results. In addition, we need to develop our method to make robust estimates. We will also to the nonparametric regression model for simultaneous outlier detection and variable selection.

13

Variable selection in Poisson HGLMs using h-likelihoood

Ha, Il Do, Cho, Geon-Ho

[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.26 No.6 2015 pp.1513-1521

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Selecting relevant variables for a statistical model is very important in regression analysis. Recently, variable selection methods using a penalized likelihood have been widely studied in various regression models. The main advantage of these methods is that they select important variables and estimate the regression coefficients of the covariates, simultaneously. In this paper, we propose a simple procedure based on a penalized h-likelihood (HL) for variable selection in Poisson hierarchical generalized linear models (HGLMs) for correlated count data. For this we consider three penalty functions (LASSO, SCAD and HL), and derive the corresponding variable-selection procedures. The proposed method is illustrated using a practical example.

14

Variable Selection with Nonconcave Penalty Function on Reduced-Rank Regression

Jung, Sang Yong, Park, Chongsun

[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.22 No.1 2015 pp.41-54

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this article, we propose nonconcave penalties on a reduced-rank regression model to select variables and estimate coefficients simultaneously. We apply HARD (hard thresholding) and SCAD (smoothly clipped absolute deviation) symmetric penalty functions with singularities at the origin, and bounded by a constant to reduce bias. In our simulation study and real data analysis, the new method is compared with an existing variable selection method using $L_1$ penalty that exhibits competitive performance in prediction and variable selection. Instead of using only one type of penalty function, we use two or three penalty functions simultaneously and take advantages of various types of penalty functions together to select relevant predictors and estimation to improve the overall performance of model fitting.

15

Variable Selection and Outlier Detection for Automated K-means Clustering

Kim, Sung-Soo

[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.22 No.1 2015 pp.55-67

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

An important problem in cluster analysis is the selection of variables that define cluster structure that also eliminate noisy variables that mask cluster structure; in addition, outlier detection is a fundamental task for cluster analysis. Here we provide an automated K-means clustering process combined with variable selection and outlier identification. The Automated K-means clustering procedure consists of three processes: (i) automatically calculating the cluster number and initial cluster center whenever a new variable is added, (ii) identifying outliers for each cluster depending on used variables, (iii) selecting variables defining cluster structure in a forward manner. To select variables, we applied VS-KM (variable-selection heuristic for K-means clustering) procedure (Brusco and Cradit, 2001). To identify outliers, we used a hybrid approach combining a clustering based approach and distance based approach. Simulation results indicate that the proposed automated K-means clustering procedure is effective to select variables and identify outliers. The implemented R program can be obtained at http://www.knou.ac.kr/~sskim/SVOKmeans.r.

16

Variable selection in censored kernel regression

Choi, Kook-Lyeol, Shim, Jooyong

[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.24 No.1 2013 pp.201-209

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

For censored regression, it is often the case that some input variables are not important, while some input variables are more important than others. We propose a novel algorithm for selecting such important input variables for censored kernel regression, which is based on the penalized regression with the weighted quadratic loss function for the censored data, where the weight is computed from the empirical survival function of the censoring variable. We employ the weighted version of ANOVA decomposition kernels to choose optimal subset of important input variables. Experimental results are then presented which indicate the performance of the proposed variable selection method.

17

Variable selection in L1 penalized censored regression

Hwang, Chang-Ha, Kim, Mal-Suk, Shi, Joo-Yong

[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.22 No.5 2011 pp.951-959

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The proposed method is based on a penalized censored regression model with L1-penalty. We use the iteratively reweighted least squares procedure to solve L1 penalized log likelihood function of censored regression model. It provide the efficient computation of regression parameters including variable selection and leads to the generalized cross validation function for the model selection. Numerical results are then presented to indicate the performance of the proposed method.

18

Variable selection in the kernel Cox regression

Shim, Joo-Yong

[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.22 No.4 2011 pp.795-801

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In machine learning and statistics it is often the case that some variables are not important, while some variables are more important than others. We propose a novel algorithm for selecting such relevant variables in the kernel Cox regression. We employ the weighted version of ANOVA decomposition kernels to choose optimal subset of relevant variables in the kernel Cox regression. Experimental results are then presented which indicate the performance of the proposed method.

19

Variable Selection with Regression Trees

Chang, Young-Jae

[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.23 No.2 2010 pp.357-366

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Many tree algorithms have been developed for regression problems. Although they are regarded as good algorithms, most of them suffer from loss of prediction accuracy when there are many noise variables. To handle this problem, we propose the multi-step GUIDE, which is a regression tree algorithm with a variable selection process. The multi-step GUIDE performs better than some of the well-known algorithms such as Random Forest and MARS. The results based on simulation study shows that the multi-step GUIDE outperforms other algorithms in terms of variable selection and prediction accuracy. It generally selects the important variables correctly with relatively few noise variables and eventually gives good prediction accuracy.

20

Variable selection for multiclassi cation by LS-SVM

Hwang, Hyung-Tae

[Kisti 연계] 한국데이터정보과학회 한국데이터정보과학회지 Vol.21 No.5 2010 pp.959-965

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

For multiclassification, it is often the case that some variables are not important while some variables are more important than others. We propose a novel algorithm for selecting such relevant variables for multiclassification. This algorithm is base on multiclass least squares support vector machine (LS-SVM), which uses results of multiclass LS-SVM using one-vs-all method. Experimental results are then presented which indicate the performance of the proposed method.

 
1 2 3 4 5
페이지 저장