Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 16
No
1

4,000원

최근 온라인 게임 이용자가 기하급수적으로 증가하는 초기 진입단계를 지나 현재 성장세가 둔화되는 현상을 나타내고 있다. 게임 시장 성장단계에서 성숙단계로 진입한 것을 단적으로 나타내는 증거로 판단된다. 성숙단계의 온라인게임시장에 대한 철저한 분석과 전략수립 없이는 시장 진입 및 기업 성과를 이끌어 내는데 걸림돌로 작용할 것이다. 이에 게임 산업 내 기업들은 온라인 게임 시장에 대한 정보를 파악하는데 시간과 자원을 투자하고 있다. 그러나 시장을 이해하기 위한 대표적 정보인 시장 세분화에 대한 연구가 미비하다. 학계나 산업계에서는 기존 게임시장세분화에 관한 연구와 보고서가 진행되고 있으나 게임 시장에 대한 전반적인 기초적 수준의 정보를 제공하고 있으며, 정보 또한 제한적인 내용이 많아 산업 내 기업들이 실무적 도움을 주고 있지 못하고 있는 것이 현실이다. 따라서 본 연구에서는 기존 연구에서 나타난 한계점을 보완하고 향후 연구에 기반이 되고자 기존 시장 세분화의 기준변수와 설명변수의 고찰하고 온라인 게임 시장 세분화 변수를 추출하여 온라인 게임 시장세분화를 수행하는데 있어 적합한 기준을 제시하고자 한다.

Recently, online game users growing exponentially over the initial steps to enter and the phenomenon indicates that growth is slowing. That is the initial stages of growth of the game entered a mature phase indicates. A thorough market analysis and strategic planning. If so, you can expect the market will not fail and corporate performance. So, companies have to learn about online game market, a lot of time and resources invested. However, the market for understanding representative incomplete information, market segmentation aggressive growth strategy that companies are now unable to lead the industry. This research will complement the existing limitations, and future studies appeared in the market for research that development would like variables. Along, online game market segmentation variables that make it suitable to perform online game based on market segmentation is expected to be utilized.

2

Functional Data Classification of Variable Stars

Park, Minjeong, Kim, Donghoh, Cho, Sinsup, Oh, Hee-Seok

[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.20 No.4 2013 pp.271-281

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper considers a problem of classification of variable stars based on functional data analysis. For a better understanding of galaxy structure and stellar evolution, various approaches for classification of variable stars have been studied. Several features that explain the characteristics of variable stars (such as color index, amplitude, period, and Fourier coefficients) were usually used to classify variable stars. Excluding other factors but focusing only on the curve shapes of variable stars, Deb and Singh (2009) proposed a classification procedure using multivariate principal component analysis. However, this approach is limited to accommodate some features of the light curve data that are unequally spaced in the phase domain and have some functional properties. In this paper, we propose a light curve estimation method that is suitable for functional data analysis, and provide a classification procedure for variable stars that combined the features of a light curve with existing functional data analysis methods. To evaluate its practical applicability, we apply the proposed classification procedure to the data sets of variable stars from the project STellar Astrophysics and Research on Exoplanets (STARE).

3

Robust Variable Selection in Classification Tree

장정이, 정광모

[Kisti 연계] 한국통계학회 한국통계학회 학술대회논문집 2001 pp.89-94

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this study we focus on variable selection in decision tree growing structure. Some of the splitting rules and variable selection algorithms are discussed. We propose a competitive variable selection method based on Kruskal-Wallis test, which is a nonparametric version of ANOVA F-test. Through a Monte Carlo study we note that CART has serious bias in variable selection towards categorical variables having many values, and also QUEST using F-test is not so powerful to select informative variables under heavy tailed distributions.

4

Tree-structured Classification based on Variable Splitting

Ahn, Sung-Jin

[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.2 No.1 1995 pp.74-88

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This article introduces a unified method of choosing the most explanatory and significant multiway partitions for classification tree design and analysis. The method is derived on the impurity reduction (IR) measure of divergence, which is proposed to extend the proportional-reduction-in-error (PRE) measure in the decision-theory context. For the method derivation, the IR measure is analyzed to characterize its statistical properties which are used to consistently handle the subjects of feature formation, feature selection, and feature deletion required in the associated classification tree construction. A numerical example is considered to illustrate the proposed approach.

5

Multivariate Procedure for Variable Selection and Classification of High Dimensional Heterogeneous Data

Mehmood, Tahir, Rasheed, Zahid

[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.22 No.6 2015 pp.575-587

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The development in data collection techniques results in high dimensional data sets, where discrimination is an important and commonly encountered problem that are crucial to resolve when high dimensional data is heterogeneous (non-common variance covariance structure for classes). An example of this is to classify microbial habitat preferences based on codon/bi-codon usage. Habitat preference is important to study for evolutionary genetic relationships and may help industry produce specific enzymes. Most classification procedures assume homogeneity (common variance covariance structure for all classes), which is not guaranteed in most high dimensional data sets. We have introduced regularized elimination in partial least square coupled with QDA (rePLS-QDA) for the parsimonious variable selection and classification of high dimensional heterogeneous data sets based on recently introduced regularized elimination for variable selection in partial least square (rePLS) and heterogeneous classification procedure quadratic discriminant analysis (QDA). A comparison of proposed and existing methods is conducted over the simulated data set; in addition, the proposed procedure is implemented to classify microbial habitat preferences by their codon/bi-codon usage. Five bacterial habitats (Aquatic, Host Associated, Multiple, Specialized and Terrestrial) are modeled. The classification accuracy of each habitat is satisfactory and ranges from 89.1% to 100% on test data. Interesting codon/bi-codons usage, their mutual interactions influential for respective habitat preference are identified. The proposed method also produced results that concurred with known biological characteristics that will help researchers better understand divergence of species.

6

신호 특성을 고려한 가변 상태 개수 HMM을 이용한 심음분류 방법에 대한 고찰

정용주

[NRF 연계] 대한의료정보학회 Healthcare Informatics Research Vol.14 No.2 2008.06 pp.179-187

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Hidden Markov model (HMM) is known to be one of the most powerful methods in the acoustic modeling of heart sound signals. Conventionally, we usually use a fixed number of states for each HMM. However, due to the various types of the heart sound signals, it seems that more accurate acoustic modeling is possible by varying the number of states in the HMM depending on the signal types to be modeled. In this paper, we propose to assign different number of states to the HMM for better acoustic modeling and consequently, improving the classification performance of the heart sound signals. Compared with when fixing the number of states, the proposed approach has shown some performance improvement in the classification experiments on various types of heart sound signals.

7

항공 라이다 데이터로부터 가변형 판정 윈도우를 이용한 지형 분류 기법

성철웅, 이성규, 박창후, 이호준, 김유성

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2010 pp.79-82

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 항공 라이다 데이터로부터 지형의 유형을 분류하는 과정에서 지형적 특성에 따라 분류에 사용하는 판정 윈도우의 크기를 가변적으로 조정하여 적용시키는 가변형 판정 윈도우를 이용한 지형 분류 기법을 제안하였다. 또한 실험을 통하여 가변형 판정 윈도우를 이용한 지형 분류 기법의 시간효율과 정확도를 분석하였다. 실험 결과에 따르면 제안된 가변형 판정 윈도우를 이용한 지형 분류 기법은 지형 분류에 사용되는 판정 윈도우의 개수를 줄여 지형 분류의 속도를 향상시켰기 때문에 빠른 분류 속도가 필요한 재해 피해 현황을 파악하기 위한 시스템에 적용 할 수 있을 것으로 기대된다.

8

항공 라이다 데이터로부터 가변형 판정 윈도우를 이용한 지형 분류 기법

성철웅, 이성규, 박창후, 이호준, 김유성

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2010 pp.79-82

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 항공 라이다 데이터로부터 지형의 유형을 분류하는 과정에서 지형적 특성에 따라 분류에 사용하는 판정 윈도우의 크기를 가변적으로 조정하여 적용시키는 가변형 판정 윈도우를 이용한 지형 분류 기법을 제안하였다. 또한 실험을 통하여 가변형 판정 윈도우를 이용한 지형 분류 기법의 시간효율과 정확도를 분석하였다. 실험 결과에 따르면 제안된 가변형 판정 윈도우를 이용한 지형 분류 기법은 지형 분류에 사용되는 판정 윈도우의 개수를 줄여 지형 분류의 속도를 향상시켰기 때문에 빠른 분류 속도가 필요한 재해 피해 현황을 파악하기 위한 시스템에 적용 할 수 있을 것으로 기대된다.

9

랜섬웨어 탐지를 위한 동적 분석 자료에서의 변수 선택 및 분류에 관한 연구

이승환, 황진수

[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.31 No.4 2018 pp.497-505

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

최근 랜섬웨어는 일반 PC 사용자에 비해 상대적으로 수준 높은 보안 체계를 갖추고 있는 기업과 정부 기관에 침입하여 상당한 피해를 입히는 등 기존 보안 체계의 허점을 찾아 진화하는 모습을 보이고 있다. 이처럼 계속해서 변화하는 랜섬웨어를 탐지하기 위해 랜섬웨어의 특징을 파악하는 정적 분석과 동적 분석과 관련된 연구가 활발히 이루어지고 있다. 본 연구에서는 582개의 랜섬웨어 샘플과 942개의 정상 샘플 프로그램을 쿠쿠 샌드박스 가상환경 내에서 실행시킨 뒤, PC에서 이루어지는 30,967가지의 행동 여부를 기록한 동적 분석 자료를 활용하여 랜섬웨어 분류에 유의한 변수를 탐색하기 위한 여러 변수 선택 방법의 적용과 랜섬웨어 분류를 위한 기계학습 모형들을 구축하고자 하였다. 변수 선택법으로 LASSO와 이항변수 만으로 이루어진 고차원 자료라는 특성을 활용하기 위한 카이제곱검정을 이용한 변수 선택, 선행 연구에서 이용된 방법인 상호정보를 이용한 변수 선택법을 적용하였으며 기계 학습 모형으로는 능형 로지스틱 회귀, 서포트 벡터 머신, 랜덤 포레스트, XGBoost가 활용되었다. 연구 결과, 정상 프로그램과 구별되는 랜섬웨어 프로그램만의 특징적인 행동을 확인할 수 있었으며 여러 변수 선택법과 기계학습 분류 모형들의 조합 중, 주어진 자료에서 카이제곱검정을 이용한 변수 선택법과 랜덤 포레스트 모형의 조합이 가장 높은 탐지율과 정분류율을 보이는 것을 확인하였다.

Attacking computer systems using ransomware is very common all over the world. Since antivirus and detection methods are constantly improved in order to detect and mitigate ransomware, the ransomware itself becomes equally better to avoid detection. Several new methods are implemented and tested in order to optimize the protection against ransomware. In our work, 582 of ransomware and 942 of normalware sample data along with 30,967 dynamic action sequence variables are used to detect ransomware efficiently. Several variable selection techniques combined with various machine learning based classification techniques are tried to protect systems from ransomwares. Among various combinations, chi-square variable selection and random forest gives the best detection rates and accuracy.

10

복지대상 사각지대 해소를 위한 점진적 변수 확장 기반 분류 모형: TVAE를 활용한 변수 생성과 재현율 최적화를 중심으로

박영식, 이동원, 이형용

[NRF 연계] 한국경영학회 경영학연구 Vol.55 No.2 2026.04 pp.583-611

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 복지사각지대의 핵심 원인인 변수 결측과 예측 누락 문제를 해결하기 위해, TVAE(Tabular Variational AutoEncoder) 기반 구조적 결측 대치 기법을 적용하고 재현율(Recall)을 중심으로 한 분류모형을 설계하였다. 2018년1월부터 2023년 11월까지의 복지행정 데이터를 활용하여, 변수 확장 및 합성데이터 결합 효과를 단계적으로 검증하였다. Phase 1에서는 실제 변수 확장을 통해 재현율이 소폭 향상(약 0.4~1.0%p)되었고, Phase 2에서는 TVAE로 대치된 데이터를 결합함으로써 재현율이 57.7%에서 94.4%로 대폭 개선되었다. 또한 합성데이터 결합비율 민감도 분석(Test 1~4)을 통해 약 10~20% 수준의 합성비율에서도 안정적인 예측 성능이 유지됨을 확인하였다. 시간 블록(Time-block) 검증 결과, ROC-AUC(0.66~0.67)와 PR-AUC(0.31~0.32)가 일정하게 유지되어 시계열적 강건성이 입증되었다. 본 연구는 TVAE 기반 결측 대치가 복지행정 데이터의 품질과 예측 신뢰성을 실질적으로 개선함을 실증적으로 제시하며, 재현율 중심의 데이터 기반 복지정책 의사결정의 타당성을 뒷받침한다.

This study aims to address the key causes of welfare blind spots?missing variables and prediction omissions?by applying a Tabular Variational AutoEncoder (TVAE)-based structural missing data imputation technique and designing a recall-centered classification model. Using welfare administrative data collected from January 2018 to November 2023, the study systematically verified the effects of variable expansion and synthetic data integration. In Phase 1, recall slightly improved by approximately 0.4?1.0 percentage points through direct variable expansion. In Phase 2, combining TVAE-imputed data led to a substantial increase in recall, from 57.7% to 94.4%. A sensitivity analysis of synthetic data combination ratios (Test 1?4) confirmed that stable predictive performance was maintained even with 10? 20% synthetic data inclusion. The time-block validation further demonstrated temporal robustness, with ROC-AUC values remaining between 0.66 and 0.67 and PR-AUC between 0.31 and 0.32 across time periods. These results empirically demonstrate that TVAE-based imputation effectively enhances the quality and predictive reliability of welfare administrative data, supporting the validity of recall-oriented, data-driven decision-making in welfare policy design.

11

단순 베이즈 분류에서의 범주형 변수의 선택

김민선, 최호식, 박창이

[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.28 No.3 2015 pp.407-415

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

단순 베이즈 분류($Na{\ddot{i}}ve$ Bayes classification)는 출력변수가 주어졌을 때 입력변수들이 조건부 독립이라는 가정에 기반한다. 단순 베이즈 가정은 비현실적이지만 고차원의 확률 추정 문제를 일련의 일차원 확률 추정 문제로 단순화 시킨다는 장점이 있으며, 특히 스팸 메일 필터링, 추천 시스템(recommendation system) 등 방대한 데이터를 다루는 분야야에서 흔히 사용된다. 본 논문에서는 입력변수와 출력변수간의 카이제곱 통계량에 기반한 변수선택법을 제안한다. 이 방법은 단순 베이즈 분류의 장점인 데이터 처리 및 계산의 단순성을 유지하면서도 설명력이 있는 변수를 선택할 수 있으며 SNP(single nucleotide polymorphism)에 의한 질병의 분류 등의 초고차원 혹은 빅데이터에서 유용할 것으로 기대된다.

$Na{\ddot{i}}ve$ Bayes Classification is based on input variables that are a conditionally independent given output variable. The $Na{\ddot{i}}ve$ Bayes assumption is unrealistic but simplifies the problem of high dimensional joint probability estimation into a series of univariate probability estimations. Thus $Na{\ddot{i}}ve$ Bayes classier is often adopted in the analysis of massive data sets such as in spam e-mail filtering and recommendation systems. In this paper, we propose a variable selection method based on ${\chi}^2$ statistic on input and output variables. The proposed method retains the simplicity of $Na{\ddot{i}}ve$ Bayes classier in terms of data processing and computation; however, it can select relevant variables. It is expected that our method can be useful in classification problems for ultra-high dimensional or big data such as the classification of diseases based on single nucleotide polymorphisms(SNPs).

12

나무구조의 분류분석에서 변수 중요도에 대한 고찰

김나영, 이은경

[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.27 No.5 2014 pp.717-729

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구에서는 나무구조의 분류분석에서 자료의 크기가 방대해짐에 따라 중요한 문제로 대두되고 있는 변수의 중요도에 대하여 사영추적분류나무를 중심으로 고찰하였다. 사영추적분류나무(projection pursuit classification tree)는 각 마디에서 사영추적을 이용하여 그룹을 잘 분리하는 변수들의 선형결합을 이용하는 방법으로 이때 사용되는 사영계수들은 각 마디에서의 분류에 대한 정보를 가지고 있다. 이를 종합하여 각 변수의 분류에 대한 중요도를 계산할 수 있다. 먼저 사영추적분류나무의 분류과정에서 계산되는 사영추적계수를 이용하여 분류를 위한 변수선택의 중요도를 계산하고 이들의 특성을 살펴보고 이를 같은 형태의 나무모형방법인 CART와 랜덤 포레스트의 결과와 비교 분석하여 사영추적분류나무의 특성을 살펴보고 비교, 분석하였다. 대부분의 자료에서 사영추적분류나무가 훨씬 좋은 성능을 보이고 있었으며 특히 상관계수가 높은 변수들이 포함되어 있는 경우에는 상대적으로 적은 수의 변수로도 잘 분류를 할 수 있음을 확인하였다. 랜덤 포레스트에서 제공하는 변수 중요도는 변수들 간의 상관관계가 높은 경우에는 사영추적분류나무의 변수중요도와 매우 다르게 나타나며 사영추적분류나무의 변수 중요도가 조금 더 나은 성능을 보이고 있음을 알 수 있다.

Projection pursuit classification tree uses a 1-dimensional projection with the view of the most separating classes in each node. These projection coefficients contain information distinguishing two groups of classes from each other and can be used to calculate the importance measure of classification in each variable. This paper reviews the variable importance measure with increasing interest in line with growing data size. We compared the performances of projection pursuit classification tree with those of classification and regression tree(CART) and random forest. Projection pursuit classification tree are found to produce better performance in most cases, particularly with highly correlated variables. The importance measure of projection pursuit classification tree performs slightly better than the importance measure of random forest.

13

의사결정나무모형을 이용한 사상체질분류함수의 개발에 관한 연구 (I) - 크론박 알파 계수에 의한 변수선택 -

김규곤

[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.6 No.3 2004.06 pp.751-766

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 한방병원에서 사상체질분류검사설문지를 이용하여 사상체질을 진단할 때 진단의 정확도를 향상시키기 위한 사상체질분류함수를 개발하기 위하여 데이터마이닝에서의 의사결정나무모형을 이용한다. 본 연구에서 사용하는 데이터는 경희대학교 강남경희한방병원과 동의대학교 부속한방병원에서 치료를 받은 환자들 중 각 병원의 사상체질전문의로부터 체질진단을 받고 최소한 4주 이상 사상체질 처방을 사용한 후 주 증상이 전반적으로 호전되어 체질이 확인된 환자 1051명을 대상으로 하고 있다. 데이터 정제 과정에서 불성실한 응답자를 제거시키기 위한 기준은 상반되는 설문의 응답 패턴과 체질별 설문의 응답 비율을 이용하며, 변수선택의 기준은 상관분석의 크론박 알파 계수와 선형판별함수의 계수를 이용한다.

In this study, when we make a diagnosis of Sasang constitution using QSCC II(Questionnaire of Sasang Constitution Classification), decision tree model in data mining is applied to seek the classification function into Sasang constitution for improving the accuracy. Data used in the analysis are the questionnaires of 1051 patients who had been treated in Dong Eui Oriental Medical Hospital and Kyung Hee Oriental Medical Hospital. The criteria for data cleansing are the response pattern in the opposite questionnaires and the positive proportion of specific questionnaires in each constitution. And the criteria for variable selection are Cronbach's alpha coefficients in correlation analysis and the coefficients in the linear discriminant function.

14

의사결정나무모형을 이용한 사상체질분류함수의 개발에 관한 연구 (II) - 도수분석에 의한 변수선택 -

김규곤, 최승배

[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.6 No.3 2004.06 pp.767-783

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 한방병원에서 사상체질분류검사설문지를 이용하여 사상체질을 진단할 때 진단의 정확도를 향상시키기 위한 사상체질분류함수를 개발하기 위하여 데이터마이닝에서의 의사결정나무모형을 이용한다. 본 연구에서 사용하는 데이터는 경희대학교 강남경희한방병원과 동의대학교 부속한방병원에서 치료를 받은 환자들 중 각 병원의 사상체질전문의로부터 체질진단을 받고 최소한 4주 이상 사상체질 처방을 사용한 후 주 증상이 전반적으로 호전되어 체질이 확인된 환자 1051명을 대상으로 하고 있다. 데이터 정제 과정에서 양질의 데이터를 확보하기 위한 기준은 상반되는 설문의 응답 패턴과 체질별 설문의 응답 비율을 이용하며, 변수선택의 기준은 도수분석의 비율차이검정과 선형판별함수의 계수를 이용한다.

In this study, when we make a diagnosis of Sasang constitution using QSCC II(Questionnaire of Sasang Constitution Classification), decision tree model in data mining is applied to seek the classification function into Sasang constitution for improving the accuracy. Data used in the analysis are the questionnaires of 1051 patients who had been treated in Dong Eui Oriental Medical Hospital and Kyung Hee Oriental Medical Hospital. The criteria for data cleansing are the response pattern in the opposite questionnaires and the positive proportion of specific questionnaires in each constitution. And the criteria for variable selection are the test of homogeneity in frequency analysis and the coefficients in the linear discriminant function.

15

판별분석모형을 이용한 사상체질분류함수의 개발에 관한 연구 (II) - 도수분석에 의한 변수선택 -

김규곤, 최승배

[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.6 No.2 2004.04 pp.505-517

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 한방병원에서 사상체질분류검사설문지를 이용하여 사상체질을 진단할 때 진단의 정확도를 향상시키기 위한 사상체질분류함수를 개발하기 위하여 데이터마이닝에서의 판별분석모형을 이용한다. 본 연구에서 사용하는 데이터는 경희대학교 강남경희한방병원과 동의대학교 부속한방병원에서 치료를 받은 환자들 중 각 병원의 사상체질전문의로부터 체질진단을 받고 최소한 4주 이상 사상체질 처방을 사용한 후 주 증상이 전반적으로 호전되어 체질이 확인된 환자 1051명을 대상으로 하고 있다. 데이터 정제 과정에서 양질의 데이터를 확보하기 위한 기준은 상반되는 설문의 응답 패턴과 체질별 설문의 응답 비율을 이용하며, 변수선택의 기준은 도수분석의 비율차이검정과 선형판별함수의 계수를 이용한다.

In this study, when we make a diagnosis of Sasang constitution using QSCC II(Questionnaire of Sasang Constitution Classification), discriminant analysis model in data mining is applied to seek the classification function into Sasang constitution for improving the accuracy. Data used in the analysis are the questionnaires of 1051 patients who had been treated in Dong Eui Oriental Medical Hospital and Kyung Hee Oriental Medical Hospital. The criteria for data cleansing are the response pattern in the opposite questionnaires and the positive proportion of specific questionnaires in each constitution. And the criteria for variable selection are the test of homogeneity in frequency analysis and the coefficients in the linear discriminant function.

16

판별분석모형을 이용한 사상체질분류함수의 개발에 관한 연구 (I) - 크론박 알파 계수에 의한 변수선택 -

김규곤, 김종원, 이의주, 최선미, 조민형, 김동준, 이소영

[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.6 No.2 2004.04 pp.493-504

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 한방병원에서 사상체질분류검사설문지를 이용하여 사상체질을 진단할 때 진단의 정확도를 향상시키기 위한 사상체질분류함수를 개발하기 위하여 데이터마이닝에서의 판별분석모형을 이용한다. 본 연구에서 사용하는 데이터는 경희대학교 강남경희한방병원과 동의대학교 부속한방병원에서 치료를 받은 환자들 중 각 병원의 사상체질전문의로부터 체질진단을 받고 최소한 4주 이상 사상체질 처방을 사용한 후 주 증상이 전반적으로 호전되어 체질이 확인된 환자 1051명을 대상으로 하고 있다. 데이터 정제 과정에서 불성실한 응답자를 제거시키기 위한 기준은 상반되는 설문의 응답 패턴과 체질별 설문의 응답 비율을 이용하며, 변수선택의 기준은 상관분석의 크론박 알파 계수와 선형판별함수의 계수를 이용한다.

In this study, when we make a diagnosis of Sasang constitution using QSCC II(Questionnaire of Sasang Constitution Classification), discriminant analysis model in data mining is applied to seek the classification function into Sasang constitution for improving the accuracy. Data used in the analysis are the questionnaires of 1051 patients who had been treated in Dong Eui Oriental Medical Hospital and Kyung Hee Oriental Medical Hospital. The criteria for data cleansing are the response pattern in the opposite questionnaires and the positive proportion of specific questionnaires in each constitution. And the criteria for variable selection are Cronbach's alpha coefficients in correlation analysis and the coefficients in the linear discriminant function.

 
페이지 저장