년 - 년
Query Optimization on Large Scale Nested Data with Service Tree and Frequent Trajectory
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.17 No.1 2021 pp.37-50
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Query applications based on nested data, the most commonly used form of data representation on the web, especially precise query, is becoming more extensively used. MapReduce, a distributed architecture with parallel computing power, provides a good solution for big data processing. However, in practical application, query requests are usually concurrent, which causes bottlenecks in server processing. To solve this problem, this paper first combines a column storage structure and an inverted index to build index for nested data on MapReduce. On this basis, this paper puts forward an optimization strategy which combines query execution service tree and frequent sub-query trajectory to reduce the response time of frequent queries and further improve the efficiency of multi-user concurrent queries on large scale nested data. Experiments show that this method greatly improves the efficiency of nested data query.
당뇨병 관리전략을 위한 혈당조절 관련 생활습관 요인 : 국민건강영양조사 활용 코호트내 환자-대조군 연구 KCI 등재
한국융합학회 한국융합학회논문지 제10권 제11호 2019.11 pp.501-510
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 당뇨병의 관리전략 수립을 위해 일상생활에서 조절가능한 혈당조절 요인을 파악하고자 하였다. 개인 이 조절하기 어려운 성별, 당뇨병 유병기간, 당뇨병 치료방법, 교육수준, 가구소득 등을 매칭한 코호트내 환자-대조군 디자인으로 분석을 하였다. 7기 국민건강영양조사(2016-2017) 원시자료를 이용하여, 983명의 당뇨병 환자를 분석한 결과, 289명(30%)만이 당화혈색소가 6.5% 미만으로 혈당이 조절되었다. 일상생활에서 조절이 가능한 당화혈색소 조절 요인 파악하기 위해 조건부 다변량 로짓스틱 회귀분석 시행한 결과, 매칭전 코호트에서는 당뇨병 유병기간, 당뇨병 치료 여부, 매칭 코호트에서는 체질량지수, 흡연, 안저검사가 당화혈색소 달성에 유의한 영향을 미치는 것으로 파악되었다. 본 결과는 집중 관리가 필요한 대상선정(장기 유병기간, 약물치료 대상자) 및 생활습관(체질량지수, 흡연, 안저검사 등) 관리전략을 마련하는데 근거자료로 활용되고, 지역사회내 국민건강 증진에 기여할 것으로 사료된다.
This study aimed to find health related lifestyle factors that influence glycemic control for diabetes mellitus (DM) management strategies. This study used nested case-control design with matching variables that were not controlled by individuals such as age, sex, insulin or oral hypoglycemic agent (OHA) use, disease duration, education level and household income. This study analyzed 983 subjects with type 2 DM who enrolled in the 7th (2016-2017) Korean National Health and Nutrition Examination Survey (KNHANES). The target HbA1c level of controlled glucose was defined as less than 6.5%, and 289 (30%) were achieved. Conditional multivariable logistic regression analysis was performed to find self-control factors associated with HbA1c levels. The results statistically significant for variables such as duration of diabetes, insulin or OHA use in overall cohort and body mass index (BMI), smoking and fundus Examination in matched cohort. These results are expected to provide as evidence for the intensive care criteria(disease duration, drug use) and lifestyle management strategy(BMI, smoking, fundus examination).
Improving the Performance of Precise Query Processing on Large-scale Nested Data with UniHash Index SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.1 2015.02 pp.111-128
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Querying nested data has become one of the most pivotal issues for seeking desired information on the Web. Unlike the traditional information retrieval, to effectively manage nested data, we generally need store the data and its structures separately, which significantly reduces the performance of data retrieving, especially when the dataset is in a large scale. More seriously, it brings a big challenge on ensuring the efficiency of processing precise queries that need to locate the exact positions of some certain values in a nested dataset. Combining the techniques of column-strip storage and inverted index, this paper defines an expression to represent the data objects’ unique location in nested records— UPath, and based on which we present a new index structure— UniHash to support precise query processing on nested datasets. In addition, this work develops the related algorithms for building UPath, establishing UniHash, performing precise queries on UniHash with MapReduce platform and maintaining UniHash as well. Compared with some existing approaches, such as XPath-based and Dremel, UniHash index is capable of supporting the execution of precise queries over nested dataset with better performance. We give the results of a group of experiments, which were conducted on different real datasets, to demonstrate the efficiency of the approach.
Latent Clas Modeling for Nested Data : Introduction to Multilevel Latent Clas Model
[NRF 연계] 경북대학교 사회과학기초자료연구소 연구방법논총 Vol.5 No.3 2020.11 pp.183-215
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The fundamental asumption in any latent variable model is that observations are independent of one another, given the latent status. However, this asumption is often inadequate when observations are nested within higher-level units because such nested data structures induce dependencies in data. The nonparametric version of the multilevel latent clas model (MLCM) is an extension of latent clas model (LCM) in which the dependencies in data are acounted for by discrete latent variables. This paper aims to review models with discrete latent variables and introduce the MLCM which integrates LCM and random efect model. The model selection isue in the MLCM is also discused with an empirical example.
Variance components for two-way nested design data
[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.25 No.3 2018 pp.275-282
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper discusses the use of projections for the sums of squares in the analyses of variance for two-way nested design data. The model for this data is assumed to only have random effects. Two different sizes of experimental units are required for a given experimental situation, since nesting is assumed to occur both in the treatment structure and in the design structure. So, variance components are coming from the sources of random effects of treatment factors and error terms in different sizes of experimental units. The model for this type of experimental situation is a random effects model with more than one error terms and therefore estimation of variance components are concerned. A projection method is used for the calculation of sums of squares due to random components. Squared distances of projections instead of using the usual reductions in sums of squares that show how to use projections to estimate the variance components associated with the random components in the assumed model. Expectations of quadratic forms are obtained by the Hartley's synthesis as a means of calculation.
Analysis of Nested CGD Infection Data Using Multilevel Frailty Model
[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.6 No.3 2004.06 pp.637-645
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Multivariate event-time data with nested structures can be modelled by random-effect survival models such as multilevel frailty models. We show how to analyze the chronic granulomatous disease data(Fleming and Harrington, 1991), which comprise recurrent infection times of patients from different hospitals, using multilevel frailty model. Inferences are developed using hierarchical-likelihood, which provides a simple unified framework and a numerically efficient fitting algorithm for various random-effect models.
The Nested Variable Model of FDI Spillover Effects: Estimation Using Hungarian Panel Data
[NRF 연계] 한국국제경제학회 International Economic Journal Vol.26 No.4 2012.12 pp.673-709
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
A new empirical model is presented that considers the productivity spillover effects of foreign direct investment (FDI) by focusing on the multi-layered structure of industrial classifications. In this model, the market presence of horizontal FDI in a host country is expressed using multiple spillover variables with a nested structure corresponding to the aggregated level of industrial classification. Using large-scale firm-level data from Hungary, we estimated the nested variable model and verified horizontal FDI spillover effects that cannot be captured with the conventional model having a single horizontal variable.
[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.16 No.4 2009 pp.659-667
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Recently Song and Cheon (2006) and Cheon and Lim (2009) developed the generalized maximum entropy(GME) estimator to solve ill-posed problems for the regression coefficients in the simple panel model. The models discussed consider the individual and a spatial autoregressive disturbance effects. However, in many application in economics the data may contain nested groupings. This paper considers a two-way error component model with nested groupings for the ill-posed data and proposes the GME estimator of the unknown parameters. The performance of this estimator is compared with the existing methods on the simulated dataset. The results indicate that the GME method performs the best in estimating the unknown parameters in terms of its quality when the data are ill-posed.
Ensemble of Nested Dichotomies 기법을 이용한 스마트폰 가속도 센서 데이터 기반의 동작 인지
[Kisti 연계] 한국지능정보시스템학회 Journal of Intelligence and Information Systems Vol.19 No.4 2013 pp.123-132
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 스마트 폰에 다양한 센서를 내장할 수 있게 되었고 스마트폰에 내장된 센서를 이용항 동작 인지에 관한 연구가 활발히 진행되고 있다. 스마트폰을 이용한 동작 인지는 노인 복지 지원이나 운동량 측정. 생활 패턴 분석, 운동 패턴 분석 등 다양한 분야에 활용될 수 있다. 하지만 스마트 폰에 내장된 센서를 이용하여 동작 인지를 하는 방법은 사용되는 센서의 수에 따라 단일 센서를 이용한 동작인지와 다중 센서를 이용한 동작인지로 나눌 수 있다. 단일 센서를 이용하는 경우 대부분 가속도 센서를 이용하기 때문에 배터리 부담은 줄지만 다양한 동작을 인지할 때에 특징(feature) 추출의 어려움과 동작 인지 정확도가 낮다는 문제점이 있다. 그리고 다중 센서를 이용하는 경우 대부분 가속도 센서와 중력센서를 사용하고 필요에 따라 다른 센서를 추가하여 동작인지를 수행하며 다양한 동작을 보다 높은 정확도로 인지할 수 있지만 다수의 센서를 사용하기 때문에 배터리 부담이 증가한다는 문제점이 있다. 따라서 본 논문에서는 이러한 문제를 해결하기 위해 스마트 폰에 내장된 가속도 센서를 이용하여 다양한 동작을 높은 정확도로 인지하는 방법을 제안한다. 서로 다른 10가지의 동작을 높을 정확도로 인지하기 위해 원시 데이터로부터 17가지 특징을 추출하고 각 동작을 분류하기 위해 Ensemble of Nested Dichotomies 분류기를 사용하였다. Ensemble of Nested Dichotomies 분류기는 다중 클래스 문제를 다수의 이진 분류 문제로 변형하여 다중 클래스 문제를 해결하는 방법으로 서로 다른 Nested Dichotomy 분류기의 분류 결과를 통해 다중 클래스 문제를 해결하는 기법이다. Nested Dichotomy 분류기 학습에는 Random Forest 분류기를 사용하였다. 성능 평가를 위해 Decision Tree, k-Nearest Neighbors, Support Vector Machine과 비교 실험을 한 결과 Ensemble of Nested Dichotomies 분류기를 사용하여 동작 인지를 수행하는 것이 가장 높은 정확도를 보였다.
As the smartphones are equipped with various sensors such as the accelerometer, GPS, gravity sensor, gyros, ambient light sensor, proximity sensor, and so on, there have been many research works on making use of these sensors to create valuable applications. Human activity recognition is one such application that is motivated by various welfare applications such as the support for the elderly, measurement of calorie consumption, analysis of lifestyles, analysis of exercise patterns, and so on. One of the challenges faced when using the smartphone sensors for activity recognition is that the number of sensors used should be minimized to save the battery power. When the number of sensors used are restricted, it is difficult to realize a highly accurate activity recognizer or a classifier because it is hard to distinguish between subtly different activities relying on only limited information. The difficulty gets especially severe when the number of different activity classes to be distinguished is very large. In this paper, we show that a fairly accurate classifier can be built that can distinguish ten different activities by using only a single sensor data, i.e., the smartphone accelerometer data. The approach that we take to dealing with this ten-class problem is to use the ensemble of nested dichotomy (END) method that transforms a multi-class problem into multiple two-class problems. END builds a committee of binary classifiers in a nested fashion using a binary tree. At the root of the binary tree, the set of all the classes are split into two subsets of classes by using a binary classifier. At a child node of the tree, a subset of classes is again split into two smaller subsets by using another binary classifier. Continuing in this way, we can obtain a binary tree where each leaf node contains a single class. This binary tree can be viewed as a nested dichotomy that can make multi-class predictions. Depending on how a set of classes are split into two subsets at each node, the final tree that we obtain can be different. Since there can be some classes that are correlated, a particular tree may perform better than the others. However, we can hardly identify the best tree without deep domain knowledge. The END method copes with this problem by building multiple dichotomy trees randomly during learning, and then combining the predictions made by each tree during classification. The END method is generally known to perform well even when the base learner is unable to model complex decision boundaries As the base classifier at each node of the dichotomy, we have used another ensemble classifier called the random forest. A random forest is built by repeatedly generating a decision tree each time with a different random subset of features using a bootstrap sample. By combining bagging with random feature subset selection, a random forest enjoys the advantage of having more diverse ensemble members than a simple bagging. As an overall result, our ensemble of nested dichotomy can actually be seen as a committee of committees of decision trees that can deal with a multi-class problem with high accuracy. The ten classes of activities that we distinguish in this paper are 'Sitting', 'Standing', 'Walking', 'Running', 'Walking Uphill', 'Walking Downhill', 'Running Uphill', 'Running Downhill', 'Falling', and 'Hobbling'. The features used for classifying these activities include not only the magnitude of acceleration vector at each time point but also the maximum, the minimum, and the standard deviation of vector magnitude within a time window of the last 2 seconds, etc. For experiments to compare the performance of END with those of other methods, the accelerometer data has been collected at every 0.1 second for 2 minutes for each activity from 5 volunteers. Among these 5,900 ($=5{\times}(60{\times}2-2)/0.1$) data collected for each activity (the data for the first 2 seconds are trashed because they do not have time window dat
ARMA-PL : 시계열 데이터에 나타나는 중첩된 주기 및 선형추세에 대한 고찰
[Kisti 연계] 한국산업경영시스템학회 Journal of the Society of Korea Industrial and Systems Engineering Vol.33 No.2 2010 pp.112-126
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
시계열데이터는 ARMA 분석에 적합지 않은 요소를 내재하고 있는 경우가 있다. 특히 선형성과 주기성을 가진 요소가 확률적인 분포와 자주 혼재되어 있다. 이 논문에서는 이런 선형적 주기적 요소를 찾아내고 분석하는 방법을 제시한다. 특히 주기적 요소는 여러 주기가 층층이 겹쳐져서 나타난다. 주기 간에는 서로 일정 정수비율을 유지하며, 한 주거 안에 다른 주기가 내포되어 있는 경우(nested periods)가 많다. 시간규모(time-scale)개념을 도입하여 이러한 주기적 요소를 개념적으로 정립하고자 했다. 선형적 요소와 주기적 요소가 제거된 후 추출된 데이터는 MA-approximation이라는 방법을 사용하여 가장 데이터에 근접한 ARMA 모텔을 찾아낸다. 마지막으로 선형적 주기적 요소와 ARMA 추정결과를 종합하여 control boundary를 결정하는 방법을 제시한다.
힐버트-황 변환을 이용한 시계열 데이터 관리한계 : 중첩주기의 사례
[Kisti 연계] 한국산업경영시스템학회 Journal of the Society of Korea Industrial and Systems Engineering Vol.37 No.4 2014 pp.35-41
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Real-life time series characteristic data has significant amount of non-stationary components, especially periodic components in nature. Extracting such components has required many ad-hoc techniques with external parameters set by users in a case-by-case manner. In this study, we used Empirical Mode Decomposition Method from Hilbert-Huang Transform to extract them in a systematic manner with least number of ad-hoc parameters set by users. After the periodic components are removed, the remaining time-series data can be analyzed with traditional methods such as ARIMA model. Then we suggest a different way of setting control chart limits for characteristic data with periodic components in addition to ARIMA components.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.