년 - 년
한국ITS학회 한국ITS학회 학술대회 Towards a Connected Future : Innovations in Mobility Technology 연결된 미래를 향하여: 모빌리티 기술의 혁신 2025.04 pp.636-646
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
보안공학연구지원센터(IJBSBT) International Journal of Bio-Science and Bio-Technology Vol.7 No.1 2015.02 pp.61-70
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The advent of the information age and the rapid development of IT skills have led to the construction of massive databases, thus the current research focus is shifting to the efficient utilization of these vast volumes of stored information. Among the data mining algorithms that have been applied to this problem, neural networks can be used with various types, qualities, distributions, or volumes of data and they have high predictive power. Thus, neural networks are known to be the most useful and extensible algorithms, whereas logistic analysis has many constraints. In addition, neural networks obtain better results when the assumptions of linear discriminant analysis cannot be satisfied. The present study evaluated a multilayer perceptron (MLP) and radial-basis function network (RBFN), and their performance levels were compared with logistic regression based on cross-validation using the same data. The experiments showed that MLP delivered better performance than other methods in medical diagnostic applications where numerical data are used. MLP also performed better with the heart disease dataset using finely specified data types compared with the diabetes dataset using simple data types.
A Simple Security Architecture for Mobile Office SCOPUS
보안공학연구지원센터(IJSIA) International Journal of Security and Its Applications Vol.9 No.1 2015.01 pp.139-146
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The Mobile Offices services based on Bring Your Own Device (BYOD) are getting more popular as the use of smart phone grows rapidly. As the free wireless services are being spread, malicious codes can be spread easily. Therefore, how we will provide security in an environment of free wireless has become an issue. This paper suggests security architecture for the better mobile office security and presents required procedures and the analysis of the expected security enhancement.
An Improved Kernel Clustering Algorithm for Mixed-Type Data in Network Forensic SCOPUS
보안공학연구지원센터(IJSIA) International Journal of Security and Its Applications Vol.10 No.1 2016.01 pp.343-354
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Clustering algorithm is a common analysis technology for network forensics, which, lacking of any prior knowledge, can effectively find out the invasions by analyzing the collected real-time communication data flowing through the network. This paper proposed an improved dynamic kernel clustering algorithm for mixed numeric and categorical network communication data. First, centroid prototype based on the mean and distribution centroid was put forward to represent the cluster center. Then by using Gaussian kernel function, the paper introduced a new dissimilarity measure between the data object and the centroid prototype in combination with the significance of different categorical values. On this basis, the objective function was defined, which took into account both the compact degree in a cluster and the discrete degree among the clusters. After that an improved kernel clustering algorithm was designed. In the process of clustering, centroid prototype and the value of the clustering parameter dynamically updated for a better description of the characteristics of clusters’ change. Finally, in order to verify the feasibility and effectiveness of the algorithm, the paper further applied it to network forensics, and the experimental results showed that the method could mine the intrusion behavior more accurately.
Parallel Coordinate Plots of Mixed-Type Data
[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.15 No.4 2008 pp.587-595
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Parallel coordinate plot of Inselberg (1985) is useful for visualizing dozens of variables, but so far the plot's applicability is limited to the variables of numerical type. The aim of this study is to extend the parallel coordinate plot so that it can accommodate both numerical and categorical variables. We combine Hayashi's (1950, 1952) quantification method of categorical variables and Hurley's (2004) endlink algorithm of ordering variables for the parallel coordinate plot. In line with our former study (Kwak and Huh, 2008), we develop Andrews' type modification of conventional straight-lines parallel coordinate plot to visualize the mixed-type data.
Cluster Analysis with Balancing Weight on Mixed-type Data
[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.13 No.3 2006 pp.719-732
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
A set of clustering algorithms with proper weight on the formulation of distance which extend to mixed numeric and multiple binary values is presented. A simple matching and Jaccard coefficients are used to measure similarity between objects for multiple binary attributes. Similarities are converted to dissimilarities between i th and j th objects. The performance of clustering algorithms with balancing weight on different similarity measures is demonstrated. Our experiments show that clustering algorithms with application of proper weight give competitive recovery level when a set of data with mixed numeric and multiple binary attributes is clustered.
[Kisti 연계] 대한원격탐사학회 대한원격탐사학회 학술대회논문집 2003 pp.467-469
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This research aimed to develop forest type classification technique for the mixed forest with coniferous and broad-leaved species using the high resolution satellite data. QuickBird data was used as satellite data. The method of this research was to extract satellite data for every single tree crown using image segmentation technique, then to evaluate the accuracy of classification by changing grouping criteria such as tree species, families, coniferous or broad-leaved species, and timber prices. As a result, the classification of tree species and families level was inaccurate, on the other hand, coniferous or broad-leaved species and timber price level was high accurate.
혼합형 데이터 보간을 위한 디노이징 셀프 어텐션 네트워크
[Kisti 연계] 한국콘텐츠학회 한국콘텐츠학회논문지 Vol.21 No.11 2021 pp.135-144
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 데이터 기반 의사결정 기술이 데이터 산업을 이끄는 핵심기술로 자리 잡고 있는바, 이를 위한 머신러닝 기술은 고품질의 학습데이터를 요구한다. 하지만 실세계 데이터는 다양한 이유에 의해 결측값이 포함되어 이로부터 생성된 학습된 모델의 성능을 떨어뜨린다. 이에 실세계에 존재하는 데이터로부터 고성능 학습 모델을 구축하기 위해서 학습데이터에 내재한 결측값을 자동 보간하는 기법이 활발히 연구되고 있다. 기존 머신러닝 기반 결측 데이터 보간 기법은 수치형 변수에만 적용되거나, 변수별로 개별적인 예측 모형을 만들기 때문에 매우 번거로운 작업을 수반하게 된다. 이에 본 논문은 수치형, 범주형 변수가 혼합된 데이터에 적용 가능한 데이터 보간 모델인 Denoising Self-Attention Network(DSAN)를 제안한다. DSAN은 셀프 어텐션과 디노이징 기법을 결합하여 견고한 특징 표현 벡터를 학습하고, 멀티태스크 러닝을 통해 다수개의 결측치 변수에 대한 보간 모델을 병렬적으로 생성할 수 있다. 제안 모델의 유효성을 검증하기 위해 다수개의 혼합형 학습 데이터에 대하여 임의로 결측 처리한 후 데이터 보간 실험을 수행한다. 원래 값과 보간 값 간의 오차와 보간된 데이터를 학습한 이진 분류 모델의 성능을 비교하여 제안 기법의 유효성을 입증한다.
Recently, data-driven decision-making technology has become a key technology leading the data industry, and machine learning technology for this requires high-quality training datasets. However, real-world data contains missing values for various reasons, which degrades the performance of prediction models learned from the poor training data. Therefore, in order to build a high-performance model from real-world datasets, many studies on automatically imputing missing values in initial training data have been actively conducted. Many of conventional machine learning-based imputation techniques for handling missing data involve very time-consuming and cumbersome work because they are applied only to numeric type of columns or create individual predictive models for each columns. Therefore, this paper proposes a new data imputation technique called 'Denoising Self-Attention Network (DSAN)', which can be applied to mixed-type dataset containing both numerical and categorical columns. DSAN can learn robust feature expression vectors by combining self-attention and denoising techniques, and can automatically interpolate multiple missing variables in parallel through multi-task learning. To verify the validity of the proposed technique, data imputation experiments has been performed after arbitrarily generating missing values for several mixed-type training data. Then we show the validity of the proposed technique by comparing the performance of the binary classification models trained on imputed data together with the errors between the original and imputed values.
[Kisti 연계] 한국산업경영시스템학회 Journal of the Society of Korea Industrial and Systems Engineering Vol.33 No.1 2010 pp.114-120
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
데이터마이닝의 사전 단계에서 데이터의 차원(Dimensionality)을 줄이기 위한 단계로서 많은 요소선택(Feature Selection) 방법들이 개발되었다. 이 방법은 결과를 예측하거나 데이터를 설명하고자 할 때 어떤 요소들이 관련이 있는지를 결정하는 과정을 포함한다. 또한 이 방법은 데이터의 크기에 대한 확장성 (Scalability)를 향상시키며 학습 모델을 더욱 이해하기 쉽도록 줄 수 있다. 이 논문에서는 NP(Nested Partition) 방법을 사용한 최적화 기반의 새로운 요소선택 방법을 NP 구조의 기본적인 이론 근거와 함께 제안한다. 또 한 편으로 많은 요소선택 방법들이 다중 형태의 데이터를 처리하는데 한계를 가지고 있는데, NP 기반의 요소선택 방법에 다중 형태의 데이터를 처리할 수 있도록 하는 요소 성능 평가도구(Evaluators)를 도입하여 이를 극복하고자 한다. 또한 어떤 평가도구가 특정 데이터 형태에서 더욱 좋은 결과를 보이는지를 실험 결과와 함께 제시하였다.
[Kisti 연계] 한국통계학회 The Korean journal of applied statistics Vol.28 No.6 2015 pp.1147-1161
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
오늘날 데이터는 p-차원의 공간에서 점들로써 표현되는 전통적인 형태를 벗어나 시그널(signal), 함수, 이미지(image), 모양(shape) 등과 같은 다양한 형태의 자료들이 데이터로써 고려되고 분석되고있다. 그러한 종류의 새로운 종류의 데이터 중 하나로 심볼릭 데이터(symbolic data)를 고려할 수 있다. 심볼릭 데이터는 구간(interval), 히스토그램(histogram), 목록(list), 통계표, 분포, 또는 모형 등과 같은 다양한 형태들을 가질 수 있다. 지금까지의 연구가 주로 심볼릭 데이터의 각각의 형태별 자료를 고려했다면, 본 연구에서는 이를 확장하여 수집된 히스토그램과 멀티모달의 혼합된 형태로 이루어진 자료에 대한 계층 분할적 군집분석방법을 소개하고 이를 업종별 산업재해자료의 분석을 위해 이용한다.
Nowadays we are considering and analyzing not only classical data expressed by points in the p-dimensional Euclidean space but also new types of data such as signals, functions, images, and shapes, etc. Symbolic data also can be considered as one of those new types of data. Symbolic data can have various formats such as intervals, histograms, lists, tables, distributions, models, and the like. Up to date, symbolic data studies have mainly focused on individual formats of symbolic data. In this study, it is extended into datasets with both histogram and multimodal-valued data and a divisive clustering method for the mixed feature-type symbolic data is introduced and it is applied to the analysis of industrial accident data.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.