Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 75
No
1

GMM 지원을 위해 k-means 알고리즘을 이용한 어휘 인식 성능 개선 KCI 등재

이종섭

한국디지털정책학회 디지털융복합연구 제13권 제2호 2015.02 pp.135-140

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

일반적인 CHMM 어휘 인식 시스템은 어휘 인식에 대한 모델들의 관측 확률 인식률이 낮고, 일부 단위 음 소 모델에만 적용되어 제한적으로 사용되는 문제점이 있다. 또한, 어휘 탐색에서 어휘의 의미가 다양하여 탐색된 어 휘가 사용자의 요구에 부합되지 않는 문제점을 가진다. 이러한 문제를 개선하기 위해 GMM(Gaussian Mixture Model) 을 이용한 음소인식을 수행하고, 개선된 k-means 알고리즘을 이용하여 어휘 특성에 따른 제한적인 탐색 문제점을 해 결하였다. 성능 실험은 기존의 시스템과 비교하여 정확도와 재현율로 대변되는 효과성을 측정하였으며, 성능 실험 결 과 정확도는 83%, 재현율은 67%로 나타났다.

General CHMM vocabulary recognition system is model observation probability for vocabulary recognition of recognition rate's low. Used as the limiting unit is applied only to some problem in the phoneme model. Also, they have a problem that does not conform to the needs of the search range to meaning of the words in the vocabulary. Performs a phoneme recognition using GMM to improve these problems. We solve the problem according to the limited search words characterized by an improved k-means algorithm. Measure the effectiveness represented by the accuracy and reproducibility as compared to conventional system performance experiments. Performance test results accuracy is 83%p, and recall is 67%p.

2

추천시스템을 위한 k-means 기법과 베이시안 네트워크를 이용한 가중치 선호도 군집 방법 KCI 등재

박화범, 조영성, 고형화

한국정보기술응용학회 JITAM Vol.20 No.3 2013.09 pp.219-230

※ 기관로그인 시 무료 이용이 가능합니다.

4,300원

Real time accessiblity and agility in Ubiquitous-commerce is required under ubiquitous computing environment. The Research has been actively processed in e-commerce so as to improve the accuracy of recommendation. Existing Collaborative filtering (CF) can not reflect contents of the items and has the problem of the process of selection in the neighborhood user group and the problems of sparsity and scalability as well. Although a system has been practically used to improve these defects, it still does not reflect attributes of the item. In this paper, to solve this problem, We can use a implicit method which is used by customer’s data and purchase history data. We propose a new clustering method of weighted preference for customer using k-means clustering and Bayesian network in order to improve the accuracy of recommendation. To verify improved performance of the proposed system, we make experiments with dataset collected in a cosmetic internet shopping mall.

3

K-means 알고리즘 기반 클러스터링 인덱스 비교 연구 KCI 등재

심요성, 정지원, 최인찬

한국경영정보학회 Asia Pacific Journal of Information Systems 제16권 제1호 2006.03 pp.127-144

※ 기관로그인 시 무료 이용이 가능합니다.

5,200원

The K-means algorithm is widely used at the initial stage of data analysis in data mining process, partly because of its low time complexity and the simplicity of practical implementation. Cluster validity indices are used along with the algorithm in order to determine the number of clusters as well as the clustering results of datasets. In this paper, we present a performance comparison of sixteen indices, which are selected from forty indices in literature, while considering their applicability to nonhierarchical clustering algorithms. Data sets used in the experiment are generated based on multivariate normal distribution. In particular, four error types including standardization, outlier generation, error perturbation, and noise dimension addition are considered in the comparison. Through the experiment the effects of varying number of points, attributes, and clusters on the performance are analyzed. The result of the simulation experiment shows that Calinski and Harabasz index performs the best through the all datasets and that Davis and Bouldin index becomes a strong competitor as the number of points increases in dataset.

4

A Novel Multilayer Data Clustering Framework based on Feature Selection and Modified K-Means Algorithm

Ganglong Duan, Wenxiu Hu, Zhiguang Zhang

보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.9 No.4 2016.04 pp.81-90

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

With the rapid development of computer science and technology, the data analysis technique has been a hottest research area in the pattern recognition research community. Cluster analysis is an important step in data mining. For clustering, various multi-objective techniques are evolved, which can automatically partition the data. In this paper, we propose a novel multilayer data clustering framework based on feature selection and modified K-Means algorithm. To facilitate the clustering, the proposed algorithm selects a representative feature subset to reduce the dimension of the raw data set. Besides, the selected feature subset has fewer missing values than the raw data set, which may improve the cluster accuracy. Another unique property of the proposed algorithm is the use of partial distance strategy. The experimental analysis and simulation indicate the feasibility and robustness of our method, in the future, we plan to conduct more mathematical analysis to modify our algorithm to achieve better result.

5

Improved K-means Algorithm with the Pretreatment of PCA Dimension Reduction

Hongtao Liu, Chen Fang, Yu Wu, Ke Xu, Tian Dai

보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.8 No.6 2015.06 pp.195-204

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

6

An Efficient K-Means Algorithm and its Benchmarking against other Algorithms SCOPUS

Anupama Chadha, Suresh Kumar

보안공학연구지원센터(IJGDC) International Journal of Grid and Distributed Computing Vol.9 No.11 2016.11 pp.119-132

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

K-Means is a widely used partition based clustering algorithm famous for its simplicity and speed. It organizes input dataset into predefined number of clusters. K-Means has a major limitation -- the number of clusters, K, need to be pre-specified as an input. Pre-specifying K in the K-Means algorithm sometimes becomes difficult in absence of thorough domain knowledge, or for a new and unknown dataset. This limitation of advance specification of cluster number can lead to “forced” clustering of data and proper classification does not emerge. In this paper, a new algorithm based on the K-Means is developed. This algorithm has advance features of intelligent data analysis and automatic generation of appropriate number of clusters. The clusters generated by the new algorithm are compared against results obtained with the original K-Means and various other famous clustering algorithms. This comparative analysis is done using sets of real data.

7

An Extended K-Means Algorithm using MapReduce Framework for Mixed Datasets SCOPUS

Anupama Chadha, Suresh Kumar

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.9 2016.09 pp.167-176

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

K-Means is a famous partition based clustering algorithm. Various extensions of K-Means have been proposed depending on the type of datasets being handled. Popular ones include K-Modes for categorical data and K-Prototype for mixed numerical and categorical data. The K-Means and its extensions suffer from one major limitation that is dependency on prior input of number of clusters K. Sometimes it becomes practically impossible to correctly estimate the optimum number of clusters in advance. Various ways have been suggested in literature to overcome this limitation for numerical data. But for categorical and mixed data work is still in progress. In this paper, we introduce a new algorithm based on the K-Means that takes mixed dataset as an input and generates appropriate number of clusters on the run using MapReduce programming style. The new algorithm not only overcomes the limitation of providing the value of K initially but also reduces the computation time using MapReduce framework.

8

An Optimized k-means Algorithm for Selecting Initial Clustering Centers SCOPUS

Jianhui Song, Xuefei Li, Yanju Liu

보안공학연구지원센터(IJSIA) International Journal of Security and Its Applications Vol.9 No.10 2015.10 pp.177-186

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Selecting the initial clustering centers randomly will cause an instability final result, and make it easy to fall into local minimum. To improve the shortcoming of the existing kmeans clustering center selection algorithm, an optimized k-means algorithm for selecting initial clustering centers is proposed in this paper. When the number of the sample’s maximum density parameter value is not unique, the distance between the plurality samples with maximum density parameter values is calculated and compared with the average distance of the whole sample sets. The k optimized initial clustering centers are selected by combing the algorithm proposed in this paper with maximum distance means. The algorithm proposed in this paper is tested through the UCI dataset. The experimental results show the superiority of the proposed algorithm.

9

An Improved K-means Algorithm based on Mapreduce and Grid

Li Ma, Lei Gu, Bo Li, Yue Ma, Jin Wang

보안공학연구지원센터(IJGDC) International Journal of Grid and Distributed Computing Vol.8 No.1 2015.02 pp.189-200

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The traditional K-means clustering algorithm is difficult to initialize the number of clusters K, and the initial cluster centers are selected randomly, this makes the clustering results very unstable. Meanwhile, algorithms are susceptible to noise points. To solve the problems, the traditional K-means algorithm is improved. The improved method is divided into the same grid in space, according to the size of the data point property value and assigns it to the corresponding grid. And count the number of data points in each grid. Selecting M(M>K) grids, comprising the maximum number of data points, and calculate the central point. These M central points as input data, and then to determine the k value based on the clustering results. In the M points, find K points farthest from each other and those K center points as the initial cluster center of K-means clustering algorithm. At the same time, the maximum value in M must be included in K. If the number of data in the grid less than the threshold, then these points will be considered as noise points and be removed. In order to make the improved algorithm can adapt to handle large data. We will parallel the improved k-mean algorithm and combined with the MapReduce framework. Theoretical analysis and experimental results show that the improved algorithm compared to the traditional K-means clustering algorithm has high quality results, less iteration and has good stability. Parallelized algorithm has a very high efficiency in data processing, and has good scalability and speedup.

10

A Network Intrusion Detection Model Based on K-means Algorithm and Information Entropy SCOPUS

Gao Meng, Li Dan, Wang Ni-hong, Liu Li-chen

보안공학연구지원센터(IJSIA) International Journal of Security and Its Applications Vol.8 No.6 2014.12 pp.285-294

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Many factors could influence the clustering performance of K-means algorithm, selection of initial cluster centers was an important one, traditional method had a certain degree of randomness in dealing with this problem, for this purpose, information entropy was introduced into the process of cluster centers selection, and a fusion algorithm combining with information entropy and K-means algorithm was proposed, in which, information entropy value was used to measure the similarity degree among records, the least similar record would be regarded as a cluster center. In addition, a network intrusion detection model was built, it could make cluster centers change dynamically along with the network changes, and the model could real-time update the cluster centers according to actual needs. Experiment results show that the improved algorithm proposed is better than the traditional K-means algorithm in detection ratio and false alarm ratio, and the network intrusion detection model is proved to be feasible.

11

Research of Customer Loyalty Based on the Improved K-means Algorithm

Li Min, Liu Wei, Chen Ming

보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.8 No.10 2015.10 pp.311-318

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

During the iteration process, the traditional K-means algorithm is easily fall into local optimal solution. In order to solve this problem, this paper proposed an improved Kmeans algorithm, and used the method of maximum distance equal division to select the initial cluster centers. We preset k cluster centers, and avoid it falling into local optimal solution. Apply this improved algorithm into e-commerce customer loyalty analysis, this paper put forward a customer loyalty analysis model using the parameters of shopping recency, shopping frequency, shopping monetary, customer satisfaction and customer attention, and used the improved K-means algorithm to analyze the RFMSA customer loyalty model. The studies show that the improved K-means algorithm and RFMSA model can effectively divide the loyalty of the e-commerce customer, it also can fully reflect the customer’s current value and potential value-added ability, and provide the basis that the e-commercial enterprises can adopt different marketing strategies for different target customers.

12

Liver Function Diagnosis Based on Artificial Bee Colony and K-Means Algorithm

Zhang Lin, Li Peng, Qiao Pei-li

보안공학연구지원센터(IJUNESST) International Journal of u- and e- Service, Science and Technology Vol.9 No.1 2016.01 pp.123-128

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The traditional K-Means clustering is sensitive to random selection of initial cluster centroids, easily into the local optimal solution. In this paper, an efficient aggregation algorithm which combined with Artificial bee colony and K-Means algorithm is proposed to apply to the diagnosis of liver function. The algorithm reduced the dependence on the initial cluster centroids and the probability to be trapped by local optimal solution, thus assigning data points to their appropriate cluster more efficient. The experimental results show that algorithm proposed in this paper is superior to the K-Means clustering in diagnosis of liver function.

13

Evaluation of the Selection of the Initial Seeds for K-Means Algorithm

ShinWon Lee, WonHee Lee

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.6 No.5 2013.10 pp.13-22

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Clustering method is divided into hierarchical clustering, partitioning clustering, and more. K-Means algorithm is one of partitioning clustering methods and is adequate to cluster a lot of data rapidly and easily. The problem is it is too dependent on initial centers of clusters and needs the time of allocation and recalculation. We compare random method, max average distance method and triangle height method for selecting initial seeds in K-Means algorithm. It reduces total clustering time by minimizing the number of allocation and recalculation.

14

A New Initialization Method to Originate Initial Cluster Centers for K-Means Algorithm

Yugal Kumar, G. Sahoo

보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology Vol.62 2014.01 pp.43-54

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

K means algorithm is most popular partition based algorithm that is widely used in data clustering. A Lot of algorithms have been proposed for data clustering using K-Means algorithm due to its simplicity, efficiency and ease convergence. In spite this K-Means algorithm has some drawbacks like initial cluster centers, stuck in local optima etc. In this study, a new method is proposed to address the initial cluster centers problem in K-Means algorithm based on binary search technique. Binary search technique is a popular searching method that is used to find an item in given list of array. So in proposed method, the initial cluster centers have obtained using binary search property and after that K-Means algorithm is applied to gain optimal cluster centers in dataset. The performance of the proposed algorithm is tested on the two benchmark dataset which are downloaded from the UCI machine learning repository and compared with Random, Hartigan and Wang, Ward, Build, Astrhan and Minkowaski ward methods. The proposed method is also applied on the Minkowaski weighted K-Means algorithm to prove its significance and effectiveness.

15

K-means Parallelization Algorithm Based on MapReduce SCOPUS

Shuguang Wang, Chao Jiang

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.8 2016.08 pp.21-30

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Spatial Cluster analysis is another important technique in the field of spatial data mining, especially the K-Means spatial clustering method, which can deal with spatial objects with geographical location and attribute. However, with the development of the information society, the spatial data grows explosively, but the serial algorithm has low computing efficiency and is difficult to process massive spatial data. Aiming at spatial with a double meaning of location and attribute, the paper designed and implemented K-Means spatial clustering parallel algorithm on Hadoop. Using Yahoo Weibo user data is to do clustering analysis. Finally, the visualization of clustering results was implemented by Google Map.

16

Research on K-MEANS Clustering Algorithm Based on HADOOP SCOPUS

Feng Hu

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.8 2016.08 pp.79-88

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

This paper proposes an improved clustering algorithm on the basis of the characteristics of sampling and density. The initial k value and initial center are determined by sampling and density, and parallel improvement is based on the HADOOP platform. Through the experiment, the improved K-Means algorithm has good parallelism.

17

An Improved Algorithm of Rough K-Means Clustering Based on Variable Weighted Distance Measure

Tengfei Zhang, Long Chen, Fumin Ma

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.7 No.6 2014.12 pp.163-174

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Rough K-means algorithm has shown that it can provides a reasonable set of lower and upper bounds for a given dataset. With the conceptions of the lower and upper approximate sets, rough k-means clustering and its emerging derivatives become valid algorithms in vague information clustering. However, the most available algorithms ignore the difference of the distances between data objects and cluster centers when computing new mean for each cluster. To solve this issue, an improved algorithm of rough k-means clustering based on variable weighted distance measure is presented in this article. Comparative experimental results of real world data from UCI demonstrate the validity of the proposed algorithm.

18

A Hybrid Mod K-Means Clustering with Mod SVM Algorithm to Enhance the Cancer Prediction

Rethina Kumar, Gopinath Ganapathy, Jeong-Jin Kang

국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.13 No.2 2021.05 pp.231-243

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

In Recent years the way we analyze the breast cancer has changed dramatically. Breast cancer is the most common and complex disease diagnosed among women. There are several subtypes of breast cancer and many options are there for the treatment. The most important is to educate the patients. As the research continues to expand, the understanding of the disease and its current treatments types, the researchers are constantly being updated with new researching techniques. Breast cancer survival rates have been increased with the use of new advanced treatments, largely due to the factors such as earlier detection, a new personalized approach to treatment and a better understanding of the disease. Many machine learning classification models have been adopted and modified to diagnose the breast cancer disease. In order to enhance the performance of classification model, our research proposes a model using A Hybrid Modified K-Means Clustering with Modified SVM (Support Vector Machine) Machine learning algorithm to create a new method which can highly improve the performance and prediction. The proposed Machine Learning model is to improve the performance of machine learning classifier. The Proposed Model rectifies the irregularity in the dataset and they can create a new high quality dataset with high accuracy performance and prediction. The recognized datasets Wisconsin Diagnostic Breast Cancer (WDBC) Dataset have been used to perform our research. Using the Wisconsin Diagnostic Breast Cancer (WDBC) Dataset, We have created our Model that can help to diagnose the patients and predict the probability of the breast cancer. A few machine learning classifiers will be explored in this research and compared with our Proposed Model “A Hybrid Modified K-Means with Modified SVM Machine Learning Algorithm to Enhance the Cancer Prediction” to implement and evaluated. Our research results show that our Proposed Model has a significant performance compared to other previous research and with high accuracy level of 99% which will enhance the Cancer Prediction.

19

Computer Aided Diagnosis Based on K-means Collaborative Filtering Algorithm

Feng Xue-yuan, Li Peng, Qiao Pei-li

보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.9 No.4 2016.04 pp.59-68

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

In computer aided diagnosis (CAD) process, one of the most challenging problems is data sparsity, which leads to the diagnosis results are not reliable. This paper proposes a clustering collaborative filtering based algorithm to solve the problem of data sparsity. In this paper, we use k-means clustering algorithm to cluster the same type of patients, and then adopt collaborative filtering method to fill the missing data values for each cluster, in this way to reduce the complexity of similarity calculation of collaborative filtering. The proposed method makes full use of the information-sharing mechanism of "similar patient population" to predict and fill the missing values. A hepatitis dataset is used for evaluating the performance of the algorithm. Results indicate that the proposed algorithm has better performance for medical record data sparsity problem.

20

An Optimizing Algorithm of Non-Linear K-Means Clustering SCOPUS

Chen Ning, Zhang Hongyi

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.4 2016.04 pp.97-106

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Kernel K-means clustering (KKC) is an effective nonlinear extension of K-means clustering, where all the samples in the initial space are mapped into the feature space and then K-means clustering is performed based on the mapped data. However, all the mapped data are expressed by the implicit form, which causes the initial cluster centers can’t be selected flexibly. Once the selected initial cluster centers aren’t suitable, it tends to fall into local optimal solutions and can’t guarantee stable result. Based on a standard orthogonal basis of the sub-space spanned by all the mapped data, a novel improving non-linear algorithm of KKC is presented in this paper. The novel algorithm can express the mapped data using the explicit form, which make it very flexible to select the initial cluster centers as the linear K-means clustering does. Moreover, the computational complexity of the presented algorithm is also significantly reduced compared to that of KKC. The results of simulation experiments illustrate the proposed method can eliminate the sensitivity to the initial cluster centers and simplify computational processing.

 
1 2 3 4
페이지 저장