년 - 년
데이터마이닝을 이용한 모바일앱 구분 별 촉진 전략에 관한 연구 KCI 등재
한국디지털정책학회 디지털융복합연구 제10권 제5호 2012.06 pp.339-349
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
모바일 컨버전스로 대표되는 스마트폰의 등장 이래 급속한 보급으로 국민의 절반이상이 사용하고 있다. 이에 따라 스마트폰에서 구동되는 모바일앱 시장도 빠른 속도로 성장하고 있다. 하지만 스마트폰과 모바일앱에 대한 기존 연구는 대부분 기술 수용 또는 기능 향상 등에만 초점을 맞춰 연구가 진행되었다. 따라서 본 연구는 모바일앱 각각에 대한 촉진 전략을 제시하기 위해 세 단계에 걸쳐 분석이 진행된다. 첫째, 빈도분석을 통해 가장 많이 사용하는 모바일앱을 도출하고 둘째, 연관성 규칙을 통해 모바일앱 간 관계에 대해 알아본다. 마지막으로 규칙유도기법은 앞선 두 단계에서 획득된 5개 모바일앱을 목표 변수로 설정하여, 각 모바일앱별 스마트폰 사용 특성을 도출하였다. 연구에 사용된 변수는 모바일앱 구분 20개와 인구통계학적 변수 및 PC, 영화, 음악, 도서, 게임의 사용 시간 그리고 지불 금액을 변수로 선정하여 총 35개가 사용되었다.
Smartphones which represent to the Mobile convergence, is the rapid spread, more than half of Korean are using. Accordingly, mobile apps that run on smartphone market is growing at a rapid pace. However, most studies on smartphone and mobile apps are focusing on the technology acceptance and improvement of function. So, this study is to suggest promotion strategy to each mobile apps, analyzed through three phases. First phase is the frequency analysis that deduct most frequently used mobile apps. Second phase is association rules that found to associate between mobile apps. Finally, to analyze deduction techniques for acquired 5 mobile apps to target variable in pre-2 phase use total 35 variables of 20 mobile apps categories, demographic variables, amount of PC, movie, music, book, game usage and fees per month.
사회적 기업의 효율성 예측 : 데이터 마이닝 기법을 이용하여 KCI 등재
한국경영컨설팅학회 경영컨설팅연구 제20권 제3호 통권 제66호 2020.08 pp.125-133
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
사회적 기업은 영리적 이익뿐만 아니라 현대의 영리 기업이 추구 할 수 없는 비영리적 가치를 실현하고자 설립된다. 사회적 기업은 사회적 목 적 실현과정에서 많은 생산손실 및 수입손실이 발생한다. 따라서 정부의 지원은 반드시 필요하나, 정부의 지원은 제한이 있다. 즉, 사회적 기업은 사회적 서비스 제공뿐만 아니라 경제적 능력까지 갖춰야한다. 정부에서는 사회적 목적과 경제적 목적을 효율적으로 달성하는 기업을 예측해서 지원하면 제한된 자원을 효과적으로 사용 가능할 것이다. 본 연구는 효율적인 사회적 기업을 판별하기 위해 자료포락분석(DEA)을 이용하였으 며, 효율적인 사회적 기업이 가지는 패턴을 찾기 위해 데이터마이닝 기법을 사용하였다. 분석결과, 약 19%의 사회적 기업이 효율적으로 분석되 었다. 분석된 사회적 기업을 종속변수로 구성하여 데이터 마이닝을 실시하였다. 우수한 데이터 마이닝 기법을 도출하기 위해 민감도, 특이도, 정 확도, G-mean, AUC를 분석하였다. 분석결과, 민감도를 제외한 특이도, 정확도, G-mean 그리고 AUC에서 효율적인 사회적 기업을 예측하는 가 장 우수한 데이터 마이닝 기법은 의사결정나무로 나타났다. 본 연구를 통해 정책 의사결정자는 효율적 사회적 기업을 예측할 수 있으며 효율적 인 운영을 하는 기업에 대해 효과적인 지원을 할 수 있을 것이다. 또한 의사결정나무의 장점인 시각화를 통해 더 쉬운 의사결정을 지원할 수 있을 것이다. 지원에 대한 차후 사회적 기업뿐만 아니라 기타 영역으로 까지 확장하여 의사결정지원에 도움을 줄 수 있을 것이다.
Many social enterprises have recently been established, and they need to be supported by the government. Governments should anticipate and support social enterprises that effectively achieve social and economic objectives in order to effectively use limited resources. This study used data envelopment analysis (DEA) to identify effective social enterprises and data mining techniques to find patterns of effective social enterprises. As a result of the analysis, about 19% of social enterprises were analyzed as efficient social enterprises. Data mining was conducted by configuring the analyzed social enterprise as a dependent variable. Sensitivity, specificity, accuracy, G-mean, and AUC were analyzed to derive excellent data mining techniques. As a result of the analysis, the decision tree was the best data mining technique for predicting efficient social enterprises in specificity, accuracy, G-mean and AUC excluding sensitivity. Through this study, policy decision makers can predict efficient social enterprises and provide effective support to enterprises that operate efficiently. In addition, it will be possible to support easier decision making through visualization, which is the advantage of decision trees. In the future, support can be extended to not only social enterprises but also other areas to help support decisionmaking.
상권별 편의점 도시락 판매 전략에 관한 연구 KCI 등재
한국디지털정책학회 디지털융복합연구 제17권 제6호 2019.06 pp.77-91
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
본 연구는 편의점 도시락 판매 전략을 수립하기 위하여 4개의 상권에 위치한 편의점에서 획득된 POS 자료를 바탕으로 연관성규칙 방법을 이용하여 분석을 진행하였다. 이를 위해 분석에 사용된 자료를 도시락이 가장 많이 판매되 는 시간대인 오전 6시부터 8시까지와 오후 17시부터 19시까지의 시간대별로 그리고 상권별편의점으로 구분하여 분석 하였다. 분석 결과 일반적으로 도시락과 함께 판매되는 상품들을 유음료, 음료 그리고 면과 같이 도시락과 함께 식사할 수 있는 제품들이 주를 이루었고 이것들은 상권이나 시간대에 차이가 없이 판매되고 있었다. 반면, 이외의 상품들을 오전 시간대와 오후 시간대 그리고 편의점이 위치한 상권별로 도시락과 함께 판매되는 상품의 종류와 개수가 차이가 있음을 확인할 수 있었다. 이와 같은 결과와 접근 방법은 편의점뿐만 아니라 사회 문화적 변화에 따라 변화하는 상품이 나 서비스의 수요를 특성을 발견하고 이에 대응하는데 기여할 수 있을 것으로 기대된다.
In order to establish a sales strategy for convenience store lunches, this study conducted analysis using association rules based on POS data obtained from convenience stores located in four commercial districts. For this purpose, the data used in the analysis were divided into the time zones from 6:00 am to 8:00 pm, 17:00 pm to 19:00 pm, and the convenience stores according to the commercial areas. As a result of the analysis, it was found that products that were sold together with a lunch box were mainly made of products that could be eaten together with lunch such as milk, beverage, and cotton. However, it was confirmed that there were differences in the types and numbers of the products that were sold together with the lunch boxes of the morning time and the afternoon hours for the other products. These results and approaches are expected to contribute to finding and responding to the needs for goods and services that change as well as convenience stores as well as sociocultural changes.
인공신경망을 이용한 청소년의 또래 애착 영향 요인 탐색 KCI 등재
한국융합학회 한국융합학회논문지 제8권 제10호 2017.10 pp.209-214
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 다층 퍼셉트론 인공신경망을 이용하여 우리나라 중학생의 또래애착에 영향을 미치는 요인을 탐색하였다. 2016년 지역아동센터의 아동패널조사에 참여한 중학교 3학년 재학생 419명(남 210명, 여 209명)을 분 석하였다. 종속변수는 또래애착 여부로 정의하였고, 설명변수는 성, 학업 성적 만족도, 주관적 가구경제수준, 학교생 활에 대한 부모-자녀대화 빈도, 주관적 건강상태, 우울증상, 자아존중감, 주관적 생활 만족도, 휴대전화의존도를 포 함하였다. 또래애착의 예측 요인은 다층 퍼셉트론 인공신경망 알고리즘을 이용하여 분석하였다. 분석 결과, 우울증 상, 성, 학교생활에 대한 부모-자녀 대화 수준, 주관적 가구 경제수준, 주관적 건강상태는 청소년의 또래애착과 관련 이 높은 요인이었다. 청소년기의 성공적인 사회관계 형성을 위해서 또래 애착에 주요한 영향을 미치는 요인들을 고려한 상담 및 교육 프로그램의 개발이 요구된다.
The aim of the present study was to analyze the factors that affects the peer attachment in Korean youth. Subjects were 419 middle school students (210 male, 209 female). Dependent variable was defined as peer attachment. Explanatory variables were included as gender, academic achievement satisfaction, subjective household economy level, parent - child dialogue frequency, subjective health status, depression symptom, self - esteem, subjective life satisfaction, and mobile phone dependency. In the multi-layer perceptron artificial neural network algorithm analysis, depression symptoms, gender, parent-child dialogue level for school life, subjective household economy level, subjective health status were significantly associated with peer attachment in Korean youth. Based on this result, systematic programs are required in order to prevention of peer attachment in Korean youth.
인공신경망 분석과 결정트리 융합에 의한 금연 프로그램 참여 결정 요인 KCI 등재후보
한국융합학회 한국융합학회논문지 제6권 제2호 2015.04 pp.25-30
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
이 연구는 신뢰성 있는 국가통계 데이터를 이용하여 지난 1년 간 금연 시도 경험 결정 요인 모형을 구축하 고 개발된 모형을 근거로 금연 프로그램 참여 표적 집단 예측에 관한 기초 자료를 제공하였다. 분석대상은 2010년 서울시복지패널조사를 완료한 19세 이상 흡연자 1,326명이다. 결과변수는 지난 1년간 금연 시도 경험으로 정의하였 고, 설명변수는 연령, 성, 최종학력, 현재 취업 상태, 가구 월 평균 총 소득, 배우자 유무, 음주 여부, 주관적 건강상태, 정기적인 운동 여부, 지난 한 달 간 우울증상 여부, 현재 순환기, 내분비계, 근골격계, 호흡기계, 이비인후 질환, 간질 환, 비뇨기계질환 등 질병 여부로 설정하였다. 분석방법은 인공신경망 분석과 결정트리모형을 이용하였다. CART 알고리즘을 이용한 금연 프로그램 참여 모형을 구축한 결과, 유의미한 요인은 질병 여부, 주관적 건강 상태, 가구 월 평균 총소득이었다. 이 결과를 기초로 금연 프로그램의 성공적인 시행을 위해서 표적 대상의 특성을 고려한 프로 그램 개발 및 교육이 요구된다.
The purpose of this study was to analyze the factors that affects the participating in a smoking cessation program. Data were from the A Study on the Seoul Welfare Panel Study 2010. Subjects were 1,326 smokers aged 19 and older living in the community. Dependent variable was defined as experience of smoking cessation. Explanatory variables were included as age, gender, level of education, employment status, household income, marital status, drinking, self-reported health status, depression, disease, and physical activity. A prediction model was developed by the use of a Decision Tree and Neural Network Algorithm. In the Prediction model, self reported health status, disease, income, household income were significantly associated with participating in a smoking cessation program. Based this study, systematic education and development of programs are required.
급성 뇌졸중 환자의 중증도 보정 재원일수 변이에 관한 연구 KCI 등재
한국디지털정책학회 디지털융복합연구 제11권 제6호 2013.06 pp.221-233
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
본 연구는 뇌졸중 환자의 효율적인 재원일수 관리를 위해 행정자료를 이용하여 우리나라 의료기관을 이용한 뇌졸중 입원환자의 중증도 보정 적정 재원일수 예측 모형을 개발하고 이를 의료기관에서 활용할 수 있는 방안을 제 시하고자 하였다. 이를 위해 2004-2009년 퇴원손상심층조사 자료 중 뇌졸중 입원환자 23,134명을 대상으로 데이터마 이닝 기법을 이용하여 뇌졸중 입원환자의 적정 재원일수 예측모형을 개발하였다. 의사결정나무 모형에 따라 뇌졸중 입원환자의 평균재원일수에 가장 큰 영향을 미치는 변수는 뇌졸중 발생유형이었으며, 의사결정나무를 이용하여 개발 된 뇌졸중 입원환자의 중증도 보정 재원일수 모형 결과, 적정 재원일수와 실제 재원일수의 차이는 진료비지불방법, 의료기관 소재지, 병상규모가 모두 통계적으로 유의하게 나타났다. 따라서 뇌졸중 입원환자의 재원일수 변이를 줄이 고 효율적으로 관리하기 위해서는 개발된 모형을 의료기관의 의료정보시스템에 적용하고 관리하는 활동을 전개해야 할 것이다.
This study aims to develop the severity-adjusted length of stay(LOS) model for acute stroke patients using data from the hospital discharge survey and propose management of length of stay(LOS) for acute stroke patients and using for Hospital management. The dataset was taken from 23,134 database of the hospital discharge survey from 2004 to 2009. The severity-adjusted LOS model for the acute stroke patients was developed by data mining analysis. From decision making tree model, the main reasons for LOS of acute stroke patients were acute stroke type. The difference between severity-adjusted LOS from the decision making tree model and real LOS was compared and it was confirmed that insurance type and bed number of hospital, location of hospital were statistically associated with LOS. And to conclude, hospitals should manage the LOS of acute stroke patients applying it into the medical information system.
신경망을 이용한 초등학생 컴퓨터 활용 능력 예측 KCI 등재후보
한국정보교육학회 정보교육학회논문지 제12권 제3호 2008.09 pp.267-274
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
신경망은 데이터로부터 반복적인 학습 과정을 통해 숨어 있는 패턴을 찾아내고, 새로운 데이터의 목표값에 대한 정확한 예측에 유용한 모델링 기법이다. 본 논문은 개인적인 특성, 가정․사회적 환경, 타 교과 성적을 이용하여 학생의 컴퓨터 활용 능력 예측을 위한 다층 인식모형(MLP) 신경망을 구축하였다. 신경망의 인식률은 예측 방법으로 널리 활용되고 있는 로지스틱 회귀분석 모델과 비교하였다. 개발한 신경망에 대한 실험 결과, 개인적인 특성이 학생들의 컴퓨터 활용 능력을 가장 잘 설명하는 요소이며, 반면 가정․사회적 환경은 가장 낮은 예측 요소임을 발견하였다. 또한 본 연구의 신경망 모델은 회귀분석보다 더욱 높은 인식률을 나타냈다.
A neural network is a modeling technique useful for finding out hidden patterns from data through repetitive learning process and for predicting target values for new data. In this study, we built multilayer perceptron neural networks for prediction of the students' computer literacy based on their personal characteristics, home and social environment, and academic record of other subjects. Prediction performance of the network was compared with that of a widely used prediction method, the regression model. From our experiments, it was found that personal characteristic features best explained computer proficiency level of a student, whereas the features of home and social environment resulted in the worse prediction accuracy among all. Moreover, the developed neural network model produced far more accurate prediction than the regression model.
데이터 스트림에서 데이터 마이닝 기법 기반의 시간을 고려한 상대적인 빈발항목 탐색 KCI 등재후보
한국정보교육학회 정보교육학회논문지 제9권 제3호 2005.09 pp.453-462
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 들어 저장장치의 발전과 네트워크의 발달로 인하여 대용량의 데이터에 내재되어 있는 정보를 빠른 시간 내에 처리하여 새로운 지식을 창출하려는 요구가 증가하고 있다. 연속적이고 빠르게 증가하는 데이터를 지칭하는 데이터 스트림에서 데이터 마이닝 기법을 이용하여 시간이 흐름에 따라 변하고, 무한적으로 증가하는 데이터 스트림에서의 빈발항목을 찾는 연구가 활발하게 진행되고 있다. 하지만 기존의 연구들은 시간의 흐름에 따른 빈발항목 탐색방법을 적절히 제시하지 못하고 있으며 단지 집계를 이용하여 빈발항목을 탐색하고 있다. 본 논문에서는 데이터 스트림에서 시간적 측면을 고려하여 상대적인 빈발항목을 탐색하기 위한 새로운 알고리즘으로 한정적인 메모리를 고려하여 빈발항목과 부분 빈발항목만을 저장하고 시간의 흐름에 따른 빈발항목의 갱신방법에 관하여 제안하였다. 논문에서 제안하는 알고리즘의 성능은 다양한 실험을 통해서 검증된다. 제안된 방법은 웹 코스웨어로 학습하는 학생들의 행동패턴을 시간대별로 파악하여 빈발항목 및 상대적인 빈발항목을 탐색함으로써 학생들의 학습효과 증진 및 지도 방향을 설정하는데 활용할 수 있다.
Recently, due to technical improvements of storage devices and networks, the amount of data increase rapidly. In addition, it is required to find the knowledge embedded in a data stream as fast as possible. Huge data in a data stream are created continuously and changed fast. Various algorithms for finding frequent itemsets in a data stream are actively proposed. Current researches do not offer appropriate method to find frequent itemsets in which flow of time is reflected but provide only frequent items using total aggregation values. In this paper we proposes a novel algorithm for finding the relative frequent itemsets according to the time in a data stream. We also propose the method to save frequent items and sub-frequent items in order to take limited memory into account and the method to update time variant frequent items. The performance of the proposed method is analyzed through a series of experiments. The proposed method can search both frequent itemsets and relative frequent itemsets only using the action patterns of the students at each time slot. Thus, our method can enhance the effectiveness of learning and make the best plan for individual learning.
데이터 마이닝을 통한 네트워크 이벤트 감사 모듈 개발 KCI 등재후보
한국융합보안학회 융합보안논문지 제5권 제2호 2005.06 pp.1-8
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 새로운 공격기법에 대한 대응방법의 하나로 네트워크 상황 즉, 네트워크 사용량을 분석을 통한 외부 공격 예방기법이 연구되고 있다. 이를 위한 네트워크 분석을 데이터 마이닝 기법을 통하여 네트워크 이벤트에 대한 연관 규칙을 주어 외부뿐만 아니라 내부 네트워크를 분석할 수 있는 기법이 제안되었다. 대표적인 데이터 마이닝 알고리즘인 Apriori 알고리즘을 이용한 네트워크 트래픽 분석은 과도한 CPU 사용시간과 메모리 요구로 인하여 효율성이 떨어진다. 본 논문에서는 이를 해결하기 위해서 새로운 연관 규칙 알고리즘을 제시하고 이를 이용하여 네트워크 이벤트 감사 모듈을 개발하였다. 새로운 알고리즘을 적용한 결과, Apriori 알고리즘을 적용한 시스템에 비해 CPU 사용시간과 메모리의 사용량에 있어 큰 향상을 보였다.
Network event analysis gives useful information on the network status that helps protect attacks. It involves finding sets of frequently used packet information such as IP addresses and requires real-time processing by its nature. Apriori algorithm used for data mining can be applied to find frequent item sets, but is not suitable for analyzing network events on real-time due to the high usage of CPU and memory and thus low processing speed. This paper develops a network event audit module by applying association rules to network events using a new algorithm instead of Apriori algorithm. Test results show that the application of the new algorithm gives drastically low usage of both CPU and memory for network event analysis compared with existing Apriori algorithm.
A Novel technique for the Detection of Mixed Noise in Medical Images using Datamining
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.8 No.11 2015.11 pp.231-242
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In Digital image processing; many researches have been done on image denoising so far. Nowadays, the noise detection from an image is the most challenging task. Though, the various algorithms introduced for the detection of noise type from a noisy image, but these algorithms work only for detection of single type of noise. To overcome the limitation of the previous built algorithms, we investigate the data mining technique called Support Vector Machine. The SVM is a powerful supervised learning method which is to be used for the detection of mixed noise models. Broadly, this technique detects the different types of noise from a mixed noise image; noise can be either single or mixed type of noise. The different parameters have combined to describe the properties of these different noise models so as to perform the detection. The detecting algorithm has been achieved by applying the SVM on the training dataset of different medical images and further extensive tests are performed on the test dataset for detection of each noise type model. This detection technique clearly outperforms various techniques with the high accuracy of results for different proposed noise models.
머신러닝 알고리즘 분석 및 비교를 통한 Big-5 기반 성격 분석 연구 KCI 등재
국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제19권 제4호 2019.08 pp.169-174
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
본 연구에서는 설문지를 이용한 데이터 수집과 데이터 마이닝에서 클러스터링 기법으로 군집하여 지도학습을 이용하여 유사성을 판단하고, 성격들의 상관 관계의 적합성을 분석하기 위해 특징 추출 알고리즘들과 지도학습을 이용하 는 것을 목표로 진행한다. 연구 수행은 설문조사를 진행 후 그 설문조사를 토대로 모인 데이터들을 정제하고, 오픈 소스 기반의 데이터 마이닝 도구인 WEKA의 클러스터링 기법들을 통해 데이터 세트를 분류하고 지도학습을 이용하여 유사성 을 판단한다. 그리고 특징 추출 알고리즘들과 지도학습을 이용하여 성격에 대해 적합한 결과가 나오는지에 대한 적합성 을 판단한다. 그 결과 유사성 판단에 가장 정확도 높게 도움을 주는 것은 EM 클러스터링으로 3개의 분류하고 Naïve Bayes 지도학습을 시킨 것이 가장 높은 유사성 분류 결과를 도출하였고, 적합성을 판단하는데 도움이 되도록 특징추출 과 지도학습을 수행하였을 때, Big-5 각 성격마다 문항에 추가되고 삭제되는 것에 따라 정확도가 변하는 모습을 찾게 되었고, 각 성격 마다 차이에 대한 분석을 완료하였다.
In this study, I use surveillance data collection and data mining, clustered by clustering method, and use supervised learning to judge similarity. I aim to use feature extraction algorithms and supervised learning to analyze the suitability of the correlations of personality. After conducting the questionnaire survey, the researchers refine the collected data based on the questionnaire, classify the data sets through the clustering techniques of WEKA, an open source data mining tool, and judge similarity using supervised learning. I then use feature extraction algorithms and supervised learning to determine the suitability of the results for personality. As a result, it was found that the highest degree of similarity classification was obtained by EM classification and supervised learning by Naïve Bayes. The results of feature classification and supervised learning were found to be useful for judging fitness. I found that the accuracy of each Big-5 personality was changed according to the addition and deletion of the items, and analyzed the differences for each personality.
Big 5 성격 요소와 머신 러닝 알고리즘을 통한 창의적인 사람들의 특징 연구 KCI 등재
국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제19권 제1호 2019.02 pp.97-102
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
창의적인 사람에 대한 정확한 기준이나 수치화를 사용하여 체계적인 분류와 분석 방법이 없었기에 정의하는 데에 어려움이 많다. 이 문제를 해결하기 위하여 본 연구에서는 창의적인 사람을 어떻게 구분 지을 수 있을지에 대한 것과 어떤 유사한 성격이 있는지 분석한다. 본 연구에서 우선 Big 5 성격 특성 기법을 이용하여 설문조사를 진행하고, 그 설문조사로 얻은 데이터 세트를 가지고 데이터 마이닝 도구인 WEKA를 이용하여 데이터 세트를 분류하고 분석한 뒤, 창의적인 사람들과 연관성 있는 성격 특징들을 다양한 머신 러닝 기법을 이용하여 분석하는 것을 목표로 진행하였다. 7개의 특징 선택 알고리즘 을 활용하고, 특징 선택 알고리즘들로 분류된 특징 집단을 선택하여 머신 러닝 알고리즘에 적용하여 정확도를 알아냈고, 서포 트 벡터 머신을 통해 나온 특징이 가장 높은 분류 결과를 도출하였다.
There are many difficulties to define because there is no systematic classification and analysis method using accurate criteria or numerical values for creative people. In order to solve this problem, this study attempts to analyze how to distinguish creative people and what kind of personality they have when distinguishing creative people. In this study, I first survey the Big 5 personality trait, classify and analyze the data set using the data mining tool WEKA, and then analyze the data set related to the creativity The goal is to analyze the features using various machine learning techniques. I use seven feature selection algorithms, select feature groups classified by feature selection algorithms, apply them to machine learning algorithms to find out the accuracy, and derive the results.
An Efficient Semantic Ranked Keyword Search of Big Data Using Map Reduce SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.6 2015.12 pp.47-56
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Information retrieval is fast becoming the prevailing form of information access, surpassing traditional database style searching. Ontologies have become the tool of choice employed in many information retrieval systems and more prominently in semantic information retrieval. In order to overcome the disadvantages in key word based information retrieval systems, which transfer irrelevant information, ontology has been designed. A system with ontology mimics the real world, where every task is laced with certain meaning as this is basic idea behind knowledge processing. Hadoop, which is an open source frame work for storing and processing large datasets, is used for pre-processing the text documents. First, a set of text documents are considered. Pre-processing is performed on a large domain of data using Hadoop MapReduce. This includes the removal of the stop words along with stemming and excluding less frequency words. Despite this pre-processing, owing to the colossal number of index terms still floating in the considered domain data, the problem of high dimensionality is encountered. Therefore the dimensionality of such a group of terms is reduced by identifying it as a concept and those concepts can be viewed as a single dimension in a ontology based information retrieval system. Now ontology is constructed by assigning synonym set to each concept in this structure using tools like word net. Thus constructed ontology can be mapped on to the processed query which gives us the relevant information from the data pool considered.
A Pipeline Model to Discover Frequent Itemsets in an Hierarchical Systems
보안공학연구지원센터(IJGDC) International Journal of Grid and Distributed Computing Vol.6 No.4 2013.08 pp.19-38
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Like all the other fields of data processing, the modern information systems have integrated the results of the advanced technologies of the last decades. These systems contain implicit data which it will be necessary to extract and exploit, by using data Mining techniques. Mining association rules which trends to find interesting association or correlation relationships among large amounts of data is one of these techniques. It is a two-steps process, the first step finds all frequent itemsets and the second step constructs association rules from these frequent sets. The overall performance of mining association rules is determined by the first step which becomes the focus problem. This step is expensive with high demands for computation and data access. Parallel computing seems to have a natural role to play since parallel computers provides scalability. In this paper, we examine the issue of mining association rules among items in large databases transactions using the algorithm Apriori proposed by Agrawal. In this context, we propose a new parallel version of the Apriori algorithm of Agrawal, that is the main algorithm of each data mining technique. Parallel computing seems to have a natural role to play since parallel computers provides scalability. In fact, our objective of our work is to have an efficient parallel execution time that requires a delicate balance between program granularity and communication latency (synchronization overhead) between the different granules. Unlike previous work on parallelization of specific data mining algorithms, our approaches consist to discover the different granularity levels of parallelism and their impact on the performance. In this paper we focus on task and data parallelism (hybrid approach) under distributed memory. In particular, if communication latency is minimal then fine grain partitioning will yield the best performance. This is the case when data parallelism is used. If communication latency is large (as in a loosely coupled system), then coarse grain partitioning is more appropriate. For the target architecture used in this work (distributed-shared memory).), the problem of load balancing among the nodes becomes a more critical issue in attempts to yield high performance. We have carried out a detailed evaluation of the parallelization techniques and the impact of combining different types of parallelism (task, data and pipeline) on the effectiveness of the system.
Travel Agency에서 CRM을 위한 DataMining, OLAP 적용
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2000 pp.152-154
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
World Wide Web(WWW) 데이터가 폭발적으로 증가하고 있는 시점에서, WWW의 데이터로부터 유용한 정보를 찾아내고 분석하는 일이 필요해졌다. 또한 WWW의 데이터만으로는 얻을 수 없는 기업의 의사결정을 위한 정보를 얻기 위해, 웹 페이지 접근 기록에서 얻어진 웹 로그기록들과 기업의 판매 트랜잭션 데이터베이스, 광고 데이터베이스 그리고 고객 정보를 통합하여 데이터 웨어하우스를 구축한다. 이러한 과정은 기업활동의 결과로 축적된 데이터 자원과 WWW의 데이터를 통합하여 체계적인 정보기반을 구축하고, 이러한 자원을 전략적으로 재활용하는 것이 목적이다. 본 논문에서는 WWW의 데이터와 기업의 데이터베이스를 통합하여 웨어하우스를 설계하고 여기에 데이터마이닝, OLAP을 적용하여 CRM에 활용하는 방안을 제안하고자 한다.
Datamining 기법을 활용한 단기 항만 물동량 예측
[Kisti 연계] 한국항만경제학회 한국항만경제학회지 Vol.37 No.3 2021 pp.1-17
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구에서는 항만의 단기 물동량을 예측하기 위해 ARIMA 모형과 CART 모형을 활용한 단기 수요예측 모형을 제시하였다. 제시한 모형은 2단계로 구성된다. 1단계에서는 시계열 예측치와 주요 교역국의 주당 근로일수를 변수로 사용하여 CART 모형을 추정하고 주별 물동량 예측을 진행한다. 2단계에서는 1단계에서 도출한 예측치와 요일 정보, 주요국 공휴일 정보, 주요국 행사 기간 정보를 설명변수로 활용하여 최종적인 일별 물동량 예측 모형을 추정한다. 제시한 수요예측 모형을 활용하여 2020년 10월 1일부터 12월 31일까지 92일의 부산항 물동량을 예측한 결과 제시한 모형의 평균 정확도가 기존 시계열 모형보다 '22.5%' 높은 것으로 나타났다. 제시 모형은 일별 물동량의 추세뿐만 아니라 물동량이 급등락하는 지점에서도 높은 정확도를 보였으며 시계열 예측 모형을 사용했을 때 비해 총 166,504(TEU)의 오차를 줄일 수 있는 것으로 나타났다. 항만의 효율적인 운영을 위해 필수적인 단기 물동량 예측에 적합한 예측 모형을 제시한 본 연구는 충분한 활용 가치가 있을 것으로 판단된다.
Forecasting the daily volume of container is important in many aspects of port operation. In this article, we utilized a machine-learning algorithm based on decision tree to predict future container throughput of Busan port. Accurate volume forecasting improves operational efficiency and service levels by reducing costs and shipowner latency. We showed that our method is capable of accurately and reliably predicting container throughput in short-term(days). Forecasting accuracy was improved by more than 22% over time series methods(ARIMA). We also demonstrated that the current method is assumption-free and not prone to human bias. We expect that such method could be useful in a broad range of fields.
Two-Step Filtering Datamining Method Integrating Case-Based Reasoning and Rule Induction
[Kisti 연계] 한국지능정보시스템학회 한국지능정보시스템학회 학술대회논문집 2007 pp.329-337
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Case-based reasoning (CBR) methods are applied to various target problems on the supposition that previous cases are sufficiently similar to current target problems, and the results of previous similar cases support the same result consistently. However, these assumptions are not applicable for some target cases. There are some target cases that have no sufficiently similar cases, or if they have, the results of these previous cases are inconsistent. That is, the appropriateness of CBR is different for each target case, even though they are problems in the same domain. Thus, applying CBR to whole datasets in a domain is not reasonable. This paper presents a new hybrid datamining technique called two-step filtering CBR and Rule Induction (TSFCR), which dynamically selects either CBR or RI for each target case, taking into consideration similarities and consistencies of previous cases. We apply this method to three medical diagnosis datasets and one credit analysis dataset in order to demonstrate that TSFCR outperforms the genuine CBR and RI.
Ananlyzing Customer Management Data by Datamining (Focused on Apartment Customer Classification)
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2004 pp.69-72
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
기업간의 경쟁이 심화되고 정보의 중요성에 대한 인식이 확대되어 가는 상황에서 다량의 데이터로부터 가치 있는 데이터를 추출하는 CRM 데이터 마이닝은 중대한 관심사가 아닐 수 없다. 본 연구는 데이터마이닝의 여러 활용 분야 중 고객세분화를 위해 최근 많이 사용되고 있는 데이터마이닝 기법인 로지스틱 회귀분석, 의사결정나무, 신경망 알고리즘 기법들을 비교하며, 이를 실제 아파트 고객의 데이터를 이용하여 검증하고자 한다. 따라서, 아파트 고객 세분화를 위한 데이터마이닝 수행시 기법 선택의 기준과 비교 평가의 기준을 제시하는 데 연구목적 있다.
Ananlyzing Customer Management Data by Datamining (Focused on Apartment Customer Classification)
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2004 pp.69-72
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
기업간의 경쟁이 심화되고 정보의 중요성에 대한 인식이 확대되어 가는 상황에서 다량의 데이터로부터 가치 있는 데이터를 추출하는 CRM 데이터 마이닝은 중대한 관심사가 아닐 수 없다. 본 연구는 데이터마이닝의 여러 활용 분야 중 고객세분화를 위해 최근 많이 사용되고 있는 데이터마이닝 기법인 로지스틱 회귀분석, 의사결정나무, 신경망 알고리즘 기법들을 비교하며, 이를 실제 아파트 고객의 데이터를 이용하여 검증하고자 한다. 따라서, 아파트 고객 세분화를 위한 데이터마이닝 수행시 기법 선택의 기준과 비교 평가의 기준을 제시하는 데 연구목적 있다.
Risk Factors of Impaired Fasting Glucose and Type 2 Diabetes Mellitus - Using Datamining -
[Kisti 연계] 대한예방의학회 대한예방의학회 학술대회논문집 2005 p.429
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.