년 - 년
Big Data Analytics in Health Care by Data Mining and Classification Techniques
[NRF 연계] 한국통신학회 ICT Express Vol.8 No.2 2022.06 pp.250-257
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Big data is the compilation of enormous data that arrives from diverse sources for instance online transaction details, social media, sensor data, etc. Such assortment of enormous data develop into tough to evaluate by conventional processing relevance’s. By the development and upcoming latent in the healthcare business field, it is essential to analyze a enormous noisy data to get significant information. In healthcare system, the aim of this work is to evaluate the medical database of diabetes patients by a mixture of innovative hierarchical decision attention network, association rules (AR) and multiclass outlier classification with MapReduce framework. The association rule apriori algorithm in a MapReduce framework considers health data to create regulations. This is employed to discover the association among disease and their signs. This examination is made by means of UCI machine learning datasets of diabetes containing 50 attributes. The results of the proposed algorithm are offered by parameters for instance precision, accuracy, recall, and F-score.
Data Mining Techniques : More Accurate Classified Algorithm For Cardiopulmonary Diseases Prediction
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 8th International Conference on Next Generation Computing 2022 2022.10 pp.245-248
Data mining techniques develop a more accurate classification algorithm for patients classified as either normotensive, prehypertensive, or hypertensive. Logistic Model Tree, NBTree, and Bagging were chosen as the three classification models with tenfold cross-validation (LMT). Over 24 hours, we collected ABP readings from 1161 patients. To analyze the data, data mining techniques were used and a tool called WEKA. The data was analyzed based on age, gender, wake-up blood pressure, medication, sleep-up blood pressure, and overall blood pressure. According to bagging results, 886 cases (76.3 percent) are correctly classified, with 270 cases classified as pre-hypertensive, 436 cases as Normotensive, and 180 cases as hypertensive. NBTree's results show that 882 (75.9%) of the 1161 instances are correctly classified. Pre-hypertensive patients make up 256, normotensive patients 442, and hypertensive patients 184. Of the 1161 instances, the LMT algorithm correctly classified 878 (75.6 percent). According to the results, 275 people are pre-hypertensive, 431 are normotensive, and 172 are hypertensive. According to our findings, bagging is the most accurate classifier for the 24 hour ABP Monitoring dataset we used. Bagging achieves less overfitting because it focuses on global accuracy. It stabilizes and improves the accuracy of unstable methods compared to single classifiers.
Knowledge Extraction from Academic Journals Using Data Mining Techniques
한국디지털정책학회 디지털융복합연구 제3권 제1호 2005.06 pp.75-88
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
최근 우리는 인접학문 간 그리고 학계와 산업계간의 연구협조가 점차 증가하고 있음을 보아오고 있다. 이러한 현상은 특히 학술저널 간 지식의존성을 촉진하는 계기를 제공하고 있다고 할 수 있다. 본 논문의 목적은 관련저널 간 지식상호 의존성을 규명하고 저널지식의 구조화를 위하여 연관성 (association), 군집화, 링크분석 등 데이터마이닝 기법을 적용하는 방법론을 제시하는 것이다. 제시된 방법을 통하여 기대되는 점들은 1) 논문의 기본 속성인 키워드, 저자, 그리고 인용데이터를 통합하는 규칙 집합을 통하여 논문지식검색기능의 향상, 2) 키워드를 기반으로 관련 저널 간 그리고 저널내부의 군집분석으로 지식동향 파악, 3) Kleinberg (1999)의 권위와 허브 개념을 인용데이터 분석에 활용하여 기존의 양적 평가 기준인 영향력지수 (impact factor)의 문제점을 보완하며, 4) 특정 논문이나 저널의 지식파급과 관련한 영향력을 산출하는 잠재적 지식파급 지수를 제안하는 것이다.
A Comparison of Implicit and Explicit Techniques in Data Mining for Financial forecasting
한국경영정보학회 한국경영정보학회 정기 학술대회 SI 산업 및 기술 발전을 위한 `97 국제컨퍼런스 1997.08 pp.837-845
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
데이터마이닝을 이용한 운행패턴 분석방법에 대한 연구 KCI 등재
한국ITS학회 한국ITS학회논문지 제8권 제6호 통권26호 2009.12 pp.1-12
※ 기관로그인 시 무료 이용이 가능합니다.
4,300원
근래에는 경제운전에 대한 중요성이 점차 부각되고 있어 운전자의 운전 행태나 성향을 자동으로 분석한 후 경제운전을 위한 방법을 운전자에게 알려줄 수 있는 연구가 필요하다. 본 논문에서는 이를 위해 차량에 대한 운행일시, 운행거리, 운행시간, 주행속도, 공회전시간, 급가속/급감속 횟수, 연료소모량 등의 운행정보를 수집하였고, 데이터마이닝을 이용하여 운전자의 운행패턴이 경제운전에 어떤 영향을 미칠 수 있는지 분석하였다. 본 연구 결과는 주행 중 운전자에게 지속적으로 공회전과 과속 정보, 급가속/급감속 횟수를 차량 단말에 표현하여 제공하고, 공회전과 과속 비율이 일정 임계치를 초과할 경우 경고 정보를 제공함으로써 경제운전에 악영향을 미칠 수 있는 운전 습관을 미리 예방할 수 있는 방안에 활용할 수 있다.
Recently, as the importance of Economical Driving has been gradually growing up, the needs for research on automatic analysis of driving patterns that will ultimately provide drivers the methods for Economical Driving have been increasingly risen. Based on this purpose, we have executed two things in this paper. First, we have collected overall driving information such as date, distance, driving time, speed, idle time, sudden acceleration/deceleration count, and the amount of fuel consumption. Second, we have analyzed the influences of driving patterns on economical driving by employing the data mining techniques. These results can be applied in preventing bad driving patterns which will have consequently bad effects on Economical Driving in two aspects: by presenting some information on the terminal of the vehicles such as idle time, over-speed time, sudden acceleration/deceleration count continuously and by providing the drivers with alert information when the idle time ratio and the over-speed time ratio are excessive.
사회적 기업의 효율성 예측 : 데이터 마이닝 기법을 이용하여 KCI 등재
한국경영컨설팅학회 경영컨설팅연구 제20권 제3호 통권 제66호 2020.08 pp.125-133
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
사회적 기업은 영리적 이익뿐만 아니라 현대의 영리 기업이 추구 할 수 없는 비영리적 가치를 실현하고자 설립된다. 사회적 기업은 사회적 목 적 실현과정에서 많은 생산손실 및 수입손실이 발생한다. 따라서 정부의 지원은 반드시 필요하나, 정부의 지원은 제한이 있다. 즉, 사회적 기업은 사회적 서비스 제공뿐만 아니라 경제적 능력까지 갖춰야한다. 정부에서는 사회적 목적과 경제적 목적을 효율적으로 달성하는 기업을 예측해서 지원하면 제한된 자원을 효과적으로 사용 가능할 것이다. 본 연구는 효율적인 사회적 기업을 판별하기 위해 자료포락분석(DEA)을 이용하였으 며, 효율적인 사회적 기업이 가지는 패턴을 찾기 위해 데이터마이닝 기법을 사용하였다. 분석결과, 약 19%의 사회적 기업이 효율적으로 분석되 었다. 분석된 사회적 기업을 종속변수로 구성하여 데이터 마이닝을 실시하였다. 우수한 데이터 마이닝 기법을 도출하기 위해 민감도, 특이도, 정 확도, G-mean, AUC를 분석하였다. 분석결과, 민감도를 제외한 특이도, 정확도, G-mean 그리고 AUC에서 효율적인 사회적 기업을 예측하는 가 장 우수한 데이터 마이닝 기법은 의사결정나무로 나타났다. 본 연구를 통해 정책 의사결정자는 효율적 사회적 기업을 예측할 수 있으며 효율적 인 운영을 하는 기업에 대해 효과적인 지원을 할 수 있을 것이다. 또한 의사결정나무의 장점인 시각화를 통해 더 쉬운 의사결정을 지원할 수 있을 것이다. 지원에 대한 차후 사회적 기업뿐만 아니라 기타 영역으로 까지 확장하여 의사결정지원에 도움을 줄 수 있을 것이다.
Many social enterprises have recently been established, and they need to be supported by the government. Governments should anticipate and support social enterprises that effectively achieve social and economic objectives in order to effectively use limited resources. This study used data envelopment analysis (DEA) to identify effective social enterprises and data mining techniques to find patterns of effective social enterprises. As a result of the analysis, about 19% of social enterprises were analyzed as efficient social enterprises. Data mining was conducted by configuring the analyzed social enterprise as a dependent variable. Sensitivity, specificity, accuracy, G-mean, and AUC were analyzed to derive excellent data mining techniques. As a result of the analysis, the decision tree was the best data mining technique for predicting efficient social enterprises in specificity, accuracy, G-mean and AUC excluding sensitivity. Through this study, policy decision makers can predict efficient social enterprises and provide effective support to enterprises that operate efficiently. In addition, it will be possible to support easier decision making through visualization, which is the advantage of decision trees. In the future, support can be extended to not only social enterprises but also other areas to help support decisionmaking.
혼합 데이터 마이닝 기법인 불일치 패턴 모델의 특성 연구 KCI 등재
한국정보기술응용학회 JITAM Vol.15 No.1 2008.03 pp.225-242
※ 기관로그인 시 무료 이용이 가능합니다.
5,200원
PM (Inconsistency Pattern Modeling) is a hybrid supervised learning technique using the inconsistence pattern of input variables in mining data sets. The IPM tries to improve prediction accuracy by combining more than two different supervised learning methods. The previous related studies have shown that the IPM was superior to the single usage of an existing supervised learning methods such as neural networks, decision tree induction, logistic regression and so on, and it was also superior to the existing combined model methods such as Bagging, Boosting, and Stacking. The objectives of this paper is explore the characteristics of the IPM. To understand characteristics of the IPM, three experiments were performed. In these experiments, there are high performance improvements when the prediction inconsistency ratio between two different supervised learning techniques is high and the distance among supervised learning methods on MDS (Multi-Dimensional Scaling) map is long.
한국경영정보학회 한국경영정보학회 정기 학술대회 2007년 추계학술대회 2007.11 pp.591-596
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구에서는 상황 추론의 기능을 추천 시스템에 접목하였다. 연구의 대상 영역은 음악 추천 분야인데, 본 연구에서 제안하는 시스템은 세 개의 모듈, 즉 Intention Module, Mood Module 그리고 Recommendation Module로 구성되어 있다. Intention Module은 사용자가 음악을 청취할 의향이 있는지 없는지를 외부 환경의 상황 데이터를 이용하여 추론한다. Mood Module은 사용자의 상황에 적합한 음악의 장르를 추론한다. 마지막으로 Recommendation Module은 사용자에게 선정된 장르의 음악을 추천한다.
개선된 데이터 마이닝 기술에 의한 웹 기반 지능형 추천시스템 구축 KCI 등재
한국정보기술응용학회 JITAM Vol.12 No.3 2005.09 pp.41-56
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
Product recommender system is one of the most popular techniques for customer relationship management. In addition, collaborative filtering (CF) has been known to be one of the most successful recommendation techniques in product recommender systems. However, CF has some limitations such as sparsity and scalability problems. This study proposes hybrid cluster analysis and case-based reasoning (CBR) to address these problems. CBR may relieve the sparsity problem because it recommends products using customer profile and transaction data, but it may still give rise to scalability problem. Thus, this study uses cluster analysis to reduce search space prior to CBR for scalability problem. For cluster analysis, this study employs hybrid genetic and K-Means algorithms to avoid possibility of convergence in local minima of typical cluster analyses. This study also develops a Web-based prototype system to test the superiority of the proposed model.
데이터 마이닝 기법을 활용한 제조 기업 물류비용 절감 방안에 대한 연구 : 도료 제조 기업 사례를 중심으로 KCI 등재
한국EA학회 정보화연구 제17권 4호 2020.12 pp.297-306
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 도료 제조 기업 B사에서 운용하고 있는 ERP(Enterprise Resources Planning) 시스템 의 경영관리 데이터 및 유통 대리점의 위치정보를 기반으로 데이터 마이닝 및 위치기반 분석 기술을 활 용하여 유통 대리점에 대한 영업 관리 방안을 수립하는데 목적이 있다. 본 연구에서는 ERP 시스템의 제 품 출하일보, 배송차량의 운행정보 및 구글(Google) Maps API(Application Programming Interface)를 이용해 대리점의 GPS 정보를 추출 후 분석에 사용한다. 연구절차는 일일 매출현황, 배차 현황 및 대리 점 위치정보의 전처리, 대리점별 물류효율 산출을 위한 데이터 마이닝, 대리점별 물류효율 등급 산정 및 영업전략 수립, 다중경유지 최적 경로 추천 프로그램 개발이다. 본 연구는 제조 기업이 보유하고 있는 ERP 시스템의 데이터를 활용해 유통 대리점별 물류효율 등급을 산출하는 한편, 등급별 차별적인 관리 를 통해 물류비용 절감 및 영업 생산성 향상 가능성을 제시했다는 점에서 의의가 있다.
This study utilizes data mining and location-based analysis technology based on the business management data of the enterprise resource planning (ERP) system operated by the paint manufacturing company and the location information of the distribution agent to develop a sales management method for the distribution agency. Purpose to establish. In this study, the ERP system's product shipment date, delivery vehicle operation information, and Google Maps API (Application Programming Interface) are used to extract and analyze the agency's GPS information. The research procedure is pre-processing of daily sales status, distribution status and agency location information, data mining for calculating
사용자 특성 별 범죄지도 구현을 위한 데이터 마이닝 기법에 관한 연구
한국경영정보학회 한국경영정보학회 정기 학술대회 빅데이터 시대의 창조 비타민 2014.06 pp.1166-1169
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
공공 DB 데이터마이닝 기법을 활용한 국내 청소년 삶의 만족도 분석에 관한 실증연구 : 의사결정나무 기법을 중심으로 KCI 등재
한국디지털정책학회 디지털융복합연구 제18권 제6호 2020.06 pp.297-309
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
본 연구는 국내 공공 DB에 데이터마이닝 기법인 로지스틱 회귀분석과 의사결정나무 분석을 적용하여 국내 청소 년의 삶의 만족도 증진에 관한 의미 있는 의사결정 규칙을 추출하는 과정을 분석한다. 분석을 위하여 한국아동·청소년패 널조사(KYCPS) 중에서 중1 패널데이터의 4~6차연도 자료인 고등학생 학년별 자료를 활용하였다. 로지스틱 회귀분석 으로 추출된 영향요인은 1학년은 전체 성적 만족도, 주의집중 문제, 우울, 자아 탄력성, 애정, 과잉간섭, 학습활동, 교사 관계, 2학년은 가정의 경제 수준, 건강상태, 전체 성적 만족도, 신뢰, 소외, 학습활동, 학교규칙, 교우관계, 교사 관계, 3학년은 가정의 경제 수준, 전체 성적 만족도, 우울, 자아 탄력성, 애정, 학대, 학교규칙, 교사 관계로 나타났다. 의사결정 나무 기법을 적용한 결과 국내 고등학생의 삶의 만족도는 개인의 정서 문제, 학교성적, 가정의 경제적 환경, 학교적응 등에 의하여 복합적으로 영향을 받는 것으로 파악되었다.
This study focuses on the application of the data mining technique logistic regression analysis and decision tree analysis to the domestic public database called Korean Children Youth Panel Survey (KCYPS) to derive a series of important factors affecting the enhancement of life satisfaction of domestic youth. As a result, the general impact factors on life satisfaction for each grade were derived from logistic regression. Using decision tree analysis, we came to conclusions that those factors such as depression, overall grade satisfaction, household economic level, and school adaptation play crucial roles in affecting high school adolesscents’ life satisfaction.
데이터마이닝 기법을 활용한 불법주차 영향요인 분석 KCI 등재
한국ITS학회 한국ITS학회논문지 제13권 제4호 통권54호 2014.08 pp.63-72
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
우리나라는 급속한 경제발전과 고속성장으로 생활수준이 향상되면서 자동차 수요가 급격히 증가함에 따라 교 통혼잡, 교통사고, 주차문제 등의 문제가 발생되고 있다. 자동차 증가로 인한 주차문제 중 불법주차는 교통혼 잡을 야기하고 주차공간으로 인한 이웃간 분쟁의 원인이 되어 사회적 문제로 대두되고 있다. 이에 본 연구에서는 지방 광역시중 승용차 수단분담률이 높음에도 불구하고 불법주차 단속건수가 상대적을 적은 대전광역시를 대상으로 주차조사를 실시하였으며 불법주차에 대한 원론적인 문제를 파악하기 위해 의사 결정나무모형 Exhaustive CHAID분석을 통하여 운전자들의 주차행위에 있어 불법주차를 선택하는 과정과 그 에 따른 영향요인을 탐색하여 불법주차의 원인을 파악하고 해결하는 방안을 제시하고자 한다. 분석결과 불법주차를 선택하는 영향요인으로는 거리, 단속경험, 직업, 이용시간대 순으로 영향을 미치는 것으 로 나타났으며 예측 모형은 최종적으로 4가지 노드가 도출되었다. 분석결과에 따른 불법주차의 해결방안으로 는 공영주차장의 추가설치와 생계유지 및 조업차량의 주차공간 확보가 우선되어야 하고 불법주차 단속강화와 시민의식 고취를 위한 캠페인의 활성화가 필요하다.
With the rapid development in the economy and other fields as well, the standard of living in South Korea has been improved, and consequently, the demand of automobiles has quickly increased. It leads to various traffic issues such as traffic congestion, traffic accident, and parking problem. In particular, this illegal parking caused by the increase in the number of automobiles has been considered one of the main reasons to bring about traffic congestion as intensifying any dispute between neighbors in relation to a parking space, which has been also coming to the fore as a social issue. Therefore, this study looked into Daejeon Metropolitan City, the city that is understood to have the highest automobile sharing rate in South Korea but with relatively few cases of illegal parking crackdowns. In order to investigate the theoretical problems of the illegal parking, this study conducted a decision- making tree model-based Exhaustive CHAID analysis to figure out not only what makes drivers park illegally when they try to park vehicles but also those factors that would tempt the drivers into the illegal parking. The study, then, comes up with solutions to the problem. According to the analysis, in terms of the influential factors that encourage the drivers to park at some illegal areas, it was learned that these factors, the distance, a driver’s experience of getting caught, the occupation and the use time in order, have an effect on the drivers’ deciding to park illegally. After working on the prediction model, four nodes were finally extracted. Given the analysis result, as a solution to the illegal parking, it is necessary to establish public parking lots additionally and first secure the parking space for the vehicles used for living and working, and to activate the campaign for enhancing illegal parking crackdown and encouraging civic consciousness.
데이터 마이닝을 이용한 입원 암 환자 간호 중증도 예측모델 구축
[NRF 연계] 대한종양간호학회 Asian Oncology Nursing Vol.5 No.1 2005.02 pp.3-11
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구에서 로지스틱 회귀분석, 의사결정나무, 신경망 분석을 이용하여 구축한 예측모델에서 진료과, 나이, 입원월이 공통적으로 환자간호중증도에 영향을 미치는 변수로 나타났다. 박정호, 조현, 박현애, 한혜라(1995) 연구에서는 내외과별, 요일별로 간호업무량에 변화를 준다고 하였는데 본 연구에서 변수로 인식된 진료과와 입원월이 영향을 주는 것으로 나타나 요일과 해당월이라는 차이는 있으나 입원 시점이 환자 간호중증도에 영향을 주는 변수로 나타났다. 박정호, 박현애, 조현, 최용선(1996)에서는 환자간호중증도로 간호업무량을 예측함에 있어 각 병원의 간호활동량 조사 연구에서 언급된 동선, 병동의 간호사 수 및 구성, 간호사의 경력, 간호전달 방법의 차이로 인한 변이가 영향을 준다고 하였으나 이번 연구에서는 그에 대한 비교가 이루어지지 못하였다. 로지스틱 회귀분석과 의사결정나무는 목표변수가 2개 이상이면 예측이 불가능하므로 환자간호중증도를 I군, II군을 묶고 III군, IV군을 묶어 분석을 실시하였다. 신경망 분석 역시 타 모델과 형평성을 맞추기 위하여 I군, II군을 묶고 III군, IV군을 묶어 분석하여 비교 하였다. 세 모델의 평가에서는 환자간호중증도를 예측하는데 신경망분석이 우수한 것으로 나타났다. 모델 평가 연구의 선행연구를 살펴보면, 김옥남 김윤희, 강성홍(2002)연구에서는 critical pathway 실행효과연구에서 로지스틱 회귀분석과 신경망분석, 의사결정기법의 3가지 기법을 비교하였고, 로지스틱 회귀분석이 대조군을 선정하는데 효과적이라고 평가 되었다. 이견직 정영철, 기미라(2003)연구에서는 의사결정나무 기법을 이용해 영향인자를 알아내어 신경망분석에 활용하였다. 정우진, 이선미, 김원훈 (2003)의 연구에서는 로지스틱 회귀분석의 구조와 최종모델에서 추정치를 이용하여 지역 보험에 속하는 각 세대의 지역 보험료 체납 확률 예측모델을 도출하였다. Mueller 등(2004)연구에서는 조산아의 기관튜브발관결과 예측 연구에서 신경망 분석이 로지스틱 회귀분석에 비해서 예측 정확도가 높은 것으로 나타났다. 박현애(1994)는 인공지능분야에서 신경망모델이 최근 분류(classification)에 많이 사용되고 있다. 특히 간호분야에서는 분류과정에 적용된 규칙을 찾아내기 어렵고 사용된 데이터가 불안정한 경우 신경망을 이용한 분류가 규칙이나 직관에 의거한 다른 방법보다 더 효율적인 것으로 알려져 있다고 하였다. 본 연구에서 간호요구에 기초한 환자간호중증도 분류를 위한 신경망 모델을 구축하여 실험한 결과 모델이 내린 중증도 분류와 실제 적용된 경험에 의해 분류한 등급 분류간의 일치율이 84.06%로 나타났다. 본 연구의 제한점으로 첫째, 자료자체가 연구목적으로 수집된 것이 아니라 환자사항과 환자분류등급간에 관계형 데이터 베이스 형태로 되어있지 않아 많은 양의 자료를 정리하여야 했다. 중환자실과 병실에서의 환자간호중증도가 통일되지 않아 중환자실 입원환자를 정리하여야 했고, 실제 환자 사항과 환자분류등급을 연결하지 못하고 자료의 결측치가 많아 7,070명의 환자 중 단지 2,228명의 환자 데이터만을 사용하였다 Goodwin 등(2003)은 데이터 마이닝을 실행하는데에 자료의 질과 결측치, 공통되지 않은 언어가 문제라고 하였다. 본 모델을 사용하여 입원한 암 환자의 간호중증도 분류를 예측하기 위해서는 자료의 질을 높이기 위해 관계형 데이터 ...
Purposes: This study was tried to make a predict model for patient classification according to nursing need. We tried to find the easier and faster method to classify nursing patients that can help efficient management of nursing manpower. Methods: The nursing patient classifications data of the hospitalized cancer patients in one of the biggest cancer center in Korea during 2003.1.1- 2003.12.31 were assessed by trained nurses. This study developed a prediction model and analyzing nursing needs by data mining techniques. Patients were classified by three different data mining techniques, (Logistic regression, Decision tree and Neural network) and the results were assessed. Results: The data set was created using 165,073 records of 2,228 patients classification database. * National Cancer Center Quality ImprovementMain explaining variables were as follows in 3 different data mining techniques. 1) Logistic regression : age, month and section. 2) Decision tree : section, month, age and tumor. 3) Neural network : section, diagnosis, age, sex, metastasis, hospital days and month. Among these three techniques, neural network showed the best prediction power in ROC curve verification. As the result of the patient classification prediction model developed by neural network based on nurse needs, the prediction accuracy was 84.06%. Conclusion: The patient classification prediction model was developed and tested in this study using real patients data. The result can be employed for more accurate calculation of required nursing staff and effective use of labor force.
데이터마이닝 기법을 이용한 기업부실화 예측 모델 개발과 예측 성능 향상에 관한 연구 KCI 등재
한국경영정보학회 경영정보학연구 제18권 제2호 2016.06 pp.173-198
※ 기관로그인 시 무료 이용이 가능합니다.
6,400원
본 연구의 목적은 비즈니스 인텔리전스 연구 관점에서 기업부실화 예측 성능을 향상키시는 것이다. 이를 위해 본 연구는 기존 연구들에서 미흡하게 다루어졌던 1) 데이터셋을 구성하는 과정에서 발생하는 바이어스 문제, 2) 거시경제위험 요소의 미반영 문제, 3) 데이터 불균형 문제, 4) 서술적 바이어스 문제를 다루어 경기순환국면을 반영한 기업부실화 예측 프레임워크를 제안하고, 이를 바탕으로 기업부실화 예측 모델을 개발하였다. 본 연구에서는 경기순환국면별로 각각의 데이터셋을 구성하고, 각 데이터셋에서 의사결정나무, 인공신경망 등 단일 분류기부터 앙상블 기법까지 다양한 데이터마이닝 알고리즘을 적용하여 실험하였다. 또한 본 연구는 데이터불균형 문제를 해결하기 위해, 오버샘플링 기법인 SMOTE(synthetic minority over-sampling technique) 기법을 통해 초기 데이터 불균형 상태에서부터 표본비율을 1:1까지 변화시켜 가며, 기업부실화 예측 모델을 개발하는 실험을 하였고, 예측 모델의 변수 선정 시에 선행연구를 바탕으로 재무비율을 추출하고, 여기서 파생된 IT 산출물인 재무상태변동성과 산업수준상태변동성을 예측 모델에 삽입하였다. 마지막으로, 본 연구는 각 순환국면에서 만들어진 기업부실화 예측 모델의 예측 성능 비교와 경기 확장기와 수축기에서의 기업부실화 예측 모델의 유용성에 대해 논의하였다. 본 연구는 비즈니스 인텔리전스 연구 측면에서 기존 연구에서 미흡하게 다루어졌던 4가지 문제점을 검토하고, 이를 해결할 프레임워크를 제안함으로써 기존 연구 대비 기업부실화 예측률을 10% 이상 향상시켰다는 점에서 연구의 의의를 찾을 수 있다.
Financial distress can damage stakeholders and even lead to significant social costs. Thus, financial distress prediction is an important issue in macroeconomics. However, most existing studies on building a financial distress prediction model have only considered idiosyncratic risk factors without considering systematic risk factors. In this study, we propose a prediction model that considers both the idiosyncratic risk based on a financial ratio and the systematic risk based on a business cycle. Ultimately, we build several IT artifacts associated with financial ratio and add them to the idiosyncratic risk factors as well as address the imbalanced data problem by using an oversampling technique and synthetic minority oversampling technique (SMOTE) to ensure good performance. When considering systematic risk, our study ensures that each data set consists of both financially distressed companies and financially sound companies in each business cycle phase. We conducted several experiments that change the initial imbalanced sample ratio between the two company groups into a 1:1 sample ratio using SMOTE and compared the prediction results from the individual data set. We also predicted data sets from the subsequent business cycle phase as a test set through a built prediction model that used business contraction phase data sets, and then we compared previous prediction performance and subsequent prediction performance. Thus, our findings can provide insights into making rational decisions for stakeholders that are experiencing an economic crisis.
데이터 마이닝 기법을 활용한 항공기 사고 및 준사고로 인한 사망 발생 요인 및 패턴 분석 KCI 등재
한국디지털정책학회 디지털융복합연구 제17권 제9호 2019.09 pp.79-88
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구에서는 데이터 마이닝 기법을 활용하여 항공기 사고와 준사고로 인한 사망 발생 요인들과 패턴들을 분석 하고자 한다. 이를 위해, 항공기 사고와 준사고 데이터를 보유하고 있는 미국연방교통안전위원회(NTSB)와 미국연방항 공청(FAA)의 데이터를 사용하였다. 다음으로 의사결정나무 알고리즘을 사용하여 항공기 사고 및 준사고에 따른 사망여 부 예측모형들을 구축하였고 이를 토대로 사망 발생에 영향을 주는 주요 요인들과 패턴들을 도출하였다. NTSB 데이터 의 경우 항공기가 완파되거나 고기동 또는 고위험 임무를 수행할 때 주로 사망이 발생하는 것을 알 수 있었다. FAA 데이터의 경우 항공기가 일부 파괴된 경우 조종사의 숙련도가 저조하거나 미인가 조종사의 경우 사망이 발생하였으며, 고공낙하점프와 지상운용단계에서 발생되는 다양한 사망관련 패턴들도 발견되었다. 또한 도출된 패턴들을 활용하여 사 망 사고 예방을 위한 실용적인 방안들을 제시한 점에서 연구의 의의를 찾을 수 있다.
This study analyzes the influential factors and patterns associated with death from aircraft accidents and incidents using data mining techniques. To this end, we used two datasets for aircraft accidents and incidents, one from the National Transportation Safety Board (NTSB) and the other from the Federal Aviation Administration (FAA). We developed our prediction models using the decision tree classifier to predict death from aircraft accidents or aircraft incidents and thereby derive the main cause factors and patterns that can cause death based on these prediction models. In the NTSB data, deaths occurred frequently when the aircraft was destroyed or people were performing dangerous missions or maneuver. In the FAA data, deaths were mainly caused by pilots who were less skilled or less qualified when their aircraft were partially destroyed. Several death-related patterns were also found for parachute jumping and aircraft ascending and descending phases. Using the derived patterns, we proposed helpful strategies to prevent death from the aircraft accidents or incidents.
데이터마이닝 분류기법을 이용한 효과적인 연구관리에 관한 연구
한국정보기술응용학회 JITAM Vol.3 No.2 2001.06 pp.1-24
※ 기관로그인 시 무료 이용이 가능합니다.
6,100원
본 연구는 R사의 대고객 만족도 향상을 위하여 고객관계관리(customer relationship management, CRM)를 수행하기 위한 목적으로 추진되었다. 연구의 주안점은 연구관리 데이터베이스로부터 연구관련 변수들의 패턴 및 상호작용을 고려하여 연구계약기관을 기관별 연 구과제의 연구유형 및 연구비에 대한 분석을 통하여 고객유형별로 분류함으로써 향후 대고객관리의 방향을 설정하기 위한 목적으로 시도되었다. 본 연구에서 의사결정나무 알고리듬을 이용하여 자료를 분석한 결과, 17개의 입력변수 중 내외부 계약기관을 분류하는 데 있어서 중요한 변수로는 연구기간, 제경비, 기술개발비의 3개 변수로 나타났다. 연구결과, R사의 고객은 6개월 이상의 연구기간, 3,000만원 이상의 제경비, 그리고 6,075만원의 기술개발비를 기준으로 연구계약기관이 분류되며, 이 연구관련 변수를 이용하여 대고객과의 연구주제 설정, 연구예산 수립 등의 고객관리방안을 수립할 수 있을 것이다.
This purpose of this study is to drive important criteria for improving customer relationship of R institute using data mining techniques. The focus of this research is to consider patterns and interactions of research variables from research management database of R institute, and to classify the outside organizations and the inside organizations for research contract organizations, and to decide the directions of customer relationship management through analyzing the research type and research cost of research topics. In order to drive criteria variables through pattern analysis of the research database, decision tree algorithm is employed. The results show that determinant variables of 17 input variables are research period, overhead cost, R & D cost as variables to classify the outside and inside contract organization.
데이터마이닝과 학습기법을 이용한 부동산가격지수 예측 KCI 등재
한국융합학회 한국융합학회논문지 제12권 제8호 2021.08 pp.47-53
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
4차 산업에 대한 관심이 증폭되면서 데이터를 활용한 과학적 방법론이 발전하고 있지만 부동산 분야에 대한 연구는 데이터 수집의 한계점을 내포하고 있다. 더불어 일반 시장 참여자들의 지식이 확장되면서 정성적인 심리가 부동 산 시장에 미치는 영향이 커지고 있다. 때문에 본 연구에서는 기존의 원천 데이터가 아닌 심리적 부분을 반영한 정량 데이터를 텍스트마이닝과 k-meas 알고리즘을 통해 수집하는 방안을 제안하고 수집된 데이터를 바탕으로 인공신경망 학습을 통해 주택 지수의 방향성을 예측하고자 한다. 2012년부터 2019년까지의 데이터를 학습 기간으로 하고 2020년 도를 예측 기간으로 설정하여 실험을 진행한 결과, 두 가지 CASE에서 예측 능력이 약 80% 이상으로 우수하였고 주택 지수의 상승 구간에서의 예측 강도 또한 우수한 결과를 보였다. 본 연구를 통해서 의사결정에 있어서 부동산 시장 참여 자들에게 인공신경망과 같은 과학적 방식의 활용도 증가 및 고전적 방식에서 벗어난 원천 데이터의 대체 데이터 확보 등에 대한 노력이 증진되기를 기대한다.
With increasing interest in the 4th industrial revolution, data-driven scientific methodologies have developed. However, there are limitations of data collection in the real estate field of research. In addition, as the public becomes more knowledgeable about the real estate market, the qualitative sentiment comes to play a bigger role in the real estate market. Therefore, we propose a method to collect quantitative data that reflects sentiment using text mining and k-means algorithms, rather than the existing source data, and to predict the direction of housing index through artificial neural network learning based on the collected data. Data from 2012 to 2019 is set as the training period and 2020 as the prediction period. It is expected that this study will contribute to the utilization of scientific methods such as artificial neural networks rather than the use of the classical methodology for real estate market participants in their decision making process.
텍스트마이닝 기법을 활용한 교육관점에서의 메타버스 관련 이슈 탐색 - 뉴스 빅데이터를 중심으로 KCI 등재
대한산업경영학회 산업융합연구(구 대한산업경영학회지) 제20권 제6호 2022.06 pp.27-35
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 뉴스 빅데이터에 나타난 메타버스 관련 이슈들을 교육관점에서 분석하여 그 특징을 탐색하고, 메타버 스의 교육적 활용가능성 및 미래교육에 대한 시사점을 제공하는데 목적이 있다. 이를 위해 포털사이트에서 검색되는 메타버스 관련 뉴스 데이터를 41,366건 수집하였고, 대표적인 용어 가중치 모델인 TF-IDF를 이용하여 추출된 모든 키워드의 가중치 값을 계산하여 순위화한 후, 워드클라우드로 시각화 분석을 수행하였다. 또한 정교한 확률기반 텍스트 마이닝 기법인 토픽모델링(LDA)을 활용하여 주요 토픽들을 분석하였다. 연구결과 교육관점에서 메타버스의 핵심 이슈 로는 플랫폼 산업, 미래인재, 기술의 확산 등과 같은 주제가 도출되었다. 또한, 기술, 직업, 교육이라는 세 개의 핵심 주제로 2차 데이터 분석을 실시한 결과 미래교육에서 메타버스는 교육플랫폼의 혁신, 미래 직업의 혁신, 미래 역량의 혁신과 관련한 이슈를 갖는 것으로 나타났다. 본 연구는 방대한 양의 뉴스 빅데이터를 단계적으로 분석하여 교육관점에 서 이슈를 도출하고 미래교육에 대한 시사점을 제공하였다는 데 의의가 있다.
The purpose of this study is to analyze the metaverse-related issues in the news big data from an educational perspective, explore their characteristics, and provide implications for the educational applicability of the metaverse and future education. To this end, 41,366 cases of metaverse-related data searched on portal sites were collected, and weight values of all extracted keywords were calculated and ranked using TF-IDF, a representative term weight model, and then word cloud visualization analysis was performed. In addition, major topics were analyzed using topic modeling(LDA), a sophisticated probability-based text mining technique. As a result of the study, topics such as platform industry, future talent, and extension in technology were derived as core issues of the metaverse from an educational perspective. In addition, as a result of performing secondary data analysis under three key themes of technology, job, and education, it was found that metaverse has issues related to education platform innovation, future job innovation, and future competency innovation in future education. This study is meaningful in that it analyzes a vast amount of news big data in stages to draw issues from an education perspective and provide implications for future education.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.