년 - 년
리뷰 데이터 마이닝을 이용한 하이브리드 추천시스템 개발 : Amazon Kindle Store 데이터 분석사례 KCI 등재
한국경영정보학회 경영정보학연구 제23권 제1호 2021.02 pp.155-172
※ 기관로그인 시 무료 이용이 가능합니다.
5,200원
최근 온라인 상품 구매의 증가로 인해 사용자의 선호에 맞는 상품을 추천해주는 시스템이 지속적으로 연구되고 있다. 추천 시스템은 사용자들에게 개인화된 상품 추천 서비스를 제공하는 시스템으로 사용자가 상품에 남긴 평점을 이용한 협업 필터링(Collaborative Filtering)이 가장 널리 쓰이는 추천 방법이다. 협업 필터링에서 상품 간의 유사도 계산은 시간이 많이 소요되는데, 특히 리뷰 데이터와 같은 빅데이터를 사용할 경우 더욱 많은 시간을 소요한다. 그래서 본 연구에서는 리뷰 데이터 마이닝을 이용하여 상품 간의 유사도 계산을 빠르게 수행할 수 있으면서 정확도를 높일 있도록 2단계(2-Phase) 방법을 이용한 하이브리드 추천시스템 방식을 제안한다. 이를 위해 온라인 전자책 상거래 상점인 아마존 킨들 스토어(Amazon Kindle Store)의 약 98만 개의 온라인 소비자 평점과 리뷰 데이터를 수집 하였다. 실험 결과 본 연구에서 제안한 사용자의 평점과 리뷰를 단계적으로 반영한 하이브리드 추천 방식이 전통적인 추천 방식과 비교하여 추천 시간은 비슷하였으나 높은 정확도를 나타내는 것을 확인하였다. 따라서 제안한 방법을 사용하면 사용자가 선호하는 상품을 빠르고 정확하게 추천함으로써 고객의 만족을 높여서 기업의 매출 증대에 기여할수 있을 것으로 기대된다.
With the recent increase in online product purchases, a recommender system that recommends products considering users' preferences has still been studied. The recommender system provides personalized product recommendation services to users. Collaborative Filtering (CF) using user ratings on products is one of the most widely used recommendation algorithms. During CF, the item-based method identifies the user's product by using ratings left on the product purchased by the user and obtains the similarity between the purchased product and the unpurchased product. CF takes a lot of time to calculate the similarity between products. In particular, it takes more time when using text-based big data such as review data of Amazon store. This paper suggests a hybrid recommendation system using a 2-phase methodology and text data mining to calculate the similarity between products easily and quickly. To this end, we collected about 980,000 online consumer ratings and review data from the online commerce store, Amazon Kinder Store. As a result of several experiments, it was confirmed that the suggested hybrid recommendation system reflecting the user's rating and review data has resulted in similar recommendation time, but higher accuracy compared to the CF-based benchmark recommender systems. Therefore, the suggested system is expected to increase the user's satisfaction and increase its sales.
텍스트 마이닝을 이용한 온라인 리뷰 데이터의 레스토랑 선택속성과 선호도에 관한 연구
관광경영학회 관광경영연구 제25권 제4호 통권 104호 2021.07 pp.41-62
※ 기관로그인 시 무료 이용이 가능합니다.
5,800원
The purpose of this study was to extract and analyze the extensive review data of tourists accumulated on the Tripadvisor travel community website to identify the restaurant selection attributes and preference factors of foreign tourists, and to derive meaningful results. In order to identify restaurant selection attributes and important preference factors in online reviews generated by restaurant customers, data was refined through text mining techniques and text network analysis was performed to reveal the structural with restaurant selection attributes. For the texts extracted by crawling, a frequency matrix was created by the word frequency list and key words using the TEXTOM program. Also, using Netdraw programs, visualized the results of the sementic network analysis and centreality of the extracted words and the structural equivalence of the words. As a result of the analysis, food, restaurant, good, place, korean, seoul, try, service and great words were found to be the main attributes in frequency and centrality. Additionally the attributes were categorized atmosphere, value, purpose and food by CONCOR analysis. Based on these findings, I would like to present theoretical and practical implications for market segmentation and marketing strategies to the restaurant industry.
정보 성분과 상대위험도를 이용한 clopidogrel의 약물상호작용 시그널 검색 : 건강보험데이터베이스를 대상으로 한 데이터마이닝 연구 KCI 등재
한국임상약학회 한국임상약학회지 제21권 제2호 2011.06 pp.90-99
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Health Insurance Review & Assessment Service (HIRA) claims database has a high potential to detect signals of new drug interactions. The aim of this study was to evaluate the usefulness of information component (IC) and relative risk (RR) as a tool for signal detection, and to analyze the possible drug interactions caused by clopidogrel using HIRA claims database. This study was performed in elderly patients over 65 years of age who administered clopidogrel from January 2005 to June 2006 in South Korea. Serious Adverse Events (SAEs) as drug interactions of clopidogrel were defined as any ambulatory hospitalization for ischemic diseases within comcomitant medication period of clopidogrel. Information Component (IC) and Relative Risk (RR) were calculated to compare the proportion of drug-SAE pairs in order to select drug specific SAEs. IC and RR signals of clopidogrel drug interaction were screened when IC’ 95% confidence interval was greater than 0 and RR’ 95% confidence interval was greater than 1 respectively. All detected signals were compared to references such as Micromedex® and 2010 Drug Interaction Facts™ Sensitivity, specificity, positive predicted value and negative predicted value were used to evaluate usefulness of this method. Among 13,252,930 cases of elderly patients who co-administered clopidogrel and other drugs, 47,485 cases were detected as SAE. Of these, one-hundred nine cases were detected by the IC-based data-mining approach and ninety one cases were detected by the RR-based data-mining approach. Total One-hundred sixty three unrecognized signals were detected by IC or RR. Twelve signals from IC-based data-mining (57.1%) were corresponded with drug interactions from references and eight signals from RR-based data-mining (38.1%) were corresponded with drug interactions from references. These signals include proton pump inhibitors, calcium channel blockers and HMG CoA reductase Inhibitors, which were known to affect CYP450 metabolism. Further studies using HIRA claims database are necessary to develop appropriate data-mining measure.
Heart Disease Prediction System Using Data Mining and Hybrid Intelligent Techniques : A Review SCOPUS
보안공학연구지원센터(IJBSBT) International Journal of Bio-Science and Bio-Technology Vol.8 No.4 2016.08 pp.139-148
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Heart disease is one of the main sources of demise around the world and it is imperative to predict the disease at a premature phase. The computer aided systems help the doctor as a tool for predicting and diagnosing heart disease. The objective of this review is to widespread about Heart related cardiovascular disease and to brief about existing decision support systems for the prediction and diagnosis of heart disease supported by data mining and hybrid intelligent techniques.
Systematic Review of Data Mining Applications in Patient-Centered Mobile-Based Information Systems
[NRF 연계] 대한의료정보학회 Healthcare Informatics Research Vol.23 No.4 2017.10 pp.262-270
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Objectives: Smartphones represent a promising technology for patient-centered healthcare. It is claimed that data mining techniques have improved mobile apps to address patients’ needs at subgroup and individual levels. This study reviewed the current literature regarding data mining applications in patient-centered mobile-based information systems. Methods: We systematically searched PubMed, Scopus, and Web of Science for original studies reported from 2014 to 2016. After screening 226 records at the title/abstract level, the full texts of 92 relevant papers were retrieved and checked against inclusion criteria. Finally, 30 papers were included in this study and reviewed. Results: Data mining techniques have been reported in development of mobile health apps for three main purposes: data analysis for follow-up and monitoring, early diagnosis and detection for screening purpose, classification/prediction of outcomes, and risk calculation (n = 27); data collection (n = 3); and provision of recommendations (n = 2). The most accurate and frequently applied data mining method was support vector machine; however, decision tree has shown superior performance to enhance mobile apps applied for patients’ self-management. Conclusions: Embedded data-mining-based feature in mobile apps, such as case detection, prediction/classification, risk estimation, or collection of patient data, particularly during self-management, would save, apply, and analyze patient data during and after care. More intelligent methods, such as artificial neural networks, fuzzy logic, and genetic algorithms, and even the hybrid methods may result in more patients-centered recommendations, providing education, guidance, alerts, and awareness of personalized output.
A Review of the Text and Data Mining (TDM) Rules on Trainable Datasets for AI in China
[NRF 연계] 한중사회과학학회 한중사회과학연구 Vol.22 No.3 2024.07 pp.192-221
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
There are contradictory cases about copyright infringement of TDM on AI training datasets in China. This inconsistency stems from the fact that the Regulation for the Implementation of the Copyright Law of the People’s Republic of China w as p romulgate d in 1991 and revised for the third time in 2013.However, no further revisions have been made in the intervening decade. Given the rapid changes in AI technology since 2013, this administrative regulation fails to keep pace with current realities. Article 21 of this administrative regulation “shall not affect the normal use of the work, nor shall it reasonably harm the legitimate interests of the copyright owner” is one of the bases for judicial decisions, so it reduces the predictability of justice and is not conducive to the development of the training dataset industry. The key focus of revising this administrative regulation is to propose countermeasures for constructing TDM rules for Chinese training datasets. This paper recommends the following amendments to the administrative regulation. Firstly, it is imperative to acknowledge that copyright owners and training datasets providers are two groups with conflicting interests, and they need to compromise. Secondly, the criterion for determining copyright infringement in the context of TDM for training datasets should be explicitly designated as the “three steps”, and the Chinese court no longer adopts the “four factors” standard. Thirdly, specify the responsible subjects involved in TDM activities, with a particular emphasis on the training dataset providers. This revised proposal is conducive to unifying the reasoning approaches of Chinese courts and promoting the development of the data industry and the AI sector. It is hoped that it will serve as a reference for China to revise this administrative regulation and formulate policies in the future.
데이터마이닝 기법을 활용한 온라인 리뷰의 해양리조트 매력속성 연구: 유추(anology)에 의한 홉스테드(Hofstede)의 문화차원이론 관점에서
[NRF 연계] 한양대학교 관광연구소 관광연구논총 Vol.37 No.3 2025.08 pp.121-151
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
이 연구는 온라인 리뷰에서 도출된 해양리조트 매력속성이 무엇인지 그리고 홉스테드의 문화차원이론의 관점에서 문화별 차이점이 무엇인지 분석하였다(연구문제1). 또한빅데이터 분석의 계량적 접근을 탐색적으로 시도하기 위하여 R 프로그램의 타이디버스 접근법을 활용하여 해양리조트 매력속성과 만족간의 관계를 분석하였다(연구문제2). 트립어드바이저의 세 리조트의 리뷰를 수집하기 위하여 웹하비를 활용하였고 총 2,981개의 리뷰를추출하여 텍스트마이닝을 수행하였다. 연구문제1의 분석 결과, 롯데호텔 제주는 다른 두 리조트들보다 부대시설과 접근성에서는 뛰어나지만, 환경 속성, 편안한 리조트분위기, F&B에서 취약한 것을 알 수 있다. 호텔 닛코 오키나와는 다른 두 리조트들보다 객실서비스, 편안한 리조트분위기 연출, F&B에서 강점을 보였다. 마지막으로 두짓타니 라구나 푸켓은 직원서비스, 환경속성, F&B에서 강점을 보였고, 고객의 긍정적인 감정 표출에서 다른 두 리조트보다 2배 넘는 표현이 도출되어 고객의 높은 만족도를 간접적으로 확인하였다. 연구문제2의분석 결과, 롯데호텔 제주는 직원서비스, 긍정적 감정이 만족에 유의미한 영향을 미쳤고, 호텔 닛코 오키나와는 직원서비스, 접근성이 만족에 유의미한 영향을 미쳤다. 두짓타니 리구나푸켓은 직원서비스, 긍정적 감정이 만족에 유의미한 영향을 미쳤다. 이 연구의 결과를 토대로 학문적 시사점과 실무적 시사점을 제시하였다.
This study examined the attractive attributes of marine resorts based on online reviews, with two research objectives: (1) to identify differences in resort attributes across countries using Hofstede’s cultural dimension theory, and (2) to apply a quantitative big data approach to analyze the relationship between attractive attributes and customer satisfaction. Online reviews were collected from TripAdvisor using WebHarvy, yielding a total of 2,981 reviews for three resorts. Text mining was conducted, and data were analyzed using the Tidyverse package in R. For Research Question 1, results showed that Lotte Hotel Jeju outperformed the other two resorts in auxiliary facilities and accessibility but was weaker in environmental attributes, a comfortable resort atmosphere, and F&B. Hotel Nikko Okinawa demonstrated strengths in room service, creating a comfortable resort atmosphere, and F&B. Dusit Thani Laguna Phuket excelled in employee service, environmental attributes, and F&B, and exhibited more than twice the frequency of positive emotional expressions compared to the other resorts, indirectly indicating higher customer satisfaction. For Research Question 2, Lotte Hotel Jeju showed significant effects of employee service and positive emotions on satisfaction. Hotel Nikko Okinawa demonstrated significant effects of employee service and accessibility on satisfaction. Dusit Thani Laguna Phuket showed significant effects of employee service and positive emotions on satisfaction. These findings provide both academic and practical implications, offering guidance for enhancing service strategies in marine resorts across different cultural contexts.
스마트워치 SNS 리뷰 데이터와 오피니언 마이닝을 통한 감성 분석 처리에 대한 연구
[Kisti 연계] 한국정보통신학회 한국정보통신학회 학술대회논문집 2015 pp.1047-1050
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
IoT(Internet of Things)에 대한 관심과 함께 웨어러블 디바이스 또한 차세대 융합 기술의 핵심으로 그 관심이 증가하고 있다. 특히, 초기 단계인 스마트워치 시장의 선점을 위하여 여러 기업들이 경쟁하고 있으며, 사용자들은 이러한 경쟁 속에서 각 기기에 대한 의견을 SNS를 통하여 공유하며 그에 대한 선호도를 표출하고 있다. 따라서 본 논문에서는 스마트워치에 관련된 속성과 감성단어들에 대한 감성사전을 먼저 구축한 뒤 이를 토대로 의견 데이터 모델을 통하여 수집된 SNS의 데이터를 속성별로 분류한다. 이후 수집된 데이터를 자연언어 처리 기법을 이용하여 전반적 극성 및 속성별 극성을 판단하고 이를 통하여 각 스마트워치 리뷰에 대한 분석을 수행하고자 한다. 그리고 수집된 자료 분석을 통하여 사용자들이 선호하는 스마트워치의 속성을 파악할 수 있도록 하고 이를 통해 각 기기별 발전방향을 판단하는데 기여하도록 한다.
Wearable device, along with IoT(Internet of Things), is considered the core of upcoming generation's convergence technology. Companies are intensely competing one another for prior occupation in the smartwatch market. Consumers that use smartwatch express their preferences by sharing their opinions through SNS(Social Networking Service). Through this study, emotions dictionary is built, which consists of attributes and emotional words related to smartwatch. Based on the emotions dictionary, SNS data has been categorized according to the attributes through opinion data model. Afterwards, overall polarity and attribute polarity of collected data are distinguished through natural language parsing, followed by an analysis of smartwatch reviews. This study will contribute to determination of which attributes of smartwatch to be improved, to arise consumer's interest for individual smartwatch.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.