년 - 년
Predicting Session Conversion on E-commerce : A Deep Learning-based Multimodal Fusion Approach KCI 등재 SCOPUS
한국경영정보학회 Asia Pacific Journal of Information Systems 제33권 제3호 2023.09 pp.737-767
※ 기관로그인 시 무료 이용이 가능합니다.
7,200원
With the availability of big customer data and advances in machine learning techniques, the prediction of customer behavior at the session-level has attracted considerable attention from marketing practitioners and scholars. This study aims to predict customer purchase conversion at the session-level by employing customer profile, transaction, and clickstream data. For this purpose, we develop a multimodal deep learning fusion model with dynamic and static features (i.e., DS-fusion). Specifically, we base page views within focal visist and recency, frequency, monetary value, and clumpiness (RFMC) for dynamic and static features, respectively, to comprehensively capture customer characteristics for buying behaviors. Our model with deep learning architectures combines these features for conversion prediction. We validate the proposed model using real-world e-commerce data. The experimental results reveal that our model outperforms unimodal classifiers with each feature and the classical machine learning models with dynamic and static features, including random forest and logistic regression. In this regard, this study sheds light on the promise of the machine learning approach with the complementary method for different modalities in predicting customer behaviors.
클릭스트림 데이터를 활용한 전자상거래에서 상품추천이 고객 행동에 미치는 영향 분석
한국경영정보학회 한국경영정보학회 정기 학술대회 지식경제 시대의 ICT와 경영혁신 한국과학기술회관 2008.06 pp.135-140
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Studies of recommender systems have focused on improving their performance in terms of error rates between the actual and predicted preference values. Also, many studies have been conducted to investigate the relationships between customer information processing and the characteristics of recommender systems via surveys and web-based experiments. However, the actual impact of recommendation on product pages for customer browsing behavior and decision-making in the commercial environment has not, to the best of our knowledge, been investigated with actual clickstream data. The principal objective of this research is to assess the effects of product recommendation on customer behavior in e-Commerce, using actual clickstream data. For this purpose, we utilized an online bookstore’s clickstream data prior to and after the web site renovation of the store. We compared the recommendation effects on customer behavior with the data. From these comparisons, we determined that the relevant recommendations in product pages have positive relationships with the acquisition of customer attention and elaboration. Additionally, the placing of recommended items in shopping cart is positively related to suggesting the relevant recommendations. However, the frequencies at which the recommended items were purchased did not differ prior to and after the renovation of the site.
[NRF 연계] 한국정책학회 한국정책학회보 Vol.20 No.2 2011.06 pp.47-80
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 네티즌들을 대상으로 중앙정부 부처 웹사이트 이용 여부에 영향을 미치는 인구통계학적 변인들과 가족구조 관련 변인들을 규명하는 것을 주목적으로 한다. 이를 위해, 코리안클릭 패널들의 2006년 이용 내역이 담긴 클릭스트림 데이터를 가공하여 부처 기능유형별로 연간 웹사이트 방문 여부를 측정한 횡단면 자료로 변환시킨 뒤 Logit과 Probit 분석을 실시하였다. 분석 결과, 인구통계학적 변인들과 관련해서는 블루칼라, 고졸, 월소득 100-300만원 그룹이 각각 화이트칼라, 대졸, 월소득 500만원이상 그룹에 비해 모든 부처 기능유형에 걸쳐 웹사이트 방문확률이 유의미하게 낮은 것으로 나타났다. 가족구조 변인들과 관련해서는 개인의 중앙정부부처 웹사이트 방문 확률은 인터넷 이용 가족수가 많을수록 감소하는 반면 정부부처 웹 사이트 이용 가족수가 많을수록 증가하는 것으로 드러났다. 아울러 사회문화 부문을 제외하고는 인터넷을 이용하는 청소년 가족수가 많을수록 웹사이트 방문확률이 유의하게 낮아지는 것으로 나타났다.
The main purpose of this study is to examine the sociodemographic characteristics and family-structure-related attributes among the Korean netizens that affects the use of website of ministries of the central government. This study processes the click-stream data of 2006 of the KoreanClick panels and transforms it into cross-sectional data measuring whether or not the individual visited the websites during 2006 as dependent variable. Logit and Probit analysis were used so as to find the main causal factors that affect the individuals' decision whether to visit the government websites or not. The result show that regardless of the functional category of ministries in terms of socio-demographic characteristics, blue-collar workers, high school graduates, and monthly income group of 1-3 million KRW are less likely to use the central government websites than white-collar workers, college graduates and monthly income group of 5 million KRW or more. Also, in terms of family-structure-related variables if more members in the family were internet users, the less likely each family member is to visit the central government websites. On the other hand, if there were many members in the family that visited these websites, it created positive externalities allowing a form of knowledge diffusion, thereby encouraging the use of government ministry websites. Also, we found that the greater the number of teenagers internet users in a family, the less likely each family member visit the central government websites.
인터넷 쇼핑몰 사용자가 구축하는 연결망 구조분석을 통한 구매행동 모델링
[NRF 연계] 대한경영학회 대한경영학회지 Vol.20 No.5 2007.10 pp.2069-2091
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
온라인 상거래 시장에 있어서 눈에 보이지 않는 방문자들의 행동을 파악하고 이에 대한 적절한 반응을 강구하는 것은 중요한 의미가 있다. 본 논문에서는 인터넷 쇼핑몰을 방문하는 사용자가 구축하는 연결망을 분석하여 이를 토대로 구매행동을 파악하고자 하였다. 연결망 분석에 유용하게 쓰이는 그래프 이론을 바탕으로 하여, 웹사이트 방문자가 구축하는 연결망의 밀도와 집중도를 중심으로 사용자의 구매행동을 살펴 보았다. 연결망 이론을 바탕으로 하여 사용자의 전체 방문기록을 의미하는 세션을 하나의 연결망으로 인식하였다. 실제 활동중인 인터넷 쇼핑몰을 방문한 사용자들의 클릭스트림 데이터를 대상으로 분석하였으며, 연구모형의 비교를 검증하기 위해 빈도 및 시간과 관련된 기타 독립변수들로 구성된 모형들과의 분류예측 성과를 로지스틱 회귀분석을 실행하여 비교검증해 보았다. 연구결과 사용자들을 구매집단 및 비구매집단으로 분류예측하는 경우, 본 연구에서 제안하는 연결망 관련 측정치인 밀도와 집중도로 구성된 모형을 이용하면 분류정확도의 성과가 높게 나타나는 것을 알 수 있었다. 즉 연관성있는 개체들을 많이 확인하는 사용자가 구매성향이 높다는 것을 알 수 있었다. 본 연구를 통해 온라인 사용자를 구매 및 비구매 기준으로 구분할 수 있는 하나의 유용한 기법을 제안할 수 있었으며, 인터넷이라는 거대한 연결망 내에서의 고객행동을 설명할 수 있는 방법론을 제시하였다.
As e-commerce based market is getting popular, classifying customers into groups and predicting the decision of whether customers buy products or not in web stores provide useful information for the web service providers. This study proposes an approach that classifies groups of customer using clickstream data and network theory. Specifically, a user’s session from clickstream dataset can be transformed into a graphical network, which represented in density and centralization based on graph theory. The history of online users' activities can be used to construct a relationship network graphically. Dataset is formed on the basis of a session which means a sequence of page views or a period of sustained web browsing. The sessions based on density and centralization of a graph are evaluated into logistic regression. The results show that the proposed model including density and centralization of the session is wellperformed in classifying groups. This study shows customers’ session of dataset with respect to a graph structure can be used to predict whether a customer will buy or not buy a product in a web store. This study proposes an approach to identify customers who show high possibility in purchasing using data set on their behaviors in the internet decision space.
정보화마을 웹사이트 이용자 특성 및 이용행태 분석: 클릭스트림 데이터를 이용하여
[NRF 연계] 한국정책분석평가학회 정책분석평가학회보 Vol.20 No.2 2010.06 pp.283-313
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
정부는 지역 간 정보격차를 해소하고 공동체 및 지역경제를 활성화하기 위해 2001년부터 정보화마을 사업을 전개해 오고 있다. 이에 대한 기존 연구들이 공급자 중심 평가에 초점을 두었다면, 본 연구는 수요자에 중점을 두고 클릭스트림 데이터를 활용하여 정보화마을 웹사이트 이용자의 인구통계학적 특징 및 이용 행태를 분석하였다. 분석 결과, 저연령층과 저학력층의 상대적 방문확률이 낮은 것으로, 연간사용시간과 방문일수에 있어서는 저학력층과 저소득층, 여성, 25~34세 연령층, 충청・강원지역민의 이용도가 높은 것으로 나타났다. 그러나 전체 패널 대비 순방문자 비율이 3.7 %이고, 이용자의 연평균 방문시간이 약 7.5 분, 연평균방문일수가 약 3.28일로 나타나 정보화마을 웹페이지 및 정보화마을 사업이 성공적으로 운영된다고 보기는 어렵다. 향후 이용자 특성 및 이용 행태 분석에 기초한 정보화마을 웹사이트 및 사업 자체의 활성화 전략이 필요하다.
Since 2001 South Korean government has been operationg 'information network village' building project to solve the digital divide between regions, and to stimulate the community and regional economies. In contrast with supplier-focused evaluation of previous research this study focused on the customers to analyse, relying upon Clickstream data, the socio-demographic attributes of 'Information Network Village' website users and their usage behaviour. The analysis found that in terms of access young aged and undereducated group have a lower percentage of log on to 'Information Network Village' website, and as for the duration of visit and the number of daily visits undereducated and low-income, female, 25-34 aged, Chungcheong and Gangwon resident group used the websites more than other groups did during the time period. In conclusion, it is hard to consider 'Information Network Village' website and the 'Information Network Village' building project as a success, because the ratio of unique visitors to total panel members is 3.7 per cent, the user's duration of visit is 7.5 minutes, and the number of his/her daily visits to the site is 3.28 days on average during the time period. It is required to develop a strategy based on the analysis of the attributes of 'Information Network Village' website users and of their behaviour to invigorate 'Information Network Village' website and the 'Information Network Village' building project itself.
클릭스트림 데이터를 이용한 온라인쇼핑몰에서의 고객 유인가격전략의 구매전환효과에 관한 실증연구
[NRF 연계] 한국인터넷전자상거래학회 인터넷전자상거래연구 Vol.7 No.3 2007.09 pp.193-219
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
…This study explores the effectiveness of the loss-leader pricing strategy in online shopping malls. It focuses on the impact of the strategy that leads customers to purchase other goods besides the loss-leader goods (purchasing conversion effects). The study empirically researched the click stream data of six online shopping malls. The results of this research are as follows: First, discount coupons was most effective in terms of purchasing conversion effect than other methods. Second, events produced a different purchasing conversion effect based on the online shopping mall type and customer’s age. Third, bargain sales produced different purchasing conversion effect based on customer’s age and gender. Finally, discount coupons revealed a different purchasing conversion effect based on online shopping mall types. This study may contribute to research on online shopping malls, because it is a pioneering study focusing on the loss-leader pricing strategy at the online business level, using web panels’ clickstream data. It may also help people working in e-commerce to understand the proper application of the loss-leader pricing strategy based on mall’s type, customers' age and gender in order to improve their financial status.
인구통계특성 기반 디지털 마케팅을 위한 클릭스트림 빅데이터 마이닝
[Kisti 연계] 한국지능정보시스템학회 Journal of Intelligence and Information Systems Vol.22 No.3 2016 pp.143-163
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
인구통계학적 정보는 디지털 마케팅의 핵심이라 할 수 있는 인터넷 사용자에 대한 타겟 마케팅 및 개인화된 광고를 위해 고려되는 가장 기초적이고 중요한 정보이다. 하지만 인터넷 사용자의 온라인 활동은 익명으로 행해지는 경우가 많기 때문에 인구통계특성 정보를 수집하는 것은 쉬운 일이 아니다. 정기적인 설문 조사를 통해 사용자들의 인구통계특성 정보를 수집할 수도 있지만 많은 비용이 들며 허위 기재 등과 같은 위험성이 존재한다. 특히, 모바일 환경에서는 대부분의 사용자들이 익명으로 활동하기 때문에 인구통계특성 정보를 수집하는 것은 더욱 더 어려워지고 있다. 반면, 인터넷 사용자의 온라인 활동을 기록한 클릭스트림 데이터는 해당 사용자의 인구통계학적 정보에 활용될 수 있다. 특히, 인터넷 사용자의 온라인 행위 특성 중 하나인 페이지뷰는 인구통계학적 정보 예측에 있어서 중요한 요인이 된다. 본 연구에서는 기존 선행 연구를 토대로 클릭스트림 데이터 분석을 통해 인터넷 사용자의 온라인 행위 특성을 추출하고 이를 해당 사용자의 인구통계학적 정보 예측에 사용한다. 또한, 1)의사결정나무를 이용한 변수 축소, 2)주성분분석을 활용한 차원축소, 3)군집분석을 활용한 변수축소의 방법을 제안하고 실험에 적용함으로써 많은 설명변수를 이용하여 예측 모델 생성 시 발생하는 차원의 저주와 과적합 문제를 해결하고 예측 모델의 정확도를 높이고자 하였다. 실험 결과, 범주의 수가 많은 다분형 종속변수에 대한 예측 모델은 모든 설명변수를 사용하여 예측 모델을 생성했을 때보다 본 연구에서 제안한 방법론들을 적용했을 때 예측 모델에 대한 정확도가 향상됨을 알 수 있었다. 본 연구는 클릭스트림 분석을 통해 추출된 인터넷 사용자의 온라인 행위는 해당 사용자의 인구통계학적 정보 예측에 활용 가능하며, 예측된 익명의 인터넷 사용자들에 대한 인구통계학적 정보를 디지털 마케팅에 활용 할 수 있다는데 의의가 있다. 또한, 제안 방법론들을 통해 어느 종속변수에 대해 어떤 방법론들이 예측 모델의 정확도를 개선하는지 확인하였다. 이는 추후 클릭스트림 분석을 활용하여 인구통계학적 정보를 예측할 때, 본 연구에서 제안한 방법론을 사용하여 보다 높은 정확도를 가지는 예측 모델을 생성 할 수 있다는데 의의가 있다.
The demographics of Internet users are the most basic and important sources for target marketing or personalized advertisements on the digital marketing channels which include email, mobile, and social media. However, it gradually has become difficult to collect the demographics of Internet users because their activities are anonymous in many cases. Although the marketing department is able to get the demographics using online or offline surveys, these approaches are very expensive, long processes, and likely to include false statements. Clickstream data is the recording an Internet user leaves behind while visiting websites. As the user clicks anywhere in the webpage, the activity is logged in semi-structured website log files. Such data allows us to see what pages users visited, how long they stayed there, how often they visited, when they usually visited, which site they prefer, what keywords they used to find the site, whether they purchased any, and so forth. For such a reason, some researchers tried to guess the demographics of Internet users by using their clickstream data. They derived various independent variables likely to be correlated to the demographics. The variables include search keyword, frequency and intensity for time, day and month, variety of websites visited, text information for web pages visited, etc. The demographic attributes to predict are also diverse according to the paper, and cover gender, age, job, location, income, education, marital status, presence of children. A variety of data mining methods, such as LSA, SVM, decision tree, neural network, logistic regression, and k-nearest neighbors, were used for prediction model building. However, this research has not yet identified which data mining method is appropriate to predict each demographic variable. Moreover, it is required to review independent variables studied so far and combine them as needed, and evaluate them for building the best prediction model. The objective of this study is to choose clickstream attributes mostly likely to be correlated to the demographics from the results of previous research, and then to identify which data mining method is fitting to predict each demographic attribute. Among the demographic attributes, this paper focus on predicting gender, age, marital status, residence, and job. And from the results of previous research, 64 clickstream attributes are applied to predict the demographic attributes. The overall process of predictive model building is compose of 4 steps. In the first step, we create user profiles which include 64 clickstream attributes and 5 demographic attributes. The second step performs the dimension reduction of clickstream variables to solve the curse of dimensionality and overfitting problem. We utilize three approaches which are based on decision tree, PCA, and cluster analysis. We build alternative predictive models for each demographic variable in the third step. SVM, neural network, and logistic regression are used for modeling. The last step evaluates the alternative models in view of model accuracy and selects the best model. For the experiments, we used clickstream data which represents 5 demographics and 16,962,705 online activities for 5,000 Internet users. IBM SPSS Modeler 17.0 was used for our prediction process, and the 5-fold cross validation was conducted to enhance the reliability of our experiments. As the experimental results, we can verify that there are a specific data mining method well-suited for each demographic variable. For example, age prediction is best performed when using the decision tree based dimension reduction and neural network whereas the prediction of gender and marital status is the most accurate by applying SVM without dimension reduction. We conclude that the online behaviors of the Internet users, captured from the clickstream data analysis, could be well used to predict their demographics, thereby being utilized to the digital marketing.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.