년 - 년
Feature Engineering Techniques based on CLV for Customer Churn Prediction
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 9th International Conference on Next Generation Computing 2023 2023.12 pp.151-156
Customer Churn Prediction is the process of identifying customers who are likely to stop using a company's products or services in the near future that is critical for the long-term financial stability of a business. Retaining existing customers is often more cost-effective than acquiring new ones, making churn prediction a key focus for customer relationship management (CRM). This study aimed to identify customer churn in data source containing 8,047 sale transactions with various features such as sales, profit, and product category. The four techniques of churn labels generation were introduced base-on features of Time, Value, and Feedback. Additionally, a combination of Time and Value also used to test. The data was split into 80% training and 20% testing subsets, focusing on seven selected features and the Multiple Criteria churn label. Four machine learning models—Random Forest (RF), Logistic Regression (LR), Gradient Boosting (GB), and Support Vector Machines (SVM) were used to create model. The results showed that LR (73.60%) and SVM (71.70%) performed a good performance in terms of accuracy. However, to compare with dataset that included churn label as E-commerce Customer Behavior and Purchase Dataset [23] the results showed that the proposed techniques can be used to impute a churn label attribute which effecting to a classification model.
산업 시계열 데이터를 활용한 Transformer 기반 딥러닝 모델의 Feature Engineering의 효용성 분석
한국경영정보학회 한국경영정보학회 정기 학술대회 Beyond AI: Building an Inclusive and Ethical Digital Economy with Web3 2025.10 pp.86-93
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 12 Number 3 2023.09 pp.89-103
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Obstructive sleep apnea (OSA) is one of the most prevalent sleep disorders that can lead to serious consequences, including hypertension and/or cardiovascular diseases, if not treated promptly. Continuous positive airway pressure (CPAP) is widely recognized as the most effective treatment for OSA, which needs the proper titration of airway pressure to achieve the most effective treatment results. However, the process of CPAP titration can be time-consuming and cumbersome. There is a growing importance in predicting personalized CPAP pressure before CPAP treatment. The primary objective of this study was to optimize the CPAP titration process for obstructive sleep apnea patients through EEG feature engineering with machine learning techniques. We aimed to identify and utilize the most critical EEG features to forecast key OSA predictive indicators, ultimately facilitating more precise and personalized CPAP treatment strategies. Here, we analyzed 126 OSA patients' PSG datasets before and after the CPAP treatment. We extracted 29 EEG features to predict the features that have high importance on the OSA prediction index which are AHI and SpO2 by applying the Shapley Additive exPlanation (SHAP) method. Through extracted EEG features, we confirmed the six EEG features that had high importance in predicting AHI and SpO2 using XGBoost, Support Vector Machine regression, and Random Forest Regression. By utilizing the predictive capabilities of EEG- derived features for AHI and SpO2, we can better understand and evaluate the condition of patients undergoing CPAP treatment. The ability to predict these key indicators accurately provides more immediate insight into the patient’s sleep quality and potential disturbances. This not only ensures the efficiency of the diagnostic process but also provides more tailored and effective treatment approach. Consequently, the integration of EEG analysis into the sleep study protocol has the potential to revolutionize sleep diagnostics, offering a time-saving, and ultimately more effective evaluation for patients with sleep-related disorders.
국제문화기술진흥원 International Journal of Advanced Culture Technology(IJACT) Volume 11 Number 4 2023.12 pp.393-405
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Accurate prediction of wind speed and power is vital for enhancing the efficiency of wind energy systems. Numerous solutions have been implemented to date, demonstrating their potential to improve forecasting. Among these, deep learning is perceived as a revolutionary approach in the field. However, despite their effectiveness, the noise present in the collected data remains a significant challenge. This noise has the potential to diminish the performance of these algorithms, leading to inaccurate predictions. In response to this, this study explores a novel feature engineering approach. This approach involves altering the data input shape in both Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) and Autoregressive models for various forecasting horizons. The results reveal substantial enhancements in model resilience against noise resulting from step increases in data. The approach could achieve an impressive 83% accuracy in predicting unseen data up to the 24th steps. Furthermore, this method consistently provides high accuracy for short, mid, and long-term forecasts, outperforming the performance of individual models. These findings pave the way for further research on noise reduction strategies at different forecasting horizons through shape-wise feature engineering.
Feature engineering with Wavelet transform for Transient detection in KMTNet Supernova Project
[Kisti 연계] 한국천문학회 한국천문학회보 Vol.42 No.2 2017 p.64
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
For the detection of transient sources in optical wide field surveys like KMTNet Supernova Project, difference imaging technique is commonly used. As this method produces a fair amount of false positives, it is also common to utilize machine learning algorithms to screen likely true positives. While deep learning methods such as a convolutional neural network has been successfully applied recently, its application can be limited if the size of the training sample is small. I will discuss a variation of more conventional method that adopts the wavelet transform for feature engineering and its performance.
[NRF 연계] 대한의료정보학회 Healthcare Informatics Research Vol.30 No.3 2024.07 pp.234-243
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Objectives: This study aimed to optimize early coronary heart disease (CHD) prediction using a genetic algorithm (GA)-based convolutional neural network (CNN) feature engineering approach. We sought to overcome the limitations of traditional hyperparameter optimization techniques by leveraging a GA for superior predictive performance in CHD detection. Methods: Utilizing a GA for hyperparameter optimization, we navigated a complex combinatorial space to identify optimal configurations for a CNN model. We also employed information gain for feature selection optimization, transforming the CHD datasets into an image-like input for the CNN architecture. The efficacy of this method was benchmarked against traditional optimization strategies. Results: The advanced GA-based CNN model outperformed traditional methods, achieving a substantial increase in accuracy. The optimized model delivered a promising accuracy range, with a peak of 85% in hyperparameter optimization and 100% accuracy when integrated with machine learning algorithms, namely naive Bayes, support vector machine, decision tree, logistic regression, and random forest, for both binary and multiclass CHD prediction tasks. Conclusions: The integration of a GA into CNN feature engineering is a powerful technique for improving the accuracy of CHD predictions. This approach results in a high degree of predictive reliability and can significantly contribute to the field of AI-driven healthcare, with the possibility of clinical deployment for early CHD detection. Future work will focus on expanding the approach to encompass a wider set of CHD data and potential integration with wearable technology for continuous health monitoring.
Ceramic tile surface defect detection with integrated feature engineering and defect fuse classifier
[NRF 연계] 한양대학교 세라믹연구소 Journal of Ceramic Processing Research Vol.25 No.4 2024.08 pp.572-588
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Introduction: Ceramic tile surface defect detection is crucial for ensuring product quality. This study proposes an integratedapproach combining feature engineering and a Defect Fuse Classifier for accurate defect detection. Methods: The proposedmodel utilizes Python and splits the collected data into 70% for training and 30% for testing. Purpose: The purposesection explicitly states the objectives of the study. It highlights the research goals, such as evaluating the effectiveness ofthe proposed methodology in detecting ceramic tile surface defects and exploring the impact of parameter variations ondetection performance. Results: Comparative analysis with state-of-the-art methods is conducted using various metrics suchas sensitivity, specificity, accuracy, precision, FPR, FNR, NPV, F-Measure, and MCC. (a) For a Training Rate of 70%: Theproposed Defect Fuse Classifier outperforms existing models with an accuracy of 97.4%, precision of 88.5%, sensitivity of88.5%, specificity of 98.5%, F-Measure of 88.5%, MCC of 87%, NPV of 98.5%, FPR of 1.4%, and FNR of 11.4%. Conclusion:This study introduces a novel deep learning approach for ceramic tile surface defect detection, encompassing data acquisition,pre-processing, feature extraction, feature selection, and deep learning-based defect detection. The proposed Defect FuseClassifier, integrating CNN, Bi-LSTM, and RNN, demonstrates superior performance, making it a promising solution fordefect detection in ceramic tile surfaces.
Automated Feature-Based Registration for Reverse Engineering of Human Models
[Kisti 연계] 대한기계학회 Journal of mechanical science and technology Vol.19 No.12 2005 pp.2213-2223
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In order to reconstruct a full 3D human model in reverse engineering (RE), a 3D scanner needs to be placed arbitrarily around the target model to capture all part of the scanned surface. Then, acquired multiple scans must be registered and merged since each scanned data set taken from different position is just given in its own local co-ordinate system. The goal of the registration is to create a single model by aligning all individual scans. It usually consists of two sub-steps: rough and fine registration. The fine registration process can only be performed after an initial position is approximated through the rough registration. Hence an automated rough registration process is crucial to realize a completely automatic RE system. In this paper an automated rough registration method for aligning multiple scans of complex human face is presented. The proposed method automatically aligns the meshes of different scans with the information of features that are extracted from the estimated principal curvatures of triangular meshes of the human face. Then the roughly aligned scanned data sets are further precisely enhanced with a fine registration step with the recently popular Iterative Closest Point (ICP) algorithm. Some typical examples are presented and discussed to validate the proposed system.
피처 엔지니어링에 의한 결과 데이터 대응 관계의 시각화에 대한 연구
[NRF 연계] 한국IT정책경영학회 한국IT정책경영학회 논문지 Vol.12 No.2 2020.04 pp.1663-1670
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 정보통신의 발달과 함께 방대한 양의 데이터들이 생산되었으며 생산된 데이터를 활용, 분석하여 가치 있는 정보를 추출하고, 현상을 예측하는 예측분석의 활용이 중요해지고 있다. 이러한 예측분석을 머신러닝을 이용할 경우 머신러닝 알고리즘에 데이터를 입력할 때 원본 데이터를 그대로 입력하는 것보다 피처 엔지니어링을 수행한 데이터를 입력하는 것이 성능이 더 좋다. 그러나 피처 엔지니어링을 수행한 데이터는 원본 데이터와 차이가 많아 머신러닝 알고리즘의 예측 결과를 설명하는데 있어 데이터의 정보를 이용하는데 한계가 있다. 본 논문에서는 이러한 한계를 극복하고자 원본 데이터와 피처 엔지니어링을 수행한 데이터를 생키 다이어그램을 활용하여 매핑하는 시각화 기능을 제안한다. 본 연구에서는 기존의 시각화와는 다르게 데이터의 변화 과정을 보이는데 강조된 시각화를 개발하여 시각화를 통해 머신러닝 예측 결과에 있어 어느 변수가 중요한 변수인지 파악할 수 있고 어느 피처 엔지니어링이 유효한지 알 수 있다. 개발한 시각화를 보스턴 주택 가격 데이터에서 주택 가격 예측에 효과적인 변수와 피처 엔지니어링이 무엇인지 알아내는 것을 통해 검증하였다.
Recently, with the development of information and communication, a huge amount of data has been produced, and the use of predictive analysis to extract valuable information and predict phenomena by using and analyzing the produced data has become important. If this predictive analysis is used, it is better to enter the data that performed feature engineering than to enter the original data when inputting data into the machine learning algorithm. However, the data that performed feature engineering differs from the original data, limiting the use of data information in describing the predicted results of machine learning algorithms. To overcome these limitations, this paper proposes a visualization function that maps the original data and the data that performed feature engineering to the sankey diagram. Unlike traditional visualizations, this study develops visualizations that emphasize the process of changing data so that visualizations can identify which variables are important in the machine learning forecast results and which feature engineering is effective. The developed visualization was validated by finding out what variables and feature engineering are effective in predicting housing prices in Boston housing price data.
[Kisti 연계] 전력전자학회 전력전자학회 논문지 Vol.27 No.4 2022 pp.332-338
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This study proposes a battery state-of-health estimation method by applying a feature extraction technique. The technique that can improve estimation performance is the process of identifying and extracting meaningful data. To apply a data-driven-based aging state estimation method to batteries, health indicators are used as training data. However, limitations occur in extracting health indicators from charge/discharge cycles. This study proposes a deep-learning-based battery state-of-health estimation method that applies feature extraction techniques to compensate for this problem. According to the performance evaluation result of the proposed method, it has a low estimation error of 0.3887% based on an absolute error evaluation method.
웨이블릿 변환 특징 공학 및 그리드 서치 최적화를 이용한 태양광 발전량 예측 모델의 성능 및 효율성 분석
[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.21 No.3 2026 pp.1053-1060
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
기후 위기 대응을 위한 전 세계적인 에너지 전환 정책에 따라 태양광 발전 설비의 보급이 급증하고 있다. 그러나 기상 조건에 의존적인 태양광 발전의 간헐성과 출력 변동성은 전력 계통의 주파수 안정성을 저해하고 예비력 운영 비용을 증가시키는 주요 원인이 된다. 본 연구에서는 태양광 발전 데이터가 가지는 비정상(Non-stationary) 시계열 특성과 급격한 출력 변동(Ramp event)을 효과적으로 예측하기 위해 이산 웨이블릿 변환(Discrete Wavelet Transform, DWT) 기반의 특징 공학을 제안한다. 또한, LightGBM 모델의 예측 성능을 극대화하기 위해 그리드 서치(Grid Search) 기법을 적용하여 하이퍼파라미터를 체계적으로 최적화하는 통합 프레임워크를 구축하였다. 전국 데이터와 경북 지역 데이터를 대상으로 한 비교 실험 결과, 웨이블릿 특징과 그리드 서치를 결합한 모델이 기존 모델 대비 획기적인 성능 향상을 보였다. 특히 기상학적 동질성이 높은 경북 지역 최적화 모델은 결정계수(R<sup>2</sup>) 0.9900, MAE 9.32, RMSE 19.89를 기록하여 실제값에 근접한 초고정밀 예측을 달성하였다. 아울러 샘플당 0.0246 ms의 초고속 추론 속도를 통해, 실시간 전력 계통 운영(EMS) 및 가상발전소(VPP) 플랫폼에의 즉각적 적용 가능성을 확인하였다.
As solar photovoltaic (PV) installations surge globally to meet carbon neutrality goals, the inherent variability and intermittency of PV power pose significant challenges to power grid stability. This study proposes a robust forecasting framework that integrates Wavelet Transform-based feature engineering to capture non-stationary time-series characteristics and utilizes Grid Search optimization to systematically tune the hyperparameters of the Light Gradient Boosting Machine (LightGBM) model. Experimental results using nationwide and Gyeongbuk regional datasets demonstrate that the proposed model significantly outperforms baseline models. Specifically, the optimized Gyeongbuk regional model achieved an R<sup>2</sup> of 0.9900, MAE of 9.32, and RMSE of 19.89. Through correlation heatmap and violin plot analysis, wavelet-derived features were verified as critical factors explaining output volatility. Furthermore, with an ultra-fast inference speed of 0.0246 ms per sample, the model's applicability to real-time Energy Management Systems (EMS) and Virtual Power Plants (VPP) was validated.
채점 자질 설계를 통한 지도 학습 기반 작문 자동 채점의 타당도 확보 방안 탐색
[NRF 연계] 청람어문교육학회 청람어문교육 Vol.69 2019.03 pp.265-295
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구의 목적은 자동 채점을 위한 단계적 절차의 일부인 자질 설계 단계에 주목하여 자동 채점의 타당성을 확보하기 위한 자질 설계 방안을 탐색하는 것이다. 이에 먼저 작문 자동 채점 시스템의 현황과 자동 채점 모델 설계의 전 과정 속에서 채점 자질의 역할을 살펴보았으며, 다음으로는 채점 자질 설계의 문제로부터 기인한 작문 자동 채점의 타당성 논의로부터 작문 자동 채점에서의 채점 자질 설계를 재개념화하였다. 또한 마지막으로 채점 자질 설계 방안을 1)증거 중심 설계에 기반한 자질 설계, 2)작문 이론을 활용한 자질 설계, 3)전산 언어학의 텍스트 자질을 활용한 자질 설계 등의 세 가지 차원으로 제시하였다. 이러한 논의는 작문 자동 채점과 관련한 기초적인 연구를 촉진하고 관련 담론을 형성하는 데에 기여할 수 있다는 점에서 의의가 있다.
The purpose of the study is to focusing on the feature engineering stage which is a part of the procedure for Automated Writing Scoring(AWS) and searching for the feature engineering method to secure the validity of AWS. For this purpose, the present status of the AWS and the whole process of designing the scoring model is examined to discuss the role of scoring features in the AWS. Next, validity problems of AWS which are derived from the inadequate feature engineering is discussed and the feature engineering in AWS is reconceptualized. Finally, the method of feature engineering is explored and presented in three dimensions, which are 1) feature engineering based on Evidence-Centered-Design(ECD) 2) feature engineering using writing theory 3) feature engineering using computational linguistics’ textual feature.
기고 - 2014년도 라오스건축사기술인협회 컨벤션 참관 기행
[Kisti 연계] 대한건축사협회 건축사 Vol.2014 No.6 2014 pp.84-87
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Software 제품계열공학에서 온톨로지에 기반한 feature의 공통성 및 가변성 분석모델
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2006 pp.139-141
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
제품계열공학에서 feature diagram(FD)은 개발자의 직관이나 도메인 전문가의 경험에 근거하여 작성되어, feature간의 공통성 및 가변성분석 기준이 불명확하며 비정형적인 feature의 공통성 및 가변성 분석으로 인한 stakeholder의 공통된 이해가 부족한 문제점을 내포하고 있다. 따라서, 본 논문에서는 이를 해결하기 위하여 공통된 feature의 이해를 위해 feature 속성리스트에 기반한 메타 feature모델과 feature간의 의미유사성관계를 이용한 온톨로지를 적용한 공통성 및 가변성 분석모델을 제안한다.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.