년 - 년
노후건축물 에너지자립률 향상을 위한 태양광 발전량 예측 알고리즘 개발 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제8권 11호 2024.11 pp.2489-2500
※ 기관로그인 시 무료 이용이 가능합니다.
4,300원
본 연구는 건물의 위치 정보만으로 태양광 발전을 예측할 수 있는 알고리즘을 개발하여 노후건축물을 대 상으로 건물에너지 효율성 향상을 위한 그린리모델링 기대효과 및 건물에너지 자립률 향상을 목표로 한다. 이를 위해, 기존 문헌에서의 태양광 발전 모델과 천공청명도 및 직산분리 방정식 등을 다중 회귀분석식으로 결합한 예 측 알고리즘을 개발하고, 테스트베드를 활용한 실증실험 데이터와의 검증을 통해 예측 알고리즘의 신뢰성을 검증 하였다. 예측 알고리즘의 신뢰성 검증결과, 회귀식의 설명력을 나타내는 R2값의 범위가 하절기(6월-8월) 기준 0.80-0.96으로 높은 예측정확도를 확보하였으며, 7곳 대상지의 건물에너지 자립률은 46%-60%로 산출되어 본 수 식이 그린리모델링 기대효과 예측 시 유용할 것으로 기대된다.
This study aims to develop a prediction algorithm that can predict solar power generation using only the location information of buildings, and to improve the expected effects of green remodeling for improving building energy efficiency and building energy self-sufficiency in aged buildings. As the results of this study, the range of the reliability verification of the prediction algorithm(R2 value) was 0.80-0.96 for the summer season (June-August) and the building energy self-sufficiency rates of the seven target sites were calculated to be 46%-60%. Through these results, this prediction algorithm were verified to be useful for predicting the effects of green remodeling.
Gradient Boosting Classifier with Zebra optimization algorithm for pregnancy risk prediction
[NRF 연계] 한국통신학회 ICT Express Vol.12 No.3 2026.06 pp.693-700
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
High-risk pregnancy endangers both mother and baby, with one maternal death every two minutes in 2023. This study has proposed three Gradient Boosting models?GB-Base, GB-SMOTE, and ZOA-GB?using the West Lombok Pregnancy Risk Prediction Dataset. GB-Base and GB-SMOTE have achieved 90.46% and 90.28% accuracy, while ZOA-GB, using 10 selected features, has reached 88.89%. GB-SMOTE has shown the best performance with an F-score of 84.41%. SHAP has identified Maternal Age, Hemoglobin, and Parity as key features, and DiCE has validated feature-driven prediction control. The study is limited by a single-source dataset, the absence of external-validation, and unexplored optimizers.
Corporate Failure Prediction Modeling Using Genetic Algorithm Technique
한국경영정보학회 한국경영정보학회 정기 학술대회 1998년 추계학술대회 1998.11 pp.599-608
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Data Mining Techniques : More Accurate Classified Algorithm For Cardiopulmonary Diseases Prediction
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 8th International Conference on Next Generation Computing 2022 2022.10 pp.245-248
Data mining techniques develop a more accurate classification algorithm for patients classified as either normotensive, prehypertensive, or hypertensive. Logistic Model Tree, NBTree, and Bagging were chosen as the three classification models with tenfold cross-validation (LMT). Over 24 hours, we collected ABP readings from 1161 patients. To analyze the data, data mining techniques were used and a tool called WEKA. The data was analyzed based on age, gender, wake-up blood pressure, medication, sleep-up blood pressure, and overall blood pressure. According to bagging results, 886 cases (76.3 percent) are correctly classified, with 270 cases classified as pre-hypertensive, 436 cases as Normotensive, and 180 cases as hypertensive. NBTree's results show that 882 (75.9%) of the 1161 instances are correctly classified. Pre-hypertensive patients make up 256, normotensive patients 442, and hypertensive patients 184. Of the 1161 instances, the LMT algorithm correctly classified 878 (75.6 percent). According to the results, 275 people are pre-hypertensive, 431 are normotensive, and 172 are hypertensive. According to our findings, bagging is the most accurate classifier for the 24 hour ABP Monitoring dataset we used. Bagging achieves less overfitting because it focuses on global accuracy. It stabilizes and improves the accuracy of unstable methods compared to single classifiers.
A Study on the Development of Product Planning Prediction Model Using Logistic Regression Algorithm KCI 등재
한국융합학회 한국융합학회논문지 제12권 제9호 2021.09 pp.39-47
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구에서는 계절적인 요인과 급변하는 상품의 트렌드를 사전예측하기 위해 로지스틱 회귀 알고리즘을 이용 한 상품기획 예측 모형을 제안하고자 수행되었다. 먼저 웹크롤링을 이용하여 포털 사이트 및 온라인 마켓의 소비자의 비정형 데이터를 수집하고 정형 데이터 변환을 위한 전처리 작업을 통해 상품에 대한 의미 있는 정보를 분석하였다. 최종 수집된 11,200개의 데이터셋은 Logistic Regression을 이용하여 상품에 대한 소비자의 만족도, 빈도분석, 상품 에 대한 장점과 단점을 분석할 수 있었다. 분석 결과 소비자의 만족도는 92%이었으며, 빈도분석을 통해 상품에 대한 불량이슈를 확인할 수 있었다. 또한, 개발된 상품 기획 예측 프로그램에 대한 사용 만족도, 시스템 효율성, 시스템 효과 성 항목에 대한 분석결과에서도 만족도가 높게 나타났다. 특히, 불량이슈는 상품에 대한 현 문제를 신속히 인지하고 개선 전략을 수립하는데 필요한 정보를 제공한다는 점에서 매우 의미 있는 자료가 된다.
This study was conducted to propose a product planning prediction model using logistic regression algorithm to predict seasonal factors and rapidly changing product trends. First, we collected unstructured data of consumers in portal sites and online markets using web crawling, and analyzed meaningful information about products through preprocessing for transformation of standardized data. The datasets of 11,200 were analyzed by Logistic Regression to analyze consumer satisfaction, frequency analysis, and advantages and disadvantages of products. The result of analysis showed that the satisfaction of consumers was 92% and the defective issues of products were confirmed through frequency analysis. The results of analysis on the use satisfaction, system efficiency, and system effectiveness items of the developed product planning prediction program showed that the satisfaction was high. Defective issues are very meaningful data in that they provide information necessary for quickly recognizing the current problem of products and establishing improvement strategies.
신체활동 비교를 통한 개인 맞춤형 신체활동 에너지 소비량 예측 알고리즘 KCI 등재
대한안전경영과학회 대한안전경영과학회지 제14권 제1호 2012.03 pp.87-93
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
The purpose of this study suggests a personalized algorithm of physical activity energy expenditure prediction through comparison and analysis of individual physical activity. The research for a 3-axial accelerometer sensor has increased the role of physical activity in promoting health and preventing chronic disease has long been established. Estimating algorithm of physical activity energy expenditure was implemented by using a tri-axial accelerometer motion detector of the SVM(Signal Vector Magnitude) of 3-axis(x, y, z). A total of 10 participants(5 males and 5 females aged between 20 and 30 years). The activities protocol consisted of three types on treadmill; participants performed three treadmill activity at three speeds(3, 5, 8 km/h). These activities were repeated four weeks.
사방댐 위치 및 규모 결정을 위한 토석류 토사유출량 예측 알고리즘 개발 KCI 등재
한국재난정보학회 한국재난정보학회논문집 제16권 3호 통권49호 2020.09 pp.586-593
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
연구목적: 이 연구는 토석류로 발생하는 토사유출량 예측 알고리즘을 개발하고, 이를 활용한 GIS 기반 사방댐 적정배치 의사결정 지원 시스템 구현을 목적으로 하였다. 연구방법: 평균 계류 폭과 길이를 이용 한 누적 토사유출량 예측 방법에 초기 붕괴량과 이에 영향하는 집수길이를 입력인자로 활용하여 토석 류로 인해 발생하는 누적 토사유출량 예측 알고리즘을 제시하였다. 연구결과: 알고리즘을 통해 산출된 예측 토사유출량과 실제 토사유출량은 평균 1.1배 차이가 나타나 정확도는 비교적 높았다. 또한 구현된 프로그램은 사방댐의 위치 및 규모를 결정하는 객관적인 지표로서 실무자의 합리적인 의사결정에 도 움을 줄 수 있다. 결론: 사방사업이 매년 시행되고 있는 상황에서 합리적인 사방댐 위치 및 규모 결정을 통해 산지토사재해 방재에 기여할 수 있을 것으로 기대된다.
Purpose: This study aims to develop an algorithm for predicting sediment discharge by debris flow, and develop GIS-based decision support system for optimal arrangement of check dam. Method: The average stream width and flow length were used to predict the cumulative sediment discharge by debris flow. At this time, the amount of slope failure on source area and average flow length were utilized as input factors. Result: The predicted sediment discharge calculated through the algorithm was 1.1 times different on average compared to the actual sediment discharge by debris flow. In addition, the program is an objective indicator that selects the location and size of the check dam, and it can help practitioners make rational decisions. Conclusion: The soil erosion control works are being implemented every year. Therefore, it is expected that the GIS-based decision support system for location and size of the check dam will contribute to the prevention of sediment-related disasters.
웨어러블 다중 파장 광용적맥파 기반 비침습 테스토스테론 예측 알고리즘 KCI 등재
대한산업경영학회 산업융합연구(구 대한산업경영학회지) 제23권 제7호 2025.07 pp.73-80
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
광용적맥파(PPG)는 빛의 흡수 변화를 통해 피부 미세혈관의 혈액량 변화를 비침습적으로 측정하는 방법으로, 웨어러 블 기기에서 심박수 및 산소포화도 측정에 널리 활용되고 있다. 특히 다양한 파장을 동시에 사용한 다중 파장 PPG는 단일 파 장 대비 인체 특성에 따른 간섭을 줄이고 측정 정확도를 높일 수 있음이 보고되었다. 본 논문에서는 이러한 광학 원리와 선행 연구를 바탕으로, 다중 파장 PPG 신호를 입력으로 하는 딥러닝 기반 테스토스테론 예측 알고리즘 구상을 제안한다. 알고리즘 설계는 다중 파장 신호의 동기화・정규화, 독립성분분석(ICA) 등을 통한 잡음 분리, 심박변이도(HRV) 등 추가 특징 추출, 그리 고 LSTM 계열 네트워크를 활용한 회귀 모델 학습 단계로 구성된다. 최종적으로 제안된 방식은 비침습적 연속 모니터링의 가 능성을 제시하나, 임상 데이터 부족과 생리적 메커니즘 미확인 등의 한계를 가지므로, 대규모 실증 연구가 뒤따라야 한다.
Photoplethysmography (PPG) is an optical biosensing technique that noninvasively monitors blood volume changes in the microvasculature by detecting variations in light absorption. Multi-wavelength PPG, which employs several light wavelengths simultaneously, has been shown to improve signal robustness and accuracy by separating contributions from different tissue layers. In this study, based on a survey of existing literature, we propose a design for a testosterone prediction algorithm using wearable multi-wavelength PPG signals. The algorithm concept includes: synchronized acquisition of multi-wavelength PPG, preprocessing such as interference removal using independent component analysis (ICA), feature extraction including heart rate variability and relevant user parameters, and a deep learning regression model to predict circadian testosterone variations. This literature-based framework highlights the potential of non-invasive hormone monitoring in wearable health devices. However, it remains a conceptual proposal: clinical validation and larger-scale data are required. Future work should focus on empirical testing, sensor optimization, and model refinement to establish practical utility.
한국 신용카드시장 기반 한 중국 신용카드 시장 미래 예측 알고리즘 개발 KCI 등재
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.13 No.2 2017.04 pp.62-70
본 연구는 한국신용 카드시장의 사용현황 및 여건을 분석하여, 중국신용카드시장의 추이를 측정할 수 있는 알고리즘개발을 통해 미래예측을 하고자한다. 한국의 신용카드 현황을 분석한 결과 개인소득증가와 신용카드사용과의 상관관계는 유의적이지 않게 나타났으며, 1인당 카드보유 수량과 신용카드사용액수는 상관관계가 있음으로 나타나고 있다. 그러나 중국 신용카드 시장에서는 개인의 소득증가가 신용카드 이용액에 영향을 미치고 있음으로 분석되었고, 1인당 카드보유수량과 신용카드사용액의 상관관계는 유의적으로 분석되었다. 또한 중국신용카드 시장의 미래 경제추이도 분석은 외부환경에서 오는 제약조건이 있는 경우와 제약조건이 없는 경우에 따른 미래예측 알고리즘을 개발하여 1인당 신용카드 소지량과 신용카드 사용액에 대한 미래예측 측정함수를 생성하였다. 본 연구의 활용방안은 중국의 경제시장의 정보미공개로 정보공유가 활발히 이루어지고 있지 않는 중국신용카드시장현황을 본 미래예측 연구를 통해 글로벌 경제시장을 예측할 수 있는 신용카드시장의 발전적 전략을 제시할 수 있을 것으로 기대한다.
In this research, we analyze the usage and conditions of the credit card market in Korea and try to predict future by developing algorithms that can measure the transition of Chinese credit card market. As a result of analyzing the current status of Korean credit cards, the correlation between the increase in personal income and the use of credit card appears is not significant, and there is a correlation between the amount of card holdings per person and the amount of credit cards used. However, in the credit card market in China, it is analyzed that the increase in personal income affects the amount of credit card usage, and the correlation between the amount of card-holding per person and the amount of credit card usage is significantly analyzed. Also, the analysis of the future economic trend of the credit card market in China has developed a future prediction algorithm according to the constraint condition and the constraint condition in the external environment. We generated future forecasting function for credit card holdings and credit card usage per person. This study suggests the development strategies of the credit card market that can predict the global economic market through the future prediction study of the current status of the Chinese credit card market..
위성통신용 적응형 전송기술 리턴링크 채널예측 알고리즘 최적화 KCI 등재후보
한국위성정보통신학회 한국위성정보통신학회논문지 제10권 제2호 2015.06 pp.19-23
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
위성통신서비스의 가용율 및 시스템 throughput 향상을 위해 사용하는 리턴링크 ACM(Adaptive Coding & Modulation)의 원리를 기술하였고, LMS(Least Mean Square) 기반 적응형 필터를 이용한 채널 예측 및 단말의 전송 MODCOD(Modulation & Code rate) 결정 알고리즘의 최적화 과정을 서술하였다. 시뮬레이션 결과 LMS 알고리즘은 필터 계수가 2차이고, μ(step size) 값이 0.00026인 경우 MMSE(Minimum Mean Square Error)가 최소임을 알 수 있다. 이때 MODCOD 결정 알고리즘을 위한 SNR 마진이 0.3dB일 경우 MODCOD 결정 오차를 최소화 할 수 있음을 확인하였다.
In this paper, we present the return link ACM method to improve the link availability and system throughput for satellite communication service. Also, we describe the optimization of an algorithm for channel prediction using the LMS (Least Mean Square) adaptive filter and the MODCOD (Modulation & Code rate) decision. The simulation results show that the optimized filter taps and step-size of adaptive filter are 2 and 0.00026, respectively. And also confirms the required SNR margin for minimization of MODCOD decision error is 0.3dB.
범죄발생 요인 분석 기반 범죄예측 알고리즘 구현 KCI 등재후보
한국위성정보통신학회 한국위성정보통신학회논문지 제10권 제2호 2015.06 pp.40-45
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문에서는 빅 데이터를 이용하여 범죄 발생 요인에 따른 범죄 예측 알고리즘을 구현했다. 제안된 알고리즘은 대검찰청에서 수집 하여 공개한 범죄관련 빅 데이터를 사용하였으며, 통계분석을 통해 서울시의 2011-2013년 범죄발생 패턴을 분석했다. 범죄예측 알고 리즘 구현을 위해 베이지안 네트워크를 적용하였으며, 범죄발생 요인으로서 공간적, 인구적, 사회적 특성 및 요일, 시간, 날씨와 같은 기타 요인으로 베이지안 네트워크의 노드를 구성하였다. 제안한 알고리즘의 구현 결과, 서울시의 각 구별로 범죄발생 패턴이 다르다 는 것을 파악할 수 있었으며, 다양한 범죄발생 패턴을 분석하고, 범죄예측 알고리즘의 정확도를 확인할 수 있었다.
In this paper, we proposed and implemented a crime prediction algorithm based upon crime influential factors. To collect the crime-related big data, we used a data which had been collected and was published in the supreme prosecutors’office. The algorithm analyzed various crime patterns in Seoul from 2011 to 2013 using the spatial statistics analysis. Also, for the crime prediction algorithm, we adopted a Bayesian network. The Bayesian network consist of various spatial, populational and social characteristics. In addition, for the more precise prediction, we also considered date, time, and weather factors. As the result of the proposed algorithm, we could figure out the different crime patterns in Seoul, and confirmed the prediction accuracy of the proposed algorithm.
고속도로에서 차량 안전 통신을 위한 교통사고 위험 예측 알고리즘 KCI 등재
한국디지털정책학회 디지털융복합연구 제10권 제9호 2012.10 pp.319-324
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
차량 안전 통신은 교통사고 예방에 중요한 기술 중 하나이다. 이를 위해 교통사고가 발생하면 연쇄 추돌 사고를 예방하기 위한 안전 메시지를 전달하는 많은 프로토콜이 연구되었다. 이러한 프로토콜은 노드가 안전 메시지를 발생시키는 시점을 교통사고가 발생 했을 경우로 가정한다. 만약, 교통사고 위험을 예측하여 안전 메시지를 전달한다면 운전자는 빠른 대응 조치가 가능하다. 그래서 본 논문에서는 통신 기법을 이용한 교통사고 위험 예측 알고리즘을 제안한다. 결과적으로, 제안된 알고리즘을 기존 프로토콜에 적용한 결과 약 4~5 %의 더 높은 프레임 수신율을 보여주었다.
Vehicle safety communications is one among the important technologies in order to protect a car accident. For this, many protocols forwarding a safe message have studied to protect a chain-reaction collision when a car accident occurs. most of these protocols assume that the time of generating a safe message is the same as an accident’s. If a node predicts some traffic hazard and forwards a safe message, a driver can response some action quickly. So, In this paper, we proposes a traffic hazard prediction algorithm using the communication technique. As a result, we show that the frame reception success rate of using our algorithm to the previous protocol improved about 4~5%.
차선인식 불가 영역에서의 주행 차선 예측 알고리즘 개발
한국정보통신설비학회 한국정보통신설비학회 학술대회 한국정보통신설비학회 2022 정보통신설비 하계학술대회 2022.08 pp.164-165
인프라 정보 융합 기반 자율주행버스의 충돌 예측 알고리즘 설계
한국ITS학회 한국ITS학회 학술대회 SMART CITY 새롭게 펼쳐지는 교통 시스템 2018.04 pp.129-135
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
통행시간 패턴인식형 버스도착시간 예측 알고리즘 개발 연구
한국ITS학회 한국ITS학회 학술대회 환상의 섬 제주도 가즈아~ 5G시대의 교통서비스 변화 2019.04 pp.341-354
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
스마트 ESP의 동작 상태 모니터링과 진단 및 고장 예측 알고리즘에 관한 연구 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제6권 5호 2022.05 pp.775-780
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문에서는 수중 펌프에 대한 동작 상태 모니터링, 고장 진단 및 예측 알고리즘에 대하여 연구하였다. 모니터링 유닛은 수중펌프의 센서 값을 모니터링 제어유닛으로 485통신을 통해서 전송되며, 모니터링 제어 유닛은 센서 값을 변환하여 서버 컴퓨터와 임베디드시스템에 송신한다. 서버 및 임베디드시스템은 데이터베이스에 센서 값을 저장하고 화면에 디스플레이 하여 수중 펌프의 동작 상태를 확인 할 수 있다. 임베디드시스템에서는 데이터 베이스에 저장된 센서 값을 읽어 들여, 개별 및 복합진단을 수행하고, 서버 및 모니터링 제어 유닛에게 송신하여 결과를 표시할 수 있게 하였다. 수중펌프의 예측진단은 23개의 개별 센서의 분류별로 나눠서 다층페셉트론을 구현 하였고, 학습 및 테스트를 통해서 가중치를 설정하였으며, 실시간 입력되는 센서 값을 통하여 알고리즘의 성능을 확인 할 수 있었다. 제안된 알고리즘은 수중 펌프 제품 경쟁력에 도움이 될 것으로 판단된다.
In this paper, operation status monitoring, fault diagnosis, and prediction algorithms for submersible pumps were studied. The monitoring unit transmits the sensor value of the submersible pump to the monitoring control unit through 485 communication, and the monitoring control unit converts the sensor value and transmits it to the server computer and embedded system. The server and embedded system store the sensor value in the database and display it on the screen to check the operation status of the submersible pump. In the embedded system, it is possible to read the sensor values stored in the database, perform individual and complex diagnosis, and display the results by sending them to the server and monitoring control unit. . For predictive diagnosis, the multi-layered peceptron was implemented by dividing by classification of 23 individual sensors, and the performance of the algorithm was confirmed through learning and testing. The proposed algorithm is judged to be helpful in the competitiveness of submersible pump products.
한국ITS학회 한국ITS학회 학술대회 Towards a Connected Future : Innovations in Mobility Technology 연결된 미래를 향하여: 모빌리티 기술의 혁신 2025.04 pp.337-339
※ 기관로그인 시 무료 이용이 가능합니다.
3,000원
미래 예측과 사건 기준 사전 예후 감지 방법의 비교 : ECG 기반 AI 알고리즘 당뇨 예후 감지 KCI 등재
한국경영컨설팅학회 경영컨설팅연구 제26권 제3호 통권 제98호 2026.06 pp.511-521
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
본 연구의 핵심 질문은 미래 시점의 질병 발생 확률을 예측하는 미래 예측(forward prediction)에 비해 사건 기준 예후 감지(event-anchored pre-diagnostic prediction)가 성능 우위를 보이는가이다. 이 질문에 대한 연구자들의 대답은 긍정적이다. 이는 전자가 본질적으로 높은 불확실성을 수반하는 반면 후자는 진단 이전에 이미 존재하는 잠재적 생리 신호를 탐지하는 접근법이기 때문이다. 이와 같이 두 접근법은 동일한 입력 변수를 사용하지만 목표 변수의 정의와 시간 구조가 상이하며, 이러한 차이가 예측 성능에 미치는 영향을 분석하는 것이 본 연구의 핵심 목적이다. 본 연구에서는 환자 10,000명과 비환자 10,000명으로 구성된 총 20,000개의 ECG 기반 데이터를 구축하였으며, 각 데이터는 진단 시점을 기준으로 한 시간 정보를 포함하도록 설계하였다. QTc 간격, 심박수, 신호 통계 특성 등 ECG 기반 변수들을 추출하여 동일한 입력 변수 집합을 구성하고, 이를 기반으로 3년, 5년, 7년, 10년, 15년의 다양한 시간 구간에 대한 미래 예측 모델과 진단 시점을 기준으로 정렬된 예후 감지모델을 구축하였다. 모델 성능은 계층적 5겹 교차검증과 ROC 곡선 아래 면적(AUC)을 통해 평가하였다. 분석 결과, 모든 시간 구간에서 예후감지 모델이 미래 예측 모델보다 일관되게 높은 성능을 보였으며, 이러한 성능 차이는 시간적으로 가까운 구간에서 더욱 크게 나타났다. 또한 모델 유형, 변수 구성, 랜덤 시드 변화에도 불구하고 결과는 강건하게 유지되었다. 추가 분석을 통해 ECG 신호는 진단 수년 이전부터 질병과 관련된 잠재적 정보를 포함하고 있음을 확인하였다. 본 연구는 이러한 성능 차이가 특정 모델에 의존하는 것이 아니라 예측 문제의 구조적 차이에 기인함을 보여준다. 이는 의료 인공지능에서 문제 정의의 중요성을 강조하며, 장기 질병 예측에서 시간 정렬 방식이 예측 성능에 미치는 영향을 이해하는 데 중요한 시사점을 제공한다.
The guiding research question of this study is whether event-anchored pre-diagnostic prediction outperforms forward prediction, which estimates the probability of future disease occurrence. The answer provided by this study is affirmative. This is because forward prediction inherently involves a high level of uncertainty, whereas pre-diagnostic prediction focuses on detecting latent physiological signals that already exist prior to clinical diagnosis. Although both approaches utilize the same set of input features, they differ fundamentally in the definition of the target variable and the temporal structure of the data. Investigating how these differences affect predictive performance constitutes the primary objective of this study. To this end, we constructed a large-scale ECG-based dataset consisting of 20,000 samples, including 10,000 patients and 10,000 non-patients. Each record was designed to include temporal information relative to the diagnosis event. ECG-derived features, such as corrected QT interval (QTc), heart rate, and signal statistical characteristics, were extracted to form a consistent set of input variables. Based on these features, forward prediction models were developed across multiple temporal horizons (3, 5, 7, 10, and 15 years), while pre-diagnostic models were aligned relative to the diagnosis event. Model performance was evaluated using stratified 5-fold cross-validation and the area under the receiver operating characteristic curve (AUC). The results consistently demonstrate that pre-diagnostic models outperform forward prediction models across all temporal horizons, with the performance gap being more pronounced at shorter temporal distances. Furthermore, the findings remain robust across different model types, feature configurations, and random seeds. Additional analyses reveal that ECG signals contain detectable latent disease-related information several years prior to diagnosis. These findings suggest that the observed performance differences are not driven by model choice, but rather by fundamental differences in problem structure. This study highlights the importance of problem formulation in medical artificial intelligence and provides important insights into how temporal alignment influences predictive performance in long-term disease risk modeling.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.