Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 65
No
1

A Study on the Drug Classification Using Machine Learning Techniques KCI 등재후보

Anmol Kumar Singh, Ayush Kumar, Adya Singh, Akashika Anshum, Pradeep Kumar Mallick

중소기업융합학회 산업과 과학 제3권 제2호 2024.06 pp.8-16

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문에서는 인구통계학적, 생리학적 특성을 기반으로 환자에게 가장 적합한 약물을 예측하는 것을 목표로 하는 약물 분류 시스템을 제시한다. 데이터 세트에는 적절한 약물을 결정하기 위한 목적으로 연령, 성별, 혈압(BP), 콜레스테롤 수치, 나트륨 대 칼륨 비율(Na_to_K)과 같은 속성들이 포함된다. 본 연구에 사용된 모델은 KNN(K-Nearest Neighbors), 로지스틱 회귀 분석 및 Random Forest이다. 하이퍼파라미터를 최적화하기 위해 5겹 교차 검증을 갖춘 GridSearchCV를 활용하였으며, 각 모델은 데이터 세트에서 훈련 및 테스트 되었다. 초매개변수 조정 유무에 관계없이 각 모델의 성능은 정확도, 혼동 행렬, 분류 보고서와 같은 지표를 사용하여 평가되었다. GridSearchCV를 적용하지 않은 모델의 정확도는 0.7, 0.875, 0.975인 반면, GridSearchCV를 적용한 모델의 정확도는 0.75, 1.0, 0.975로 나타났다. GridSearchCV는 로지스틱 회귀 분석을 세 가지 모델 중 약물 분류에 가장 효과적인 모델로 식별했으며, K-Nearest Neighbors가 그 뒤를 이었고 Na_to_K 비율은 결과를 예측하는 데 중요한 특징인 것으로 밝혀졌다.

This paper shows the system of drug classification, the goal of this is to foretell the apt drug for the patients based on their demographic and physiological traits. The dataset consists of various attributes like Age, Sex, BP (Blood Pressure), Cholesterol Level, and Na_to_K (Sodium to Potassium ratio), with the objective to determine the kind of drug being given. The models used in this paper are K-Nearest Neighbors (KNN), Logistic Regression and Random Forest. Further to fine-tune hyper parameters using 5-fold cross-validation, GridSearchCV was used and each model was trained and tested on the dataset. To assess the performance of each model both with and without hyper parameter tuning evaluation metrics like accuracy, confusion matrices, and classification reports were used and the accuracy of the models without GridSearchCV was 0.7, 0.875, 0.975 and with GridSearchCV was 0.75, 1.0, 0.975. According to GridSearchCV Logistic Regression is the most suitable model for drug classification among the three-model used followed by the K-Nearest Neighbors. Also, Na_to_K is an essential feature in predicting the outcome.

2

DNN 기법을 활용한 지하공동 데이터기반의 지반침하 위험 지도 작성 KCI 등재

김한응, 김창헌, 김태건, 박정준

한국재난정보학회 한국재난정보학회논문집 제19권 2호 통권60호 2023.06 pp.334-343

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

연구목적: 본 연구에서는 지반공동탐사로 발견된 공동자료를 지하시설물과의 원인별 상관관계로 분석 하고, AI 알고리즘 기반으로 지반침하 예측지도를 검증하여 시민에게 안전한 도로 환경을 제공하고자 한다. 연구방법: 위험도 평가 관련 데이터조사와 빅데이터 수집, AI분석을 위한 데이터 전처리, 그리고 AI 알고리즘을 이용하여 지반침하 위험도 예측지도 검증 등 3가지 단계로 연구를 수행하였다. 연구결과: 작성한 지반침하 위험 예측지도를 분석하여 부산시 부산진구와 사하구에 대해 긴급, 우선, 일반 3단계의 공동관리 위험등급 분포를 확인 할 수 있었다. 또한, 지반침하 위험 등급 예측 값을 도로노선의 구간별로 정리하여 긴급 등급이 포함된 도로가 부산진구는 총 61개구간 중 3개소, 사하구는 총 68개구간 중 7개소 임을 확인하였으며 각 도로노선별 지반침하 위험 예측 순위를 파악하였다. 결론: 도출된 지반침하 위험 예측지도를 바탕으로 효율적으로 탐사구간을 설정하여 우선 조사, 선제 조치함으로써 시민들의 불안을 해소하고 효율적인 도로유지관리 및 보수, 제도의 개선 등의 부수적인 효과를 얻을 수 있다.

Purpose: In this study, the cavity data found through ground cavity exploration was combined with underground facilities to derive a correlation, and the ground subsidence prediction map was verified based on the AI algorithm. Method: The study was conducted in three stages. The stage of data investigation and big data collection related to risk assessment. Data pre-processing steps for AI analysis. And it is the step of verifying the ground subsidence risk prediction map using the AI algorithm. Result: By analyzing the ground subsidence risk prediction map prepared, it was possible to confirm the distribution of risk grades in three stages of emergency, priority, and general for Busanjin-gu and Saha-gu. In addition, by arranging the predicted ground subsidence risk ratings for each section of the road route, it was confirmed that 3 out of 61 sections in Busanjin-gu and 7 out of 68 sections in Sahagu included roads with emergency ratings. Conclusion: Based on the verified ground subsidence risk prediction map, it is possible to provide citizens with a safe road environment by setting the exploration section according to the risk level and conducting investigation.

3

빅 데이터를 활용한 애완동물 상품 추천 시스템 구현 KCI 등재

김삼택

한국융합학회 한국융합학회논문지 제11권 제11호 2020.11 pp.19-24

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

최근, 애완동물의 급격한 증가로 애완동물의 건강상태 체크와 다양하게 수집된 데이터를 활용하여 사료 추천 등 통합적인 애완동물관련 개인화 상품 추천 서비스가 요구된다. 본 논문은 빅 데이터 기술을 활용하여 애완동물관련 데이터 수집, 전처리, 분석, 관리등 다양한 개인화서비스를 할 수 있는 상품 추천시스템을 구현한다. 먼저, 애완동물이 착용하고 있는 센서 정보와 고객의 구매 패턴, SNS 정보를 수집해 데이터베이스에 저장하고 통계적 분석을 활용하여 사료제작, 애완동물 건강관리 등 맞춤형 개인화 추천 서비스가 가능한 플랫폼을 구현한다. 본 플랫폼은 유사도가 분석 될 상품과 상품정보에 대한 유사도 상품 정보를 출력하고 최종적으로 추천 분석한 결과를 출력하여 고객에게 정보를 제공 할 수 있다.

Recently, due to the rapid increase of pets, there is a need for an integrated pet-related personalized product recommendation service such as feed recommendation using a health status check of pets and various collected data. This paper implements a product recommendation system that can perform various personalized services such as collection, pre-processing, analysis, and management of pet-related data using big data. First, the sensor information worn by pets, customer purchase patterns, and SNS information are collected and stored in a database, and a platform capable of customized personalized recommendation services such as feed production and pet health management is implemented using statistical analysis. The platform can provide information to customers by outputting similarity product information about the product to be analyzed and information, and finally outputting the result of recommendation analysis.

4

Data Preprocessing for Machine-Learning-Based Adaptive Data Center Transmission

Kamran Keykhosravi, Ahad Hamednia, Houman Rastegarfar, Erik Agrell

[NRF 연계] 한국통신학회 ICT Express Vol.8 No.1 2022.03 pp.37-43

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

To enable optical interconnect fluidity in next-generation data centers, we propose adaptive transmission based on machine learning in a wavelength-routing network. We consider programmable transmitters that can apply possible code rates to connections based on predicted bit error rate (BER) values. To classify the BER, we employ a preprocessing algorithm to feed the traffic data to a neural network classifier. We demonstrate the significance of our proposed preprocessing algorithm and the classifier performance for different values of and switch port count.

5

4,500원

비대면 서비스의 확산과 원격근무 등의 생활이 일상화되면서 메타버스의 이용자가 급증하였으며, 세컨드 라이프를 위한 금 융, 부동산 등 자산 실거래로써 사람들의 생활반경이 자연스럽게 확대되고 있다. 하지만 실세계에서와 같이 경제 활동을 자유 롭게 하기에는 전혀 다른 쟁점의 보안 이슈들이 발생할 수 있으며, 기존의 사이버 공간과 비교해도 완전히 새로운 유형의 문 제들이 발생할 가능성이 크다. 본 논문에서는 이러한 문제를 해결하기 위해 인증 및 프라이버시를 중점으로 방안을 제안한다. 메타버스에서 수집되는 데이터의 특수성에 기반하여 데이터 전처리 기능을 개선하고, W3C의 DID의 표준 방식을 준수하는 가 운데 NFT를 활용한 메타버스 내 새로운 인증 서비스를 정의하며, 이에 대한 시스템 설계 및 개발 방안을 제시한다. 제안한 시스템은 하이퍼레저 인디 블록체인을 활용하여 구현하고, 결과를 분석하여 본 연구의 적용 가능성 및 우수성을 검증한다.

As remote services and remote work become commonplace, the use of the Metaverse has grown. This allows transactions like real estate and finance in virtual Second Life. However, conducting economic activities in the Metaverse presents unique security challenges compared to the physical world and conventional cyberspace. To address these, the paper proposes solutions centered on authentication and privacy. It suggests improving data preprocessing based on Metaverse data's uniqueness and introduces a new authentication service using NFTs while adhering to W3C's DID framework. The system is implemented using Hyperledger Indy blockchain, and its success is confirmed through implementation analysis.

6

4,000원

7

최근 몇 년간 딥 오토 인코더 또는 UNet과 같은 딥러닝 모델이 이상 상황 탐지를 위한 재구성 기반 모델로 사용되었다. 이러한 재구성 기반 모델은 재구성 손실을 기반으로 비정상 지수를 측정한다. 그러나 재구성 기반 모델의 탐지 성능은 입력 데이터의 형식에 따라 크게 좌우되므로 데이터 전처리 과정은 재구성 기반 이상 탐지 모델의 탐지 성능을 향상하는 데 매우 중요한 요소이다. 따라서 본 연구에서는 이상 탐지를 위한 새로운 데이터 처리 방법을 제안한다. 우리가 제안한 데이터 전처리 과정에서는 우선 현재 프레임과 이전 프레임 간의 차이를 계산하여 새로운 이미지를 생성한다. 그런 다음 새로 구축된 이미지를 UNet 모델에 대한 입력 이미지로 사용한다. 그런 다음 원래 이미지와 UNet 모델을 통해 재구성된 이미지 간의 차이를 계산하여 이상 영역 위치를 탐지한다. 실험에서 우리가 제안한 데이터 전처리 방법과 기존의 다른 데이터 전처리 기법의 탐지 성능을 비교하였다. 실험 결과 우리가 제안한 데이터 처리 방법을 사용한 UNet 기반 이상 탐지 모델이 다른 데이터 전처리 방법을 사용한 모델보다 이상 영역을 더 잘 탐지할 수 있음을 보여준다.

8

4,500원

데이터 기반 의사결정 방법론, 고도화된 빅데이터 처리 기법의 발달로 데이터를 처리하는 방법에 대한 정보의 수요가 늘어나고 있다. 데이터를 활용하는 거의 모든 작업과 연구에서 데이터 전처리 과정이 포함 되나, 이러한 과정은 주장하고자 하는 내용이나 결과물을 도출하기 위한 수단으로써 언급될 뿐 실질적인 과정에 대해서 자세하게 설명하고 있는 연구는 부족하였다. 실질적인 분석 기법을 활용하기 이전의 단계 로 간단하게 언급되는 경우가 많아 데이터 처리에 대한 인사이트를 획득하기 어려운 경우가 많았다. 따라 서 해당 연구에서는, raw data에서부터 데이터를 처리하는 과정, 즉 데이터 처리 파이프라인에 대해서 자세하게 작성하고자 하였다. 특히 수입식품 수입 절차에 대한 설명을 구체화하여 해당 상황에서 데이터 의 필드들이 어떻게 해석될 수 있고 어떠한 필드들을 왜 활용하게 되었는지에 대한 상황과 해당 도메인 지식을 공유하면서 흐름을 기술하고자 하였다.

With the development of data-driven decision-making and sophisticated big data processing technique, there is a growing demand for information on how to process data. However, recent studies with data preprocessing mentioned only as a means to achieve a result. Therefore, in this study, we aimed to write in detail about the data processing pipeline, include prerocessing data. In particular, we shares the context and domain knowledge to aid fluent understand of the research.

9

자율주행 데이터 전처리 기법 최적화 방안

이호운, 오철

한국ITS학회 한국ITS학회 학술대회 AI-powered Innovations in ITS 2025.10 pp.519-524

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

10

양자 컴퓨터 기반 탄성파 자료 전처리 기법의 가능성 및 한계

이주완, 조용채

[NRF 연계] 한국자원공학회 한국자원공학회지 Vol.62 No.2 2025.04 pp.169-178

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

양자 컴퓨터는 정보를 처리하기 위해 양자역학의 원리를 활용하는 컴퓨터 방식으로, 특정 알고리듬에 대해서는 기존 CPU 기반 컴퓨터에 비해 효율적인 연산비용을 가진다. 가장 범용적인 양자 컴퓨터인 양자 게이트 컴퓨터는 트랜지스터로 만들어진 게이트에 대응되는 양자 게이트를 연산 도구로 사용한다. 반면 양자 어닐러는 입자가 가진 파동 형태의 특성상 입자의 에너지보다 높은 장벽을 통과할 수 있다는 이론인 양자 터널링을 이용한 양자 컴퓨터로, 수학적 최적화 수행에 특화되어있다. 본 연구는 탄성파 탐사 자료의 처리 기법 및 지하구조 영상화 알고리듬 과정을 양자 컴퓨터로수행할 수 있는 이론적 가능성을 제시한다. 또한 해당 수행방법의 한계점을 함께 제시한다. 결과적으로, 본 연구는 양자 처리 장치가 추후 수치모델링, 탄성파 탐사 자료 전처리 및 자료 품질 향상 등의 기법을 시간효율적으로 수행할 수 있다는 가능성을 시사한다.

Quantum computers utilize the principles of quantum mechanics to process information and have efficient computational costs compared with central-processing-unit-based computers for specific algorithms. The most general-purpose quantum computer, the quantum gate computer, uses quantum gates corresponding to gates made of transistors as computational tools. Quantum annealers, however, utilize quantum tunneling, based on the theory that particles can pass through barriers higher than the particle energy owing to the particle waveform, and it is specialized in performing mathematical optimizations. This study investigated the theoretical possibility of applying seismic exploration data processing techniques and subsurface structure imaging algorithms using quantum computers. However, this study also reveals the limitations of these methods. The results suggest that quantum computers have the potential to perform numerical modeling, seismic data preprocessing, and data enhancement in a time-efficient manner.

11

4,000원

효율적인 데이터베이스 마케팅을 위하여 고객들을 세분화하고, 새로운 지식을 탐색할 수 있는 데이터마이닝 의 필요성이 증대되고 있다. 데이터마이닝 도구를 구축하기 위해서는 단계별 구현이 요구되어 지는데, 본 연구에서는 데이터마이닝 을 위한 분산 환경에 적응 가능한 데이터 전처리 도구를 구성하였다. 기존의 데이터마이닝 도구인 앤 서 트리, 클레멘타인, 엔터프라이즈 마이너, 캔싱턴, 웨카의 전처리 부분을 고찰하고, 분산 환경에서 효율적으로 사용 할 수 있는 데이터 마이닝 전처리 도구를 구성하였다. 새로이 제안된 시스템은 엔터프라이즈 자바 빈즈와 XML을 기반으로 하였다.

This paper is to construction of the data mining preprocessing tool for efficient database marketing. We compare and evaluate the often used data mining tools based on the access method to local and remote databases, and on the exchange of information resources between different computers. The evaluated preprocessing of data mining tools are Answer Tree, Climentine, Enterprise Miner, Kensington, and Weka. We propose a design principle for an efficient system for data preprocessing for data mining on the distributed networks. This system is based on Java technology including EJB(Enterprise Java Beans) and XML(eXtensible Markup Language).

12

차량검지기 이력자료 활용을 위한 전처리과정 개선

이환필, 남궁성, 김수희, 김진

한국ITS학회 한국ITS학회 학술대회 2012년 한국ITS학회 춘계학술대회 2012.04 pp.353-359

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

13

고속도로 차량검지기 이력자료 활용을 위한 전처리과정 개선 KCI 등재

이환필, 남궁성, 김수희, 김진

한국ITS학회 한국ITS학회논문지 제12권 제1호 통권45호 2013.02 pp.15-27

※ 기관로그인 시 무료 이용이 가능합니다.

4,500원

그간 차량검지기로부터 수집되는 다양한 정보는 주로 실시간 자료로 이용되었으나 최근 교통데이터 이력자료의 활용방안에 대한 중요성이 증대되고 있다. 이러한 배경에서 본 연구는 차량검지기자료의 이력자료 활용을 위한 전처리 개선에 대한 연구를 수행하였다. 실제 교통현상과 가장 가까운 데이터 처리를 목적으로 세부처리로직을 개선하였다. 평가결과 기존 전처리 과정보다 개선 전처리 과정이 실제값에 가까운 결과를 나타내는 것으로 분석되었다.

While the vehicle detector is collected from a variety of information was mainly used as a real-time data. Recently scheme of application for archived traffic data has become increasingly important. In this background, this research were conducted on the improvement of the preprocessing for archived traffic data application. The purpose of improving specific preprocessing was reflect transportation phenomena by traffic data. As evaluation result, improvement preprocessing was close to the actual value than exist preprocessing

14

4,000원

15

통행시간 추정을 위한 TCS 데이터의 전처리 모형 개발 KCI 등재

이현석, 남궁성

한국ITS학회 한국ITS학회논문지 제8권 제5호 통권25호 2009.10 pp.1-11

※ 기관로그인 시 무료 이용이 가능합니다.

4,200원

TCS (Toll Collection System) 데이터는 원시 데이터 자체로서도 구간의 교통상황을 어느 정도 반영할 수 있는 교통특성을 내포하고 있다. 그러나 TCS 데이터에는 이상치가 포함되어 있어 이러한 데이터는 해당 구간의 통행시간을 대표한다고 볼 수 없으므로 만약 이러한 이상치들이 포함되어 있음에도 불구하고 제거하지 않고 집락을 한다면 이상치들로 인해 통행시간은 크게 왜곡 될 가능성이 있다. 특히 장거리 구간일수록 통행시간의 분산이 증가하여 동일구간 동일시간대라도 다양한 통행시간이 분포하고 있다. 구간이 길어질수록 통행시간의 변동이 심하여 적절한 통행시간 대푯값을 구하기가 어렵다. 따라서 TCS 자료를 이용하여 통행시간의 대푯값을 산정하기 위해서는 통행시간의 변동 특성을 파악하는 것이 중요하다. 본 연구에서는 TCS 데이터의 전처리 기법을 개선하되 구간의 길이와 교통상황에 따른 통행시간의 변동을 고려하여 TCS 원시데이터로부터 시・공간적 통행패턴을 파악할 수 있는 의미 있는 통행시간을 추출하고자 한다.

TCS Data imply characteristics of traffic conditions. However, there are outliers in TCS data, which can not represent the travel time of the pertinent section, if these outliers are not eliminated, travel time may be distorted owing to these outliers. Various travel time can be distributed under the same section and time because the variation of the travel time is increase as the section distance is increase, which make difficult to calculate the representative of travel time. Accordingly, it is important to grasp travel time characteristics in order to compute the representative of travel time using TCS Data. In this study, after analyzing the variation ratio of the travel time according to the link distance and the level of congestion, the outlier elimination model and the smoothing model for TCS data were proposed. The results show that the proposed model can be utilized for estimating a reliable travel time for a long-distance path in which there are a variation of travel times from the same departure time, the intervals are large and the change in the representative travel time is irregular for a short period.

16

기계학습 기반 근감소증 예측을 위한 데이터 전처리 기법 KCI 등재

최윤, 윤유림

국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.9 No.3 2023.05 pp.737-744

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

근감소증은 노인들 사이에서 점점 더 흔하게 발생하고 있어, 최근 주목을 받고 있는 질병이다. 근감소증의 원 인은 매우 다양하게 나타나지만, 노화, 식습관, 운동 부족등이 주요한 원인들 중 하나이다. 근감소증은 원인이 다양한 만큼 예방 및 치료에 전략을 개발하는 것이 중요하다. 하지만 요인이 다양한 만큼 사람이 근감소증을 정확하게 예측 하기는 어렵다. 여기서 기계학습을 이용해 근감소증 예측의 정확도와 편의를 크게 높일 수 있다. 그러나 생활습관과 생체 데이터의 양은 방대한 만큼, 전처리 없이 데이터를 쓰기에는 시간복잡도와 정확성 측면에서 부적절할 수 있다. 본 논문에서는 근감소증과 그 원인에 대한 최신 문헌을 검토하고, 그에 맞게 기계학습 기만 근감소증 예측에 활용할 데이터를 전처리하는데 초점을 맞춘다.

Sarcopenia is an increasingly common disease among the elder that has recently received attention. Although the causes of sarcopenia are diverse, aging, dietary habits, lack of exercise are the one of the major factors. As the causes of sarcopenia are diverse, it is important to develop strategies for prevention and treatment. However, predicting sarcopnia accuartely is difficult due to the variety of factors involved. Here, machine learning can significantly improve the accuracy and convenience of predicting sarcopenia. However, since lifestyle habits and biological data are vast, using data without preprocessing may be inappropriate in terms of time complexity and accuracy. This paper reviews recent literature on sarcopnia and its causes, focusing on preprocessing the data to be used in sarcopnia prediction machine learning accrodingly.

17

다양한 노인 생활 지표를 활용한 기계학습 기반 노인 건강 요인 예측 KCI 등재

아잠, 이재형, 윤유림

국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.10 No.6 2024.11 pp.677-689

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

노인들의 삶의 질, 악력, 경제활동 등 다양한 지표들은 그들의 종합적인 복지와 건강 상태를 반영한다. 이러한 정보를 활용한 종합적인 평가는 노인의 건강 상태를 예측하는 데 유용하다. 본 연구에서는 지역사회 거주 노인의 건 강을 예측하는 종합적인 지표에 대해 기계학습 기반 예측 모델을 적용하고 비교하는 것을 목표로 한다. 고령화연구패 널에서 제공하는 4652명의 데이터를 활용하여 예측 변수에 맞게 다양한 머신러닝 기법을 사용하여 각 모델을 평가 하였다. 그 결과, 악력 예측에는 LightGBM Regression 모형이 RMSE 5.082, MSE 25.83로 가장 우수한 성능을 보였으며, 현재 건강 상태 예측에서는 Gradient Boosting이 RMSE 0.588과 R-Square 0.456로 가장 좋은 성과를 보였다. 한편 고령층의 경제활동 참여에 대한 예측 결과는 Random Forest 모델이 우수함을 드러냈다. 이러한 기계 학습 기반 예측 모델은 노인의 건강 상태 평가와 경제활동 참여 예측에 대한 방향성을 제시하며, 종합적인 예측을 위 해 다양한 방법론을 수행하여야 함을 시사한다.

The quality of life, frailty, economic activity, and other indicators are crucial for assessing older adults' overall well-being and health status. A comprehensive evaluation using this information helps predict the health status of older adults. This study aims to apply and compare machine learning-based prediction models for comprehensive health indicators of community-dwelling older adults. Utilizing data from 4,652 individuals provided by the Aging Research Panel, we assessed various machine learning techniques to fit the predictor variables. Our findings reveal that the LightGBM Regression model performed the best, with an RMSE of 5.082 and an MSE of 25.83. The Gradient Boosting model best predicted current health status, with an RMSE of 0.588 and an R-Square of 0.456. Additionally, the Random Forest model showed strong performance in predicting economic activity participation among older adults. These machine learning-based models offer valuable insights for evaluating health status and predicting economic activity participation, highlighting the importance of employing diverse methodologies for comprehensive predictions.

18

안정적 데이터 수집을 위한 지능형 IIoT 플랫폼 개발 KCI 등재

조우진, 이형아, 김동주, 구재회

국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.10 No.4 2024.07 pp.687-692

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

전 세계적으로 에너지 위기가 심각한 문제로 대두되고 있다. 대한민국의 경우 전체 에너지의 53% 이상 사용하 며, 온실 가스 배출량 또한 대한민국 전체의 45% 이상을 차지하고 있는 산업 단지 관련 에너지 효율화 연구에 높은 관심을 가지고 있다. 그 연구 중 하나로 가상 에너지 네트워크 플랜트라는 산업 단지 내 동일 유틸리를 사용하는 공 장 들 간의 공유 설비와 에너지 생산 공장과 수요 공장 간의 거래로 에너지를 절감하는 연구를 제시한다. 이러한 에 너지 절감 연구에서는 분석, 예측 등 데이터의 활용처가 다양하기 때문에 데이터의 수집이 굉장히 중요하다. 하지만, 시계열 데이터를 안정적으로 수집하는데는 기존의 시스템들은 여러 부족함이 있었다. 본 연구에서는 그를 개선하기 위해 지능형 IIoT 플랫폼을 제안한다. 지능형 IIoT 플랫폼은 비정상 데이터를 식별하고 적시에 처리하기 위한 전처리 시스템을 포함하며, 이상과 결측 데이터를 분류하고 안정적인 시계열 데이터를 유지하기 위한 보간 기법을 제시한다. 또한 데이터베이스 최적화를 통해 시계열 데이터 수집을 효율화한다. 본 논문은 안정적 데이터 수집과 신속한 문제 대응을 통해 산업 환경에서의 데이터 활용성을 높이는데 기여하며, 다양한 챗봇 알림 시스템을 도입하여 데이터 수집 부담을 줄이고 모니터링 부하를 최적화하는데 기여한다.

The energy crisis is emerging as a serious problem around the world. In the case of Korea, there is great interest in energy efficiency research related to industrial complexes, which use more than 53% of total energy and account for more than 45% of greenhouse gas emissions in Korea. One of the studies is a study on saving energy through sharing facilities between factories using the same utility in an industrial complex called a virtual energy network plant and through transactions between energy producing and demand factories. In such energy-saving research, data collection is very important because there are various uses for data, such as analysis and prediction. However, existing systems had several shortcomings in reliably collecting time series data. In this study, we propose an intelligent IIoT platform to improve it. The intelligent IIoT platform includes a preprocessing system to identify abnormal data and process it in a timely manner, classifies abnormal and missing data, and presents interpolation techniques to maintain stable time series data. Additionally, time series data collection is streamlined through database optimization. This paper contributes to increasing data usability in the industrial environment through stable data collection and rapid problem response, and contributes to reducing the burden of data collection and optimizing monitoring load by introducing a variety of chatbot notification systems.

19

A Survey on Pre-processing and Post-processing Techniques in Data Mining

Divya Tomar, Sonali Agarwal

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.7 No.4 2014.08 pp.99-128

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Knowledge Discovery in Databases (KDD) covers various processes of exploring useful information from voluminous data. These data may contain several inconsistencies, missing records or irrelevant features, which make the knowledge extraction, a difficult process. So, it is essential to apply pre-processing techniques to these data in order to enhance its quality. Detailed description of data cleaning, imbalanced data handling and dimensionality reduction pre-processing techniques are depicted in this paper. Another important aspect of Knowledge Discovery is to filter, integrate, visualize and evaluate the extracted knowledge. In this paper, several visualization techniques such as scatter plots, parallel co-ordinates and pixel oriented technique are explained. The paper also includes detail descriptions of three visualization tools which are DBMiner, Spotfire and WinViz along with their comparative evaluation on the basis of certain criteria. It also highlights the research opportunities and challenges of Knowledge Discovery process.

20

A Theoretical Approach towards Data Preprocessing Techniques in Data Mining – A Survey

Lilly Florence M, Swamydoss D

보안공학연구지원센터(IJUNESST) International Journal of u- and e- Service, Science and Technology Vol.9 No.11 2016.11 pp.421-428

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Nowadays, a huge amount of data is stored in system from web, mobile, social media, and health centre or from any production company. Unfortunately these data cannot be used as it is. For knowledge discovery, we must use the quality input. To convert this data into quality input, we have many algorithms and methods in data mining. The preprocessing techniques are used for convert the noise data into quality data. So preprocessing steps are the most important steps in data mining. This paper presents all data preprocessing methods. Here the authors describes the four main process namely data cleaning, data integration, data reduction and data transformation.

 
1 2 3 4
페이지 저장