Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 19
No
1

데이터 전 처리를 통한 의료패널 데이터 분석 : 건강검진의 개인특성을 중심으로 KCI 등재

최영진, 박현춘, 윤영배

한국EA학회 정보화연구 제16권 2호 2019.06 pp.179-187

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 연구는 개인 특성에 따른 건강검진과 민간보험 이용 행태를 파악하기 위한 것이다. 한국보 건사회연구원과 건강보험관리공단에서 제공하는 한국의료패널 데이터를 전 처리 과정을 거쳐 변환한 후, 개인의 사회적, 경제적, 건강요인에 따른 이용 행태를 STATA 10.0으로 분석하였다. 연구결과, 만 성질환자와 경제적으로 여유있는 계층이 건강검진을 많이 이용하고, 민간보험 가입에는 가구소득 및 경제활동 등 경제적 요인이 주요 영향요인으로 분석되었다. 그리고 건강을 우려하는 그룹이 건강검진 을 적극적으로 이용할 것이라는 예측과 달리 건강하다고 인식하는 그룹에서 건강검진을 많이 이용하 는 것으로 파악되었다.

Purpose: This study is to examine the personal characteristics, such as economic, health, social prospects and helath-scan or private insurance usage. Methods: The data used in this study were collected from Korea Health Panel, and testified using regression in STATA 10.0. Results: According to the results of the study, chronic illness and wealth group more use healthscan service than the opposite, and economic factors are related to taking of private health insurance. Also, it show that the healthy group more use the health-scan service than unhealthy group.

2

사회재난의 위험성이 증가함에 따라 리스크 평가에 대한 중요성이 증가하면서 이에 대한 연구가 활발하게 수행되고 있다. 리스크 평가를 바탕으로 위험지도를 제작하여 이를 바탕으로 정책 결정과 시민들에게 공표하는 등 재난안전 측면에 있어 위험 지도는 매우 중요한 연구이다. 그러나 위험지도를 제작하는 과정에서 데이터의 속성마다 적절한 표준화가 이루어지지 않을 경 우 최종 위험지도의 결과를 신뢰하기 어렵다. 사회재난의 경우 지역 특성에 따라서 일부 지역에만 발생하거나 계절적 요인으로 인해 특정 기간에 집중되어있다. 따라서 특정 재난은 이상치 속성을 보이는 데이터 분포를 가지고 있어 위험지도를 작성하기 전 전처리 과정이 필요하다. 본 논문에서는 이상치 속성을 갖는 재난을 대상으로 새로운 표준화 기법을 제안하였으며 이상치 탐색을 위해 통계적 기법을 적용하였다. 또한 이상치 탐색 후 제안한 방법을 통해 표준화를 위한 범위설정을 재검토하였다. 이 상치를 고려하기 전과 후의 결과를 비교하였으며 이를 바탕으로 데이터 전처리의 필요성을 제시하고 가축질병을 중심으로 적 용한 결과를 나타내었다.

3

개인정보 보호를 고려한 딥러닝 데이터 자동 생성 방안 연구 KCI 등재

장성봉

국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.10 No.1 2024.01 pp.435-441

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

수집된 대량의 데이터셋이 딥러닝 학습데이터로 사용되기 위해서는 주민번호, 질병 정보등과 같이 민감 한 개인정보는 해커에게 노출되지 않도록 값을 변경하거나 암호화해야 하고 구축된 딥러닝 모델의 구조와 일 치 하도록 데이터를 재구성 해주어야 한다. 현재, 이러한 작업은 전문가에 의해 수동으로 이루어지기 때문에, 시간과 비용이 많이 소요 된다. 이러한 문제점을 해결하기 위해, 본 논문에서는 딥러닝 과정에서 개인정보 보 호를 위한 데이터 처리 작업을 자동으로 수행할 수 있는 기법을 제안한다. 제안된 기법에서는 데이터 일반화에 기반한 개인정보 보호 작업을 수행하고 원형큐를 사용하여 데이터 재구성 작업을 수행한다. 제안된 기법의 타 당성을 검증하기 위해, C언어를 사용하여 직접 구현하였다. 검증 결과, 데이터 일반화가 정상적으로 수행되고 딥러닝 모델에 맞는 데이터 재구성이 제대로 수행됨을 확인 할 수 있었다.

In order for the large amount of collected data sets to be used as deep learning training data, sensitive personal information such as resident registration number and disease information must be changed or encrypted to prevent it from being exposed to hackers, and the data must be reconstructed to match the structure of the built deep learning model. Currently, these tasks are performed manually by experts, which takes a lot of time and money. To solve these problems, this paper proposes a technique that can automatically perform data processing tasks to protect personal information during the deep learning process. In the proposed technique, privacy protection tasks are performed based on data generalization and data reconstruction tasks are performed using circular queues. To verify the validity of the proposed technique, it was directly implemented using C language. As a result of the verification, it was confirmed that data generalization was performed normally and data reconstruction suitable for the deep learning model was performed properly.

4

HEMS: Automated Online System for SEGAK Analysis and Reporting SCOPUS

Fadzli Syed Abdullah, Nor Saidah Abd Manan, Aryati Ahmad, Sharifah Wajihah Wafa, Mohd Razif Shahril, Nurzaime Zulaily, Rahmah Mohd Amin, Amran Ahmed

보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.10 No.10 2016.10 pp.89-104

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The Ministry of Education Malaysia (MOE) implemented the National Physical Fitness Standard (SEGAK) for Malaysian School Children Assessment Program. Ever since, the SEGAK assessment data had been collected by the respective teachers in every school twice a year, then its summary is being submitted to the State Department of Education manually through email. This creates problems such as lack of a standardized report format, complex formula in calculating SEGAK score and different data interpretation. In this paper, an integrated and automated SEGAK submission and analysis system is proposed. The system, which is known as Health Monitoring System (HEMS), is a web based system developed with an automated pre-processing method and implemented three tier architecture. HEMS have a centralized database that collects the assessment data from seven districts in Terengganu. A total of 35,681 data was collected from 213 primary schools, and 27,201 data from 44 secondary schools, giving a big total of 67,519 data. During the pre-processing, 4,637 data or 6.9% of the collected data were excluded due to wrong and incomplete information. Using HEMS template, the submitted data have a consistent format of data types. HEMS generates an automated analysis and reporting for the use of related authorities.

5

A Survey on Pre-processing and Post-processing Techniques in Data Mining

Divya Tomar, Sonali Agarwal

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.7 No.4 2014.08 pp.99-128

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Knowledge Discovery in Databases (KDD) covers various processes of exploring useful information from voluminous data. These data may contain several inconsistencies, missing records or irrelevant features, which make the knowledge extraction, a difficult process. So, it is essential to apply pre-processing techniques to these data in order to enhance its quality. Detailed description of data cleaning, imbalanced data handling and dimensionality reduction pre-processing techniques are depicted in this paper. Another important aspect of Knowledge Discovery is to filter, integrate, visualize and evaluate the extracted knowledge. In this paper, several visualization techniques such as scatter plots, parallel co-ordinates and pixel oriented technique are explained. The paper also includes detail descriptions of three visualization tools which are DBMiner, Spotfire and WinViz along with their comparative evaluation on the basis of certain criteria. It also highlights the research opportunities and challenges of Knowledge Discovery process.

6

Development of Terra MODIS data pre-processing system on WWW

Takeuchi, W., Nemoto, T., Baruah, P.J., Ochi, S., Yasuoka, Y.

[Kisti 연계] 대한원격탐사학회 대한원격탐사학회 학술대회논문집 2002 pp.569-572

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Terra MODIS is one of the few space-borne sensors currently capable of acquiring radiometric data over the range of view angles. Institute of Industrial Science, University of Tokyo, has been receiving Terra MODIS data at Tokyo since May 2001 and Asian Institute of Technology at Bangkok since May 2001. They can cover whole East Asia and is expected to monitor environmental changes regularly such as deforestation, forest fires, floods and typhoon. Over eight hundred scenes have been archived in the storage system and they occupy 2 TB of disk space so far. In this study, MODIS data processing system on WWW is developed including following functions: spectral subset (250m, 500m, 1000m channels), radiometric correction to radiance, spatial subset of geocoded data as a rectangular area with latitude-longitude grid system in HDF format, generation of a quick look file in JPEG format. Users will be notified just after all the process have finished via e-mail. Using this system enables us to process MODIS data on WWW with a few input parameters and download the processed data by FTP access. An easy to use interface is expected to promote the use of MODIS data. This system is available via the Internet on the following URL from September 1 2002, "http : //webmodis.iis.u-tokyo.ac.jp/".

7

Pre-processing of load data of agricultural tractors during major field operations

Ryu, Myong-Jin, Kabir, Md. Shaha Nur, Choo, Youn-Kug, Chung, Sun-Ok, Kim, Yong-Joo, Ha, Jong-Kyou, Lee, Kyeong-Hwan

[Kisti 연계] 충남대학교 농업과학연구소 농업과학연구 Vol.42 No.1 2015 pp.53-61

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Development of highly efficient and energy-saving tractors has been one of the issues in agricultural machinery. For design of such tractors, measurement and analysis of load on major power transmission parts of the tractors are the most important pre-requisite tasks. Objective of this study was to perform pre-processing procedures before effective analysis of load data of agricultural tractors (30, 75, and 82 kW) during major field operations such as plow tillage, rotary tillage, baling, bale wrapping, and to select the suitable pre-processing method for the analysis. A load measurement systems, equipped in the tractors, were consisted of strain-gauge, encoder, hydraulic pressure, and radar speed sensors to measure torque and rotational speed levels of transmission input shaft, PTO shaft, and driving axle shafts, pressure of the hydraulic inlet line, and travel speed, respectively. The entire sensor data were collected at a 200-Hz rate. Plow tillage, rotary tillage, baling, wrapping, and loader operations were selected as major field operations of agricultural tractors. Same or different farm works and driving levels were set differently for each of the load measuring experiment. Before load data analysis, pre-processing procedures such as outlier removal, low-pass filtering, and data division were performed. Data beyond the scope of the measuring range of the sensors and the operating range of the power transmission parts were removed. Considering engine and PTO rotational speeds, frequency components greater than 90, 60, and 60 Hz cut off frequencies were low-pass filtered for plow tillage, rotary tillage, and baler operations, respectively. Measured load data were divided into five parts: driving, working, implement up, implement down, and turning. Results of the study would provide useful information for load characteristics of tractors on major field operations.

8

Pre-processing Method of Raw Data Based on Ontology for Machine Learning

황치곤, 윤창표

[Kisti 연계] 한국정보통신학회 한국정보통신학회논문지 Vol.24 No.5 2020 pp.600-608

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

머신러닝은 학습 데이터로부터 목적함수를 구성하고, 테스트 데이터를 통해 목적함수의 확인함으로써 발생하는 데이터에 대한 예측을 수행한다. 머신러닝에서 입력데이터는 전처리 과정을 통해 정규화 과정을 거친다. 이런 정규화는 입력데이터의 평균과 표준편차를 이용하여 표준화하거나, 수치 데이터가 아닌 nominal value는 one-hot 코드 형태로 변환하는 방식을 이용한다. 그러나 이 전처리 과정만으로 문제를 해결할 수 없다. 이러한 이유로 본 논문에서 입력데이터의 정규화를 위해 온톨로지를 이용하는 방법을 제안한다. 이를 위한 테스트 데이터는 모바일 기기로부터 수집된 와이파이 장치의 RSSI값을 이용하고, 수집된 데이터의 노이즈와 이질적 문제는 온톨로지를 이용하여 정제하는 방법을 제시한다.

Machine learning constructs an objective function from learning data, and predicts the result of the data generated by checking the objective function through test data. In machine learning, input data is subjected to a normalisation process through a preprocessing. In the case of numerical data, normalization is standardized by using the average and standard deviation of the input data. In the case of nominal data, which is non-numerical data, it is converted into a one-hot code form. However, this preprocessing alone cannot solve the problem. For this reason, we propose a method that uses ontology to normalize input data in this paper. The test data for this uses the received signal strength indicator (RSSI) value of the Wi-Fi device collected from the mobile device. These data are solved through ontology because they includes noise and heterogeneous problems.

9

빅데이터 및 고성능컴퓨팅 프레임워크를 활용한 유전체 데이터 전처리 과정의 병렬화

변은규, 곽재혁, 문지협

[Kisti 연계] 한국정보처리학회 정보처리학회논문지/컴퓨터 및 통신 시스템 Vol.8 No.10 2019 pp.231-238

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

차세대 염기 서열 분석법이 생성한 유전체 원시 데이터를 기존의 방식대로 하나의 서버에서 분석하기 위해서는 데이터 크기에 따라 수십 시간이 필요할 수 있다. 그러나 응급 환자의 진단처럼 수 시간 내에 결과를 알아야 하는 상황이 존재하기 때문에 단일 유전체 분석의 성능을 향상시킬 필요가 있다. 본 연구에서는 빅데이터 기술의 병렬화 기법과 고속의 네트워크로 연결되고 병렬파일시스템을 공유하는 고성능컴퓨팅 클러스터를 적극적으로 활용하여 분석 시간을 크게 단축시킬 수 있는 유전체 데이터 분석의 전처리 프로세스의 병렬화 방법을 제안한다. 분석 데이터의 신뢰성을 위해 기존의 검증된 분석 도구 및 알고리즘을 새로운 환경에 맞게 병렬화 하는 전략을 선택하였다. 프로세스의 병렬화, 데이터의 분배 및 병렬 병합 기법을 개발하였고 실험을 통해 성능 향상을 확인하였다.

Analyzing next-generation genome sequencing data in a conventional way using single server may take several tens of hours depending on the data size. However, in order to cope with emergency situations where the results need to be known within a few hours, it is required to improve the performance of a single genome analysis. In this paper, we propose a parallelized method for pre-processing genome sequence data which can reduce the analysis time by utilizing the big data technology and the highperformance computing cluster which is connected to the high-speed network and shares the parallel file system. For the reliability of analytical data, we have chosen a strategy to parallelize the existing analytical tools and algorithms to the new environment. Parallelized processing, data distribution, and parallel merging techniques have been developed and performance improvements have been confirmed through experiments.

10

빅데이터 및 고성능컴퓨팅 프레임워크를 활용한 유전체 데이터 전처리 과정의 병렬화

변은규, 곽재혁, 문지협

[NRF 연계] 한국정보처리학회 KIPS Transactions on Computer and Communication Systems Vol.8 No.10 2019.10 pp.231-238

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

차세대 염기 서열 분석법이 생성한 유전체 원시 데이터를 기존의 방식대로 하나의 서버에서 분석하기 위해서는 데이터 크기에 따라 수십 시간이 필요할 수 있다. 그러나 응급 환자의 진단처럼 수 시간 내에 결과를 알아야 하는 상황이 존재하기 때문에 단일 유전체 분석의 성능을 향상시킬 필요가 있다. 본 연구에서는 빅데이터 기술의 병렬화 기법과 고속의 네트워크로 연결되고 병렬파일시스템을 공유하는 고성능컴퓨팅 클러스터를 적극적으로 활용하여 분석 시간을 크게 단축시킬 수 있는 유전체 데이터 분석의 전처리 프로세스의 병렬화 방법을 제안한다. 분석 데이터의 신뢰성을 위해 기존의 검증된 분석 도구 및 알고리즘을 새로운 환경에 맞게 병렬화 하는 전략을 선택하였다. 프로세스의 병렬화, 데이터의 분배 및 병렬 병합 기법을 개발하였고 실험을 통해 성능 향상을 확인하였다.

Analyzing next-generation genome sequencing data in a conventional way using single server may take several tens of hours depending on the data size. However, in order to cope with emergency situations where the results need to be known within a few hours, it is required to improve the performance of a single genome analysis. In this paper, we propose a parallelized method for pre-processing genome sequence data which can reduce the analysis time by utilizing the big data technology and the high- performance computing cluster which is connected to the high-speed network and shares the parallel file system. For the reliability of analytical data, we have chosen a strategy to parallelize the existing analytical tools and algorithms to the new environment. Parallelized processing, data distribution, and parallel merging techniques have been developed and performance improvements have been confirmed through experiments.

11

특허문서의 IPC 분류기 생성을 위한 데이터 전처리

박수현, 김진

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2024 pp.542-543

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

특허심사절차는 짧지 않은 과정으로 이루어져 있는데, 현재 모든 절차가 사람이 직접 관여하여 진행되고 있다. 특허심사절차의 효율적 시간 분배를 위해, 특허문서 분류 과정의 자동화 처리 필요성을 느끼게 되었다. 따라서, 본 논문에서는 해당 분류기 생성을 위한 데이터의 전처리 과정을 다루었다.

12

에너지신산업을 위한 에너지 빅데이터 전처리 시스템

양수영, 김요한, 김상현, 김원중

[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.16 No.5 2021 pp.851-858

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

재생에너지 및 분산자원의 증가로 에너지신산업에서는 전통적인 데이터뿐만 아니라 다양한 에너지 관련 데이터들이 생성되고 있다. 즉 다양한 재생에너지 설비와 발전 데이터, 계통 운영 데이터, 계량 및 요금 관련 데이터뿐만 아니라 새로운 서비스와 분석을 위해 필요한 기상 및 에너지 효율화 데이터 등이 있다. 에너지 빅데이터 처리 기술은 분산자원, 계통, AMI(: Advanced Metering Infrastructure)를 포함한 전력 생산·소비 인프라의 전반기에서 발생하는 데이터를 체계적으로 분석 ·진단할 수 있다. 이를 통해 ICT(: Information and Communications Technology)산업과 에너지 산업 간 융복합의 새로운 비즈니스 창출을 지원하는 기술이 될 수 있을 것이다. 이를 위해서 수집된 데이터의 항목별 특성 분석 및 연관관계 표본 추출과 각 특징들의 범주화 및 요소 정의 등 데이터 분석 시스템에 대한 연구가 필요하다. 또한 데이터의 손실 및 이상 상태 처리를 위한 데이터 정제 기술에 대한 연구가 이루어져야 한다. 그리고 에너지 데이터를 실시간으로 저장 및 관리할 수 있도록 Apache NIFI, Spark, HDFS(: Hadoop Distributed File System)에 대한 개발 및 구축이 필요하다. 본 연구에서는 위와 같은 다양한 전력거래를 위한 전반적인 에너지 데이터 처리 기술과 시스템를 제안하였다.

Due to the increase in renewable energy and distributed resources, not only traditional data but also various energy-related data are being generated in the new energy industry. In other words, there are various renewable energy facilities and power generation data, system operation data, metering and rate-related data, as well as weather and energy efficiency data necessary for new services and analysis. Energy big data processing technology can systematically analyze and diagnose data generated in the first half of the power production and consumption infrastructure, including distributed resources, systems, and AMI. Through this, it will be a technology that supports the creation of new businesses in convergence between the ICT industry and the energy industry. To this end, research on the data analysis system, such as itemized characteristic analysis of the collected data, correlation sampling, categorization of each feature, and element definition, is needed. In addition, research on data purification technology for data loss and abnormal state processing should be conducted. In addition, it is necessary to develop and structure NIFI, Spark, and HDFS systems so that energy data can be stored and managed in real time. In this study, the overall energy data processing technology and system for various power transactions as described above were proposed.

13

CCTV 영상 기반 강수량 산정을 위한 데이터 전처리 방안 연구

변종윤, 전창현, 이진욱, 김현준, 차호영

[Kisti 연계] 한국수자원학회 한국수자원학회 학술대회논문집 2022 p.167

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

최근 빅데이터에 관련된 연구에 있어 데이터의 품질관리에 대한 논의가 꾸준히 이뤄져 오고 있다. 특히 이미지 처리 및 분석에 활용되어온 딥러닝 기술의 경우, 분류 작업 및 패턴인식 등으로부터 데이터의 특징을 추출함으로써 비지도학습(Unsupervised Learning)을 가능하게 한다는 장점이 있음에도 불구하고 빅데이터를 다루는 과정에 있어 용량, 다양성, 속도 및 신뢰성 측면에서의 한계가 있었다. 본 연구에서는 CCTV 영상을 활용한 강수량 산정 모델 개발에 있어 예측 정확도 향상 및 성능 개선을 도모할 수 있는 데이터 전처리 방법을 제안하였다. 서울 근린 AWS 4개소 지역(김포장기, 하남덕풍, 강동, 성남) 및 중앙대학교 지점 내 CCTV를 설치한 후, 최대 9개월의 영상을 확보하여 강수량 산정을 위한 딥러닝 모델을 개발하였다. 배경분리, 조도조정, 영역설정, 데이터증진, 이상데이터 분류 등이 가능한 알고리즘을 개발함으로써 데이터셋 자체에 대한 전처리 작업을 수행한 후, 이에 대한 결과를 기존 관측자료와 비교·분석하였다. 본 연구에서 제안한 전처리 방법들을 적용한 결과, 강수량 산정 모델의 예측 정확도를 평가하는 지표로 선정한 평균 제곱근 편차(Root Mean Square Error; RMSE)가 약 30% 감소함을 확인하였다. 본 연구의 결과로부터 CCTV 영상 데이터를 활용한 강수량 산정의 가능성을 확인할 수 있었으며 특히, 딥러닝 모델 개발시 필요한 적정 전처리 방법들에 대한 기준을 제시할 수 있을 것으로 판단된다.

14

지진에 의한 측지학적 지각변동 분석을 위한 GNSS 자료 전처리 연구

손동효, 김두식, 박관동

[Kisti 연계] 대한공간정보학회 한국지형공간정보학회지 Vol.23 No.1 2015 pp.47-54

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

이 논문에서는 지진에 의한 지각변동 분석에서 측지학적 요소만을 구분하고자 하는 목적으로 GNSS 자료를 전처리하는 전략을 연구하였다. 이를 위해 GNSS 자료처리 결과의 해석에 앞서 GNSS 좌표 시계열에서 나타나는 위신호들을 검출하고 제거하였다. GNSS 관측소는 한반도가 포함된 큰 지각판 위에 위치하므로 판의 운동으로 인한 속도가 좌표 시계열에 포함된다. 그리고 일부 관측소 주변에 위치한 나무들은 계절에 따라 성장변화가 일어나기 때문에 계절적 신호특성이 GNSS 좌표 시계열에 반영된다. 따라서 오일러축에 의한 지각판 운동효과를 정확히 제거하기 위해 축의 위치와 각속도를 한반도 지각판에 맞게 새롭게 추정하였고 이에 대한 검증을 수행하였다. 그리고 1년 주기로 나타나는 계절변동 신호를 추정해 각 관측소의 좌표시계열에 반영하였다. 두 효과를 제거함으로써 지진에 의한 영향을 측지학적으로 분석할 수 있다. 이를 이용해 2011년 동일본 대지진에 의한 지각변위 예비 분석을 수행하였다.

In this study, we developed strategies for pre-processing GNSS data for the purpose of separating geodetic factors from crustal deformation due to the earthquakes. Before interpreting GNSS data analysis results, we removed false signals from GNSS coordinate time series. Because permanent GNSS stations are located on a large tectonic plate, GNSS position estimates should be affected by the tectonic velocity of the plate. Also, stations with surrounding trees have seasonal signals in their three-dimensional coordinate estimates. Thus, we have estimated the location of an Euler pole and angular velocities to deduce the plate tectonic velocity and verified with geological models. Also, annual amplitudes and initial phases were estimated to get rid of those false annual signals showing up in the time series. By considering the two effects, truly geodetic analysis was possible and the result was used as preliminary data for analyzing post-seismic deformation of the Korean peninsula due to the Tohoku-oki earthquake.

15

에이전트 이동성을 이용한 효율적인 전자태그 데이터 전처리 가능한 분산 소프트웨어 도구

안용선, 안진호

[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.12 No.4 2009 pp.608-615

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

RFID 기술의 발전으로 태그 가격이 급속도로 낮아짐에 따라, 더 세밀하게 제품을 관리하기 위해 각 태그가 포장상자에만 부착되는 것이 아니라 개별 제품에 부착된다. 그러나 RFID 데이터를 처리하는 판독기와 미들웨어가 한정된 하드웨어 자원을 가지고 있기 때문에, 초대용량의 태그 데이터를 신속하게 처리하기 위한 방법들이 필수적이다. 본 논문에서는 이러한 요구조건을 효율적으로 충족시키는 새로운 이동 에이전트 기반 분산형 소프트웨어 도구를 설계하고 구현한다. 이 도구는 명시된 데이터 수집 정책을 포함한 이동 에이전트를 다수의 이동식 판독기들로 전송함으로써, 필요한 데이터가 제품의 운송 중 반복적으로 전 처리 될 수 있도록 하는 편리한 환경을 제공한다. 이러한 수행 형태는 목적지에 도착한 후 매우 많은 양의 태그 데이터를 고정식 판독기에서 처리하는 기존 방식에 비해, 판독기와 미들웨어에서 매우 높은 인식률로 태그 데이터를 처리하기 위해 요구되는 소요시간을 상당히 줄일 수 있다.

As RFID tag prices have rapidly been declining because of the advance of RFID technology, each tag is attached to an individual item, not a packing box only, for managing the item much more precisely. However, some mechanisms are essential to handle a very large amount of tag data quickly because readers and middlewares processing RFID data have limited hardware resources. In this paper, we design and implement a new mobile agent-based distributed software tools to satisfy this requirement efficiently. These tools provide a convenient environment enabling required data to be pre-processed repeatedly in transit by transferring a mobile agent including its specified data collection policy to numerous mobile readers. This behavior can significantly reduce the elapsed time required for processing huge volumes of tag data at the readers and middlewares with their very high recognition rates compared with the existing one to process the data by fixed readers after having arrived at the destination

16

데이터 전처리를 고려한 하수처리장 머신러닝 모델 개발

심규대, 김효상, 박찬수, 김동균, 김신걸

[Kisti 연계] 한국수자원학회 한국수자원학회 학술대회논문집 2023 p.495

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 하수처리장 운영시스템 자료를 활용하여, 머신러닝 기반의 예측 모델을 개발하고, 모델 정확도 향상에 대하여 검토하였다. 하수처리장에 설치된 각종 센서를 통해 실시간으로 자료가 모니터링되고 있으며, 수집된 자료는 운영시스템에 저장된다. 하수처리장 시스템은 설정된 값과 센서의 측정값을 비교해 이상치가 발생하면 운영자가 즉각적으로 조치하여 문제를 해결하고 있으나, 비정상적인 상황 발생시 이를 대처할 시간이 부족하여 적절한 조치가 이루어지지 못하는 경우가 발생 되고 있다. 따라서, 이러한 문제점을 해결하기 위해 A 하수처리장 운영자료를 활용하여 결과 예측이 신속하고 신뢰도 높은 머신러닝 기반의 예측 모델을 개발하고자 하였다. 모델의 예측 정확도 및 신뢰성을 향상하기 위하여 결과에 영향을 미치는 주요 영향 인자를 분석하고, 이를 기반으로 모델의 추가 분석 및 개선을 수행하여 모델의 예측력을 평가하였다. 금회 연구는 데이터 전처리를 과정을 통한 인사이트를 도출하고 이를 활용하여 하수처리장 운영자료 예측 정확도를 높일 수 있었으며, 이 결과를 바탕으로 다른 하수처리장의 모델 개발시에도 유용하게 활용이 가능할 것으로 검토되었다.

17

CT영상에서의 AlexNet과 VggNet을 이용한 간암 병변 분류 연구

최보혜, 김영재, 최승준, 김광기

[Kisti 연계] 대한의용생체공학회 Journal of biomedical engineering research : the official journal of the Korean Society of Medical & Biological Engineering Vol.39 No.6 2018 pp.229-236

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Liver cancer is one of the highest incidents in the world, and the mortality rate is the second most common disease after lung cancer. The purpose of this study is to evaluate the diagnostic ability of deep learning in the classification of malignant and benign tumors in CT images of patients with liver tumors. We also tried to identify the best data processing methods and deep learning models for classifying malignant and benign tumors in the liver. In this study, CT data were collected from 92 patients (benign liver tumors: 44, malignant liver tumors: 48) at the Gil Medical Center. The CT data of each patient were used for cross-sectional images of 3,024 liver tumors. In AlexNet and VggNet, the average of the overall accuracy at each image size was calculated: the average of the overall accuracy of the $200{\times}200$ image size is 69.58% (AlexNet), 69.4% (VggNet), $150{\times}150$ image size is 71.54%, 67%, $100{\times}100$ image size is 68.79%, 66.2%. In conclusion, the overall accuracy of each does not exceed 80%, so it does not have a high level of accuracy. In addition, the average accuracy in benign was 90.3% and the accuracy in malignant was 46.2%, which is a significant difference between benign and malignant. Also, the time it takes for AlexNet to learn is about 1.6 times faster than VggNet but statistically no different (p > 0.05). Since both models are less than 90% of the overall accuracy, more research and development are needed, such as learning the liver tumor data using a new model, or the process of pre-processing the data images in other methods. In the future, it will be useful to use specialists for image reading using deep learning.

18

과실류의 비파괴 내부품질 인자의 예측 성능 향상을 위한 VIS/NIR 스펙트럼 전처리 기법 개발

류동수, 황인근, 노상하

[Kisti 연계] 한국농업기계학회 한국농업기계학회 학술대회논문집 2000 pp.451-456

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

19

전-후 처리 과정을 포함한 거대 구조물의 유한요소 해석을 위한 효율적 데이터 구조

박시형, 박진우, 윤태호, 김승조

[Kisti 연계] 한국전산구조공학회 한국전산구조공학회 학술대회논문집 2004 pp.389-395

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

We consider the interface between the parallel distributed memory multifrontal solver and the finite element method. We give in detail the requirement and the data structure of parallel FEM interface which includes the element data and the node array. The full procedures of solving a large scale structural problem are assumed to have pre-post processors, of which algorithm is not considered in this paper. The main advantage of implementing the parallel FEM interface is shown up in the case that we use a distributed memory system with a large number of processors to solve a very large scale problem. The memory efficiency and the performance effect are examined by analyzing some examples on the Pegasus cluster system.

 
페이지 저장