Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 19
No
1

최근 보안 업계에서 새롭게 화두가 되고 있는 XDR(eXtended Detection and Response)은 엔드포인트 와 네트 워크 영역의 모든 정보에 대한 통합 연계 분석을 통하여, 기업 내 보안 목표를 보다 효과적 으로 달성할 수 있는 솔 루션이다. 그러나, XDR에서 추구하는 통합 연계 분석을 위해서는 기업내 에서 운영되는 이기종 보안 솔루션들 간 의 서로 다른 프로토콜과 정보 저장 방식의 문제를 해결해야 한다. 서로 다른 형태의 데이터를 통합하고 이를 분석 하기 위한 얼라이언스와 통합 레이블링 등의 표준화를 제안한 연구들이 있지만, 연동하는 모든 보안솔루션이 표준 API를 기반으로 개발되어야 하는 현실적인 어려움이 존재한다. 이 문제를 해결하기 위해 본 논문에서는 다양한 보 안솔루션 데이터 중에서 가장 기본적이고 필수적인 데이터인 비정형 로그를 중심으로 계층적 구조 기반의 데이터 정 규화 방법을 제안한다. 공통적인 CODE와 Corpus로 구성한 정규화 모델 알고리즘으로 정제된 비정형로그 데이터 는 계층적 구조를 가지는 정규화된 데이터 형태로 출력된다.

XDR(eXtended Detection and Response) is a new topic in the security industry. It is a solution to achieve security goals within the enterprise more effectively through integrated linkage analysis of all information in the endpoint and network area. However, for the integrated linkage analysis of XDR, it is necessary to solve the problem of different protocols and information storage methods between heterogeneous security solutions operated within the enterprise. Standardization approaches such as alliance and unified labeling to integrate and analyze different types of data have been proposed. However, those approaches suffer from some practical difficulties that all interworking security solutions are developed based on standard APIs. To overcome the difficulties, in this paper, we propose a new method to normalize the hierarchical structure-based data of the unstructured log, which is the most basic and essential data in the security solution data. The unstructured log data refined through the normalization model algorithm composed of common CODE and Corpus is output in the form of normalized data having a hierarchical structure.

2

Background: Identification of radioisotopes for plastic scintillation detectors is challenging because their spectra have poor energy resolutions and lack photo peaks. To overcome this weakness, many researchers have conducted radioisotope identification studies using machine learning algorithms; however, the effect of data normalization on radioisotope identification has not been addressed yet. Furthermore, studies on machine learning-based radioisotope identifiers for plastic scintillation detectors are limited. Materials and Methods: In this study, machine learning-based radioisotope identifiers were implemented, and their performances according to data normalization methods were compared. Eight classes of radioisotopes consisting of combinations of 22Na, 60Co, and 137Cs, and the background, were defined. The training set was generated by the random sampling technique based on probabilistic density functions acquired by experiments and simulations, and test set was acquired by experiments. Support vector machine (SVM), artificial neural network (ANN), and convolutional neural network (CNN) were implemented as radioisotope identifiers with six data normalization methods, and trained using the generated training set. Results and Discussion: The implemented identifiers were evaluated by test sets acquired by experiments with and without gain shifts to confirm the robustness of the identifiers against the gain shift effect. Among the three machine learning-based radioisotope identifiers, prediction accuracy followed the order SVM >ANN>CNN, while the training time followed the order SVM>ANN>CNN. Conclusion: The prediction accuracy for the combined test sets was highest with the SVM. The CNN exhibited a minimum variation in prediction accuracy for each class, even though it had the lowest prediction accuracy for the combined test sets among three identifiers. The SVM exhibited the highest prediction accuracy for the combined test sets, and its training time was the shortest among three identifiers.

3

발전 제어시스템의 융합보안 연구 KCI 등재

이대성

한국융합보안학회 융합보안논문지 제18권 제5호 2018.12 pp.93-98

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

한국수력원자력, 한국전력공사, 한국남동발전 등의 발전 제어시스템은 전력을 공급하는 국가의 주요 인프라 시설로 악의적인 해킹 공격이 진행될 경우 그 피해는 상상을 초월한다. 실제로 한국수력원자력은 해킹 공격을 당하여 내부 정보가 유출되는 등 사회적인 큰 문제를 야기하였다. 본 연구에서는 최근 이슈가 되고 있는 융합보안 연구에 대 해 전력회사 중심의 발전 제어시스템을 대상으로 그 환경을 분석하고, 현황을 분석하여 다양한 발전 제어시스템 의 안정화를 위한 전략체계 수립과 대응책을 제시하고자 한다. 다양한 물리적 보안시스템(시설), IT 보안시스템, 출입통제시스템 등에서 나오는 데이터 형태를 정규화하고 통합하여 융합인증을 통해 전체 시스템을 통제하고, 융합관제를 통해 위험을 탐지하는 방법을 제안한다.

Korea Hydro & Nuclear Power Co., Ltd., Korea Electric Power Corporation, and Korea South-East Power Corporat ion are major infrastructure facilities of power supplying countries. If a malicious hacking attack occurs, the damage is beyond the imagination. In fact, Korea Hydro & Nuclear Power has been subjected to a hacking attack, causing internal information to leak and causing social big problems. In this paper, we propose a strateg y and countermeasures for stabilization of various power generation control systems by analyzing the en vironment and the current status of power generation control system for convergence security research, which is becoming a hot issue. We propose a method to normalize and integrate data types from various physical security systems (facilities), IT security systems, access control systems, to control the whole s ystem through convergence authentication, and to detect risks through fusion control.

4

LiDAR 센서를 활용한 배회 동선 검출 알고리즘 개발 KCI 등재

정은비, 유소영

한국ITS학회 한국ITS학회논문지 제16권 제6호 통권74호 2017.12 pp.1-15

※ 기관로그인 시 무료 이용이 가능합니다.

4,800원

최근 국제적인 테러 위협이 불특정 다수를 대상으로 발생하고 있으며, 이러한 위협에서 시민을 보호하기 위한 다양한 대책이 논의 중이다. 저렴해진 센서 기술을 활용한 사전 감시 시스템에 대한 요구가 높아지고 있으나, 보행 궤적의 고유 특성 검출 및 상세 분석 연구가 미비한 실정이다. 본 연구에서는 상용화된 보행 동선 솔루션을 활용하여, 삼성역 개찰구에서 코엑스와 직접 연결되는 연결 통로 (3-6번 출구 근처) 일대의 보행 동선 궤적 조사를 수행하 였다. 조사된 궤적 자료를 바탕으로, 궤적 자료의 정규화 기법, Clustering 방법을 중심으로 보 행 궤적을 유형화하고 배회 동선을 추출하는 분석 방법론을 제시하였다. 분석 결과, 동일 군 집내에서 유사성이 크게 떨어지는 보행 궤적의 검출 가능성을 검증하였다.

Recently terrorism targets unspecified masses and causes massive destruction, which is so-called Super Terrorism. Many countries have tried hard to protect their citizens with various preparation and safety net. With inexpensive and advanced technologies of sensors, the surveillance systems have been paid attention, but few studies associated with the classification of the pedestrians’ trajectories and the difference among themselves have attempted. Therefore, we collected individual trajectories at Samseoung Station using an analytical solution (system) of pedestrian trajectory by LiDAR sensor. Based on the collected trajectory data, a comprehensive framework of classifying the types of pedestrians’ trajectories has been developed with data normalization and “trajectory association rule-based algorithm.” As a result, trajectories with low similarity within the very same cluster is possibly detected.

5

4,900원

상수관망 내 무수율은 관로의 파손, 운영 손실, 물리적 요소 등에 의해 발생하는 수도공급량에 대한 손실 비율을 나타낸다. 국내에서 시행하는 관망정비사업, 유수율 제고 사업에 있어서 무수율은 시⋅ 도 및 소블록에 대한 비교 지표로써 무수율에 영향을 주는 인자의 발굴 및 무수율 추정기법에 대한 연구는 수도공급시설의 경제성과 연관되어 점차 중요해지고 있다. 본 연구에서는 무수율 추정하기 위한 방법으로 통계분석 기법 중 주성분 분석(PCA), 인공신경망(ANN)을 이용하여 무수율을 추정하 였다. 연구를 위하여 상수관망 주요 영향인자에 대한 데이터를 수집하였고, Z-score방법을 통하여 데이터를 정규화 한 후 PCA-ANN 모형에 적용하여 무수율을 추정하였다. 무수율 모의 결과와 실측 무수율을 비교하기 위하여 정확도 평가를 수행하였다. 연구결과 무수율 예측에 있어 PCA-ANN기법 이 원데이터를 이용하여 ANN을 단독으로 모의하는 조건보다 정확도가 높은 것으로 나타났다. 또한 ANN 모형 구축시 은닉층 내의 뉴런수에 따라 추정결과가 상이하며 본 연구에서 사용된 6개의 독립 변수에 대하여, 12개의 뉴런을 이용한 조건에서 무수율 예측정확도가 높은 것으로 나타났다.

The non-revenue water (NRW) ratio in water distribution systems is an index of the loss of water supply caused by pipe burst, operational loss and physical factors. NRW ratio is a comparative index for city, province and DMA (district metering area) in the domestic water supply maintenance project. An investigation of the factors affecting the NRW as well as its estimation have become increasingly important in an economic sense. In this study, PCA (principal component analysis) and ANN (artificial neural network) are used as statistical methods to estimate the NRW ratio. The normalized data were obtained through the Z-score method, and then the PCA-ANN model was constructed for the NRW ratio estimation. Accuracy assessment was performed to compare the observed NRW ratio with the estimated ratio from the ANN model. The results show that the PCA-ANN model is more accurate than the single ANN and the estimation results differ by the number of neurons in the hidden layer of ANN. As for the six independent variables used in this study, the accuracy of NRW ratio prediction was found highest when 12 neurons were used.

6

4,500원

최대 리아푸노프 지수(Maximum Lyapunov Exponent, MLE)는 보행 안정성을 평가하는 핵심 지표로 부상하 였으나, 연구 간 방법론적 불일치로 인해 비교 분석에 어려움이 있었다. 본 연구는 지면 보행 중 MLE 추정에 영향을 미치 는 보폭 수(50, 100, 150, 200 보폭), 데이터 정규화(비정규화 vs. 정규화된 시계열), 그리고 Average Mutual Information-False Nearest Neighbor, AMI-FNN 매개변수화의 효과를 체계적으로 분석하였다. 주요 신체 마커 (C7, 오른쪽 무릎, 오른쪽 발목)의 모션 캡처 데이터를 Wolf 및 Rosenstein 알고리즘을 이용해 분석하였다. 선형 혼합 효과 모형 결과, 보폭 수가 MLE 값에 유의한 영향을 미치는 것으로 나타났으며, 긴 데이터 세트(100-200 보폭)가 짧은 세트(50 보폭)보다 더 안정적인 추정치를 산출하였다. 데이터 정규화는 보행의 핵심 역학 특성을 유지하면서 계산적 변동 성을 줄였다. AMI 기반의 시간 지연(time delay)은 보폭 수에 따라 일정하게 유지되었고, FNN으로 계산된 매개변수(임 베딩 차원)는 100-150 보폭 이후 5에서 안정화되었다. Wolf 알고리즘은 Rosenstein 알고리즘보다 일관되게 낮은 MLE 값을 산출하였다. 그러나 표준화된 매개변수를 적용한 사후 분석에서는 두 알고리즘 간 유의한 차이가 나타나지 않았다. 본 연구는 MLE 계산의 표준화에 대한 근거 기반 권고안을 제시하며, 보행 안정성 평가의 방법론적 부분에서 일관성의 중요성을 강조한다고 판단된다.

The Maximum Lyapunov Exponent (MLE) has emerged as a critical metric for assessing gait stability, yet methodological inconsistencies across studies have hindered comparative analyses. This study systematically investigates the effects of stride count (50, 100, 150, 200 strides), data normalization (raw vs. normalized time series), and Average Mutual Information–False Nearest Neighbor (AMI–FNN) parameterization on MLE estimation during overground walking. Motion capture data from key anatomical markers (C7,right knee, right ankle) were analyzed using Wolf and Rosenstein algorithms. Linear mixed-effects models revealed significant effects of stride count on MLE values, with longer datasets (100-200 strides) producing more stable estimates than shorter ones (50 strides). Data normalization preserved essential gait dynamics while reducing computational variability. AMI-derived time delays remained consistent across stride counts, while FNN-determined embedding dimensions stabilized at 5 after 100-150 strides. The Wolf algorithm consistently produced lower MLE estimates than Rosenstein , though post-hoc analyses revealed no significant differences when standardized parameters were applied. This study presents evidence-based recommendations for the standardization of MLE calculation and emphasizes the importance of methodological consistency in gait stability assessment.

7

Integrating Normalization with Random Tree Data Mining Approach for Mining Cloud Based Big Data’s of Asthma Patient’s SCOPUS

Abhinav Hans, Sheetal Kalra

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.6 2015.12 pp.165-174

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The potential of cloud computing for overriding the needs for deploying various infrastructures for running a server based services brought up a revolutionary change in the way the traditional demands of the people use to be handled. Cloud computing provides the rental service for the user in which a user can use the particular software by paying for that on the cloud server. Since the whole scenario is beneficial to big industries like Facebook, Google, Orkut etc., various other fields are also getting dependent on cloud computing. Since tons of data is uploading every second to the cloud server does need to be mined properly for efficient data storage. In this paper we try to integrate the data preprocessing technique with data classification technique to mine big data’s of asthma based patients. We have used simulation tool called eclipse to run the API’s of weka and cloudsim for setting up the experimental environment.

8

Effect of local background intensities in the normalization of cDNA microarray data with a skewed expression profiles

김진혁, 신동미, 이용성

[NRF 연계] 생화학분자생물학회 Experimental and Molecular Medicine Vol.34 No.3 2002.07 pp.224-232

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Normalization of the data of cDNA microaray is anobligatory step during microarray experiments dueto the relatively frequent non-specific erors. Gener-ally, normalization of microaray data is based onthe null hypothesis and variance model. In theYangs model (Yang et al., 2001), at least two types ofnoises are included. The one is additive noise andthe other is multiplicative noise. Usually, backgroundis considered as one of additive noise to the signaland the variation between the signal pixels is therepresentative multiplicative noise. In this study, therelation between the signal (spot intensity minusbackground intensity) and background was observedand the influence of background on normalizationas a representative additive factor was investigated.Although the relation has not been considered as afactor affecting the normalization, it could improvethe accuracy of microaray data when the normaliza-tion was carried out considering signal/backgroundratio. The background dependent normalization de-creased the number of genes whose expression lev-els were changed significantly and it could maketheir distribution more consistent through the wholerange of signal intensities. In this study, printing pindependent normalization was also carried outregarding the printing pin as a representative multi-plicative noise. It improved the distribution of spotsin the Cy3-Cy5 scatter plot, but its effect was slight.These studies suggest that there are some influ-ences of the signals on the local backgrounds andthey must be considered for the normalization ofcDNA microarray data.

9

An Iterative Normalization Algorithm for cDNA Microarray Medical Data Analysis

Kim, Yoonhee, Park, Woong-Yang, Kim, Ho

[Kisti 연계] 한국유전체학회 Genomics & informatics Vol.2 No.2 2004 pp.92-98

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

A cDNA microarray experiment is one of the most useful high-throughput experiments in medical informatics for monitoring gene expression levels. Statistical analysis with a cDNA microarray medical data requires a normalization procedure to reduce the systematic errors that are impossible to control by the experimental conditions. Despite the variety of normalization methods, this. paper suggests a more general and synthetic normalization algorithm with a control gene set based on previous studies of normalization. Iterative normalization method was used to select and include a new control gene set among the whole genes iteratively at every step of the normalization calculation initiated with the housekeeping genes. The objective of this iterative normalization was to maintain the pattern of the original data and to keep the gene expression levels stable. Spatial plots, M&A (ratio and average values of the intensity) plots and box plots showed a convergence to zero of the mean across all genes graphically after applying our iterative normalization. The practicability of the algorithm was demonstrated by applying our method to the data for the human photo aging study.

10

Comparison of Normalization Methods for Defining Copy Number Variation Using Whole-genome SNP Genotyping Data

Kim, Ji-Hong, Yim, Seon-Hee, Jeong, Yong-Bok, Jung, Seong-Hyun, Xu, Hai-Dong, Shin, Seung-Hun, Chung, Yeun-Jun

[Kisti 연계] 한국유전체학회 Genomics & informatics Vol.6 No.4 2008 pp.231-234

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Precise and reliable identification of CNV is still important to fully understand the effect of CNV on genetic diversity and background of complex diseases. SNP marker has been used frequently to detect CNVs, but the analysis of SNP chip data for identifying CNV has not been well established. We compared various normalization methods for CNV analysis and suggest optimal normalization procedure for reliable CNV call. Four normal Koreans and NA10851 HapMap male samples were genotyped using Affymetrix Genome-Wide Human SNP array 5.0. We evaluated the effect of median and quantile normalization to find the optimal normalization for CNV detection based on SNP array data. We also explored the effect of Robust Multichip Average (RMA) background correction for each normalization process. In total, the following 4 combinations of normalization were tried: 1) Median normalization without RMA background correction, 2) Quantile normalization without RMA background correction, 3) Median normalization with RMA background correction, and 4) Quantile normalization with RMA background correction. CNV was called using SW-ARRAY algorithm. We applied 4 different combinations of normalization and compared the effect using intensity ratio profile, box plot, and MA plot. When we applied median and quantile normalizations without RMA background correction, both methods showed similar normalization effect and the final CNV calls were also similar in terms of number and size. In both median and quantile normalizations, RMA backgroundcorrection resulted in widening the range of intensity ratio distribution, which may suggest that RMA background correction may help to detect more CNVs compared to no correction.

11

The Classification System of Microarray Data Using Adaptive Simulated Annealing based on Normalization.

박수영, 정채영

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2006 pp.69-72

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

최근 생명 정보학 기술의 발달로 마이크로 단위의 실험조작이 가능해짐에 따라 하나의 chip상에서 전체 genome의 expression pattern을 관찰할 수 있게 되었고, 동시에 수 만개의 유전자들 간의 상호작용도 연구가능하게 되었다. 이처럼 DNA 마이크로어레이 기술은 복잡한 생물체를 이해하는 새로운 방향을 제시해주게 되었다. 따라서 이러한 기술을 통해 얻어진 대량의 유전자 정보들을 효과적으로 분석하는 방법이 시급하다. 본 논문에서는 마이크로어레이 실험에서 다양한 원인에 의해 발생하는 잡음(noise)을 줄이거나 제거하는 과정인 정규화과정을 거쳐 특징 추출방법인 SVM(Support Vector Machine) 방법을 이용하여 데이터를 2개의 클래스로 나누고, 표준화 방법들의 성능 비교를 위해 Adaptive Simulated Annealing 알고리즘으로 정확도를 평가하는 분류 시스템을 설계 구현하였다.

12

네트워크 데이터 정형화 기법을 통한 데이터 특성 기반 기계학습 모델 성능평가

이우호, 노봉남, 정기문

[Kisti 연계] 한국정보보호학회 정보보호학회논문지 Vol.29 No.4 2019 pp.785-794

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

최근 4차 산업 혁명 기술 중 하나인 딥러닝(Deep Learning) 기술은 보안 분야에서는 탐지하기 어려운 네트워크 데이터의 숨겨진 의미를 식별하고 공격을 예측하는 데 사용되고 있다. 침입탐지에 사용될 딥러닝 알고리즘을 선택하기 전에 데이터의 속성과 품질 분석이 필요하다. 학습에 사용되는 데이터의 오염여부에 따라 탐지 방법에 영향을 주기 때문이다. 따라서 데이터의 특징을 파악하고 특성을 선정해야 한다. 본 논문에서는 네트워크 데이터 셋을 이용하여 악성코드의 단계적 특징을 분석하고 특성을 추출하여 딥러닝 모델을 적용하였을 때 각 특성이 성능에 미치는 영향을 분석하였다. 네트워크 특징에 따른 특성들의 비교에 대한 트래픽 분류 실험을 진행하였으며 선정한 특성을 기반으로 96.52% 정확도를 분류하였다.

Recently Deep Learning technology, one of the fourth industrial revolution technologies, is used to identify the hidden meaning of network data that is difficult to detect in the security arena and to predict attacks. Property and quality analysis of data sources are required before selecting the deep learning algorithm to be used for intrusion detection. This is because it affects the detection method depending on the contamination of the data used for learning. Therefore, the characteristics of the data should be identified and the characteristics selected. In this paper, the characteristics of malware were analyzed using network data set and the effect of each feature on performance was analyzed when the deep learning model was applied. The traffic classification experiment was conducted on the comparison of characteristics according to network characteristics and 96.52% accuracy was classified based on the selected characteristics.

13

시큐어 코딩을 적용한 입력데이터 정규화 검증 연구

이지선, 최진영

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2013 pp.644-647

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

인터넷과 정보기술의 발전으로 정보시스템들이 보편화 되고, 편리함을 제공하고 있다. 반면에 시스템은 더욱 복잡해지고, 프라이버시 침해, 개인정보 수집 등 사이버공격은 계속적으로 증가하고 있으며 이로 인한 피해가 심각하다. 사이버 공격을 예방하기 위해서는 정보시스템 제품출시 이전 단계에서 제품의 보안 취약점을 제거하는 것이 중요하다. 따라서 개발단계부터 보안을 고려한 소프트웨어를 개발하는 것은 향후 발생 가능한 보안취약점을 예방하고 피해를 최소화 하여 보다 안전한 소프트웨어를 개발하는 근본적인 해결책이 된다. 본 논문에서는 소프트웨어 개발과정에서 발생할 수 있는 보안약점을 최소화 하여 안전한 소프트웨어를 개발하기 위한 시큐어 코딩(secure coding)과 입력 데이터 값(문자열)을 정규화 함으로써 크로스 사이트 스크립팅(XSS)의 공격을 사전에 예방할 수 있는 방법을 제시한다.

14

정량지표의 정규화 작업에 따른 순위 역전 현상과 해결방안

박영선

[NRF 연계] 글로벌경영학회 글로벌경영학회지 Vol.14 No.6 2017.12 pp.215-237

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

국내·외적으로 발표되는 대학순위는 그 영향력이 매우 크지만 순위를 결정짓는 평가지표와 가중치 그리고 점수 산출방법에 대한 학문적 연구는 미흡한 실정이다. 최근에 다기준 의사결정 방법을적용하여 대학 순위를 결정하는 연구가 진행되고 있으나 그 결과에 대한 타당성을 입증해 보이지못하고 있다. 따라서 본 연구는 순위 역전이 발생하는 비율을 타당성 척도로 보고 이 비율이 낮게나타나는 대학 순위 평가모형의 개발을 목적으로 진행하였다. TOPSIS와 VIKOR을 대학 순위 결정 문제에 적용해 보았고 이들의 장점을 살려 새로운 평가모형을 개발하였다. 영국 CUG에서 발표한 자료를 이용하여 본 연구에서 개발한 평가모형으로 전체 대학의 순위를 구했고, 이상점이 있는 대학을 제외한 경우와 일부 대학만을 임의로 선택한 경우에 대해서도 순위를 구했다. TOPSIS 와 VIKOR 그리고 z-변형 모형으로도 위의 세 가지 경우에 대해 순위를 구하였다. 본 연구에서 개발한 평가모형으로 구한 순위에서 TOPSIS와 VIKOR로 구한 순위보다 순위 역전 현상이 덜 발생한다는 사실을 발견하였다. 본 연구에서 개발한 평가모형은 기존의 대학 순위 평가에서 많이 사용하는 z-변형 정규화를 대체할 수 있을 것으로 기대한다.

College and university rankings are conducted by various organizations using various combinations of multiple factors. Recently, the multi-criteria decision making methods, such as TOPIS and VIKOR, have been used to calculate the rankings, but the validity of those methods are still in question. The current study focused on the rank reversal phenomenon, which occurs when the number of universities in consideration is different. This study aimed to develop a new method that reduces the rank reversal phenomenon prevalent in multi-criteria decision problems. Specifically, we modified TOPSIS by using a linear normalization instead of a vector normalization. Using the published data in the Complete University Guide 2017, we calculated university rankings by applying TOPSIS, VIKOR, z-transform, and modified TOPSIS developed in this study. These four methods were used to obtain the rankings considering the following two cases: when only the outliers were removed from the data and when 30 universities were randomly selected from the data. The modified TOPSIS showed less rank reversal compared to TOPSIS and VIKOR in both cases. The results suggest that the modified TOPSIS could be a substitute for the z-transform, which is a method currently widely used in the university ranking calculation.

15

식품 데이터 정규화를 위한 쌀 음식의 건물중 기반 영양 편차 고찰

김상철, 이운용, 박우풍, 윤기오, 김종린

[Kisti 연계] 한국스마트미디어학회 스마트미디어저널 Vol.11 No.7 2022 pp.76-84

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

동일한 재료를 사용하고, 식품명이나 음식명이 같음에도 불구하고 동일한 중량에서 식품의 영양성분이 편차를 나타내는 경우가 많이 있다. 그 원인은 조리 방법과 조리 공정에 따른 음식의 수분함량과 깊은 관계가 있다. 개인의 건강 맞춤형 식단을 설계하고, 정확한 열량과 양분을 공급하기 위해서는 조리 공정이나 조리 방법에 영향을 받지 않는 음식 데이터의 표시 방법이 필요하다. 이 연구에서는 동일한 식자재나 식품이 함수율의 차이로 인해 다른 식자재나 식품으로 분류됨으로 데이터베이스의 복잡성과 활용측면의 어려움이 증가하는 문제를 개선하기 위해 건물중(乾物重) 기반의 식품 데이터 표시를 제안하고자 하며, 이를 위해 식품재료로서 쌀의 특징과 쌀을 재료로 한 다양한 쌀 가공 식품의 물성에 대하여 수분의 변화에 따른 주요 영양성분의 변화를 고찰하고, 이를 통해 식품 데이터를 정규화 하기 위한 예시로서 쌀의 건물중 기반 영양 표시를 제안하고자 하였다. 동일한 재료로 가공된 32종의 쌀 가공 식품 데이터는 수분 분포에 있어 1.1~95%, 에너지량은 20~415kcal, 단백질은 0.3~9.1g, 지질은 0.1~3.9g, 탄수화물은 4.4~91.0g의 범위로 매우 넓은 영역에 분포하고 있다. 그러나 수분영향을 제거하고 고형물로 환산한 쌀가공 식품의 100g 당 영양성분은 에너지량의 최대값과 최소값의 범위는 376.9~421.1kcal, 단백질의 최대값과 최소값의 범위는 4.3~12.6g, 지질의 최대값과 최소값의 범위는 0.1~4.1g, 탄수화물의 최대값과 최소값의 범위는 80.5~95.1g 로 나타났다. 수분 중량을 포함한 음식의 영양성분 데이터에 비해 최대값과 최소값, 데이터의 표준편차가 90%이상 감소하고, 정규화되는 경향을 나타내었다.

In Korea, where rice is the staple food, there are many cases in which the nutritional composition of food is different at the same weight, even though the same ingredients are used and the food or food name is the same. The cause is closely related to the moisture content of the food according to the cooking method and cooking process. In order to design a diet tailored to individual health and supply accurate calories and nutrients, a method of expressing food data that is not affected by the cooking process or cooking method is required. Usually, the same ingredients or foods show a lot of deviation from the nutritional components presented in the standard food database due to the difference in moisture content. For this reason, there are problems that increase the complexity of the food ingredient database and the difficulty in using it. As a method to improve these problems, we would like to propose a food data expression method based on dry weight. As an example of this, the characteristics of rice as a food material and changes in major nutritional components according to the change in moisture of various rice-processed foods made from rice were considered. In addition, as an example of how to normalize food data through this, the dry weight-based nutrition label of rice was presented.

16

LED 열화 데이터에 대한 정규화 방법에 대한 연구

정의효, 임홍우, 형재필, 정창욱, 조정하, 장중순

[Kisti 연계] 한국신뢰성학회 신뢰성응용연구 Vol.18 No.1 2018 pp.49-55

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Purpose: To propose improved method for normalization, compare to the de facto international standard which is IESNA TM-21 or conventional normalization methods. Methods: Firstly, we analysed conventional methods and specified the problem of normalization method which is based on first measured data. Secondly, we proposed our approach which is based on the design specification. Lastly, we studied a real degradation data which is conducted for 15,000 hours. Conclusion: Proposed normalization method is better approach because it can reflect real data and design specification, and reduce distortion when analysing degradation data. Also, It is appliable to other long-life reliability items.

17

한반도 식생에 대한 MODIS 250m 자료의 BRDF 효과에 대한 반사도 정규화

염종민, 한경수, 김영섭

[Kisti 연계] 대한원격탐사학회 대한원격탐사학회지 Vol.21 No.6 2005 pp.445-456

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

지표변수는 지면 근처의 기후변화에 중요한 역할을 하기 때문에, 충분히 높은 정확성을 가진 값이 산출되어야 한다 하지만 지표 반사도는 강한 이방성(non-Lambertian) 특징을 가지고 있기 때문에, 위성 천저각으로부터 멀어질수록 태양-지점-위성과의 기하학적 영향을 더욱 강하게 받는 효과를 가져온다. 또한 지표 각 영향을 포함하고 있는 지표 반사도는 노이즈를 가지게 된다. 따라서 본 연구의 목적은 한반도 지역의 MODIS 반사도 자료(250m)를 이용하여 각 영향이 제거된 보다 정확한 반사도 값에 대한 데이터베이스를 제공하는 것이다. 본 연구에서는 매일 2회씩 제공하는 MODIS(Moderate Resolution Imaging Spoctroradiometer) 센서의 가시영역과 근적외영역의 반사도(250m)자료를 이용하였다. 먼저 구름화소를 제거하기 위해서 연속적인 물리과정을 통하여 각각의 구름 화소를 제거하였다. 그리고 지리보정은 MODIS 센서에서 제공하는 지리정보자료를 이용하여 2차 다항회귀식을 통한 최근접 내삽법을 사용하였다. 본 연구에서는 지표 이방성 효과를 보정하기 위해서 반 경험적 양방향성분포함수(BRDF) 모델을 사용하였다. 이 알고리즘은 위성으로부터 관측된 위성천정각, 태양천정각, 위성방위각, 태양방위각과 같은 각 성분을 이용하여 Kernel-deriven 모델의 역변환을 통하여 지표 반사도를 재생산한다. 먼저 우리는 BRDF 모델을 수행하기 위해 총 31일 모델 관측 실행기간을 고려하였다. 다음 단계로 각각의 화소 및 밴드에 대해서 BRDF 모델을 통하여 분리된 각 성분들을 변조함으로써 위성 직하점 반사도 정규화가 수행되었다. 모델을 이용하여 산출된 반사도 값은 실제 위성 반사도 값과 잘 일치하였고, RMSE(Root Mean Square Error)값은 전체적으로 약 0.01(최고값=0.03)이였다. 마지막으로, 우리는 한반도 지역에 대해서 2001년 동안 총 36개로 구성된 정규화 지표반사도 값의 데이터베이스를 구축하였다.

The land surface parameters should be determined with sufficient accuracy, because these play an important role in climate change near the ground. As the surface reflectance presents strong anisotropy, off-nadir viewing results a strong dependency of observations on the Sun - target - sensor geometry. They contribute to the random noise which is produced by surface angular effects. The principal objective of the study is to provide a database of accurate surface reflectance eliminated the angular effects from MODIS 250m reflective channel data over Korea. The MODIS (Moderate Resolution Imaging Spectroradiometer) sensor has provided visible and near infrared channel reflectance at 250m resolution on a daily basis. The successive analytic processing steps were firstly performed on a per-pixel basis to remove cloudy pixels. And for the geometric distortion, the correction process were performed by the nearest neighbor resampling using 2nd-order polynomial obtained from the geolocation information of MODIS Data set. In order to correct the surface anisotropy effects, this paper attempted the semiempirical kernel-driven Bi- directional Reflectance Distribution Function(BRDF) model. The algorithm yields an inversion of the kernel-driven model to the angular components, such as viewing zenith angle, solar zenith angle, viewing azimuth angle, solar azimuth angle from reflectance observed by satellite. First we consider sets of the model observations comprised with a 31-day period to perform the BRDF model. In the next step, Nadir view reflectance normalization is carried out through the modification of the angular components, separated by BRDF model for each spectral band and each pixel. Modeled reflectance values show a good agreement with measured reflectance values and their RMSE(Root Mean Square Error) was totally about 0.01(maximum=0.03). Finally, we provide a normalized surface reflectance database consisted of 36 images for 2001 over Korea.

18

항만효율성 측정 자료의 정규성과 변환 불변성 검증소고;DEA접근

박노경

[Kisti 연계] 한국항만경제학회 한국항만경제학회 학술대회논문집 2007 pp.391-405

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 항만효율성 측정 시 문제가 되었던 두 가지 문제점(첫째, 각기 상이한 기본단위를 갖는 투입변수와 산출변수의 정규화문제, 둘째, DEA분석의 기본가정인 비음수조건에 벗어난 자료, 즉, 음수를 갖는 투입-산출자료의 변환불변성)를 해결하기 위해서 국내 26개항만의 자료를 이용하여 실증분석을 한 후에 검증을 함으로써 항만효율성 측정방법을 부분적으로 확장시켰다. 본 논문의 실증분석의 핵심적인 결과는 다음과 같다. 첫째, 항만효율성 측정 시 사용되는 자료의 정규성과 변환불변성은 실증분석 결과 분명하게 있는 것으로 검증되었다. 둘째, 항만효율성 측정 시 사용되는 자료가 마이너스(-)인 경우에 가장 큰 음수보다 더 큰 양수를 더해 주는 이른바 자료의 변환를 검증하는 변환불변성은 투입지향-산출지향 BCC 모형에서 확인되었다. 위와 같은 실증분석 결과는 다음과 같은 정책적인 함의를 갖고 있다. 즉, 효율성 측정 시 사용되는 자료의 정규성과 변환불변성이 실증적으로 검증되었으므로, 국내 항만의 정책입안가들은 항만효율성 측정 시 이용되는 자료의 정규성과 변환불변성과 같은 사항을 고려하여 보다 세부적인 항만통계자료를 수집 ${\cdot}$ 정리 ${\cdot}$ 공표하는 것이 매우 필요하다. 예를 들면 항만사고와 같은 통계도 해역별이 아닌 항만별로 세부적으로 통계를 발행하도록 관련된 정책적인 지원이 필요하다.

The purpose of this paper is to verify the two problems(normalization for the different inputs and outputs data, and translation invariant for the negative data) which will be occurred in measuring the seaport DEA(data envelopment analysis) efficiency. The main result is as follow: Normalization and translation invariant in the BCC model for measuring the seaport efficiency by using 26 Korean seaport data in 1995 with two inputs(berthing capacity, cargo handling capacity) and three outputs(import cargo throughput, export cargo throughput, number of ship calls) was verified. The main policy implication of this paper is that the port management authority should collect the more specific data and publish these data on the inputs and outputs in the seaports with consideration of negative(ex. accident numbers in each seaport) and positive value for analyzing the efficiency by the scholars, because normalization and translation invariant in the data was verified.

19

데이터베이스 정규화 이론을 이용한 국민건강영양조사 중 다년도 식이조사 자료 정제 및 통합

권남지, 서지혜, 이헌주

[Kisti 연계] 한국환경보건학회 한국환경보건학회지 Vol.43 No.4 2017 pp.298-306

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Objectives: Since 1998, the Korea National Health and Nutrition Examination Survey (KNHANES) has been conducted in order to investigate the health and nutritional status of Koreans. The food intake data of individuals in the KNHANES has also been utilized as source dataset for risk assessment of chemicals via food. To improve the reliability of intake estimation and prevent missing data for less-responded foods, the structure of integrated long-standing datasets is significant. However, it is difficult to merge multi-year survey datasets due to ineffective cleaning processes for handling extensive numbers of codes for each food item along with changes in dietary habits over time. Therefore, this study aims at 1) cleaning the process of abnormal data 2) generation of integrated long-standing raw data, and 3) contributing to the production of consistent dietary exposure factors. Methods: Codebooks, the guideline book, and raw intake data from KNHANES V and VI were used for analysis. The violation of the primary key constraint and the $1^{st}-3rd$ normal form in relational database theory were tested for the codebook and the structure of the raw data, respectively. Afterwards, the cleaning process was executed for the raw data by using these integrated codes. Results: Duplication of key records and abnormality in table structures were observed. However, after adjusting according to the suggested method above, the codes were corrected and integrated codes were newly created. Finally, we were able to clean the raw data provided by respondents to the KNHANES survey. Conclusion: The results of this study will contribute to the integration of the multi-year datasets and help improve the data production system by clarifying, testing, and verifying the primary key, integrity of the code, and primitive data structure according to the database normalization theory in the national health data.

 
페이지 저장