Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 81
No
1

OLAP(On-Line Analytical Processing)은 데이터 큐브 또는 큐브라고 불리는 다차원 데이터 구조를 이용하여 복잡한 질의를 고속으로 처리하는 데이터 분석 기술이다. 전통적 방식의 OLAP은 디스크 기반 DBMS 환경으로 데 이터를 선 저장한 후 사용자의 질의에 응답하는 일회성 질의(One-Time Query) 수행 방식이었다. 하지만 지속적 으로 방대한 양의 데이터가 생성되는 데이터 스트림 환경에서 기존 처리 방식은 질의를 반복적으로 수행해야 하기 때문에 우수한 성능을 기대하기 어려우며, 동시적으로 다차원 계층 데이터에 질의를 수행하는데 한계가 존재한다. 본 연구에서는 이러한 문제점을 극복하기 위해 연속질의 기반 다차원 계층 큐브 처리 기법을 제안한다. 본 연구 모 델은 계층 데이터를 처리하는 하이퍼 데이터 큐브를 구축한다. 각 큐보이드들은 이전 집계된 데이터 큐보이드 중 가 장 작은 계산 비용 큐보이드를 집계하는 최소 비용 트리를 형성하여 성능적 향상을 기대한다. 본 연구 모델의 성능 을 검증하기 위해서 다양한 실험을 진행하였다.

OLAP(On-Line Analytical Processing) is one of the data analysis techniques that processes a complex query in a fast time using multi-dimensional data structure called the‘data cube’ or simply‘cube’. However, conventional OLAP system is not applicable to data streams because of these reasons: low performance; limitation of execution of a number of queries simultaneously. This paper proposes continuous query based evaluation of multi-dimensional hierarchical data cube. Minimal cost cube tree that computing a cuboid from the smallest cost, previously computed cuboid is constructed. Finally, the proposed method is verified by a series of experiments.

2

효과적인 데이터 공유를 위한 계층적 구조를 갖는 사이버 보안 데이터 공유시스템 모델 연구 KCI 등재

유호제, 김찬희, 조예림, 임성식, 오수현

한국융합보안학회 융합보안논문지 제22권 제1호 2022.03 pp.39-54

※ 기관로그인 시 무료 이용이 가능합니다.

4,900원

최근 지능화·고도화되는 사이버 위협에 효과적으로 대응하기 위해 다양한 사이버 보안 데이터의 수집, 분석, 실시간 공유의 중요성이 대두되고 있다. 이러한 상황에 대응하기 위해 국내에서는 사이버 보안 데이터 공유시스템의 확대를 위해 노력하고 있지만, 많은 민간 기업들은 사이버 보안 데이터를 수집하기 위한 예산과 전문 인력의 부족으로 인해 사이버 보안 데이터 공 유시스템에 참여가 어려운 상황이다. 이러한 문제를 해결하기 위해 본 논문에서는 현존하는 국·내외 사이버 보안 데이터 공유 시스템의 연구·개발 동향을 분석하고 이를 기반으로 조직 규모를 고려하여 계층적 구조를 갖는 사이버 보안 데이터 공유시스 템 모델과 해당 모델에 적용할 수 있는 단계별 보안정책을 제안한다. 본 논문에서 제안하는 모델을 적용할 때 다양한 민간 기 업들이 사이버 보안 데이터 공유시스템에 참여하는 것을 확대할 수 있으며, 지능화되고 있는 보안 위협에 신속하게 대처할 수 있는 대응체계 마련에 활용할 수 있을 것으로 기대한다.

Recently, the importance of collecting, analyzing, and real-time sharing of various cybersecurity data has emerged in order to effectively respond to intelligent and advanced cyber threats. To cope with this situation, Korea is making efforts to expand its cybersecurity data sharing system, but many private companies are unable to participate in the cybersecurity data sharing system due to a lack of budget and professionals to collect cybersecurity data. In order to solve such problems, this paper analyzes the research and development trends of existing domestic and foreign cyber security data sharing systems, and based on that, propose a cybersecurity data sharing system model with a hierarchical structure that considers the size of the organization and a step-by-step security policy that can be applied to the model. In the case of applying the model proposed in this paper, it is expected that various private companies can expand their participation in cybersecurity data sharing systems and use them to prepare a response system to respond quickly to intelligent security threats.

3

이중 해쉬체인을 이용한 계층적 다중 처리를 위한 효율적인 데이터 관리 기법 KCI 등재

정윤수, 김용태, 박길철

한국디지털정책학회 디지털융복합연구 제13권 제10호 2015.10 pp.271-278

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

현재 인터넷을 통해 수집되는 빅 데이터는 데이터의 종류와 크기에 따라 데이터가 수집되는 시간보다 데이터가 증가하는 속도가 높아 사용자가 원하는 데이터를 원활하게 수집하는 것이 어려운 상황이다. 특히, 데이터의 사용 목적 및 종류에 따라 다르게 처리되기 때문에 데이터의 정확성과 계산비용이 빅 데이터 관리에 중요한 항목 중 하나이다. 본 논문에서는 인터넷에 존재하는 수많은 서로 다른 종류의 데이터를 사용자가 원할 때, 데이터를 정확하게 추출하는 동시에 데이터의 계산비용을 최소화하기 위해서 이중 해쉬체인을 이용한 계층적 다중처리 기반의 데이터 처리기법을 제안한다. 제안 기법은 다양한 종류의 데이터를 추출하기 위해서 데이터를 사용 목적 및 방법에 따라 계층적으로 분류한다. 이때, 데이터의 정확도를 높이기 위해서 데이터를 이중 해쉬체인으로 묶어 다중 처리한다. 또한, 제안 기법은 계층적으로 분류된 데이터를 손쉽게 접근하기 위해서 해쉬체인으로 데이터를 구성하여 데이터의 처리 비용을 줄였다. 실험결과, 제안 기법은 기존 기법보다 데이터의 정확도는 평균 7.8% 높았고, 데이터의 처리 비용은 4.9% 단축시켰다.

Recently, bit data is difficult to easily collect the desired data because big data is collected via the Internet. Big data is higher than the rate at which the data type and the period of time for which data is collected depending on the size of data increases. In particular, since the data of all different by the intended use and the type of data processing accuracy and computational cost is one of the important items. In this paper, we propose data processing method using a dual-chain in a manner to minimize the computational cost of the data when data is correctly extracted at the same time a multi-layered process through the desired number of the user and different kinds of data on the Internet. The proposed scheme is classified into a hierarchical data in accordance with the intended use and method to extract various kinds of data. At this time, multi-processing and tie the data hash with the double chain to enhance the accuracy of the reading. In addition, the proposed method is to organize the data in the hash chain for easy access to the hierarchically classified data and reduced the cost of processing the data. Experimental results, the proposed method is the accuracy of the data on average 7.8% higher than conventional techniques, processing costs were reduced by 4.9% of the data.

4

최근 이미지 분류 작업의 기술들은 잘 정제되어 있는 대규모 데이터 세트에서 강력한 성능을 보여주지만, 긴꼬리 분 포에서는 약한 모습을 보인다. 또한, 계층적 데이터 세트에 대한 연구는 제한적이다. 그러나 트리 형태의 구조로 조 직된 계층적 데이터 세트에서는 상위 클래스 정보를 추가적으로 활용할 수 있다. 본 논문에서는 계층적 특징 활용을 위해 문화유산 데이터 세트를 수집하였으며, 해당 데이터 세트에서 분류 실패의 두 가지 원인을 발견하였다. 1) 클 래스 불균형으로 인한 오 분류, 2) 상위 클래스 오 분류. 우리는 이러한 문제를 해결하기 위해 두 가지의 손실 함수 를 활용한 학습 전략을 제안한다. 유니폼 샘플링과 클래스별 유니폼 샘플링, 두 가지 샘플링 방법의 이점을 모두 가 져가기 위한 consistency 손실 함수와 상위 클래스의 특성을 위해 하위 클래스 예측에 기반한 손실 함수를 사용하 여 문화유산 데이터 세트에서 베이스라인 대비 5.4%의 성능 향상을 달성한다.

Recent techniques in image classification show strong performance on large, well-refined datasets, but are weak on long-tailed distributions. In addition, there is limited research on hierarchical datasets; however, hierarchical datasets organized in a tree-like structure can further exploit higher-level class information. In this paper, we collected a cultural heritage dataset to utilize hierarchical features, and found two causes of classification failure in the dataset. 1) misclassification due to class imbalance, and 2) superclass misclassification. To solve these problems, we propose a learning strategy utilizing two loss functions. We use a consistency loss function to take advantage of both sampling methods, uniform sampling and class-specific uniform sampling, and a loss function based on subclass prediction for the characteristics of the parent class. We achieve a performance improvement of 5.4% over baseline on a cultural heritage dataset using our proposed learning technique.

5

실세계의 산업 및 공공 데이터는 정제된 공개 데이터셋과 달리 계층 구조가 불규칙하거나 특정 단계의 정보가 생략되는 등 비정형적인 특성을 지닌다. 이러한 데이터의 불완전성은 결과적으로 분류 모델이 계층적 일관성을 유지하며 정교한 추론을 수행하는 데 큰 제약이 된다. 본 연구에서는 이러한 한계를 극복하기 위해 실세계의 복잡하고 불완전한 계층 구조를 효과적으로 학습할 수 있는 계층적 지식 증류 알고리즘을 제안한다. 제안하는 시스템은 특히 레이블 누락 문제를 해결하기 위해 가용 레이블에 대해서만 손실 함수를 계산하는 레이블 마스킹 전략과 상위 계층의 지식을 바탕으로 누락된 정보를 보완하는 의사 레이블링 기법을 도입하였다. 또한, 하위 클래스의 확률이 부모 클래스의 확률을 초과하지 않도록 강제하는 계층 일관성 제약 조건을 적용하여 구조적 모순을 차단하였다. 제안된 방법은 기존의 계층 분류 알고리즘 대비 벤치마크 데이터셋에서 우수한 성적을 거뒀다. 특히, iNaturalist2018에서는 기존 대비 Accuracy가 약 7.2% 향상된 성능을 달성했으며, 이는 본 연구에서 제안하는 방법론의 범용성과 구조적 일관성 유지 능력을 실증적으로 확인하였음을 시사한다.

Unlike refined public datasets, real-world industrial and public data often exhibit irregular hierarchical structures and missing information at specific levels. This data incompleteness acts as a significant constraint, preventing classification models from maintaining hierarchical consistency and performing sophisticated reasoning. To overcome these limitations, this study proposes a Hierarchical Knowledge Distillation algorithm designed to effectively learn from complex and incomplete real-world hierarchies. To specifically address the problem of missing labels, the proposed system introduces a label masking strategy that calculates loss functions only for available labels, alongside a pseudo-labeling technique that compensates for missing information based on knowledge from higher hierarchies. Furthermore, hierarchical consistency constraints are applied to prevent structural contradictions by ensuring that the probability of a sub-class does not exceed that of its parent class. The proposed method outperformed existing hierarchical classification algorithms on standard benchmark datasets. Notably, it achieved an approximately 7.2% improvement in accuracy on the iNaturalist 2018 dataset, empirically validating the proposed methodology's versatility and its ability to maintain structural consistency.

6

최근 보안 업계에서 새롭게 화두가 되고 있는 XDR(eXtended Detection and Response)은 엔드포인트 와 네트 워크 영역의 모든 정보에 대한 통합 연계 분석을 통하여, 기업 내 보안 목표를 보다 효과적 으로 달성할 수 있는 솔 루션이다. 그러나, XDR에서 추구하는 통합 연계 분석을 위해서는 기업내 에서 운영되는 이기종 보안 솔루션들 간 의 서로 다른 프로토콜과 정보 저장 방식의 문제를 해결해야 한다. 서로 다른 형태의 데이터를 통합하고 이를 분석 하기 위한 얼라이언스와 통합 레이블링 등의 표준화를 제안한 연구들이 있지만, 연동하는 모든 보안솔루션이 표준 API를 기반으로 개발되어야 하는 현실적인 어려움이 존재한다. 이 문제를 해결하기 위해 본 논문에서는 다양한 보 안솔루션 데이터 중에서 가장 기본적이고 필수적인 데이터인 비정형 로그를 중심으로 계층적 구조 기반의 데이터 정 규화 방법을 제안한다. 공통적인 CODE와 Corpus로 구성한 정규화 모델 알고리즘으로 정제된 비정형로그 데이터 는 계층적 구조를 가지는 정규화된 데이터 형태로 출력된다.

XDR(eXtended Detection and Response) is a new topic in the security industry. It is a solution to achieve security goals within the enterprise more effectively through integrated linkage analysis of all information in the endpoint and network area. However, for the integrated linkage analysis of XDR, it is necessary to solve the problem of different protocols and information storage methods between heterogeneous security solutions operated within the enterprise. Standardization approaches such as alliance and unified labeling to integrate and analyze different types of data have been proposed. However, those approaches suffer from some practical difficulties that all interworking security solutions are developed based on standard APIs. To overcome the difficulties, in this paper, we propose a new method to normalize the hierarchical structure-based data of the unstructured log, which is the most basic and essential data in the security solution data. The unstructured log data refined through the normalization model algorithm composed of common CODE and Corpus is output in the form of normalized data having a hierarchical structure.

7

본 연구에서는 비정형 벡터 타일 데이터의 실시간 GPU 시각화를 위해, 동적 메모리 관리와 공간 계층 구조, 그리고 적응적 LOD 제어를 통합한 시스템을 제안하였다. 기존의 정적 버퍼 기반 그래픽스 파이프라인은 데이터의 비정형 성과 실시간 갱신 특성을 처리하기 어렵다는 한계를 가진다. 이를 해결하기 위해 GPU 내부에서 페이지 단위로 메 모리를 동적으로 관리하고, 세대 식별자를 이용한 안정적인 메모리 회수 구조를 설계하였다. 또한 도로·건물·지형과 같은 레이어별 특성을 고려한 다중 루트 기반 공간 계층 구조를 구성하여, 스트리밍 데이터의 삽입과 제거를 효율적 으로 처리할 수 있도록 하였다. 실험 결과, 제안된 시스템은 CPU 기반 파이프라인에 비해 평균 프레임 시간이 약 9~12% 감소하였고, 메모리 파편화율은 약 30% 낮게 유지되었다. 또한 프레임 최장 렌더링 시간(p95) 개선으로 인한 안정성이 10~15% 향상되어 대규모 스트리밍 환경에서도 실시간성과 예측 가능성을 확보하였다. 본 연구는 GPU 내부에서 데이터의 생성·관리·LOD 처리를 일관된 구조로 수행함으로써, 향후 디지털 트윈·실시간 지도 서비 스·자율주행 시각화 등 대규모 공간 데이터 처리 분야에 활용 가능한 기반 기술을 제시한다.

This paper presents a unified GPU-based system that integrates dynamic memory management, hierarchical spatial structuring, and adaptive level-of-detail (LOD) control for real-time visualization of unstructured vector tile data. Conventional static-buffer rendering pipelines suffer from inefficiencies when handling irregular and continuously streamed spatial data. To address this limitation, the proposed framework introduces a page-based dynamic memory management scheme within the GPU, utilizing generation identifiers for safe memory reclamation. A multi-root spatial hierarchy is constructed to manage heterogeneous layers—such as roads, buildings, and terrains— allowing efficient insertion and removal of streaming data without CPU intervention. Experimental results show that, compared with a CPU-driven baseline, the proposed system reduces average frame time by approximately 9–12%, lowers memory fragmentation by about 30%, and improves p95 frame-time stability by 10–15%. These improvements demonstrate that the proposed architecture enables predictable and stable real-time rendering even under large-scale streaming workloads. The system provides a structural foundation for GPU-resident rendering engines applicable to digital twins, real-time mapping, and autonomous driving visualization.

8

4,000원

주가지표처럼 동적이며 시간흐름을 따르는 시계열자료들을 이해하는 효과적인 방법은 주어진 시계열자료들 에 대하여 모델을 결정함으로서 이해하는 것이 좋다. 주어진 자료들에 대한 모델 결정과정은 수집되어진 대용량 시계 열자료 전체를 한 번에 다 살펴보는 것보다 자료를 특정의 중요한 몇 개의 하위그룹으로 군집화하여 각 군집별 모델 결정을 통해 자료 전체를 이해하는 것이 효율적이다. 본 연구에서는 주어진 시계열자료들에 대하여 하위그룹으로의 효율적 군집화 과정 그리고 각 군집별 모델결정의 두 과정 중 첫 번째 과정인 하위집단으로 군집화 과정에 자료의 구간특징화 기법과 휴리스틱 베이지안기법의 융합을 이용하여 시간 및 계산비용을 감소시킬 수 있는 기법을 제안하 였으며 실제적인 주가지표를 이용한 실험을 통해 제안하는 기법의 유효성을 확인하였다.

An effective way to understand the dynamic and time series that follows the passage of time, as valuation is to establish a model to analyze the phenomena of the system. Model of the decision process is efficient clustering information of the total mass of the time series data of the relevant population been collected in a particular number of sub-groups than to look at all a time to an understand of the overall data through each community-specific model determination. In this study, a sub-grouping of the group and the first of the two process model of each cluster by determining, in the following in sub-population characterized by a fusion with heuristic Bayesian clustering techniques proposed a process which can reduce calculation time and cost was confirmed by experiments using actual effectiveness valuation.

9

glmmLasso 기계학습 기법을 통한 일반고와 특성화고 고등학생의 학교소속감 예측 변수 탐색

유진은, 구미령

한국교원대학교 교육연구원 교원교육 제37권 제2호 2021.04 pp.467-486

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

학교소속감이 중도탈락을 막을 수 있는 중요한 변수임에도 불구하고 특히 고등학교 학생에 대 한 학교소속감 연구는 그 수가 많지 않다. 아직 충분한 연구가 이루어지지 않은 분야에서 기계학습 기법을 활용하는 탐색적 연구가 학문적 기여를 할 수 있다. 본 연구의 주요 목적은 일반고 및 특성 화고 학생들의 학교소속감을 예측하는 중요한 변수를 파악하는 것이다. 이를 위하여 부산교육종단 연구 2016 4차년도 데이터의 일반고와 특성화고 학생, 교사, 교장, 학교 변수를 모두 활용하였다. 구체적으로 75개 일반고 1,775명의 824개 변수 자료와 36개 특성화고 739명의 854개 변수 자료를 기계학습 기법으로 분석한 결과, 일반고와 특성화고에서 각각 20개와 21개의 학생, 교사, 학교 관 련 변수가 선택되었다. 학교소속감을 개인의 심리적 차원에 초점을 맞추어 분석한 선행연구와 달 리, 본 연구는 교사, 교장, 학교 변수까지 모두 모형에 투입함으로써 학교 현장에서의 변화를 꾀하 였다. 기계학습 기법 중 벌점회귀모형으로 분류되는 glmmLasso를 활용하여 변수 선택 시 자료의 위계적 구조를 반영한 점 또한 연구 의의라 하겠다. 특히 특성화고 자료는 사례 수보다 변수 수가 더 많은 고차원 자료였으므로 기계학습 기법을 활용하는 것이 필수적이었다. 연구 결과를 토대로 고등학생의 학교소속감을 향상시키기 위하여 필요한 정책적 제언을 제시하고, 후속 연구주제 또한 논하였다.

Although students’ sense of belonging to school is an important indicator of at-risk students including early dropouts, there has not been sufficient research on high school students’ sense of belonging. While confirmatory research requires solid theoretical background, exploratory research particularly via machine learning is unbounded by established theory and thus can make contributions to the existing body of literature. This study aimed at exploring and identifying important variables to predict students’ sense of belonging enrolled in general and specialized vocational high schools. BELS (Busan Educational Longitudinal Study) 2016 fourth wave data (824 variables of 1,775 general high school students and 854 of 739 specialized vocational high school students) were analyzed with glmmLasso, a machine learning technique. In particular, vocational high school data were high-dimensional and thus machine learning was a necessary tool. Specifically, glmmLasso, grouped as penalized regression among machine learning, is known to consider the hierarchical data structure resulting from multistage sampling schemes in variable selection. Among student, teacher, principal, and school variables explored, a total of 20 and 21 variables were selected as important for general and vocational high schools, respectively. Policy suggestions and future research topics were discussed based on the results of the study.

10

Job Scheduling and Data Replication in Hierarchical Data Grid

Somayeh Abdi, Sayyed Mohsen Hashemi

보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology Vol.47 2012.10 pp.91-100

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Data Grid environment is a geographically distributed that deal with date-intensive application in scientific and enterprise computing. In data-intensive applications data transfer is a primary cause of job execution delay. Data access time depends on bandwidth, especially when hierarchy of bandwidth appears in network. Effective job scheduling can reduce data transfer time by considering hierarchy of bandwidth and also dispatching a job to where the needed data are present. Additionally, replication of data from primary repositories to other locations can be an important optimization step to reduce the frequency of remote data access. Objective of dynamic replica strategies is reducing file access time which leads to reducing job runtime. In this paper we develop a job scheduling policy, called SS (Scheduling strategy), and a dynamic data replication strategy, called DRS (Distributed Replication Strategy), to improve the data access efficiencies in a cluster grid. We study our approach and evaluate it through simulation. The results show that combination of SS and DRS has improved 17% over other combinations.

11

Hierarchical Data Gathering Scheme for Energy Efficient Wireless Sensor Network

Jinping Chen, Zinan Chang

보안공학연구지원센터(IJFGCN) International Journal of Future Generation Communication and Networking Vol.9 No.3 2016.03 pp.189-200

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Compressive Data Gathering (CDG) has recently emerged to provide energy efficient data aggregation in wireless sensor networks (WSNs). Existing CDG is done by forming a span-tree from all sensors to the data sink, which leads to a lot of energy consumption in maintain this tree. In this paper, we present a clustering based CDG scheme for large-scale WSNs. This scheme reduces traditional network communication cost by introducing in-network compression and recovers sensory data accurately by using compress sensing theory. With extensive simulation, we demonstrate that our scheme prolongs nearly 45% network’s lifetime compared with the Hybrid compressive data gathering schemes.

12

Hierarchical Multi-sensors Data Fusion for Enhanced Context Inference SCOPUS

Jeongbong You

보안공학연구지원센터(IJCA) International Journal of Control and Automation Vol.7 No.2 2014.02 pp.189-196

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Dempster-Shafer Evidence Theory can infer the context only with some native sensed signals but its number of assumptions is limited, compared to the number of evidence. In this research, we proposed a method that can increase the number of sensors collecting symptoms for more dynamic contextual inference yet with reduced calculation load increase. This is possible by clustering sensors with relevance to a situation and fusing data hierarchically.

13

Low Level Segmentation of Motion Capture Data based on Hierarchical Clustering with Cosine Distance SCOPUS

Yang Yang, Jinfu Chen, Zhanzhan Liu, Yongzhao Zhan, Xinyu Wang

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.4 2015.08 pp.231-240

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

3D motion capture is to track and record human movements. In recent years, it has been applied into many fields, such as human computer interaction, animation, etc. Low-level segmentation of motion capture data is of significance to the various applications of 3D motion capture; however, due to the high dimensionality of motion capture data, traditional low-level segmentation methods can hardly work out a suitable segmentation for motion capture data. In order to solve this problem, a low-level temporal segmentation algorithm based on cosine distance is proposed, hierarchical clustering is explored so that similar velocity vectors are clustered together according to the cosine distance in a progressive way, the center of each cluster is updated as the vector derived with linear regression, the segment boundaries are determined as the point when the cosine distance between adjacent velocity vectors is greater than 1 (angle>90 degrees). We have conducted experiments on the motion capture database provided by Carnegie Mellon University (CMU), the experiment results show that the performance of the proposed method is optimistic.

14

Detecting anomalies in human movement is an important task in industrial applications, such as monitoring industrial disasters or accidents and recognizing unauthorized factory intruders. In this paper, we propose an abnormal worker movement detection system based on data stream processing and hierarchical clustering. In the proposed system, Apache Spark is used for streaming the location data of people. A hierarchical clusteringbased anomalous trajectory detection algorithm is designed for detecting anomalies in human movement. The algorithm is integrated into Apache Spark for detecting anomalies from location data. Specifically, the location information is streamed to Apache Spark using the message queuing telemetry transport protocol. Then, Apache Spark processes and stores location data in a data frame. When there is a request from a client, the processed data in the data frame is taken and put into the proposed algorithm for detecting anomalies. A real mobility trace of people is used to evaluate the proposed system. The obtained results show that the system has high performance and can be used for a wide range of industrial applications.

15

Data Compression Algorithm based on Hierarchical Cluster Model for Sensor Networks

Wang Lei, Wang Tongsen, Yang Ronghua

보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology vol.2 2009.01 pp.71-84

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

A new distributed algorithm of data ccompression based on hierarchical cluster model for sensor networks s proposed, the basic ideas of which are as follows, firstly the whole sensor network is mapped into a kind f hierarchical clusters model, and then different wavelet transform models are used to commit data ompression in inner and super clusters respectively, according to the relative regularity of sensor nodes eployed in the inner clusters, and the relative irregularity of sensor nodes deployed in super cluster. heoretical analyses and simulation results show that, the above new methods have good performance of pproximation, and can compress data and reduce the amount of data efficiently. So, it can prolong the ifetime of the whole sensor network to a greater degree.

16

Data Clustering and Analyzing Techniques Using Hierarchical Clustering Method

Wen Hu, Qing he Pan

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.6 No.2 2013.04 pp.109-118

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Data clustering and analyzing techniques are studied by using hierarchical clustering method. A matrix of words is constructed with a randomly chosen RSS list. By collecting data from this list a matrix is built. In the matrix each row corresponds to a article and each column represents a word. Based on the matrix a hierarchical clustering algorithm is designed. In this algorithm the Pearson correlation coefficient is used to compute the distances among different contents. The dendrogram is used to describe the hierarchical relationship of contents and words. And the 2-D graph also is used to represent the dendrogram in another format.

17

A Hierarchical Dynamic Load Balancing Strategy for Distributed Data Mining

Raja Tlili, Yahya Slimani

보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology Vol.39 2012.02 pp.29-48

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Extracting useful knowledge from data sets measuring in gigabytes and even terabytes is a challenging research area for the data mining community. Sequential approaches suffer from a performance problem due to the fact that they have to mine voluminous databases. Parallelism is introduced as an important solution that could improve the response time and the scalability of these approaches. However, parallelization process is not trivial and still facing many challenges including the workload balancing problem. In this paper, we propose a hierarchical dynamic load balancing strategy for parallel association rule mining algorithms in the context of a Grid computing environment. The French research grid “Grid’5000” is used as our experimental test-bed. Through a detailed experimental study, we show that our strategy improves the performance and helps the parallel algorithm to scale very well with the number of computational nodes available.

18

Optimization of the Data Traffic Distribution based on Hierarchical Mobile Networks SCOPUS

Ronnie D. Caytiles

보안공학연구지원센터(IJCA) International Journal of Control and Automation Vol.8 No.11 2015.11 pp.415-420

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Data traffic distribution across heterogeneous mobile networks requires a control management in order to optimize the handoffs between its entities. The more the technologies are evolving, the more complex the distribution of data traffic becomes. The standard MIPv6 and hierarchical MIPv6 mobility management were introduced in order to address the flow mobility of data traffic. This paper exploits the multi-connectivity option for mobile nodes in HMIPv6 in order to utilize the optimum route for the data traffic distribution.

19

구조적 지역특성이 성범죄에 미치는 영향 : 전국 읍면동과 시군구를 대상으로 한 위계적선형모형 분석 KCI 등재

정진성, 홍성욱, 이가을

한국치안행정학회 한국치안행정논집 제12권 제1호 2015.05 pp.119-144

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

본 연구는 구조적 지역특성이 성범죄에 어떤 영향을 미치는 지 살펴보기 위해 읍면동 단위와 시군구 단위로 구성된 위계적선형모형(HLM)을 구축·분석했다. 국내 최초로 전국의 3,468개 읍면동과 251개 시군구를 대상으로 최근 3년간(2011-2013) 발생한 성범죄율에 대한 분석을 수행한 결과, 성범죄율의 분산 가운데 77.97%는 시군구 내 분산이, 22.04%는 시군구 간 분산이 차지하였다. 읍면동 단위에서는 인구이동비율, 외국인비율, 이혼율, 1인가구비율, 숙박음식업비율이 정적 효과를 보였고 비아파트거주비율이 부적 효과를 보였다. 시군구 단위에서는 숙박음식업비율과 비아파트거주비율이 부적 효과를 보였으며, 읍면동 단위 변수들과의 층위간 상호작용효과는 인구밀도와 인구이동비율, 1인가구비율, 비아파트거주비율 간에 발견되었다. 비아파트거주비율은 가설과 달리 읍면동과 시군구 단위에서 모두 성범죄를 감소시켰는데, 층위간 상호작용효과를 고려하면 이러한 현상은 인구밀도가 낮은 시골 지역에는 해당되지 않는 것으로 판단되었다. 숙박음식업비율은 읍면동 단위에서는 정적 효과를 보인 반면, 시군구 단위에서는 부적 효과를 보였는바, 숙박음식업비율의 의미가 읍면동 수준에서는 유흥과 향락이 발달한 도시적 특성을 반영하지만 시군구 수준에서는 관광과 생업으로서의 시골적 특성을 반영하는 것으로 사료되었다. 인구밀도가 인구이동비율, 1인가구비율과 보인 층위간 상호작용효과를 비롯하여 본 연구의 전반적인 결과는 보다 세분화된 지역 특성(예, 도시 vs. 시골, 사회경제적 지위)에 따른 추가적인 연구와 정책개발이 필요함을 시사하였다. 경제적 열악성과 같은 주요 변수가 모형에서 제외된 점 등의 한계가 있었지만, 전국의 모든 읍면동과 시군구를 대상으로 연구를 진행한 점, 지역 수준의 변수들이 성범죄에 미치는 영향력의 범위를 입증한 점, 보다 세부적인 지역사회 연구의 필요성을 제시한 점 등에서 본 연구의 의의를 찾을 수 있었다.

This study attempted to examine the effect of structural neighborhood characteristics on sex crime. To this end, a hierarchical linear model was constructed for sex crime of 2011-2013 using national data of 3,468 Eup-Myon-Dong(level 1) and 251 Si-Gun-Gu(level 2). The results showed that the total variance of sex crime rates was divided into within variance of Si-Gun-Gu, 77.97% and between variance of it, 22.04%. At the level of Eup-Myon-Dong, rate of population mobility, rate of foreigners, divorce rate, rate of single family, and rate of food and lodging had a positive effect on sex crime rates. At the Si-Gun-Gu level, rate of food and lodging and rate of non-apartment residency showed a negative effect. Also, cross-level interaction effects were discovered in population density with rate of population mobility, rate of single family, and rate of non-apartment residency. Contrary to the hypothesis, the rate of non-apartment residency lowered sex crime rates at both levels. However, its cross-level interaction effect with population density showed that the negative effect might be limited to rural areas with low population density. The different direction of food and lodging’s influences indicated that it reflects urban feature of adult entertainment and hedonism at level 1, while reflecting rural feature of sightseeing and means of living at level 2. The overall patten of results including the cross-level interaction effects suggested that more subspecialized research and policy development are necessary depending on the regional characteristics such as urban vs. rural and socioeconomic status. Despite some limitations such as the exclusion of important variables like economic disadvantage, this study made important contributions to current neighborhood crime research at least in three aspects: First, it made the first attempt to reveal the causal relationship between structural characteristics and sex crime using nation-wide data of all Eup-Myon-Dong and Si-Gun-Gu. Second, it discovered the extent of influences that each level’s characteristics have on sex crime. And lastly, it suggested that future research have to focus more on subspecialized neighborhood characteristics.

20

운영 체제와 컴파일러에 따른 Geospatial Data Abstraction Library의 Hierarchical Data Format 형식 원격 탐사 자료 추출 속도 비교

유병현, 김광수, 이지혜

[Kisti 연계] 한국농림기상학회 한국농림기상학회지 Vol.21 No.1 2019 pp.65-73

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

지역이나 전구 규모의 농업 생태계를 감시하기 위해 HDF 형식으로 제공되는 MODIS 원격 탐사자료가 사용되어 왔다. 대개의 경우, 다량의 영상자료들이 처리되어야 하기 때문에, 이들 자료의 처리 성능을 향상시키는 것이 유리하다. 본 연구는 HDF 파일을 처리할 수 있는 GDAL과 같은 라이브러리가 운영 체제나 배포 방식 등에 따른 처리속도의 차이를 확인하여 원격 탐사 자료 처리 시스템 구축을 지원하고자 하였다. 이를 위해, GDAL이 시스템에 설치되는 주요 조건들에 따라 MODIS 영상자료 처리 시간을 측정하고 비교하였다. 운영 체제(Ubuntu 및 openSUSE), 컴파일러(GNU 및 Intel), 설치 옵션 및 바이너리 패키지 조건을 조합하여 GDAL성능 비교가 이루어졌다. 각 조건에 따라 설치된 GDAL을 사용하여 MODIS 영상 중 대기측정 자료(MOD07)의 2차원 변수와 3차원 변수에 해당하는 총 10 종의 자료를 추출하였다. 자료처리에 소요된 구동 시간은 각 변수 값을 시스템 메모리에 저장하는 작업이 끝난 직후 측정되었다. 가장 좋은 성능을 보인 설치 조건은 Ubuntu에서 Intel Compiler를 사용하여 컴파일 된 GDAL을 사용하는 것이었다. OpenSUSE에서는 GNU와 Intel 컴파일러가 각각 2차원 자료와 3차원 자료를 처리하기 위한 작업에 효과적인 것으로 나타났다. 한편 "--with-hdf4=no" 옵션으로 컴파일 된 GDAL과 RPM package manager 버전의 GDAL의 경우, 다른 조건에 비해 상당히 낮은 성능을 보였다. 이러한 결과는 운영 체제나 컴파일러, 설치 옵션 등을 조정하여 원격 탐사자료 처리 도구의 속도를 개선할 수 있다는 것을 암시하였다. 특히, 원격 탐사 자료의 경우 다양한 형식으로 배포되므로, 이를 처리하는 라이브러리들이 최고의 성능을 발휘할 수 있는 조건을 탐색하고 이러한 결과의 공유가 후속연구에서 진행되어야 할 것으로 보인다.

The MODIS (Moderate Resolution Imaging Spectroradiometer) data in Hierarchical Data Format (HDF) have been processed using the Geospatial Data Abstraction Library (GDAL). Because of a relatively large data size, it would be preferable to build and install the data analysis tool with greater computing performance, which would differ by operating system and the form of distribution, e.g., source code or binary package. The objective of this study was to examine the performance of the GDAL for processing the HDF files, which would guide construction of a computer system for remote sensing data analysis. The differences in execution time were compared between environments under which the GDAL was installed. The wall clock time was measured after extracting data for each variable in the MODIS data file using a tool built lining against GDAL under a combination of operating systems (Ubuntu and openSUSE), compilers (GNU and Intel), and distribution forms. The MOD07 product, which contains atmosphere data, were processed for eight 2-D variables and two 3-D variables. The GDAL compiled with Intel compiler under Ubuntu had the shortest computation time. For openSUSE, the GDAL compiled using GNU and intel compilers had greater performance for 2-D and 3-D variables, respectively. It was found that the wall clock time was considerably long for the GDAL complied with "--with-hdf4=no" configuration option or RPM package manager under openSUSE. These results indicated that the choice of the environments under which the GDAL is installed, e.g., operation system or compiler, would have a considerable impact on the performance of a system for processing remote sensing data. Application of parallel computing approaches would improve the performance of the data processing for the HDF files, which merits further evaluation of these computational methods.

 
1 2 3 4 5
페이지 저장