년 - 년
VANET에서 효율적인 분산적 데이터 복제본 할당 기법 KCI 등재
한국ITS학회 한국ITS학회논문지 제9권 제2호 통권28호 2010.04 pp.87-95
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
차량 간의 애드훅 네트워크 (VANET)은 모바일 애드훅 네트워크 (MANET)의 일종으로 차량 간의 무선 링크를 통하여 임시적인 정보 전송을 가능케 한다. VANET에서는 차량이 하나의 노드로 참여하면서 정보를 전송한다. 하지만 차량은 빠르게 이동하기 때문에 빈번하게 연결이 끊기게 된다. 이렇게 빈번하게 연결이 끊기기 때문에 차량의 데이터 접근성이 떨어지게 된다. 기존 연구에서는 모바일 애드훅 네트워크 환경에서 데이터 복제본 할당 기법을 통하여 노드의 데이터 접근성을 향상시키고자 하였다. 하지만 기존 연구에서 제안한 중앙 집중식 그룹 방법은 차량 간의 애드훅 네트워크에서는 빠르게 이동하는 차량 때문에 적합하지 않다. 본 연구에서는 데이터 복제본 할당을 분산적인 그룹 방법으로 할 수 있는 TBG (Tree Based Grouping)기법을 제안한다. 노드 자신 만의 TBG를 형성하여 연결의 안정성을 기반으로 데이터 복제본을 할당하여 데이터 접근성을 향상시킨다. 실험 평가를 통해 기존 기법에 비해 데이터 접근성이 크게 향상됨을 보여준다.
Vehicular Ad-Hoc Network (VANET) is form of the Mobile Ad-hoc Network (MANET) to provide temporary communication among vehicles via wireless links. In VANET, the vehicle is one of the nodes in networks and communicates with each other. However, the wireless links disconnect very frequently, because vehicles have mobility and move freely. The reason why data accessibility degrades is that disconnection occurs frequently. To improve data accessibility, data replica allocation methods that made group to allocate data replica have proposed in MANET. However, those are not suitable because it is difficult to maintain stable links among the nodes moving fast by centralized group. In this paper, we proposed TBG (Tree Based Grouping) to allocate data replica with the distributed grouping method. Each node has own TBG and allocates data replica based on stability of links to improve data accessibility. The experiment demonstrates that the proposed method outperforms traditional methods in term of data accessibility.
VANET환경에서의 효율적인 데이터 접근성 향상기법 KCI 등재
한국ITS학회 한국ITS학회논문지 제8권 제1호 통권21호 2009.03 pp.65-75
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
차량 애드 혹 네트웍 (VANET)은 모바일 애드혹 네트워크의 일종으로써 차량 간의 임시적인 통신을 가능케 한다. VANET의 모바일 노드는 네트웍의 일원으로 참여하면서 에너지와 자원을 사용한다. 몇몇의 노드는 다른 노드와 협력하는 대신에 자신의 이득만을 취하려고 하는 이기성을 가지고 있다. 이러한 이기성을 가진 노드들로 인해 데이터 접근성과 네트웍의 효율성이 감소하게 된다. 이 논문에서는 새로운 기법인 VANET 그룹에서 이기성을 갖고있는 노드를 제거하는 Friendship-VaR을 제안한다. Friendship-VaR은 이기성을 갖고있는 노드를 제거하고 믿을 수 있는 노드들간에 데이터 공유를 가능케 한다. Friendship-VaR는 간단한 데이터 교환을 통해 노드의 이기성 여부를 결정한다. 실험 결과는 제안한 기법이 기존의 기법에 비해 데이터 접근성이 좋아짐을 보여준다.
A Vehicular Ad-Hoc Network (VANET) is a form of Mobile ad-hoc network, to provide temporary communications among nearby vehicles. Mobile node of VANET consumes energy and resource with participating in the member of network. Some node tends to have a selfishness to place one's own profits above cooperation with others. As result of selfish node, it reduces data accessibility and the efficiency of networks. In this paper, we propose noble method, Friendship-VaR that excludes selfish nodes from a group of VANET. Friendship-VaR enables to improve data accessibility by eliminating selfish nodes and sharing data among reliable nodes. Friendship-VaR determines selfishness of nodes by simple data exchange. The experiments shows proposed method outperform existing method in terms of data accessibility.
사물인터넷 데이터를 위한 안정성-인지 클러스터드 복제 기법 KCI 등재후보
국제차세대융합기술학회 차세대융합기술학회논문지 제4권 5호 2020.10 pp.444-452
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문에서는 사물인터넷 센서 데이터베이스 기술을 분석하였고, 센서 노드 장애와 통신 장애에 대처하 기 위한 효율적인 데이터 관리 및 저장 기술에 대해 논의하였다. 센서 데이터베이스 환경에서 데이터 수집 및 모니 터링을 목적으로, 데이터 서비스 신뢰성 및 성능을 개선하기 위하여 새로운 안정성-인지 클러스터링 방식과 압축 쉐도우 복제 방식을 제안하였다. 또한, 데이터 미러링을 이용한 수정된 동기화 기술을 적용하여, 스토리지 효율성을 위한 압축된 쉐도우 데이터를 저장하였다. 쉐도우 데이터 복제 기술은 낮은 클러스터에 있는 안정성이 낮은 센서 노드의 배터리 및 통신 장애를 최소화한다. 결과적으로 불필요한 통신비용 및 배터리 소모를 줄일 수 있으며, 클러 스터 내 센서 노드를 동적으로 관리하여 전체적인 배터리 에너지 밸런스를 유지할 수 있다. 온도와 상대 습도 데이 터에 대하여 실험한 결과 센서 데이터의 크기가 14%이하로 감소되고 성능이 향상됨을 확인할 수 있었다.
This paper analyzed IoT sensor database technology and discussed efficient data management and storage technologies to cope with sensor nodes and communication failures. For data collection and monitoring in sensor database environment, a new stability-aware clustering scheme and compressed shadow replication scheme were proposed to improve the data service reliability and performance. In addition, a modified synchronizing skill using data mirroring was applied to store compressed shadow data for storage efficiency. Shadow data replication skill minimizes battery and communication failures of unreliable sensor nodes in the low clusters. Consequently, unnecessary communication costs and battery consumption can be reduced, and overall battery energy balance can be maintained by dynamically managing sensor nodes in the clusters. Experiments with temperature and relative humidity data have shown that the size of the sensor data is reduced to less than 14% and performance is improved.
해시 트리 기반의 대규모 데이터 서명 시스템 구현 KCI 등재
한국융합보안학회 융합보안논문지 제18권 제1호 2018.03 pp.19-31
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
ICT기술이 발전함에 따라 산업 전분야에 걸쳐 이전보다 훨씬 많은 디지털 데이터들이 생성, 이동, 보관, 활용되고 있다. 산출되는 데이터의 규모가 커지고 이를 활용하는 기술들이 발전함에 따라 대규모 데이터 기반의 신 서비스들이 등장하여 우리의 생활을 편리하게 하고 있으나 반대로 이들 데이터를 위변조 하거나 생성 시간을 변경하는 사이버 범죄 또한 증가하고 있다. 이에 대한 보안을 위해서는 데이터에 대한 무결성 및 시간 검증 기술이 필요한데 대표적인 것이 공개키 기반의 서명 기술이다. 그러나 공개키 기반의 서명 기술의 사용은 인증서와 키 관리 등에 필요한 부가적인 시스템 자원과 인프라 소요가 많아 대규모 데이터 환경에서는 적합하지 않다. 본 연구에서는 해시 함수와 머클 트리를 기반으로 시스템 자원의 소모가 적고, 동시에 대규모 데이터에 대해 서명을 할 수 있는 데이터 서명 기법을 소개하고, 서버 고장 등 장애 상황에서도 보다 안정적인 서비스가 가능하도록 개선한 해시 트리 분산 처리 방법을 제안하였다. 또한, 이 기술을 구현한 시스템을 개발하고 성능분석을 실시하였다. 본 기술은 클라우드, 빅데이터, IoT, 핀테크 등 대량의 데이터가 산출되는 분야에서 데이터 보안을 담보하는 효과적인 기술로써 크게 활용될 수 있다.
As the ICT technologies advance, the unprecedently large amount of digital data is created, transferred, stored, and utilized in every industry. With the data scale extension and the applying technologies advancement, the new services emerging from the use of large scale data make our living more convenient and useful. But the cybercrimes such as data forgery and/or change of data generation time are also increasing. For the data security against the cybercrimes, the technology for data integrity and the time verification are necessary. Today, public key based signature technology is the most commonly used. But a lot of costly system resources and the additional infra to manage the certificates and keys for using it make it impractical to use in the large-scale data environment. In this research, a new and far less system resources consuming signature technology for large scale data, based on the Hash Function and Merkle tree, is introduced. An improved method for processing the distributed hash trees is also suggested to mitigate the disruptions by server failures. The prototype system was implemented, and its performance was evaluated. The results show that the technology can be effectively used in a variety of areas like cloud computing, IoT, big data, fin-tech, etc., which produce a large-scale data.
A Game-theoretic Approach for Efficient Clustering in Wireless Sensor Networks
보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.7 No.1 2014.01 pp.67-80
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In wireless sensor networks, clustering divides network into clusters and makes cluster heads (CHs) responsible for data aggregation. CHs play a significant role in such topology and focus should be fixed on the CH selection. Due to the constraints on available resources, however, a sensor node is likely to be selfish and refuse to serve as a CH. Based on game theory, this paper models the problem and discusses the condition of Nash Equilibrium. Moreover, in case of disconnection between a CH and the sink, data replication is adopted. We set candidate CHs that replicate data in the original CHs respectively under the scenario of a second price sealed auction. Simulation results show that nodes have tendency to cooperate due to the reduction in delay and loss rate. Moreover the throughput of the sink can still be guaranteed if any CH fails to work.
Improving the Performance of Cluster Architecture Through Asynchronous Replication Technique
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application vol.4 no.2 2011.06 pp.77-86
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In Web Server Cluster (WSC), replication techniques are necessary when providing accesses to data at different locations. In WSC, when designing the cluster architecture, the elements of availability and reliability for cases of data sharing or backing-up data are given the top priority. This is to ensure that the response time and throughput is still within the acceptable range. In ensuring the values of response time and throughput have not surpassed the limit, thus the replication technique must be employed intelligently. The replication technique that is been proposed is based on asynchronous replication technique. In doing so, we have managed to improve the reliability of the web cluster system. We have also managed to control the response time to a very nominal value but at the same time able to provide utmost throughput.
Job Scheduling and Data Replication in Hierarchical Data Grid
보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology Vol.47 2012.10 pp.91-100
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Data Grid environment is a geographically distributed that deal with date-intensive application in scientific and enterprise computing. In data-intensive applications data transfer is a primary cause of job execution delay. Data access time depends on bandwidth, especially when hierarchy of bandwidth appears in network. Effective job scheduling can reduce data transfer time by considering hierarchy of bandwidth and also dispatching a job to where the needed data are present. Additionally, replication of data from primary repositories to other locations can be an important optimization step to reduce the frequency of remote data access. Objective of dynamic replica strategies is reducing file access time which leads to reducing job runtime. In this paper we develop a job scheduling policy, called SS (Scheduling strategy), and a dynamic data replication strategy, called DRS (Distributed Replication Strategy), to improve the data access efficiencies in a cluster grid. We study our approach and evaluate it through simulation. The results show that combination of SS and DRS has improved 17% over other combinations.
Efficient Data Replication Scheme based on Hadoop Distributed File System SCOPUS
보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.9 No.12 2015.12 pp.177-186
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Hadoop distributed file system (HDFS) is designed to store huge data set reliably, has been widely used for processing massive-scale data in parallel. In HDFS, the data locality problem is one of critical problem that causes the performance decrement of a file system. To solve the data locality problem, we propose an efficient data replication scheme based on access count prediction in a Hadoop framework. By the previous data access count, the existing data replication scheme predicts the next access count of data files using Lagrange’s interpolation. Then, the proposed data replication scheme determines the replication factor with the predicted data access count, whether it generates a new replica or it uses the loaded data as cache selectively. Finally, the proposed scheme provides improvement of data locality. By performance evaluation, proposed efficient data replication scheme is compared with default data replication setting of Hadoop that shows proposed scheme reduces averagely 8.9% of the task completion time in the map phase. Regarding the data locality, proposed scheme provides the increase of node locality by 6.6% and the decrease of rack and rack-off locality by 38.9% and 56.5%.
A Study on Effect of Code Distribution and Data Replication for Multicore Computing Architectures KCI 등재
국제문화기술진흥원 International Journal of Advanced Culture Technology(IJACT) Volume 9 Number 4 2021.12 pp.282-287
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
A multicore system must be able to take full advantage of the program's instruction and data parallelism. This study introduces the data replication technique as a support technique to maximize the program's instruction and data parallelism. Instruction level parallelism can be limited by data dependency. In this case, if data is replicated to each processor core and used, instruction level parallelism can be used to the maximum. The technique proposed in this study can maximize the performance improvement effect when applied to scientific applications such as matrix multiplication operation.
A Novel Replication Model with Enhanced Data Availability in P2P Platforms SCOPUS
보안공학연구지원센터(IJGDC) International Journal of Grid and Distributed Computing Vol.9 No.4 2016.04 pp.151-160
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
With the development of P2P technology, more and more distributed systems are applying it for deploying large-scale storage infrastructure. However, maintaining desirable data availability in P2P environments is still an opening issue because of the vulnerability of participated peers. In this paper, a novel replication model is proposed to improving the data availability of P2P storage platforms. In the proposed replication model, the probability of peer’s failure is estimated by semi-Markov chain, which enables us to significantly improving the prediction accuracy of peer’s failure in a given period. Massive experiments are conducted in a real-world P2P platform to examining the performance of the proposed replication model, and the results indicate that it can achieve better tradeoffs between performance and data availability comparing with other replication schemes.
보안공학연구지원센터(IJBSBT) International Journal of Bio-Science and Bio-Technology Vol.4 No.1 2012.03 pp.49-64
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Recent advancement in the field of life science data mining has inspired researchers and healthcare professionals to apply this novel technology to obtain descriptive patterns and predictive models from biomedical and healthcare databases. The discovery of hidden biomedical patterns from large clinical database can uncover potential knowledge to support prognosis and diagnosis decision makings. However, clinical application of data mining algorithms has a severe problem of low predictive accuracy rate that hampers their wide usage in the clinical environment. We thus focus our study on the improvement of predictive accuracy of the models created from the data mining algorithms. Our main research interest concerns the problem of learning a classification model from a multiclass data set with low prevalence rate of some minority classes. With such data characteristics, directly applying classification data mining techniques such as decision tree induction, regression analysis, neural networks, or support vector machines yields a suboptimal model in terms of predictive accuracy rate. To remedy the imbalanced class distribution among data instances, we apply random over-sampling and synthetic minority over-sampling (SMOTE) techniques to increase the predictive performance of the learned model. In our preliminary study, we consider specific kinds of primary tumors occurring at the frequency rate less than one percent as rare and minority classes. From the experimental results, the SMOTE technique gave a high specificity model, whereas the random over-sampling produced a high sensitivity classifier. The precision performance of a classification model obtained from the random over-sampling technique is on average much better than the model learned from the original imbalanced data set. We then extend our study by designing the heuristic based method to cope with the abundance of irrelevant feature that causes the decrease in learning time and sometimes lower the accuracy rate. The over-sampling technique and the heuristic-based feature selection are coupled as a data preparation method to deal with imbalanced data sets with many irrelevant features. The experimental results on arrhythmia and communities-and-crime data sets show significant improvement on the predicting accuracy, specificity, and sensitivity of the induced models.
Based on the Correlation of the File Dynamic Replication Strategy in Multi-Tier Data Grid SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.1 2015.02 pp.75-86
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The data grid is one of the infrastructures for data and storage resource management. Data replication can improve data access performance, fault tolerance, and reduce network transmission bandwidth. As the site storage space is limited in the grid, how to effectively use grid resources become an important challenge. This paper introduces a dynamic grid replication algorithm based on popularity support and confidence (BPSC). Through the algorithm, data and its associated copy can be placed on a suitable site, with reduced access latency. Through Optorsim, simulation results show that the algorithm can provide better performance compared with other algorithms
엔터프라이즈 재해복구시스템 구축을 위한 데이터 복제 방안 연구 KCI 등재
국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.10 No.1 2024.01 pp.411-417
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
재해 발생 시 주요 IT 인프라 중단을 최소화하고 연속적인 비즈니스 서비스를 제공하기 위한 재해복구 계획 및 재해복구시스템의 구축은 반드시 필요하다. 재해복구시스템 구축과정에서 데이터 복제는 재해 발생 시 중단 없는 연속적인 비즈니스 서비스 제공을 위한 데이터 복구의 핵심요소로 데이터 복제방식은 시스템 구성환경과 재해복구 목표수준에 따라 결정할 수 있다. 본 논문에서는 재해복구시스템 구축에서 구성환경과 재해복구 목표수준에 적합한 데이터 복제방식 결정 방안에 대해 제시한다. 또한 복제방식 결정 절차를 적용하여 재해복구시스템을 구축하고 구축 결과를 분석한다. 재해복구시스템 구축 후 재해 상황에서 재해복구센터로 서비스가 전환, 정상적인 서비스가 진행되 는지를 판단하기 위한 모의 테스트를 진행하고 결과를 분석하였다. 그 결과 재해복구시스템 구축 단계에서는 체계적 으로 최적의 데이터 복제방식의 선정이 가능했다. 구축된 재해복구시스템은 연속적인 비즈니스 서비스 제공을 위해 재해복구센터로 서비스 전환되는 시간 RTO는 3.7시간으로, Tier 2였던 재해복구 수준이 목표수준 RTO 4시간 이내, RPO=0으로 개선되었다.
In the event of a disaster, it is essential to establish a disaster recovery plan and disaster recovery system to minimize disruption to major IT infrastructure and provide continuous business services. In the process of building a disaster recovery system, data replication is a key element of data recovery to provide uninterrupted and continuous business services in the event of a disaster. The data replication method can be determined depending on the system configuration environment and disaster recovery goal level. In this paper, we present a method for determining a data replication method suitable for the configuration environment and disaster recovery target level when building a disaster recovery system. In addition, the replication method decision procedure is applied to build a disaster recovery system and analyze the construction results. After establishing the disaster recovery system, a test was conducted to determine whether the service was transferred to the disaster recovery center in a disaster situation and normal service was provided, and the results were analyzed. As a result, it was possible to systematically select the optimal data replication method during the disaster recovery system construction phase. The established disaster recovery system has an RTO of 3.7 hours for service conversion to the disaster recovery center to provide continuous business services, and the disaster recovery level, which was Tier 2, has been improved to the target level within 4 hours of RTO and RPO=0.
데이터그리드 환경에서의 전문가 시스템을 이용한 Data Replication 시스템의 개선
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2003 pp.173-176
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
기존 데이터그리드에서는 대용량의 데이터 서비스를 위해 Peer-to-Peer 기반의 Data Replication 을 사용하였다. 하지만 기존의 방법은 로컬 사이트의 정보만을 가지고 Data Replication 을 수행하므로, 데이터그리드 전체의 Data Replication 을 수행하는데 비효율적이다. 이에 본 논문에서는 Rule 기반 Forward Chaining 을 수행하는 전문가 시스템을 이용하여, 데이터그리드 전체의 Data Replication 을 수행하는 방법을 제안하고 이를 구현하였다.
OGSA 에서의 그리드 정보 서비스를 기반으로한 대용량 Data Replication 관리
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2003 pp.193-196
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
그리드 환경에서 OGSA(Open Grid Service Architecture)는 분산된 서비스의 이용 편의를 위한 시스템 독립적인 인터페이스를 제공한다. 하지만 OGSA 에서 사용자가 작업 수행시 필요로 하는 QoS 와 서비스의 신뢰성을 보장하기 위해, 동일한 그리드 정보 제공자를 이용하는 여러 서비스간의 공유 자원에 대한 경쟁 문제를 해결해야 한다. 본 논문에서는 OCSA 에서 여러 서비스의 효율적인 자원 할당을 보장하는 다이나믹 Data Replication 관리를 위한 의사결정 알고리즘을 제안한다.
Variance Estimation for Imputed Survey Data using Balanced Repeated Replication Method
[Kisti 연계] 한국통계학회 Communications for statistical applications and methods Vol.12 No.2 2005 pp.365-379
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Balanced Repeated Replication(BRR) is widely used to estimate the variance of linear or nonlinear estimators from complex sampling surveys. Most of survey data sets include imputed missing values and treat the imputed values as observed data. But applying the standard BRR variance estimation formula for imputed data does not produce valid variance estimators. Shao, Chen and Chen(1998) proposed an adjusted BRR method by adjusting the imputed data to produce more accurate variance estimators. In this paper, another adjusted BRR method is proposed with examples of real data.
Replication Server: An answer to distributed data base
[Kisti 연계] 한국경영과학회 한국경영과학회 학술대회논문집 1995 p.45
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
데이터 접근 기록 정보를 이용한 적응적 데이터 복제 기법 제안
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2004 pp.937-940
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
프로세서 자원, 데이터 저장장치 자원을 제공하면서 가상기관(Virtual Organization)을 구성하는 각 사이트는 사용할 수 있는 네트윅 자원이 한정된 상황에서 애플리케이션 처리량을 극대화하는 최적화된 데이터그리드 시스템을 기대한다. 본 논문에서는 크기가 제한적이며 지리적으로 분산된 데이터 저장공간에서 적응적 데이터 복제 기법을 제안하고 Replica의 지리적 분배를 위한 평가 모델을 제안한다. 이를 위해 논리 시간 데이터 접근 기록 및 통계를 적용하여 복제할 파일들을 구분 하는 이산적 결정 모델을 제안하고 삭제할 Replica 결정에 논리 시간 접근 기록을 적용한다.
데이터 접근 기록 정보를 이용한 적응적 데이터 복제 기법 제안
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2004 pp.937-940
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
프로세서 자원, 데이터 저장장치 자원을 제공하면서 가상기관(Virtual Organization)을 구성하는 각 사이트는 사용할 수 있는 네트윅 자원이 한정된 상황에서 애플리케이션 처리량을 극대화하는 최적화된 데이터그리드 시스템을 기대한다. 본 논문에서는 크기가 제한적이며 지리적으로 분산된 데이터 저장공간에서 적응적 데이터 복제 기법을 제안하고 Replica의 지리적 분배를 위한 평가 모델을 제안한다. 이를 위해 논리 시간 데이터 접근 기록 및 통계를 적용하여 복제할 파일들을 구분 하는 이산적 결정 모델을 제안하고 삭제할 Replica 결정에 논리 시간 접근 기록을 적용한다.
클라우드 스토리지 시스템에서 데이터 접근빈도와 Erasure Codes를 이용한 데이터 복제 기법
[Kisti 연계] 대한전자공학회 Journal of the Institute of Electronics Engineers of Korea Vol.51 No.2 2014 pp.85-91
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
클라우드 스토리지 시스템은 데이터의 저장과 관리를 위해서 분산 파일시스템을 사용한다. 기존 분산 파일시스템은 데이터 디스크의 손실 발생시 이를 복구하기 위해서 3개의 복제본을 만든다. 그러나 데이터 복제 기법은 저장공간을 원본 파일의 복제 횟수만큼 필요로하고 복제과정에서 입출력 발생이 증가하는 문제가 있다. 본 논문에서는 SSD 기반 클라우드 스토리지 시스템에서 저장공간 효율성 향상과 입출력 성능 향상을 위하여 Erasure Codes를 이용한 데이터 복제 기법을 제안한다. 특히, 데이터 접근 빈도에 따라 복제 횟수를 줄이더라도 Erasure Codes를 사용하여 데이터 복구 성능을 동일하게 유지하였다. 실험 결과 제안한 기법이 HDFS 보다 저장공간 효율성은 최대 약40% 향상되었으며, 읽기성능은 약11%, 쓰기성능은 약10% 향상됨을 확인하였다.
Cloud storage system uses a distributed file system for storing and managing data. Traditional distributed file system makes a triplication of data in order to restore data loss in disk failure. However, enforcing data replication method increases storage utilization and causes extra I/O operations during replication process. In this paper, we propose a data replication method using erasure codes in cloud storage system to improve storage space efficiency and I/O performance. In particular, according to data access frequency, the proposed method can reduce the number of data replications but using erasure codes can keep the same data recovery performance. Experimental results show that proposed method improves performance in storage efficiency 40%, read throughput 11%, write throughput 10% better than HDFS does.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.