년 - 년
커넥티드 카–분산 엣지 환경에서 실시간 대용량 데이터 처리를 위한 엣지 노드간 웹소켓 통신 기법 및 디지털 트윈 시뮬레이션 KCI 등재
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.20 No.3 2024.06 pp.1-11
커넥티드 카에서 발생하는 실시간 데이터 처리를 위한 엣지–커넥티드 카 환경에서 웹소켓 통신 연구가 활발히 진행 되고 있다. 그러나 웹소켓은 연결 유지, 고정된 연결의 특성으로 대용량 데이터 처리를 위한 분산 엣지 환경을 구성 하는데, 구조적 한계가 있다. 기존 연구는 세션 공유, 아파치 카프카 통신을 활용했으나, 이는 추가 리소스와 종속 성, 확장성 문제가 존재한다. 본 연구는 실시간성, 확장성, 낮은 리소스를 보장하는 새로운 웹소켓 기반 데이터 통신 기법을 제안하며. 이를 Digital Twin 모델에 적용하였고 차량 충돌 이벤트 처리를 통해 제안된 기법의 효율성을 검증했다.
Research on WebSocket communication in edge-connected car environments is ongoing. WebSockets have structural limits for large data handling due to persistent and fixed connections. Previous approaches using session sharing and Apache Kafka faced scalability and resource dependency issues. This study proposes a new WebSocket-based technique that ensures real-time, scalable, and resource-efficient data communication. It has been applied to the Digital Twin model, proving its effectiveness through vehicle collision event handling.
디지털 포렌식 증거 채택 기준의 한계와 개선 방안 KCI 등재
한국융합보안학회 융합보안논문지 제18권 제4호 2018.10 pp.35-43
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
현재 디지털 증거는 증거 채택에 있어 자유심증주의를 택하는 바, 채택의 기준이 명확하지 못하며, 분산처리 기법으로 디지 털 증거의 분석시간을 단축시킬 수 있으나 암호기술의 발전 등으로 영장청구 제한시간인 48시간에 적합하지 못한다는데 문제 점이 있다. 본 논문에서는 민/형사 소송절차에서의 증거능력과 증명력을 판례를 중심으로 분석하여, 자유심증주의를 대체할 객관적이고 세부적인 채택기준의 필요성에 대해 논하고자 한다. 또한, 포렌식 관점에서의 영장 청구 제한시간에 대한 문제점 과 이에 대한 개선안으로 디지털 증거분석 사전신청제를 제시하고자 한다.
Currently, digital evidence takes judicial discretion in adopting it, which does not clarify the criteria for adoption, and it can shorten the analysis time of digital evidence with distributed processing techniques. However, due to the development of cryptographic techniques, there is a problem in that it is not suitable for the 48 hour limit of the warrant request. In this paper, we analyze the precedents for admissibility of evidence and the probative power in the civil/criminal proceedings, and discuss the need for objective and detailed adoption criteria to replace judicial discretion. In addition, we'd like to propose a preliminary application form for analysis of digital evidence as a problem for limit time for warrant claims from the perspective of forensics and a solution to the problem.
분산 처리(Distributed Processing)를 위한 표준화 - 개방형 분산 처리(Open Distributed Processing)
한국컴퓨터통신연구회 OSIA Standards & Technology Review Journal 제5권 제2호 통권12호 1991.09 pp.32-43
※ 기관로그인 시 무료 이용이 가능합니다.
4,300원
위치기반 서비스란 휴대폰과 같은 모바일 단말기 속에 위성 항법장치(GPS)와 연결되는 칩을 부착하여 위치추적 서비스, 공공안전 서비스, 위치기반 정보 서비스 등 위치와 관련된 각종 정보를 제공하는 서비스를 일컫는다. 그리고 클라우드 컴퓨팅이란 개인용 컴퓨터 또는 기업의 서버에 개별적으로 저장해 두었던 자료와 소프트웨어들을 클라우드 클러스터로 구축하여 필요할 때 PC나 휴대폰 같은 각종 단말기를 이용하여 원격 작업을 수행할 수 있는 환경을 의미한다. 본 논문에서는 구글의 연구로부터 시작된 클라우드 컴퓨팅 기술에 주목하고 클라우드 컴퓨팅 환경을 위한 대용량 이동 객체의 정보를 저장하고 질의하는 시스템에 대하여 연구하였다. 그리고 위치 기반 서비스를 위한 대용량 이동객체에 대하여 효율적인 질의를 수행할 수 있는 저장 구조를 제안하였다. 본 논문에서 제안한 방법을 적용한 시스템의 효율성을 증명하기 위해 저장 및 질의 성능을 다양한 측면에서 실험하고, 실험 결과를 다른 분산처리 시스템과의 비교하여 성능을 입증하였다.
The location-based service is referred to as the service for providing various types of information regarding such as the location-tracing service, the public-safety service and the location-based information service by attaching chips connected with the GPS to the mobile phone. In addition, cloud computing means the environment which makes it possible to carry out remote operations by establishing cloud clusters with the data and software which have been individually stored in personal computer or corporate servers and using devices such as personal computers and cellular phones whenever it is necessary. Throughout this paper, the cloud computing technology which originates from the research by Google the system has been focused on and the system for storing a large volume of information for moving object and making query for cloud computing environment has been studied. Besides, the storing structure for executing efficient query regarding the high-capacity mobile devices for the location-based service has been suggested. In order to prove the efficiency of the system with the application of the method suggested in this paper, the storage and qualitative functions have been experimented in various aspects and compared with other distributed processing system's performance.
멀티 플랫폼 게임을 위한 메시지 기반의 분산처리 시스템 KCI 등재후보
한국컴퓨터게임학회 컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) 제8호 2006.06 pp.29-37
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
빅데이터 환경에서 분산 처리 기반의 데이터 마이닝 기법
한국인공지능교육학회 인공지능연구 논문지 Vol.2 No.3 2021.12 pp.29-36
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본격적인 4차 산업혁명 시대가 도래하면서 인공지능 기반의 기술 개발과 활용은 그 어느 때보다 중요해지고 있다. 데이터 마이닝은 4차 산업혁명과 인공지능 시대의 초석이라고 할 수 있는 전통적인 정보 처리 분야이며, 이러한 데이터 마이닝을 기반으로 다양한 정보 분석을 수행하고 인공지능 기술에 접목하기 위해서는 빅데이터 환경에서 효율적으로 실행 가능한 체계 또는 방법론이 필요하 다. 빅데이터 환경에서는 데이터의 발생 속도가 매우 빠르고 용량이 크기 때문에 일반적인 단일 서버 방식으로는 데이터 마이닝이 불가능하다. 이러한 문제를 해결하기 위해, 본 논문에서는 빅데이터 처리를 할 수 있는 분산 병렬 처리 기반의 데이터 마이닝 기법 을 제안한다. 제안하는 방법은 트랜잭션을 여러 개의 샤드로 분할하고, 빈발 항목집합 탐사 과정을 내부 및 외부 패턴으로 구분하여 분산 병렬 처리한 뒤, 이전에 생성된 집합과 새로 갱신된 결과를 병합하여 최종 결과를 도출하는 3단계로 구성된다. 이 과정에서 상대적으로 정확도는 약간 감소하지만 다중 스캔 없이 빠르게 근사적으로 클러스터 내의 빈발 항목집합을 탐사한다. 제안하는 방법 을 통해 하둡 맵리듀스 환경에서 점진적으로 마이닝 결과를 탐사하고 단일 서버에서는 제공하지 못하는 안정적인 유연성과 확장성 을 확보할 수 있다.
With the advent of the era of the 4th Industrial Revolution, the development and utilization of the artificial intelligence(AI) technologies becomes more important than ever. Data mining is a traditional information processing field which is the basis of the 4th Industrial Revolution and AI era. It is necessary to develop a system or methodology that can be efficiently executed in a big data environment in order to perform various information analysis and incorporate into artificial intelligence technology. In a big data environment, since data is generated very fast and its volume is large, it is impossible to process the method with a general single server approach. To solve this problem, this paper proposes a distributed parallel processing approach for data mining process. The proposed method consists of three steps: dividing the transaction into several shards, distributed frequent itemsets mining process with the internal and external patterns in parallel, and merging the previously generated set with newly updated results. Through this process, the accuracy is slightly decreased, however the set of frequent itemsets can be quickly and approximately explored without multiple scans. The proposed method can gradually find the resulting sets using the Hadoop MapReduce technique and provide flexibility and scalability.
IP보안 카메라의 보급으로 원격에서 얼굴인식을 수행함에 있어 서버의 부하를 줄이기 위한 여러 가지 방법들이 구현되고 있다. 본 논문에서는 원격지에 있는 IP 보안 카메라 영상을 얼굴검출기능이 탑재된 DSP 보드를 통해 입력받아 얼굴검출을 수행 한 후 해당 얼굴영역 이미지를 서버로 전송하여 이를 얼굴인식 분산 처리를 통해 얼굴인식 기능을 수행한다. 결과적으로 전체적인 서버시스템 로드를 상당히 줄이는 성과와 실시간 얼굴 인식처리를 최대 256 대의 카메라를 연동하면서 수행할 수 있는 장점을 가지고 있다. 이를 수행할 수 있는 기술은 분산처리 서버기술을이용하여 한 서버 당 64채널 얼굴인식을 수행하며, 4개 분산처리 서버를 운영할 경우 250여개 카메라 채널을 통한얼굴검출 결과를 처리하는 성과를 가져올 수 있었다.
Various methods for reducing the load on the server have been implemented in performing face recognition remotely by the spread of IP security cameras. In this paper, IP surveillance cameras at remote sites are input through a DSP board equipped with face detection function, and then face detection is performed. Then, the facial region image is transmitted to the server, and the face recognition processing is performed through face recognition distributed processing . As a result, the overall server system load and significantly reduce processing and real-time face recognition has the advantage that you can perform while linked up to 256 cameras. The technology that can accomplish this is to perform 64-channel face recognition per server using distributed processing server technology and to process face search results through 250 camera channels when operating four distributed processing servers there was.
AM60 합금 내 MWCNT 용탕 분산처리를 통한 기계적 물성 거동에 관한 연구 KCI 등재
한국기계항공기술학회(구 한국기계기술학회) 한국기계항공기술학회지(구 한국기계기술학회지) 제18권 제6호 2016.12 pp.918-923
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Magnesium alloy is becoming known for the lightest material in the metallic materials. Recently the automotive industry has a variety application to the light weight parts replacement. This study focuses on the mechanical property improving through a tiny amount’s CNT addition into the magnesium alloy as AM60. The CNT material is an arduous combination of the metallic materials. Therefore this study is concentrating on the contact force growth for the CNT material. Consequently, the made CNT is produced by the CVD process using the magnesium catalyst. The CNT material has dispersive with mechanical process into the molten AM60 alloy. The mechanical experiment result that hardness is 18% increasing and tensile strength is 13% increasing, better than the raw AM60 alloy on this investigation.
RGBD 카메라와 분산처리 기법을 이용하여 볼류메트릭 시스템에서 사용 가능한 실시간 3D 캐릭터를 생성하는 방법에 관한 연구 KCI 등재
대한산업경영학회 산업융합연구(구 대한산업경영학회지) 제23권 제12호 2025.12 pp.59-72
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
본 연구는 다중 RGBD 센서를 실시간 볼류메트릭 시스템에 적용하여, 대역폭・동기화・센서 간 간섭 문제를 극복하고 안정적으로 초당 30프레임 이상을 구현하기 위한 기술적 해법을 제시한다. 총 9대의 RGBD 카메라를 데이지 체인 방식으로 동기화하고, 3대의 영상 데이터 생성기에서 분산 전처리를 수행한 뒤, 볼류메트릭 서버에서 3차원 데이터를 합성・렌더링하여 실시간 서비스를 제공한다. 분산 전처리된 3차원 영상 데이터는 이더넷을 통해 전송되며 데이터를 수신한 볼류메트릭 서버는 초당 30 프레임 이상 속도로 렌더링을 수행한다. 데이지 체인 동기화로 센서 간 적외선 간섭을 줄여 깊이 정보 품질을 높였으 며, 레지스트레이션 과정을 통해 서로 다른 좌표계를 단일 글로벌 좌표계로 정합하여 볼류메트릭 모델을 실시간으로 생성하 였다. 대역폭・연산 부담・센서 간 간섭 문제를 효과적으로 해소함으로써 실시간성과 정확도를 모두 충족시켰으며, 향후 딥러 닝 후처리 또는 센서 확장을 통해 고해상도・저지연 볼류메트릭 구현이 가능할 것으로 기대한다.
This study proposes a technical solution for integrating multiple RGBD sensors into a real-time volumetric system, overcoming bandwidth constraints, synchronization challenges, and sensor interference while maintaining a stable frame rate of over 30 fps. Nine RGBD cameras are synchronized in a daisy chain method, distributed preprocessing is performed on three video data generators, and 3D data is synthesized and rendered on a volumetric server to provide real-time services. Distributed preprocessed 3D image data is transmitted via Ethernet, and the volumetric server that receives the data performs rendering at a speed of more than 30 fps. Daisy-chain synchronization reduces infrared interference among sensors, thereby improving depth data quality. In addition, a registration process aligns disparate coordinate systems into a single global reference frame for real-time volumetric model generation. This approach effectively addresses bandwidth, computational overhead, and sensor interference issues, ensuring both real-time performance and high accuracy. Future work includes deep-learning-based post-processing or sensor expansion to achieve higher-resolution, lower-latency volumetric content.
분산 시맨틱 웹 환경에서의 효율적인 시공간 RDF 데이터 처리 KCI 등재
한국EA학회 정보화연구 제14권 4호 2017.12 pp.367-373
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
빅데이터 시대가 도래하면서, 시맨틱 웹 환경에서도 다양한 건물, 도로, 시설물 등에서 생성되는 다양한 빅데이터를 시공간 RDF(Resource Description Framework) 형태로 저장 및 교환하는 방식 이 증가하고 있다. 시공간 RDF 데이터는 일반 및 시공간 데이터가 같이 저장되어 있으므로, 시공간 정보를 효율적으로 관리하기 위한 시공간 데이터 타입, 연산자, 인덱스 등이 지원되어야 한다. 그러나 기존 시맨틱웹 기반 시스템들은 시공간 RDF 데이터를 저장하거나 시공간 특성에 따른 연산 처리를 효율적으로 지원하지 못하며, MongoDB, HBase 등의 분산 데이터베이스 시스템들 또한 시공간 연 산 및 인덱스의 부재로 시공간 정보의 특성을 반영한 데이터 검색이 어렵다. 따라서, 본 논문에서는 분산 시맨틱 웹 환경에서 기존에는 지원하지 않던 시공간 RDF 데이터의 통합 처리가 가능하고 효율 적인 시공간 연산을 지원하도록 시공간 데이터 처리 기술을 제안한다. 시공간 데이터 처리 기술은 OGC의 공간 표준을 따르는 공간 데이터 타입 및 연산자를 지원하고, ISO의 시간 표준을 따르는 시 간 데이터 타입 및 시간 연산자를 지원하며, 이를 통합한 시공간 데이터 타입, 시공간 연산자, 시공간 인덱스를 지원한다. 그리고 관련 연구와의 검색 시간을 비교하여 연구의 우수성을 검증하였다.
분산 스트림 처리 시스템의 데이터 손실방지를 위한 적응적 Upstream Backup 알고리즘 KCI 등재후보
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.6 No.5 2010.10 pp.28-35
분산 스트림 처리 시스템을 위한 고가용성 알고리즘에는 Passive Standby, Active Standby, Upstream Backup 알고리즘 등이 있다. 기존의 고가용성 알고리즘은 복구 시 필요한 데이터의 백업을 위한 bandwidth overhead가 크며 노드의 연산결과를 다수의 downstream 노드들이 공유하며 데이터의 유입률이 폭발적으로 증가하는 경우에 출력 큐의 오버플로우로 인한 데이터의 손실 문제가 발생 할 수 있다. 본 논문은 이러한 문제들을 해결하기 위해 데이터 스트림의 유입량과 노드들의 연산처리율 모니터링을 통해 백업 방법을 유동적으로 변경시키는 적응적 Upstream Backup 알고리즘을 제안한다.
There have been High-Availability algorithms for Distributed Stream Processing System such as Passive Standby, Active Standby and Upstream Backup. Existing High Availability Algorithms have high bandwidth overhead when they backup data needed for restoring from the system failure and data can be lost by overflow of output queue in case of explosive increasing of input data rates. In this paper, we suggest adaptive Upstream Backup algorithm which changes fluidly the way to backup using monitor input rates of data stream and throughput of operation to solve those problems.
대용량 민감 데이터 보호를 지원하는 비트맵 기반 분산 색인구조 및 암호화 질의처리 기법 KCI 등재
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.11 No.2 2015.04 pp.24-40
대용량 민감 데이터에 대한 아웃소싱이 각광받음에 따라, 이를 보호하기 위한 데이터 암호화 기법이 요구되고 있다. 이에 따라, 대용량 암호화 데이터 관리를 지원하는 분산 색인 구조 및 암호화된 데이터 상에서의 질의처리 알고리즘 이 요구되고 있다. 그러나 기존 분산 색인 구조 중 암호화 데이터의 특성을 고려한 연구는 존재하지 않는다. 또한, 기존 암호화 질의처리 알고리즘은 지원 가능한 질의 타입이 한정적이며, 상이한 방식으로 암호화된 컬럼 간 연산을 지원하지 못하는 문제점이 존재한다. 이를 해결하기 위해, 본 논문에서는 비트맵 기반 분산 색인 구조 및 암호화 질 의처리 기법을 제안한다. 제안하는 분산 암호화 색인 구조는 데이터 프라이버시를 보장하며, 다양한 종류의 질의에 대해 성능 향상을 제공한다. 아울러, 제안하는 암호화 질의처리 기법은 복호화를 수행하지 않고 질의처리를 수행함 으로써 데이터 보호 수준을 향상시키며, 높은 질의 처리 성능 및 정확도를 보장한다. 아울러 성능평가를 통해 제안 하는 색인구조 및 암호화 질의처리 기법이 대용량 민감 데이터 보호에 적합함을 보인다.
As the outsourcing of the large sensitive data has been highlighted, data encryption schemes to protect the sensitive data are required. Accordingly, it is necessary to develop not only a distributed index structure to manage the large amount of encrypted data, but also a query processing scheme over the encrypted data. However, there has been no index structure considering the encrypted data. Existing query processing schemes over the encrypted data can support limited types of queries. In addition, the schemes cannot support operations among data with different columns because they use different types of encryption schemes depending on their attribute type. To solve these problems, in this paper, we propose a bitmap-based distributed index structure and a query processing scheme for the encrypted data. The proposed distributed index structure guarantees data privacy preservation and performance improvement for the various types of queries. In addition, by processing a query over the encrypted data without data decryption, the proposed query processing scheme guarantees the high query performance and accuracy while preserving the data privacy. Finally, we show from our performance evaluation that our proposed index structure and query processing scheme are suitable for protecting the data privacy of the large sensitive data.
Distributed learning 에서의 통신 상황을 고려한 학습과 추론 연산 분배에 대한 분석 연구
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 2021 한국차세대컴퓨팅학회 춘계학술대회 2021.05 pp.59-62
클라우드 뿐 아니라 모바일 기기들이 함께 학습과 추론을 위한 연산들을 나누어 수행하고 이를 다시 종합하는 distributed learning은 주목 받고 있는 learning 기법들 중 한 형태이다. 이러한 형태의 learning의 경우, 모바일 기기들을 활용하여 학습하고 추론 연산을 수행할 때, 응용의 성능뿐 아니라 통신 상황과 모바일 기기들에서 소모되는 전력을 고려하는 것은 매우 중요하다. 따라서 본 논문에서는 엣지 클라우드와 연결된 멀티 홉 기반의 네트워크에서 distributed learning을 수행할 때, 통신 상황과 소모 전력 관점에서 학습 및 추론 연산 분배에 따른 주요 case들에 대한 분석을 수행한다.
빅데이터 수집 처리를 위한 분산 하둡 풀스택 플랫폼의 설계 KCI 등재
한국융합학회 한국융합학회논문지 제12권 제7호 2021.07 pp.45-51
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
급속한 비대면 환경과 모바일 우선 전략에 따라 해마다 많은 정형/비정형 데이터의 폭발적인 증가와 생성은 모든 분야에서 빅데이터를 활용한 새로운 의사 결정과 서비스를 요구하고 있다. 그러나 매년 급속히 증가하는 빅데이터 를 활용하여 실무 환경에서 적용 가능한 표준 플랫폼으로 빅데이터를 수집하여 적재한 후, 정재한 빅데이터를 관계형 데이터베이스에 저장하고 처리하는 하둡 에코시스템 활용의 참조 사례들은 거의 없었다. 따라서 본 연구에서는 스프링 프레임워크 환경에서 3대의 가상 머신 서버를 통하여 하둡 2.0을 기반으로 쇼셜 네트워크 서비스에서 키워드로 검색한 비정형 데이터를 수집한 후, 수집된 비정형 데이터를 하둡 분산 파일 시스템과 HBase에 적재하고, 적재된 비정형 데이 터를 기반으로 형태소 분석기를 이용하여 정형화된 빅데이터를 관계형 데이터베이스에 저장할 수 있게 설계하고 구현하 였다. 향후에는 데이터 심화 분석을 위한 하이브나 머하웃을 이용하여 머신 러닝을 이용한 클러스터링과 분류 및 분석 작업 연구가 지속되어야 할 것이다.
In accordance with the rapid non-face-to-face environment and mobile first strategy, the explosive increase and creation of many structured/unstructured data every year demands new decision making and services using big data in all fields. However, there have been few reference cases of using the Hadoop Ecosystem, which uses the rapidly increasing big data every year to collect and load big data into a standard platform that can be applied in a practical environment, and then store and process well-established big data in a relational database. Therefore, in this study, after collecting unstructured data searched by keywords from social network services based on Hadoop 2.0 through three virtual machine servers in the Spring Framework environment, the collected unstructured data is loaded into Hadoop Distributed File System and HBase based on the loaded unstructured data, it was designed and implemented to store standardized big data in a relational database using a morpheme analyzer. In the future, research on clustering and classification and analysis using machine learning using Hive or Mahout for deep data analysis should be continued.
대용량 분산 클러스터 기반의 시계열 빅-데이터 처리 및 수집 시스템
한국정보통신설비학회 한국정보통신설비학회 학술대회 2016년도 정보통신설비 학술대회 2016.09 pp.232-233
The technology of IoT which began in the sensor networks generate numerous sensing data from a variety of physical environments to virtual world. There are many types of IoT device used the sensors based on application. They usually create time series data that people can use to make more sensible decision. Things produced various data which is generated and stored as time based and then saved them to the existing database. In this paper, we introduce the system architecture for the large distributed cluster based system for the processing and collecting of big-data of time-series to meet the goal of performance using Redis Cluster and OpenTSDB.
보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.8 No2 2013.03 pp.213-224
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In this paper, we propose a Hadoop-based Distributed Video Transcoding System in a cloud computing environment that transcodes various video codec formats into the MPEG-4 video format. This system provides various types of video content to heterogeneous devices such as smart phones, personal computers, television, and pads. We design and implement the system using the MapReduce framework, which runs on a Hadoop Distributed File System platform, and the media processing library Xuggler. Thus, the encoding time to transcode large amounts of video content is exponentially reduced, facilitating a transcoding function. For performance evaluation, we focus on measuring the total time to transcode a data set into a target data set for three sets of experiments. We also analyze the experimental results, providing optimal Hadoop Distributed File System and MapReduce options suitable for video transcoding. Based on the experiments performed on a 28-node cluster, the proposed distributed video transcoding system provides excellent performance in terms of speed and quality.
Efficient Processing Distributed Joins with Bloomfilter using MapReduce
보안공학연구지원센터(IJGDC) International Journal of Grid and Distributed Computing Vol.6 No.3 2013.06 pp.43-58
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The MapReduce framework has been widely used to process and analyze largescale datasets over large clusters. As an essential problem, join operation among large clusters attracts more and more attention in recent years due to the utilization of MapReduce. Many strategies have been proposed to improve the efficiency of distributed join, among which bloomfilter is a successful one. However, the bloomfilter’s potential has not yet been fully exploited, especially in the MapReduce environment. In this paper, three strategies are presented to build the bloomfilter for the large datasets using MapReduce. Based on these strategies, we design two algorithms for two-way join and one algorithm for multi-way join. The experimental results show that our algorithms can significantly improve the efficiency of current join algorithm. Moreover, cost models of these algorithms are characterized in order to find out the way of improving the performance of two-way and multi-way joins.
보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.10 No.1 2016.01 pp.51-58
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In this paper, we propose an in-memory distributed processing method that can rapidly process vehicle location and traffic event data using Spark Streaming. The proposed system enables to share information about surrounding vehicles, pedestrians, and traffic events in real time with drivers who use the WEVING service. In the proposed method, vehicle location and traffic event streams are indexed using the grid indexing technique according to time, and the continuous range query method is processed based on the index. Also, traffic events are grouped based on occurrence time, location, content, and road segment of the traffic event transferred in real time in order to avoid duplicated traffic events. Through experiments, we show that the proposed method is able to deduplicate similar traffic events efficiently.
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.5 2015.10 pp.15-26
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Hadoop–a popular open-source implementation of MapReduce is widely used for the analysis of large datasets. The current Hadoop implementation assumes that computing nodes in a cluster are homogeneous in nature. In this paper we evaluate performance of Hadoop Platform and Oracle for Distributed Parallel Processing in large datasets. For evaluation, we implement a prototype of a virtual datacenter using distributed and parallel computing technology. The purpose of this paper is to reduce datacenter implementation cost using commodity hardware and provide high performance. Hadoop is installed on a commodity Linux cluster the distributed processing of large data sets across clusters of computers using distributed and parallel computing architecture. This paper also helps to explain about some new technology and framework which are open source; that can easily utilize those technologies for our complex data analysis which resembling structured, semi structured and non-structured data. Here we tried to demonstrate a performance comparison by executing some queries between distributed parallel computing system and traditional single computing system. For the simulation of the infrastructure Hadoop cluster has been used for distributed parallel processing and Oracle 11g is used for traditional single processing system. We prepare three virtual host for Hadoop cluster and a high-end hardware for Oracle 11g.
Study on the Distributed Crawling for Processing Massive Data in the Distributed Network Environment SCOPUS
보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.10 No.10 2015.10 pp.375-384
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Due to the development of IT, distribution of smart phone, and an increase of use of SNS, various types of contents are being produced and consumed in Internet. Therefore, information searching technology has become important due to a sharp rise in data. However, information searching technology requires much of background knowledge and hence has been recognized as what was difficult to access to. Issues with previous search engine were how many of qualified personnel with background knowledge along with huge amount of development expenses were required. Therefore, search engines have been recognized as what was exclusively possessed by leading IT companies or specialized organizations. This study is intended to suggest a search engine with an index structure for making it convenient to effectively search information by distributed crawling massive amount of websites and web-documents in the distributed environment. Search engine suggested in this study has been realized by Hadoop structure for supporting the distributed processing.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.