Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 104
No
1

순위다중패턴매칭문제는 길이가 n인 텍스트 T와 패턴들의 집합 가 주어졌을 때, 에 속한 패턴들과 순위동형인 의 모든 부분문자열들을 찾는 문제이다. 개의 연속적인 문자인 -그램, 에서 가장 짧은 패턴의 길이를 , 가장 긴 패턴의 길이를 , 모든 패턴들의 길이의 합을 이라 할 때, -그램과 계승수체계를 이용하여 시간에 순위다중패턴매칭문제를 해결하는 알고리즘이 제시되었다. 본 논문에서는 이진인코딩된 텍스트의 핑거프린트를 이용하여 시간에 순위다중패턴매칭문제를 해결하는 알고리즘을 제시한다. 또한, 시간에 탐색 과정을 수행하는 병렬 구현 방법을 제시한다. 실험 결과, 가 커질수록 제시한 알고리즘의 기존 알고리즘보다 수행시간이 느려지지만 공간사용량은 적어졌다. 제시한 병렬 구현 방법은 기존 알고리즘보다 공간효율적이면서 수행시간은 유사하다.

순위다중패턴매칭문제는 길이가 인 텍스트 와 패턴들의 집합 가 주어졌을 때, 에 속한 패턴들과 순위동형인 의 모든 부분문자열들을 찾는 문제이다. 개의 연속적인 문자인 -그램, 에서 가장 짧은 패턴의 길이를 , 가장 긴 패턴의 길이를 , 모든 패턴들의 길이의 합을 이라 할 때, -그램과 계승수체계를 이용하여 시간에 순위다중패턴매칭문제를 해결하는 알고리즘이 제시되었다. 본 논문에서는 이진인코딩된 텍스트의 핑거프린트를 이용하여 시간에 순위다중패턴매칭문제를 해결하는 알고리즘을 제시한다. 또한, 시간에 탐색 과정을 수행하는 병렬 구현 방법을 제시한다. 실험 결과, 가 커질수록 제시한 알고리즘의 기존 알고리즘보다 수행시간이 느려지지만 공간사용량은 적어졌다. 제시한 병렬 구현 방법은 기존 알고리즘보다 공간효율적이면서 수행시간은 유사하다.

2

Parallel Implementation of Scrypt: A Study on GPU Acceleration for Password-Based Key Derivation Function

SeongJun Choi, DongCheon Kim, Seog Chung Seo

[Kisti 연계] 한국정보통신학회 Journal of information and communication convergence engineering Vol.22 No.2 2024 pp.98-108

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Scrypt is a password-based key derivation function proposed by Colin Percival in 2009 that has a memory-hard structure. Scrypt has been intentionally designed with a memory-intensive structure to make password cracking using ASICs, GPUs, and similar hardware more difficult. However, in this study, we thoroughly analyzed the operation of Scrypt and proposed strategies to maximize computational parallelism in GPU environments. Through these optimizations, we achieved an outstanding performance improvement of 8284.4% compared with traditional CPU-based Scrypt computations. Moreover, the GPU-optimized implementation presented in this paper outperforms the simple GPU-based Scrypt processing by a significant margin, providing a performance improvement of 204.84% in the RTX3090. These results demonstrate the effectiveness of our proposed approach in harnessing the computational power of GPUs and achieving remarkable performance gains in Scrypt calculations. Our proposed implementation is the first GPU implementation of Scrypt, demonstrating the ability to efficiently crack Scrypt.

3

Parallel implementation of GCM on GPUs

JaeSeok Lee, 김동천, 서석충

[NRF 연계] 한국통신학회 ICT Express Vol.11 No.2 2025.04 pp.310-316

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper presents the first fully parallelized optimization of GCM in a GPU environment. As the era of IoT emerges, a large number of clients communicate with servers, necessitating encrypted communications for security. GCM is a type of AEAD and is currently used in various security protocols, including TLS 1.3 and IPsec. Due to the burden of performing encrypted communication with numerous clients, there has been significant research on utilizing GPUs for high-speed parallel processing in encryption. However, to date, there has been no fully parallelized implementation of GCM on GPUs. This paper proposes a method for parallelizing the challenging GHASH computation in GCM mode, leading to a high-speed parallel implementation of AES-GCM that can exceed 400Gb/s, meeting the requirements of next-generation communication systems. The proposed approach is algorithm-independent and can be applied to any block ciphers. Our implementation on an RTX 4090 demonstrates a performance improvement of *15.38 compared to the maximum processing throughput of a multi-threaded Intel(R) Core(TM) i7-13700K. It also achieves a *17.87 improvement compared to a hybrid CPU?GPU system. Compared to the most researched FPGA implementation for GCM, specifically Xilinx Ultrascale FPGA, our implementation achieves *1.11 better performance. For not only throughput but also power efficiency also better than other implementation, it achieves *3.33 compared to CPU implementation on Intel Xeon E3-1220, also it achieves *21.09 compared to FPGA implementation for AES on Xilinx Virtex 7 series, which is not including full GCM.

4

Parallel implementation of CRYSTALS-Dilithium for effective signing and verification in autonomous driving environment

Seog Chung Seo, SangWoo An

[NRF 연계] 한국통신학회 ICT Express Vol.9 No.1 2023.02 pp.100-105

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In the autonomous driving environment, each vehicle performs numerous signing and verification while sending and receiving BSMs(Basic Security Messages) in real-time. We present an optimized CRYSTALS-Dilithium software which can quickly process each signingand verification in parallel using embedded Graphic Processing Unit. For efficiency, we propose several optimization techniques such asdummy operation-based warp divergence reducing technique, parallel implementation NTT (Number Theoretic Transform)-based polynomialmultiplication, optimization of rejection sampling process using rejection sequence table, and so on. The proposed CRYSTALS-Dilithiumsoftware provides a performance improvement of up to 19.41 times compared to the Dilithium Software on CPUs on NVIDIA Jetson AGXXavier which is an off-the-shelf autonomous vehicle OBU (On-Board Unit).

5

Parallel and Sequential Implementation to Minimize the Time for Data Transmission Using Steiner Trees

Anand, V., Sairam, N.

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.13 No.1 2017 pp.104-113

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, we present an approach to transmit data from the source to the destination through a minimal path (least-cost path) in a computer network of n nodes. The motivation behind our approach is to address the problem of finding a minimal path between the source and destination. From the work we have studied, we found that a Steiner tree with bounded Steiner vertices offers a good solution. A novel algorithm to construct a Steiner tree with vertices and bounded Steiner vertices is proposed in this paper. The algorithm finds a path from each source to each destination at a minimum cost and minimum number of Steiner vertices. We propose both the sequential and parallel versions. We also conducted a comparative study of sequential and parallel versions based on time complexity, which proved that parallel implementation is more efficient than sequential.

6

Novel Parallel Approach for SIFT Algorithm Implementation

Le, Tran Su, Lee, Jong-Soo

[Kisti 연계] 한국정보통신학회 Journal of information and communication convergence engineering Vol.11 No.4 2013 pp.298-306

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The scale invariant feature transform (SIFT) is an effective algorithm used in object recognition, panorama stitching, and image matching. However, due to its complexity, real-time processing is difficult to achieve with current software approaches. The increasing availability of parallel computers makes parallelizing these tasks an attractive approach. This paper proposes a novel parallel approach for SIFT algorithm implementation using a block filtering technique in a Gaussian convolution process on the SIMD Pixel Processor. This implementation fully exposes the available parallelism of the SIFT algorithm process and exploits the processing and input/output capabilities of the processor, which results in a system that can perform real-time image and video compression. We apply this implementation to images and measure the effectiveness of such an approach. Experimental simulation results indicate that the proposed method is capable of real-time applications, and the result of our parallel approach is outstanding in terms of the processing performance.

7

5,400원

There has been a snappy forward advancement in cloud. With the developing measures of affiliations, turning number of affiliations are being utilized as assets in the cloud. Issues arise regarding the certainties of various clients utilizing these assets like round measure of room associations which helps in keeping from the cost. Putting away these associations evades the high cost while posting information during the building of machine orders. This collection of information gives better execution whilst less “far point cost” and ready to make prepared modifications. Cloud can yield better possibilities through net where security weaknesses are restricted. Despite of “wellbeing” being one of the brimming with risk weakness, it adjusts unmistakable connection to go into assumed control handling general condition. This paper uses “MapReduce” library that parallelizes the figuring, and handles perplexed issues like data spread, stack modifying and adjustment to non-basic disappointment. Enormous information, spread across finished many machines, need to parallelize. Moves the data, and gives booking, adjustment to non-basic disappointment. In this paper, a graph of MapReduce programming model along with the applications are explored. The maker has delineated here the work procedure of MapReduce process. Some fundamental issues, like adjustment to non-basic disappointment, are considered in more detail. Without a doubt, even the outline of working of Map Reduce is given.

8

NVIDIA CUDA PTX를 활용한 SPECK, SIMON, SIMECK 병렬 구현

장경배, 김현준, 임세진, 서화정

[Kisti 연계] 한국정보보호학회 정보보호학회논문지 Vol.31 No.3 2021 pp.423-431

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

SPECK과 SIMON은 NSA(National Security Agency)에서 개발한 경량블록암호이며 SIMECK은 SPECK과 SIMON의 장점을 결합하여 만든 새로운 경량블록암호이다. 본 논문에서는 SPECK, SIMON, SIMECK을 사용한 대용량 암호화를 구현 하는데 있어 병렬 처리에 용이한 GPU를 활용하였다. NVIDIA에서 제공하는 CUDA 라이브러리를 활용하였으며 불필요한 연산들을 제거하기 위해 CUDA 어셈블리 언어 PTX를 사용하여 성능을 극대화 하였다. 단순 CPU 구현과 GPU를 활용한 구현 결과를 비교해보았을 때, 더 빠른 속도로 대용량 암호화를 수행할 수 있었다. 또한 GPU 구현 시, C언어를 사용한 구현과 PTX를 사용한 구현을 비교해 보았을 때, PTX 사용 시, 성능이 더욱 증가하는 것을 확인하였다.

SPECK and SIMON are lightweight block ciphers developed by NSA(National Security Agency), and SIMECK is a new lightweight block cipher that combines the advantages of SPECK and SIMON. In this paper, a large-capacity encryption using SPECK, SIMON, and SIMECK is implemented using a GPU with efficient parallel processing. CUDA library provided by NVIDIA was used, and performance was maximized by using CUDA assembly language PTX to eliminate unnecessary operations. When comparing the results of the simple CPU implementation and the implementation using the GPU, it was possible to perform large-scale encryption at a faster speed. In addition, when comparing the implementation using the C language and the implementation using the PTX when implementing the GPU, it was confirmed that the performance increased further when using the PTX.

9

Apache Spark을 이용한 병렬 DNA 시퀀스 지역 정렬 기법 구현

김보성, 김진수, 최도진, 김상수, 송석일

[Kisti 연계] 한국콘텐츠학회 한국콘텐츠학회논문지 Vol.16 No.10 2016 pp.608-616

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Smith-Waterman(SW) 알고리즘은 DNA 시퀀스 분석에서 중요한 연산 중 하나인 지역 정렬을 처리하는 알고리즘이다. SW 알고리즘은 동적 프로그래밍 방법으로 최적의 결과를 도출할 수 있지만 수행시간이 매우 길다는 문제가 있다. 이를 해결하기 위해서 다수의 노드를 이용한 병렬 분산 처리 기반의 SW 알고리즘이 제안되었다. Apache Spark을 기반으로 하는 병렬 분산 DNA 처리 프레임워크인 ADAM에서도 SW 알고리즘을 병렬로 처리하고 있다. 하지만, ADAM의 SW 알고리즘은 Smith-Waterman 이 동적프로그래밍 기법이라는 특성을 고려하지 않고 있어 최대의 성능을 얻지 못하고 있다. 이 논문에서는 ADAM의 병렬 SW 알고리즘을 개선한다. 제안하는 병렬 SW 기법은 두 단계에 걸쳐 실행된다. 첫 번째 단계에서는 지역정렬 대상인 DNA 시퀀스를 다수의 파티션(partition)으로 분할하고 분할된 각 파티션에 대해서 SW 알고리즘을 병렬로 수행한다. 두 번째 단계에서는 파티션 각각에 대해서 독립적으로 SW를 적용함으로써 발생하는 오류를 보완하는 과정을 역시 병렬로 수행한다. 제안하는 병렬 SW 알고리즘은 ADAM을 기반으로 구현하고 기존 ADAM의 SW와 비교를 통해서 성능을 입증한다. 성능 평가 결과 제안하는 병렬 SW 알고리즘이 기존의 SW에 비해서 2배 이상의 좋은 성능을 내는 것을 확인하였다.

The Smith-Watrman (SW) algorithm is a local alignment algorithm which is one of important operations in DNA sequence analysis. The SW algorithm finds the optimal local alignment with respect to the scoring system being used, but it has a problem to demand long execution time. To solve the problem of SW, some methods to perform SW in distributed and parallel manner have been proposed. The ADAM which is a distributed and parallel processing framework for DNA sequence has parallel SW. However, the parallel SW of the ADAM does not consider that the SW is a dynamic programming method, so the parallel SW of the ADAM has the limit of its performance. In this paper, we propose a method to enhance the parallel SW of ADAM. The proposed parallel SW (PSW) is performed in two phases. In the first phase, the PSW splits a DNA sequence into the number of partitions and assigns them to multiple nodes. Then, the original Smith-Waterman algorithm is performed in parallel at each node. In the second phase, the PSW estimates the portion of data sequence that should be recalculated, and the recalculation is performed on the portions in parallel at each node. In the experiment, we compare the proposed PSW to the parallel SW of the ADAM to show the superiority of the PSW.

10

병렬 Shifted Sort 알고리즘의 Warp 단위 CUDA 구현 최적화

박태정

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.18 No.4 2017 pp.739-745

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 GPU 병렬 처리 하드웨어 아키텍처 내 최소 물리적 스레드 실행 단위(warp) 내에서 shifted sort 기반 k개 최근접 이웃 검색 기법을 구현하는 방법을 논의하고 일반적으로 동일한 목적으로 널리 사용되는 GPU 기반 kd-tree 및 CPU 기반 ANN 라이브러리와 비교한 결과를 제시한다. 또한 많은 애플리케이션에서 k가 비교적 작은 값이 필요한 경우가 많다는 사실을 고려해서 k가 warp 내부에서 직접 처리 가능한 2, 4, 8, 16개일 때 최적화에 집중한다. 구현 세부에서는 사용한 CUB 공개 라이브러리의 루프 내 메모리 관리 방법, GPU 하드웨어 직접 명령 적용 방법 등의 최적화 방법을 논의한다. 실험 결과, 제안하는 방법은 기존의 GPU 기반 유사 방법에 비해 데이터 지점과 질의 지점의 개수가 각각 $2^{23}$개 일 때 16배 이상의 빠른 처리 속도를 보였으며 이러한 경향은 처리해야 할 데이터의 크기가 커지면 더욱 더 커지는 것으로 판단된다.

This paper presents and discusses an implementation of the GPU shifted sorting method to find approximate k nearest neighbors which executes within "warp", the minimum execution unit in GPU parallel architecture. Also, this paper presents the comparison results with other two common nearest neighbor searching methods, GPU-based kd-tree and ANN (Approximate Nearest Neighbor) library. The proposed implementation focuses on the cases when k is small, i.e. 2, 4, 8, and 16, which are handled efficiently within warp to consider it is very common for applications to handle small k's. Also, this paper discusses optimization ways to implementation by improving memory management in a loop for the CUB open library and adopting CUDA commands which are supported by GPU hardware. The proposed implementation shows more than 16-fold speed-up against GPU-based other methods in the tests, implying that the improvement would become higher for more larger input data.

11

병렬형 역진자 시스템 제작 및 분리제어

김주호, 박운식, 최재원

[Kisti 연계] 한국정밀공학회 한국정밀공학회지 Vol.17 No.7 2000 pp.162-169

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, we develop a parallel inverted pendulum system that has the characteristics of the strongly coupled dynamics of motion by an elastic spring, the time-variant system parameters, and inherent instability, and so on. Hence, it is possible to approximate some kinds of a physical system into this representative system and to apply the various control theories to this system in order to verie their fidelity and efficiency. For this purpose, an experimental system of the parallel inverted pendulum has been implemented, and a control scheme using the eigenstructure assignment for decoupling control is presented in comparison with the conventional LQR optimal control method. Furthermore, this system can be utilized as a testbed to develop and evaluate new control algorithms through various setups. Finally, in this paper, the results of the experiment are compared with those of numerical simulations for validation.

12

GPU 하드웨어 아키텍처 기반 sub-warp 단위 병렬 프리픽스(prefix) 연산의 정확한 구현

박태정

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.18 No.3 2017 pp.613-619

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 대규모 데이터를 길이가 32 미만인 로컬 세그먼트 단위로 구분하고 이 로컬 세그먼트 내에서 정확한 GPU 병렬 프리픽스(prefix) 연산 결과를 출력하는 CUDA (Compute Unified Device Architecture) 코드를 제시한다. 이미 Mark Harris와 Michael Garland가 이러한 목적을 수행하기 위한 CUDA 코드를 이미 발표한 바 있으나 본 논문에서는 로컬 세그먼트의 길이가 32 미만일 때 기존 코드의 결과가 정확하지 않다는 사실을 살펴 보고 그 원인을 논의한 후, 정확한 결과를 출력하는 코드를 제안한다. 본 논문에서 다루는 로컬 세그먼트 단위의 병렬 프리픽스 연산은 최인접 요소 탐색(k-nearest neighbor search) 등은 물론 다양한 대규모 병렬 처리 알고리즘을 구성하는 기본 연산으로 활용 가능하다.

This paper presents a CUDA (Compute Unified Device Architecture) code to achieve correct GPU parallel segmented prefix operation results with less than 32 segment length for large data arrays. Mark Harris and Michael Garland had published CUDA code to address the tasks. This paper shows that their code does not generate correct results when the local segment length is less than 32, discusses the cause of the problem, and presents a CUDA code that generates correct results. The segmented parallel prefix operation presented in this paper can be applied as a building block to various large parallel processing algorithms including the k-nearest neighbor search problems.

13

공유 메모리 병렬 컴퓨터 환경에서 Bitonic Sorting 알고리즘 설계와 효율적인 통신의 구현

이재동, 권경희, 박용범

[Kisti 연계] 한국정보처리학회 정보처리학회논문지 Vol.4 No.11 1997 pp.2690-2700

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구에서는 공유메모리 병렬 컴퓨터 환경에서 N개의 key를 $O(log^2N)$ 시간에 정렬 할 수 있는 병 알고리즘인 SARED-MEMORY-BS와 REDUCED-BS를 설계하였다. REDICED-BS 알고리즘은 각각 로세서에 있는 local memory를 효율적으로 사용할 수 있도록 제안한 parity전략을 사용하였다. 각각의 프로세서에 있는 local memo교를 효율적으로 사용함으로써 REDUCED-BS 알고리즘은 SHARED-MEMORY-BS 알고리즘에 비햐여 통신의 빈도수가 약 1/2정도 감소된 것으로 나타났다. 결과적으로 REDUCED-BS 알고리즘은 병렬 정렬시 통신을 감소시킴으로써 컴퓨터의 사용 효율을 향상시킬 수 있다.

This paper presents parallel sorting algorithm, SHARED-MEMORY-BS and REDUCED-BS, which are implemented on shared-memory parallel computers. These algorithm sort N keys in $O(log^2N)$ time. REDUCED-BS users a parity strategy which gives an idea for the efficient usage of the local memory associated with each processor. By taking advantage of the local memory associated with each processor, the communication of REDUCED-BS is decreased by approximately half that of SHARED-MEMORY-BS. On the basis of alleviating the communication, the algorithm REDUCED-BS results in a significant improvement of performance.

14

Grain Stream Cipher 기반의 Light Weight 병렬 알고리즘 연산 방식으로 구현하였다. 제안된 병렬 연산 방식의 알고리즘을 FPGA로 구현하여 동작을 검증하였고, 0.18um CMOS 공정으로 설계하여 UHF RFID 시스템에 적용 가능성을 타진하였다. 제안된 병령 연산 방식의 보안 알고리즘의 설계 및 ASIC 구현 결과를 실제 UHF RFID 표준 프로토콜에 적용하여 현재의 보안 문제 및 복제 방지 기능이 내장된 태그 칩 제작에 활용할 수 있음을 실험을 통해 확인하였다. 0.18um CMOS 공정의 경우에 Radix-16의 경우에도 ASIC 구현 결과 면적 및 전력소비에서도 UHF RFID 시스템에 적용 가능함을 확인하였다.

We have implemented a parallel operation of the Light Weight algorithm based on Grain Stream Cipher authentication algorithm for UHF RFID systems. We also verified parallel operation algorithm by Verilog coding and ASIC implementation of the Radix-1, 2, 4, 8 and 16 structures using a 0.18um CMOS process. The implemented results show that proposed parallel architecture can be used for design of UHF RFID tag for anti-cloning protection or authentication. We have verified that chip area and power consumption of the Radix-16 designed by 0.18um CMOS process can be applied to the UHF RFID tag design.

15

OpenCL을 이용한 JPEG2000 4K 초고화질 영상처리의 병렬고속화 구현 KCI 등재후보

박대승, 김정길

한국위성정보통신학회 한국위성정보통신학회논문지 제10권 제1호 2015.02 pp.1-5

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

멀티미디어 기술의 급속한 발전과 사용자의 대형 화면에 대한 선호도가 높아지는 가운데 새로운 영상 압축 기술인 HEVC(HighEfficiency Video Coding) 고화질 영상 압축 표준을 탄생시켰으며, 그 결과 기존의 HD급 영상보다 4배 이상, 16배까지 선명한 초고화질UHD(Ultra High Definition) 영상 서비스가 새롭게 주목받고 있다. 또한 JPEG 2000 압축도 기존 처리되던 픽셀 이미지를 넘어 초고화질 해상도 이미지(4K : 3,840 x 2,160 또는 8K : 7680 x 4320)를 처리 지원을 하고 있다. 따라서 초고화질 이미지의 획득 및 저장을위해서는 고속의 처리 기술이 필요하다. 이에 본 논문은 초고화질 해상도 이미지의 고속 처리를 위한 병렬처리 기술에 대한 연구를위하여, JPEG 2000의 처리 과정을 살펴보고 전처리 단계인 색공간 변환 알고리즘 적용을 위하여 GPU환경에서 병렬 컴퓨팅을 통해처리속도를 향상시키는 방법을 제안한다. 병렬화한 알고리즘의 구현은 OpenCL(Open Computing Language)을 이용하였다. 실험 결과사용자 정의 쓰레드 기반 고속 처리와 비교하여 초고화질 해상도 이미지(UHD 4K : 3,840 x 2,160)를 기준으로 최대 5배의 성능 향상의결과를 보여주었다.

With the help of fast growing multimedia technology and high preference for users of large screens, the newest video codingstandard, HEVC (High Efficiency Video Coding) high-quality video compression), has been introduced. Therefore, the highdefinition image services which are four times more clear than conventional HD video, are getting popular. JPEG 2000 alsohas stated to support 4K and 8K UHD. As a result, it requires fast processing technology to read and write UHD images. This paper introduces a study on fast parallel processing technology for UHD images. For this purpose, first, JPEG 2000is reviewed and a GPU based parallel implementation is proposed for a preprocessing of color conversion stage. The parallelledalgorithm is implemented with OpenCL (Open Computing Language). The simulation results show that the proposed methodshows 5 times performance improvements on processing speed for 4K UHD over the method using threads.

16

병렬 객체지향 프로그래밍을 위한 시각 환경의 설계 및 구현

최숙영

[Kisti 연계] 한국정보처리학회 정보처리학회논문지 Vol.6 No.2 1999 pp.485-496

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

병렬 프로그래밍은 프로세스간의 통신과 동기화 문제, 병렬 시스템의 구성 형태등을 고려해야 하기 때문에 순차 프로그래밍에 ? 많은 노력을 필요로 한다. 효율적인 병렬 프로그램을 작성하기 위해서는 사용자와 컴파일러간의 상호 지원이 이루어져야 한다. 이러한 관점에서 본 연구는 선행 연구로써 병렬 객체지향 표기언어 POOSL을 개발하였다. 그러나, 사용자 입장에서 볼 때 병렬 프로그램을 작성하기 위해 POOSL의 문법 구조를 염두에 두고 텍스트 중심의 프로그램을 작성한다면 여전히 부담스러운 작업이 될 것이다. 사용자에게 보다 편리함을 제공하기 위해서는 텍스트보다는 시각적인 프로그래밍 환경이 더욱 효율적이고 바람직할 것이다. 따라서, 본 논문에서는 POOSL을 기초로 하여 사용자가 좀더 쉽고, 편리하게 병렬 프로그래밍 할 수 있는 시각 환경으로써 VEPO(Visual Environment for Parallel Object-Oriented Programing)를 제안하고 있다. 본 논문의 목적은 사용자가 병렬 프로그램을 작성하는데 있어 문제에 내재된 병렬성을 객체지향 개념에 입각하여 시각적으로 자연스럽게 표현하도록 하고, 병렬 프로그램 개발에 관련된 과정들을 하나의 환경을 통합시킴으로써 편리한 프로그램 환경을 제공하는 것이다. 본 연구에서 제안하고 있는 VEPO는 병렬 프로그램을 개발하는데 필요한 기본적인 단계들로써 프로그램 기술 단계, 실행 단계, 실행 과정의 시각화등을 지원하고 있으며, 시각 프로그래밍의 장점을 충분히 살릴 수 있도록 여러 개념들이 지원되고 있다. 특히, 병렬 프로그램에서 복잡하고 까다로운 통신과 동기화에 관련된 코드 등은 번역 과정에서 여러 개념들이 생성되도록, 함으로써 사용자로 하여금 병렬 프로그램을 작성하는데 따르는 부담감을 줄 일 수 있도록 한다. 본 시스템은 PC를 호스트로 연결한 트랜스퓨터들로 구성된 병렬 컴퓨터 MC-3에서 구현되었다. VEPO 그래픽 사용자 인터페이스는 Visual C++로 구현되었고, VEPO에서 작성된 시각 프로그램은 Inmos C 코드로 번역되어 MC-3에서 수행된다.

Comparing with sequential programming, parallel programming has additional complexity due to the consideration of parallelism, communication and synchronization of processes. A synergism between users and compliers should be established, each assisting the other to produce high quality parallel programs. On the above underlying philosophy, we developed a parallel Object-Oriented specification language, POOSL, as preliminary works. However, it is still likely to hard for users to write parallel program because users have to consider grammar of POOSL and to write text-based parallel program. It would be more desirable to provide users wit visual environment for effective parallel programming. Therefore, we propose a visual programming environment. VEPO(Visual environment for Parallel Object-Oriented Programming), based on POOSL in order that users can develop parallel programs more easily and conveniently. It aims at supporting a programming environment in which users can represent their programs more naturally and visually I parallel manner with object-oriented concept and essential steps during parallel program development such as program specification, compilation, execution and animation of execution are integrated. VEPO has useful features for parallel processing. Especially, complicated parallel codes for synchronization and communication of processes are automatically generated in the translation phase, so users can be relieved of writing error-prone parallel codes. The system is targeted to the transputer-based parallel system, MC-3. The graphic user interface of VEPO was implemented using Visual C++. Visual programs descirbed on VEPO are translated into Inmos C and executed on MC-3.

17

병렬 프로세서 연산 기법을 사용한 완전 음역, 다성분 저류층 시뮬레이터에서의 물 주입정 및 석유 생산정 모사

한충용, 카미 세퍼누리

[NRF 연계] 한국자원공학회 한국자원공학회지 Vol.41 No.3 2004.06 pp.242-252

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

병렬 프로세서 연산이 가능한 완전 음역, 다성분 저류층 시뮬레이터에서 물 주입정 및 석유 생산정을 모사할 수 있는 방법을 제시하였다. 물 주입정에서는 물을 정량 또는 정압으로 주입할 수 있도록 하였고, 석유 생산정에서는 석유를 정량으로 생산할 수 있게 하였다. 이를 위하여 유정으로 출입하는 각 성분의 시간당 몰량을 물질 균형 방정식 잔차 및 자코비안 행렬을 구성할 때 고려해주었다. 개발된 모델을 기존 다성분 저류층 시뮬레이터인 ECLIPSE 300 및 UTCOMP를 이용하여 검증하였으며, 병렬 프로세서 연산 결과가 동일한 유정 입력 자료에 대한 단일 프로세서 연산 결과와 잘 일치함을 확인할 수 있었다.

Water injection and oil production wells are implemented in parallel, fully implicit, compositional reservoir simulator. Water can be injected at constant volumetric rate or pressure and oil can be produced at constant volumetric rate. These well constraints are simulated by taking molar rate of each component injected or produced in a well into account when constructing residual and Jacobians of material balance equation of the component. The implemented model is verified using other compositional reservoir simulators, ECLIPSE 300 and UTCOMP. Results from parallel processors runs are in good agreement with those from a single processor run for a same well input data.

18

구형 3자유도 병렬 메커니즘의 기구학 해석 및 구현

이석희, 김희국, 오세민, 소병록, 이병주

[Kisti 연계] 한국정밀공학회 한국정밀공학회지 Vol.22 No.11 2005 pp.72-81

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

A new spherical-type 3-degree-of-freedom parallel mechanism consisting of a two degree-of-freedom parallel module and a serial module is proposed. Two alternative designs for the serial sub-chain are suggested and compared. The first design employs RU joint arrangement for the serial sub chain structure. The second design incorporates a gear chain to drive the distal revolute joint of the serial sub-chain from the base platform of the mechanism. This modification significantly improves kinematic characteristics of the mechanism within its workspace. Firstly, the closed-form solutions of both the forward and the reverse position analysis are derived. Secondly, the first-order kinematic model with respect to three inputs which are located at the base is derived. Thirdly, it is confirmed through simulation that the modified mechanism has much more improved isotropic characteristic throughout the workspace of the mechanism. Lastly, the proposed mechanism is implemented to verify the results from this analysis.

19

한국어학 전공자를 위한 전문 용어 병렬 데이터베이스 구축 KCI 등재후보

박병선

고려대학교 언어정보연구소 언어정보 제27호 2018.09 pp.30-44

※ 기관로그인 시 무료 이용이 가능합니다.

4,800원

The aim of this research is two-fold. I first attempt to compare linguistic terminologies in Chinese and Korean, and then implement a Korean-Chinese parallel database, which contain linguistic terms in Korean juxtaposed with their corresponding Chinese expressions. To do so, I extracted major terminologies used in the field of Korean linguistics based on their frequency and identified how they are translated in Chinese. I go one step further to demonstrate a three-way comparison: comparing how two different countries, China and Taiwan, use linguistic terminologies against their Korean counterparts. The outcome of this result has practical import; it can be used for researchers and students alike in Greater China, who are interested in linguistics and/or the Korean language as it is. I also discuss the demand of building a three-language- database, Korean-English-Chinese, to enhance the researcher’s understanding of terminologies in these languages, which often can be hindrance for researchers who are not familiar with two of the three aforementioned languages.

20

4,000원

This paper presents a parallel kd-tree traversal algorithm based on the parallel binary radix tree construction scheme proposed by Tero Karras in 2012. In his paper, Tero Karras proposed parallel tree construction algorithm which can maximize the utilization of GPU threads, but implementation and analysis of kd-tree are not fully discussed. This paper aims to fill the gap for the specific kd-tree cases. As an application for the kd-tree traversal method proposed in this paper, the proposed method has been implemented with NVIDIA’s CUDA framework and tested on NVIDIA’s realtime raytracing library, OptiX. As a result, the proposed method can construct tree structures within the requirement for realtime process, but still needs specialized spatial caching data structure like the “primitive tree” for highly detailed meshes to handle spatial queries as fast as to visualize implicit surfaces under OptiX framework. However, the proposed method can be applied to dynamic collision detection and manage scene and object information for game applications thanks to its fast tree construction process. Also, realtime raytracing can be applied to game applications based on this study.

 
1 2 3 4 5
페이지 저장