Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 51
No
1

A Parallel-Architecture Approach to Preposition-Noun Compounds in English KCI 등재

Kyungchul Chang

한국언어연구학회 언어학연구 제23권 3호 2018.12 pp.1-19

※ 기관로그인 시 무료 이용이 가능합니다.

5,400원

This article proposes a parallel-architecture approach to the class of preposition-noun (P-N) compounds in English such as downstairs. In this approach, it is argued that P-N compounds are not words but phrases, more specifically not adverbs but prepositional phrases, particularly when used with some verbs of motion such as He brought the suitcase downstairs. They are also differentiated from so-called particles like down and prepositional phrases like down the stairs, though sharing the same head down. The article argues for a class into which all these three cases can be unified. The resulting unified class is basically a preposition and can be divided into intransitive and transitive types. The article puts forward a sub-categorisation of the transitive preposition into two types; one is lexical, and the other phrasal. A classification of P-N compounds as the lexical type is suggested together with a lexical PP analysis using the notion of representational modularity.

2

GPU 하드웨어 아키텍처 기반 sub-warp 단위 병렬 프리픽스(prefix) 연산의 정확한 구현

박태정

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.18 No.3 2017 pp.613-619

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 대규모 데이터를 길이가 32 미만인 로컬 세그먼트 단위로 구분하고 이 로컬 세그먼트 내에서 정확한 GPU 병렬 프리픽스(prefix) 연산 결과를 출력하는 CUDA (Compute Unified Device Architecture) 코드를 제시한다. 이미 Mark Harris와 Michael Garland가 이러한 목적을 수행하기 위한 CUDA 코드를 이미 발표한 바 있으나 본 논문에서는 로컬 세그먼트의 길이가 32 미만일 때 기존 코드의 결과가 정확하지 않다는 사실을 살펴 보고 그 원인을 논의한 후, 정확한 결과를 출력하는 코드를 제안한다. 본 논문에서 다루는 로컬 세그먼트 단위의 병렬 프리픽스 연산은 최인접 요소 탐색(k-nearest neighbor search) 등은 물론 다양한 대규모 병렬 처리 알고리즘을 구성하는 기본 연산으로 활용 가능하다.

This paper presents a CUDA (Compute Unified Device Architecture) code to achieve correct GPU parallel segmented prefix operation results with less than 32 segment length for large data arrays. Mark Harris and Michael Garland had published CUDA code to address the tasks. This paper shows that their code does not generate correct results when the local segment length is less than 32, discusses the cause of the problem, and presents a CUDA code that generates correct results. The segmented parallel prefix operation presented in this paper can be applied as a building block to various large parallel processing algorithms including the k-nearest neighbor search problems.

3

병렬 출력을 갖는 LFSR 구조를 적용한 HIGHT 프로세서 설계 KCI 등재

이제훈, 김상춘

한국융합보안학회 융합보안논문지 제15권 제2호 2015.03 pp.81-89

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

HIGHT (HIght security and light weighHT) 암호는 기밀성을 요구하는 네트워크 환경에서 사용할 수 있도록 국내 에서 개발된 저전력 경량화 64비트 블록 암호 알고리즘이다. 본 논문은 키스케쥴러에 사용되는 LFSR 및 역 LFSR의 4 개의 병렬 출력을 허용할 수 있는 구조를 제안하였다. 또한, 각 라운드 연산에 필요한 4개의 서브키를 동일 클럭 사이 클에 생성할 수 있도록 구성하였다. 따라서, 전체 HIGHT 암호 프로세서가 단일 시스템 클럭에 의해 제어할 수 있다. VHDL을 이용하여 회로를 합성한 후, 검증한 결과 제안된 키 스케쥴러의 회로 크기는 기존 키 스케쥴러에 비해 9% 감 소되었다.

HIGHT is an 64-bit block cipher, which is suitable for low power and ultra-light implementation that are used in the network that needs the consideration of security aspects. This paper presents a key scheduler that employs the presented LFSR and reverse LFSR that can generate four outputs simultaneously. In addition, we construct new key scheduler that generates 4 subkey bytes at a clock since each round block requires 4 subkey bytes at a time. Thus, the entire HIGHT processor can be controlled by single system clock with regular control mechanism. We synthesize the HIGHT processor using the VHDL. From the synthesis results, the logic size of the presented key scheduler can be reduced as 9% compared to the counterpart that is employed in the conventional HIGHT processor.

4

4,000원

5

Parallel Architecture for High-Speed Block Cipher, HIGHT SCOPUS

Je-Hoon Lee, Duk-Gyu Lim

보안공학연구지원센터(IJSIA) International Journal of Security and Its Applications Vol.8 No.2 2014.03 pp.59-66

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

This paper presents the implementation of high-speed block, cipher, HIGHT. The proposed architecture employs parallel architecture to enhance throughput. In addition, it shares key scheduling block for encryption and decryption to reduce hardware complexity. It also introduces an efficient protocol applicable to RFID systems, implementing the HIGHT block cipher algorithm. The new HIGHT structure yields a size small enough to afford tag applications and twice as high performance with respect to conventional HIGHT implementation. The proposed protocol overcomes the security vulnerability of RFID tags, and reduces energy consumption per transaction by sharing key generation.

6

A Novel Parallel Architecture Design of Information Retrieval System for Scientific Papers SCOPUS

Aziz Murtazaev, Sanggil Kang, Sangyoon Oh

보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.6 No.2 2012.04 pp.107-112

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Indexing allows converting raw document collection into easily searchable representation. Bigger scale indexing poses some challenges such as how to distribute indexing computation efficiently on a cluster of nodes. MapReduce framework can be an effective tool for parallelizing such tasks as inverted index construction. We propose SciPDFindexer, distributed information retrieval system for scientific articles in PDF. For given large collection of scientific articles in PDF our system parses and extracts metadata from articles, and then indexes extracted content using our proposed scheme. Our contribution is the design of distributed IR system and indexing scheme that improve the overall indexing performance.

7

High-Speed Parallel Architecture of the Whirlpool Hash Function

Deen Kotturi, Seong-Moo Yoo

보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology vol.7 2009.06 pp.21-26

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Whirlpool hash function should be capable of processing the input data streams at high speeds. We propose a fully synchronous parallel pipelined architecture with ten stages of pipelining between rounds, and internal pipelining within each round stage. The proposed architecture can tremendously improve the performance of Whirlpool hash function. Our final implementation can encrypt continuous bit streams seamlessly, achieving a throughput of 56.89 G bps, which is considerably high compared to existing implementations in literature.

8

An advanced Parallel FPGA Architecture for Bi-directional Motion Estimation

Yangfan Huang, Minjun Deng, Donglian Li, Zhenzhen Li, Mingyan Yu, Cailan Zeng, Yu Zhang, Zhuo Chen, Pengcheng Cao, Ran Liu

보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.8 No.9 2015.09 pp.235-244

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Motion estimation (ME) and motion compensation (MC) are the key elements for frame rate up-conversion (FRUC) system. Fast and accurate motion estimation is the premise of high quality motion compensation. Unlike conventional unidirectional motion estimation which brings holes, overlaps and blocking artifacts, the bi-directional motion estimation does not produce any overlapped pixel or hole in the interpolated frames. As a result, the bi-directional motion estimation has better performance than conventional unidirectional motion estimation. This paper presents an efficient FPGA architecture targeting bi-directional motion estimation hardware implementation. This proposed architecture can achieve real-time processing for 1280x720@60Hz under 200MHz operating frequency. The design is described in Verilog HDL, verified in Virtex5 FPGA platform. Experimental results show that the proposed architecture has high performance and low cost for bi-directional motion estimation algorithm. This architecture can be used for video post-processing system.

9

A Parallel Algorithm of Multiple String Matching Based on Set-Partition in Multi-core Architecture SCOPUS

Jiahui Liu, Fangzhou Li, Guanglu Sun

보안공학연구지원센터(IJSIA) International Journal of Security and Its Applications Vol.10 No.4 2016.04 pp.267-278

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

With the coming of the big data era, the data processing in large scale comes out with a new challenge. However, string matching still plays an important role in the network security and information retrieval fields, because of the large size of pattern set with the overhead of memory and access memory time. Improving the string matching algorithm to adapt to the large scale tasks is desirable and meaningful. In this paper, we present and implement a parallel algorithm of multiple string matching based on multi-core platform. In addition, this work focuses on the partition of pattern set by using genetic algorithm through the internal relation of the patterns to reduce the memory overhead and execution performance. Compared with the classical ones, our experiments on both high and low hit-rate data demonstrate that the performance of algorithm enhances about on average by 20%-40% in general. Besides, the proposed algorithm reduces the memory cost on average by 4%-20%.

10

A 32nm EGPU Parallel Multiprocessor Based on Co-issue and Multi-Dimensional Parallelism Architecture SCOPUS

Yang Wang, Li Zhou, Tao Sun, Yanhu Chen, Jia Wang, Yuanzhi Zhang, Yuanyuan Gao

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.10 No.5 2015.05 pp.343-354

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

In this paper, a Parallel Multiprocessor (PM) based on SIMT (Single Instruction and Multiple Threads) architecture is proposed. With co-issue architecture and multi-dimensional parallelism implemented in high-effective PM, Embedded Graphics Processing Unit (EGPU) provides great performance for various situations, such as general purpose computing, 3D scene rendering, and graphics processing. Application programs are departed into separated threads. Allocated by Thread Processing Unit (TPU), separated threads can be executed in parallel. Parallelism in different hierarchy and dimension are implemented by Multi-Dimensional Parallelism Processor (MD-PP), which has made a proper trade-off between performance and cost. Additionally, PM improves the hardware occupancy with its co-issue architecture and internal bus accessing mechanism to meet the demand of processing capability. Its unified shading architecture also helps to hide processing latency. PM can execute 4 basic operations in the best case and 2 in the worst case within each clock cycle. With 32nm process technology and 200MHz clock frequency, PM’s area is about 5104494um2, power consumption is about 101.838mW, and it can process nearly 28M vertices or fragments in average. Experimental results show that the MD-PP based PM can process data with high performance and get a balance between efficiency and hardware consumption simultaneously.

11

Time Complexity Measurement on CUDA-based GPU Parallel Architecture of Morphology Operation

Izmantoko, Yonny S., Choi, Heung-Kook

[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.16 No.4 2013 pp.444-452

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Operation time of a function or procedure is a thing that always needs to be optimized. Parallelizing the operation is the general method to reduce the operation time of the function. One of the most powerful parallelizing methods is using GPU. In image processing field, one of the most commonly used operations is morphology operation. Three types of morphology operations kernel, na$\ddot{i}$ve, global and shared, are presented in this paper. All kernels are made using CUDA and work parallel on GPU. Four morphology operations (erosion, dilation, opening, and closing) using square structuring element are tested on MRI images with different size to measure the speedup of the GPU implementation over CPU implementation. The results show that the speedup of dilation is similar for all kernels. However, on erosion, opening, and closing, shared kernel works faster than other kernels.

12

K-Nearest Neighbor Associative Memory with Reconfigurable Word-Parallel Architecture

An, Fengwei, Mihara, Keisuke, Yamasaki, Shogo, Chen, Lei, Mattausch, Hans Jurgen

[Kisti 연계] 대한전자공학회 Journal of semiconductor technology and science Vol.16 No.4 2016 pp.405-414

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

IC-implementations provide high performance for solving the high computational cost of pattern matching but have relative low flexibility for satisfying different applications. In this paper, we report an associative memory architecture for k nearest neighbor (KNN) search, which is one of the most basic algorithms in pattern matching. The designed architecture features reconfigurable vector-component parallelism enabled by programmable switching circuits between vector components, and a dedicated majority vote circuit. In addition, the main time-consuming part of KNN is solved by a clock mapping concept based weighted frequency dividers that drastically reduce the in principle exponential increase of the worst-case search-clock number with the bit width of vector components to only a linear increase. A test chip in 180 nm CMOS technology, which has 32 rows, 8 parallel 8-bit vector-components in each row, consumes altogether in peak 61.4 mW and only 11.9 mW for nearest squared Euclidean distance search (at 45.58 MHz and 1.8 V).

13

Retina-Motivated CMOS Vision Chip Based on Column Parallel Architecture and Switch-Selective Resistive Network

Kong, Jae-Sung, Hyun, Hyo-Young, Seo, Sang-Ho, Shin, Jang-Kyoo

[Kisti 연계] 한국전자통신연구원 ETRI journal Vol.30 No.6 2008 pp.783-789

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

A bio-inspired vision chip for edge detection was fabricated using 0.35 ${\mu}m$ double-poly four-metal complementary metal-oxide-semiconductor technology. It mimics the edge detection mechanism of a biological retina. This type of vision chip offer several advantages including compact size, high speed, and dense system integration. Low resolution and relatively high power consumption are common limitations of these chips because of their complex circuit structure. We have tried to overcome these problems by rearranging and simplifying their circuits. A vision chip of $160{\times}120$ pixels has been fabricated in $5{\times}5\;mm^2$ silicon die. It shows less than 10 mW of power consumption.

14

The Idiomatic Give-Get Alternation in English: A Parallel-Architecture Approach

장경철

[NRF 연계] 현대문법학회 현대문법연구 Vol.124 2024.12 pp.55-75

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In English there are idioms that alternate between give and get, as in They gave him the boot (‘They dismissed him’) and He got the boot (‘He was/became dismissed’). They are not isolated but related in terms of the properties they share, such as the same noun phrase(NP) the boot and the verbal meaning ‘to dismiss/to become dismissed’. This paper argues that the two alternantes are better characterised as the level of predicate-argument structure which mediates between semantics and syntax. It suggests that the clause-final NP combines with either of the verbs to form a predicate. The resulting complex predicate is either transitive or intransitive, depending on whether the internal argument is present or absent. Though the former appears syntactically more complex than the latter, both are basically semantic and grammatical units that are internally complex. The paper also offers a unified analysis of the two complex predicate types using the framework of parallel architecture of grammar.

15

An FPGA-based Parallel Hardware Architecture for Real-time Eye Detection

Kim, Dong-Kyun, Jung, Jun-Hee, Nguyen, Thuy Tuong, Kim, Dai-Jin, Kim, Mun-Sang, Kwon, Key-Ho, Jeon, Jae-Wook

[Kisti 연계] 대한전자공학회 Journal of semiconductor technology and science Vol.12 No.2 2012 pp.150-161

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Eye detection is widely used in applications, such as face recognition, driver behavior analysis, and human-computer interaction. However, it is difficult to achieve real-time performance with software-based eye detection in an embedded environment. In this paper, we propose a parallel hardware architecture for real-time eye detection. We use the AdaBoost algorithm with modified census transform(MCT) to detect eyes on a face image. We parallelize part of the algorithm to speed up processing. Several downscaled pyramid images of the eye candidate region are generated in parallel using the input face image. We can detect the left and the right eye simultaneously using these downscaled images. The sequential data processing bottleneck caused by repetitive operation is removed by employing a pipelined parallel architecture. The proposed architecture is designed using Verilog HDL and implemented on a Virtex-5 FPGA for prototyping and evaluation. The proposed system can detect eyes within 0.15 ms in a VGA image.

16

MediaFrame: Parallel multimedia system architecture through HTTP redirection

김성기, 한상영

[Kisti 연계] 한국정보처리학회 정보처리학회논문지 A Vol.a14 No.1 2007 pp.15-24

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

한 대의 비디오 서버가 확장성, 용량, 고장 감내(Fault-tolerance), 비용 효율성(Cost-efficiency)에서 한계들을 드러냄에 따라 해결책들이 나타났다. 그러나, 이러한 해결책들도 각각의 문제점들을 지닌다. 우리는 이러한 문제점들을 해결하고 다양한 비디오 서버들을 활용하기 위해, HTTP 레벨 리디렉션(Redirection)을 통하여 이질적인 퍼스널 컴퓨터(Personal Computer), 운영 체제(Operating System) 비디오 서버들로 내용에 따른 라우팅(Routing)을 지원하는 병렬 멀티미디어 시스템 구조를 디자인하였다. 또한 프로토타입을 개발하였으며 서로 다른 비디오서버 제품들을 프로토타입에 추가하였고 오버헤드를 측정하였다.

As a single video server exposes its limitation in scalability, capability, fault-tolerance, and cost-efficiency, solutions of this limitation emerge. However, these solutions have their own problems that will be discussed in this paper. To solve these problems and exploit various video silvers, we designed a parallel multimedia system architecture that supported a content-aware routing to heterogeneous personal computer (PC), operating system (OS), video servers through a HTTP-level redirection. We also developed a prototype, added different video servers into the prototype, and measured its overheads.

17

Massively Parallel Processor based on Systolic Architecture for High-Performance Computation of Difference Schemes

Sano, Kentaro, Iizuka, Taknnori, Yamamoto, Satoru

[Kisti 연계] 한국전산유체공학회 한국전산유체공학회 학술대회논문집 2006 pp.174-177

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

18

Transputer-based Pyramidal Parallel Array Computer(TPPAC) architecture (Prelimineary Version)

정창성, 정철환

[Kisti 연계] 대한전기학회 대한전기학회 학술대회논문집 1988 pp.647-650

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper proposes and sketches out a new parallel architecture of transputer-based pyramidal parallel array computer (TPPAC) used to process computationally intensive problems for geometric processing applications such as computer vision, image processing etc. It explores how efficiently the pyramid computer architecture is designed using transputer chips, and poses a new interconnection scheme for TPPAC without using additional transputers.

19

Exhibition Analysis of Skin + Bones : Parallel Practices in Fashion and Architecture

Ahn, Ji-Won

[Kisti 연계] 한국복식학회 International journal of costume Vol.9 No.1 2009 pp.59-69

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The exhibition would have benefited from a more sustained examination of the contemporary meanings and historic meanings behind fashion ideas and architecture as a communication vehicles, which reflect public preferences as an art or design. Both are based on structure, shape, and th ornament basic necessities. Skin+Bones pools contemporary exemplars and cultural capital - providing resources, creating the opportunity for new hybrids, and advancements for fashionistas who are much more interested in fashion. The overall aim of this research is to understand both fashion and architecture by analyzing exhibition and interpreting the meaning of objects that have been shown and studying the problems and obstacles to be overcome in presenting a significant meaning of fashion and architecture.

20

A Study on Efficient Executions of MPI Parallel Programs in Memory-Centric Computer Architecture

Lee, Je-Man, Lee, Seung-Chul, Shin, Dongha

[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.25 No.1 2020 pp.1-11

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 프로세서 중심 컴퓨터 구조에서 개발된 MPI 병렬 프로그램을 수정하지 않고 메모리 중심 컴퓨터 구조에서 더 효율적으로 수행시키는 기술을 제안한다. 본 연구에서 제안하는 기술은 메모리 중심 컴퓨터 구조가 가지는 빠른 대용량 공유 메모리 특징을 이용하여 MPI 표준 라이브러리 함수가 수행하는 네트워크 통신을 통한 느린 데이터 전달을 공유 메모리를 통한 빠른 데이터 전달로 대체하여 효율성을 얻는다. 본 연구에서 제안한 기술은 두 개의 프로그램에 구현되었다. 첫 번째 프로그램은 MC-MPI-LIB라고 불리는 수정된 MPI 라이브러리인데 이는 기존 MPI 표준 라이브러리 함수의 의미를 유지하면서 메모리 중심 컴퓨터 구조에서 더 효율적으로 수행한다. 두 번째 프로그램은 MC-MPI-SIM이라고 불리는 시뮬레이션 프로그램인데 이는 프로세서 중심 컴퓨터 구조 상에서 메모리 중심 컴퓨터 구조의 수행을 시뮬레이션한다. 본 논문에서 제안한 기술은 도커 가상화 상에서 구현된 분산 시스템 환경에서 개발하고 시험하였다. 다수의 MPI 병렬 프로그램을 이용하여 제안한 기술의 성능을 측정한 결과 메모리 중심 컴퓨터 구조에서 더 높은 성능으로 수행 가능함을 보였으며, 특히 통신 오버헤드 비율이 높은 MPI 병렬 프로그램의 경우 매우 높은 성능으로 수행 가능하다는 점을 확인하였다.

In this paper, we present a technique that executes MPI parallel programs, that are developed on processor-centric computer architecture, more efficiently on memory-centric computer architecture without program modification. The technique we present here improves performance by replacing low-speed data communication over the network of MPI library functions with high-speed data communication using the property called fast large shared memory of memory-centric computer architecture. The technique we present in the paper is implemented in two programs. The first program is a modified MPI library called MC-MPI-LIB that runs MPI parallel programs more efficiently on memory-centric computer architecture preserving the semantics of MPI library functions. The second program is a simulation program called MC-MPI-SIM that simulates the performance of memory-centric computer architecture on processor-centric computer architecture. We developed and tested the programs on distributed systems environment deployed on Docker based virtualization. We analyzed the performance of several MPI parallel programs and showed that we achieved better performance on memory-centric computer architecture. Especially we could see very high performance on the MPI parallel programs with high communication overhead.

 
1 2 3
페이지 저장