년 - 년
퀀텀 에스프레소와 제온 파이 프로세서의 융합을 이용한 분산컴퓨팅 성능에 대한 연구 KCI 등재
한국융합학회 한국융합학회논문지 제7권 제5호 2016.10 pp.15-21
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 프로세서의 집적도는 급속도로 발전하고 있으나 클락 스피드는 증가하지 않는 대신에 프로세서 내의 코어 수가 늘어나고 있는 실정으로 프로그래밍 속도 향상을 위한 방법에 대한 연구가 필수적이라 할 수 있다. 이에 본 논문에서는 현재 연산 가속화를 위해 사용되는 매니 코어 프로세서의 대표적인 인텔 제온 파이의 성능 분석을 위하여 퀀텀 에스프레소를 활용하였다. 또한 제온 파이에서 MPI 실행시 랭크의 수를 변화시키면서 성능 벤치마킹 을 수행하여 하드웨어적인 성능 특성을 연구하였다. 그 결과 물리 코어가 57개인 제온파이 프로세서의 하나의 코어 당 4개의 작업을 처리할 때 가장 좋은 성능을 나타내고 있으며, 물리 코어 하나에 MPI 랭크수를 4개 이상 확장하면 성능향상이 거의 일어나지 않는다. 이러한 융합 기술을 통하여 퀀텀 에스프레소의 성능 향상과 제온 파이의 하드웨 어적인 특성을 확인할 수 있다.
Recently the degree of integration of processor and developed rapidly. However, clock speed is not increased, a situation that increases the number of cores in the processor. In this paper, we analyze the performance of a typical Intel Xeon Phi of many core process used for the current operation accelerate. Utilizing the Quantum ESPRESSO, which was calculated using the FFTW library. By varying the number of ranks in MPI when running the benchmarks the performance Xeon Phi. The result shows a good performance in the handling of four job on one physical core. However, four or more to expand the number of MPI Rank is degraded. Through this convergence it was found to improve the performance of Quantum ESPRESSO. It is possible to check the hardware characteristics of the Xeon Phi.
The Status of the Coordinate Structure Constraint KCI 등재
한국언어과학회 언어과학 제18권 2호 2011.05 pp.175-193
※ 기관로그인 시 무료 이용이 가능합니다.
5,400원
This paper raises a question whether the Coordinate Structure Constraint should maintain its status as a general constraint. There are some reported counterexamples in the literature. Considering the undeniable presence of counterexamples, there are two possible ways to deal with those documented counterexamples. One is to decide to abandon the Coordinate Structure Constraint altogether. The other is to find a plausible way to circumvent the Coordinate Structure Constraint to a restrictive extent. Given the general acceptance of the Coordinate Structure Constraint, this paper pursues to find a feasible solution to relax the Coordinate Structure Constraint so that it is allowed to be violable only in a well-controlled circumstance. Instead of seeing the Coordinate Structure Constraint as a condition on movement, it may well be better understood as a well-formenness condition at LF
The mainstream of the RTA case‐law in the WTO Dispute Settlement System KCI 등재
원광대학교 법학연구소 원광법학 제25권 제2호 2009.06 pp.165-191
지역무역협정은 WTO의 출범시기부터 각국의 체결이 급증하여, 최근에는 국제무역규범의 대안으로까지 평가되기도 한다. 이러한 상황 속에서, WTO는 회원국들의 지역무역협정 체결에 대한 규제를 보다 강화하기 위해 지역무역협정위원회의 설치를 비롯한 다양한 방법을 시도하고 있다. 지역무역협정위원회의 역할에 더하여, WTO 분쟁해결절차에서 지역무역협정 관련 분쟁의 해결을 통한 일종의 판례법을 축적하여 지역무역협정에 관련한 해석의 기준을 마련하는 것도 WTO 체계 내에서 지역무역협정의 규율을 강화하기 위한 방법이 될 수 있다. WTO 분쟁해결절차에서는 몇 차례의 지역무역협정 관련 사례가 논의된 바 있다. 이러한 사례들 중 많은 경우가 FTA 또는 관세동맹과 같은 형태의 지역무역협정과 세이프가드 적용에 관한 것이었다. 또한 사례들의 분석 결과, 분쟁해결기구는 상당수의 분쟁에 있어서 지역무역협정에 관한 논점을 적극적으로 해결하기 보다는 이른바 ‘병행주의’(Parallelism) 등과 같은 전혀 다른 이론을 적용하여 가급적 피하려고 한다는 점을 발견할 수 있다. 이러한 내용은, 지역무역협정에 관한 WTO 규범들의 해석에 있어서 모호한 부분을 해결하기 위해 축적된 ‘WTO 판례법’이 이용될 수 있을 것이라는 기대에는 미치지 못하는 결과이다. 지역무역협정에 관한 WTO의 적극적인 규제와 분쟁 발생의 방지를 위해서는, 이와 관련된 판례의 축적이 더욱 요구된다. 아울러 분쟁해결절차에서 지역무역협정의 논의가 보다 적극적으로 진행되어야 할 것이다.
[NRF 연계] 한국언론학회 Asian Communication Research Vol.12 No.1 2015.06 pp.5-36
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Relying on ‘press-party parallelism’ (Seymour-Ure, 1974) as a theoretical and analytical framework, this study examines the extent to which newspapers ideologically coincided with political parties during three South Korean political regimes (1993-1998, 1998-2003, and 2003-2008). In so doing, this study analyzes editorials by three Korean dailies with ideological diversity and different ownership regarding a long-time controversial issue, ‘openness of agriculture’ embedded in globalization since the1990s. Results reveal that the ideologically conservative newspaper, Donga Ilbo, was parallel to conservative political parties, which is opposite for the progressive newspaper, Hankyoreh. With the somewhat moderate newspaper, Hankook, being inbetween, the two contrasting newspapers, Donga Ilbo and Hankyoreh, confirmed ‘press-party parallelism’ across the political regimes, although Korean dailies have become excessively commercialized since the 1990s. The implications for press-party/state relations are discussed.
Enhancing network function parallelism in mobile edge computing using Deep Reinforcement Learning
[NRF 연계] 한국통신학회 ICT Express Vol.11 No.1 2025.02 pp.41-46
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper introduces a Deep Reinforcement Learning (DRL)-based framework to enhance Network Function Parallelism (NFP) in Mobile Edge Computing (MEC). Leveraging Network Function Virtualization (NFV), the proposed framework optimizes service delay by solving a fairness-aware throughput maximization problem for service function chain placement. It aims to maximize the long-term cumulative reward while satisfying Quality of Service (QoS) requirements. The framework also preserves resources for future requests by efficiently managing the initialized network functions distribution. Simulation results demonstrate the superior performance of the proposed framework across various metrics. Specifically, our framework improves the average delay and deployment rate by 1.2% and 2.4% compared to the existing best method.
Parallelism is required in the case of FS(Form Set)
한국언어과학회 한국언어과학회 학술대회 문법 지식의 역동성: 이론ㆍ습득ㆍ사용의 접점 2025.08 pp.35-38
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Parallelism in Labeling between (Free) Relativitization and Coordinate Conjunction KCI 등재
한국중앙영어영문학회 영어영문학연구 제57권 1호 2015.03 pp.247-266
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
This paper examines the long-standing issue of syntactic projection/ labeling in free relatives (FRs) of English. Pointing out some drawbacks on Chomsky’s (2013) recent labeling analysis of FRs, we suggest that labeling of FRs can be accounted for more effectively on a par with his (2013) analysis of coordinate structure. Specifically, both the morphologically simple elements that head FRs, such as what, where, when, and the morphologically complex -ever forms like whoever, serve the ambiguous role of relative head and relative operator. They thus determine the label of the resulting construction they build merging with the following FR-forming complementizer. In essence, this complementizer in FRs behaves like Conj, which as Chomsky suggests is ‘not available as a label,’ and therefore the moved element occupying its Spec position in FRs projects itself up to the resulting construction.
A Harmonic Parallelism Account of English Spirantization KCI 등재
한국중앙영어영문학회 영어영문학연구 제61권 3호 2019.09 pp.249-266
※ 기관로그인 시 무료 이용이 가능합니다.
5,200원
This paper accounts for English spirantization without using any intermediate stages between the input and output within the framework of harmonic parallelism. We have reviewed the previous analyses of English spirantization, especially a harmonic serialism analysis. We have found that, unlike an alternative multi-stratal model in particular, it provides a good tool of analysis in that it uses a single constraint hierarchy. However, it still needs to introduce intermediate steps or derivation-like steps. On the contrary, the harmonic parallelism approach provides a good account of phonological opacity shown in spirantization, which has been problematic for classic OT, by using the notion of ‘trace’. It is more advantageous in economy than the harmonic serialism approach in that it does not accept the notion of derivation-like serial application, maintaining the fundamental stance of classic OT by using the notion of ‘trace’. In doing so, we have revised the previous constraints and presented a more generalized constraint *[+lax, -high -mid]wd instead of a very specific *[ɛ]wd. The overall constraint ranking is *[t/d+v˘ /v˘ NLF], Id-IO(pl) ≫ Id-IO(cont), *[z+ɪv/ɪvCI] ≫ Id-IO(voice), *[+lax, -high -mid]wd ≫ Max-IO.
Accommodation and Scope Parallelism in Ellipsis
고려대학교 언어정보연구소 언어정보 제10호 2009.03 pp.21-43
※ 기관로그인 시 무료 이용이 가능합니다.
6,000원
4,000원
Many previous studies have taken the view that conditional clauses and cause-reason clauses differ only in whether their propositions are hypothetical or factual, and that they share other characteristics in common. This paper shows that there are differences between the two clauses in terms of clause size, rewriting into two sentences, etc. This difference is not automatically derived from the hypothetical/factual difference between the two clauses, which indicates that the conditional clause and the cause-reason clause are not completely parallel.
강원 영서, 영동지역 청동기시대 편년 병행 관계 -석촉 형식 분류를 중심으로- KCI 등재
숭실대학교 역사문물연구소(구 숭실사학회) 숭실사학 제29집 2012.12 pp.5-36
※ 기관로그인 시 무료 이용이 가능합니다.
7,300원
이 글은 강원 영서, 영동지역 청동기시대 편년 틀을 재구성하기 위해 작성되었다. 기존의 편년안은 주거지의 변화와 토기, 석기 구성의 변화를 토대로 전기와 중기로 구분되었지만, 중기 시점이 송국리유형과의 병행 관계에서 도출된 것은 아니었다. 게다가 영서지역의 북한강유역과 남한강유역, 그리고 영동지역 등 3개 지역의 병행 관계도 일정한 기준에 의해 결정된 것이 아니었다. 따라서 지역 차이가 크지 않으면서 시간 차이를 민감하게 반영하는 석촉을 세부 형식으로 분류하고 이들의 공반 관계를 통해 청동기시대 편년 틀을 재구성하였다. 석촉은 크게 무경촉(Ⅰ유형), 이단경촉(Ⅱ유형), 일단납작경촉(Ⅲ유형)으로 분류되며, Ⅰa1식, Ⅰa2식, Ⅰb식, Ⅱa식, Ⅱb식, Ⅱc식, Ⅲa식, Ⅲb식, Ⅲc식, Ⅲd식 등 모두 10개의 형식으로 세분된다. 또한 유적에서의 공반 관계를 통해 Ⅰa1 → Ⅰa2․Ⅱa → Ⅱb → Ⅰb․Ⅱc → Ⅲa → Ⅲb → Ⅲc → Ⅲd 석촉 순으로 발생했다고 판단된다. 강원 영서, 영동지역에서의 청동기시대 중기는 송국리형 무덤이 분포한 논산 마전리유적과의 비교를 통해 Ⅲa식 석촉이 출현하는 단계로 파악하였다. 이러한 편년안에 따르면 기존의 청동기시대 전기 후엽으로 편년되었던 유적의 일부 주거지들은 중기로 편년된다. 또한 청동기시대 전기는 전엽, 중엽, 후엽으로 중기는 전반과 후반으로 세분된다. 강원 영서, 영동지역의 청동기시대 중기는 기원전 8세기를 중기의 상한으로 보아왔지만, 기원전 10세기 무렵까지 소급될 가능성이 높은데, Ⅲa식 석촉이 출현하는 단계부터 석관묘, 석곽묘에서 지석묘로의 묘제 변화가 관찰되기 때문에 지석묘 축조는 강원 영서, 영동지역 청동기시대 중기의 가장 큰 특징이다.
This article aims to reconstruct the chronological frame of the Bronze Age of Yeongseo and Yeongdong regions in Gangwondo Province, Korea. Exiting chronology divides this age into two of the early and the middle periods based on changes of dwellings and in compositions of earthen and stone wares. However, starting point of the middle period was not drawn from its parallelism with Songguk-ri(松菊里) type. In addition parallelism of three regions of the North-Han and South-Han river valleys within Yeoungseo, and of Yeoungdong was not set by a certain standards either. Therefore I reconstruct chronological frame of the Bronze Age with the stone arrowheads, which precisely reflect time difference. To do so, I divided them into detailed forms and investigated if they are excavated together. The stone arrowheads are classified into categories of type Ⅰ, type Ⅱ, and type Ⅲ, which are subdivided into ten types: Ⅰa1, Ⅰa2, Ⅰb, Ⅱa, Ⅱb, Ⅱc, Ⅲa, Ⅲb, Ⅲc, Ⅲd. Their chronological order is Ⅰa1 → Ⅰa2․Ⅱa → Ⅱb → Ⅰb․Ⅱc → Ⅲa → Ⅲb → Ⅲc → Ⅲd according to the arrowheads unearthed together with within the same site. The mid-Bronze Age of Yeongseo and Yeongdong of Gangwondo Province is a period when IIIa type arrowheads emerged by being compared with Majeon-ri(麻田里) site where Songguk-ri type tombs are distributed. According to this new chronology, some dwellings having been regarded as the latter part of the early Bronze Age bythe exiting chronology could be changed into the mid-Bronze age. In addition the early Bronze Age can be subdivided into two parts of the former and latter, and the middle Bronze Age into three of the early, middle and late parts. The earliest middle Bronze age of Yeongseo and Yeongdong can’t be dated to earlier than the mid-8th century B.C., but it could be dated back to 10thcentury B.C. Building dolmens is the most distinct characteristics of the mid-Bronze age of Yeongseo and Yeongdong because tombs were changed from stone-coffin tombs and stone-chamber toms to dolmens, as IIIa type emerged.
스피노자 표상개념의 문제점 -표상론적 해석과 평행론적 해석의 불일치 KCI 등재
범한철학회 범한철학 제44집 2007.03 pp.91-115
※ 기관로그인 시 무료 이용이 가능합니다.
6,300원
스피노자 표상개념에는 정리 13과 정리 16 사이에 명백한 모순이 존재한다. 정리 13에서처럼 인간정신을 구성하고 있는 관념의 대상이 인간신체만을 의미하는지, 아니면 정리 16의 진술처럼 인간신체뿐만 아니라 인간신체에 작용하는 외부물체까지 포함하는지에 대해 스피노자는 우리에게 상충되는 사실을 알려주고 있다. 이 문제에 관심 있는 대부분의 학자들은 스피노자가 관념이라는 용어를 이중적인 의미로 사용함으로써 표상론과 평행론 사이에서 심각하게 혼동했다고 불만을 토로한다. 그러나 델라 로카의 말처럼 우리는 스피노자가 자기모순을 범했다고 결론내리기 이전에 모든 가능한 시도를 해 보아야 한다. 스피노자를 자기모순으로부터 구하기 위한 델라 로카의 재치 있는 시도는 아깝게도 성공하지 못했다. 필자 역시 문제 해결을 위한 하나의 가능한 시도를 해보았다. 정리 13에서의 라틴어 ‘corpus’ 를 ‘인간신체(the body)’가 아닌 ‘물체(a body)’라고 번역함으로써 정리 16과의 모순을 해결하고자 하였다. 그러나 필자의 부정관사(a body) 번역 시도는 매력적이기는 하지만, 여러 진술들에서 스피노자는 이를 거부하고 있다. 그러나 문제 해결을 포기하기에는 아직 이르다고 생각한다. 필자는 스피노자를 자기모순으로부터 구하기 위해 또 다른 해결 방안을 모색해 볼 것이다. 이러한 작업을 통해 우리는 스피노자 표상개념의 문제점을 해결하는 데 한 걸음 더 가까이 갈 수 있을 것이다.
병행 교육과정을 활용한 소설의 인물 이해 교육 연구- 핵심 교육과정과 정체성 교육과정의 병행을 중심으로 KCI 등재후보
한국언어문학교육학회 한어문교육 제33집 2015.08 pp.115-138
※ 기관로그인 시 무료 이용이 가능합니다.
6,100원
This paper is aimed at investigating into possibility of education of understanding characters in novels using a parallel curriculum. For this, the stages and methods of the education of understanding characters in novels which fits for each characteristic of students were examined through parallelism of core curriculum and identity curriculum. The current education of understanding characters in schools only hands down the knowledge of understanding and appreciation which are organized in reference books to most students, not leading to a profound understanding of character. Accordingly, this study aims to identify the correlationship between parallel curriculum model and education of understanding characters through focusing questions of PCM; to connect internal factors of the education of understanding characters with contents of core curriculum; and to design the plans for the education of understanding characters by relating external factors of the education of understanding characters to identity curriculum. The internal factors of the education of understanding characters includes objective information of characters, viewing attitude towards characters, and roles of characters in fictional works, and as external factor, there are learners who interpret the works. This study is intended to contribute in helping learners create character-related knowledge and actively appreciate novels through parallelism of core curriculum and identity curriculum where concept discovery for an understanding of characters progresses along with education fit for learners' characteristics at the same time.
루이스 캐럴의 그림책 ‘앨리스’ 시리즈의 상상력을 중심으로 살펴본 화문대구성 연구 KCI 등재
한국디자인트렌드학회 한국디자인포럼 Vol. 20 2008.08 pp.307-316
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
오랫동안 서구문명에서 어린이는 어른의 축소판으로 간주되어 순수하게 어린이를 대상으로 쓰인 현존하는 고전문학은 거의 없다시피 하다. 영국은 18세기에 이르러 산업혁명의 영향으로 중간계급의 가족중심 라이프스타일과 인쇄술이 발달하면서 교육용 아동 문학이 출판되기 시작하였다. 이후 19세기에는 상상력의 재평가와 낭만주의 운동 등으로 인하여 교육뿐만 아니라 읽는 재미와 즐거움을 목적으로 한 어린이 문학이 등장하게 된다. 수학자였던 루이스 캐럴은 《이상한 나라 앨리스의 모험》, 《앨리스의 땅속 모험》, 《거울 속으로》, 《유아용 앨리스》 등 ‘앨리스‘ 시리즈를 출판하였고 이중 《이상한 나라 앨리스의 모험》은 읽는 즐거움과 기쁨을 위해 출판된 최초의 판타지 그림책이라 할 수 있다. ‘앨리스‘ 시리즈는 아크로스틱, 동음어 등의 어문학, 역방향 사고체계의 논리학, 거울을 이용한 과학적, 그리고 병리학적 지식 등 다양한 학문의 지식을 이용하여 상상의 세계를 표현하고 있다. 《이상한 나라 앨리스의 모험》, 《거울 속으로》 등을 중심으로 살펴본 ’앨리스‘ 시리즈는 현실세계의 이론을 바탕으로 구축된 상상력을 텍스트로 표현하고 이와 상호보완관계를 이루는 존 테니얼의 일러스트레이션을 통해 비가시적인 상상의 세계를 이미지화함으로써 성공적인 화문대구를 이룬 판타지 그림책의 정전이라 할 수 있다.
In the western society, children was conceived as a miniature adult throughout the Middle Ages. So literature that was originally intended for children hardly existed. Children were recognized as an independent being and children's book as a teaching device were published under the influence of the rise of the middle class in the middle of the 18th century. In the 19th century, children's book for pleasing and fun were introduced due to the revaluation of imaginative power and romanticism. Lewis Carroll wrote 《Alice's Adventures in Wonderland》, 《Alice's Adventures Under Ground》,《Through the Looking-Glass and What Alice Found There》, and《Nursery Alice》.《Alice's Adventures in Wonderland》 is considered a landmark of fantasy picture book for the beginning of pleasure reading for children. There are linguistic, logical, scientific and pathological concepts with fantasy in 'series for Alice'. Parallelism between Carroll's text and Tenniel's illustration in 'series for Alice' did significant the role for developing imaginative power.
Multi-Core Program Optimization: Parallel Sorting Algorithms in Intel Cilk Plus
보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.7 No.2 2014.03 pp.151-164
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
New performance leaps has been achieved with multiprogramming and multi-core systems. Present parallel programming techniques and environment needs significant changes in programs to accomplish parallelism and also constitute complex, confusing and error-prone constructs and rules. Intel Cilk Plus is a C based computing system that presents a straight forward and well-structured model for the development, verification and analysis of multi-core and parallel programming. In this article, two programs are developed using Intel Cilk Plus. Two sequential sorting programs in C/C++ language are converted to multi-core programs in Intel Cilk Plus framework to achieve parallelism and better performance. Converted program in Cilk Plus is then checked for various conditions using tools of Cilk and after that, comparison of performance and speedup achieved over the single-core sequential program is discussed and reported.
Applying Hessian Curves in Parallel to improve Elliptic Curve Scalar Multiplication Hardware. SCOPUS
보안공학연구지원센터(IJSIA) International Journal of Security and Its Applications Vol.4 No.2 2010.04 pp.27-38
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
As a public key cryptography, Elliptic Curve Cryptography (ECC) is well known to be the most secure algorithms that can be used to protect information during the transmission. ECC in its arithmetic computations suffers from modular inversion operation. Modular Inversion is a main arithmetic and very long-time operation that performed by the ECC crypto-processor. The use of projective coordinates to define the Elliptic Curves (EC) instead of affine coordinates replaced the inversion operations by several multiplication operations. Many types of projective coordinates have been proposed for the elliptic curve E: y2 =x3 +ax+b which is defined over a Galois field GF(p) to do EC arithmetic operations where it was found that these several multiplications can be implemented in some parallel fashion to obtain higher performance. In this work, we will study Hessian projective coordinates systems over GF (p) to perform ECC doubling operation by using parallel multipliers to obtain maximum parallelism to achieve maximum gain.
미국의 철강 세이프가드 조치에 대한 WTO 분쟁 해결 사건
한국국제경제법학회 국제경제법연구 제3호 2005.12 pp.7-31
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Monitoring of Programs with Nested Parallelism using Efficient Thread Labeling
보안공학연구지원센터(IJUNESST) International Journal of u- and e- Service, Science and Technology vol.4 no.1 2011.03 pp.21-36
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
It is difficult and troublesome to detect data races occurred in an execution of parallel programs. On-the-fly race detection techniques use monitoring to threads and events, which access to shared variables. Any of them using Lamport's happened-before relation needs a thread labeling scheme for generating unique identifiers which maintain logical concurrency information of the parallel threads. Previous labeling schemes for on-the-fly race detection in parallel programs with nested parallelism depend on the nesting depth on every fork and join operation. This paper presents e-NR (Extended Nest Region) labeling in which every thread generates and maintains its label in a constant amount of time and space. Experimental results using OpenMP programs show that presented labeling detects data races more efficiently, because it requires reduced time overhead about 10% and it 3.2 times smaller than previous labeling scheme in average space for race detection.
A Study of Leveraging Memory Level Parallelism for DRAM System on Multi-Core Architecture
보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.9 No.4 2016.04 pp.319-338
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
DRAM system has been more and more critical on modern multi-core architecture where the Moore’s law has been made effect on increasing the number of cores integrated in a processor chip. The performance of DRAM system is usually measured in term of bandwidth and latency, which are regarded as inherently depending on Row Buffer Hit Rate (RBHR) according to previous studies. In this paper, we find that Memory Level Parallelism (MLP) exhibits a stronger correlation with the performance of DRAM system on multi-core/many-core architecture than RBHR, and promoting MLP significantly improves DRAM system performance. In order to exploit the MLP, we have evaluated various approaches including multi-bank, multi-row-buffers, multi-memory-controllers and the obsolete Virtual Channel Memory (VCM). The experimental results show that VCM is a better alternative to traditional DRAM chip on multicore/many-core architecture than the other three approaches because VCM has almost all the advantages of the others: 1) it can improve homogeneous workloads’ IPC by 2.21X on a 16-core system with 32 virtual channels due to leveraging unexploited MLP. 2) It can also promote Quality-of-Service (QoS) of DRAM system by removing unfairness while memory controllers serve memory requests. 3) It can save energy and has low area costs. Unfortunately, VCM, which was proposed in the late 1990s, faded away before multi-core/manycore became dominated. Therefore, we suggest memory chip vendors reconsider the VCM technology for multi-core architecture.
A 32nm EGPU Parallel Multiprocessor Based on Co-issue and Multi-Dimensional Parallelism Architecture SCOPUS
보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.10 No.5 2015.05 pp.343-354
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In this paper, a Parallel Multiprocessor (PM) based on SIMT (Single Instruction and Multiple Threads) architecture is proposed. With co-issue architecture and multi-dimensional parallelism implemented in high-effective PM, Embedded Graphics Processing Unit (EGPU) provides great performance for various situations, such as general purpose computing, 3D scene rendering, and graphics processing. Application programs are departed into separated threads. Allocated by Thread Processing Unit (TPU), separated threads can be executed in parallel. Parallelism in different hierarchy and dimension are implemented by Multi-Dimensional Parallelism Processor (MD-PP), which has made a proper trade-off between performance and cost. Additionally, PM improves the hardware occupancy with its co-issue architecture and internal bus accessing mechanism to meet the demand of processing capability. Its unified shading architecture also helps to hide processing latency. PM can execute 4 basic operations in the best case and 2 in the worst case within each clock cycle. With 32nm process technology and 200MHz clock frequency, PM’s area is about 5104494um2, power consumption is about 101.838mW, and it can process nearly 28M vertices or fragments in average. Experimental results show that the MD-PP based PM can process data with high performance and get a balance between efficiency and hardware consumption simultaneously.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.