Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 114
No
1

5,500원

This study aims to describe and analyze Korean and Indonesian object pronouns based on a Korean-Indonesian parallel corpus. Both Korean and Indonesian are categorized as ‘discourse pro-drop languages’ that allow arguments(subjects and objects) to be dropped. However, the object elliption in Indonesian is more restrained because in independent clauses the object cannot be ommited. The results of the Korean-Indonesian parallel corpus showed that there were instances when Korean object pronouns corresponded with Indonesian object pronouns and when they did not. The Indonesian object pronoun is controlled by its verb’s transitivity and therefore can appear even when the object pronoun in the Korean text was expressed using zero anaphor. The Korean object pronoun when overtly expressed, displayed new information, focus or emphasis. In addition, we also observed cases on the account of different language expression, Korean and Indonesian preference on expression was discovered.

2

병렬말뭉치 기반 한국어와 인도네시아어의 구어 정보구조 대조 KCI 등재

미쉘 알렉산드라 자자, 김선정

한국국제문화교류학회 문화교류와 다문화교육(구 문화교류연구) 제12권 제3호 2023.05 pp.349-372

※ 기관로그인 시 무료 이용이 가능합니다.

6,100원

본 연구는 한국어와 인도네시아어의 정보구조(information structure)를 비교·대조하 여 공통점과 차이점을 밝히는 데 그 목적이 있다. 한국어와 인도네시아어 드라마 병렬말 뭉치를 구축한 후 이를 바탕으로 두 언어의 구어 정보구조의 특징을 분석한다. 본 연구 에서 사용하는 정보구조의 개념은 ‘주어짐(givenness)’으로 화자와 청자가 이미 알고 있 는지의 사실 여부에 따라 신정보(new information)와 구정보(given information)로 분류된다. 연구 결과에 따르면, 신정보는 한국어에서는 상황에 따라 주격조사 ‘이/가’, 부정 (indefiniteness) 표현으로 표시되는 반면 인도네시아어에서는 어순과 보문소 ‘yang’으 로 표시된다. 구정보는 한국어에서는 보조사 ‘은/는’, 대명사, 용언 ‘있잖아/말이야’로 표시되고, 인도네시아어에서는 대명사와 접어 ‘=nya’로 표시된다. 구정보의 경우 한국 어와 인도네시아어 두 언어 모두에서 맥락을 통해 쉽게 드러날 수 있는 경우 해당하는 문장 성분을 생략할 수 있다. 또 한 가지 공통점은 수사구의 경우 핵 명사가 구정보이 면 이를 생략할 수 있는 있다는 점이다. 본 연구는 한국어와 인도네시아어의 구어 정보구조 대조에 관한 최초의 연구라는 데 그 의의가 있다. 본 연구의 결과는 인도네시아인 학습자를 위한 한국어 교육과 한국인 학습자를 위한 인도네시아어 교육의 기초 자료로 활용될 수 있을 것으로 기대한다.

The aim of this study is to reveal the commonalities and differences in Korean and Indonesian by comparing and contrasting the two languages’ spoken language information structures based on a parallel corpus. The notion of information structure used in this study is givenness, which is classified into new and given information according to what the speaker already knows. The results of the study show that new information in Korean is indicated by the subjective case marker ‘이/가’ and indefinite expressions, while in Indonesian, it is indicated by word order and the complementizer ‘yang’. Given information in Korean is indicated by the topic marker ‘은/는’, pronouns, and the terms ‘있 잖아/말이야’, while in Indonesian, it is indicated by pronouns and the clitic ‘-nya’. In the case of given information, the constituent can be ellipted in both Korean and Indonesian. Both languages also show that in numeral phrases, if the noun is a given information, it can be ellipted. This study is meaningful in that it is the first study on Korean and Indonesian spoken language information structure. It is hoped that this study will serve as useful material for Indonesian Korean learners in the field of Korean language education.

3

병렬말뭉치 기반 한국어와 인도네시아어의 구어 정보구조 대조 KCI 등재

미쉘 알렉산드라 자자, 김선정

인하대학교 다문화융합연구소 문화교류와 다문화교육 제12권 제3호 2023.05 pp.349-372

※ 기관로그인 시 무료 이용이 가능합니다.

6,100원

본 연구는 한국어와 인도네시아어의 정보구조(information structure)를 비교·대조하 여 공통점과 차이점을 밝히는 데 그 목적이 있다. 한국어와 인도네시아어 드라마 병렬말 뭉치를 구축한 후 이를 바탕으로 두 언어의 구어 정보구조의 특징을 분석한다. 본 연구 에서 사용하는 정보구조의 개념은 ‘주어짐(givenness)’으로 화자와 청자가 이미 알고 있 는지의 사실 여부에 따라 신정보(new information)와 구정보(given information)로 분류된다. 연구 결과에 따르면, 신정보는 한국어에서는 상황에 따라 주격조사 ‘이/가’, 부정 (indefiniteness) 표현으로 표시되는 반면 인도네시아어에서는 어순과 보문소 ‘yang’으 로 표시된다. 구정보는 한국어에서는 보조사 ‘은/는’, 대명사, 용언 ‘있잖아/말이야’로 표시되고, 인도네시아어에서는 대명사와 접어 ‘=nya’로 표시된다. 구정보의 경우 한국 어와 인도네시아어 두 언어 모두에서 맥락을 통해 쉽게 드러날 수 있는 경우 해당하는 문장 성분을 생략할 수 있다. 또 한 가지 공통점은 수사구의 경우 핵 명사가 구정보이 면 이를 생략할 수 있는 있다는 점이다. 본 연구는 한국어와 인도네시아어의 구어 정보구조 대조에 관한 최초의 연구라는 데 그 의의가 있다. 본 연구의 결과는 인도네시아인 학습자를 위한 한국어 교육과 한국인 학습자를 위한 인도네시아어 교육의 기초 자료로 활용될 수 있을 것으로 기대한다.

The aim of this study is to reveal the commonalities and differences in Korean and Indonesian by comparing and contrasting the two languages’ spoken language information structures based on a parallel corpus. The notion of information structure used in this study is givenness, which is classified into new and given information according to what the speaker already knows. The results of the study show that new information in Korean is indicated by the subjective case marker ‘이/가’ and indefinite expressions, while in Indonesian, it is indicated by word order and the complementizer ‘yang’. Given information in Korean is indicated by the topic marker ‘은/는’, pronouns, and the terms ‘있 잖아/말이야’, while in Indonesian, it is indicated by pronouns and the clitic ‘-nya’. In the case of given information, the constituent can be ellipted in both Korean and Indonesian. Both languages also show that in numeral phrases, if the noun is a given information, it can be ellipted. This study is meaningful in that it is the first study on Korean and Indonesian spoken language information structure. It is hoped that this study will serve as useful material for Indonesian Korean learners in the field of Korean language education.

4

공공 한영 병렬 말뭉치를 이용한 기계번역 성능 향상 연구 KCI 등재

박찬준, 임희석

한국디지털정책학회 디지털융복합연구 제18권 제6호 2020.06 pp.271-277

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

기계번역이란 소스언어를 목적언어로 컴퓨터가 번역하는 소프트웨어를 의미하며 규칙기반, 통계기반 기계번역 을 거쳐 최근에는 인공신경망 기반 기계번역에 대한 연구가 활발히 이루어지고 있다. 인공신경망 기계번역에서 중요한 요소 중 하나로 고품질의 병렬 말뭉치를 뽑을 수 있는데 이제까지 한국어 관련 언어쌍의 고품질 병렬 코퍼스를 구하기 쉽지 않은 실정이었다. 최근 한국정보화진흥원의 AI HUB에서 고품질의 160만 문장의 한-영 기계번역 병렬 말뭉치를 공개하였다. 이에 본 논문은 AI HUB에서 공개한 데이터 및 현재까지 가장 많이 쓰인 한-영 병렬 데이터인 OpenSubtitles와 성능 비교를 통해 각각의 데이터의 품질을 검증하고자 한다. 테스트 데이터로 한-영 기계번역 관련 공식 테스트셋인 IWSLT에서 공개한 테스트셋을 이용하여 보다 객관성을 확보하였다. 실험결과 동일한 테스트셋으로 실험한 기존의 한-영 기계번역 관련 논문들보다 좋은 성능을 보임을 알 수 있었으며 이를 통해 고품질 데이터의 중요성 을 알 수 있었다.

Machine translation refers to software that translates a source language into a target language, and has been actively researching Neural Machine Translation through rule-based and statistical-based machine translation. One of the important factors in the Neural Machine Translation is to extract high quality parallel corpus, which has not been easy to find high quality parallel corpus of Korean language pairs. Recently, the AI HUB of the National Information Society Agency(NIA) unveiled a high-quality 1.6 million sentences Korean-English parallel corpus. This paper attempts to verify the quality of each data through performance comparison with the data published by AI Hub and OpenSubtitles, the most popular Korean-English parallel corpus. As test data, objectivity was secured by using test set published by IWSLT, official test set for Korean-English machine translation. Experimental results show better performance than the existing papers tested with the same test set, and this shows the importance of high quality data.

5

6,400원

이 연구는 병렬말뭉치 기반 대조 연구의 일환으로 성경 자료를 활용하여 한․베 병렬 말뭉치를 구축하고, 이를 바탕으로 한국어와 베트남어의 차이점을 문장의 계량적 특성 에 초점을 맞추어 살피고자 한 연구이다. 이 연구에서는 절 단위로 정렬된 병렬말뭉치를 직접 구축하였으며, 형태소 분석을 통하여 형태 주석 말뭉치로 가공한 후 체언 및 용언 빈도, 대응 문장 수를 추출하였다. 이를 바탕으로 한국어와 베트남어의 체언 및 용언의 빈도를 비교해 본 결과, 한․베 문장의 체언 및 용언의 빈도에는 유의미한 차이가 없어 양 언어 모두 동사 지향적 언어 일 가능성이 높음을 알 수 있었다. 그리고 한국어와 베트남어의 문장 대응 수를 비교하여 양 언어의 복문 구성의 차이 를 확인하였다. 베트남어에서는 특정 표지 없이도 절의 연결이 가능하여 절과 절의 의 미 관계를 맥락에 의해 파악해야 하는 경우가 많은 반면, 한국어에서는 절과 절이 연결 될 때 연결어미의 사용이 활발하다는 양 언어의 차이를 문장 대응 수를 통하여 실증하 였다. 또한 한국어의 연결어미가 ‘병렬’, ‘부가’, ‘시간’, ‘배경’, ‘이유’, ‘대립’, ‘조건’, ‘전환’, ‘양보’의 의미 관계를 나타낼 때 베트남어에서는 독립적인 두 문장으로 표현하는 것이 가능하기에 베트남인 모어 화자가 한국어 학습이나 통번역 시 한국어 연결어미 선택에 어려움을 겪을 가능성이 있음을 언급하였다.

This study is part of a parallel corpus-based contrastive study that examines the differences between Korean and Vietnamese focusing on the quantitative characteristics of sentences based on a Korean-Vietnamese parallel corpus using biblical data. In this study, a parallel corpus arranged by clause was constructed, and after processing it into a morphological annotation corpus through morpheme analysis, the frequency of nouns and verbs and the number of corresponding sentences were extracted. Based on this, a comparison of noun and verb frequencies between Korean and Vietnamese was conducted. The results showed that there were no significant differences in the frequencies of nouns and verbs in Korean and Vietnamese sentences, suggesting that both languages are likely to be verb-oriented languages. Additionally, a comparison of sentence correspondences between Korean and Vietnamese revealed differences in sentence structure, particularly in complex sentences. Vietnamese allows for the connection of clauses without specific markers, often requiring contextual understanding of the relationship between clauses. On the other hand, Korean employs connective endings more actively when linking clauses, and this difference in sentence structure was empirically demonstrated through the comparison of sentence correspondences. Furthermore, it was noted that while Korean connective endings can convey the semantic relationships of “parallel”, “addition”, “time”, “background”, “reason”, “contrast”, “condition”, “transition”, and “concession”, Vietnamese may express these relationships with two independent sentences. Thus we argue that Vietnamese native speakers might face challenges in learning Korean, interpreting or translating due to the complexities of selecting appropriate Korean connective endings.

6

6,400원

이 연구는 병렬말뭉치 기반 대조 연구의 일환으로 성경 자료를 활용하여 한․베 병렬 말뭉치를 구축하고, 이를 바탕으로 한국어와 베트남어의 차이점을 문장의 계량적 특성 에 초점을 맞추어 살피고자 한 연구이다. 이 연구에서는 절 단위로 정렬된 병렬말뭉치를 직접 구축하였으며, 형태소 분석을 통하여 형태 주석 말뭉치로 가공한 후 체언 및 용언 빈도, 대응 문장 수를 추출하였다. 이를 바탕으로 한국어와 베트남어의 체언 및 용언의 빈도를 비교해 본 결과, 한․베 문장의 체언 및 용언의 빈도에는 유의미한 차이가 없어 양 언어 모두 동사 지향적 언어 일 가능성이 높음을 알 수 있었다. 그리고 한국어와 베트남어의 문장 대응 수를 비교하여 양 언어의 복문 구성의 차이 를 확인하였다. 베트남어에서는 특정 표지 없이도 절의 연결이 가능하여 절과 절의 의 미 관계를 맥락에 의해 파악해야 하는 경우가 많은 반면, 한국어에서는 절과 절이 연결 될 때 연결어미의 사용이 활발하다는 양 언어의 차이를 문장 대응 수를 통하여 실증하 였다. 또한 한국어의 연결어미가 ‘병렬’, ‘부가’, ‘시간’, ‘배경’, ‘이유’, ‘대립’, ‘조건’, ‘전환’, ‘양보’의 의미 관계를 나타낼 때 베트남어에서는 독립적인 두 문장으로 표현하는 것이 가능하기에 베트남인 모어 화자가 한국어 학습이나 통번역 시 한국어 연결어미 선택에 어려움을 겪을 가능성이 있음을 언급하였다.

This study is part of a parallel corpus-based contrastive study that examines the differences between Korean and Vietnamese focusing on the quantitative characteristics of sentences based on a Korean-Vietnamese parallel corpus using biblical data. In this study, a parallel corpus arranged by clause was constructed, and after processing it into a morphological annotation corpus through morpheme analysis, the frequency of nouns and verbs and the number of corresponding sentences were extracted. Based on this, a comparison of noun and verb frequencies between Korean and Vietnamese was conducted. The results showed that there were no significant differences in the frequencies of nouns and verbs in Korean and Vietnamese sentences, suggesting that both languages are likely to be verb-oriented languages. Additionally, a comparison of sentence correspondences between Korean and Vietnamese revealed differences in sentence structure, particularly in complex sentences. Vietnamese allows for the connection of clauses without specific markers, often requiring contextual understanding of the relationship between clauses. On the other hand, Korean employs connective endings more actively when linking clauses, and this difference in sentence structure was empirically demonstrated through the comparison of sentence correspondences. Furthermore, it was noted that while Korean connective endings can convey the semantic relationships of “parallel”, “addition”, “time”, “background”, “reason”, “contrast”, “condition”, “transition”, and “concession”, Vietnamese may express these relationships with two independent sentences. Thus we argue that Vietnamese native speakers might face challenges in learning Korean, interpreting or translating due to the complexities of selecting appropriate Korean connective endings.

7

5,700원

Multimedia corpus enables researchers to observe both spoken sentence materials and the nonverbal expressions of speakers. To undertake a contrastive study of the language behaviors of Japanese and Korean speakers, I have established the multimedia corpora of Japanese and Korean talk shows. Improving upon the previous settings of the corpora (e.g.,these featured inconsistency in the number of transcribed spoken sentences materials in one unit cell, or the length of time of the video record capturing each speaker varied), I analyzed the ways the demonstrative pronouns of deixis, kore and yi-geot, were used in real dialogues. The comparative analysis of the two corpora revealed that kore is used to demonstrate the very things that are indicated, whereas yi-geot is used to demonstrate the gestures and the action movements of the speakers who use it. As pointed out, the scenes that included the speech acts of kore and yi-geot are not identical in the corpus of Japan or of Korea. In other words, the scenes that feature the two words in each corpus draw upon different situations and themes in the talk shows. Therefore, one should consider such differences in the corpora that were studied for this paper when researching the language behaviors of Japanese and Korean speakers. In order to reach a fuller understanding of the differences in the scenes of these languages as they were spoken, I further investigated the usages of words other than those studied in this paper.

8

서로 다른 문장 구조의 병렬 말뭉치 통합을 통한 기계번역 모델 품질의 향상 KCI 등재

김호경, 김건우, 최근호

한국경영정보학회 경영정보학연구 제26권 제3호 2024.08 pp.103-117

※ 기관로그인 시 무료 이용이 가능합니다.

4,800원

최근 AI 기술이 빠르게 발전하면서 이전에는 개발하기 어려웠던 번역기를 민간에서도 비교적 쉽게 만들 수 있게 되었고, 일반적으로 학습 데이터의 양을 늘릴 경우 번역 품질은 향상되는 경향을 보였다. 하지만 뉴스 데이터로 학습된 기계번역 모델은 동일한 뉴스 데이터를 추가 학습해도 정형화되어 있지 않은 뉴스 데이터의 특성으로 인해 번역 모델의 품질 향상 폭이 크지 않다. 이에 본 연구에서는 이러한 뉴스 데이터가 가진 구조적 한계점을 보완하기 위해 정형화된 문장 구조를 가진 특허 데이터를 기계학습 시 학습 데이터에 추가하여 번역 품질을 향상시키고자 하였다. 현재 다양한 문장 구조를 가진 학습 데이터를 조합하여 기계번역 품질을 향상시키는 연구는 많이 이루어지지 않았으며, 대부분의 연구는 학습 데이터 자체의 품질이나 오류율을 최소화하는 데 중점을 두고 있다. 이를 위해 본 연구는 다양한 문장 구조를 가진 뉴스 학습 데이터와 정형화된 문장 구조를 가진 특허 학습 데이터의 비율을 조정하여 다양한 번역 모델을 생성하였고, 생성된 번역 모델의 품질 변화에 대한 분석을 수행하였다. 실험 결과, 뉴스 데이터와 특허 데이터의 비율을 2:8로 조정한 학습 데이터로 생성한 모델의 품질이 가장 좋게 나타났으며, 뉴스 데이터로만 학습한 모델 대비 66.7% 높은 품질을 보이는 것으로 나타났다.

Recent advances in AI technology have rapidly made it relatively easy for the public to develop translation systems that were previously difficult to create. Generally, increasing the amount of training data has tended to improve translation quality. However, machine translation models trained on news data do not show significant improvements in translation model quality even when additional news data is used for training, due to the unstructured nature of news data. In this study, we aimed to enhance translation quality by supplementing training data with patent data that has structured sentence patterns to address these structural limitations of news data. Research on improving machine translation quality by combining training data with various sentence structures is not extensively conducted, with most focusing on minimizing the quality or error rate of the training data itself. To address this, we generated various translation models by adjusting the ratio of news training data with structured patent training data and analyzed the quality changes of the generated translation models. Experimental results showed that the model trained with a 2:8 ratio of news data to patent data exhibited the highest quality, demonstrating a 66.7% improvement compared to models trained only on news data.

9

6,000원

이 연구는 한국어와 인도네시아어 병렬 말뭉치를 바탕으로 한국어 문장의 인도네시 아어 대응 양상을 고찰하는 데 그 목적이 있다. 고찰 결과, 연결어미가 사용된 복합문 중 ‘-는데, -고, -아서, -니까’가 사용된 경우, 해당 연결어미에 대응하는 접속사가 출현 하지 않는 경우가 대부분이어서 인도네시아 학습자들이 이들 연결어미가 사용된 복합 문의 생성에 어려움을 겪을 가능성이 예측된다. 특히 ‘-아서’, ‘-니까’보다는 ‘-는데’, ‘-고’ 의 접속사 대응 비율이 상대적으로 낮아 ‘-는데’, ‘-고’를 사용한 복합문 생성에 더욱 어 려움이 예상된다. 다음으로 전성어미가 사용된 복합문의 경우 한국어 문장에서 관형사 절의 사용이 훨씬 빈번하고 더욱이 한국어 복합문의 관형사절과 주절의 순서가 인도네 시아어 문장에서 도치되는 경우가 많기 때문에 인도네시아 학습자는 관형사절을 안은 복합문의 해석에 어려움을 겪을 것으로 예측된다. 덧붙여 한국어 단문의 비대칭 대응 관계에서는 ‘에’가 조건적 상황 또는 부가의 의미를 지닐 때, ‘로’가 부가적 설명의 의미 를 지닐 때 그리고 동격 구문, 병렬 구문, 서술명사 구문의 경우에 한국어 단문이 인도 네시아어에서 두 문장에 대응하는 예가 관찰되었는데, 인도네시아 학습자들이 습득에 어려움을 느끼는 구조일 가능성이 있다. 이 연구는 병렬말뭉치를 바탕으로 문법 항목 의 형태 대조가 아닌 구문적 차원의 대조를 시도해 보고자 하였다는 데 의의가 있을 것 이다

The purpose of this study is to examine the Indonesian correspondences of Korean complex sentences based on the parallel corpus of Korean and Indonesian. The first set of findings show that in the cases where ‘-는데, -고, -아서’, and ‘-니까’ are used as connective endings, the conjunctions corresponding to the connective endings mostly do not appear in Indonesian sentences, so this study predicts that Indonesian learners may have difficulties in constructing complex sentences using ‘-는데, -고, -아서’, and ‘니까’. Furthermore, the data of '-는데' and '-고' that correspond to Indonesian conjunctions are relatively lower than ‘-아서’ and ‘-니까’. Hence, this study predicts that constructing complex sentences using ‘-는데’ and ‘-고’ will be more difficult rather than using ‘-아서’ and ‘-니까’ for Indonesian learners. The second findings show that in the case of complex sentences using transformative endings, the use of adnominal clauses are more frequent in Korean sentences. Moreover, the order of the adnominal clauses and main clauses in Korean complex sentences are frequently realized in inverted sentences in Indonesian language. Consequently, this study predicts that Indonesian learners will have difficulties in analyzing Korean complex sentences containing adnominal clauses. The last findings show that some data related to Korean simple sentences correspond to two sentences in Indonesian language in which they are asymmetric correspondences. The findings refer to sentences containing ‘에’ which has the meaning of ‘conditional situation’ or ‘addition’, ‘로’ which has the meaning of ‘additional explanation’, and Korean sentences related to appositive structures, parallel structures, and predicate noun structures. This study would be meaningful in that it attempts to contrast the structures of sentences rather than the form of grammatical items based on a parallel corpus.

10

6,000원

이 연구는 한국어와 인도네시아어 병렬 말뭉치를 바탕으로 한국어 문장의 인도네시 아어 대응 양상을 고찰하는 데 그 목적이 있다. 고찰 결과, 연결어미가 사용된 복합문 중 ‘-는데, -고, -아서, -니까’가 사용된 경우, 해당 연결어미에 대응하는 접속사가 출현 하지 않는 경우가 대부분이어서 인도네시아 학습자들이 이들 연결어미가 사용된 복합 문의 생성에 어려움을 겪을 가능성이 예측된다. 특히 ‘-아서’, ‘-니까’보다는 ‘-는데’, ‘-고’ 의 접속사 대응 비율이 상대적으로 낮아 ‘-는데’, ‘-고’를 사용한 복합문 생성에 더욱 어 려움이 예상된다. 다음으로 전성어미가 사용된 복합문의 경우 한국어 문장에서 관형사 절의 사용이 훨씬 빈번하고 더욱이 한국어 복합문의 관형사절과 주절의 순서가 인도네 시아어 문장에서 도치되는 경우가 많기 때문에 인도네시아 학습자는 관형사절을 안은 복합문의 해석에 어려움을 겪을 것으로 예측된다. 덧붙여 한국어 단문의 비대칭 대응 관계에서는 ‘에’가 조건적 상황 또는 부가의 의미를 지닐 때, ‘로’가 부가적 설명의 의미 를 지닐 때 그리고 동격 구문, 병렬 구문, 서술명사 구문의 경우에 한국어 단문이 인도 네시아어에서 두 문장에 대응하는 예가 관찰되었는데, 인도네시아 학습자들이 습득에 어려움을 느끼는 구조일 가능성이 있다. 이 연구는 병렬말뭉치를 바탕으로 문법 항목 의 형태 대조가 아닌 구문적 차원의 대조를 시도해 보고자 하였다는 데 의의가 있을 것 이다.

The purpose of this study is to examine the Indonesian correspondences of Korean complex sentences based on the parallel corpus of Korean and Indonesian. The first set of findings show that in the cases where ‘-는데, -고, -아서’, and ‘-니까’ are used as connective endings, the conjunctions corresponding to the connective endings mostly do not appear in Indonesian sentences, so this study predicts that Indonesian learners may have difficulties in constructing complex sentences using ‘-는데, -고, -아서’, and ‘니까’. Furthermore, the data of '-는데' and '-고' that correspond to Indonesian conjunctions are relatively lower than ‘-아서’ and ‘-니까’. Hence, this study predicts that constructing complex sentences using ‘-는데’ and ‘-고’ will be more difficult rather than using ‘-아서’ and ‘-니까’ for Indonesian learners. The second findings show that in the case of complex sentences using transformative endings, the use of adnominal clauses are more frequent in Korean sentences. Moreover, the order of the adnominal clauses and main clauses in Korean complex sentences are frequently realized in inverted sentences in Indonesian language. Consequently, this study predicts that Indonesian learners will have difficulties in analyzing Korean complex sentences containing adnominal clauses. The last findings show that some data related to Korean simple sentences correspond to two sentences in Indonesian language in which they are asymmetric correspondences. The findings refer to sentences containing ‘에’ which has the meaning of ‘conditional situation’ or ‘addition’, ‘로’ which has the meaning of ‘additional explanation’, and Korean sentences related to appositive structures, parallel structures, and predicate noun structures. This study would be meaningful in that it attempts to contrast the structures of sentences rather than the form of grammatical items based on a parallel corpus.

11

한국어 상표지 ‘-어 오-’의 중국어 번역 연구 KCI 등재

유쌍옥, 김지혜

한중인문학회 한중인문학연구 제82집 2024.03 pp.79-107

※ 기관로그인 시 무료 이용이 가능합니다.

6,900원

본 연구는 신문 번역텍스트에서 선행동사의 동사부류가 한국어 상표지 ‘-어 오-’의 중국 어 번역에 어떤 영향을 미칠 수 있는지를 밝히는 데 목적을 둔다. 본 연구에서 구축한 신문 텍스트 한·중 병렬 말뭉치에서 ‘-어 오-’는 주로 중국어 시간부사 ‘一直’, 상표지 ‘在’, ‘着’, ‘下來’로 번역되는 양상을 보였다. 선행 동사의 동사부류에 따라 ‘-어 오-’의 중국어 번역 양 상이 다르게 나타났다. 달성동사와 결합하는 ‘-어 오-’는 ‘在’, 下來’, ‘着²’와의 결합 시 제약 을 받으며 주로 ‘一直’, ‘着¹’으로 번역되었다. 상표지 ‘着¹’과 ‘下來’는 완수동사와 결합은 가능 하지만, 완수동사와의 결합으로 나타내는 상적 의미가 ‘-어 오-’와 일치하지 않는다. 따라서 완수동사와 결합하는 ‘-어 오-’는 ‘下來’, ‘着¹’으로 번역되지 않고 주로 중국어 시간부사 ‘一 直’, 상표지 ‘在’로 번역되었다. 마지막으로 행위동사와 결합하는 ‘-어 오-’는 ‘着¹’와 결합 제 약을 가져 ‘着¹’을 제외한 ‘在’, ‘着²’, ‘下來’, 시간부사 ‘一直’ 등으로 다양하게 번역되는 양상 을 보였다.

This study aims to elucidate how the verb category of the preceding verb in actual translated texts may influence the Chinese translation of the Korean grammatical form ‘-e o-.’ To achieve this goal, the study first collected newspaper texts to build a Korean-Chinese parallel corpus. The analysis of the Chinese translation patterns of ‘-e o-’ revealed that it is primarily translated into Chinese using the temporal adverb ‘yizhi’ and the markers ‘zai,’ ‘zhe,’ and ‘xialai.’Subsequently, the study analyzed the combination patterns of achievement verbs, completion verbs, and action verbs with ‘-e o-’, ‘yizhi,’ and the markers ‘zai,’ ‘zhe,’ ‘xialai’ through actual translation cases. The results of the study indicate that the Chinese translation patterns of ‘-e o-’ vary depending on the verb category. ‘-e o-’ combined with achievement verbs, which have combinatory constraints with ‘zai,’ ‘xialai,’ and ‘zhe²,’ are primarily translated into ‘yizhi’ and ‘zhe¹.’ However, ‘-e o-’ combined with completion verbs, which can be combined with ‘zai’ and ‘xialai,’ is not translated into ‘xialai’ and ‘zhe¹’ due to the lack of semantic congruence between completion verbs and the conveyed meaning of ‘-e o-.’ Instead, it is mainly translated into the Chinese temporal adverb ‘yizhi’ and the marker ‘zai.’ ‘-e o-’ combined with action verbs, which only has combinatory constraints with ‘zhe¹,’ exhibits diverse translations into Chinese as ‘zhe¹,’ ‘zai,’ ‘xialai,’ and ‘yizhi.’

12

5,700원

본고는 대규모 병렬말뭉치를 이용하여 ‘-ㄹ 수 {있/없}-’ 구성의 각 의미 범주가 말뭉치에서 사용되는 빈도를 살펴보고 각 의미 범주에 대응되는 중국어 표현을 분석하는데에 목적이 있다. 말뭉치를 분석한 결과, ‘-ㄹ 수 {있/없}-’ 구성은 [능력]과 [가능성]의 두 가지의 기본의미를 나타내지만 말뭉치에서 고빈도로 [능력]의 의미로 쓰이고 [가능성]으로의 쓰임이 비교적 적음을 확인하였다. 특히 ‘-ㄹ 수 없-’ 구성의 경우 절대다수는 [능력 부정]의의미로 쓰이고 매우 드물게 [가능성 부정]의 의미로 쓰이고 있음을 확인하였다. ‘-ㄹ 수 있-’ 구성은 [능력]으로 쓰이는 경우 가장 높은 빈도로 중국어 능원동사와 대응되며, 이외에 가능보어 그리고 [능력]을 나타내는 구절이나 단어와의 대응 관계도 찾을 수 있다. [가능성]으로 쓰이는 경우 [추측]을 나타내는 중국어 능원동사, 추측성 어기부사 그리고 이들로 구성된 구조 표현과 대응되며 ‘가능성이 있음’의 의미를 나타내는 구절과도 비교적 높은 빈도로 대응된다. 한편, ‘-ㄹ 수 없-’ 구성은 ‘-ㄹ 수 있-’ 구성의부정 형식이지만 서로 다른 대응 양상을 보인다. ‘-ㄹ 수 없-’ 구성은 [능력 부정]으로 쓰이는 경우 능원동사의 부정 형식 및 가능보어와의 대응 관계를 찾을 수 있으나 ‘방법이없음’을 나타내는 단어나 구절과 더 빈번하게 대응됨을 확인하였다. 또한 [가능성 부정] 으로 쓰이는 경우 [추측]을 나타내는 능원동사의 부정 형식과의 대응 관계만 찾을 수있었다.

This research uses the parallel corpus of TV dramas to study the Chinese expressions corresponding to the structure of Korean ‘-l swu {iss/eps}-’. When the ‘-l swu iss-’ structure is used as the meaning of [ability], its high frequency corresponds to the Chinese modal verbs. In addition, it can also correspond to the Chinese potential complements, as well as phrases or words that express the meaning of [ability]. When it used as the meaning of [possibility], it can correspond to the modal verbs, modal adverbs, and the construction composed of them that express the meaning of [speculation]. And it can also form a corresponding relationship with the phrase that means ‘it’s possible’. On the other hand, although the ‘-l swu eps-’ structure is the negative form of the ‘-l swu iss-’ structure, there are different corresponding characteristics between the two. When the ‘-l swu eps-’ structure is used as the meaning of [Negation of Ability], although it can correspond to the negative forms of the Chines modal verbs and potential complements, corresponding to words and phrases that express the meaning of [there is a way] have a higher rating. And when it is used as the meaning of [possibility negation], we can only find the corresponding relationship with the negative form of the modal verbs that expresses the meaning of [speculation].

13

5,700원

본고는 대규모 병렬말뭉치를 이용하여 ‘-ㄹ 수 {있/없}-’ 구성의 각 의미 범주가 말뭉치에서 사용되는 빈도를 살펴보고 각 의미 범주에 대응되는 중국어 표현을 분석하는데에 목적이 있다. 말뭉치를 분석한 결과, ‘-ㄹ 수 {있/없}-’ 구성은 [능력]과 [가능성]의 두 가지의 기본의미를 나타내지만 말뭉치에서 고빈도로 [능력]의 의미로 쓰이고 [가능성]으로의 쓰임이 비교적 적음을 확인하였다. 특히 ‘-ㄹ 수 없-’ 구성의 경우 절대다수는 [능력 부정]의의미로 쓰이고 매우 드물게 [가능성 부정]의 의미로 쓰이고 있음을 확인하였다. ‘-ㄹ 수 있-’ 구성은 [능력]으로 쓰이는 경우 가장 높은 빈도로 중국어 능원동사와 대응되며, 이외에 가능보어 그리고 [능력]을 나타내는 구절이나 단어와의 대응 관계도 찾을 수 있다. [가능성]으로 쓰이는 경우 [추측]을 나타내는 중국어 능원동사, 추측성 어기부사 그리고 이들로 구성된 구조 표현과 대응되며 ‘가능성이 있음’의 의미를 나타내는 구절과도 비교적 높은 빈도로 대응된다. 한편, ‘-ㄹ 수 없-’ 구성은 ‘-ㄹ 수 있-’ 구성의부정 형식이지만 서로 다른 대응 양상을 보인다. ‘-ㄹ 수 없-’ 구성은 [능력 부정]으로 쓰이는 경우 능원동사의 부정 형식 및 가능보어와의 대응 관계를 찾을 수 있으나 ‘방법이없음’을 나타내는 단어나 구절과 더 빈번하게 대응됨을 확인하였다. 또한 [가능성 부정] 으로 쓰이는 경우 [추측]을 나타내는 능원동사의 부정 형식과의 대응 관계만 찾을 수있었다.

This research uses the parallel corpus of TV dramas to study the Chinese expressions corresponding to the structure of Korean ‘-l swu {iss/eps}-’. When the ‘-l swu iss-’ structure is used as the meaning of [ability], its high frequency corresponds to the Chinese modal verbs. In addition, it can also correspond to the Chinese potential complements, as well as phrases or words that express the meaning of [ability]. When it used as the meaning of [possibility], it can correspond to the modal verbs, modal adverbs, and the construction composed of them that express the meaning of [speculation]. And it can also form a corresponding relationship with the phrase that means ‘it’s possible’. On the other hand, although the ‘-l swu eps-’ structure is the negative form of the ‘-l swu iss-’ structure, there are different corresponding characteristics between the two. When the ‘-l swu eps-’ structure is used as the meaning of [Negation of Ability], although it can correspond to the negative forms of the Chines modal verbs and potential complements, corresponding to words and phrases that express the meaning of [there is a way] have a higher rating. And when it is used as the meaning of [possibility negation], we can only find the corresponding relationship with the negative form of the modal verbs that expresses the meaning of [speculation].

14

인공신경망 기계번역에서 말뭉치 간의 균형성을 고려한 성능 향상 연구 KCI 등재

박찬준, 박기남, 문현석, 어수경, 임희석

한국융합학회 한국융합학회논문지 제12권 제5호 2021.05 pp.23-29

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

최근 딥러닝 기반 자연언어처리 연구들은 다양한 출처의 대용량 데이터들을 함께 학습하여 성능을 올리고자 하는 연구들을 진행하고 있다. 그러나 다양한 출처의 데이터를 하나로 합쳐서 학습시키는 방법론은 성능 향상을 막게 될 가능성이 존재한다. 기계번역의 경우 병렬말뭉치 간의 번역투(의역, 직역), 어체(구어체, 문어체, 격식체 등), 도메인 등의 차이로 인하여 데이터 편차가 발생하게 되는데 이러한 말뭉치들을 하나로 합쳐서 학습을 시키게 되면 성능의 악영 향을 미칠 수 있다. 이에 본 논문은 기계번역에서 병렬말뭉치 간의 균형성을 고려한 Corpus Weight Balance (CWB) 학습 방법론을 제안한다. 실험결과 말뭉치 간의 균형성을 고려한 모델이 그렇지 않은 모델보다 더 좋은 성능을 보였다. 더불어 단일 말뭉치로도 고품질의 병렬 말뭉치를 구축할 수 있는 휴먼번역 시장과의 상생이 가능한 말뭉치 구축 프로세 스를 추가로 제안한다.

Recent deep learning-based natural language processing studies are conducting research to improve performance by training large amounts of data from various sources together. However, there is a possibility that the methodology of learning by combining data from various sources into one may prevent performance improvement. In the case of machine translation, data deviation occurs due to differences in translation(liberal, literal), style(colloquial, written, formal, etc.), domains, etc. Combining these corpora into one for learning can adversely affect performance. In this paper, we propose a new Corpus Weight Balance(CWB) method that considers the balance between parallel corpora in machine translation. As a result of the experiment, the model trained with balanced corpus showed better performance than the existing model. In addition, we propose an additional corpus construction process that enables coexistence with the human translation market, which can build high-quality parallel corpus even with a monolingual corpus.

15

5,500원

Although there are many previous studies that examined quantitative fuzzy semantics, the contrastive study of hedges between the Chinese and Korean languages are still undetermined. Whether one is able to accurately use the hedge expressions reflects the pragmatic competence of the foreign language student. The purpose of this paper is to investigate the corresponding expression of the Korean hedges expression ‘-(으)ㄹ 것이다’ in official Chinese speech texts, and the application of the results from this research as a pragmatic approach to the teaching scene of Korean as a foreign language education to Chinese students. In this paper, the author selected eleven official speeches by the heads of the Chinese government and their Korean translations as the base of parallel corpus. By using the corpus, statistical methods and the pragmatic approach, we found 104 corpus pairs of Korean hedge expression ‘-(으)ㄹ 것이다’ in these eleven official speeches. The results of this study show that the Korean hedges expression ‘-(으)ㄹ 것이다’ is the translation of various Chinese expressions from the original speeches. Moreover, 51.92% of them are unmarked Chinese expressions. It shows that not only were the suggestive expressions translated into Korean euphemism, but also the translators used the Korean hedges expression ‘-(으)ㄹ 것이다’ consciously in order to make it suitable for Korean readers. It suggests that we need to demonstrate the hedges expression ‘-(으)ㄹ 것이다’ in Korean teaching context, in order to improve Chinese students’ pragmatic competence and increase the Korean learners and readers’ acceptance of such expressions.

16

6,600원

Nam, Won Jun. (2006). Towards a Corpus-based Experiment in the into-English Translation Classroom: an Action Research Approach. Interpreting and Translation Studies 10-1, pp.29-55 본 연구는 우선 번역학의 맥락 안에서 실행 연구(action research)가 차지하는 또는 차지할 수 있는 잠재적 위치를 규명하고, 한영 번역 강의에서의 코퍼스 기반 실험을 제안한다. 통번역 전공 석사 과정생들이 피험자로 참여하는 이 실험에서는, 관광 브로슈어에서 발췌한 네 개의 원천 텍스트(source text)를 네 가지의 다른 환경[병렬 코퍼스(parallel corpus), 유관 코퍼스(comparable corpus), 병렬 + 유관 코퍼스, 코퍼스 없음]에서 번역한다. 본 연구에서 제안하는 실행 연구 접근법의 실험이 성공적으로 실행될 경우, 한영 번역 강의에서의 코퍼스의 역할을 어느 정도 규명할 수 있을 것으로 예상한다. 즉, 실험을 통해 1) 코퍼스가 번역 교강사/학생에게 유의한 번역 보조 도구인지, 2) 병렬, 유관 코퍼스 중, 어떤 것이 번역 교육에 보다 유의한지, 나아가 양자 간의 차이는 무엇인지, 3) 코퍼스를 사용하는 학생들의 번역을 보다 향상시키기 위해서는 어떻게 해야 하는지 등을 알아볼 수 있을 것으로 내다보인다.

17

A Parallel Corpus Based Contrastive Study on Negation in Korean and Indonesian Languages KCI 등재

Darmanto, Seon Jung Kim

인하대학교 다문화융합연구소 다문화와 교육 Vol.8 No.1 2023.04 pp.45-69

※ 기관로그인 시 무료 이용이 가능합니다.

6,300원

The current research aims to uncover the similarities and differences between Korean ‘an (안)’ and ‘-ji anhta (-지 않다)’, and Indonesian ‘tidak’ and ‘belum’, as well as to formulate their characteristics. This research is an applied contrastive study and the data are derived from the Korean-Indonesian parallel corpus made from Korean drama subtitles. Based on the results, the first prominent findings show that ‘an’ and ‘-ji anhta’, and ‘tidak’ and ‘belum’ are equivalent in regards to their usage as negative markers in negative sentences, thus indicating that they correspond to each other. The findings show that among 1,104 data of Korean negative utterances, only 770 data (69.75%) are translated into Indonesian negative utterances, and 334 data (30.25%) are translated into affirmative utterances by applying alternative translation strategies. The second prominent findings reveal that they are non-equivalent due to the shifts that happen, thus showing that they do not correspond to each other. There are three types of shifts causing the non-equivalence, namely (1) shift from negative verbal into negative nominal forms (32 data, 0.90%), (2) shift from negative declarative into negative imperative forms (36 data, 3.26%), and (3) shift from negative into affirmative forms (334 data, 30.25%). In general, the shifts occur in the subtitle translation because Korean and Indonesian languages belong to different categories in terms of language family and have different ways to express negative meanings.

18

Using a Parallel Corpus to Study the Translation of Personal Pronouns

WEN Ting-hui

한국통역번역학회 FORUM Volume.10 No.2 2012.12 pp.187-204

※ 기관로그인 시 무료 이용이 가능합니다.

5,200원

Cette recherche actuelle essaie d’utiliser un corpus parallel pour étudier le phenomène que le pourcentage de pronoms personnels est relativement plus élevé dans le texte chinois traduit que non-traduit. Il y a deux hypothèses pour expliquer cette phenomème. La première est l’ingérence du texte de source. Le chinois est un language du « pro-drop », les pronoms, que ce soient sujets ou objets, peut etre négliges dans des proposition définitives ou dans d’autre phrases. En anglais, on ne peut pas négliger un sujet ou un object d’une propositions définitives. On pense que le subcorpus traduit du texte de source anglais a eu quelque influences sur leur traductions. La seconde hypothèse est la stradégie d’explicationadoptée par les traducteurs. Hu Xianyao(2006) a remarquéle même phenomène dans sa recherche, mais il réclame que le pourcentage plus élevé de pronoms dans les textes traduits est une manifestation d’explicitation. Ce projet y compris un subcorpus de textes de source englais a former un parallel corpus pour mesurer le pourcentage de différents types de pronom personels dans the textes traduit et non-traduits. L’analyse de resultats montre que c’est au niveau d’explicitation et de language du « pro-drop ».

19

A Comparative Study on Chinese ‘有(you)’ and Korean ‘있다(issta)’ with Korean-Chinese Parallel Corpus KCI 등재후보

Jiang Linghui

한국언어연구학회 언어학연구 제19권 3호 2014.12 pp.255-280

※ 기관로그인 시 무료 이용이 가능합니다.

6,400원

Jiang Linghui. 2014. A Comparative Study on Chinese ‘有(you)’ and Korean ‘있다(issta)’ with Korean-Chinese Parallel Corpus. Journal of Linguistic Studies 19(3), 255-280. This paper attempts to make a comparative study on Chinese ‘有(you)’ and Korean ‘있다(issta)’ through Korean-Chinese parallel corpus by implementing the usage analysis of sentences. In Chinese, ‘有(you)’ has many meanings, and the most frequently used is ‘to possess' or 'to own’ and ‘to exist’. Generally, Chinese-studying students think Chinese ‘有(you)’ is the same as Korean ‘있다(issta)’. It is because in many cases Chinese ‘有(you)’ is directly interpreted or translated to Korean ‘있다(issta)’ as a one-to-one correspondence. This paper reorganized the meanings of Chinese ‘有(you)’ through the existing dictionaries and grammar books, observed and analyzed the meanings by the sentences which were extracted from Korean-Chinese parallel corpus, and contrasted Chinese ‘有(you)’ with Korean ‘있다(issta)’ based on these sentences. (Yonsei University)

20

한국어 과거시제 표지의 병렬 말뭉치 기반 분석 KCI 등재

송상헌

고려대학교 언어정보연구소 언어정보 제20호 2015.03 pp.75-104

※ 기관로그인 시 무료 이용이 가능합니다.

7,000원

In human language, a single linguistic form can be used to convey different meanings. One of the most representative forms which involve such a relation between form and meaning in Korean is the verbal inflectional morpheme (e/a)ss. This morpheme is responsible for the past tense by default, but in more than a few cases it does not necessarily denote an event that happened in the past. This article probes into the past tense morpheme in Korean (e/a)ss in a way of comparative corpus linguistics. In order to create the findings using a data-based method, the present study explores a bilingual parallel corpus in which a sentence in one language is aligned to the corresponding sentence in the other language. The parallel data the current work makes use of is the Sejong English-Korean Bilingual Corpus. Exploring the parallel data, this corpus study provides a quantitative analysis of (i) which linguistic form in English the past tense morpheme in Korean corresponds to and (ii) which verbs are more frequently associated with the non-past meanings when they are combined with (e/a)ss.

 
1 2 3 4 5
페이지 저장