년 - 년
Ambiguity Resolution of Complement Transfer in Chinese‐to‐Korean Machine Translation
한국어정보학회 한국어정보학 제9권 1호 2007.06 pp.21-29
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
A Chinese complement construction composed of a head and its following complement contains distinctive linguistic features, while most other languages like Korean, English, etc. have no such features. The general classification of a complement construction by Chinese linguists is helpful but not enough to be completely transferred to a target language such as Korean. In this paper, we propose an approach where complement constructions are analyzed and classified based on the contrastive and computational viewpoint of the target language in Chinese‐to‐Korean machine translation. Further, the translation ambiguities of a head‐ndcomplement dependent case, which is the most complex type of complement construction in terms of translation, are resolved by k‐means clustering using mutual information. Some experimental results of such resolved ambiguities are provided from random samples. The proposed method improved the accuracy of disambiguation by 17.3 % with respect to the baseline.
기계번역ㆍ인간번역ㆍ트랜스크리에이션의 문체 비교 : 광고 번역을 중심으로 KCI 등재
한국외국어대학교 통번역연구소 통번역학연구 제21권 2호 2017.04 pp.163-188
※ 기관로그인 시 무료 이용이 가능합니다.
6,400원
Neural Machine Translation(NMT) has greatly improved machine translation quality, further signaling wider use of MT in the human translation process. As the localization industry evolves, translation production process becomes segmented into draft translation by translator, review by external reviewer and editing by client. In particular, changes in translation ecosystem have led to switching translation service object from technical contents to creative web contents. Among text types that currently flow in, commercial advertizement is assumed to shed light on style difference between machine translation and human translation. Therefore, this study explores style shifts in cosmetic commercials by translation mode such as machine translation, human translation and trans-creation, employing methodological approach from systemic functional language(SFL) - transitivity and register. Except for information freshly added or deleted, trans-created version finalized by a client organization shows the same features as human translation performed by an individual translator within a range that TT can be compared to ST. Meanwhile, machine translation and human translation are distinctly different in transitivity and register, indicating that both linguistic features are two important factors to determine translation style.
생성형 AI와 기계번역 - 챗GPT 번역을 통한 한일통역교육 고찰 KCI 등재
한국외국어대학교 통번역연구소 통번역학연구 제27권 3호 2023.08 pp.27-56
※ 기관로그인 시 무료 이용이 가능합니다.
7,000원
With the popular uptake of ChatGPT, the interpreter community is set to undergo significant changes in its discussion and directions on topics related to machine translation. This study investigates the features and differences of ChatGPT, a natural language processing (NLP) tool, in comparison to Google Translate’s translation engine, specifically between the Korean and Japanese language pair, with the purpose of exploring means to leverage ChatGPT’s capabilities in the context of interpreter training in said pair. For this purpose, we will first analyze ChatGPT’s translation of idioms and proverbs to identify whether it is capable of paraphrasing polysemous connotations and implications. We will next analyze ChatGPT’s ability to translate unstructured colloquial utterances. Finally, we will analyze whether ChatGPT’s translation output displays translation universals such as simplification and clarification. Based on these analyses, this study will discuss how ChatGPT may be used toward interpreter training.
인공지능 시대의 한국어 교육에서 기계번역 활용 가능성-중한 번역을 중심으로 KCI 등재
한중인문학회 한중인문학연구 제69집 2020.12 pp.111-132
※ 기관로그인 시 무료 이용이 가능합니다.
5,800원
본고는 한국어 교육의 관점에서 인공지능 시대 신속하게 발전하는 기계번역 기술의 활용가능성을 탐구하고자 하였다. 한국어 교육에서 번역은 교수법으로서의 번역 방법과 교육 대상으로서의 번역 교육으로 나눌 수 있는데, 전자는 언어 인식의 고양, 후자는 번역 수행에 수반되는 종합적인 능력 함양을 지향한다. 교육 현장에서 지향하고 있는 이러한 번역의 교육적 목표는 기계번역의 결과물에 대한 후속 편집 작업을 통해 효과적으로 성취할 수 있는 것이다. 왜냐하면 기계번역은 어느 정도의 정확성을 확보할 수 있는 결과물을 신속하게 산출할 수 있지만 현재까지는 명백한 한계를 보여주고 있기 때문이다. 또한 기계번역을 비판적으로 수용하고 창조적으로 활용할 수 있는 능력의 함양은 인공지능과의 공존 관계를 구성해야 하는 지금의 시대적 요구이기도 하다. 대상 텍스트 유형에 따른 기계번역의 특성과 한계점을 살펴보기 위해 본고는 정보적 텍스트, 호소적 텍스트, 표현적 텍스트에 대한 기계번역과 인간번역의 결과물을 어휘적, 문법적, 화용적 층위에서 비교 분석하였다. 정보적 텍스트의 경우, 기계번역은 인간번역의 결과물과 유사한 어휘, 문법 등을 사용하여 비교적 높은 수준의 번역문을 생산하였으나 유창성과 전달력이 다소 부족한 면을 보여주었다. 호소적 텍스트의 경우, 기계번역은 다양한 맥락이 존재하는 텍스트의 의미를 제대로 파악하지 못해 오류가 여러 층위에서 발생하였다. 표현적 텍스트의 경우, 기계번역은 결과물의 완성도 측면에서 인간번역과 비교적 큰 차이를 보여주었으며 전반적인 수정 작업이 요구된다. 이와 같은 결과를 고려하여 기계번역의 교육적 활용에 대해 번역 목적, 텍스트 유형, 학습자 수준 등 세 가지 측면에서 제안하였다.
This study aims to explore the possibility of educational use of Korean language in machine translation technology that develops rapidly in the age of artificial intelligence. Translation in Korean language can be divided into translation methods as a teaching method and translation education as a subject of education, with the former aiming to enhance language awareness and the latter to foster comprehensive ability involved in translation performance. the former aims to enhance language awareness and the latter to develop comprehensive skills involved in translation performance. The educational objective of translation is to be achieved effectively through post-editing work on the results of machine translation. This is because machine translation can quickly produce results that can ensure some degree of accuracy, but so far it shows clear limitations. Furthermore, the cultivation of the ability to accept machine translation critically, and utilize it creatively is also a demand of the present era to form a coexistence relationship with artificial intelligence. To look at the characteristics and limitations of machine translation according to the type of source text, this study compared and analyzed the results of machine translation and human translation on lexical, grammatical and pictorial layers for informative text, operative text and expressive text. Considering the result, suggestions were made on the use according to the translation purpose, text type and learners’ level.
Coronavirus Disease-19(COVID-19)에 특화된 인공신경망 기계번역기 KCI 등재
한국융합학회 한국융합학회논문지 제11권 제9호 2020.09 pp.7-13
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 세계보건기구(WHO)의 Coronavirus Disease-19(COVID-19)에 대한 팬데믹 선언으로 COVID-19는 세계적인 관심사이며 많은 사망자가 속출하고 있다. 이를 극복하기 위하여 국가 간 정보 교환과 COVID-19 관련 대응 방안 등의 공유에 대한 필요성이 증대되고 있다. 하지만 언어적 경계로 인해 원활한 정보 교환 및 공유가 이루어지지 못하고 있는 실정이다. 이에 본 논문은 COVID-19 도메인에 특화 된 인공신경망 기반 기계번역(Neural Machine Translation(NMT)) 모델을 제안한다. 제안한 모델은 영어를 중심으로 프랑스어, 스페인어, 독일어, 이탈리아어, 러시아 어, 중국어 지원이 가능한 Transformer 기반 양방향 모델이다. 실험결과 BLEU 점수를 기준으로 상용화 시스템과 비교 하여 모든 언어 쌍에서 유의미한 높은 성능을 보였다.
With the recent World Health Organization (WHO) Declaration of Pandemic for Coronavirus Disease-19 (COVID-19), COVID-19 is a global concern and many deaths continue. To overcome this, there is an increasing need for sharing information between countries and countermeasures related to COVID-19. However, due to linguistic boundaries, smooth exchange and sharing of information has not been achieved. In this paper, we propose a Neural Machine Translation (NMT) model specialized for the COVID-19 domain. Centering on English, a Transformer based bidirectional model was produced for French, Spanish, German, Italian, Russian, and Chinese. Based on the BLEU score, the experimental results showed significant high performance in all language pairs compared to the commercialization system.
서로 다른 문장 구조의 병렬 말뭉치 통합을 통한 기계번역 모델 품질의 향상 KCI 등재
한국경영정보학회 경영정보학연구 제26권 제3호 2024.08 pp.103-117
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
최근 AI 기술이 빠르게 발전하면서 이전에는 개발하기 어려웠던 번역기를 민간에서도 비교적 쉽게 만들 수 있게 되었고, 일반적으로 학습 데이터의 양을 늘릴 경우 번역 품질은 향상되는 경향을 보였다. 하지만 뉴스 데이터로 학습된 기계번역 모델은 동일한 뉴스 데이터를 추가 학습해도 정형화되어 있지 않은 뉴스 데이터의 특성으로 인해 번역 모델의 품질 향상 폭이 크지 않다. 이에 본 연구에서는 이러한 뉴스 데이터가 가진 구조적 한계점을 보완하기 위해 정형화된 문장 구조를 가진 특허 데이터를 기계학습 시 학습 데이터에 추가하여 번역 품질을 향상시키고자 하였다. 현재 다양한 문장 구조를 가진 학습 데이터를 조합하여 기계번역 품질을 향상시키는 연구는 많이 이루어지지 않았으며, 대부분의 연구는 학습 데이터 자체의 품질이나 오류율을 최소화하는 데 중점을 두고 있다. 이를 위해 본 연구는 다양한 문장 구조를 가진 뉴스 학습 데이터와 정형화된 문장 구조를 가진 특허 학습 데이터의 비율을 조정하여 다양한 번역 모델을 생성하였고, 생성된 번역 모델의 품질 변화에 대한 분석을 수행하였다. 실험 결과, 뉴스 데이터와 특허 데이터의 비율을 2:8로 조정한 학습 데이터로 생성한 모델의 품질이 가장 좋게 나타났으며, 뉴스 데이터로만 학습한 모델 대비 66.7% 높은 품질을 보이는 것으로 나타났다.
Recent advances in AI technology have rapidly made it relatively easy for the public to develop translation systems that were previously difficult to create. Generally, increasing the amount of training data has tended to improve translation quality. However, machine translation models trained on news data do not show significant improvements in translation model quality even when additional news data is used for training, due to the unstructured nature of news data. In this study, we aimed to enhance translation quality by supplementing training data with patent data that has structured sentence patterns to address these structural limitations of news data. Research on improving machine translation quality by combining training data with various sentence structures is not extensively conducted, with most focusing on minimizing the quality or error rate of the training data itself. To address this, we generated various translation models by adjusting the ratio of news training data with structured patent training data and analyzed the quality changes of the generated translation models. Experimental results showed that the model trained with a 2:8 ratio of news data to patent data exhibited the highest quality, demonstrating a 66.7% improvement compared to models trained only on news data.
병렬 말뭉치 필터링을 적용한 Filter-mBART기반 기계번역 연구 KCI 등재
한국융합학회 한국융합학회논문지 제12권 제5호 2021.05 pp.1-7
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최신 기계번역 연구 동향을 살펴보면 대용량의 단일말뭉치를 통해 모델의 사전학습을 거친 후 병렬 말뭉치로 미세조정을 진행한다. 많은 연구에서 사전학습 단계에 이용되는 데이터의 양을 늘리는 추세이나, 기계번역 성능 향상을 위해 반드시 데이터의 양을 늘려야 한다고는 보기 어렵다. 본 연구에서는 병렬 말뭉치 필터링을 활용한 mBART 모델 기반의 실험을 통해, 더 적은 양의 데이터라도 고품질의 데이터라면 더 좋은 기계번역 성능을 낼 수 있음을 보인다. 실험결과 병렬 말뭉치 필터링을 거친 사전학습모델이 그렇지 않은 모델보다 더 좋은 성능을 보였다. 본 실험결과를 통 해 데이터의 양보다 데이터의 질을 고려하는 것이 중요함을 보이고, 해당 프로세스를 통해 추후 말뭉치 구축에 있어 하나의 가이드라인으로 활용될 수 있음을 보였다.
In the latest trend of machine translation research, the model is pretrained through a large mono lingual corpus and then finetuned with a parallel corpus. Although many studies tend to increase the amount of data used in the pretraining stage, it is hard to say that the amount of data must be increased to improve machine translation performance. In this study, through an experiment based on the mBART model using parallel corpus filtering, we propose that high quality data can yield better machine translation performance, even utilizing smaller amount of data. We propose that it is important to consider the quality of data rather than the amount of data, and it can be used as a guideline for building a training corpus.
언어적 특성과 서비스를 고려한 딥러닝 기반 한국어 방언 기계번역 연구 KCI 등재
한국융합학회 한국융합학회논문지 제13권 제2호 2022.02 pp.21-29
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 방언 연구, 보존, 의사소통의 중요성을 바탕으로 소외될 수 있는 방언 사용자들을 위한 한국어 방언 기계번역 연구를 진행하였다. 사용한 방언 데이터는 최상위 행정구역을 기반으로 배포된 AIHUB 방언 데이터를 사용하였다. 방언 데이터를 바탕으로 Transformer 기반의 copy mechanism을 적용하여 방언 기계번역기의 성능 향상을 도모하는 모델링 연구와 모델 배포의 효율성을 도모하는 Many-to-one 기반의 방언 기계 번역기를 제안한다. 본 논문은 one-to-one 모델과 many-to-one 모델의 성능을 비교 분석하고 이를 다양한 언어학적 시각으로 분석 하였다. 실험 결과 BLEU점수를 기준으로 본 논문이 제안하는 방법론을 적용한 one-to-one 기계번역기의 성능 향상과 many-to-one 기계번역기의 유의미한 성능을 도출하였다.
Based on the importance of dialect research, preservation, and communication, this paper conducted a study on machine translation of Korean dialects for dialect users who may be marginalized. For the dialect data used, AIHUB dialect data distributed based on the highest administrative district was used. We propose a many-to-one dialect machine translation that promotes the efficiency of model distribution and modeling research to improve the performance of the dialect machine translation by applying Copy mechanism. This paper evaluates the performance of the one-to-one model and the many-to-one model as a BLEU score, and analyzes the performance of the many-to-one model in the Korean dialect from a linguistic perspective. The performance improvement of the one-to-one machine translation by applying the methodology proposed in this paper and the significant high performance of the many-to-one machine translation were derived.
한국어 인공신경망 기계번역의 서브 워드 분절 연구 및 음절 기반 종성 분리 토큰화 제안 KCI 등재
한국융합학회 한국융합학회논문지 제12권 제3호 2021.03 pp.1-7
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
인공신경망 기계번역(Neural Machine Translation, NMT)은 한정된 개수의 단어만을 번역에 이용하기 때문 에 사전에 등록되지 않은 단어들이 입력으로 들어올 가능성이 있다. 이러한 Out of Vocabulary(OOV) 문제를 완화하고자 고안된 방법이 서브 워드 분절(Subword Tokenization)이며, 이는 문장을 단어보다 더 작은 서브 워드 단위로 분할하여 단어를 구성하는 방법론이다. 본 논문에서는 일반적인 서브 워드 분절 알고리즘들을 다루며, 나아가 한국어의 무한한 용언 활용을 잘 다룰 수 있는 사전을 만들기 위해 한국어의 음절 중 종성을 분리하여 서브 워드 분절을 학습하는 새로운 방법론을 제안한다. 실험결과 본 논문에서 제안하는 방법론이 기존의 서브 워드 분리 방법론보다 높은 성능을 거두었다.
Since Neural Machine Translation (NMT) uses only a limited number of words, there is a possibility that words that are not registered in the dictionary will be entered as input. The proposed method to alleviate this Out of Vocabulary (OOV) problem is Subword Tokenization, which is a methodology for constructing words by dividing sentences into subword units smaller than words. In this paper, we deal with general subword tokenization algorithms. Furthermore, in order to create a vocabulary that can handle the infinite conjugation of Korean adjectives and verbs, we propose a new methodology for subword tokenization training by separating the Jongsung(coda) from Korean syllables (consisting of Chosung-onset, Jungsung-neucleus and Jongsung-coda). As a result of the experiment, the methodology proposed in this paper outperforms the existing subword tokenization methodology.
공공 한영 병렬 말뭉치를 이용한 기계번역 성능 향상 연구 KCI 등재
한국디지털정책학회 디지털융복합연구 제18권 제6호 2020.06 pp.271-277
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
기계번역이란 소스언어를 목적언어로 컴퓨터가 번역하는 소프트웨어를 의미하며 규칙기반, 통계기반 기계번역 을 거쳐 최근에는 인공신경망 기반 기계번역에 대한 연구가 활발히 이루어지고 있다. 인공신경망 기계번역에서 중요한 요소 중 하나로 고품질의 병렬 말뭉치를 뽑을 수 있는데 이제까지 한국어 관련 언어쌍의 고품질 병렬 코퍼스를 구하기 쉽지 않은 실정이었다. 최근 한국정보화진흥원의 AI HUB에서 고품질의 160만 문장의 한-영 기계번역 병렬 말뭉치를 공개하였다. 이에 본 논문은 AI HUB에서 공개한 데이터 및 현재까지 가장 많이 쓰인 한-영 병렬 데이터인 OpenSubtitles와 성능 비교를 통해 각각의 데이터의 품질을 검증하고자 한다. 테스트 데이터로 한-영 기계번역 관련 공식 테스트셋인 IWSLT에서 공개한 테스트셋을 이용하여 보다 객관성을 확보하였다. 실험결과 동일한 테스트셋으로 실험한 기존의 한-영 기계번역 관련 논문들보다 좋은 성능을 보임을 알 수 있었으며 이를 통해 고품질 데이터의 중요성 을 알 수 있었다.
Machine translation refers to software that translates a source language into a target language, and has been actively researching Neural Machine Translation through rule-based and statistical-based machine translation. One of the important factors in the Neural Machine Translation is to extract high quality parallel corpus, which has not been easy to find high quality parallel corpus of Korean language pairs. Recently, the AI HUB of the National Information Society Agency(NIA) unveiled a high-quality 1.6 million sentences Korean-English parallel corpus. This paper attempts to verify the quality of each data through performance comparison with the data published by AI Hub and OpenSubtitles, the most popular Korean-English parallel corpus. As test data, objectivity was secured by using test set published by IWSLT, official test set for Korean-English machine translation. Experimental results show better performance than the existing papers tested with the same test set, and this shows the importance of high quality data.
6,100원
In this study, we discuss the basic technology of machine learning of the deep neural network for natural language processing(NLP). We explain the distributed vector representation of words. Distributed vector representation is proved to be able to carry semantic meanings and are useful in various NLP tasks. The recurrent neural network(RNN) is employed to get the vector representation of sentences. We discuss the RNN encoder-decoder model and some modifications of the RNN structure to improve the accuracy of the machine translations. To test and verify the accuracy of Google translator, we performed the translation among Korean, English, and Japanese, and examined the meaning change between the original and the translated sentence. In neural network translation, we showed some inaccuracies of the translation such as wrong relation between subject and object, or some omission or repetition of the original meaning. In order to increase the performance and accuracy of machine translation, it is necessary to acquire more data for training.
AI기계번역을 활용한 외국어교육의 방향성 - 포스트에디팅 수업 사례를 통해 - KCI 등재
한국언어연구학회 언어학연구 제26권 1호 2021.04 pp.23-42
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
Recently, the improvement of artificial intelligence network translators has emerged as a social issue, and vague anxiety and expectations are coexisting due to AI that goes far beyond human capabilities. Research on the accuracy and quality of machine translation is underway due to the increase in learners learning foreign languages using artificial intelligence machine translation. Existing research focused on error analysis, but this study was conducted as a post-editing activity that learners correct and improve while comparing the results of machine translation of Google and Papago with human translation. This study focuses on how translation education should be changed as a foreign language learning, and look at the direction of AI machine translation and foreign language education, such as whether post-editing education is possible as a foreign language learning and whether it helps improve foreign language skills.
인공신경망 기계번역에서 말뭉치 간의 균형성을 고려한 성능 향상 연구 KCI 등재
한국융합학회 한국융합학회논문지 제12권 제5호 2021.05 pp.23-29
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 딥러닝 기반 자연언어처리 연구들은 다양한 출처의 대용량 데이터들을 함께 학습하여 성능을 올리고자 하는 연구들을 진행하고 있다. 그러나 다양한 출처의 데이터를 하나로 합쳐서 학습시키는 방법론은 성능 향상을 막게 될 가능성이 존재한다. 기계번역의 경우 병렬말뭉치 간의 번역투(의역, 직역), 어체(구어체, 문어체, 격식체 등), 도메인 등의 차이로 인하여 데이터 편차가 발생하게 되는데 이러한 말뭉치들을 하나로 합쳐서 학습을 시키게 되면 성능의 악영 향을 미칠 수 있다. 이에 본 논문은 기계번역에서 병렬말뭉치 간의 균형성을 고려한 Corpus Weight Balance (CWB) 학습 방법론을 제안한다. 실험결과 말뭉치 간의 균형성을 고려한 모델이 그렇지 않은 모델보다 더 좋은 성능을 보였다. 더불어 단일 말뭉치로도 고품질의 병렬 말뭉치를 구축할 수 있는 휴먼번역 시장과의 상생이 가능한 말뭉치 구축 프로세 스를 추가로 제안한다.
Recent deep learning-based natural language processing studies are conducting research to improve performance by training large amounts of data from various sources together. However, there is a possibility that the methodology of learning by combining data from various sources into one may prevent performance improvement. In the case of machine translation, data deviation occurs due to differences in translation(liberal, literal), style(colloquial, written, formal, etc.), domains, etc. Combining these corpora into one for learning can adversely affect performance. In this paper, we propose a new Corpus Weight Balance(CWB) method that considers the balance between parallel corpora in machine translation. As a result of the experiment, the model trained with balanced corpus showed better performance than the existing model. In addition, we propose an additional corpus construction process that enables coexistence with the human translation market, which can build high-quality parallel corpus even with a monolingual corpus.
AI시대 강의실의 코끼리 : 기계번역 문해력 교육을 위한 제언 KCI 등재
동국대학교 번역학연구소 번역∙언어∙기술 제6권 2025.02 pp.67-86
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
Today human-machine collaboration for a translation is urgently needed. Due to the ease and convenience of using machine translation, its use has become routine for the general public, and as a result, translation is no longer recognized as a specialized field for the professional translators. In addition to translation, learners in all fields are increasingly using convenient and free online translators to translate, read, and write. However, despite being fully aware of the situation, educators in all fields have been mostly negative and passive in their responses. Therefore, it is clear that machine translation has become a taboo topic in today’s university classrooms, especially in graduate schools of translation and interpreting, and are marginalized from formal discussions. In this context, the machine translation literacy instruction needs to be introduced into translation education for a sustainable translation profession in the age of AI. And future education of machine translation literacy should cultivate professional translators who are both informed and critical about various machine translation technologies and machine translation consultants who can utilize their expertise in the use of machine translation to help citizens and clients use machine translation wisely and responsibly.
한국어 뉴스 헤드라인의 구글 번역기 페르시아어 번역 오류에 관한 연구 KCI 등재
국제한국언어문화학회 한국언어문화학 제22권 제1호 2025.03 pp.69-106
※ 기관로그인 시 무료 이용이 가능합니다.
8,200원
본 연구는 한국어 뉴스 헤드라인의 구글 번역기 페르시아어 번역 결과물 에 나타나는 오류를 유형별로 정리하고, 빈번하게 발생하는 오류 유형 및 패턴을 규명하는 것을 목적으로 하였다. 최근 신경망 기계 번역 기술 의 급속한 발전과 함께 기계 번역의 풀질이 과거에 비해 상당히 향상되었 으며, 이에 따른 기계 번역 수요도 증가하였다. 그러나 번역 결과물에서 사소한 오류부터 중대한 오류까지 다양한 문제점이 여전히 남아 있다. 기존 국내 기계 번역 오류 연구는 주로 주요 언어 쌍을 대상으로 진행되 었다. 이에 본 연구에서는 한국어 뉴스 헤드라인의 구글 번역기의 페르 시아어 번역 오류 유형을 분석하였다. 연구 결과, 페르시아어 뉴스 헤드 라인의 특성을 반영하여 인간 번역자 포스트에디팅 수행의 필요성을 시 사하였다.
This study aims to classify the types of errors found in Google Translate’s Persian translations of Korean news headlines and identify the most frequent error types and patterns. With the rapid advancements in neural machine translation technology in recent years, the quality of machine translation has improved significantly, leading to increased demand and usage. However, various issues, ranging from minor inaccuracies to critical errors, continue to persist in machine translation outputs. Previous domestic studies on machine translation errors have primarily focused on major language pairs. In this context, the present study examines the types of errors observed in Google Translate’s Persian translations of Korean news headlines. The findings highlight the necessity of post-editing by human translators to better reflect the linguistic characteristics of Persian news headlines.
4,500원
This study explores the basic concepts and research trends of PEMT (pre-editing and machine translation), and its application in related industries. Although academic and practical attention has been focused on MTPE (machine translation and post-editing), PEMT is also worth exploring as it aims to improve the quality of source text (ST) which can in turn affect that of target text (TT). PEMT simplifies the original text and clarifies unclear representations, making it easier for computers to understand. Most previous studies so far have revealed the quality and effective use of PEMT, supporting its significant impact on the quality and efficiency of machine translation. In translation practice, PEMT is widely used and mainly takes three forms: entrancement of ST expression, grammar, and terminology; production of ST in controlled languages; and technical writing that can improve of the quality of TT by making ST machine-translatable. Given the importance of PEMT, further in-depth research and practical efforts based on developments to date are required to ultimately improve the quality of machine translation.
Natural Language Processing and Machine Linguistic Interpretation
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 7th International Conference on Next Generation Computing 2021 2021.11 pp.235-237
Natural language processing, as an integral part of artificial intelligence technology, has foundations in a variety of disciplines, including linguistics, computer science, and mathematics. Rapid advances in natural language processing provide solid backing for machine translation research. This document first sets out the key concepts and key points of computational linguistics, followed by a brief review of the history and progress of NLP research in the United States and abroad. The document then summarizes the three stages of machine translation as well as the current state of research. Historically, the advancement curves of natural language processing and machine translation have almost coincided, as well as the two complement each other. On this premise, the paper examines NLP applications in machine translation and highlights problems and trends in the fields of artificial intelligence. Finally, the authors examine the link between machine translation and human interpretation in the era of artificial intelligence and speculate on machine translations long term prospects.
한국외국어대학교 통번역연구소 한국외국어대학교 통번역연구소 학술대회 The Roles and Identities of Interpreters and Translators in an Ever-Changing World 2021.01 pp.215-223
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
During the last decades of the 21st century, the increasing discussions about technological advancement, machine learning, artificial intelligence and post-humanism have triggered new areas of research for translation. So to speak, the nature of Translation Studies, as a multidisciplinary field, smooths its path to adapt to the new technological conditions, especially by engaging with issues pertaining to automatic translation and the tools it offers to provide accurate renditions, and therefore perform the same, if not better, role of the human translator. Such a speculation is mostly based on the predominant traditional thinking of translation as the condition that presupposes an inter-lingual transfer of some inherent qualities of the source text. Obviously, there is more to translation than to simply redirect the attention of both the human and machine translator to look as faithfully as possible for the closest match, equivalence or equivalent effect in the target language and culture. Translation is beyond transfer and reproduction; it is intrinsically grounded at the very heart of Hermeneutics, as there will be no translation without hermeneutic understanding/interpretation: an understanding that is explicitly guided by the subjectivity of the translator and his own experiences in his own world. Therefore, the present paper attempts to negotiate the new aesthetics of automatic translation, partly by arguing that machine translation can never account for the problematic that arises as a result of translating, say, religious and literary texts that are always amenable to be interpreted anew. The human translator is the only agent who is able to perform the task of translating texts that are open to different yet conflicting interpretations. To verify this assumption, the paper draws on the problem of (un)translating the Quran, as a religious script, into English. The case of the Quran shows that translation is, par excellence, a matter of human interpretation that is beyond the grasp of machine translation programs.
기계번역 사후교정(Automatic Post Editing) 연구 KCI 등재
한국융합학회 한국융합학회논문지 제11권 제5호 2020.05 pp.1-8
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
기계번역이란 소스문장(Source Sentence)을 타겟문장(Target Sentence)으로 컴퓨터가 번역하는 시스템을 의 미한다. 기계번역에는 다양한 하위분야가 존재하며 APE(Automatic Post Editing)이란 기계번역 시스템의 결과물을 교정하여 더 나은 번역문을 만들어내는 기계번역의 하위분야이다. 즉 기계번역 시스템이 생성한 번역문에 포함되어 있 는 오류를 수정하여 교정문을 만드는 과정을 의미한다. 기계번역 모델을 변경하는 것이 아닌 기계번역 시스템의 결과 문장을 교정하여 번역품질을 높이는 연구분야이다. 2015년부터 WMT 공동 캠페인 과제로 선정되었으며 성능 평가는 TER(Translation Error Rate)을 이용한다. 이로 인해 최근 APE에 모델에 대한 다양한 연구들이 발표되고 있으며 이에 본 논문은 APE 분야의 최신 동향에 대해서 다루게 된다.
Machine translation refers to a system where a computer translates a source sentence into a target sentence. There are various subfields of machine translation. APE (Automatic Post Editing) is a subfield of machine translation that produces better translations by editing the output of machine translation systems. In other words, it means the process of correcting errors included in the translations generated by the machine translation system to make proofreading. Rather than changing the machine translation model, this is a research field to improve the translation quality by correcting the result sentence of the machine translation system. Since 2015, APE has been selected for the WMT Shaed Task. and the performance evaluation uses TER (Translation Error Rate). Due to this, various studies on the APE model have been published recently, and this paper deals with the latest research trends in the field of APE.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.