Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 30
No
1

Principles and Methods of Chinese Proofreading System

Song Rou

한국어정보학회 한국어정보학 제7ㆍ8집 2002.12 pp.37-46

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

2

Word Segmentation on English Learning : Phonetic or Lexical? KCI 등재

Soojung Kim

한국언어과학회 언어과학 제23권 4호 2016.11 pp.227-238

※ 기관로그인 시 무료 이용이 가능합니다.

4,300원

This study investigates what the key factors are on word segmentation on L2(English) learning: phonetic or lexical? Previous researches on the second language learners' abilities to segment L2 speech in terms of phonetic cues focused on the presence or absence of allophonic cues. For example, presence or absence of aspiration or glottalization determines the location of word boundaries. Not taking into consideration other potential cues such as lexical information or word frequency, it is not clear which cues of linguistic subparts play a main role in word segmentation on English learners. To control potential lexical influences, sequences of non-words are examined with strong phonetic cues across word boundaries. If listeners rely on phonetic cues in segmenting words, non-word sequences should be perceived as real word sequences are perceived. We compare the perception rate of non-words with that of real words. It is hypothesized that different perception rates may be resulted from cues other than phonetic ones. 40 Korean college learners of English participated in an English perception task for the strong phonetic cue of aspiration in both real and non-words. The results show that learners' perception was better on real words than on non-words, suggesting lexical information is contributing to English word segmentation on L2 learning.

3

A Well-formed Chinese Lexicon for Word Segmentation and Part-of-speech Tagging

Sun Maosong, Kang Shiyong

한국어정보학회 한국어정보학 제7ㆍ8집 2002.12 pp.27-36

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

4

Ternary Decomposition and Dictionary Extension for Khmer Word Segmentation KCI 등재

Thaileang Sung, Insoo Hwang

한국정보기술응용학회 JITAM Vol.23 No.2 2016.06 pp.11-28

※ 기관로그인 시 무료 이용이 가능합니다.

5,200원

In this paper, we proposed a dictionary extension and a ternary decomposition technique to improve the effectiveness of Khmer word segmentation. Most word segmentation approaches depend on a dictionary. However, the dictionary being used is not fully reliable and cannot cover all the words of the Khmer language. This causes an issue of unknown words or out-of-vocabulary words. Our approach is to extend the original dictionary to be more reliable with new words. In addition, we use ternary decomposition for the segmentation process. In this research, we also introduced the invisible space of the Khmer Unicode (char\u200B) in order to segment our training corpus. With our segmentation algorithm, based on ternary decomposition and invisible space, we can extract new words from our training text and then input the new words into the dictionary. We used an extended wordlist and a segmentation algorithm regardless of the invisible space to test an unannotated text. Our results remarkably outperformed other approaches. We have achieved 88.8%, 91.8% and 90.6% rates of precision, recall and F-measurement.

5

韩国和泰国学习者汉语阅读中断词单位大小的初步研究 KCI 등재

韩美爱, 江新

한국언어연구학회 언어학연구 제23권 1호 2018.04 pp.317-330

※ 기관로그인 시 무료 이용이 가능합니다.

4,600원

This paper investigates the influence of different text presentation on the size of Chinese reading unit with word segmentataion task by Chinese as a second language learners. The participants were 30 Korean and 27 Thai, who were learning Chinese in Beijing Language and Culture University. They were divided into two language proficiency levels (intermediate and advanced) according to their Chinese words size and were required to add spaces in four Chinese passages. The results showed that: (1) the Chinese reading unit was smaller for Korean intermediate learners than it was for Thai learners. However, it was the same for advanced learners of both language backgrounds; (2) for Korean learners, the reading unit increased with the improvement of Chinese proficiency, while it would not change for Thai learners.

6

彝文自动分词系统的设计模型研究

沙马拉毅, 王成平

한국어정보학회 한국어정보학 제10권 1호 2008.06 pp.36-38

※ 기관로그인 시 무료 이용이 가능합니다.

3,000원

Yi Word segmentation has not been well studied yet. This paper introduces the research of language’s automatic word segmentation at first , on this foundation and according to Yi language’s characteristic , puts forward the way which are suiting and designed algorithm choosing for the Yi automatic word segmentation ‘s structure frame and realizing the method , and expounds the appraisal of every function module of Yi automatic word segmentation, test parameter, and discusses the realistic meaning on the development trend of Yi automatic word segmentation .

7

한국어 단어 및 문장 분류 태스크를 위한 분절 전략의 효과성 연구 KCI 등재

김진성, 김경민, 손준영, 박정배, 임희석

한국융합학회 한국융합학회논문지 제12권 제12호 2021.12 pp.39-47

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

효과적인 분절을 통한 양질의 입력 자질 구성은 언어모델의 문장 이해력을 향상하기 위한 필수적인 단계이다. 입력 자질의 품질 제고는 세부 태스크의 성능과 직결된다. 본 논문은 단어와 문장 분류 관점에서 한국어의 언어적 특징을 효과적으로 반영하는 분절 전략을 비교 연구한다. 분절 유형은 언어학적 단위에 따라 어절, 형태소, 음절, 자모 네 가지로 분류하며, RoBERTa 모델 구조를 활용하여 사전학습을 진행한다. 각 세부 태스크를 분류 단위에 따라 문장 분류 그룹과 단어 분류 그룹으로 구분 지어 실험함으로써, 그룹 내 경향성 및 그룹 간 차이에 대한 분석을 진행한다. 실험 결과에 따르면, 문장 분류에서는 자모 단위의 언어학적 분절 전략을 적용한 모델이 타 분절 전략 대비 최대 NSMC: +0.62%, KorNLI: +2.38%, KorSTS: +2.41% 높은 성능을, 단어 분류에서는 음절 단위의 분절 전략이 최대 NER: +0.7%, SRL: +0.61% 높은 성능을 보임으로써, 각 분류 그룹에서의 효과성을 보여준다.

The construction of high-quality input features through effective segmentation is essential for increasing the sentence comprehension of a language model. Improving the quality of them directly affects the performance of the downstream task. This paper comparatively studies the segmentation that effectively reflects the linguistic characteristics of Korean regarding word and sentence classification. The segmentation types are defined in four categories: eojeol, morpheme, syllable and subchar, and pre-training is carried out using the RoBERTa model structure. By dividing tasks into a sentence group and a word group, we analyze the tendency within a group and the difference between the groups. By the model with subchar-level segmentation showing higher performance than other strategies by maximal NSMC: +0.62%, KorNLI: +2.38%, KorSTS: +2.41% in sentence classification, and the model with syllable-level showing higher performance at maximum NER: +0.7%, SRL: +0.61% in word classification, the experimental results confirm the effectiveness of those schemes.

8

A Survey on the Methods for Uyghur Stemming SCOPUS

Abdurahim Mahmoud, Abdusalam Dawu, Peride Tursun, Askar Hamdulla

보안공학연구지원센터(IJCA) International Journal of Control and Automation Vol.9 No.11 2016.11 pp.143-158

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Stemming is a very important natural language processing task that to remove all inflectional affixes from a word, and to lemmatize the remaining part of the word. It is very important in most of the Information Retrieval systems. In this paper we have introduced different stemming algorithms that commonly used in English and other English-Like European languages, and also introduced methods used in Chinese word segmentation, finally we have carried out a more detailed discussions about the methods for Uyghur stemming.

9

Text Recognition Algorithm Based on Text Features SCOPUS

De Li, XueZhe Jin, LiHua Cui

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.11 No.5 2016.05 pp.209-220

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

It is difficult to realize the text watermarking algorithm on natural language, and the format of text watermarking algorithm has poor robustness against format attacks. This paper presents the new text recognition algorithm based on the text feature. The words are segmented and extracted according to the text feature. The feature dimensions are reduced with the technology of LSA and stop-words database. The new similarity method is also defined to determine the threshold in order to detect the watermarking. The experimental results indicate that the proposed algorithm has better operating efficiency and stronger robustness than the previous researches. This algorithm can also handle the text document written in both Chinese and English effectively.

10

In this study, we propose a flexible template-matching algorithm for word segmentation, and structural analysis of features extraction is used for character recognition in the printed Arabic text. The input text image is preprocessed by the binarization and then by morphological operations. A vector quantization of the thinned image (VQTM) is created based on the idea of a freeman chain code tracking method. In the segmentation process, 113 character templates are compared for partially/completely existence in the VQTM. A non-linear filter is applied on the segmented regions to extract the termination and bifurcation features. The spatial distribution of the extracted features and other statistical characteristics are analyzed for the verification of recognition. Experimental results show that the overall recognition rate of the three fonts: Arabic transparent, simplified Arabic and traditional Arabic is 98.63%.

11

On-line Handwritten Uyghur Word Recognition Using Segmentation-Based Techniques

Mayire Ibrayim, Askar Hamdulla

보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.8 No.6 2015.06 pp.51-60

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

An approach to online handwriting word recognition using segmentation-based techniques is presented in this paper. This approach is referred to as lexicon-driven approach because an optimal segmentation is generated for each string in the lexicon. Word recognition problem is transformed into matching optimization problems between the dictionary entry and the handwritten word image. The segmentation processes use these steps such as removing delayed strokes, shape analysis of the stroke trajectory, reconstructing delayed strokes and combining adjacent fragments. Dynamic matching is used to ranking the lexicon entries in order to get best match. A match score is assigned to a segmentation and string by matching each segment to the corresponding character in the string with a character recognition algorithm that returns confidence value for each character class. As a result the performance for lexicons of size 10, 100, 500 and 1000are 93.17%, 70.33%, 59.79%,51.20% and 94.85%, 79.75%, 74.42%, 62.19% for adding distance and normalizing distance respectively.

12

A Methodology for Urdu Word Segmentation using Ligature and Word Probabilities

Khan, Yunus, Nagar, Chetan, Kaushal, Devendra S.

[Kisti 연계] 한국해양공학회 International journal of ocean system engineering Vol.2 No.1 2012 pp.24-31

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper introduce a technique for Word segmentation for the handwritten recognition of Urdu script. Word segmentation or word tokenization is a primary technique for understanding the sentences written in Urdu language. Several techniques are available for word segmentation in other languages but not much work has been done for word segmentation of Urdu Optical Character Recognition (OCR) System. A method is proposed for word segmentation in this paper. It finds the boundaries of words in a sequence of ligatures using probabilistic formulas, by utilizing the knowledge of collocation of ligatures and words in the corpus. The word identification rate using this technique is 97.10% with 66.63% unknown words identification rate.

13

Chinese Word Segmentation

Li, Haizhou, Yuan, Baosheng

[Kisti 연계] 한국언어정보학회 한국언어정보학회 학술대회논문집 1998 pp.212-217

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

14

Ambiguity Resolution in Chinese Word Segmentation

Maosong, Sun, T'sou, Benjamin-K.

[Kisti 연계] 한국언어정보학회 한국언어정보학회 학술대회논문집 1995 pp.121-126

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

A new method for Chinese word segmentation named Conditional F'||'&'||'BMM (Forward and Backward Maximal Matching) which incorporates both bigram statistics (ie., mutual infonllation and difference of t-test between Chinese characters) and linguistic rules for ambiguity resolution is proposed in this paper The key characteristics of this model are the use of: (i) statistics which can be automatically derived from any raw corpus, (ii) a rule base for disambiguation with consistency and controlled size to be built up in a systematic way.

15

Error-Driven Learning of Chinese Word Segmentation

Hockenmaier, Julia, Brew, Chris

[Kisti 연계] 한국언어정보학회 한국언어정보학회 학술대회논문집 1998 pp.218-229

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

16

Ternary Decomposition and Dictionary Extension for Khmer Word Segmentation

Sung, Thaileang, Hwang, Insoo

[Kisti 연계] 한국데이타베이스학회 Journal of information technology applications & management Vol.23 No.2 2016 pp.11-28

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, we proposed a dictionary extension and a ternary decomposition technique to improve the effectiveness of Khmer word segmentation. Most word segmentation approaches depend on a dictionary. However, the dictionary being used is not fully reliable and cannot cover all the words of the Khmer language. This causes an issue of unknown words or out-of-vocabulary words. Our approach is to extend the original dictionary to be more reliable with new words. In addition, we use ternary decomposition for the segmentation process. In this research, we also introduced the invisible space of the Khmer Unicode (char\u200B) in order to segment our training corpus. With our segmentation algorithm, based on ternary decomposition and invisible space, we can extract new words from our training text and then input the new words into the dictionary. We used an extended wordlist and a segmentation algorithm regardless of the invisible space to test an unannotated text. Our results remarkably outperformed other approaches. We have achieved 88.8%, 91.8% and 90.6% rates of precision, recall and F-measurement.

17

The Role of Prosodic Boundary Cues in Word Segmentation in Korean

Kim, Sa-Hyang

[Kisti 연계] 한국음성과학회 말소리와 음성과학 Vol.13 No.1 2006 pp.29-41

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This study investigates the degree to which various prosodic cues at the boundaries of prosodic phrases in Korean contribute to word segmentation. Since most phonological words in Korean are produced as one Accentual Phrase (AP), it was hypothesized that the detection of acoustic cues at AP boundaries would facilitate word segmentation. The prosodic characteristics of Korean APs include initial strengthening at the beginning of the phrase and pitch rise and final lengthening at the end. A perception experiment utilizing an artificial language learning paradigm revealed that cues conforming to the aforementioned prosodic characteristics of Korean facilitated listeners' word segmentation. Results also indicated that duration and amplitude cues were more helpful in segmentation than pitch. Nevertheless, results did show that a pitch cue that did not conform to the Korean AP interfered with segmentation.

18

The Role of Post-lexical Intonational Patterns in Korean Word Segmentation

Kim, Sa-Hyang

[Kisti 연계] 한국음성과학회 말소리와 음성과학 Vol.14 No.1 2007 pp.37-62

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The current study examines the role of post-lexical tonal patterns of a prosodic phrase in word segmentation. In a word spotting experiment, native Korean listeners were asked to spot a disyllabic or trisyllabic word from twelve syllable speech stream that was composed of three Accentual Phrases (AP). Words occurred with various post-lexical intonation patterns. The results showed that listeners spotted more words in phrase-initial than in phrase-medial position, suggesting that the AP-final H tone from the preceding AP helped listeners to segment the phrase-initial word in the target AP. Results also showed that listeners' error rates were significantly lower when words occurred with initial rising tonal pattern, which is the most frequent intonational pattern imposed upon multisyllabic words in Korean, than with non-rising patterns. This result was observed both in AP-initial and in AP-medial positions, regardless of the frequency and legality of overall AP tonal patterns. Tonal cues other than initial rising tone did not positively influence the error rate. These results not only indicate that rising tone in AP-initial and AP_final position is a reliable cue for word boundary detection for Korean listeners, but further suggest that phrasal intonation contours serve as a possible word boundary cue in languages without lexical prominence.

19

Intonational Pattern Frequency of Seoul Korean and Its Implication to Word Segmentation

Kim, Sa-Hyang

[Kisti 연계] 한국음성과학회 말소리와 음성과학 Vol.15 No.2 2008 pp.21-30

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The current study investigated distributional properties of the Korean Accentual Phrase and their implication to word segmentation. The properties examined were the frequency of various AP tonal patterns, the types of tonal patterns that are imposed upon content words, and the average number and temporal location of content words within the AP. A total of 414 sentences from the Read speech corpus and the Radio corpus were used for the data analysis. The results showed that the 84% of the APs contained one content word, and that almost 90% of the content words are located in AP-initial position. When the AP-initial onset was not an aspirated or tense consonant, the most common AP patterns were LH, LHH, and LHLH (78%), and 88% of the multisyllabic content words start with a rising tone in AP-initial position. When the AP-initial onset was an aspirated or tense consonant, the most common AP patterns were HH, HHLH, and HHL (72%), and 74% of the multisyllabic content words start with a level H tone in AP-initial position. The data further showed that 84.1% of APs end with the final H tone. The findings provide valuable information about the prosodic pattern and structure of Korean APs, and account for the results of a previous study which showed that Korean listeners are sensitive to AP-initial rising and AP-final high tones (Kim, 2007). This is in line with other cross-linguistic research which has revealed the correlation between prosodic probability and speech processing strategy.

20

확장된 음절 bigram을 이용한 자동 띄어쓰기 시스템

임동희, 전영진, 김형준, 강승식

[Kisti 연계] 한국정보과학회언어공학연구회 한국정보과학회언어공학연구회 학술대회논문집 2005 pp.189-193

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문은 통계 기반 방법인 음절 bigram을 이용한 자동 띄어쓰기를 기본 방법으로 하고 경우의 수를 세분화한 확장된 음절 bigram을 이용한 공백 확률, 띄어쓰기 통계를 바탕으로 최종 띄어쓰기 임계치 차등 적용, 에러 사전 적용 3가지 방법을 추가로 사용하는 경우 기본적인 방법만을 쓴 경우보다 띄어쓰기 정확도가 향상된다는 것을 확인하였다. 그리고 해당 음절에 대한 bigram이 없는 경우 확장된 음절 unigram을 통해 근사적으로 계산해 데이터부족 문제를 개선하였다. 한국어 말뭉치와 중국어 말뭉치에 대한 실험을 통해 본 논문에서 제안하는 방법이 한국어 자동 띄어쓰기뿐만 아니라 중국어 단어 분리에 적용할 수 있다는 것도 확인하였다.

 
1 2
페이지 저장