년 - 년
Gradation Image Processing for Text Recognition in Road Signs Using Image Division and Merging KCI 등재
한국ITS학회 한국ITS학회논문지 제13권 제2호 통권52호 2014.04 pp.27-33
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
This paper proposes a gradation image processing method for the development of a Road Sign Recognition Platform (RReP), which aims to facilitate the rapid and accurate management and surveying of approximately 160,000 road signs installed along the highways, national roadways, and local roads in the cities, districts (gun), and provinces (do) of Korea. RReP is based on GPS(Global Positioning System), IMU(Inertial Measurement Unit), INS(Inertial Navigation System), DMI(Distance Measurement Instrument), and lasers, and uses an imagery information collection/classification module to allow the automatic recognition of signs, the collection of shapes, pole locations, and sign-type data, and the creation of road sign registers, by extracting basic data related to the shape and sign content, and automated database design. Image division and merging, which were applied in this study, produce superior results compared with local binarization method in terms of speed. At the results, larger texts area were found in images, the accuracy of text recognition was improved when images had been gradated. Multi-threshold values of natural scene images are used to improve the extraction rate of texts and figures based on pattern recognition.
Output, Noticing, and Uptake: Processing of Reformulation and a Model Text in L2 Writing KCI 등재
한국외국어교육학회 외국어교육 제20권 제4호 2013.12 pp.61-89
※ 기관로그인 시 무료 이용이 가능합니다.
6,900원
As a replication study of Hanaoka and Izumi (2012), the present research investigated the process of learners’ noticing of problems in their interlanguage (IL) while they were producing the second language (L2) and their process of relevant input provided in the form of two types of feedback (i.e., reformulation and a model text) in L2 writing. The data were collected from twenty-five university students in Korea, and they engaged in a four-staged picture description task consisting of the writing stage (Stage 1), the comparison stage (Stage 2), the immediate revision stage (Stage 3), and the delayed revision stage (Stage 4). It was found that output triggered learners’ noticing of problems; however, the problems did not always show up overtly in their writing. Some problems were manifested in writing (i.e., overt problems), while some were hidden (i.e., covert problems). A model text and reformulation played somewhat different roles in offering the solutions to the problems: a model text provided more solutions to covert problems while reformulation provided more solutions to overt problems. And, the extent of learners’ noticing was associated with the types of problems (overt vs. covert problems) and feedback (i.e., models vs. reformulation). However, the types of problems and feedback were not related to the rate of incorporation of feedback into subsequent writing.
텍스트에 대한 내용 및 구조 친숙도가 EFL 학습자의 텍스트 처리과정에 미치는 영향 KCI 등재
한국중앙영어영문학회 영어영문학연구 제61권 3호 2019.09 pp.393-415
※ 기관로그인 시 무료 이용이 가능합니다.
6,000원
This study aimed to investigate how Korean college readers attempted to process texts when encountering texts written in their first language or English, using their schematic knowledge. More specifically, it attempted to figure out how those readers utilized their content and formal schema when reading texts which were familiar or unfamiliar with them, in terms of its topic and experience. This research made use of students’ English language proficiency test scores, and two kinds of data from a questionnaire regarding how they perceived new information in the text, and their summaries after they’re done with reading those two texts. The results of data analysis reinforced that background knowledge of Korean EFL readers help themselves understand the Korean text better and more easily than English texts. It has also indicated that all of those students did mention the Korean text was easy to understand because they already had experiences relating with the text and they could use their background knowledge to process it. On the other hand, they struggled while reading the English text with its vocabulary, and topic. It has also been shown that their summaries showed differences in terms of its average length, and the average number of words used in a sentence. Finally, pedagogical implications in using readers’ schema in the reading class were suggested.
자연어처리(NLP) 기반 텍스트 마이닝을 활용한 스타트업 경영성과 연구동향 분석 : RISS DB 중심으로 KCI 등재
대한산업경영학회 산업융합연구(구 대한산업경영학회지) 제23권 제4호 2025.04 pp.1-16
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
본 연구는 자연어 처리(NLP) 기반 텍스트 마이닝 기법을 활용하여 스타트업 경영성과에 대한 국내연구 동향을 체계적 으로 분석하는 것을 목적으로 한다. 한국교육학술정보원(RISS)에서 ‘스타트업 경영성과’, ‘스타트업 경영’을 키워드로, 전 기간의 초록을 웹 크롤링하여 수집하였으며 최종 323편의 문헌이 분석에 활용되었다. 분석과정에서 형태소분석, 단어빈도분석(TF), 역 문서 빈도분석(TF-IDF), N-gram 분석, 네트워크 분석, CONCOR 기법이 적용되었다. 분석결과, ‘창업’, ‘규제’, ‘영향’, ‘혁신’, ‘기술’과 같은 키워드가 핵심 주제로 도출되었으며, 자금 조달, 기술 혁신, 정책 지원, 기업가 정신이 주 연구 분야로 확인되었다. 키워드 간 구조적 연관성을 파악하였으며 7개의 군집을 도출하였다. 본 연구는 데이터 기반 분석을 통해 기존 문헌과의 차별성을 제시하였다. 향후 연구에서는 시계열 분석, 해외연구 동향분석, 머신러닝 기법을 활용한 연구가 필요할 것이다.
This study aims to systematically analyze domestic research trends on startup business performance using text mining techniques based on natural language processing (NLP). Abstracts were collected through web crawling from the RISS (Korea Education and Research Information Sharing Service) database using the keywords “startup business performance” and “startup management” across all available years, resulting in a final dataset of 323 academic papers. During the analysis process, various methods were applied, including morphological analysis, term frequency (TF), term frequency–inverse document frequency (TF-IDF), N-gram analysis, network analysis, and CONCOR clustering. The analysis revealed that core keywords such as “entrepreneurship,” “regulation”, “impact”, “innovation”, and “technology” emerged as central themes. Major research areas included funding, technological innovation, policy support, and entrepreneurship. Furthermore, the structural relationships among keywords were identified, and seven clusters were derived. This study contributes to the literature by offering a data-driven perspective that distinguishes itself from prior qualitative research. Future studies should consider time-series analysis, comparative analysis of international research trends, and the application of machine learning techniques.
감시 응용 프로그램에서 텍스트 인식을 위한 시각적 콘텐츠의 강력한 처리 KCI 등재
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.14 No.1 2018.02 pp.65-77
자연 장면 이미지의 텍스트 정보는 몇 가지 이미지 기반 응용 프로그램에서 필수적 단서입니다. 그러나 자연 장면 이미지에서 정확한 텍스트를 인식하는 것은 조명, 텍스트 비대칭성 및 색상, 크기 글꼴 및 카메라 정렬 변화로 인해 어려운 작업입니다. 이 논문에서는 시각적 감시 응용 프로그램에서 효율적이고 효과적인 텍스트 인식을 위한 최적의 전처리 파이프 라인을 제안합니다. 전처리 단계에는 이미지 디 샘플링, 노이즈 제거, 색상에서 회색 변환 및 적응형 임계 값이 포함됩니다. 또한 최대 안정적 극한 영역(MSER)텍스트 검출 알고리즘을 사용하여 텍스트 인식 모듈로 전달되는 텍스트 영역을 검출했습니다. OCR에 의해 인식된 텍스트는 검색 테이블과 일치되어 가장 가능성 있는 값 이 추출됩니다. 다양한 데이터 세트를 사용하여 Android및 PC플랫폼 모두에 대해 수행된 실험을 통해 제안된 방법 이 감시 응용 프로그램에서 텍스트 인식에 더 적합하다고 결론을 내렸습니다.
Textual information in natural scene images is an essential clue for several image-based applications. However, recognizing accurate text from natural scene images is a challenging task due to variation in illumination, text skewness, and variation in color, size, font, and camera alignment. In this paper, an optimal pre-processing pipeline is proposed for efficient and effective text recognition in visual surveillance applications. The pre-processing steps involve image de-sampling, de-noising, color to gray conversion, and adaptive thresholding. Furthermore, maximally stable extremal region (MSER) text detection algorithm has been used to detect textual regions which are then forwarded to the text recognition module. The recognized text by OCR is matched with a lookup table and the most likely value is extracted. Through experiments conducted for both Android and PC platforms using various datasets, it was concluded that the proposed method is more suitable for text recognition in surveillance applications.
다중 회귀 분석을 이용한 한자 난이도 예측 기법 연구 KCI 등재
국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제19권 제6호 2019.12 pp.219-225
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
한자 급수와 같이 기존 한자 난이도 선정 방식에 문제점이 있다. 실생활에서 쓰이는 한글 단어와 차이가 나며 해당 급수가 실제로 얼마나 많이 쓰이는지 알 수가 없다. 이러한 문제를 해결하기 위해 빈도수를 이용하여 다중 회귀 분석을 이용하여 한자 난이도를 측정한다. 초등 교과서를 기반으로 한자활용빈도수와 한글의미빈도수를 집계한다. 두 빈도수와 획수를 함께 사용하여 설문지를 작성하여 해당 한자의 학습 적정 시기를 답변 받아 이를 회귀에서 사용할 타겟 변수로 이용한다. 단계별 회귀분석을 이용하여 적절한 피처를 선택하고 다중 선형 회귀 분석을 한다. 모델의 R2는 0.1105가 나왔으며 RMSE는 0.1105의 결과가 나왔다.
There is a problem with the existing method of selecting the difficulty levels of Hanja characters. Some Hanja characters selected by the existing methods are different from Sino-Korean words used in real life and it is impossible to know how many times the Hanja characters are used. To solve this problem, we measure the difficulty of Hanja characters using the multiple regression analysis with the frequency as the features. Based on the elementary textbooks, FWS and FHU are counted. A questionnaire is written using the two frequencies and stroke together to answer the appropriate timing of learning the Hanja characters and use them as target variables for regression. Use stepwise regression to select the appropriate features and perform multiple linear regression. The R2 score of the model was 0.1105 and the RMSE was 0.1105.
A Semantic Logic for Noun Interpretation for Automatic Text Processing SCOPUS
보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.8 No.1 2014.01 pp.377-384
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The main object of this study is to offer a clue for semantic interpretation of nouns to facilitate automatic text processes. Unlike thematic roles lexically registered on verbs, the criterions needed to interpret nouns are not definable by distinctive features. Since nouns in same form can bear both generic and existential meaning, its interpretation rather depends upon logical reasoning. This paper starts with the assumption that nouns can bear generic meaning by themselves thanks to its double functions: Feature Set and Sum of individual objects. This means that bare singular or bare plural nouns are principle carriers of generic meaning if there is no hindering constraint. The determiners like articles, in this sense, play a role of semi-conductive barriers that may block the percolation of generic meaning on nouns to the upper projections. The relevant attributes that activate the barrier hood are [Predicate ± permanent], [Article ±strong Numb, ±limited Numb], [Time ±confined], [Space ±confined]. The combination of bare nouns with these features adjoined at syntax decides whether a noun properly bears generic meaning or not.
Pre/post-processing of Text Mining Techniques Improved through Referring to the Trend Dictionary SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.10 2016.10 pp.177-188
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Needs for the collection of data for text mining of the case with many protolanguages or emotional words distributed in SNS and following trends, and the treatment process method of purification of improved data in a previous step are raised. In data collection for mining, the online trend dictionary based on tag was referred and semi-structured data was effectively parsing processed based on tags of dictionaries according to domains of treating languages, and data for analysis was collected. Additionally, there were the cases to show inefficiency in the text processing of the general genre or the limitation of noun extraction, however, it can be suggested as an alternative on searching trend vocabularies which requires the timeliness or the class processing for corpus work of sentiment dictionary.
A Robust and Real Time Approach for Scene Text Localisation and Recognition in Image Processing
보안공학연구지원센터(IJSIP) International Journal of Signal Processing, Image Processing and Pattern Recognition Vol.9 No.1 2016.01 pp.43-50
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The text localization and recognition in real time scene text images is still a big issue in current application. Mobile application and digitization in real world gives a vital and broad impact on real time scene text images. However, the efficiency of recognition rate depends upon the text localization, i.e., higher the purity of text background segmentation and decomposition, higher the rate of accuracy for the image recognition. In this paper, we present a new scene text detection algorithm based on Stroke detection and Hog Transform method. The method introduces an approach for character detection and recognition which combines the advantages of Hog Transform and Connected Component methods. Characters are detected and recognized on the basis of image regions which contain strokes of specific orientations in a specific relative position, where the strokes are efficiently detected by convolving the image gradient field with a set of oriented bar filters. The method was evaluated on a standard dataset consisting mostly real time images where it achieves state-of-the-art results in both text localization and recognition. The results clearly depict the higher bit of accuracy in terms of localization and recognition for a collected dataset.
불용어 시소러스를 이용한 비정형 텍스트 데이터 후처리 방법론에 관한 연구 KCI 등재
국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.9 No.6 2023.12 pp.935-940
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
인공지능과 빅데이터 분석을 위해 웹 스크래핑으로 수집된 대부분의 텍스트 데이터들은 일반적으로 대용량이 고 비정형이기 때문에 빅데이터 분석을 위해서는 정제과정이 요구된다. 그 과정은 휴리스틱 전처리 정제단계와 후처 리 머시인 정제단계를 통해서 분석이 가능한 정형 데이터가 된다. 따라서 본 연구에서는 후처리 머시인 정제과정에서 한국어 딕셔너리와 불용어 딕셔너리를 이용하여 워드크라우드 분석을 위한 빈도분석을 위해 어휘들을 추출하게 되는 데 이 과정에서 제거되지 않은 불용어를 효율적으로 제거하기 위한 “사용자 정의 불용어 시소러스” 적용에 대한 방 법론을 제안하고 R의 워드클라우드 기법으로 기존의 “불용어 딕셔너리” 방법의 문제점을 보완하기 위해 제안된 “사 용자 정의 불용어 시소러스” 기법을 이용한 사례분석을 통해서 제안된 정제방법의 장단점을 비교 검증하여 제시하고 제안된 방법론의 실무적용에 대한 효용성을 제안한다.
Most text data collected through web scraping for artificial intelligence and big data analysis is generally large and unstructured, so a purification process is required for big data analysis. The process becomes structured data that can be analyzed through a heuristic pre-processing refining step and a post-processing machine refining step. Therefore, in this study, in the post-processing machine refining process, the Korean dictionary and the stopword dictionary are used to extract vocabularies for frequency analysis for word cloud analysis. In this process, “user-defined stopwords” are used to efficiently remove stopwords that were not removed. We propose a methodology for applying the “thesaurus” and examine the pros and cons of the proposed refining method through a case analysis using the “user-defined stop word thesaurus” technique proposed to complement the problems of the existing “stop word dictionary” method with R’s word cloud technique. We present comparative verification and suggest the effectiveness of practical application of the proposed methodology.
교양교육에서의 근대 국한문 혼용 신문자료 활용을 위한 텍스트 처리(text processing) 자동화 방식의 고안과 추가적인 과제
[NRF 연계] 한국교양교육학회 교양교육연구 Vol.17 No.5 2023.10 pp.41-52
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
기존에 한국사 교양교육 현장에서 사료 활용의 필요성은 인식되어 왔지만 사료를 곧바로 활용하기에는 난점이 있었다. 특히 학생들에게 직접 사료를 검색하며 살펴보도록 하는 방식을 취하기 위해서는 그 사료가 국역이 제공되는 것이어야 했다. 본 연구에서는 교양교육에서의 근대 국한문 혼용 신문자료 활용을 위한 텍스트 처리(text processing) 자동화 방식의 고안과 추가적인 과제에 관해 검토해보았다. 한국 근대 신문자료는 역사학뿐 아니라 정치학, 사회학 등의 제 영역에 있어 좋은 자료가 되며, 본 연구에서는 대량으로 내려 받은 텍스트를 일괄 처리하는 것에 주안점을 두었다. 근대 국한문 혼용의 신문자료를 교양교육에 활용하고자 할 때 대부분의 체언이 한자로 되어 있고 아래아 등 옛한글 요소가 다수 포함되어 있는 등, 별도의 사료 가공이 전제되지 않는다면 학생들이 독해하기 어려울 수 있다. 자동화 처리가 가능하다면, 특정 토픽과 관련된 기사들을 뽑아낸 후 가독성을 높인 기사를 제공하여 함께 읽어 나가는 방식으로 활용할 수 있고, 또한 툴 자체를 추출하여 학생들에게 활용하도록 제공함으로써 직접 필요한 사료를 찾아보고 활용하게 할 수 있다. 본 연구에서 고안한 자동화 처리 방식은 다음과 같다. 소스 데이터에서 옛한글 부분의 문자코드를 표준화하고 옛 방식의 어조사 등을 현대어 표기로 대체해주는 작업을 수행하도록 했다. 다음으로 한자 부분에 음독 한글 표기를 괄호로 병기해주는 처리를 하였다. 각기 순서들이 차례대로 수행되어 별도의 파일로 저장되도록 하였으며, 결과물을 검토해보면 상당한 가독성의 향상을 확인할 수 있다. 그러나 현 상황에서는 한계도 존재한다. 예를 들어 어떤 한자가 복수의 독음을 가지는 경우 유니코드 대표 음가가 독음으로 달리게 되는데, 그 한자 단어에 맞지 않는 독음일 경우 오히려 독해를 방해하는 경우도 있을 수 있다. 이러한 경우는 수정을 해준 후 교육에 활용해야 한다. 또한 가공하고자 하는 텍스트가 국한문 혼용이기는 하나 한문 표현의 빈도가 높을 경우 독음이 달린다고 하더라도 한문 지식이 없는 학생이 독해하기는 어려울 수 있다. 이러한 경우까지도 포함하여 매끄럽게 변환하기 위해서는 한문 형태소를 분석하는 기술의 개발이 진전될 것을 요한다.
This study aims to develop an automated text processing method for the use of modern Korean-Chinese mixed newspaper materials in liberal arts education. This article describes the process, its results, and additional tasks. In particular, the focus was placed on the batch processing of texts downloaded in large quantities. The problem with the existing computerized and serviced Korean-Chinese mixed texts is that most of the old Korean texts were computerized with PUA codes, which are not currently in Unicode standards. To process or analyze texts in computer language, it is necessary to convert these characters into standard code methods. Based on the standardized data brought in, the work of replacing the old form of words with the current Korean notation was carried out. Finally, in the text, phonological and Korean notation of Chinese characters are added in parentheses. Reviewing the results shows a significant improvement in readability. If you want to use modern Korean and Chinese newspaper materials for liberal arts education, it may be difficult unless a separate feed process is premised. When trying to use modern Korean-Chinese newspaper materials for liberal arts education, it can be difficult due to the problem of notations. Furthermore, this automated form of processing allows instructors to extract articles related to specific topics and to read articles with students that increase their readability. But as things stand, there are limits. For example, if a Chinese character has multiple consonants, there may be cases in which a Chinese character has a reading sound that does not fit the Chinese character word and rather interferes with the reading of it. These should be used for education purposes after correction. In addition, even if it is a Korean-Chinese mixed text and pronunciation is provided here, it is difficult for a student without knowledge of classical Chinese grammar to read that text if a lot of classical Chinese expressions are mixed in. This case ultimately becomes a problem that can be solved by the development of a classical Chinese morpheme analyzer. If progress is made on such things as the construction of the classic Chinese Corpus, it will be of great help in the development of history and liberal arts teaching tools.
Text Processing in a Second Language
[NRF 연계] 연세대학교 인문학연구원 인문과학 Vol.85 2003.12 pp.255-262
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
[NRF 연계] 서울대학교 인지과학연구소 Journal of Cognitive Science Vol.12 No.3 2011.12 pp.261-278
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
We constructed an Associative Concept Dictionary based on the results from large-scale association experiments with participants in order to develop simulation models and systems of understanding natural language. In this paper, we first briefly explain the construction of this dictionary and describe its unique features; we then show how to apply it to text summarization and word sense disambiguation. The text summarization method uses this dictionary for calculating the importance scores of sentences. A Contextual Semantic Network that includes the semantic relations and quantitative distance information among words is constructed as a model of human contextual understanding using this dictionary. We compare the quality of the summarization with that of human participants and that using conventional methods such as term frequencies. Our method shows that the quality of summarization is better than that of conventional methods. The word sense disambiguation method uses a Dynamic Contextual Network Model constructed using this dictionary. An interactive activation method on the network is used in this system as a model for the human dynamic contextual understanding to identify a word’s meaning. The results show that such dynamic features are effective.
[NRF 연계] 미래영어영문학회 영어영문학 Vol.21 No.4 2016.11 pp.367-387
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The present study aimed to investigate young EFL learners’ text recall processing in L1 and L2 and language proficiency, with conducting the on-line test measuring their response time. The research questions focused on young EFL learners’ language processing through the text recall process with several variants such as L1 and L2, song and prose input, correct answers and incorrect answers, and higher language proficiency and lower language proficiency. Forty young EFL learners participated. As a result, the difference of response time between in the correct and incorrect answer conditions when the song input was provided. The facilitation effects were more salient in L1 than in L2 processing, which was supported by investigation into the language proficiency factor, where it was found that the facilitation effects in recall processing were only observed in the higher proficient group. The results indicate that language processing can be influenced by language proficiency.
Text-Driven Multiple-Path Discourse Processing for Descriptive Texts
[Kisti 연계] 한국정보과학회 Journal of electrical engineering and information science Vol.1 No.2 1996 pp.1-8
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper presents a text-driven discourse analysis system, called DPAS. DPAS constructs a discourse structure by weaving together clauses in the text by finding discourse relations between a clause and the clauses in a context. The basic processing model of DPAS is based on the stack based model of discourse analysis suggested by Grosz and Sidner. We extend the model with dynamic programming method to handle various discourse ambiguities effectively and efficiently. We develop the idea of a context space to keep all information of a context. DPAS parses a text by considering all possible discourse relations between a clause and a context. Since different discourse relations may result in different states of a context, DPAS maintains multiple context spaces for an ambiguous text. Since maintaining all interpretations until the whole text is processed requires too much computing resources, DPAS uses the idea of depth-limited search to limit the search space. If there is more than one discourse relation between an input clause and a context, DPAS constructs context spaces one context space for each discourse relation. Then, DPAS applies heuristics to choose the most desirable context space after it processes some more input clauses. Since the basic idea of DPAS is domain independent, although we used descriptive texts to demonstrate DPAS, we believe the idea of DPAS can be extended to understand other styles of texts.
[NRF 연계] 국어교육학회 국어교육학연구 Vol.58 No.5 2023.12 pp.5-40
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This study comprehensively and empirically examined a hypothesized model that incorporates the task model (i.e., planning and revision), documents model (i.e., intertext and integrated mental models), consideration of counterarguments as predictive variables, and text quality of essays as the dependent variable. Data were collected from 421 South Korean high school students who wrote argumentative essays using multiple documents and responded to a series of questionnaires measuring processing variables. Structural equation models were applied to analyze the data. The results indicated that the documents model predicted the participants’ text quality, while the effect of the task model was mediated by the documents model and consideration of counterargument. Notably, the paths in which planning and revision affect text quality are different. Furthermore, the intertext model and consideration of counterarguments, as opposed to the mental model, did not directly affect text quality. These findings not only contribute to the literature on multipledocument processing but also provide empirical evidence to suggest an interdisciplinary approach applicable in both language arts and content subject domains.
The role of topic familiarity in text-based lexical input processing tasks
[NRF 연계] 글로벌영어교육학회 Studies in English Education Vol.14 No.2 2009.12 pp.149-171
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The purpose of this study was to examine whether topic familiarity affects the vocabulary acquisition of L2 learners while reading texts followed by lexical input processing tasks. For this purpose, 98 Korean EFL middle school students were divided into three groups for the lexical input processing tasks, and performed their assigned vocabulary tasks after reading a text for the two experiments. In Experiment 1, the text whose content was familiar to the subjects was employed, and in Experiment 2, the text whose content was not familiar to the subjects was employed. Immediate and delayed posttests which were comprised of word recognition and translation product items, were administered to assess vocabulary acquisition of the subjects. The scores were submitted to an one-way and two-way ANOVA. The results suggest that (a) the gap-fill exercise using sentence context was the most effective for vocabulary gains, (b) there were significant effects of topic familiarity on vocabulary gains through lexical input processing tasks. Expecially the effect of topic familiarity was clear in the delayed posttests, and (c) the effect of topic familiarity is consistent through different lexical input processing tasks. The discussion concerns the significance of findings for the role of topic familiarity.
[NRF 연계] 한국사회과학협의회 Korean Social Science Journal Vol.31 No.1 2004.10 pp.113-136
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This article examines the hegemonic trend in the social treatment and processing of Korean immigration in Buenos Aires, Argentina. It is suggested that Korean immigrants gained visibility --especially during the ‘90s-- through the circulation of public and private discourses that articulated biology, culture and class in a shadowy operation of local racism: the ethnicization of class conflict. A corpus of contemporary everyday discourses on Korean immigrants --namely news stories published in the main national newspapers and informal interviews with non-Korean residents-- is therefore analyzed against the specific background of Argentina’s ethnic formation, its migration history, and the political and economic context informing discursive production. Taking discourse as a situated social practice, this article attempts to contribute not only to the study of racist phenomena, but also to a facet of Korean studies that has lately attracted academic interest: that of the Korean diaspora in Latin America.
고전 텍스트 DB의 활용을 위한 텍스트 처리와 가공 방법 연구
[NRF 연계] 성신여자대학교 인문과학연구소 人文科學硏究 Vol.51 2025.02 pp.1-23
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
이 연구는 기존의 고전 텍스트 데이터베이스(DB)가 가진 디지털 이미지 중심의 서비스가 가지는 한계를 극복하고 전산적인 처리와 분석이 가능한 디지털 고전 텍스트중심의 데이터베이스를 구축하기 위한 다양한 이슈들을 논의하는 데 그 목적이 있다. 즉 디지털 원본 이미지, 디지털 원본 텍스트, 디지털 변환 텍스트, 디지털 주석 텍스트등과 같이 디지털 텍스트의 다양한 이형태들을 종합적으로 구축하여 DB에 포함할 것을 제안하였다. 이를 위해 기존의 DB 현황에 대한 기초적이고 포괄적인 조사를 수행하고, 우선적으로 고전 활자본 소설 DB를 대상으로 구체적인 전처리 원칙과 세부 지침에 따라 디지털 변환, 주석 텍스트를 처리하는 과정에 대해서 논의하였다. 그 결과로 전산적 처리가 가능한, 신뢰할 수 있는 디지털 정본 및 다양한 이본 자료를 확보할수 있음을 보였으며 이를 통해 디지털인문학의 성공적인 사례를 창출할 수 있다고 제안하였다.
This study aims to discuss various issues in overcoming the limitations of existing classical text databases (DB), which are primarily focused on digital images, by establishing a database centered on digital classical texts that can be computationally processed and analyzed. It proposes to comprehensively construct various forms of digital texts, such as digital original images, digital original texts, digital transformed texts, and digital annotated texts, to be included in the database. For this purpose, a basic and comprehensive investigation of the current state of existing DB was conducted. The study specifically targeted the classical printed novel DB to discuss the process of handling digital transformed and annotated texts according to specific preprocessing principles and detailed guidelines. As a result, it demonstrated that reliable digital originals and various versioned materials that are amenable to computational processing could be secured, suggesting that this could create successful examples in the field of digital humanities.
[Kisti 연계] 한국정보과학회언어공학연구회 한국정보과학회언어공학연구회 학술대회논문집 2006 pp.189-196
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문은 시맨틱 웹 응용 서비스를 구현함에 있어 필수적으로 요구되는 온톨로지 인스턴스 구축을 효율적으로 처리하는 데 있어 텍스트 처리 기술이 어떤 역할을 수행할 수 있는 가를 $OntoFrame-K^{(R)}$라는 시맨틱 웹 기반 정보 유통 체계에의 적용 사례를 통해 살펴본다. 본 논문에서 소개하는 텍스트 처리 기술은 개체 확인물 통한 개념 사례화, 주제 분야 할당을 통한 메타데이터 확장에, 그리고 인용 정보 추출 및 인용 관계 구축을 통한 객체 관계속성 구축에 적용된다. 개체 확인에서는 메타데이터 비교 잊 병합을 사용하였으며 이를 기반으로 한 수작업 구축을 통해 8,543명의 인력 URI를 확보하였다. 주제 및 분야 할당에서는 색인어와 분야분류명이 매핑된 시소러스 개념어의 매칭을 통해 색인어 별 TF (Term Frequency), 색인어와 매칭된 개념어 별 TF, 색인어와 매칭된 개념어 별 시소러스에서의 깊이, 색인어와 매칭된 개념어 별 개념 패싯, 색인어와 매칭된 각 개념어에 부착된 분야분류명 목록 등 할당을 위한 다양한 자질을 확보 적용하였다. 인용 정보 추출과 인용 관계 구축에서는 객체 URI와 인력 URI를 기반으로 하여 자동 추출된 인용 정보를 반영하는 방식으로 7,237개 문헌으로부터 총 135개의 인용 네트워크 그룹을 자동으로 확보하였다. 본 연구를 통해 제시된 텍스트 처리 기술의 활용 방안이 향후 시맨틱 웹 응용 서비스 및 인프라 구현에서 다각적으로 활용될 수 있기를 기대한다.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.