Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 88
No
1

자연 장면 이미지의 텍스트 정보는 몇 가지 이미지 기반 응용 프로그램에서 필수적 단서입니다. 그러나 자연 장면 이미지에서 정확한 텍스트를 인식하는 것은 조명, 텍스트 비대칭성 및 색상, 크기 글꼴 및 카메라 정렬 변화로 인해 어려운 작업입니다. 이 논문에서는 시각적 감시 응용 프로그램에서 효율적이고 효과적인 텍스트 인식을 위한 최적의 전처리 파이프 라인을 제안합니다. 전처리 단계에는 이미지 디 샘플링, 노이즈 제거, 색상에서 회색 변환 및 적응형 임계 값이 포함됩니다. 또한 최대 안정적 극한 영역(MSER)텍스트 검출 알고리즘을 사용하여 텍스트 인식 모듈로 전달되는 텍스트 영역을 검출했습니다. OCR에 의해 인식된 텍스트는 검색 테이블과 일치되어 가장 가능성 있는 값 이 추출됩니다. 다양한 데이터 세트를 사용하여 Android및 PC플랫폼 모두에 대해 수행된 실험을 통해 제안된 방법 이 감시 응용 프로그램에서 텍스트 인식에 더 적합하다고 결론을 내렸습니다.

Textual information in natural scene images is an essential clue for several image-based applications. However, recognizing accurate text from natural scene images is a challenging task due to variation in illumination, text skewness, and variation in color, size, font, and camera alignment. In this paper, an optimal pre-processing pipeline is proposed for efficient and effective text recognition in visual surveillance applications. The pre-processing steps involve image de-sampling, de-noising, color to gray conversion, and adaptive thresholding. Furthermore, maximally stable extremal region (MSER) text detection algorithm has been used to detect textual regions which are then forwarded to the text recognition module. The recognized text by OCR is matched with a lookup table and the most likely value is extracted. Through experiments conducted for both Android and PC platforms using various datasets, it was concluded that the proposed method is more suitable for text recognition in surveillance applications.

2

Text Extraction from Complex Natural Images

Kumar, Manoj, Lee, Guee-Sang

[Kisti 연계] 한국콘텐츠학회 International journal of contents Vol.6 No.2 2010 pp.1-5

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The rapid growth in communication technology has led to the development of effective ways of sharing ideas and information in the form of speech and images. Understanding this information has become an important research issue and drawn the attention of many researchers. Text in a digital image contains much important information regarding the scene. Detecting and extracting this text is a difficult task and has many challenging issues. The main challenges in extracting text from natural scene images are the variation in the font size, alignment of text, font colors, illumination changes, and reflections in the images. In this paper, we propose a connected component based method to automatically detect the text region in natural images. Since text regions in mages contain mostly repetitions of vertical strokes, we try to find a pattern of closely packed vertical edges. Once the group of edges is found, the neighboring vertical edges are connected to each other. Connected regions whose geometric features lie outside of the valid specifications are considered as outliers and eliminated. The proposed method is more effective than the existing methods for slanted or curved characters. The experimental results are given for the validation of our approach.

3

Text Location and Extraction for Business Cards Using Stroke Width Estimation

Zhang, Cheng Dong, Lee, Guee-Sang

[Kisti 연계] 한국콘텐츠학회 International journal of contents Vol.8 No.1 2012 pp.30-38

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Text extraction and binarization are the important pre-processing steps for text recognition. The performance of text binarization strongly related to the accuracy of recognition stage. In our proposed method, the first stage based on line detection and shape feature analysis applied to locate the position of a business card and detect the shape from the complex environment. In the second stage, several local regions contained the possible text components are separated based on the projection histogram. In each local region, the pixels grouped into several connected components based on the connected component labeling and projection histogram. Then, classify each connect component into text region and reject the non-text region based on the feature information analysis such as size of connected component and stroke width estimation.

4

Grammatical Structure Oriented Automated Approach for Surface Knowledge Extraction from Open Domain Unstructured Text

Tissera, Muditha, Weerasinghe, Ruvan

[Kisti 연계] 한국정보통신학회 Journal of information and communication convergence engineering Vol.20 No.2 2022 pp.113-124

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

News in the form of web data generates increasingly large amounts of information as unstructured text. The capability of understanding the meaning of news is limited to humans; thus, it causes information overload. This hinders the effective use of embedded knowledge in such texts. Therefore, Automatic Knowledge Extraction (AKE) has now become an integral part of Semantic web and Natural Language Processing (NLP). Although recent literature shows that AKE has progressed, the results are still behind the expectations. This study proposes a method to auto-extract surface knowledge from English news into a machine-interpretable semantic format (triple). The proposed technique was designed using the grammatical structure of the sentence, and 11 original rules were discovered. The initial experiment extracted triples from the Sri Lankan news corpus, of which 83.5% were meaningful. The experiment was extended to the British Broadcasting Corporation (BBC) news dataset to prove its generic nature. This demonstrated a higher meaningful triple extraction rate of 92.6%. These results were validated using the inter-rater agreement method, which guaranteed the high reliability.

5

HTML 논리적 구조분석을 통한 본문추출 알고리즘

전현지, 고찬

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.16 No.3 2015 pp.445-455

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

인터넷과 컴퓨터 기술이 발전함에 따라 정보의 양이 폭발적으로 증가하였으며, 이로 인해 다양한 웹 저작 도구 및 새로운 웹 표준의 출현과 웹에 대한 접근성이 보다 편리해지면서 매우 다양한 종류의 웹 콘텐츠들이 아주 빠르게 생산되고 있다. 하지만 웹 문서는 여러 블록으로 나누어 다양한 주제를 담아내고 있으며, 각각의 블록들이 서로 연관성이 없는 주제를 다루는 경우가 많을 뿐만 아니라 네비게이션, 단순한 장식물, 광고, 저작권 정보 등과 같이 콘텐츠로 볼 수 없는 블록들도 존재한다. 이러한 문제를 해결하기 위해 HTML 웹 문서의 정확한 본문영역만을 추출하여 사용자 요구조건을 충족하고 효과적으로 정보를 학습할 수 있도록 하며, 추후에는 문서를 체계적으로 관리할 수 있게 최적화된 웹 검색 시스템으로서의 재구성 방법을 제안하고자 한다.

According as internet and computer technology develops, the amount of information has increased exponentially, arising from a variety of web authoring tools and is a new web standard of appearance and a wide variety of web content accessibility as more convenient for the web are produced very quickly. However, web documents are put out on a variety of topics divided into some blocks where each of the blocks are dealing with a topic unrelated to one another as well as you can not see with contents such as many navigations, simple decorations, advertisements, copyright. Extract only the exact area of the web document body to solve this problem and to meet user requirements, and to study the effective information. Later on, as the reconstruction method, we propose a web search system can be optimized systematically manage documents.

6

색상 단순화와 윤곽선 패턴 분석을 통한 이미지에서의 글자추출 KCI 등재

양재호, 박영수, 이상훈

한국융합학회 한국융합학회논문지 제8권 제8호 2017.08 pp.33-40

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문은 이미지에서 효과적인 문자검출을 위해 색상단순화 및 윤곽선에서의 패턴 분석을 통한 문자 검출방법을 제안한다. 윤곽선 기반방법을 사용하는 문자검출 알고리즘은 단순한 배경의 이미지에서는 우수한 성능을 보이지만, 복잡한 배경의 이미지에서는 성능이 떨어지는 단점이 있다. 따라서 제안하는 방법은 복잡한 배경에서의 비문자영역을 최소화하기 위해 이미지 단순화 및 패턴분석을 통한 문자 검출 알고리즘을 제안한다. 먼저 이미지에서의 문자영역 부분을 검출하기 위하여 전처리 과정으로 K-means 군집화를 사용하여 이미지의 색상을 단순화하고, 색상 단순화 과정에서의 물체의 경계의 흐릿해짐을 개선하기 위해 고주파통과필터를 통해 물체의 경계를 강화한다. 그 후 모폴로지 기법의 팽창과 침식의 차이를 이용하여 물체의 윤곽선을 검출하고, 획득한 영역의 윤곽선 부분의 정보(높이, 너비 면적)를 구한 후 패턴분석을 통해 조건을 줌으로써 문자 후보영역을 판별하여 문자가 아닌 불필요한 영역(그림, 배경)을 제거한다. 최종 결과로 라벨링을 통해 불필요한 영역이 제거된 결과를 보여준다.

In this paper, we propose a text extraction method by pattern analysis on contour for effective text detection in image. Text extraction algorithms using edge based methods show good performance in images with simple backgrounds, The images of complex background has a poor performance shortcomings. The proposed method simplifies the color of the image by using K-means clustering in the preprocessing process to detect the character region in the image. Enhance the boundaries of the object through the High pass filter to improve the inaccuracy of the boundary of the object in the color simplification process. Then, by using the difference between the expansion and erosion of the morphology technique, the edges of the object is detected, and the character candidate region is discriminated by analyzing the pattern of the contour portion of the acquired region to remove the unnecessary region (picture, background). As a final result, we have shown that the characters included in the candidate character region are extracted by removing unnecessary regions.

7

교통 표지판에서 텍스트 추출 및 기울기 추출

최규담, 김성동, 최기호

한국ITS학회 한국ITS학회 학술대회 2003년 한국ITS학회 정기총회 및 추계학술대회 2003.11 pp.81-84

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

8

자연영상에서 교통 표지판의 기울기 보정 및 텍스트 추출

최규담, 김성동, 최기호

한국ITS학회 한국ITS학회논문지 제3권 제2호 통권5호 2004.12 pp.19-28

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문은 자연영상에서 얻은 교통표지판의 기울기를 보정하고 텍스트를 추출하는 방법을 제안한다. 본 연구는 명도 이미지를 대상으로 모든 과정이 4단계로 이루어진다. 첫째, 자연 영상에서 에지 검출을 위한 전처리 및 Canny 에지 추출을 수행하며, 둘째, 영상의 기울기를 추출하기 위해 허프 변환에 대한 전처리와 후처리를 한 후, 셋째로 잡음영상과 선을 제거하고 텍스트가 가지고 있는 특징을 이용하여 후보영역 검출을 한다 마지막으로 검출된 텍스트 후보영역 안에서 지역적 이진화를 수행한 후, 불필요한 비텍스트 연결 요소를 추려내기 위해 텍스트와 비텍스트 간의 연결요소에 나타나는 특징 차이를 이용하여 텍스트 추출을 수행한다 100장의 샘플영상을 대상으로 실험한 결과 82.54% 텍스트 추출률과 79.69% 추출 정확도를 가짐으로써 기존의 런 길이 평활화 방법이나 퓨리어 변환을 이용한 방법보다 더 정확한 텍스트 추출 향상을 보였다. 또한 기울어진 각도 추출에서도 94.3%의 추출률로 기존의 Hough 변환만을 이용한 방법보다 약 26%의 향상을 보였다. 본 연구는 시각 장애인 보행 보조 시스템이나 무인 자동차 운행에 있어 위치 정보를 제공하는데 활용할수 있을 것이다.

This paper shows how to compensate the skew from the traffic sign included in the natural image and extract the text. The research deals with the Process related to the array image. Ail the process comprises four steps. In the first fart we Perform the preprocessing and Canny edge extraction for the edge in the natural image. In the second pan we perform preprocessing and postprocessing for Hough Transform in order to extract the skewed angle. In the third part we remove the noise images and the complex lines, and then extract the candidate region using the features of the text. In the last part after performing the local binarization in the extracted candidate region, we demonstrate the text extraction by using the differences of the features which appeared between the tett and the non-text in order to select the unnecessary non-text. After carrying out an experiment with the natural image of 100 Pieces that includes the traffic sign. The research indicates a 82.54 percent extraction of the text and a 79.69 percent accuracy of the extraction, and this improved more accurate text extraction in comparison with the existing works such as the method using RLS(Run Length Smoothing) or Fourier Transform. Also this research shows a 94.5 percent extraction in respect of the extraction on the skewed angle. That improved a 26 percent, compared with the way used only Hough Transform. The research is applied to giving the information of the location regarding the walking aid system for the blind or the operation of a driverless vehicle

9

장면 텍스트 영역 추출을 위한 적응적 에지 강화 기반의 기울기 검출 및 보정

백재경, 장재혁, 서영건

[Kisti 연계] 한국디지털콘텐츠학회 디지털콘텐츠학회 논문지 Vol.18 No.4 2017 pp.777-785

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

실세계에서 텍스트가 포함 된 장면은 텍스트를 추출하고 인식하여 많은 정보를 얻을 수 있으므로, 장면의 텍스트 영역을 추출하고 인식하는 기술들은 꾸준히 발전하고 있다. 장면에서 텍스트 영역을 추출하는 기술은 크게 텍스쳐를 기반으로 하는 방법과 연결요소방법, 그리고 이 둘을 적절히 혼합하는 방법들로 구분 할 수 있다. 텍스처를 기반으로 하는 방법은 영상의 색상, 명도 등의 정보를 이용하여 텍스트가 다른 요소와는 다른 값을 갖는다는 것을 기반으로 한다. 연결 요소 방법은 장면의 각 화소마다 인접해 있는 유사 화소를 연결 요소로 만들어 기하학적인 특성을 이용하여 판별한다. 본 논문에서는 텍스트 영역 추출의 정확도를 높이기 위해 영상의 기울기를 검출하고 보정한 후 에지를 적응적으로 변경하는 방법을 제안한다. 제안 방법은 영상의 기울기를 보정한 후 텍스트가 포함 된 정확한 영역만 추출하기 때문에 MSER보다 15%, EEMSER보다 10% 더 정확하게 영역을 얻었다.

In the modern real world, we can extract and recognize some texts to get a lot of information from the scene containing them, so the techniques for extracting and recognizing text areas from a scene are constantly evolving. They can be largely divided into texture-based method, connected component method, and mixture of both. Texture-based method finds and extracts text based on the fact that text and others have different values such as image color and brightness. Connected component method is determined by using the geometrical properties after making similar pixels adjacent to each pixel to the connection element. In this paper, we propose a method to adaptively change to improve the accuracy of text region extraction, detect and correct the slope of the image using edge and image segmentation. The method only extracts the exact area containing the text by correcting the slope of the image, so that the extracting rate is 15% more accurate than MSER and 10% more accurate than EEMSER.

10

자연영상에서 문자의 형태 분석을 이용한 문자영역 추출에 관한 연구 KCI 등재

양재호, 한현호, 김기봉, 이상훈

한국융합학회 한국융합학회논문지 제9권 제11호 2018.11 pp.61-68

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문에서는 일상에서 획득할 수 있는 자연 영상에서 문자를 검출하기 위해 영상 개선 및 문자의 형태를 분석하여 문자를 검출하는 방법을 제안한다. 제안하는 방법은 자연 영상에서 문자로 인식될 영역의 검출률을 향상시키기 위해 객체부 분의 경계를 언샤프 마스크를 사용하여 강조하였다. 향상된 객체의 경계 부분을 이용하여 영상의 문자 후보영역을 MSER(Maximally Stable Extermal Regions)을 이용하여 검출하였다. 검출된 문자 후보영역에서 실제 문자로 판단될 영역 을 검출하기 위해 각 영역들의 형태를 분석하여 글자의 특성을 갖는 영역외의 비 문자영역을 제거하여 실제 문자영역 검출 률을 높였다. 본 논문의 정량적 평가를 위해 문자 영역의 검출률과 정확도를 이용하여 기존의 방법들과 비교하였다. 실험 결과 기존의 문자 검출 방법보다 제안하는 방법이 비교적 높은 문자영역의 검출률 및 정확도를 보였다.

In this paper, we propose a method of character detection by analyzing image enhancement and character type to detect characters in natural images that can be acquired in everyday life. The proposed method emphasizes the boundaries of the object part using the unsharp mask in order to improve the detection rate of the area to be recognized as a character in a natural image. By using the boundary of the enhanced object, the character candidate region of the image is detected using Maximal Stable Extermal Regions (MSER). In order to detect the region to be judged as a real character in the detected character candidate region, the shape of each region is analyzed and the non-character region other than the region having the character characteristic is removed to increase the detection rate of the actual character region. In order to compare the objective test of this paper, we compare the detection rate and the accuracy of the character region with the existing methods. Experimental results show that the proposed method improves the detection rate and accuracy of the character region over the existing character detection method.

11

감정 표현 로봇을 위한 SMS 문장의 감정 추출 방법

최동엽, 박진규, 김태정

한국특허학회 특허학연구 : 한국특허학회지 Vol.18 No.1 통권 44호 2016.06 pp.5-8

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문에서는 감정 표현 로봇에서 사용하기 위한 SMS 문장에 내포되어 있는 감정을 추출하기 위한 방법을 제시하였다. 제시된 방법은 철자 오류나, 동사 형용사의 어미 활용에 대처하기 위한 것으로서, 감 정 추출은 두 단계로 이루어졌다. 첫 번째는, 사용되는 각 단어에 내포되어 있는 감정을 결정하는 단계이 고, 두 번째는 그 결과를 이용하여 문장의 감정을 결정하는 단계이다. 각각의 단계를 위하여 신경 회로망 과 히든 마르코프 모델을 사용하였으며, 16비트 한글 완성형 코드를 사용하였다. 하위 3비트 오차 및 하위 5 비트 오차의 경우에 대하여 감정을 추출하기 위한 시스템을 학습시켰으며, 감정 추출의 성공률은 3비트의 경우 92.25%, 5비트의 경우 95.30%였고, HMM의 추출 성공률은 80.0%였다.

An emotion extraction algorithm from SMS text was developed for the emotion expression robot. The emotion extraction is performed in two steps. First, the emotion of each word is determined in accordance with the data base on the basis of emotion classification system. Secondly, the emotion of a sentence is determined according to the emotions extracted from the words of the sentence. For the first step, artificial neural network was used for the training of words with spelling errors. And for the second step, hidden markov model was used. For the artificial neural network and hidden markov model processing, Korea graphic character set for information interchange was used which is a South Korean 16bit coded character set standard to represent Korean characters on a computer. Two cases were simulated. First, lower 3 bits of 16bits were used to make spelling errors. Second case was with 5bits spelling errors. The matching rates of training data were 98.75% with 3 bits error and 95.74% with 5 bits error with 20K training iteration.. The matching rate of sentence emotion was 80.0% with the trained HMM.

12

특허 문서로부터 키워드 추출을 위한 위한 텍스트 마이닝 기반 그래프 모델 KCI 등재

이순근, 임영문, 엄완섭

대한안전경영과학회 대한안전경영과학회지 제17권 제4호 2015.12 pp.335-342

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

The increasing interests on patents have led many individuals and companies to apply for many patents in various areas. Applied patents are stored in the forms of electronic documents. The search and categorization for these documents are issues of major fields in data mining. Especially, the keyword extraction by which we retrieve the representative keywords is important. Most of techniques for it is based on vector space model. But this model is simply based on frequency of terms in documents, gives them weights based on their frequency and selects the keywords according to the order of weights. However, this model has the limit that it cannot reflect the relations between keywords. This paper proposes the advanced way to extract the more representative keywords by overcoming this limit. In this way, the proposed model firstly prepares the candidate set using the vector model, then makes the graph which represents the relation in the pair of candidate keywords in the set and selects the keywords based on this relationship graph.

13

합성곱 신경망(CNN)은 다수의 2차원 필터가 계층적으로 중첩된 구조를 가지므로, 그 내부 구조와 데이터 흐름을 직관적으로 시각화하는 데 본질적인 어려움이 있다. 기존의 시각화 도구들은 각기 다른 표현 방식과 추상화 수준을 사용함으로써 서로 다른 모델 간 구조를 비교하거나 해석하는 데 한계를 가지며, 이러한 비표준화된 시각화 환경은 연구자가 모델을 이해하고 비교하는 과정에서 인지적 부담을 초래한다. 본 연구에서는 이러한 문제를 개선하기 위해 CNN 아키텍처 설명으로부터 구조 정보를 추출하고, 이를 구조화된 표현으로 변환한 뒤, 상호작용 가능한 웹 기반 3차원 시각화 프레임워크를 제안한다. 제안하는 시스템은 대형 언어 모델(LLM)을 활용하여 논문 텍스트를 JSON 형태의 중간 표현으로 변환하고, 사용자가 해당 프레임워크 상에서 CNN 아키텍처를 확대·축소 및 회전 등의 상호작용을 통해 탐색하고 비교할 수 있도록 한다. 이러한 접근은 모델 구현 파일에 의존하지 않고 논문에 기술된 아키텍처 구조를 직관적으로 분석할 수 있는 환경을 제공하며 복잡한 모델에 대한 이해를 지원하고 연구의 접근성을 제고하는 데 기여한다.

Convolutional Neural Networks (CNNs) consist of hierarchically stacked two-dimensional filters, which makes their internal structures and data flows inherently difficult to visualize intuitively. Existing visualization tools employ diverse representation styles and abstraction levels, limiting consistent comparison and interpretation across different models and increasing cognitive load for researchers. To address this issue, we propose an interactive web-based 3D visualization framework that extracts architectural information from CNN descriptions in academic papers and converts it into a structured representation. The proposed system utilizes a Large Language Model (LLM) to transform unstructured paper text into a JSON-based intermediate format, enabling users to explore and compare CNN architectures through interactions such as zooming and rotation within the framework. This approach supports intuitive analysis of architectures described in papers without relying on model implementation files and contributes to improving the accessibility of deep learning research.

14

A Robust Approach for Overlay Text Localization and Extraction in Complex Video Scene SCOPUS

Jingfan Tang, Zhitao Li, Xingqi Wang, Ming Jiang, Ziyang Li

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.11 No.10 2016.10 pp.407-420

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Overlay text in video carries important semantic clues for video information retrieval and summarization. In this paper, we propose a robust method that is able to accurately locate text lines and extract text even in complex video scene. In the text localization stage, this paper adopts the method based on corner point. First, corner detection is used to extract corners as text features from video frames. Then multi-layer filtering mechanism (MLFM) is used to locate the text lines, which consists of corners clustering, corners horizontal projection, background filtering and heuristic rules. This MLFM can effectively remove the isolated corners, locate the text lines accurately and remove the background or pseudo text lines automatically. In the text extraction stage, this paper proposed a twice binarization method that combines with polarity judgment on image. The polarity judgment was used as a guide to adjust the first binarization threshold when we perform the first binarization. After the first binarization, a main proportion of the image has been processed, and the rest will be processed by the second binarization. Experimental results show that this approach can fast and robustly locate text lines and extract text in video even under complex background.

15

Image Retrieval System for Composite Images using Directional Chain Codes

Akriti Nigam, Rupesh Yadav, R. C. Tripathi

보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology Vol.58 2013.09 pp.51-64

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Most of the image retrieval systems published and implemented have focused on basic features like color, shape and texture of an image with little or no consideration of the included text region of the image. In this work, we have developed an image retrieval system that works on finding similar composite images containing graphical shapes as well as text from a database of thousands of images. By proposing a novel method for text localization, extraction followed by detection, we have demonstrated how this method outperforms commercial OCR tools. The significant feature of this work is its handling the requirements of invariance to font size, design, text region orientation and its ability to give accurate result even in the presence of complex background and graphical elements. The methodology has been tested for English text but is capable to handle any other language.

16

Ancient Cuneiform Text Extraction Based on Automatic Wavelet Selection SCOPUS

Raed Majeed, Zou Beiji, Hiyam Hatem, Jumana Waleed

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.10 No.6 2015.06 pp.253-264

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Ancient Iraq was the home of a major urban civilization which developed during 4000-3000 BCE. The Sumerians, who lived in Mesopotamia in southern Iraq, invented the cuneiform system of writing, which was an essential element of Sumerians culture. The translation of cuneiform is a highly complicated process. It is only in comparatively recent years that the grammar has been scientifically established, while the lexical problems are still numerous and far from resolved. Furthermore, most of the Sumerians tablets lost only few old images left, some of it saved in a special collection or worldwide museums. In this paper, we present a novel method used to obtain the cuneiform text from old Sumerian clay tablets, proposed method based on automatically select wavelet bases which it is essential and critical issues for wavelet algorithm implementation. Our procedure offers the archaeological and Cuneiformest an easy, fast and active method for extracting the cuneiform sentences. Experimental results of sample images show that the proposed system has superior result.

17

This study uses a large language model (LLM) to identify Aristotle's rhetorical principles (ethos, pathos, and logos) in COVID-19 information on Naver Knowledge-iN, South Korea's leading question-and-answer community. The research analyzed the differences of these rhetorical elements in the most upvoted answers with random answers. A total of 193 answer pairs were randomly selected, with 135 pairs for training and 58 for testing. These answers were then coded in line with the rhetorical principles to refine GPT 3.5-based models. The models achieved F1 scores of .88 (ethos), .81 (pathos), and .69 (logos). Subsequent analysis of 128 new answer pairs revealed that logos, particularly factual information and logical reasoning, was more frequently used in the most upvoted answers than the random answers, whereas there were no differences in ethos and pathos between the answer groups. The results suggest that health information consumers value information including logos while ethos and pathos were not associated with consumers’ preference for health information. By utilizing an LLM for the analysis of persuasive content, which has been typically conducted manually with much labor and time, this study not only demonstrates the feasibility of using an LLM for latent content but also contributes to expanding the horizon in the field of AI text extraction.

18

The traditional information extraction methods based on specific domain usually depend on the domain dictionaries to discover the text feature. It is inconvenient for reproducing and difficult to transplant in multi-domain environment. The application scope is limited seriously. Oriented to the deficiencies above, a multi-domain web text feature extraction model for e-Science is proposed (named e-FTM). This model adopts the Chinese split words technology without dictionary into the process of multi-domain text feature discovery and avoids the dependency of domain dictionaries effectively. With the help of classification of common and individual features, the model tracks the generation and the development trend of domain events dynamically, and forms a couple of local data centers eventually. Through cooperative scheduling the domain knowledge between different local data centers, the knowledge utilization efficiency of the domain information in the global scope is improved sharply. To validate the performance, the experiments on the multi-domain text feature extraction, topic features dynamical tracking and the domain knowledge cooperative scheduling demonstrate that the model has higher application validity and practicality in e-Science environment.

19

Text Mining : Extraction of Interesting Association Rule with Frequent Itemsets Mining for Korean Language from Unstructured Data SCOPUS

Irfan Ajmal Khan, Junghyun Woo, Ji-Hoon Seo, Jin-Tak Choi

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.10 No.11 2015.11 pp.11-20

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Text mining is a specific method to extract knowledge from structured and unstructured data. This extracted knowledge from text mining process can be used for further usage and discovery. This paper presents the method for extraction information from unstructured text data and the importance of Association Rules Mining, specifically for of Korean language (text) and also, NLP (Natural Language Processing) tools are explained. Association Rules Mining (ARM) can also be used for mining association between itemsets from unstructured data with some modifications. Which can then, help for generating statistical thesaurus, to mine grammatical rules and to search large data efficiently. Although various association rules mining techniques have successfully used for market basket analysis but very few has applied on Korean text. A proposed Korean language mining method calculates and extracts meaningful patterns (association rules) between words and presents the hidden knowledge. First it cleans and integrates data, select relevant data then transform into transactional database. Then data mining techniques are used on data source to extract hidden patterns. These patterns are evaluated by specific rules until we get the valid and satisfactory result. We have tested on Korean news corpus and results have shown that it has worked well, and the results were adequate enough to research further.

20

Simultaneous Entities and Relationship Extraction from Unstructured Text SCOPUS

Jingtai Zhang, Jin Liu, Xiaofeng Wang

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.6 2016.06 pp.151-160

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Entity recognition and entity relationship extraction are two very important tasks in information extraction. Most research work in the literature treats these two work independently when processing the text. This paper proposes a novel method for performing entity recognition and entity relationship extraction simultaneously from unstructured text based on Conditional Random Fields (CRFs). This method makes use of entity features, entity relationship features and features of the triples which is composed of entities and their relationship to conduct the model training. Experiment results show that this method can recognize entity and extract entity relationship effectively.

 
1 2 3 4 5
페이지 저장