년 - 년
A Novel Crawler Based on Loginning Simulation for Weibo Social Network SCOPUS
보안공학연구지원센터(IJGDC) International Journal of Grid and Distributed Computing Vol.9 No.1 2016.01 pp.135-144
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
With the rapid development of Weibo, which is the most popular microblog in china, more and more attention was paid to relative studies about it. With the objective of gathering precise information data from Weibo, which is the groundwork of these researches, a novel high efficient Weibo crawler (WCrawler) based on loginning simulation is designed. The priority evaluation is described to ensure the correlation between entires. MD5 is introduced to check for duplicates of URL crawled. Experiments demonstrate that the novel crawler has an efficiency and integrity of information collecting compared with API crawler. In addition, we present a summary of the data that collected from Weibo social network by WCrawler.
국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.16 No.4 2024.12 pp.203-208
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The exponential increase in publications and the interconnected nature of sub-domains make traditional methods of information extraction and organization inadequate. This inefficiency can impede scientific progress and innovation. To address these challenges, this research leverages the ability of Bidirectional Encoder Representations from Transformers for keyword extraction (KeyBERT) and integrates with K-Means clustering to organize topics from large datasets effectively. Analyzing a dataset of 47,627 articles from SCOPUS in the domains of Reinforcement Learning and Computer Vision. An ablation study demonstrates the generalizability of the approach across these fields, with the optimal number of clusters determined to be three using the Elbow Method. The results demonstrate that KeyBERT is effective in extracting and organizing topics within these domains, with a particular focus on applications such as medical imaging, autonomous driving, and real-time detection systems. This methodology offers a scalable solution for organizing vast academic datasets, enabling researchers to extract meaningful insights efficiently and apply this approach to other domains.
Research on Model of Network Information Currency Evaluation Based on Web Semantic Extraction Method
보안공학연구지원센터(IJFGCN) International Journal of Future Generation Communication and Networking Vol.7 No.2 2014.04 pp.103-116
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
With the arrival of the big data era, it has become a spot for research to evaluate currency of network information so far. This paper proposes a model of network information currency evaluation based on Web semantic extraction method taking Web news as object of study. The author elaborates the method, technology and main functions on every layer of the model in detail, which have been used or completed, and focus on how to extract semantic information efficiently from the contents of Web news, in order to explore a research method for network information currency evaluation. The experimental results show the validity of the model design that plays a very important role in leading network users pay attention to more valuable network information and helping Web site managers build a higher currency Web site.
웹 문서 정보추출과 자연어처리를 통한 온톨로지 자동구축에 관한 연구 KCI 등재후보
국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제9권 제3호 2009.06 pp.61-67
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
인터넷의 발달로 전자문서가 증가함에 따라, 정보검색기술의 중요성도 함께 증가하게 되었다. 본 연구는 비정형 텍스트 웹 문서로부터 사용자가 요구하는 핵심 의미 지식을 추출하기 위하여 LGG (Local Grammar Graph) 구축에 기반 하여 보다 효율적이고 정확한 지식구축을 가능하게 한다. 주가등락이라는 특정 분야의 패턴을 추출하여 만든 패턴 문법을 사용해서 OWL(Web Ontology Language) 기반의 온톨로지를 구축하였다. 특정 분야의 온톨로지를 구축함으로써 기존 검색에서 할 수 없었던 지식의 의미 검색이 가능하며 나아가 사용자가 원하는 질의에 대한 정보의 추론이 가능할 것 이다.
The proliferation of the Internet grows, according to electronic documents, along with increasing importance of technology in information retrieval. This research is possible to build a more efficient and accurate knowledge-base with unstructured text documents from the Web using to extract knowledge of the core meaning of LGG (Local Grammar Graph). We have built a ontology based on OWL(Web Ontology Language) using the areas of particular stocks up/down patterns created by the extraction and grammar patterns. It is possible for the user can search for meaning and quality of information about the user wants.
[Kisti 연계] 한국전자거래학회 한국전자거래학회 학술대회논문집 2005 pp.55-61
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
We will introduce a new web-based knowledge management system in progress, in which XML-based web information extraction and our structuring knowledge technologies are combined using ontology-based natural language processing. Our aim is to provide efficient access to heterogeneous information on the web, enabling users to use a wide range of textual and non textual resources, such as newspapers and databases, effortlessly to accelerate knowledge acquisition from such knowledge sources. In order to achieve the efficient knowledge management, we propose at first an XML-based Web information extraction which contains a sophisticated control language to extract data from Web pages. With using standard XML Technologies in the system, our approach can make extracting information easy because of a) detaching rules from processing, b) restricting target for processing, c) Interactive operations for developing extracting rules. Then we propose a structuring knowledge system which includes, 1) automatic term recognition, 2) domain oriented automatic term clustering, 3) similarity-based document retrieval, 4) real-time document clustering, and 5) visualization. The system supports integrating different types of databases (textual and non textual) and retrieving different types of information simultaneously. Through further explanation to the specification and the implementation technique of the system, we will demonstrate how the system can accelerate knowledge acquisition on the Web even for novice users of the field.
특정 영역 정보 에이전트의 지식베이스 확장을 위한 웹 정보추출
[Kisti 연계] 한국지능정보시스템학회 한국지능정보시스템학회 학술대회논문집 2002 pp.336-341
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
현재 연구개발 중인 웹 정보 에이전트는 Agent Manager와 KB Manager. Web Manager로 구성되어 있다. 이 시스템은 동물영역에 관련된 정보를 영어로 서비스하고 있어 국내 접근보다는 외국에서의 접근이 더 많았다. 그러므로 국내 사용을 높이기 위해 애완용 동물을 위주로 한 정보추출(IE)을 수행하여 지식베이스(KB)의 확장을 시도하고 있다. 이를 위하여 태그(tag) 및 심볼(symbol)의 패턴(pattern) 유사성 정보를 찾아내고, 기존 KB와 연계하여 KB의 확장 및 수정에 이용하기 위한 유효 정보 패턴 결정에 활용함으로써 정보 추출의 새로운 방법을 고찰하고 그 가능성을 제시하고자 한다.
[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.11 No.5 2008 pp.567-578
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
인터넷에 있는 방대한 양의 웹 페이지들을 분석하기 위해서는 웹 페이지에 내재된 정보를 추출하는 것이 필요하다. 본 논문에서는 웹 페이지로부터 정보를 추출하고 이를 XML 문서로 변환하여 다차원적으로 분석하는 방법을 제안한다. 웹 페이지로부터 정보를 추출하기 위하여 두 종류의 언어를 제안한다. 하나는 객체지향 모델에 의거하여 웹 정보 추출 규칙을 기술하기 위한 것이고, 다른 하나는 추출하고자 하는 정보를 찾기 위한 HTML 태그 패턴을 정규식으로 기술하기 위한 것이다. XML 문서에 대한 다차원 분석을 위하여 관계형 데이터에 대해 하는 것처럼 웨어하우스를 구축하고 이로부터 다양한 큐브를 생성하는 방법을 제안한다. 마지막으로 본 논문에서 제안한 방법을 미국특허 웹 페이지에 적용한 예를 통해 그 타당성을 보인다.
For analyzing a huge amount of web pages available in the Internet, we need to extract the encoded information in web pages. In this paper, we propose a method to extract and convert web information from web pages into XML documents for multidimensional analysis. For extracting information from web pages, we propose two languages: one for describing web information extraction rules based on the object-oriented model, and another for describing regular expressions of HTML tag patterns to search for target information. For multidimensional analysis on XML documents, we propose a method for constructing an XML warehouse and various XML cubes from it like the way we do for relational data. Finally, we show the validness of our method through the application to US patent web pages.
[Kisti 연계] 한국정보시스템학회 한국정보시스템학회 학술대회논문집 2005 pp.79-92
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
To query the vast amount of web pages which are available i]l the Internet, it is necessary to extract the encoded information in the web pages for converting it into structured data (e.g. relational data for SQL) or semistructured data (e.g. XML data for XQuery), In this paper, we propose a new web information extraction system, PIES, to convert web information into XML documents. PIES is based on a user-specified target schema and HTML tag pattern descriptions. The web information is extracted by the pattern descriptions and validated by the target schema. We designed a new language to describe extraction rules, and a new regular expression to describe HTML tag patterns. We implemented PIES and applied it to the US patent web site to evaluate its correctness. It successfully extracted more than thousands of US patent data and converted them into XML documents.
웹 정보추출의 성능향상을 위한 사용자 관심 부분 추출기의 구현
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2005 pp.673-675
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
인터넷이 발전할수록 정보의 양이 늘어나게 되어 방대한 양의 데이터 속에서 적합한 정보를 추출하는 방법이 필요하다. 그리고 같은 데이터라 하더라도 유용한 정보라고 판단하는 것은 개인의 관심도에 따라 다르다. 따라서 우리는 사용자 관심 정보 추출이라는 목표 아래에서 개인간의 차이에도 명확히 정보를 추출할 수 있는 방법의 필요성을 인지하여 정보추출의 사전 단계에서 사용자가 원하는 정보가 있는 블록을 식별하는 방법에 대해서 연구하였다. 사용자가 선호하는 정보가 들어있는 블록들에 대해서만 정보 추출 기법을 적용하면 정확성과 속도면에서 좋은 결과를 얻을 수 있을 것으로 예상된다. 또한 XML-QL[7]형식의 질의를 통해 사용자의 요구 변화에 유연하게 대처하는 방법을 제안한다.
[Kisti 연계] 한국산업정보학회 한국산업정보학회논문지 Vol.6 No.2 2001 pp.7-15
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문은 웹사이트에서 문서를 삽입하거나 삭제할 경우, <A HREF>태그를 생성하는 수동적인 방법을 사용하지 않고, 자동적인 방법으로 문서를 삽입·삭제하는 알고리즘을 제안한다. 자동적인 방법으로 문서를 삽입·삭제하기 위하여 문서의 HTML 태그 중 문서와 문서를 연결하여 주는 <A HREF>태그를 사용한다. 그래프 구조의 상하계층은 한 문서의 <A HREF>...</A>태그사이의 텍스트들을 추출하여 추출된 텍스트를 이 문서의 하위노드의 노드명으로 링크를 생성하는 방법을 사용한다. 웹문서들간의 링크를 새로이 설정하고자 하는 경우 <A HREF>태그를 생성하지 않고 구조화된 그래프 형태를 이용하여, 그래프에서 노드를 삽입하거나 삭제한 후 새로운 구조를 웹에 적용한다. 문서를 삭제할 때에는 삭제될 노드와 링크되어 있는 노드들에 대하여 삭제되는 노드의 부모노드와의 링크를 새로이 선정해 줌으로써 단절 링크를 없애준다.
In this paper, we suggest the algorithm that inserts or deletes documents into web sites without creating <A HREF> tag. This algorithm uses <A HREF> tag which links between documents to automatically inset or delete the web documents. This study extracts the texts in the <A HREF>...</A> tag of the document to put into the structure as a type of graph creating the link as a node m of sub-node. That is, in this case of configurating new link between web documents, this algorithm allows to insert or delete the node to or from the graph without creating the <A HREF> tag. In the case of deleting the document, it removes the broken link connecting the sub-nodes of deleted node newly to its parent node.
분산된 웹 정보의 효과적 통합$\cdot$추출을 위한 동적 Wrapper 조합
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2005 pp.676-678
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
웹 정보 통합은 사용자 질의에 적합한 정보를 분산된 웹에서 추출하여 제공하는 방법으로 질의응답 속도의 향상을 위해 질의처리 방식을 주로 사용한다. 질의 처리는 Wrapper를 이용해 웹으로부터 제약조건을 만족하는 정보를 추출하고 사용자가 원하는 형태로 결합하는 방식인데, 통합과정에서 제거될 정보까지 미리 추출하는 문제가 있다. 본 논문에서는 이를 해결하기 위해 튜플 단위 웹 정보 추출 방법을 제안한다. 제안하는 방법은 F-Logic으로 표현된 도메인 모델과 CHR(Constraint Handling Rule)로 정의한 규칙을 이 용해 질의를 확장하고 적절한 Wrapper들을 선택한 뒤 추출에 필요한 Wrapper를 동적으로 조합한다. 쇼핑몰 사이트에 분산된 웹 정보 획득에 제안하는 방법을 적용하여 유용성을 확인하였다.
텍스트 정보와 시각 특징 정보를 이용한 효과적인 웹 이미지 캡션 추출 방법
[Kisti 연계] 한국정보과학회 한국정보과학회 학술대회논문집 2006 pp.346-348
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
기존의 웹 이미지 검색 시스템들은 웹 페이지에 포함된 텍스트들의 출현빈도, 태그유형 등을 고려해 각 키워드들의 중요도를 평가하고 이를 이용해 이미지의 캡션을 결정한다. 하지만 텍스트 정보만으로 캡션을 결정할 경우, 키워드와 이미지 사이의 관련성을 평가할 수 없어 부적절한 캡션의 배제가 어렵고, 사람의 인지와 맞지 않는 캡션이 추출되는 문제점이 있다. 본 논문에서는 기존의 웹 이미지 마이닝 방법을 통해 웹 페이지로부터 캡션 후보 키워드를 추출하고, 자동 이미지 주석 방법을 통해 이미지의 개념 부류 키워드를 결정한 후, 두 종류의 키워드를 결할하여 캡션을 선택한다. 가능한 결합 방법으로는 키워드 병합 방법, 공통 키워드 추출 방법, 개념 부류 필터링 방범 캡션 후보 필터링 방법 등이 있다. 실험에 의하면 키워드 병합 방법은 높은 재현율을 가져 이미지에 대한 다양한 주석이 가능하고 공통 키워드 추출 방법과 개넘 부류 키워드 필터링 방법은 정확률이 높아 이미지에 대한 정확한 기술이 가능하다. 특히, 캡션 후보 키워드 필터링 방법은 기존의 방법에 비해 우수한 재현율과 정확률을 가지므로 기존의 방법에 비해 적은 개수의 캡션으로도 이미지를 정확하게 기술할 수 있으며 일반적인 웹 이미지 검색 시스템에 적용할 경우 효과적인 방법이다.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.