년 - 년
Automatic Extraction of Semi-structured Web Data
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.6 No.4 2013.08 pp.131-144
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
As a huge data source the internet contains a large number of valuable information, and the data of information is usually in the form of semi-structured in HTML web pages. In order to extract the web data and organize the data with the relationships which are similar to the real world, this paper has proposed a method for automatic data extraction from the web. With the combination of keywords and database content matching, the target web pages which contain valuable data will be crawled. Via HTML structure and visual features, extracting the data from the web pages crawled. Eventually, the data been extracted will be integrated to the structure of information network model. Experimental results indicate that this method can be able to apply to semi-structured data extraction in the web, and this paper has provided positive significance to extraction and manage semi- structured web data.
Efficient Storage Construction for Semi-Structured Microarray Data Exploiting Structural Similarity SCOPUS
보안공학연구지원센터(IJBSBT) International Journal of Bio-Science and Bio-Technology Vol.5 No.1 2013.02 pp.13-26
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
To promote molecular biology studies, public repositories for microarray data need to be constructed; the minimum contents for analysis of microarray experiment have been defined and standardized. Public repositories have been constructed by some researches which follow the standards such as MIAME-compliant data and MAGE-OM/ML. However, enough consideration has not been taken into the design of storage structure for the hierarchy of microarray data. In this paper, we propose alternative mapping strategy to mine the structural similarity and an advanced mapping rule from the algorithm. Object-relational mapping technique is used for extracting advanced storage design schema for microarray data and structural similarity of elements is evaluated for efficient storage construction. The mapping strategy reduced the number of relational tables remarkably. The strategy will contribute to design of the storage structure of microarray data and performance enhancement of a public repository.
보안공학연구지원센터(IJSEIA) International Journal of Software Engineering and Its Applications Vol.6 No.2 2012.04 pp.179-184
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Public repositories for microarray data have been constructed by some researches which follow the standards such as MIAME-compliant data and MAGE-OM/ML. However, enough consideration has not been taken into the design of storage structure for the hierarchy of microarray data. In this paper, we propose alternative mapping strategy to mine the structural similarity and an advanced mapping rule from the algorithm. Object-relational mapping technique is used for extracting advanced storage design schema for microarray data and structural similarity of elements is evaluated for efficient storage construction. The mapping strategy reduced the number of relational tables remarkably.
마이크로 데이터에서 익명화된 정형, 반정형, 비정형 정보의 유용성 측도 비교 분석
[NRF 연계] 사단법인 미래융합기술연구학회 아시아태평양융합연구교류논문지 Vol.10 No.10 2024.10 pp.465-476
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 익명화된 정형, 반정형, 비정형 데이터의 유용성 측도에 대해 고찰하며, 특히 익명화가 데이터 품질에 미치는 영향을 중점적으로 다루고자 한다. 익명화는 개인정보보호법 제3조 원칙과 제58조의2에 따라 개인정보를 보호하는 중요한 방법이지만, 종종 데이터 안전성과 유용성 사이에 상충 관계를 초래한다. 본 연구는 데이터 유용성 즉, 정보 손실 측면에서 익명화가 정형 데이터(예: 수치 및 텍스트 값), 반정형 트랜잭션 데이터, 그리고 비정형 이미지 데이터에 미치는 영향을 조사하고자 하는 것이 목적이다. 정형 데이터의 경우, 공분산, 상관계수, 평균 제곱오차 및 절대 오차 등 기존의 유용성 측도와 k-익명성과 같은 프라이버시 보호 모델에서의 유용성 측도에 대해 논의한다. 반정형 트랜잭션 데이터의 경우, km-익명성 모델과 여기에서 사용하는 NCP(Normalized Certainty Penalty)를 주요 측도로 다루고자 한다. 비정형 이미지 데이터의 경우 현재 프레셰 인셉션 거리(FID), 학습된 지각 이미지 패치 유사도(LPIPS), 구조적 유사성 지수(SSIM) 등의 측도들을 통하여 익명화된 이미지와 원본 이미지 간의 유사성을 유용성 측도로 평가하고 있다. 본 논문에서는 이러한 측도들을 상세히 고찰함으로써, 익명화와 데이터 유용성 사이의 균형 유지가 여전히 향후 연구들에 있어 중요한 도전 과제임을 보여주고자 하며, 아울러 익명화된 데이터를 활용하려는 기관 및 기업이 보다 높은 품질의 데이터를 생성하여 의미 있는 분석을 수행할 수 있도록 기여하기를 기대한다. 향후 연구로는 음성 및 텍스트와 같은 비정형 데이터에서의 유용성 측도를 확장해 고찰할 예정이다.
This paper examines the utility measures of anonymized structured, semi-structured, and unstructured data, with a particular focus on the impact of anonymization on data quality. While anonymization is a key method for protecting personal information under the principles of Article 3 and Article 58-2 of the Personal Information Protection Act, it often creates a trade-off between data security and utility. The purpose of this study is to investigate the effect of anonymization on data utility, or information loss, in structured data (e.g., numerical and textual values), semi-structured transactional data, and unstructured image data. For structured data, this paper discusses traditional utility measures such as covariance, correlation, mean squared error, and absolute error, as well as utility measures within privacy protection models like k-anonymity. For semi-structured transactional data, it focuses on the km-anonymity model and the key utility measure used in this model, the Normalized Certainty Penalty (NCP). For unstructured image data, the utility is evaluated using contemporary measures such as Frechet Inception Distance (FID), Learned Perceptual Image Patch Similarity (LPIPS), and Structural Similarity Index Measure (SSIM) to assess the similarity between anonymized and original images. By thoroughly reviewing these measures, this paper aims to demonstrate that balancing anonymization and data utility remains a significant challenge for future research. Moreover, it seeks to contribute to organizations and businesses aiming to generate higher quality anonymized data for meaningful analysis. Future research will extend the discussion on utility measures to other forms of unstructured data, such as audio and text.
분산된 준구조적 데이터 검색을 위한 경로 질의 처리 기법
[Kisti 연계] 한국정보과학회 정보과학회논문지:데이타베이스 Vol.28 No.1 2001 pp.95-103
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문에서는 분산된 준구조적 데이터에 대한 질의 처리 문제를 다룬다. 분산된 준구조적 데이터는 루트가 있고 간선에 레이블이 있는 그래프 모델로 표현될 수 있으며, 그래프의 조드들은 한 사이트 또는 여러 사이트들에 위치할 수 있다. 분산된 준구조적 데이터의 효율적인 검색을 위해 ‘질의 단축 및 확산’ 방법에 기반을 둔 질의 처리 모델을 제안한다. 이 방법은 사용자 질의가 사이트 내부에서 단축되고 다른 사이트로 분산되는 과정을 통해 데이터를 검색한다. 또한, 제안된 모델에 필요한 알고리즘들을 제시하고 정확성을 증명한다.
집합값을 갖는 반정형 트랜잭션 데이터 익명화에서 정보손실 측도 개선에 관한 연구
[NRF 연계] 사단법인 미래융합기술연구학회 아시아태평양융합연구교류논문지 Vol.10 No.9 2024.09 pp.45-54
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
정형데이터란 일반적으로 테이블에서 하나의 셀 내에 하나의 문자나 혹은 숫자로 구성된 데이터를 말한다. 만일 여기서 하나의 셀 내에 하나의 문자나 혹은 숫자가 아니라 여러 문자나 혹은 숫자들의 집합 값으로 구성된 경우, 우리는 이것을 반정형 데이터라 부른다. 예로 우리가 마트에서 구매한 일련의 물품 아이템 들이나 혹은 병원에서 환자에 대한 여러 병명들에 대한 목록이 그러하며, 이때 한 명에 대한 집합 값들로 구성된 목록을 하나의 트랜잭션이라 부른다. 우리는 이러한 일련의 집합값들로 구성된 반정형 트랜잭션 데이터셋에서 개인정보보호를 위한 익명화 문제를 다룬다. 즉, 개인이 구매한 상품 목록이나 혹은 환자의 병명에 대한 사항들은 민감정보로서 개인의 프라이버시 차원에서 보호되어야 할 개인정보들이기 때문이다. 이러한 익명화 문제와 관련하여 우리는 기존의 LG (Local Generalization) 알고리즘을 개선하여 LGR (Local Generalization & Reallocation)이라는 새로운 알고리즘을 제안한바 있다. 그러나 만일 우리가 익명화된 개인정보들을 활용하거나 분석하는 관점에서 바라보면 안전하기도 해야하지만 이에 못지 않게 데이터가 쓸모가 있어야 한다. 즉 데이터 품질이 분석에 용이하도록 유용해야한다. 이는 반대로 익명화 과정에서 정보 손실이 최소화 되어야한다는 것과 동일한 의미를 갖는다. 기존 우리의 LGR 알고리즘은 정보 손실을 계산하기 위한 측도로 기존 LG 알고리즘과 동일한 NCP (Normalized Certainty Penalty)를 사용하였다. NCP 측도는 전체 아이템 수 대비 일반화된 아이템의 비율로 계산된다. 따라서 계산이 단순하고, 다양한 데이터셋에 쉽게 적용할 수 있는 장점이 있지만, 그 반대로 정보 손실이 많아 데이터의 유용성을 떨어뜨릴 수 있는 단점이 있다. 우리는 이러한 단점을 개선하고자 새로운 IGH (Information Gain-based Heuristic) 측도를 새로이 제안하고 이를 이론적으로 검증해 보고자 한다. 제안하는 측도는 기존 NCP 방식에 비하여 정보 손실을 최소화하고 데이터의 유용성을 최대한 보존할 수 있는 장점이 있다.
Structured data typically refers to data composed of a single character or number within a single cell of a table. If, instead, a single cell contains a set of multiple characters or numbers, we refer to this as semi-structured data. For example, the list of items purchased at a supermarket or diagnoses for a patient in a hospital represents such cases. In these cases, the list of values for an individual is called a transaction. We address the anonymization problem for privacy protection in semi-structured transaction datasets composed of these values. Specifically, lists of purchased items or diagnoses contain sensitive information that must be protected for individual privacy. Regarding this anonymization problem, we previously proposed an improvement to the existing Local Generalization (LG) algorithm, resulting in the new Local Generalization & Reallocation (LGR) algorithm. However, the data must be secure and valuable from the perspective of utilizing or analyzing anonymized personal data. This implies that data quality must be preserved to facilitate analysis, which means that information loss during anonymization must be minimized. Our existing LGR algorithm used the Normalized Certainty Penalty (NCP) as a measure for calculating information loss, the same as the existing LG algorithm. The NCP measure calculates the ratio of generalized items to the total number of items. While this measure is simple to compute and easily applicable to various datasets, it has the drawback of potentially high information loss, which can reduce data utility. To address this drawback, we propose a new Information Gain-based Heuristic (IGH) measure and aim to verify its effectiveness theoretically. The proposed measure has the advantage of minimizing information loss and maximizing data utility compared to the existing NCP method.
반정형 트랜잭션 데이터를 위한 새로운 익명화 알고리즘 설계 및 구현에 관한 연구
[NRF 연계] 사단법인 미래융합기술연구학회 아시아태평양융합연구교류논문지 Vol.9 No.11 2023.11 pp.13-22
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
트랜잭션 데이터는 다양한 고객이 슈퍼마켓에서 구매한 품목에 대한 정보를 저장하는 관계형 데이터베이스를 고려한 것으로 하나의 셀 내에 여러 개의 아이템들이 일종의 집합으로 구성된 것을 말한다. 본 본문은 이러한 트랜잭션 데이터에서의 익명화에 관한 문제를 다루고자 한다. Y. He와 J. Naughton은 이 문제를 해결하기 위해 일명 ‘로컬 일반화’라고 부르는 k-익명성 기반 기법을 제안한 바 있다. 그러나 이 기법은 강력한 프라이버시를 제공하고 실행시간이 짧다는 장점이 있지만 정보손실이 크다는 단점이 있다. 한편 J. Liu와 K. Wang은 이러한 단점을 해결하기 위하여 HgHs(Heuristic generalization with Heuristic suppression) 알고리즘을 제안한 바 있다. 그러나 이 기법은 그 반대로 정보손실은 적지만 실행시간이 오래 걸린다는 단점이 있다. 본 저자는 기존에 이러한 단점들을 개선하기 위하여 상기 두가지 기법들에 비해 정보 손실은 최소화 하면서도 강력한 프라이버시를 보장할 수 있는 새로운 기법을 제안한 바 있다. 본 논문에서는 이를 실제 상품화할 수 있도록 시스템을 새로이 구성하고 기존 제안 알고리즘의 설계 및 구현과정과 테스트 결과를 제시하고자 한다. 테스트는 기존 로컬일반화와 HgHs 기법에서 사용한 데이터셋(BMS-WebView2와 BMS-POS)을 동일하게 사용하였으며, 테스트 결과 제안 알고리즘의 정확도를 확인할 수 있었다.
The transaction data refers to information stored in a relational database, considering various customers' purchases at a supermarket, where multiple items are grouped together as a kind of set within a single cell. This paper aims to address the issue of anonymization in such transaction data. Y. He and J. Naughton have previously proposed a technique known as 'Local Generalization,' which is based on k-anonymity to tackle this problem. However, this technique offers strong privacy protection and short execution times but suffers from significant information loss. On the other hand, J. Liu and K. Wang have proposed the HgHs (Heuristic generalization with Heuristic suppression) algorithm to mitigate these drawbacks, with reduced information loss but longer execution times. In order to improve these existing shortcomings, we have proposed a new technique that can guarantee strong privacy while minimizing information loss compared to the above two techniques. In this paper, we will reconfigure the system so that it can be commercialized and present the design, implementation process, and test results of the existing proposed algorithm. The test used the datasets (BMS-WebView2 and BMS-POS) used in the existing local generalization and HgHs techniques, and the test results confirmed the accuracy of the proposed algorithm.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.