년 - 년
Personal Health Data Anonymization in China : The Legal Framework and Its Refinements KCI 등재
원광대학교 법학연구소 의생명과학과 법 제31권 2024.06 pp.135-174
최근 중국정부는 개인에 대한 의료 및 건강 데이터의 익명화 방안에 대하여 다양한 방안을 내놓고 있다. 그럼에도 불구하고, 다양한 법률문제가 발생함에 따라 개인의료데이터의 수집부터 새로운 논의가 필요하다는 목소리가 높다. 중 국은 개인의료데이터의 ‘수집-저장-이용-이동’ 등의 과정에 있어 데이터의 익 명화를 위한 표준도 아직 부족한 상태이다. 더구나 데이터 익명화를 위한 ‘식별 불가’ 및 ‘가역불가’의 기준이 너무 높아 의료데이터를 활용하는 기술분야에서는 어려움을 겪고 있는 것이 사실이다. 즉, 의료데이터의 효과적 사용과 개인정보 보호가 서로 충돌하고 있어 입법을 통한 적절한 규제가 필요하다. 유럽의 경우 는 의료데이터의 ‘위험관리’ 접근방식을 채택하고, 미국은 ‘비식별화’ 방식을 운 영하고 있는데, 이는 현재 중국이 안고 있는 의료데이터 익명화 표준에 있어 좋 은 시사점을 제공할 것이라 사료된다. 즉, ‘위험관리’ 접근 방식과 ‘비식별화’ 운 영 방식을 최적화할 경우 익명화를 통한 개인정보보호를 개선하면서 규제의 불 확실성을 줄일 수 있을 것이라 본다. 본 연구는 상기의 문제의식에 기초하여 기술 변화와 다양한 정보활용의 요구 를 충족시키기 위해 중국의 의료데이터 익명화 표준에 대한 구체적 의견을 제 시하였다. 이를 통해 개인정보를 보호함은 물론, 데이터 활용을 촉진할 수 있을 것이라 본다.
China has adopted a dual approach to managing the anonymization of personal health data. In the context of the increasing collection of personal health information, china's current legal standards and practices of data anonymization are inadequate for data protection and utilization. In particular, the requirements of “unidentifiable” and “irreversible” for anonymized data are too high to achieve, and cannot properly address the challenges posed by technological progress. There are also inherent conflicts between the standards of “unidentifiable” and “irreversible” with the guidelines of “de-identification”. These issues affect the effective use of health data and the protection of individual privacy in China. Through a comparative analysis of the EU's “risk management” approach and the US's “de-identification” method in data anonymization, this paper explores the possibility of optimizing the “risk management” approach and the “de-identification” operation to reduce the regulatory uncertainty while enhancing data anonymization protection. The paper proposes suggestions for revising and improving the requirements for personal health data anonymization in China. To better adapt to technological changes and the diverse needs for information use, this paper argues for the establishment of a dynamic legal framework that protects privacy and promotes data utilization, as well as draws on international practices to provide more flexible and adaptable legal responses.
本文深入探讨了中国在个人医疗健康数据匿名化管理方面的法律框架及其存在的 问题。文章指出,在中国个人健康信息收集日益增长的整体背景下,当前的数据匿 名化标准和实践存在不足。特别是,数据匿名化“不可识别”和“不可逆转”的标准过 高,无法妥善处理技术进步所带来的挑战,并与“去标识”的操作指南之间存在严重 冲突。这些问题也影响了中国健康数据的有效利用和个人隐私保护。通过比较法上 对欧洲健康数据“风险管理”进路和美国“去标识化”操作的分析,本文探索了优化“风 险管理”进路和“去标识化”操作的可能,以便减少在提高数据匿名化保护的同时,减 少规制的不确定。 本文最后提出对中国健康数据匿名化标准的修订与完善意见,以适应技术变革和 多样化的信息利用需求。这一建议包括建立一个动态的法律框架,同时保护隐私和 促进数据利用,并借鉴国际惯例,提供更灵活和适应性更强的法律对策。
Bigdata Anonymization Using One Dimensional and Multidimensional Map Reduce Framework on Cloud SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.6 2015.12 pp.253-262
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
Data privacy preservation is one of the most disturbed issues on the current industry. Data privacy issues need to be addressed urgently before data sets are shared on cloud. Data anonymization refers to as hiding complex data for owners of data records. In this paper investigate the problem of big data anonymization for privacy preservation from the perspectives of scalability and time factor etc. At present, the scale of data in many cloud applications increases tremendously in accordance with the big data trend. Here propose a scalable Two Phase Top-Down Specialization (TPTDS) approach to anonymize large-scale data sets using the MapReduce framework on cloud. For the data anonymization-45,222 records of adults information with 15 attribute values was taken as the input big data. With the help of multidimensional anonymization on map reducing framework, here implemented the proposed Two-Phase Top-Down Specialization anonymization algorithm on hadoop will increases the efficiency of the big data processing system. In both phases of the approach, deliberately design multidientional MapReduce jobs to concretely accomplish the specialization computation in a highly scalable way. Data sets are generalized in a top-down manner and the better result was shown in multidmientional MapReduce framework by compairing the onedimentional MapReduce framework anonymization job. The anonymization was performed with specialization operation on the taxonomy tree. The experiment demonstrates that the solutions can significantly improve the scalability and efficiency of big data privacy preservation compared to existing approaches. This work has great applications to both public and private sectors that share information to the society.
Balancing patient privacy and predictive accuracy through data anonymization in healthcare
[Kisti 연계] 테크노프레스 Advances in computational design Vol.11 No.1 2026 pp.47-62
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Data anonymization in healthcare is essential for protecting sensitive patient information while enabling secure usage for research, analytics, and AI-driven clinical decision-making. In this study, the MIMIC-III - Deep Reinforcement Learning dataset was used, which contains comprehensive electronic health records (EHRs) of ICU patients. Data preprocessing was performed using Min-Max Normalization to scale numerical features and ensure consistency. Anonymization techniques such as pseudonymization, generalization, suppression, data masking, and statistical methods like k-anonymity, l-diversity, and t-closeness were applied to safeguard patient privacy. The anonymized dataset was then utilized for predictive modelling using AI techniques including Random Forest and LSTM. Results demonstrated that privacy was maintained with 0% PII leakage, while predictive accuracy remained high, achieving accuracy of 94.6%, precision of 93.8%, recall of 92.5%, and F1-score of 93.1%. This study highlights that effective data anonymization ensures compliance with HIPAA and GDPR while retaining the utility of healthcare data for advanced analytics and AI applications.
실시간 스트림 데이터 익명화에서 정보 손실 측도의 개선에 관한 연구
[NRF 연계] 사단법인 미래융합기술연구학회 아시아태평양융합연구교류논문지 Vol.9 No.11 2023.11 pp.23-33
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
스트림 데이터는 각종 범죄 신고 정보, 온라인 판매 거래 정보, 병원 내 환자 모니터링 기기 정보 등 실시간으로 수집되는 정보를 말한다. 본 논문은 스트림 데이터의 프라이버시 문제인 익명화 문제를 다룬다. 일반적으로 익명화의 주요 고려 사항은 안전한 동시에 유용한 방식으로 데이터를 처리하는 것이다. 데이터 유용성은 곧 데이터의 품질을 말하며, 소위 정보 손실 척도로 측정된다. 본 논문에서는 실시간 스트림 데이터 익명화에서 그동안의 정보 손실 지표들을 검토하고 단점을 개선한 새로운 지표를 제안한다. 우리는 실시간 스트림 데이터 익명화에서 기존 Goldberger 등이 제안한 기법을 적용하여 단점을 개선하였다. Goldberger 등이 제안한 측도는 k-익명성 모델이 적용되어 일반화된 데이터셋에서 기존에 적용되지 않았던 데이터 테이블의 전체 동질집합과 동질집합 내 레코드 수를 함께 고려한 것이 특징이다. 기존에는 하나의 클러스터 내 하나의 동질집합만을 가정하였기 때문에 고려사항에 포함되지 않았다. 그 이유는 하나의 클러스터 내에서 k-익명성을 만족하지 못하는 해당 레코드들은 이동 가능한 즉, 차순위 유사 식별자로 클러스터링이 가능한 다른 클러스터에 할당하거나 할당 가능한 클러스터가 없는 경우 삭제되기 때문이다. 그러나 정보 손실 측면에서 해당 레코드들의 삭제는 곧 손실 증가로 이어지기 때문에 다른 클러스터에 배정하는 것이 타당하다. 따라서 이들 레코드들을 이동 가능한 타 클러스터로 배정하거나 배정할 클러스터가 없을 경우 삭제 대신 이들 레코드 모두를 별도의 독립된 클러스터에 할당하여 정보 손실을 최소화하는 것이 바람직하다. 제안 아이디어는 전자의 경우보다는 후자에 주목한다. 이 경우 독립된 클러스터 내에는 유사 식별자들 갖는 여러 동질집합들이 존재할 수 있기 때문이다. 따라서 제안하는 Goldberger 등의 측도를 이용함으로서 동질집합과 동질집합 내 레코드 수를 감안하여 정보 손실을 측정한 후 손실 값이 최소화 되도록 값들을 일반화할 필요가 있다.
Stream data refers to information collected in real time, such as crime report information, online sales transaction information, and information from patient monitoring devices in hospitals. This paper deals with the anonymization problem, which is the privacy issue of stream data. Usually, the main consideration in anonymization is how to process the data in such a way that it is secure and useful at the same time. Data usefulness refers to the quality of data and is measured by a so-called loss of information measure. In this paper, we review the current information loss metric in real-time stream data anonymization and propose a new metric that improves the disadvantages by applying Goldberger et al.'s scheme. The measure proposed by Goldberger et al. is characterized by considering the total equivalent class of the data table and the number of records in the equivalent class, which were not previously applied in the generalized dataset to which the k-anonymity model was applied. In the past, only one equivalent class within one cluster was assumed, so it was not included in the consideration. The reason is that records that do not satisfy k-anonymity within one cluster are allocated to other clusters that can be moved, that is, clustered with the next-rank quasi-identifier, or are deleted if there are no clusters that can be allocated. However, in terms of information loss, it is reasonable to assign them to other clusters because deleting those records leads to increased loss. Therefore, it is desirable to minimize information loss by allocating these records to another cluster that can be moved or, if there is no cluster to be assigned, allocating all of these records to a separate independent cluster instead of deleting them. The proposed idea focuses on the latter rather than the former. In this case, it is because several equivalent classes with quasi-identifiers can exist in an independent cluster. Therefore, by using the proposed metric of Goldberger et al., it is necessary to generalize to minimize the measured loss value after measuring the information loss by considering the equivalent class and the number of records in the equivalent class.
빅데이터 환경에서 개인정보보호를 위한 익명화된 데이터의 비익명화를 통한 데이터 안전성 테스트 방법론에 관한 연구
[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2013 pp.684-687
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
빅데이터 환경은 수많은 데이터의 조합으로 가치를 발견하여 이를 활용하는 것이다. 이러한 환경의 전제조건은 데이터의 공개 및 공유 개방이 될 것이다. 하지만 데이터 공개 시 개인정보와 같은 정보가 포함되어 법적 도덕적인 문제나 공개된 정보의 범죄 활용 등 2차적인 피해가 발생할 수 있어 데이터 공개 시 개인정보에 대한 익명화가 반드시 필요하다. 하지만 익명화된 데이터는 다른 정보와 결합을 통하여 재식별되어 비익명화 될 가능성이 항상 존재한다. 따라서 본 논문에서는 데이터 공개 시 익명화된 데이터를 공개하기 전에 재식별성에 대한 위험을 평가하는 테스트 방법론을 제안한다. 제안하는 방법론은 실제 테스트를 수행하는 3가지 과정 및 테스트 레벨 설정과 익명화 시 고려해야 할 부분으로 이루어져 있다. 제안하는 방법론을 통하여 안전한 데이터 공개 환경이 조성되어 빅데이터 시대에 개인정보에 안전한 데이터 공유와 개방이 이루어질 것으로 기대한다.
결정트리 기반의 기계학습을 이용한 동적 데이터에 대한 재익명화기법
[Kisti 연계] 한국정보과학회 정보과학회논문지 Vol.44 No.1 2017 pp.21-26
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
사물인터넷, 클라우드 컴퓨팅, 빅데이터 등 새로운 기술의 도입으로 처리하는 데이터의 종류와 양이 증가하면서, 개인의 민감한 정보가 유출되는 것에 대한 보안이슈가 더욱 중요시되고 있다. 민감정보를 보호하기 위한 방법으로 데이터에 포함된 개인정보를 공개 또는 배포하기 전에 일부를 삭제하거나 알아볼 수 없는 형태로 변환하는 익명화기법을 사용한다. 그러나 준식별자의 일반화 수준을 계층화하여 익명화를 수행하는 기존의 방법은 데이터 테이블의 레코드가 추가 또는 삭제되어 k-익명성을 만족하지 못하는 경우에 더 높은 일반화 수준을 필요로 한다. 이와 같은 과정으로 인한 정보의 손실이 불가피하며 이는 데이터의 유용성을 저해하는 요소이다. 따라서 본 논문에서는 결정트리 기반의 기계학습을 적용하여 기존의 익명화방법의 정보손실을 최소화하여 데이터의 유용성을 향상시키는 익명화기법을 제안한다
In recent years, new technologies such as Internet of Things, Cloud Computing and Big Data are being widely used. And the type and amount of data is dramatically increasing. This makes security an important issue. In terms of leakage of sensitive personal information. In order to protect confidential information, a method called anonymization is used to remove personal identification elements or to substitute the data to some symbols before distributing and sharing the data. However, the existing method performs anonymization by generalizing the level of quasi-identifier hierarchical. It requires a higher level of generalization in case where k-anonymity is not satisfied since records in data table are either added or removed. Loss of information is inevitable from the process, which is one of the factors hindering the utility of data. In this paper, we propose a novel anonymization technique using decision tree based machine learning to improve the utility of data by minimizing the loss of information.
데이터 스트림의 프라이버시 보호를 위한 지연 없는 익명화 기법
[Kisti 연계] 한국정보과학회 정보과학회논문지:데이타베이스 Vol.40 No.6 2013 pp.411-422
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
개인정보가 포함된 데이터 스트림은 프라이버시 보호를 위해 익명화 과정을 거쳐 활용되어야 한다. 기존의 데이터 스트림을 위한 익명화 기법은 일정량의 데이터를 저장하였다가 익명화하는 방법을 사용한다. 이러한 기존의 데이터 스트림을 위한 익명화 기법은 배포의 실시간성을 보장하지 못하고, 큰 정보 손실을 발생시킨다. 본 논문에서는 이러한 기존 기법의 문제점을 해결하기 위하여, 무지연 익명화(Delay-free) 기법을 제안한다. 무지연 익명화는 레코드를 입력되는 즉시 Anatomy 방법으로 배포하여 데이터 활용의 실시간성을 보장한다. 또한 배포된 레코드의 정보를 유지하여 높은 정보 유용성을 보인다. 몇 가지 성능 향상 기법은 익명화 과정에서 발생하는 정보 손실을 크게 감소시킨다. 알려진 바로, 무지연 익명화는 데이터 스트림의 실시간성을 보장하는 최초의 기법이다. 실험을 통하여 제안하는 기법이 높은 정보유용성을 보임을 확인한다.
Data streams which contain private information should be utilized after the anonymization process. Existing anonymization methods for data streams utilize the manner which stores a certain amount of data and then anonymizes them. These methods suffer from no guarantee of real-time publishing and generate huge information loss. To solve the defects of existing studies, in this paper, we propose a delay-free anonymization method. The delay-free anonymization guarantees real-time application by publishing records immediately using the Anatomy approach just after the input. It also shows high utility by managing the information of published records. Several performance improvement techniques highly decrease the information loss generated during the anonymization process. To our best knowledge, the delay-free anonymization is the first method for guaranteeing the real-time property of data streams. We demonstrate that the proposed anonymization achieves high utility by experiments.
집합값을 갖는 반정형 트랜잭션 데이터 익명화에서 정보손실 측도 개선에 관한 연구
[NRF 연계] 사단법인 미래융합기술연구학회 아시아태평양융합연구교류논문지 Vol.10 No.9 2024.09 pp.45-54
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
정형데이터란 일반적으로 테이블에서 하나의 셀 내에 하나의 문자나 혹은 숫자로 구성된 데이터를 말한다. 만일 여기서 하나의 셀 내에 하나의 문자나 혹은 숫자가 아니라 여러 문자나 혹은 숫자들의 집합 값으로 구성된 경우, 우리는 이것을 반정형 데이터라 부른다. 예로 우리가 마트에서 구매한 일련의 물품 아이템 들이나 혹은 병원에서 환자에 대한 여러 병명들에 대한 목록이 그러하며, 이때 한 명에 대한 집합 값들로 구성된 목록을 하나의 트랜잭션이라 부른다. 우리는 이러한 일련의 집합값들로 구성된 반정형 트랜잭션 데이터셋에서 개인정보보호를 위한 익명화 문제를 다룬다. 즉, 개인이 구매한 상품 목록이나 혹은 환자의 병명에 대한 사항들은 민감정보로서 개인의 프라이버시 차원에서 보호되어야 할 개인정보들이기 때문이다. 이러한 익명화 문제와 관련하여 우리는 기존의 LG (Local Generalization) 알고리즘을 개선하여 LGR (Local Generalization & Reallocation)이라는 새로운 알고리즘을 제안한바 있다. 그러나 만일 우리가 익명화된 개인정보들을 활용하거나 분석하는 관점에서 바라보면 안전하기도 해야하지만 이에 못지 않게 데이터가 쓸모가 있어야 한다. 즉 데이터 품질이 분석에 용이하도록 유용해야한다. 이는 반대로 익명화 과정에서 정보 손실이 최소화 되어야한다는 것과 동일한 의미를 갖는다. 기존 우리의 LGR 알고리즘은 정보 손실을 계산하기 위한 측도로 기존 LG 알고리즘과 동일한 NCP (Normalized Certainty Penalty)를 사용하였다. NCP 측도는 전체 아이템 수 대비 일반화된 아이템의 비율로 계산된다. 따라서 계산이 단순하고, 다양한 데이터셋에 쉽게 적용할 수 있는 장점이 있지만, 그 반대로 정보 손실이 많아 데이터의 유용성을 떨어뜨릴 수 있는 단점이 있다. 우리는 이러한 단점을 개선하고자 새로운 IGH (Information Gain-based Heuristic) 측도를 새로이 제안하고 이를 이론적으로 검증해 보고자 한다. 제안하는 측도는 기존 NCP 방식에 비하여 정보 손실을 최소화하고 데이터의 유용성을 최대한 보존할 수 있는 장점이 있다.
Structured data typically refers to data composed of a single character or number within a single cell of a table. If, instead, a single cell contains a set of multiple characters or numbers, we refer to this as semi-structured data. For example, the list of items purchased at a supermarket or diagnoses for a patient in a hospital represents such cases. In these cases, the list of values for an individual is called a transaction. We address the anonymization problem for privacy protection in semi-structured transaction datasets composed of these values. Specifically, lists of purchased items or diagnoses contain sensitive information that must be protected for individual privacy. Regarding this anonymization problem, we previously proposed an improvement to the existing Local Generalization (LG) algorithm, resulting in the new Local Generalization & Reallocation (LGR) algorithm. However, the data must be secure and valuable from the perspective of utilizing or analyzing anonymized personal data. This implies that data quality must be preserved to facilitate analysis, which means that information loss during anonymization must be minimized. Our existing LGR algorithm used the Normalized Certainty Penalty (NCP) as a measure for calculating information loss, the same as the existing LG algorithm. The NCP measure calculates the ratio of generalized items to the total number of items. While this measure is simple to compute and easily applicable to various datasets, it has the drawback of potentially high information loss, which can reduce data utility. To address this drawback, we propose a new Information Gain-based Heuristic (IGH) measure and aim to verify its effectiveness theoretically. The proposed measure has the advantage of minimizing information loss and maximizing data utility compared to the existing NCP method.
반정형 트랜잭션 데이터를 위한 새로운 익명화 알고리즘 설계 및 구현에 관한 연구
[NRF 연계] 사단법인 미래융합기술연구학회 아시아태평양융합연구교류논문지 Vol.9 No.11 2023.11 pp.13-22
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
트랜잭션 데이터는 다양한 고객이 슈퍼마켓에서 구매한 품목에 대한 정보를 저장하는 관계형 데이터베이스를 고려한 것으로 하나의 셀 내에 여러 개의 아이템들이 일종의 집합으로 구성된 것을 말한다. 본 본문은 이러한 트랜잭션 데이터에서의 익명화에 관한 문제를 다루고자 한다. Y. He와 J. Naughton은 이 문제를 해결하기 위해 일명 ‘로컬 일반화’라고 부르는 k-익명성 기반 기법을 제안한 바 있다. 그러나 이 기법은 강력한 프라이버시를 제공하고 실행시간이 짧다는 장점이 있지만 정보손실이 크다는 단점이 있다. 한편 J. Liu와 K. Wang은 이러한 단점을 해결하기 위하여 HgHs(Heuristic generalization with Heuristic suppression) 알고리즘을 제안한 바 있다. 그러나 이 기법은 그 반대로 정보손실은 적지만 실행시간이 오래 걸린다는 단점이 있다. 본 저자는 기존에 이러한 단점들을 개선하기 위하여 상기 두가지 기법들에 비해 정보 손실은 최소화 하면서도 강력한 프라이버시를 보장할 수 있는 새로운 기법을 제안한 바 있다. 본 논문에서는 이를 실제 상품화할 수 있도록 시스템을 새로이 구성하고 기존 제안 알고리즘의 설계 및 구현과정과 테스트 결과를 제시하고자 한다. 테스트는 기존 로컬일반화와 HgHs 기법에서 사용한 데이터셋(BMS-WebView2와 BMS-POS)을 동일하게 사용하였으며, 테스트 결과 제안 알고리즘의 정확도를 확인할 수 있었다.
The transaction data refers to information stored in a relational database, considering various customers' purchases at a supermarket, where multiple items are grouped together as a kind of set within a single cell. This paper aims to address the issue of anonymization in such transaction data. Y. He and J. Naughton have previously proposed a technique known as 'Local Generalization,' which is based on k-anonymity to tackle this problem. However, this technique offers strong privacy protection and short execution times but suffers from significant information loss. On the other hand, J. Liu and K. Wang have proposed the HgHs (Heuristic generalization with Heuristic suppression) algorithm to mitigate these drawbacks, with reduced information loss but longer execution times. In order to improve these existing shortcomings, we have proposed a new technique that can guarantee strong privacy while minimizing information loss compared to the above two techniques. In this paper, we will reconfigure the system so that it can be commercialized and present the design, implementation process, and test results of the existing proposed algorithm. The test used the datasets (BMS-WebView2 and BMS-POS) used in the existing local generalization and HgHs techniques, and the test results confirmed the accuracy of the proposed algorithm.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.