공용 PC 환경에서 잔존 파일 기반 개인정보 노출 방지를 위한 경량 PDF 개인정보 사전 탐지 시스템 설계 및 구현
Design and Implementation of a Lightweight PDF-Based Personal Information Pre-Detection System for Preventing Privacy Exposure from Residual Files in Shared PC Environments
Shared PC environments, such as public service kiosks, university computer labs, libraries, and self-service certificate issuance systems, frequently expose users to privacy risks caused by residual PDF files left on systems after viewing or printing documents. Conventional Data Loss Prevention (DLP) solutions primarily focus on preventing data leakage during external transmission processes, including email transfers, USB storage copying, and cloud uploads, and thus have limitations in addressing privacy exposure caused by residual files within systems. This study proposes a lightweight PDF-based personal information pre-detection system to prevent privacy exposure from residual files in shared PC environments. The proposed system extracts text from PDF documents and employs regular expression–based detection techniques to identify sensitive information, including phone numbers, email addresses, resident registration numbers, and student identification numbers, in real time. Furthermore, the system is designed with an extensible architecture that can be integrated with automatic encryption and access control mechanisms. Experimental evaluations using 100 PDF documents achieved a Precision of 0.96, Recall of 0.93, and F1-score of 0.94, with an average processing time of 12 ms per PDF document. Comparative experiments with a BERT-based detection model demonstrated superior processing speed and resource efficiency of the proposed approach. The results indicate that the proposed system can serve as a practical and lightweight security solution for enhancing privacy protection in shared computing environments.
한국어
공공기관, 대학 전산실, 도서관 및 무인민원발급기 등 공용 PC 환경에서는 사용자가 문서를 열람하거나 출력한 후 PDF 파일을 삭제하지 않아 개인정보가 노출되는 사례가 지속적으로 발생하고 있다. 기존 데이터 유출 방지(DLP) 기술은 이메일 전송, USB 저장매체 복사 및 클라우드 업로드 등 외부 반출 과정의 정보 유출 방지에 초점을 두고 있어 시스템 내부에 남겨진 잔존 파일에 의한 개인정보 노출 문제에는 효과적으로 대응하지 못한다. 본 연구에서는 공용 PC 환경에서 잔존 PDF 파일로 인한 개인정보 노출을 방지하기 위한 경량 개인정보 사전 탐지 시스템을 제안한다. 제안 시스템은 PDF 문서로부터 텍스트를 추출한 후 정규표현식 기반 탐지 기법을 적용하여 전화번호, 이메일, 주민등록번호 및 학번 등의 개인정보를 실시간으로 분석한다. 또한 향후 자동 암호화 및 접근통제 기능과 연계 가능한 확장형 구조를 적용하였다. 100개의 PDF 문서를 대상으로 성능을 평가한 결과, Precision 0.96, Recall 0.93, F1-score 0.94를 기록하였으며 평균 처리시간은 12ms/PDF로 나타났다. 또한 BERT 기반 탐지 모델과 비교하여 우수한 처리 속도와 자원 효율성을 확인하였다. 연구 결과는 공용 PC 환경의 개인정보 보호 수준 향상을 위한 실용적인 경량 보안 기술로 활용될 수 있음을 보여준다.
목차
요약 ABSTRACT 1. 서론 2. 관련 연구 2.1 개인정보 보호 기술 2.2 개인정보 탐지 기술 2.3 PDF 문서 보안 및 개인정보 보호 연구 2.4 기존 연구와의 차별성 3. 위협 모델 및 시스템 요구사항 3.1 위협 시나리오 3.2 시스템 요구사항 4. 제안 시스템 설계 및 구현 4.1 시스템 구조 4.2 시스템 아키텍처 4.3 개인정보 탐지 절차 4.4 텍스트 추출 모듈 4.5 개인정보 탐지 모듈 4.6 자동 보호 확장 구조 5. 실험 및 평가 5.1 실험 환경 5.2 성능 평가 지표 5.3 개인정보 탐지 성능 평가 5.4 Regex 기반 방식과 BERT 기반 방식 비교 5.5 기존 DLP와 제안 시스템 비교 5.6 결과 분석 6. 결론 및 향후 연구 참고문헌