년 - 년
NFT/C2PA 메타데이터를 활용한 AI 학습데이터 저작권 관리 방안 연구 KCI 등재
한국지식재산학회 산업재산권 제80호 2025.04 pp.331-368
※ 기관로그인 시 무료 이용이 가능합니다.
8,200원
인공지능 기술 발전 초기에는 혁신 촉진을 위해 저작물의 학습데이 터 활용 규제가 완화되었으나, 생성형 AI가 상용화되고 수익모델로 확 립되면서 저작권자들의 적정 보상 요구가 심화되고 있다. EU와 미국을 포함한 주요국에서는 저작물 이용의 체계적 관리와 투명성 확보를 위한 법제를 도입하는 추세이며, 기술 혁신과 저작권자 보상이라는 두 가치 의 균형적 발전을 위한 법적·기술적 프레임워크 구축이 시급하다. 옵트아웃 방식을 통한 AI 학습데이터 관리는 저작권자 권리 보호를 위한 직접적 메커니즘이나 실질적 적용에 여러 제약이 존재한다. EU AI Act는 텍스트 및 데이터 마이닝(TDM) 면책 규정과 함께 저작권자의 권리 유보 원칙을 명시하여 창작자가 기계 판독 가능한 방식으로 옵트 아웃할 수 있게 제도화했으나, 표준화된 실행 절차 부재와 명확한 기 준 미비로 규제 집행의 일관성이 결여되어 있다. 더욱이 규제 회피 가 능성과 옵트아웃 조치를 무력화하는 기술적·법적 우회 전략은 이 접 근법의 실효성에 의문을 제기한다. AI 학습데이터 투명성 확보를 위한 미국의 입법적 노력은 AI 블랙박 스 문제 해소를 통해 저작권자와 AI 개발자 간 협의 기반을 구축하는 데 중점을 둔다. 2024년 제안된 TRAIN Act는 저작권자가 자신의 저작 물이 AI 학습에 활용되었는지 확인할 수 있도록 AI 개발자에 대한 행 정적 소환장 발부 요청권을 부여한다. TRAIN Act의 행정적 소환장은 주관적 선의의 믿음이라는 완화된 발급 요건으로 신속한 정보 접근이 가능하나, 이러한 낮은 수준의 심사 기준은 소환장 남용 가능성을 증 가시키고 AI 개발자에게 불균형적 부담을 초래할 위험이 있어 입법적 보완이 요구된다. AI 학습데이터는 전통적 저작물 이용과 근본적으로 다른 방식으로 작동한다. 원저작물을 직접 소비하지 않고 정제, 토큰화, 정규화 등의 전처리를 거쳐 학습에 적합한 형태로 변환하여 이용하는 복합적 프로 세스는 기존 저작권 관리 체계로는 효과적으로 추적하거나 규제하기 어려운 구조적 한계를 보인다. AI 학습데이터 및 산출물에 대한 저작권 관리의 복잡성을 고려할 때, 기술적 지원 방안으로 NFT 및 C2PA 기반 메타데이터를 활용한 학 습데이터 관리 방안을 검토할 수 있다. NFT와 블록체인 기술은 저작물 의 출처, 이용 내역, 라이선스 계약을 불변의 분산원장에 기록함으로써 투명하고 추적 가능한 저작권 관리 메커니즘을 제공한다. C2PA 표준의 메타데이터 서명 및 검증 기술은 콘텐츠의 진위성을 보장하고, AI 학 습데이터의 출처와 변형 이력을 암호화된 방식으로 추적할 수 있게 한 다. 이는 저작권자의 사전 통제권과 사후 보상체계의 실효성을 제고하 는 수단이 될 수 있으며, NFT 기반의 추적 시스템과 결합될 경우 보다 정밀한 권리 관리가 가능하다. 다만, 이러한 기술적 보호조치가 현행 저작권법 체계 내에서 법적 효력을 갖기 위해서는, 그 적용 방식과 권 리구조에 대한 정합성 검토가 선행되어야 하며, 기술의 표준화와 산업 내 확산을 위한 제도적 지원이 함께 마련되어야 한다.
In the early stages of AI technology development, regulations on the use of copyrighted works as training data were relaxed to promote innovation. However, as generative AI becomes commercialized and establishes itself as a viable business model, demands from copyright holders for fair compensation have intensified. Major jurisdictions, including the EU and the United States, are introducing legal frameworks to systematically manage copyrighted content usage and ensure transparency. It is now imperative to establish a legal and technological framework that balances technological innovation with fair compensation for copyright holders. The opt-out mechanism for managing AI training data serves as a direct method to protect copyright holders' rights, but its practical implementation faces several limitations. The EU AI Act institutionalizes the opt-out process by specifying the principle of rights reservation, allowing creators to opt out in a machine-readable format alongside the text and data mining (TDM) exception. However, the lack of standardized implementation procedures and clear guidelines has resulted in inconsistent regulatory enforcement. U.S. legislative efforts to ensure AI training data transparency primarily aim to address the AI black box problem and establish a foundation for negotiations between copyright holders and AI developers. The TRAIN Act, proposed in 2024, grants copyright holders the right to request administrative subpoenas against AI developers to verify whether their works have been used in AI training. While the TRAIN Act facilitates rapid access to information through a reduced issuance threshold based on subjective good faith belief, this lower evidentiary standard increases the risk of subpoena abuse and imposes a disproportionate burden on AI developers, necessitating legislative refinements. Given the complexity of copyright management for AI training data and outputs, the use of NFT and C2PA-based metadata may be considered as a technical support measure. NFT and blockchain technologies provide a transparent and traceable mechanism by recording the provenance, usage history, and licensing terms of works on an immutable distributed ledger. The C2PA standard offers metadata signing and verification technologies that ensure content authenticity and enable encrypted tracking of the origin and modification history of AI training data. These technologies may enhance both the ex-ante control and ex-post compensation mechanisms for copyright holders. When combined with NFT-based tracking systems, they can support more precise rights management. However, for such technical protection measures to have legal effect under the current copyright framework, a prior legal assessment of their compatibility with existing rights structures is required, along with institutional support for standardization and broader industry adoption.
인공지능 학습용 데이터 품질에 대한 연구 : 퍼지셋 질적비교분석 KCI 등재
한국경영정보학회 경영정보학연구 제26권 제1호 2024.02 pp.19-56
※ 기관로그인 시 무료 이용이 가능합니다.
8,200원
본 연구는 한국의 인공지능 학습용 데이터 구축 사업과 데이터의 공공 개방에 관한 정책 수행 기관, 데이터 구축 기업, 그리고 이를 활용하는 다양한 기관의 데이터 품질에 대해 이해를 제고하고, 신뢰할 수 있는 인공지능 알고리즘 개발에 있어 가장 중요한 학습용 데이터 품질에 대한 이론적 토대를 만들기 위한 실증적 연구이다. 이를 위해, 데이터의 속성 요인, 데이터 구축환경 요인, 데이터 타입 관련 요인 등 인공지능 학습용 데이터 품질과 관련된 중요 선행요인을 도입하여 이론적 모형을 제안한다. 본 연구는 393명의 인공지능 학습용 데이터 구축 기업과 인공지능 서비스 개발 기업의 실무 담당자를 대상으로 설문조사를 실시하여 데이터를 수집하였다. 데이터 분석은 퍼지셋 질적비교분석 방법과 인공신경망 분석을 통해 이루어졌으며, 분석 결과를 통해 인공지능 학습용 데이터 관련 학술적 및 실무적 시사점을 도출했다.
This study is empirical research to enhance understanding of AI (artificial intelligence) training data project in South Korea. It primarily focuses on the various concerns regarding data quality from policy-executing institutions, data construction companies, and organizations utilizing AI training data to develop the most reliable algorithm for society. For academic contribution, this study suggests a theoretical foundation and research model for understanding AI training data quality and its antecedents, as well as the unique data and ethical aspects of AI. For this purpose, this study proposes a research model with important antecedents related to AI training data quality, such as data attribute factors, data building environmental factors, and data type-related factors. The study collects 393 sample data from actual practitioners and personnel from companies building artificial intelligence training data and companies developing artificial intelligence services. Data analysis was conducted through Fuzzy Set Qualitative Comparative Analysis (fsQCA) and Artificial Neural Network analysis (ANN), presenting academic and practical implications related to the quality of AI training data.
생성형 인공지능의 저작권 침해 예외 규정으로서 공정이용 법리의 유효성에 관한 시론적 고찰 KCI 등재
유럽헌법학회 유럽헌법연구 제44호 2024.04 pp.269-303
※ 기관로그인 시 무료 이용이 가능합니다.
7,800원
생성형 인공지능이 ‘일반인공지능’(generative AI)으로 사회 전반에 걸쳐 상용화됨에 따라 지적 재산권 및 저작권 보호에 관한 논쟁이 새로 운 국면을 맞이하고 있다. 생성형 인공지능은 기존의 인공지능과 달리 코딩이 아닌 인간의 언어로 사용자가 직접 명령을 내리고, 다양하게 개 발되고 있는 멀티모달 프로그램을 통해 그림이나 사진, 동영상 등의 영 상자료를 바로 인식하여 하나의 인터페이스에서 인간의 언어로 표현된 명령의 결과물을 생산한다. 또한 특정분야에서 전문가적 지식을 활용했 던 기존의 인공지능과는 달리, 사회 일반에 걸쳐 통합적으로 활용할 수 있는 파운데이션 모델로서 범용인공지능으로 기능하고 있다는 점에서 기 존의 인공지능과 차별성을 가진다. 생성형 인공지능의 저작권 침해 문제는 크게 두 가지 측면에서 논의 되고 있는데, 첫째, 생성형 인공지능이 개발 단계에서 학습한 방대한 데 이터에 대한 저작권 침해 이슈이고, 둘째, 방대한 데이터 학습을 통해 개발된 생성형 인공지능이 만들어 내는 다양한 생성물과 이차 저작물에 의한 저작권 침해 이슈이다. 물론 기존 인공지능 모델들도 학습에 데이 터를 사용했고, 이로 인한 저작권 침해 갈등은 존재해 왔다. 그리고 이 때 법은 인공지능산업 진흥의 관점에서 저작권 침해의 예외 적용을 위해 공정이용의 법리, TDM 면책규정, 이용약관 등의 다양한 방식들을 고민 해 왔다. 그런데 생성형 인공지능의 등장과 함께 그간 저작권 침해의 예외 적 용을 위해 논의되었던 다양한 법제들의 적용과 해석이 새로운 국면을 맞 이하고 있다. 세계 각국이 앞다투어 발표했던 TDM 면책규정을 유예하 였고, 미국의 경우 Campbell 판례 이후 저작권 침해 예외 인정을 위해 유지되어 오던 공정이용의 심사 강도가 최근 판례에서 변경되고 있다. 한편, 우리의 경우 TDM 면책규정과 관련하여, 외국의 사례들과 유사 하게 2021년 말 인공지능산업의 진흥을 위해 저작권법 개정안에 포함되 어 발표되었으나 2022년 생성형 인공지능의 등장과 함께 유예되었다. 그런데 최근 공정이용의 법리를 적용한 우리의 판례에서는 생성형 인공 지능 개발 이전으로 회기하여 미국 구글의 공공도서관 판례에서의 공정 이용 판단기준과 유사한 판결을 내린바 있다. 그렇다면 생성형 인공지능 의 시대를 맞이하여 우리의 판례에도 변경이 필요한 것인지, 아니면 변 형적 이용을 폭넓게 인용하여 이차 저작물을 만들 수 있는 포괄적인 권 리를 부여하고 이를 통해 새로운 산업분야의 진흥을 도모할 것인지에 대 한 논의가 필요할 것이다. 이에 본고에서는 그간 우리나라와 미국에서 저작권 침해의 예외 사유 로 적용되어 온 공정이용의 법리와 최근 미국판례에서의 공정이용 법리의 변화를 분석하고, 공정이용의 법리가 범용인공지능으로서 생성형 인 공지능의 등장이라는 새로운 환경에서도 여전히 유효하게 적용될 수 있는지에 대한 시론적 고찰을 통해 공정이용의 법리가 생성형 인공지능의 데이터 학습 등 다양한 저작권 침해 사례에서 일정 부분 여전히 유효성 을 가짐을 논의하고자 한다.
As generative artificial intelligence is commercialized throughout society as ‘generative AI’, the debate over intellectual property rights and copyright protection is entering a new phase. Unlike existing artificial intelligence, generative artificial intelligence allows users to directly issue commands using human language rather than coding, and recognizes visual data such as pictures, photos, and videos directly through a variety of multimodal programs being developed. It is possible to express it in human language in the interface. In addition, unlike existing artificial intelligence that utilizes expert knowledge in a specific field, it has the greatest differentiation in that it functions as a ‘general artificial intelligence’ (generative AI) that can be utilized comprehensively throughout society. The problem of copyright infringement in generative artificial intelligence is being discussed from two aspects. First, the issue of copyright infringement on the vast amount of data learned by generative artificial intelligence during the development stage, and second, the issue of copyright infringement on the vast amount of data learned by generative artificial intelligence during the development stage. This is an issue of copyright infringement by various products. Of course, existing artificial intelligence models also used data for learning, and there have been conflicts over copyright infringement as a result. And from the perspective of promoting the artificial intelligence industry, the law has been considering various methods, such as the legal principles of fair use, TDM exemption provisions, and terms of use, to apply exceptions to copyright infringement. However, with the emergence of generative artificial intelligence, various laws that have been discussed for exceptions to copyright infringement are entering a new phase. The TDM exemption regulations that countries around the world had rushed to announce were postponed, and in the case of the United States, the strength of fair use review, which has been maintained to recognize exceptions to copyright infringement since the Campbell precedent, has been strengthened in recent precedents. Meanwhile, in our case, the TDM exemption regulation was announced as included in the “Copyright Act Amendment” at the end of 2021, similar to foreign cases, but was postponed with the emergence of generative artificial intelligence in 2022. However, in our recent case law applying the legal principle of fair use, a ruling was made similar to the fair use judgment criteria in the US case law before the development of generative artificial intelligence. So, is it necessary to change our precedents in the era of generative artificial intelligence, or should we grant comprehensive rights to create secondary works by broadly citing transformative use and promote new industrial fields through this? Discussion will be needed. Accordingly, in this paper, we analyze the legal principle of fair use that has been applied as an exception to copyright infringement in Korea and the United States and the changes in the fair use legal principle in recent U.S. cases, and the legal principle of fair use is used to apply generative artificial intelligence as general artificial intelligence. Through a theoretical consideration of whether it can still be applied effectively even in the new environment of its emergence, we would like to discuss that the legal principle of fair use still has some validity in various copyright infringement cases, such as data learning of generative artificial intelligence.
AI 학습용 데이터의 이용 행태 및 구조적 소비 패턴 분석 - AI 허브 로그 데이터 기반 연관 규칙 및 사회 연결망 분석 - KCI 등재
한국EA학회 정보화연구 제22권 4호 2025.12 pp.385-394
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 초거대 AI의 부상으로 고품질 학습용 데이터의 중요성이 증대됨에 따라, 본 연구는 기존 설문조사 기반 연구의 한계를 넘어 대규모 로그 데이터를 활용해 AI 허브 이용자의 실제 데이터 소비 행태를 실증적으로 규명하고자 하였다. 이를 위해 2024년과 2025년의 이용 로그를 바탕으로 이용 그 룹별 방식 변화를 통계적으로 검증하고, 연관 규칙 및 사회 연결망 분석(SNA)을 통해 데이터 간의 구 조적 관계를 시각화하였다. 분석 결과, 개인 이용자는 웹 중심의 탐색을 유지하는 반면, 기업 및 기관 이용자는 API 기반의 자동화된 이용 방식으로 급격히 전환되고 있음이 확인되었다. 또한 데이터 소 비 패턴 분석 시, 텍스트 데이터는 네트워크 전반에 걸쳐 높은 연결 중심성을 가지며, 영상 및 이미지 데이터는 높은 향상도(Lift)를 기반으로 특정 도메인 내에서 강하게 결합된 형태의 소비 특성을 보였 다. 본 연구는 이러한 결과를 바탕으로 이용자 유형에 따른 이원화된 플랫폼 운영 정책과 데이터 특 성을 고려한 큐레이션 전략을 제안한다.
This study empirically analyzes AI Hub user behavior using large-scale log data from 2024 to 2025, addressing the growing demand for high-quality training data. We applied Association Rule Mining and Social Network Analysis (SNA) to visualize structural consumption patterns. Results confirm a divergence: individual users prefer web-based exploration, while organizations are shifting toward API-based automation. Network analysis reveals that text data acts as a central “anchor” with broad connectivity, whereas image and video data exhibit strong, domain-specific clustering based on high lift values. Consequently, we propose a dual-track strategy: enhancing web UX for individuals and API infrastructure for organizations, alongside purpose- driven data curation to optimize the AI ecosystem.
생성형 AI 기반 언어모델 성능 최적화를 위한 학습데이터 구축 전략 KCI 등재
한국융합보안학회 융합보안논문지 제25권 제1호 2025.03 pp.217-224
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
생성형 언어모델은 최근 AI 기술의 발전과 함께 다양한 산업 분야에서 혁신을 주도하고 있다. 모델이 최적의 성능을 발휘하기 위해서는 충분한 양의 고품질 학습데이터 확보가 필수적이며, 이는 모델의 일반화 성능을 향상시키고 신뢰성 높은 결과를 제공하는 데 중요한 역할을 한다. 이에 본 연구에서는 학습데이터의 양과 품질이 생성형 AI 기반 언어모델 의 성능에 미치는 영향을 분석하고, 기존의 데이터 양(Quantity) 중심 접근 방식에서 벗어나 데이터 품질(Quality)을 반 영한 손실함수(loss function) 확장 모델을 제안하였다. 제안된 모델의 유효성을 정량․실증적으로 검증 후 데이터의 양 과 질을 기반으로 한 2×2 모델을 적용하여 충분한 고품질 학습데이터의 구축이 AI 성능 최적화에 필수적인 요소임을 입증하고 모델의 성능을 극대화하기 위한 최적의 학습데이터 구축 전략을 제시한다.
Generative language models have been driving innovation across various industries with the rapid advancement of AI technology. To achieve optimal performance, these models require a sufficient amount of high-quality training data, which plays a crucial role in enhancing generalization capabilities and ensuring reliable outputs. This study analyzes the impact of training data quantity and quality on the performance of generative AI-based language models and proposes an extended loss function model that incorporates data quality, moving beyond the traditional quantity-centric approach. The proposed model is quantitatively and empirically validated, followed by the application of a 2×2 model based on data quantity and quality. Through this approach, the research demonstrates that constructing high-quality training data is essential for optimizing AI performance and presents a strategic framework for maximizing language model performance.
[NRF 연계] 사단법인 미래융합기술연구학회 아시아태평양융합연구교류논문지 Vol.12 No.2 2026.02 pp.17-41
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper explores the derivative implications of the absence of responsibility in AI music training datasets for the copyright ownership of AI-generated music. The research context stems from the disruptive impact of generative AI technology on the music industry in 2025, where the massive training data it relies on carries significant copyright risks. During the construction of AI music generation model training datasets, the unauthorized use of copyrighted content means that accompanying licensing agreements cannot fully mitigate infringement risks. The study finds that this lack of accountability at the upstream of the "data supply chain" is amplified by technological "black boxes" and complex data processing workflows, posing a fundamental challenge to determining copyright ownership of downstream AI-generated musical works. To systematically analyze this issue, this study employs grounded theory methodology and conducts a three-tier coding analysis of multi-source policy and industry texts. This ultimately constructs the "Conflict, Response, and Reconciliation" (CRR) theoretical model for AI data governance. The model reveals a dynamic evolutionary path: from "technological and legal conflicts" to "systemic risks," then through "multi-level governance responses" toward "economic reconciliation". The study identifies "training data traceability" and "training data lifecycle management" as key mediating variables linking governance and reconciliation. Building upon this foundation, a systematic governance strategy centered on a trusted data foundation has been proposed. This strategy encompasses four dimensions: a technical compliance system (resolving conflicts), risk buffer mechanisms (managing risks), a rule-of-law regulatory environment (implementing governance), and market allocation models (achieving reconciliation). It provides both theoretical underpinnings and a practical roadmap for resolving AI music copyright challenges and fostering the industry's healthy, sustainable development.
부정경쟁방지법상 AI 학습데이터 보호 체계에 관한 연구 - 영업비밀·데이터 부정사용·성과도용의 기능 분담을 중심으로 -
[NRF 연계] 강원대학교 비교법학연구소 강원법학 Vol.83 2026.05 pp.295-372
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
생성형 인공지능의 확산으로 학습데이터를 둘러싼 사적 권리관계의 법리적 정립이 시급한 과제로 대두되었다. 본 논문은 부정경쟁방지법의 영업비밀, 카목(데이터 부정사용), 파목(성과도용)의 세 조항이 AI 학습데이터에 어떻게 적용되는지를 분석한다. 본 논문의 핵심 명제는 세 조항이 단순한 보충 관계가 아니라 ‘기능 분담’ 관계로 작동한다는 것이다. 영업비밀은 학습데이터의 비밀성을, 카목은 거래성을, 파목은 창출성을 각각 보호한다. 그러나 영업비밀은 비밀관리성 요건이 거래·유통 환경과, 비공지성 요건이 공개 의무와 각각 충돌하는 구조적 한계를 지니며, 카목은 거래성 요건으로 인해 자체 학습용 데이터에는 적용되지 않는다. 파목은 두 규정 어디에도 포섭되지 않는 영역, 즉 외부 콘텐츠를 수집하여 자체 학습에 사용하는 행위를 포착한다. 학습데이터의 비공지성 판단에서는 대법원 2024년 맥주제조기 판결의 ‘공지된 정보의 조합과 비공지성’ 법리가 큐레이션의 영업비밀성을 인정하는 토대가 된다. 파목 적용에서는 야놀자 v. 여기어때 사건의 민사 판결이 확립한 “공개된 정보의 무단 수집·사용도 파목 위반이 될 수 있다”는 법리가 핵심 근거가 되며, BTS 사건이 제시한 위법성 판단의 세 징표가 학습데이터 분쟁의 판단 척도로 기능한다. 각국의 보호 구조를 비교하면, 일본은 영업비밀과 한정제공데이터의 두 층을, 미국은 영업비밀을, EU는 영업비밀과 데이터베이스권의 두 층을 두고 있으나, 세 나라 모두 데이터의 창출성을 보호하는 성과도용 일반조항을 두지 않는다. 한국은 학습데이터의 세 보호 차원에 모두 대응하는 보호 수단을 갖춘 유일한 법체계로서 비교법적 강점을 가진다. 이러한 '기능 분담' 명제는 진행 중인 방송 3사 v. 네이버 사건을 비롯한 학습데이터 분쟁에 직접 적용될 수 있다. 가령 네이버는 자신의 학습데이터셋을 영업비밀로 주장하면서도, 그 구축 과정에서의 방송 3사 콘텐츠 무단 사용에 대해서는 파목 위반의 책임을 질 수 있다. 나아가 본 논문은 학습데이터 공개 의무와 영업비밀 보호의 균형을 위한 운용 기준을 제시하여 AI 거버넌스의 실무적 과제에 시사점을 제공한다.
With the proliferation of generative artificial intelligence, the doctrinal articulation of private rights over training data has emerged as an urgent task. This article analyzes how three provisions of the Korean Unfair Competition Prevention Act?trade secret protection (Article 2(2)), the data misappropriation clause (Article 2(1)(ka)), and the catch-all clause on misappropriation of substantial outputs (Article 2(1)(pa))?apply to AI training data. The article advances a central thesis: rather than operating in a merely supplementary relationship, the three provisions operate in a relationship of functional division. Trade secret protection covers the secrecy dimension of training data; the data misappropriation clause covers its tradability dimension; and the catch-all clause covers its creative-output dimension. Trade secret protection faces two structural limitations: the secrecy-management requirement conflicts with the realities of training data trading and distribution, while statutory transparency (disclosure) obligations erode the requirement that the information remain not generally known (the non-public-knowledge requirement). The data misappropriation clause does not reach data used solely for an operator's own model training, due to its tradability requirement. The catch-all clause occupies the residual space, governing the unauthorized collection and use of external content for in-house training?the core fact pattern of AI training data disputes. The Korean Supreme Court's 2024 ruling on combinations of publicly available information provides the foundation for recognizing the trade-secret status of training data curation. For the catch-all clause, the civil ruling in Yanolja v. GoodChoice?holding that unauthorized scraping of publicly available data may also violate it?supplies the key basis, while the three indicia of unlawfulness established in the BTS case serve as the operative criteria for AI training data disputes. Comparing each jurisdiction's protective structure, Japan has two layers (trade secrets and limitedly provided data), the United States has trade secret protection, and the European Union has trade secrets and the sui generis database right; yet none of the three provides a catch-all clause protecting the creative-output dimension of data. Korea is the only legal system that corresponds to all three protective dimensions of training data, which constitutes its comparative advantage. This functional-division thesis applies directly to ongoing disputes such as KBS·MBC·SBS v. Naver, in which Naver may simultaneously claim trade-secret protection over its training dataset while bearing liability under the catch-all clause for unauthorized use of the broadcasters' content during dataset construction. The article further offers operational criteria for balancing training data transparency obligations with trade secret protection, providing concrete guidance for AI governance practice.
AI 학습데이터와 개인정보 권리의 경계 : 에이닷(A.) 사례를 통해 본 통제와 거버넌스의 과제
[NRF 연계] 한양대학교 제3섹터연구소 시민사회와 NGO Vol.23 No.1 2025.05 pp.195-245
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
생성형 인공지능(Generative AI)의 확산은 방대한 양의 고품질 학습 데이터를 필수 자원으로 만들고 있으며, 이에 따라 개인정보 수집·활용을 둘러싼 법적·윤리적 문제는 더욱 복잡하고 구조화된 형태로 전개되고 있다. 특히 통신 기반 AI 서비스는 실시간으로 민감 데이터를 수집할 수 있는 구조를 갖추었으나, 이러한 데이터가 인공지능 학습에 활용되는 과정에서 그 정당성과 책임 구조는 여전히 불투명하다. 본 연구는 SK텔레콤의 AI 비서 서비스 ‘에이닷(A.)’ 사례를 통해 통화 데이터, 대화 기록, 제3자 연동 정보 등 고위험 개인정보 활용 방식과 이에 따른 사용자 통제권 침해, 제3자 권리 미보장, 알고리즘 불투명성의 문제를 분석하였다. 특히 통화 데이터의 처리 과정은 개인정보보호법(PIPA)상 목적 제한성 원칙 및 사전 동의 체계와 충돌하며, 기존 법제도가 AI 학습 데이터의 복합성과 재사용 가능성을 충분히 포섭하지 못하고 있음을 실증적으로 드러낸다. 또한 Stack Overflow 사례를 통해, 공개된 데이터라도 정보 주체의 권리 고지, 활용 목적의 명확성, 저작권 보호 등 최소한의 규범 요건이 충족되지 않으면 법적·윤리적 위반으로 전환될 수 있음을 밝혔다. 이러한 분석을 바탕으로 GNU GPL 라이선스의 핵심 원칙 - ‘공개’, ‘책임 공유’, ‘권리 연속성’ - 을 AI 데이터 거버넌스 구조에 적용할 수 있는 가능성을 탐색하였다. 결론적으로 본 연구는 기술·법·윤리 통합적 관점에서 새로운 데이터 규범 설계 필요성을 제시하며, AI 생태계의 투명성과 책임성, 디지털 시민사회의 정보주권 강화를 위한 기반을 제공하고자 한다.
The rapid expansion of generative AI has significantly increased the demand for large-scale, high-quality training data, raising critical legal and ethical concerns regarding the use of personal information. Telecommunication-based AI services, such as SK Telecom’s “A.” assistant, have structural access to sensitive data, including call logs and voice content. This study examines how such data is utilized for AI training, highlighting challenges related to user control, third-party rights, and algorithmic transparency. Through a case analysis of A., and a comparison with the Stack Overflow incident, this study highlights how even publicly available datasets can cause harm when proper consent, attribution, and legal compliance are absent. Existing legal frameworks, such as Korea’s Personal Information Protection Act (PIPA), are found to be inadequate in addressing AI-specific risks, particularly concerning high-risk data types. As a normative response, this paper explores the applicability of governance principles derived from the GNU General Public License (GPL), including openness, shared responsibility, and continuity of rights. The findings indicate a need for hybrid governance models that integrate legal, technical, and ethical mechanisms to ensure transparency, accountability, and data sovereignty in the era of artificial intelligence(AI).
생성형 AI 학습데이터로서 박물관 소장품 데이터의 신뢰성 연구: 이원적 분석틀(L․Q)을 통한 e뮤지엄 진단
[NRF 연계] 한국박물관학회 박물관학보 Vol.51 2026.06 pp.57-86
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 생성형 AI 시대 박물관 소장품 데이터의 신뢰성을 진단하기 위해 L․Q 이원적 분석틀을 제시하였다. 이는 법적 개방성(Legal Openness, L)과 데이터 품질(Data Quality, Q)의 동시 충족 여부로 신뢰성을 판단하는 틀로, 국내 대표 디지털 문화유산 플랫폼인 e뮤지엄에 실증적으로 적용하였다. 분석에 앞서 스미스소니언과 유로피아나의 L․Q 구현 방식을 비교 준거로 검토하였으며, 문화공공데이터광장(KCISA)이 제공하는 e뮤지엄 국립중앙박물관 유물정보 Open API에서 소장품 메타데이터 100건을 무작위 추출하여 항목별 기재율을 탐색적으로 분석하였다. 분석 결과, 유물명 등 기초 식별 정보의 기재율은 100%로 안정적인 반면, 관련 인물(person)․참조자원(sourceTitle)은 API에서 제공되지 않았고(0%), 권리정보(rights)는 API에서 4%만 기재되었으며, 최종수정일(lastModifyDate) 항목은 API 제공 목록 자체에 존재하지 않았다. 또한 e뮤지엄 내 동일 유형 유물이 기관마다 상이한 명칭으로 등록되어 AI가 유사 자료를 일관되게 식별․분류하는 데 구조적 한계가 확인되었다. 이러한 Q 차원의 결핍은 법적으로 개방된 데이터라 하더라도 AI 학습데이터로서의 신뢰성을 담보하기 어렵게 하며, 박물관의 제도적 공신력이 오히려 오류 정보의 확산 가능성을 높이는 ‘권위의 역설(Authority Paradox)’로 이어질 수 있음을 개념화하였다. 이를 바탕으로 Open API의 기계가독형 구조 재설계, CIDOC CRM․통제어휘 기반 의미 구조화, 신 공공누리 AI 유형 적용을 통한 L․Q 병행 정비를 정책적으로 제언한다.
This study proposes an L?Q Dual-Condition Framework for assessing the reliability of museum collection data in the era of generative AI. The framework evaluates reliability by examining whether legal openness (L) and data quality (Q) are simultaneously satisfied, and it is empirically applied to eMuseum, Korea’s leading digital cultural heritage platform. Prior to the analysis, the L?Q implementation models of the Smithsonian Institution and Europeana were reviewed as comparative reference cases. One hundred metadata records were randomly sampled from the eMuseum Open API for National Museum of Korea collection data provided by the Korea Culture Information Service Agency (KCISA), and field completion rates were analyzed in an exploratory manner. The analysis revealed that while basic identifying information, such as object titles, showed a stable completion rate of 100%, several critical fields were deficient: related persons (person) and reference resources (sourceTitle) were not provided through the API (0%), rights information (rights) appeared in only 4% of records, and the last modification date (lastModifyDate) field was absent from the API specification itself. In addition, objects of the same type within eMuseum were registered under different names across institutions, indicating structural limitations that may prevent AI systems from consistently identifying and classifying similar materials. These deficiencies in the Q dimension make it difficult to ensure the reliability of legally open data as generative AI training data. This study conceptualizes the risk that the institutional authority of museums may paradoxically amplify the spread of erroneous information as the “Authority Paradox.” Based on these findings, the study proposes parallel L?Q reforms: redesigning the Open API for machine readability, semantic structuring based on CIDOC CRM and controlled vocabularies, and applying the new KOGL AI type.
AI 학습데이터를 둘러싼 분쟁과 한국법체계의 통합적 적용 - 저작권법, 부정경쟁방지법, 국제사법의 중첩 적용을 중심으로 -
[NRF 연계] 숭실대학교 법학연구소 법학논총 Vol.65 2026.05 pp.35-70
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
지상파 3사가 네이버와 OpenAI를 상대로 각각 제기한 소송은 인공지능 학습데이터를 둘러싼 저작권 분쟁이 한국에서 사실심리의 대상으로 진입하였음을 보여준다. 본 논문은 한국법체계가 이러한 분쟁에 어떠한 통합적 분석틀을 제공하는지를 검토한다. 한국법상 인공지능 학습데이터에 대한 권리자 보호는 저작권법, 부정경쟁방지법, 국제사법이라는 세 보호 층의 중첩 구조로 형성되어 있다. 저작권법은 공정이용(제35조의5)과 데이터베이스 제작자 권리(제93조)를 통해 보호의 일차 통로를 제공하고, 부정경쟁방지법은 데이터 부정사용(카목)과 성과 도용(파목)을 통해 저작권법의 보호 공백을 보완하며, 국제사법은 관할 조항(제39조)과 보호국법주의(제40조)를 통해 외국 인공지능 기업에 대한 한국법 적용의 매개가 된다. 야놀자 v. 여기어때 사건의 대법원 판결을 비롯한 한국 판례 흐름은 데이터 보호의 무게중심이 저작권법상 데이터베이스 제작자 권리에서 부정경쟁방지법으로 이동하는 경향을 보여주며, 이 경향은 인공지능 학습 분쟁에 동일하게 투영될 가능성이 크다. GEMA v. OpenAI 판결은 인공지능 학습 분쟁에 보호국법주의를 본격적으로 적용한 첫 유럽 사례로서, 외국 인공지능 기업에 대한 한국법 적용의 비교법적 근거를 제공한다. 본 논문은 이러한 분석틀을 지상파 3사가 제기한 두 사건에 적용하여 피고가 국내 기업인지 외국 기업인지에 따라 적용 법역과 항변 구조가 어떻게 달라지는지를 비교 평가한다. 이상의 분석을 토대로 본 논문은 텍스트ㆍ데이터마이닝 면책의 도입과 부정경쟁방지법상 권리와의 정합성, 학습데이터 출처 공개 의무의 도입, 보호국법주의 적용 매개의 입법적 명확화에 관한 시사점을 도출한다.
The lawsuits filed by the three Korean broadcasters against Naver and OpenAI, respectively, indicate that copyright disputes over AI training data have entered the stage of substantive adjudication in Korea. This article examines the integrated analytical framework that the Korean legal system provides for such disputes. Under Korean law, the protection of rightholders against the unauthorized use of works as AI training data is structured as the overlapping application of three layers: the Copyright Act, the Unfair Competition Prevention Act, and the Private International Law Act. The Copyright Act offers the primary route of protection through the fair use clause (Article 35-5) and the database producer's right (Article 93); the Unfair Competition Prevention Act supplements the gaps through the prohibition of unauthorized use of data (Article 2(1)(ka)) and misappropriation of results (Article 2(1)(pa)); and the Private International Law Act mediates the application of Korean law to foreign AI companies through the jurisdictional clause covering acts directed at Korea (Article 39) and the principle of lex loci protectionis (Article 40). The trajectory of Korean case law, including the Supreme Court's decision in the Yanolja v. Yeogieottae case, demonstrates a tendency for the center of gravity of data protection to shift from the database producer's right under the Copyright Act to the protective regime of the Unfair Competition Prevention Act, a tendency likely to be projected onto AI training disputes. With respect to the application of Korean law to the training activities of foreign AI companies, the GEMA v. OpenAI decision rendered by the Munich Regional Court I in November 2025, as the first European case to substantively apply lex loci protectionis to AI training disputes, provides comparative grounds for applying Korean law to such activities. Applying this framework to the two cases filed by the Korean broadcasters, the article comparatively evaluates how the scope of applicable layers and the structure of defenses differ when domestic and foreign companies are sued. On the basis of this analysis, the article derives implications for legislative reforms regarding the introduction of a text and data mining exception and its coherence with rights under the Unfair Competition Prevention Act, the introduction of disclosure obligations on training data sources, and the legislative clarification of the mediating criteria for lex loci protectionis.
[NRF 연계] 한국정보법학회 정보법학 Vol.30 No.1 2026.04 pp.168-217
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
생성형 인공지능(AI) 기술의 비약적인 발전은 기존 저작권 법제에 전례 없는 도전을 제기하고 있다. AI 모델 학습을 위한 대규모 데이터 이용 과정에서는 기술 혁신의필요성과 창작자 권리 보호라는 상충하는 가치가 충돌하고 있음에도 불구하고, 현행한국 저작권법은 이를 조정할 명확한 규범적 기준을 충분히 제시하지 못하고 있다. 사법 영역에서는 공정이용 법리에 따른 사후적⋅사안별 판단이 축적되고 있으나 예측가능성이 낮고, 입법 영역에서는 생성형 AI 학습을 전제로 한 체계적 규율이 부재하여 법적 불확실성과 권리 보호의 공백이 지속되고 있다. 이에 본 연구는 AI 학습 데이터 문제를 공정이용의 확장이나 텍스트⋅데이터 마이닝(TDM) 예외 도입 여부라는 이분법적 구도에서 벗어나, 사전 통제권, 사후 보상, 이용 목적 제한, 투명성 의무라는 네 가지 규제 축을 시장 대체 위험과 교섭력 불균형이라는 두 구조적 기준에 따라 차등적으로 조합⋅설계해야 할 입법적 설계 문제로 파악한다. EU, 미국, 일본, 싱가포르, 중국의 규제 모델을 위 네 가지 분석 축으로 유형화하여 비교⋅분석하되, 비교법을 규범의 직접적 이식 대상이 아닌 국내 규율 설계를위한 분석 도구로 활용하였다. 본 연구는 헌법상 비례의 원칙에 기초하여 뉴스, 출판, 시각예술, 음악, 영상, 교육등 주요 산업을 개별적으로 분석하고, 시장 대체 위험이 높고 교섭력 불균형이 심각한 영역에는 보다 강한 사전 통제와 집단적 보상 논의가 필요할 수 있는 반면, 공익성이 높고 시장 대체 위험이 낮은 영역에는 보다 제한적이고 신중한 예외 설계가 가능하다는 점을 논증한다. 나아가 각 산업별 규제 방향을 구현하기 위한 저작권법 개정및 조문 설계의 기본 방향을 제시함으로써, 향후 한국형 AI 저작권 규율 논의를 위한입법적⋅정책적 판단 기준을 제시하는 데 목적이 있다.
The rapid advancement of generative artificial intelligence (AI) poses unprecedented challenges to existing copyright regimes. In the process of training AI models on large-scale datasets, the need for technological innovation increasingly conflicts with the protection of creators’ rights. Despite this structural tension, Korean copyright law still lacks sufficiently clear normative standards for regulating AI training, resulting in persistent legal uncertainty and gaps in the protection of rights. Judicial responses relying on ex post, case-specific fair use analysis remain fragmented and unpredictable, while legislative frameworks specifically tailored to generative AI training have yet to be fully developed. This Article reframes the problem of AI training data not as a binary choice between expanding fair use and introducing a text-and-data-mining (TDM) exception, but as a question of regulatory design: how four key regulatory instruments?prior control, ex post remuneration, purpose-based limitations, and transparency obligations?should be differentially combined according to two structural criteria, namely market substitution risk and bargaining power asymmetry. It conducts a comparative analysis of regulatory approaches in the European Union, the United States, Japan, Singapore, and China, treating foreign models not as templates for direct transplantation but as analytical tools for designing a Korean regulatory framework. Grounded in the constitutional principle of proportionality, the Article examines major sectors?including news media, publishing, visual arts, music, audiovisual works, and education?on a sector-specific basis. It argues that sectors characterized by high market substitution risk and severe bargaining asymmetry may justify stronger forms of prior control and collective remuneration, whereas sectors marked by stronger public interest and lower substitution risk may justify more limited and carefully tailored exceptions. Rather than presenting a complete and final legislative solution, this Article offers a differentiated framework and a set of drafting directions for future Korean legislative and policy discussions on copyright regulation of generative AI training.
인공지능 학습용데이터 활용을 위한 입법 과제 ― 인공지능기본법과 개인정보보호법을 중심으로 ―
[NRF 연계] 한국공법학회 공법연구 Vol.53 No.4 2025.06 pp.101-123
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문은 인공지능 산업 발전에 필수적인 학습용데이터의 활용 활성화와 관련하여, 학습용데이터와 개인정보의 관계와 현행 「개인정보 보호법」과 「인공지능기본법」의 한계를 분석함으로써 입법방안을 제시하고자 한다. 특히 학습용데이터에 기존 개인정보보호 체계를 일률적으로 적용하는 데에 근본적인 한계가 존재한다는 것을 지적한다. 아울러 「인공지능기본법」의 법적 성격과 기능을 고찰하고, 실질적인 기본법으로 기능하기 위한 입법적 개선방안을 다음과 같이 제안한다. 첫째, 학습용데이터의 활용에 관한 법적 근거로서 「인공지능기본법」에 학습용데이터 활용의 특례를 신설하는 방안을 제시하고, 장기적 입법과제로서 학습데이터 활용에 관한 개별법을 제정하는 방안도 함께 제안한다. 둘째, 학습용데이터와 관련해 「개인정보 보호법」과의 정합성 확보를 위해 「인공지능기본법」상 다른 법률과의 관계를 개정하고, 학습용데이터의 특수성을 반영한 입법기술적 조율 방안을 제시한다. 또한, 유럽의 인공지능법(AI Act)과 유럽 일반 개인정보보호법(GDPR)의 관계에 관하여 비교법적 관점에서 고찰한다. 이를 통해 궁극적으로는 ‘개인정보 보호’와 ‘인공지능 산업 발전’의 조화를 토대로, 학습용데이터의 활용 활성화를 위한 입법 과제를 실현하고자 한다.
This study aims to propose legislative improvements centered on the Korean AI Act, by analyzing the relationship between AI training data and personal information, as well as identifying limitations in the current regulatory structure of the Personal Information Protection Act(PIPA). In particular, it points out that there are fundamental limitations in uniformly applying the existing personal information protection system to training data. First, it proposes a plan to newly establish special cases for the use of training data in Korean AI Act as a legal basis for the use of training data, and also proposes a plan to enact a separate law on the use of training data as a long-term legislative measure. Second, in order to secure systematic consistency with the PIPA in relation to training data, the relationship with other laws under the Framework Act on Artificial Intelligence is revised, and a legislative and technical coordination plan reflecting the special nature of learning data is proposed. In addition, the relationship between the EU AI Act and the GDPR is examined from a comparative law perspective.
[NRF 연계] 한국IT정책경영학회 한국IT정책경영학회 논문지 Vol.18 No.2 2026.06 pp.4399-4404
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
생성형 AI의 급속한 발전은 인간이 축적해 온 방대한 창작물을 학습데이터로 사용하는 과정에서 저작권 분쟁을 야기하고 있다. 기존 연구들은 공정이용 법리 분석 또는 블록체인·C2PA 기반 기술 시스템을 제안해 왔으나, AI 기업의 자발적 참여를 전제하고, 소급 적용이 불가능하며, 다세대 합성 파이프라인을 통한 저작권 세탁 문제를 해결하지 못한다는 공통 한계를 지닌다. 본 연구는 이 문제를 단순한 저작권 침해 여부가 아니라, 창작자 보상 협상력이 시간의 흐름에 따라 구조적으로 약화되는 시간적 비대칭성의 문제로 재정의한다. 이를 해소하기 위해 법적 의무 확립을 선행 조건으로 하고, 블록체인 기반 온체인 레지스트리를 통해 학습 데이터 귀속(attribution)·AI-FOPT 계보 추적·소급 보상 분배를 기술적으로 구현하는 통합 아키텍처를 제안한다.
The rapid development of generative AI has intensified copyright disputes over the use of vast human-created works as training data. Prior studies have proposed either legal analyses of fair use doctrine or blockchain/C2PA-based technical systems; however, these approaches share three common limitations: they assume voluntary AI-developer participation, are incapable of retroactive application to already-trained models, and fail to address copyright laundering through multi-generation synthetic pipelines. This study reframes the issue not merely as copyright infringement, but as one of temporal asymmetry, in which creators' bargaining power for compensation structurally diminishes over time. To resolve this, we propose an integrated architecture in which legal obligation is established as a prerequisite, and a blockchain-based on-chain registry technically implements training-data attribution, AI-FOPT lineage tracking, and retroactive compensation distribution.
AI 학습용 데이터의 보호에 관한 소고 - 지식재산법상의 보호를 중심으로 -
[NRF 연계] 조선대학교 법학연구원 법학논총 Vol.28 No.1 2021.04 pp.65-102
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
데이터는 AI 기술과 밀접한 관계를 가진다. 데이터가 축적되면 될수록, 데이터가 정확하면 할수록 좋은 분석결과가 나온다. 이러한 분석에는 AI 기술(AI 소프트웨어)이 활용된다. 딥러닝과 머신러닝과 같은 AI 기술의 발전은 데이터를 더욱 가치 있게 만들어 왔다. 데이터는 AI 기술과 함께 거래되는 경우가 많아지게 되었고, 데이터와 AI 기술을 활용한 비즈니스도 더욱 활기를 띠고 있다. 이러한 관점에서 AI 학습용 데이터는 적절히 보호될 필요가 있다.
Data is closely related to AI technology. The more data is accumulated and the more accurate the data is, the better the analysis results come out. AI technology (AI software) is used for this analysis. Advances in AI technologies such as deep learning and machine learning have made data more valuable. Data is often traded with AI technology, and businesses using data and AI technology are also becoming more vibrant. From this point of view, AI learning data needs to be adequately protected. In the case of the domestic data industry, data creation and utilization is evaluated as relatively inadequate. Data required for data construction and utilization (distribution) is insufficient, and industrial and social use is poor due to a closed distribution system. This phenomenon is believed to be due to the limited use of data due to restrictions on personal information, and the lack of manpower responding to corporate demand. In particular, learning data essential for AI-related inventions is no exception. In this study, the following measures were proposed to protect AI learning data. First, a plan to strengthen protection under the patent law, second, a plan to protect data through the Unfair Competition Prevention Act, third, a plan to introduce a 'data patent' application system in preparation for the era of big data, and fourth, Fourth, similar to the microbial donation system, the introduction of the AI system for depositing data and learning completion models was suggested.
생성형 AI 학습 데이터 무단 이용의 위법성 판단과 민법 제750조의 보충적 적용 - 저작권법, 부정경쟁방지법의 검토와 민법 제750조의 보충적 적용을 중심으로 -
[NRF 연계] 한국재산법학회 재산법연구 Vol.43 No.1 2026.02 pp.131-169
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
생성형 인공지능의 발전에 따라 뉴스, 블로그, SNS 등 방대한 데이터가 자동으로 수집・가공되어 모델 학습에 이용되고 있으며, 이 과정에서 데이터 제공자의 동의 여부, 저작권 및 데이터베이스권 침해 가능성, 공정한 경쟁질서 유지 등 다양한 법적 쟁점이 논의되고 있다. 이에 대하여 저작권법, 부정경쟁방지법, 데이터산업법, 인공지능기본법 등은 각기 저작물, 데이터베이스, 한정제공 데이터, 고영향 AI에 대한 규율을 통하여 일정한 보호와 이용 통제의 틀을 마련해 왔다. 그러나 생성형 AI 학습 단계에서 일어나는 대량의 웹 크롤링과 비정형 데이터 축적은 기존 규범이 예정하지 않았던 새로운 이용 형태를 전제로 하고 있어, 현행 특별법 체계만으로 위법성・책임의 범위를 선명하게 그려내기에는 한계가 있다는 지적도 제기된다. 본 논문은 개인정보에 해당하지 않는 비개인 데이터를 중심으로, 생성형 AI 학습 단계에서의 데이터 이용에 대해 이들 특별법이 어떠한 범위에서 보호를 제공하는지를 먼저 검토한 뒤, 그로써도 남게 되는 영역에서 민법 제750조 일반불법행위 규정이 보충적으로 어떤 역할을 수행할 수 있는지를 살펴본다. 특히 저작권법과 부정경쟁방지법의 적용 범위를 분석하여, 특별법이 이미 충분한 보호를 제공하는 영역과 그렇지 못한 영역을 구분하고, 후자의 경우를 중심으로 민법 제750조의 적용 가능성과 위법성 판단 기준을 제시한다. 이 과정에서 데이터 보호를 물권적 권리 부여가 아닌 행위규제 중심으로 설계해 온 우리 법제의 흐름과 EU 데이터베이스・TDM 규정 및 일본 한정제공 데이터 법제를 참고하여, 일반불법행위법리가 이러한 법제 하에서 어떤 보충적 기능을 수행할 수 있는지 이론적으로 정립하고자 한다. 저작권법은 창작성 있는 저작물과 상당한 투자를 통해 구축된 데이터베이스에 대해 배타적 권리를 부여함으로써 창작・투자 인센티브를 보호하는 기능을 수행한다. 부정경쟁방지법 또한 데이터 부정사용행위 규정과 성과 도용 일반조항을 통하여, 업으로서 한정 제공되는 데이터나 경쟁자의 성과에 대한 무임승차를 규율한다. 이와 같이 상당 부분에서 기존 특별법이 이미 데이터 투자와 공정경쟁질서를 실질적으로 보호하고 있다는 점을 전제로 하면서도, 웹상에 공개된 단순 사실 데이터, 창작성이 없는 비정형 데이터, 배열・구조화를 통해 독립된 데이터베이스로 구축되었다고 보기 어려운 자료 등은 여전히 보호 범위 밖에 놓일 여지가 있다. 본 논문은 이러한 잔여 영역에서 데이터 자체에 대한 배타적 소유권이 아니라 데이터의 수집・선별・갱신・관리 과정에 투입된 ‘상당한 투자와 노력의 성과’를 민법 제750조상 법률상 보호할 가치가 있는 이익으로 파악할 수 있는지를 검토한다. 그리고 상관관계설에 기초하여, ① 기술적 보호조치(robots.txt, IP 차단, 캡차 등)의 명시적 제한을 무력화하는 경우, ② 과도한 크롤링으로 시스템 장애를 야기하는 경우, ③ 학습 결과가 원본 콘텐츠의 시장을 실질적으로 대체하는 경우, ④ 데이터 출처・취득 경로・처리 과정에 관한 기록을 전혀 남기지 않아 투명성과 설명가능성을 현저히 결여한 경우 등 일정한 요건 하에서는 일반불법행위 성립 가능성이 인정될 수 있음을 논증한다. 다만 비영리 학술 연구, 데이터 보유자의 명시적・묵시적 승인, 저작권법 제35조의5 공정이용 및 EU・일본의 TDM 면책 규정 취지와 ...
With the rapid development of generative artificial intelligence, massive volumes of data from news, blogs, social media and other sources are automatically collected and processed for model training. In this process, numerous legal issues have been raised, including the requirement for consent from data providers, the risk of copyright and database rights infringement, and the maintenance of fair competition. In response, the Copyright Act, the Unfair Competition Prevention and Trade Secret Protection Act, the Framework Act on the Promotion of Data Industry and Utilization, and the Framework Act on Artificial Intelligence each have established frameworks for protection and use control through regulations addressing, respectively, works of authorship, databases, data provided on a limited basis, and high??impact AI systems. However, large??scale web crawling and the accumulation of unstructured data in the training phase of generative AI presuppose new forms of use that were not contemplated when these statutory provisions were enacted. Accordingly, it has been observed that the existing framework of special statutes alone faces limitations in clearly delineating the scope of unlawfulness and responsibility. Focusing on non-personal data that does not constitute personal information under applicable law, this article first examines the extent to which these special statutes provide protection for data use in the training phase of generative AI, and subsequently considers what supplementary role the general tort liability provision of Article 750 of the Civil Act can perform in addressing the remaining gaps. In particular, by analyzing the scope of application of the Copyright Act and the Unfair Competition Prevention and Trade Secret Protection Act, the article distinguishes between areas where these special laws already provide adequate protection and those where they do not, and, only with respect to the latter category, proposes conditions for the applicability of Article 750 of the Civil Act and a framework for assessing unlawfulness. In doing so, it situates the analysis within the broader legal framework that has designed data protection primarily through conduct-based regulatory approach rather than through the conferral of proprietary rights, and, with reference to EU regulations on databases and text-and-data- mining (TDM) as well as the Japanese “limited data provision” regime, seeks to establish theoretically how the general law of torts can perform supplementary functions within this paradigm. The Copyright Act protects incentives for creativity and investment by conferring exclusive rights in original works of authorship and in databases created through substantial investment. The Unfair Competition Prevention and Trade Secret Protection Act, through provisions addressing unlawful use of data and the general provision against misappropriation of achievements, regulates both the free-riding on data made available on a limited basis as part of a commercial enterprise and the appropriation of a competitor's accumulated achievements. While the article takes as a premise that much of this existing special??law framework already provides substantive protection for data investment and the maintenance of fair competition, it observes that certain categories of data remain outside the scope of protection: simple factual data publicly available on the web, unstructured data lacking in creativity, and materials that are difficult to characterize as constituting an organized database by virtue of their selection or structure. In these residual areas, the article examines whether one can identify, instead of exclusive property rights in data itself, the “fruits of substantial investment and effort” expended in the collection, selection, updating and management of data as constituting an “interest worthy of legal protection” under Article 750 of the Civil Act. Building on the correlational approach to u...
한국 전통문양의 생성형 AI 학습 데이터 구축과 활용 사례 및 전략 연구
[NRF 연계] 한국경영학회 경영학연구 Vol.55 No.1 2026.02 pp.103-128
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 초거대 인공지능(Artificial Intelligence, AI) 시대에 대응하여 한국 전통문양의 디지털 자산화와 문화콘텐츠산업 활용을 위한 생성형 AI 학습 데이터 구축 사례를 분석하고, 그 활용 전략을 제시하는 것을 목적으로 한다. 2024년 국가유산진흥원이 수행한 「한국 전통문양 생성형 AI 학습 데이터 구축 사업」을 탐색적 사례 연구(Exploratory Case Study)로 선정하여, 해당 사례의 데이터 수집, 정제, 가공, 품질관리 등 전 과정을 분석하고, 실무적 성과를 구조화하여구체적 방법론과 활용 가능성을 제시하였다. 특히 전통문양의 형태별․용도별․시대별 분류체계와 한․영 이미지 캡셔닝방식의 적용 사례를 고찰하고, 게임, 메타버스, 디지털 아트, AR/VR 등 문화콘텐츠 산업 분야의 적용 가능성과 공공데이터 개방 및 생태계 조성 전략을 제안하였다. 이러한 분석은 한국 전통문양의 문화적 정체성과 독창성을 보존하면서 디지털문화유산의 활성화 및 산업적 활용 가능성을 제고하는 데 실증적으로 기여할 것이다.
This study examines the construction of generative AI training datasets for the digital assetization and industrial application of traditional Korean patterns within the contemporary era of hyper-scale artificial intelligence (AI). Utilizing the 2024 “Korean Traditional Pattern Data Establishment Project” by the Korea Heritage Agency as an exploratory case study, this research analyzes the end-to-end lifecycle of data development, including collection, refinement, processing, and quality control. By structuralizing these empirical outcomes, the study proposes concrete methodologies and versatile application frameworks for the digital cultural content industry. Specifically, the research evaluates classification systems categorized by morphology, historical utility, and era, alongside the implementation of bilingual (Korean-English) image captioning techniques for multimodal AI training. Furthermore, it explores the integration of these datasets into emerging sectors such as the metaverse, digital art, and Extended Reality (AR/VR/XR), while outlining strategic pathways for public data dissemination and the cultivation of a sustainable data ecosystem. This analysis provides an empirical foundation for promoting the digital revitalization of traditional cultural heritage while ensuring the preservation of its cultural identity and aesthetic authenticity.
AI 학습데이터 관련 저작권 소송 현황과 공정이용에 대한 정량적 위험 분석
[NRF 연계] 강원대학교 비교법학연구소 강원법학 Vol.82 2026.02 pp.39-79
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 생성형 AI 기술의 급격한 발전에 따른 저작권 침해 논란을 법학적 관점과 정량적 데이터 분석 기법을 결합하여 고찰하였다. 특히 미국 저작권법상의 공정이용 법리를 매개로, 파편화된 소송 양상을 체계적으로 구조화하고 개별 사건의 법적 리스크를 객관적으로 산출하는 모델을 제시하고자 하였다. 이를 위해 2025년 9월 기준 미국 내 주요 AI 저작권 소송 50건을 분석 대상으로 선정하였으며, 연구의 정밀도를 높이기 위해 최신 생성형 AI 도구를 방법론에 도입하였다. 소송 데이터 수집, 공정이용 4요소를 세분화한 12개 정량 평가 지표의 수립, 그리고 유클리드 거리 공식을 이용한 다차원 벡터 공간상의 리스크 산출 과정에는 Gemini 2.5 Pro Deep Research를 활용하였다. 또한, 산출된 데이터를 바탕으로 한 K-평균(K-Means) 군집 분석과 주성분 분석(PCA)을 통한 소송 지형도 시각화 과정에서는 Gemini 2.5 Pro Canvas를 활용하여 분석하였다. 분석 결과, AI 소송 사례는 리스크 프로필에 따라 ‘직접 경쟁자(초고위험군)’, ‘콘텐츠 생성자(고위험군)’, ‘LLM 학습 코퍼스(중위험군)’, ‘기능적 도구(저위험군)’의 4개 클러스터로 유형화되었다. ‘직접 경쟁자’ 그룹은 원본 저작물의 시장 대체 효과가 극대화되어 공정이용 주장이 취약한 반면, 코드 생성 AI와 같은 ‘기능적 도구’ 그룹은 변형적 이용 성격이 강해 법적 리스크가 상대적으로 낮은 것으로 나타났다. 본 연구는 추상적인 법리를 계량 가능한 지표로 전환함으로써 AI 개발자에게는 데이터 수집 및 모델 설계 단계에서의 ‘위험 완화 매뉴얼’을, 법률 전문가에게는 소송 전략 수립을 위한 ‘정량적 사례 평가 프레임워크’를 제공한다는 점에서 실무적 의의를 지닌다.
This study examines the escalating legal disputes surrounding copyright infringement in the development of generative AI through a synthesized approach of legal doctrine and quantitative data analysis. Focusing on the "Fair Use" doctrine under U.S. copyright law, this research aims to systematically structure fragmented litigation patterns and propose a model for objectively assessing the legal risks associated with individual cases. To achieve this, fifty AI-related copyright lawsuits in the United States as of September 2025 were selected for analysis. The methodology integrates advanced generative AI tools to enhance analytical precision. Specifically, Gemini 2.5 Pro Deep Research was employed for systematic data collection, the establishment of twelve quantitative evaluation indicators derived from the four factors of fair use, and the calculation of total risk through Euclidean distance within a multi-dimensional vector space. Furthermore, Gemini 2.5 Pro Canvas was utilized to perform K-Means clustering and Principal Component Analysis (PCA) to visualize the "litigation landscape". The results of the analysis classify the lawsuits into four distinct risk profiles: 'Direct Competitors' (Extreme-risk), 'Content Creators' (High-risk), 'LLM Training Corpus' (Moderate-risk), and 'Functional Tools' (Low-risk). The 'Direct Competitors' group exhibits the highest legal vulnerability due to significant market substitution effects, whereas 'Functional Tools', such as code generation AI, demonstrate relatively lower risk due to their transformative nature. This research contributes to the field by translating abstract legal principles into measurable indicators, providing AI developers with a 'risk mitigation manual' and legal professionals with a 'quantitative case evaluation framework' for navigating the complex legal environment of AI technology.
인공지능 학습데이터 자발적 확보를 위한 멀티모달 마이데이터 유통시스템 설계
[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.19 No.5 2024 pp.895-902
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
인공지능의 활용을 위해서는 학습이 필요하고, 학습을 위해서는 데이터가 필요하다. 데이터는 산, 바다, 지형 등 저작권이 없는 것도 있고, 개인정보보호법 및 저작권법 등 각종 법률에 제한받는 데이터도 있다. 본 논문은 정보 주체가 자발적으로 데이터의 수집, 활용, 유통에 동의하고 참여하여, 법률적 제한을 극복하는 방법을 연구한다. 공공장소에 특정한 공간을 만들고, 기업이 참여하여 학습에 필요한 데이터를 정의하고, 국민은 특정한 공간에서 자발적으로 멀티모달 마이데이터의 수집에 참여하여 보상받는 시스템을 설계한다. 또한, 정부에서 운영 중인 마이데이터 플랫폼과 연계해서 생성된 데이터의 인증, 유통, 판매/재판매가 가능한 시스템을 구현한다. 이를 정부 주도로 진행한다면, 학습 도메인별로 학습용 데이터를 법률에 제재받지 않고 새로운 방법으로 데이터의 수집을 할 수 있어, 인공지능 기술 발전과 활용 방안이 더욱 활성화되는 계기가 될 것이다.
AI requires learning, and learning requires data. Some data is copyright-free, such as mountains, oceans, and terrain, while others are restricted by various laws, such as privacy and copyright laws. This thesis investigates how data subjects can voluntarily consent and participate in the collection, utilization, and distribution of their data, overcoming legal restrictions. We design a system that creates specific spaces in public places, engages businesses to define the data needed for learning, and rewards citizens for voluntarily participating in the collection of Multimodal MyData in specific spaces. In addition, a system that enables authentication, distribution, and sale/resale of generated data in connection with the government's MyData platform will be implemented. If this is led by the government, it will be possible to collect data for learning in a new way without legal sanctions for each learning domain, which will further revitalize the development and utilization of AI technology.
인공지능 산업의 데이터 수집과 활용에 대한 韓⋅美 규제 비교 연구
[NRF 연계] 한국질서경제학회 질서경제저널 Vol.28 No.3 2025.09 pp.91-104
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
2022년 11월 OpenAI의 ChatGPT 출시는 인공지능 기술의 혁신적 발전을 상징하며 글로벌 IT 산업에 큰 파장을 일으켰다. ChatGPT는 출시 2개월 만에 월간 활성 사용자(MAU) 1억 명을 돌파하며 기존 플랫폼 대비 압도적인 성장 속도를 보였다. 그러나 한국은 2000년대 초반 IT 강국으로 명성을 얻었음에도 불구하고, 대규모 언어 모델(LLM) 기반 대화형 AI 개발 분야에서 미국에 크게 뒤처져 있다. 이 연구는 한국의 AI 생태계가 ChatGPT와 같은 혁신적 기술을 선도하지 못한 원인을 데이터 수집 단계에서의 규제뿐만 아니라 수집된 데이터의 활용에 대한 규제로 인한 것임을 밝히고 있다. 이에 대한 심층적 분석을 바탕으로 제도적 환경 차원에서 정책적 개선 방안을 제시하는 것을 목표로 한다.
The release of ChatGPT by OpenAI in November 2022 marked a significant milestone in the history of artificial intelligence, making s substantial splash across the global IT industry. Achieving over 100 million monthly active users (MAU) within just two months, ChatGPT demonstrated an unprecedented growth trajectory compared to existing global digital platform companies. However, despite its reputation as an IT powerhouse in the early 2000s, South Korea lags far behind the United States in the development of conversational AI based on large-scale language models (LLM). This study examines the structural and regulatory factors regarding data collection and the use of collected data that have hindered Korea’s capacity to lead new innovated IT industry whish is AI. Based on the depth analysis on this structural and regulatory factors this paper proposes policy improvement measures at the institutional environment level, creating a more conducive environment for the advancement of LLM-based technologies.
Andy Warhol 케이스의 변형적 이용의 해석과 AI 학습데이터의 공정이용
[NRF 연계] 한국경영법률학회 경영법률 Vol.34 No.2 2024.01 pp.87-126
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
현재 미국을 중심으로 저작권자들이 Open AI, Midjourney, Stability AI 등 생성형 AI 개발자 내지 공급자들을 상대로 소송을 제기하고 있다. 현재까지는 피고측에서 공정이용과 관련된 항변은 이루어지지 않는 것으로 보이는데, 이는 저작권자들이 자신의 저작물이 학습데이터로 사용되었다는 것을 입증하기 어려운 것에 기인하는 것으로 보인다. 한국에서 학습데이터 이용에 대한 소송이 제기되는 경우, TDM 예외는 아직 존재하지 않으므로, 피고측이 제기할 수 있는 항변이 바로 공정이용이다. 다만 공정이용의 항변이 이루어지기 위해서는 저작물의 복제 등이 입증되어야 하지만, 이미 저작물을 학습데이터로 사용한 것이 인정된 경우도 존재한다. 따라서 이러한 분야에 대하여 소송이 제기된다면 당장 공정이용 여부가 쟁점이 된다. 이 글은 AI와 관련된 여러 쟁점들 중에서 생성형 AI가 학습데이터를 이용하는 것이 저작권법상의 공정이용에 해당하는지 여부를 분석하는 것을 목적으로 한다. 공정이용 해당 여부는 법원에 의한 판결이 이루어질 때까지는 알 수 없는 것이므로, 이 글은 학습데이터 이용이 공정이용 해당 여부를 판단할 고려요소를 분석하는데 초점을 맞추고자 한다. 특히 지난 2023.5.18. 미국 연방대법원은 공정이용에 관한 판결을 하였는데, 이 글은 연방대법원의 판결을 중심으로 논의하고자 한다. 한국의 공정이용에 관한 규정(§35조의5)은 사실상 미국 저작권법상의 공정이용(§107)을 받아들여 입법한 것이고, 미국 연방대법원이 몇십 년 만에 공정이용에 대하여 판시한 것은 상당히 중요한 의미를 가지고 있고, 공정이용을 판단하는데 가장 중요한 첫 번째 요소인 저작물 이용의 성격과 목적, 특히 변형적 이용과 관련하여 판단하였으며, 공정이용의 나머지 3개 판단요소는 생성형 AI에 대하여 공정이용을 부정하는 방향으로 작용한다는 것을 고려한다면, 연방대법원의 판시는 AI 학습데이터 이용의 공정이용 여부 판단에 대하여 상당한 영향을 미칠 것으로 보인다.
Currently, copyright holders in the United States are suing generative AI developers and providers, including Open AI, Midjourney, and Stability AI. So far, defendants do not seem to raise the fair use defense, which is likely due to the difficulty for copyright holders to prove that their works were used as training data. If a lawsuit is filed in Korea over the use of learning data, fair use is a defense that defendants can raise. While it needs to be shown that the work was reproduced, but there are cases where it has already been recognized that the work was used as learning data. Therefore, if a lawsuit is filed in this area, the issue of fair use becomes an immediate one. This paper aims to analyze whether the use of training data by generative AI constitutes fair use under copyright law. Since it is not possible to know whether it is fair use until the court makes a decision, this paper focuses on analyzing the factors to determine whether the use of training data is fair use. In particular, on May 18, 2023, the U.S. Supreme Court issued a ruling on fair use, and this article focuses on the Supreme Court's ruling. Korea's fair use provision (§35(5)) was actually adopted from the U.S. Copyright Act's fair use (§107), and the Supreme Court decision will have much impact on whether the use of training data is fair use.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.