년 - 년
Design and Implementation of Game 3D Character Creation Framework Using Stable Diffusion KCI 등재
한국컴퓨터게임학회 컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) 제36권 제3호 2023.09 pp.19-25
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 기술의 발전으로 인해 게임 산업은 급격한 성장을 보이고 있습니다. 게임 엔진으로는 유니티, 언리얼 등이 적극적으로 활용되고 있다. 이러한 게임 엔진에서 다양한 플러그인을 제공함으로써 엔진 기능을 자연스럽게 확 장하고 있다. 이 중에서도 3D 그래픽 기술의 혁신은 게임 캐릭터의 표현력을 더욱 높여 사용자의 경험을 향상시 키는 역할을 하고 있다. 그러나 기존의 3D 모델링 접근법은 때로 제약과 어려움을 가져올 수 있다. 특히 AI 스 캔을 통해 일반 사진을 활용할 때 미세한 세부 요소의 표현 문제가 발생하는 경우가 있다. 이에 본 논문에서는 Stable Diffusion을 활용한 게임 3D 캐릭터 생성 Framework를 설계하고 구현하여 이러한 어려움을 극복하고자 한 다. Hair Style가 없는 Concept Art 이미지를 활용하여 제작 프로세스를 개선함으로써, 빠르고 효율적인 캐릭터 제 작을 가능하게 한다. 이 Framework의 적용은 제작 프로세스 단계를 감소시키면서 더욱 신속하고 일관된 캐릭터 제작을 허용하며, 게임 개발자들이 엔진에 캐릭터를 보다 쉽고 빠르게 적용할 수 있게 도와준다. 이를 통해 게임 개발자들이 보다 빠르고 효율적으로 다양한 캐릭터를 엔진에 적용할 수 있게 했다. 또한, 게임엔진에서 애니메이 션 동작을 확인하여 Stable Diffusion으로 생성한 3D 게임 캐릭터의 데이터의 실효성과 확장성을 확인했다.
The gaming industry is experiencing rapid growth due to recent advances in technology. Game engines such as Unity and Unreal are being actively utilized. These game engines offer a variety of plugins to naturally extend the functionality of the engine. Among them, innovations in 3D graphics technology have made game characters more expressive and enhanced the user experience. However, traditional 3D modeling approaches can sometimes bring limitations and challenges. In particular there are problems with the representation of fine details when utilizing general photos through AI scanning. This paper aims to overcome these difficulties by designing and implementing a game 3D character generation framework using Stable Diffusion. By improving the creation process by utilizing concept art images without hairstyles, it enables fast and efficient character creation. The application of this framework allows for faster and more consistent character creation with fewer steps in the creation process and helps game developers to adapt characters to the engine more easily and quickly. In addition, we verified the effectiveness and scalability of the 3D game character data generated by Stable Diffusion by checking the animation behavior in the game engine.
A Study on Effective Unreal Interactive Level Implementation Using Stable Diffusion KCI 등재
한국컴퓨터게임학회 컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) 제36권 제2호 2023.06 pp.89-94
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 언리얼 엔진(Unreal Engine) 기반의 가상환경에서 인공지능 시각화의 대표적인 시스템인 Stable Diffusion을 이용하여 다양한 형태의 인터렉티브 레벨 디자인 구현 방법을 분석하고 제안한다. Stable Diffusion 은 Midjourney와 함께 대표적인 인공지능 시각화 시스템이며 다양한 프로그램 및 툴에 응용되고 있으며 프롬 프트 입력 기반으로 Text to Image, Image to Image는 물론이며 간단한 추가 작업으로 바디 모션 변화까지 자 연스럽게 적용되는 인공지능 시각화 시스템이다. Unreal engine은 현재 5.2까지 업데이트 된 대표적인 게임 엔 진이며 본 연구에서는 5.1을 사용하였다. Unreal Engine은 게임은 물론 애니메이션, 영화CG, 건축, 디지털 트윈 까지 광범위하게 사용되고 있으며 게임 엔진이지만 게임 이외의 콘텐츠 저작 툴로써 점점 더 많이 이용되고 있는 프로그램이다. 본 논문은 이 Stable Diffusion과 Unreal engine을 접목시켜 효과적인 인터랙티브 레벨 구현 을 할 수 있는 프로세스에 관해 분석하고 제안한다. 이미 인공지능은 다양한 2D, 3D콘텐츠에 접목이 시도되 어 실제 퍼블리싱가지 이어지고 있는 시대라고 할 수 있다. 본 연구는 현재 Unreal Engine과 인공지능을 효과 적으로 융합하여 시각적 결과물을 구현하는 방법론을 분석하였으며 효과적으로 구축할 수 있는 토대를 확인 하였다.
This study analyzes and proposes various types of interactive level design implementation methods using Stable Diffusion, a representative system of artificial intelligence visualization in a virtual environment based on Unreal Engine. Stable Diffusion is a representative artificial intelligence visualization system along with Midjourney and is applied to various programs and tools. It is an artificial intelligence visualization system that naturally applies text to image and image to image based on prompt input, as well as body motion changes with simple additional work. . Unreal engine is a representative game engine that has been updated up to 5.2, and 5.1 was used in this study. Unreal Engine is widely used not only for games but also for animation, film CG, architecture, and digital twins. Although it is a game engine, it is a program that is increasingly being used as a content authoring tool other than games. This paper analyzes and proposes a process that can implement an effective interactive level by combining this Stable Diffusion with the Unreal engine. It can be said that artificial intelligence is already being applied to various 2D and 3D contents, leading to actual publishing. This study analyzed the methodology of implementing visual results by effectively integrating the current Unreal Engine and artificial intelligence, and confirmed the foundation for effective construction.
생성형 인공지능 도구 활용 창작 미술 주제 교과 연계 프로그램 개발 및 타당성 검증 KCI 등재
한국정보교육학회 정보교육학회논문지 제27권 제6호 2023.12 pp.655-664
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 교육 현장에서는 ChatGPT가 대표적인 생성형 인공지능에 대한 관심이 뜨겁다. 생성형 인공지능의 교육적 활용에 대한 의견은 찬반으로 분분한 실정이지만, 최근 인간의 고유 영역이라고 생각했던 창작 영역에까지 생성형 인공지능이 활용되고 있는 상황이다. 교육 관련 연구자들은 향후 생성형 인공지능의 교육적 활용 측면에서 가장 중요한 학습자들의 역량으로 생성형 인공지능과 협업할 수 있는 역량이라고 보고 있다. 그러나 아직까지 생성형 AI와 협업 역량을 신장할 수 있는 예술교육 관련 콘텐츠는 전무한 실정이다. 이에 이 연구에서는 창작 미술교육 주제로 생성형 AI 도구 중 하나인 스테이블 디퓨전(Stable Diffusion)을 활용하여 학습자가 생성형 AI 도구와 협업 하여 창의적인 미술 작품을 만들 수 있는 교과 연계 교수·학습과정안을 개발하였다. 프로그램의 타당성 검증을 위 해서 Lawshe(1975)의 내용타당도 비율(Content Validity Ratio: CVR) 계산 공식을 활용하였다. 검증 결과는 학교 현장에서 활용하기에 적합한 것으로 분석되었다.
Recently, there has been a lot of interest in generative artificial intelligence, of which ChatGPT is a representative example, in the educational field. Opinions on the educational use of generative artificial intelligence are divided into pros and cons, but generative artificial intelligence has recently been used even in creative areas that were thought to be the domain of humans. Education-related researchers believe that the ability to collaborate with generative artificial intelligence is the most important learner competency in terms of educational use of generative artificial intelligence in the future. However, there is still no content related to art education that can enhance generative artificial intelligence and collaboration capabilities. Accordingly, in this study, we used Stable Diffusion, one of the generative artificial intelligence tools, as a topic for creative art education to develop a curriculum-linked teaching and learning curriculum that allows learners to create creative works of art by collaborating with generative artificial intelligence tools. did. For validity verification, Lawshe (1975)'s content validity ratio (CVR) calculation formula was used. The verification results were analyzed as suitable for use in school education.
색상 제어 MLP 어댑터를 적용한 Stable Diffusion 기반 2D 실내 리라이팅 시스템 KCI 등재
한국컴퓨터게임학회 컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) 제39권 제2호 2026.02 pp.69-79
※ 기관로그인 시 무료 이용이 가능합니다.
4,200원
최근 인테리어 디자인 플랫폼들이 AR 및 3D 렌더링 기술을 빠르게 도입하고 있으나, 기 존 온라인 시뮬레이터와 AR 서비스가 제공하는 조명 미리보기 기능은 광학 효과가 반영되 지 않아 실제 공간의 분위기를 체감하는 데 한계가 있다. 본 논문에서는 이러한 현실감의 격차를 해소하기 위해 AI 기반 2D 실내 리라이팅 시스템을 제안한다. 제안한 시스템은 이 미지 전체의 스타일을 변환하는 대신 실제 조명 기구 위치에 광원을 적용하는 조명 기구 기반 리라이팅을 수행하며, 사용자 경험 향상을 위해 다중 색상 조명 변경을 지원한다. Latent-Intrinsics와 확산 모델을 기반으로 램프 객체에 특화되도록 U-Net의 Cross-attention을 미세 조정하고, 안정적인 색상 주입을 위해 학습된 ControlNet과 Color Adaptor를 적용하였다. 실험 결과, 기존 베이스라인 모델 대비 우수한 색상재현력과 RMSE와 LPIPS 성능이 약 2.5배, SSIM은 약 1.2배 향상되었다. 연구 결과는 AR 기반 인테리어 서비 스를 위한 정교한 조명 미리보기 기술로 활용되기를 기대한다.
While interior design platforms are rapidly adopting AR and 3D rendering technologies, the lighting preview features in existing online simulators and AR services often fail to incorporate accurate optical effects, limiting users' ability to experience the actual atmosphere of a space. To bridge this realism gap, this paper proposes an AI-based 2D indoor relighting system. The proposed system performs luminaire-grounded relighting—applying the light source directly to the actual fixture location—and supports multi-color lighting changes to enhance the user experience. Built upon Latent-Intrinsics and diffusion models, we fine-tuned the U-Net cross-attention to specialize in lamp objects and employed a trained ControlNet and Color Adaptor for stable color injection. Experimental results demonstrate superior color reproduction compared to baseline models, with performance improvements of approximately 2.5× in RMSE and LPIPS, and 1.2× in SSIM. We anticipate that these findings will serve as a sophisticated lighting preview solution for future AR-based interior services.
스테이블 디퓨전 (Stable Diffusion)인공지능 생성 이미지를 통한 디자인 접근방법 연구 : 칭저우(Qingzhou)박물관 롱싱사(⿓興寺) 불상 피규어 디자인을 중심으로 KCI 등재후보
한국공공디자인학회 공공디자인연구 Vol. 4 No. 4 2024.12 pp.86-101
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
본 연구는 Stable Diffusion 기술을 활용하여 핸드피규어 디자인을 지원하고, 이미지 생성 AI를 통해 혁신적인 방법을 탐구하여 독창적인 불상 핸드피규어를 제작하는 데 중점을 둔다. 연구는 칭저우(Qingzhou) 롱싱사(龙兴寺) 불상 핸드피규어를 대상으로 하여 문헌 및 사례 연구를 통해 기술 원리와 시각적 창작 특성을 설명하고, 핸드피규어 개념을 분석하며 불상 미학과 핸드피규어 시장의 SWOT을 탐구하여 디자인의 어려움을 제시한다. 동종 조사와 사용자 분석을 바탕으로 요구 사항을 도출하고, 불상 미학 특성을 의미 변환하여 LoRA 모델 훈련 집합을 작성하고 훈련하여 그 학습 효과를 검증한다. 이후 프롬프트와 LoRA 계층 제어를 통해 이미지를 생성하고 그 사용 가능성을 평가한다. 연구 결과에 따르면, 이 핸드피규어 디자인 방법은 미학적 특성을 현대적으로 전환하고, 트렌드 요소를 융합하여 효율성과 효과를 향상시키는 데 기여하며, 핸드피규어 디자인 개발에 새로운 방향을 제시할 것으로 기대된다.
This study focuses on utilizing Stable Diffusion technology to assist hand figurine design, exploring innovative methods through image generation AI to create unique Buddhist statue hand figurines. The research centers on the hand figurine of the Longxing Temple Buddha from the Cheongju Museum, explaining the principles of the technology and the characteristics of visual creation through literature and case studies. It analyzes the concept of hand figurines and explores the aesthetics of Buddhist statues and the SWOT analysis of the hand figurine market, identifying design challenges. Based on similar research and user analysis, the study derives user requirements and transforms the aesthetic characteristics of the Buddhist statue into a LoRA model training set, verifying its learning effectiveness. Subsequently, it generates images using prompts and LoRA hierarchical control, evaluating their usability. The results show that this hand figurine design method contributes to modernizing aesthetic characteristics and integrating trendy elements, enhancing efficiency and effectiveness, and is expected to offer a new direction for hand figurine design development.
생성형 인공지능을 활용한 무형문화유산의 문화 상품 이미지 생성에 관한 연구 - Stable Diffusion와 Midjourney을 이용한 중국 공예품 경태람을 중심으로 KCI 등재후보
한국공공디자인학회 공공디자인연구 Vol. 5 No. 2 2025.06 pp.56-69
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
경태람은 중국의 무형문화유산으로서 풍부한 역사적·문화적 함의와 독창적인 예술적 가치를 지니고 있다. 그러나 현대화의 급속한 진전에 따라 전통 공예의 단절 현상이 심화되고 있으며, 경태람 예술 역시 전승과 혁신이라는 이중적 과제에 직면하고 있다. 본 연구는 Stable Diffusion과 Midjourney 기술을 중심으로 생성형 인공지능(AI)을 활용하여 시각 이미지 생성의 혁신적인 방법을 탐구하고, 경태람의 문화상품 이미지를 설계하고 생성하는 데 목적을 두었다. 문헌 연구, 디자인 실험과 전문가 평가의 방법론을 적용하였으며, 연구 범위는 생성형 AI 도구, 무형문화유산과 문화상품의 개념, 무형문화유산 디지털화의 새로운 경로 탐색 및 디자인 실천을 포함한다. 실험 단계에서는 경태람 관련 지식 그래프를 구축하여 이론적 토대를 마련하고, LoRA 모델을 구성하고 학습시킨 후, 이를 기반으로 Stable Diffusion을 제어하여 경태람 문화상품의 시각 이미지를 생성하였다. 생성된 이미지는 Midjourney를 활용하여 최적화하였다. 실험 결과, Stable Diffusion과 Midjourney는 창의성과 예술성을 갖춘 문화상품 시각 이미지를 효과적으로 생성하였으며, 이는 경태람에 새로운 시대적 활력과 창조적 가능성을 부여하고, 무형문화유산의 디지털 전승 및 혁신에 기여할 수 있음을 입증하였다.
As an intangible cultural heritage of China, Cloisonne embodies rich historical and cultural significance alongside unique artistic value. However, rapid modernization has led to a growing disconnection in traditional crafts, posing challenges for both preservation and innovation. This study explores innovative visual image generation methods using generative AI technologies—specifically Stable Diffusion and Midjourney—to design and create cultural product images inspired by Cloisonne. Employing literature review, design practice and expert evaluation, the research covers generative AI tools, concepts of intangible cultural heritage and cultural goods, and new approaches to digitalization and design practice. Experimentally, a Cloisonne knowledge graph was constructed as a theoretical basis, followed by building and training a LoRA model. Stable Diffusion was then used to generate visual images of Cloisonne cultural products, which were further refined with Midjourney. Results demonstrate that these AI tools effectively produce creative and artistic cultural product visuals, revitalizing Cloisonne with fresh vitality and innovation, while advancing the digital preservation and modernization of intangible cultural heritage.
AI 기반 제품 배경 생성 : Stable Diffusion과 Image Segmentation의 기술 결합
한국경영정보학회 한국경영정보학회 정기 학술대회 지속 가능한 미래를 위한 디지털 기술의 통합과 혁신 2024.05 pp.63-69
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Urban Landscape Game Scene Sketch Generation Framework with Stable Diffusion KCI 등재
한국컴퓨터게임학회 컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) 제36권 제4호 2023.12 pp.101-105
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구에서는Stable Diffusion 프레임워크를 활용하여 게임 스타일의 스케치, 특히 도시 장면을 생성하는 방법 을 소개한다. 확산 기반의 모델인Stable Diffusion은 쉬운 접근성과 뛰어난 성능으로 많은 연구자와 일반인들에 게 선호되며, 텍스트-스케치, 이미지-스케치의 생성이 가능하다. Stable Diffusion의 몇 가지 문제는 이미지의 국 소성 보존 문제 및 미세 조정인데, 이를ControlNet과DreamBooth를 사용하여 해결한다. 결과적으로, 본 연구를 통 해 게임 제작에 사용될 수 있는 텍스트-스케치, 이미지-스케치 생성이 가능하며, 더 나아가 아티스트를 돕는 툴 로도 활용될 수 있다.
Games are a vital part of our culture, leading the generation of tools such as Adobe Photoshop and Unity for game developers. Although sketches are a fundamental form that can be stylized in various ways, there is a scarcity of tools capable of generating images into sketch. To bridge this gap in the artistic sketch domain of gaming and the field of deep generative models, we propose a multimodal sketch generation framework with Stable Diffusion, focusing especially on urban scenes. Stable Diffusion, a model within the diffusion-based category, has gained notable attention in the open-source community and is user-friendly. Thus, we have chosen to utilize Stable Diffusion in our approach. This model processes input prompts and images through the CLIP encoder and effectively generates images. However, Stable Diffusion faces challenges such as a losing locality of input image and difficulties in fine-tuning. To overcome these issues, we incorporate ControlNet and DreamBooth into our framework. We conclude with a demonstration of promising results, urban landscape sketch, in both text-to-sketch and image-to-sketch generation.
자전거도로 영상 데이터 합성을 위한 스테이블 디퓨전 모델 미세 조정 기법 KCI 등재
한국ITS학회 한국ITS학회논문지 제24권 제6호 통권122호 2025.12 pp.79-93
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
자전거 이용 증가와 함께 자전거도로의 사고 예방 및 위험 상황에 대한 영상 기반 모니터링 의 중요성이 더욱 부각되고 있다. 그러나 계절·조도·기상 변화가 충분히 반영된 영상 데이터를 확보하는 데에는 많은 시간이 소요되고, 설상 가상으로 라벨링 비용 또한 높아 객체 탐지 모델 개발에 제약이 발생한다. 본 연구는 이러한 문제를 해결하기 위해 구조적 제약과 스타일 적응 을 동시에 반영하는 영상 합성 기법을 제안하였다. 제안한 방법은 Stable Diffusion 기반 모델에 ControlNet과 Low-Rank Adaptation을 결합하여, 마스크 영상을 통한 구조 제어와 스타일 미세조 정을 통합적으로 수행한다. 실제 CCTV 영상을 기반으로 데이터세트를 구축하고, 세 가지 Stable Diffusion 계열 기저 모델을 대상으로 합성 성능을 비교하였다. 성능 평가는 Fréchet Inception Distance와 CLIP-score를 활용하였으며, 그 결과 제안한 방법이 사실성과 텍스트 정합 성 측면에서 우수한 합성 품질을 달성함을 확인하였다. 또한 텍스트 프롬프트 조작만으로 계 절 및 기상 조건을 반영한 영상 생성이 가능함을 검증하였다. 본 연구는 촬영이 어려운 다양한 환경 조건의 데이터를 효율적으로 생성할 수 있어 자전거도로 모니터링을 위한 데이터 부족 문제 해결에 기여하며, 향후 객체 탐지 및 안전관리 기술의 고도화에 효과적으로 활용될 수 있다.
The growing use of bicycles has heightened the importance of video-based monitoring for preventing accidents and detecting hazardous situations on bicycle roads. However, collecting video data that adequately reflects variations in season, illumination, and weather requires substantial time, and the high cost of data labeling further limits the development of effective object-detection models. To address these challenges, this study proposes an image synthesis method that simultaneously incorporates structural constraints and style adaptation. The proposed approach integrates Stable Diffusion with ControlNet and Low-Rank Adaptation (LoRA), enabling unified control of scene structure through mask images and fine-grained style adjustment. A dataset was constructed using real CCTV footage, and three Stable Diffusion–based backbone models were evaluated for their synthesis performance. Fréchet Inception Distance and CLIP-score were used for quantitative assessment, demonstrating that the proposed method achieves superior realism and semantic alignment between images and text. Furthermore, the model successfully generated images reflecting seasonal and weather variations solely through prompt manipulation. This research provides an efficient solution for generating diverse environmental conditions that are difficult to capture in practice, thereby alleviating data scarcity in bicycle-road monitoring and supporting the advancement of nextgeneration object-detection and safety-management technologies.
A Study on AI Softwear [Stable Diffusion] ControlNet plug-in Usabilities
국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.15 No.4 2023.12 pp.166-171
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
With significant advancements in the field of artificial intelligence, many novel algorithms and technologies have emerged. Currently, AI painting can generate high-quality images based on textual descriptions. However, it is often challenging to control details when generating images, even with complex textual inputs. Therefore, there is a need to implement additional control mechanisms beyond textual descriptions. Based on ControlNet, this passage describes a combined utilization of various local controls (such as edge maps and depth maps) and global control within a single model. It provides a comprehensive exposition of the fundamental concepts of ControlNet, elucidating its theoretical foundation and relevant technological features. Furthermore, combining methods and applications, understanding the technical characteristics involves analyzing distinct advantages and image differences. This further explores insights into the development of image generation patterns.
국제문화기술진흥원 International Journal of Advanced Culture Technology(IJACT) Volume 11 Number 2 2023.06 pp.403-408
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The widespread discussion of AI-generated content, fueled by the emergence of consumer applications like ChatGPT and Midjourney, has attracted significant attention. Among various AI applications, AI painting has gained popularity due to its mature technology, user-friendly nature, and excellent output quality, resulting in a rapid growth in user numbers. Midjourney and Stable Diffusion are two of the most widely used AI painting tools by users. In this study, the author adopts a perspective that represents the general public and utilizes case studies and comparative analysis to summarize the distinctive features and differences between Midjourney and Stable Diffusion in the context of AI character illustration. The aim is to provide informative material for those interested in AI painting and lay a solid foundation for further in-depth research on AI-generated content. The research findings indicate that both software can generate excellent character images but with distinct features.
A Research on Aesthetic Aspects of Checkpoint Models in [Stable Diffusion]
국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 13 Number 2 2024.06 pp.130-135
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
The Stable diffsuion AI tool is popular among designers because of its flexible and powerful image generation capabilities. However, due to the diversity of its AI models, it needs to spend a lot of time testing different AI models in the face of different design plans, so choosing a suitable general AI model has become a big problem at present. In this paper, by comparing the AI images generated by two different Stable diffsuion models, the advantages and disadvantages of each model are analyzed from the aspects of the matching degree of the AI image and the prompt, the color composition and light composition of the image, and the general AI model that the generated AI image has an aesthetic sense is analyzed, and the designer does not need to take cumbersome steps. A satisfactory AI image can be obtained. The results show that Playground V2.5 model can be used as a general AI model, which has both aesthetic and design sense in various style design requirements. As a result, content designers can focus more on creative content development, and expect more groundbreaking technologies to merge generative AI with content design.
Research on AI Painting Generation Technology Based on the [Stable Diffusion]
국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 12 Number 2 2023.06 pp.90-95
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
With the rapid development of deep learning and artificial intelligence, generative models have achieved remarkable success in the field of image generation. By combining the stable diffusion method with Web UI technology, a novel solution is provided for the application of AI painting generation. The application prospects of this technology are very broad and can be applied to multiple fields, such as digital art, concept design, game development, and more. Furthermore, the platform based on Web UI facilitates user operations, making the technology more easily applicable to practical scenarios. This paper introduces the basic principles of Stable Diffusion Web UI technology. This technique utilizes the stability of diffusion processes to improve the output quality of generative models. By gradually introducing noise during the generation process, the model can generate smoother and more coherent images. Additionally, the analysis of different model types and applications within Stable Diffusion Web UI provides creators with a more comprehensive understanding, offering valuable insights for fields such as artistic creation and design.
ProGen: 점진적 커리큘럼 학습 기반 문화 간 스타일 융합 생성
국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.12 No.3 2026.05 pp.741-753
※ 원문이용 방식은 연계기관[Kisti]의 정책을 따르고 있습니다.
텍스트-투-이미지 생성 분야에서 의미적으로 거리가 먼 두 미적 영역을 하나의 이미지로 통합하는 작업은 여전히 해결하기 어려운 과제로 남아 있다. 기존의 LoRA 기반 접근법은 대부분 고정된 데이터셋을 활용한 단일 단계 학습에 의존하기 때문에, 한국 전통 건축 양식인 한옥과 서양의 스테인드글라스처럼 구조적으로 이질적인 두 전통을 결합할 때 시각적으로 파편화된 결과물을 생성하는 한계를 보인다. 본 연구에서는 이러한 문제를 해결하기 위해 점진적 파인튜닝 프레임워크인 ProGen을 제안한다. ProGen은 커리큘럼 학습 원리를 바탕으로 학습 데이터를 단계적으로 확장해 나가는 방식을 채택하며, 누적 LoRA 가중치 통합을 포함하는 4단계 학습 구조로 설계되었다. 초기 단계에서는 건축 구조에 대한 이해를 형성하고, 이후 단계에서는 스타일적 복잡성을 점진적으로 심화시킨다. 10개의 수작업 예시 이미지를 출발점으로 삼아 모델 생성과 연구자의 선별 과정을 반복함으로써, 최종적으로 40개의 학습 데이터셋을 구성한다. 실험 결과, ProGen은 소규모 데이터 기반의 단일 단계 파인튜닝 방식과 확장된 데이터셋을 활용한 단일 단계 학습 방식 모두를 성능 면에서 상회하였다. 특히 제거 실험(ablation study)을 통해 커리큘럼 기반의 점진적 학습 구조가 단순한 데이터 양 증가보다 성능 향상에 더 결정적인 역할을 한다는 사실을 입증하였다. IS, LPIPS, CLIP 기반의 다중 평가를 통해 점진적 데이터 확장이 스타일 융합 이미지의 다양성과 표현력을 효과적으로 향상시킴을 검증하였다.
Cross-cultural style fusion in text-to-image generation faces significant challenges when integrating semantically distant aesthetic domains. Existing LoRA-based approaches rely on one-shot training with static datasets, producing fragmented results when applied to structurally incompatible traditions such as Korean Hanok architecture and Western stained glass. We introduce ProGen, a progressive fine-tuning framework that expands training datasets through curriculum learning. The method employs four-stage training with cumulative LoRA weight integration, where early stages establish architectural understanding before later stages introduce stylistic complexity. Starting from 10 manually created exemplars, the dataset grows to 40 through iterative model generation and human-guided selection. Experimental validation demonstrates that ProGen substantially outperforms both single-stage fine-tuning with limited data and single-stage training on expanded datasets. Critically, ablation studies reveal that curriculum structure-not just data quantity-drives performance. Multi-metric evaluation (IS, LPIPS, CLIP) provides convergent evidence that curriculum-based progressive expansion enables superior exploration of fusion variations.
Stable Diffusion을 활용한 철근 및 배근 영상의 암맹 초해상화 성능 향상 방안
[Kisti 연계] 한국구조물진단유지관리공학회 한국구조물진단유지관리공학회 논문집 Vol.29 No.5 2025 pp.29-39
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 건설 현장에서 촬영되는 철근 및 배근 영상은 낮은 해상도, 촬영 흔들림에 따른 블러나 압축 손실과 같은 다양한 열화 요인뿐만 아니라, 촬영 환경에서 발생하는 무작위 잡음까지 포함하는 경우가 많다. 이러한 요인들은 영상의 구조적 정보 손실을 유발하여, 이를 기반으로 한 인공지능 분석의 정확도와 신뢰성을 저하시킨다. 특히 건설 현장의 특성상 고품질 참조 영상을 충분히 확보하기 어려워, 초해상화 알고리즘 학습에 필요한 훈련 데이터가 부족하다는 문제가 존재한다. 이러한 한계를 해결하기 위하여 본 연구에서는 Stable Diffusion X4 모델을 활용하여 저해상도 영상으로부터 선명하고 고품질의 합성 영상을 생성하고, 이를 철근 및 배근 영상과 함께 학습 데이터로 사용하는 새로운 암맹초해상화 기법을 제안하였다. 제안된 방법은 저해상도 영상과 합성된 고해상도 영상을 동시에 입력으로 사용하여 초해상화 알고리즘을 개발하는 것이다. 본 방법을 최신 초해상화 신경망 모델 3건에 적용하여 실험을 수행한 결과, 합성 영상 없이 실제 저해상도-고해상도 쌍만을 이용한 기존 학습 방식 대비 평균 0.45 dB의 해상도 지표 향상과 0.85%의 구조 유사성 지표 향상을 달성하였으며, 철근 및 배근의 형태와 세부 구조를 보다 정밀하게 복원하는 성능을 확인하였다. 본 연구는 영상 품질이 저하된 건설 현장 환경에서도 합성 데이터를 효과적으로 활용하여 초해상화 성능을 개선할 수 있음을 입증하였으며, 향후 철근 및 배근 상태 분석의 신뢰성 강화와 영상 기반 구조물 관리 자동화 기술의 고도화에 기여할 수 있을 것으로 기대된다.
In recent construction sites, images of rebar and reinforcement often suffer from various degradation factors, such as low resolution, motion blur, and compression loss, as well as random noise introduced by the surrounding environment. These factors lead to the loss of structural information in the images, thereby reducing the accuracy and reliability of artificial intelligence-based analysis. In particular, due to the nature of construction sites, it is difficult to secure sufficient high-quality reference images, resulting in a lack of training data necessary for super-resolution algorithm development. To address this limitation, this study proposes a novel blind super-resolution method that utilizes the Stable Diffusion X4 model to generate sharp, high-quality synthetic images from low-resolution images, which are then used together with real rebar images as training data. The proposed method develops a super-resolution algorithm by simultaneously using low-resolution images and synthesized high-resolution images as inputs. Experiments were conducted by applying the method to three state-of-the-art super-resolution neural network models. As a result, experiments conducted by applying the proposed method to three state-of-the-art super-resolution neural network models demonstrated that, compared to the conventional training approach using only real low- and high-resolution image pairs without synthetic data, the proposed strategy achieved an average improvement of 0.45 dB in resolution metrics and 0.85% in structural similarity metrics, while more accurately restoring the shapes and fine structural details of rebar and reinforcement. This study demonstrates that even in construction site environments where image quality is degraded, synthetic data can be effectively utilized to improve super-resolution performance. Furthermore, it is expected to contribute to enhancing the reliability of rebar condition analysis and advancing automation technologies for image-based structural management in the future.
Stable diffusion의 기저 모델에 따른 콘크리트 손상 영상의 생성 품질 비교 연구
[Kisti 연계] 한국구조물진단유지관리공학회 한국구조물진단유지관리공학회 논문집 Vol.28 No.4 2024 pp.55-61
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 들어 노후화된 콘크리트 구조물의 비중이 점차 늘어나는 추세다. 이는 대다수의 구조물이 기대수명에 근접하고 있기 때문이다. 이 같은 구조물은 정확한 점검과 지속적인 관리가 필수적으로 요구되며, 철저한 점검이 이루어지지 않을 경우 본래의 기능과 성능이 저하되어 안전사고로 이어질 수 있음은 자명한 사실이다. 따라서 딥러닝과 컴퓨터 비전을 이용한 객관적인 점검 기술에 대한 연구가 활발하기 이뤄지고 있다. 특히 고해상도는 미세한 균열뿐만 아니라 박락과 철근 노출까지 정확하게 관찰할 수 있으며, 딥러닝을 통해서 자동화 탐지가 가능하다는 장점이 있다. 딥러닝은 다양하고 다수의 훈련 데이터가 있어야지만 높은 탐지 성능을 보장할 수 있지만, 콘크리트의 표면 손상은 비정상 장면으로 일반적으로 촬영하여 확보할 수 있는 데이터가 아니므로 훈련 데이터의 수는 부족할 수밖에 없다. 이러한 한계를 극복하기 위해서 이 연구에서는 stable diffusion을 통해 균열, 박락, 철근 노출을 포함하고 있는 콘크리트 표면 손상 영상을 생성하는 방법을 제안했다. 이는 문자열과 영상이 쌍을 이룬 데이터로 새로운 손상 영상을 합성하는 방법이다. 이를 위해서 총 678장의 훈련 데이터 세트를 구축했고, low rank adaptation을 통해서 fine-tuning을 수행했다. 이때 stable diffusion의 세 가지 기저 모델에 따른 생성 영상의 품질을 비교했다. 결과적으로 가장 다양하고 고품질의 콘크리트 손상 영상을 합성하는 방법을 완성했다. 이 연구는 향후 데이터 부족 문제 해결에 기여하여 딥러닝 기반 손상 탐지 알고리즘의 정확도 향상에 긍정적인 영향을 미칠 것으로 기대한다.
Recently, the number of aging concrete structures is steadily increasing. This is because many of these structures are reaching their expected lifespan. Such structures require accurate inspections and persistent maintenance. Otherwise, their original functions and performance may degrade, potentially leading to safety accidents. Therefore, research on objective inspection technologies using deep learning and computer vision is actively being conducted. High-resolution images can accurately observe not only micro cracks but also spalling and exposed rebar, and deep learning enables automated detection. High detection performance in deep learning is only guaranteed with diverse and numerous training datasets. However, surface damage to concrete is not commonly captured in images, resulting in a lack of training data. To overcome this limitation, this study proposed a method for generating concrete surface damage images, including cracks, spalling, and exposed rebar, using stable diffusion. This method synthesizes new damage images by paired text and image data. For this purpose, a training dataset of 678 images was secured, and fine-tuning was performed through low-rank adaptation. The quality of the generated images was compared according to three base models of stable diffusion. As a result, a method to synthesize the most diverse and high-quality concrete damage images was developed. This research is expected to address the issue of data scarcity and contribute to improving the accuracy of deep learning-based damage detection algorithms in the future.
가상 면접관 생성을 위한 Stable Diffusion 기반의 프롬프트 엔지니어링
[Kisti 연계] 한국컴퓨터정보학회 한국컴퓨터정보학회 학술대회논문집 2023 pp.23-24
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
가상 면접은 현대 사회에서 필수적인 기술이지만, 상호작용의 부족으로 인해 한계가 존재한다. 현실적이고 현장감 있는 가상 면접을 구현하기 위해서는 면접관을 자동으로 생성하여 다양한 상황에서의 면접을 진행할 필요가 있다. 본 논문은 Stable Diffusion 기반의 프롬프트 엔지니어링을 통해 가상 면접관 생성에 대한 연구 결과를 제시한다. 프롬프트 엔지니어링은 Stable Diffusion 모델이 생성하는 결과의 품질을 향상시킬 수 있으며 다양한 조건에 따른 실험 결과를 제시한다.
Stable Diffusion 모형의 객체 개수 불일치 개선을 위한 프롬프트 엔지니어링
[NRF 연계] 한국자료분석학회 Journal of The Korean Data Analysis Society Vol.28 No.2 2026.04 pp.575-587
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Stable Diffusion은 텍스트 프롬프트에 명시된 객체 수량을 정확히 재현하지 못해 지속적인 객체 개수 불일치(object count mismatch) 문제가 발생하는 경우가 많다. 본 연구는 모델을 수정하지 않고 프롬프트 엔지니어링만으로 이러한 오류를 완화할 수 있는지를 분석하였다. 이를 위해 기본적인 수량 명시 프롬프트, 사람이 직접 설계한 구조화 프롬프트, 그리고 LLM 기반 자기 정제(self-refine) 프롬프트의 세 가지 유형을 비교하였다. Text-to-Image 생성은 Stable Diffusion 1.5를 사용하여 수행되었으며, 생성된 이미지는 YOLOv8을 활용해 평가하였다. 수량 재현 성능은 수량 정확도(count accuracy), 평균절대오차(MAE), 과소 생성률(under-rate), 과대 생성률(over-rate) 지표를 통해 측정하였다. 분석 결과, 수량 명시 프롬프트에 비해 구조화 프롬프트는 일부 개선 효과를 보인 반면, 자기 정제 프롬프트는 수량 정확도를 크게 향상시키는 성과를 나타냈다. 이는 LLM을 활용한 프롬프트 정제가 모델이 객체 수량 정보를 보다 명확하게 반영하도록 유도할 수 있음을 시사한다. 다만, 프롬프트만을 활용한 접근이 객체 수량 오류를 완전히 제어할 수는 없어서 근본적인 한계 또한 여전히 존재함을 보여주었다. 따라서 향후 연구에서는 프롬프트 설계뿐 아니라 모델 구조나 학습 방식과 결합한 접근이 필요할 것으로 판단된다.
Stable Diffusion often fails to accurately reproduce the object quantities explicitly specified in text prompts, leading to persistent count mismatch errors. This study investigates whether such errors can be mitigated solely through prompt design, without modifying the underlying model. To this end, we compare three types of prompts: a basic quantitative prompt that explicitly specifies object counts, a manually designed structured prompt, and an LLM-based self-refine prompting strategy. Text-to-Image generation was conducted using Stable Diffusion 1.5, and the generated images were evaluated using YOLOv8. Quantitative reproduction performance was measured using the exact match rate (count accuracy), mean absolute error (MAE), under-generation rate (under-rate), and over-generation rate (over-rate). The results show that, compared to simple quantitative prompts, structured prompts yield modest improvements, while the self-refine prompting strategy significantly enhances the exact match rate. However, the findings also reveal that prompt-only approaches cannot completely eliminate object count errors, indicating the persistence of fundamental limitations.
Stable Diffusion 유해 콘텐츠 생성 억제 연구 동향
[Kisti 연계] 한국정보보호학회 정보보호학회지 Vol.35 No.3 2025 pp.79-96
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 인공지능(AI) 기술의 발전으로 텍스트에서 이미지로 변환하는 생성형 AI 모델의 성능이 획기적으로 향상되었으며, Stable Diffusion과 같은 확산 모델은 예술, 광고, 의료, 교육 등 다양한 분야에서 활용되고 있다. 그러나 이러한 기술이 폭력적이거나 성적으로 부적절한 콘텐츠 생성에 악용될 가능성이 제기되면서, 유해 이미지 생성을 억제하는 기술의 필요성이 더욱 강조되고 있다. 특히, Stable Diffusion 1.x 버전에서는 NSFW(Not Safe For Work) 콘텐츠가 쉽게 생성될 수 있다는 문제가 보고되었으며, 이를 해결하기 위한 다양한 연구가 진행되고 있다. 본 고에서는 Stable Diffusion을 중심으로 생성형 AI 모델의 유해 콘텐츠 억제 기술을 분석하고, 기존 필터링 기법과 그 한계를 고찰한다. 그러나 데이터셋의 한계, 분류기의 성능 부족, 프롬프트 기반 필터링의 취약성과 같은 문제는 여전히 해결해야 할 과제로 남아 있다. 이에 본고에서는 보다 정교한 데이터셋 구축, 고도화된 필터링 모델 개발, AI 윤리 강화를 위한 연구의 필요성을 강조하며, 생성형 AI의 안전성을 확보하기 위한 향후 연구 방향을 제안한다.
[NRF 연계] 한국지식정보기술학회 (사)한국지식정보기술학회논문지 Vol.19 No.3 2024.06 pp.401-412
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The landscape of artificial intelligence (AI) has witnessed a rapid and remarkable evolution, particularly within artistic applications. The integration of AI in artistic endeavors has been a longstanding practice. However, it attained widespread recognition with user-friendly workflows exemplified by Midjourney, DALL-E, and Stable Diffusion. The transformative capabilities of AI have empowered individuals to expeditiously generate artistic paintings, marking a paradigm shift in the temporal dynamics of artistic production. Consequently, this study undertakes a comprehensive comparative analysis of Midjourney, Stable Diffusion, and Krea within the context of storyboard development in animation pre-production. The primary focus centers on elucidating effective methodologies for leveraging these generators in storyboard creation. The analytical framework scrutinizes their distinctive approaches, output quality, user-friendliness, and overall impact on workflow dynamics. Crucially, the findings of this study underscore that these generators, while collectively enhancing efficiency and creative possibilities in storyboard development, each proffers unique advantages: Midjourney excels in the realms of artistic quality and visual allure, Stable Diffusion demonstrates superiority in user control and customization, while Krea distinguishes itself with its noteworthy real-time AI image generation capabilities.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.