Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 31
No
1

4,900원

본 연구는 Stable Diffusion 기술을 활용하여 핸드피규어 디자인을 지원하고, 이미지 생성 AI를 통해 혁신적인 방법을 탐구하여 독창적인 불상 핸드피규어를 제작하는 데 중점을 둔다. 연구는 칭저우(Qingzhou) 롱싱사(龙兴寺) 불상 핸드피규어를 대상으로 하여 문헌 및 사례 연구를 통해 기술 원리와 시각적 창작 특성을 설명하고, 핸드피규어 개념을 분석하며 불상 미학과 핸드피규어 시장의 SWOT을 탐구하여 디자인의 어려움을 제시한다. 동종 조사와 사용자 분석을 바탕으로 요구 사항을 도출하고, 불상 미학 특성을 의미 변환하여 LoRA 모델 훈련 집합을 작성하고 훈련하여 그 학습 효과를 검증한다. 이후 프롬프트와 LoRA 계층 제어를 통해 이미지를 생성하고 그 사용 가능성을 평가한다. 연구 결과에 따르면, 이 핸드피규어 디자인 방법은 미학적 특성을 현대적으로 전환하고, 트렌드 요소를 융합하여 효율성과 효과를 향상시키는 데 기여하며, 핸드피규어 디자인 개발에 새로운 방향을 제시할 것으로 기대된다.

This study focuses on utilizing Stable Diffusion technology to assist hand figurine design, exploring innovative methods through image generation AI to create unique Buddhist statue hand figurines. The research centers on the hand figurine of the Longxing Temple Buddha from the Cheongju Museum, explaining the principles of the technology and the characteristics of visual creation through literature and case studies. It analyzes the concept of hand figurines and explores the aesthetics of Buddhist statues and the SWOT analysis of the hand figurine market, identifying design challenges. Based on similar research and user analysis, the study derives user requirements and transforms the aesthetic characteristics of the Buddhist statue into a LoRA model training set, verifying its learning effectiveness. Subsequently, it generates images using prompts and LoRA hierarchical control, evaluating their usability. The results show that this hand figurine design method contributes to modernizing aesthetic characteristics and integrating trendy elements, enhancing efficiency and effectiveness, and is expected to offer a new direction for hand figurine design development.

2

4,000원

본 연구에서는Stable Diffusion 프레임워크를 활용하여 게임 스타일의 스케치, 특히 도시 장면을 생성하는 방법 을 소개한다. 확산 기반의 모델인Stable Diffusion은 쉬운 접근성과 뛰어난 성능으로 많은 연구자와 일반인들에 게 선호되며, 텍스트-스케치, 이미지-스케치의 생성이 가능하다. Stable Diffusion의 몇 가지 문제는 이미지의 국 소성 보존 문제 및 미세 조정인데, 이를ControlNet과DreamBooth를 사용하여 해결한다. 결과적으로, 본 연구를 통 해 게임 제작에 사용될 수 있는 텍스트-스케치, 이미지-스케치 생성이 가능하며, 더 나아가 아티스트를 돕는 툴 로도 활용될 수 있다.

Games are a vital part of our culture, leading the generation of tools such as Adobe Photoshop and Unity for game developers. Although sketches are a fundamental form that can be stylized in various ways, there is a scarcity of tools capable of generating images into sketch. To bridge this gap in the artistic sketch domain of gaming and the field of deep generative models, we propose a multimodal sketch generation framework with Stable Diffusion, focusing especially on urban scenes. Stable Diffusion, a model within the diffusion-based category, has gained notable attention in the open-source community and is user-friendly. Thus, we have chosen to utilize Stable Diffusion in our approach. This model processes input prompts and images through the CLIP encoder and effectively generates images. However, Stable Diffusion faces challenges such as a losing locality of input image and difficulties in fine-tuning. To overcome these issues, we incorporate ControlNet and DreamBooth into our framework. We conclude with a demonstration of promising results, urban landscape sketch, in both text-to-sketch and image-to-sketch generation.

3

자전거도로 영상 데이터 합성을 위한 스테이블 디퓨전 모델 미세 조정 기법 KCI 등재

심승보, 이유화, 문재필

한국ITS학회 한국ITS학회논문지 제24권 제6호 통권122호 2025.12 pp.79-93

※ 기관로그인 시 무료 이용이 가능합니다.

4,800원

자전거 이용 증가와 함께 자전거도로의 사고 예방 및 위험 상황에 대한 영상 기반 모니터링 의 중요성이 더욱 부각되고 있다. 그러나 계절·조도·기상 변화가 충분히 반영된 영상 데이터를 확보하는 데에는 많은 시간이 소요되고, 설상 가상으로 라벨링 비용 또한 높아 객체 탐지 모델 개발에 제약이 발생한다. 본 연구는 이러한 문제를 해결하기 위해 구조적 제약과 스타일 적응 을 동시에 반영하는 영상 합성 기법을 제안하였다. 제안한 방법은 Stable Diffusion 기반 모델에 ControlNet과 Low-Rank Adaptation을 결합하여, 마스크 영상을 통한 구조 제어와 스타일 미세조 정을 통합적으로 수행한다. 실제 CCTV 영상을 기반으로 데이터세트를 구축하고, 세 가지 Stable Diffusion 계열 기저 모델을 대상으로 합성 성능을 비교하였다. 성능 평가는 Fréchet Inception Distance와 CLIP-score를 활용하였으며, 그 결과 제안한 방법이 사실성과 텍스트 정합 성 측면에서 우수한 합성 품질을 달성함을 확인하였다. 또한 텍스트 프롬프트 조작만으로 계 절 및 기상 조건을 반영한 영상 생성이 가능함을 검증하였다. 본 연구는 촬영이 어려운 다양한 환경 조건의 데이터를 효율적으로 생성할 수 있어 자전거도로 모니터링을 위한 데이터 부족 문제 해결에 기여하며, 향후 객체 탐지 및 안전관리 기술의 고도화에 효과적으로 활용될 수 있다.

The growing use of bicycles has heightened the importance of video-based monitoring for preventing accidents and detecting hazardous situations on bicycle roads. However, collecting video data that adequately reflects variations in season, illumination, and weather requires substantial time, and the high cost of data labeling further limits the development of effective object-detection models. To address these challenges, this study proposes an image synthesis method that simultaneously incorporates structural constraints and style adaptation. The proposed approach integrates Stable Diffusion with ControlNet and Low-Rank Adaptation (LoRA), enabling unified control of scene structure through mask images and fine-grained style adjustment. A dataset was constructed using real CCTV footage, and three Stable Diffusion–based backbone models were evaluated for their synthesis performance. Fréchet Inception Distance and CLIP-score were used for quantitative assessment, demonstrating that the proposed method achieves superior realism and semantic alignment between images and text. Furthermore, the model successfully generated images reflecting seasonal and weather variations solely through prompt manipulation. This research provides an efficient solution for generating diverse environmental conditions that are difficult to capture in practice, thereby alleviating data scarcity in bicycle-road monitoring and supporting the advancement of nextgeneration object-detection and safety-management technologies.

4

본 논문은 Stability AI의 최신 모델인 Stable Diffusion 3 Medium을 활용하여 고서의 이미지를 생성하는 방법 에 대해 제안한다. 고서의 텍스트를 기반으로 한 이미지 생성 과정에서 Stable Diffusion 3 Medium과 OpenAI 의 DALL-E 모델을 비교 평가하였다. 실험에서는 고서의 원문, 한글 번역문, 영어 번역문을 각각 텍스트 입력으로 사용하여 두 모델의 이미지 생성 결과를 분석하였다. Stable Diffusion 3 Medium은 고서의 내용을 보다 정확히 반영한 이미지를 생성하는 데 있어 우수한 성능을 보였다. 특히, 한글 번역문을 사용한 경우, 고서의 내용을 충실히 반영한 이미지를 생성하였다. 반면, DALL-E는 텍스트의 특정 부분을 강조하여 이미지를 생성하였다. Stable Diffusion 3 Medium은 DALL-E에 비해 보다 일반적인 해석을 통해 이미지를 생성하는 경향이 있었으며, 고서의 시각적 세부 사항을 더 잘 재현하였다.

This paper proposes a method for generating images of historical texts using Stability AI’s latest model, Stable Diffusion 3 Medium. The paper compares the image generation capabilities of Stable Diffusion 3 Medium and OpenAI’s DALL-E model based on the text of historical documents. Experiments involved using the original text, Korean translation, and English translation of the historical document as input for both models, analyzing the results. Stable Diffusion 3 Medium demonstrated superior performance in accurately reflecting the content of the historical texts, particularly excelling with the Korean translation. In contrast, DALL-E highlighted specific parts of the text in its images. Stable Diffusion 3 Medium generally provided a more comprehensive interpretation of the text, better reproducing visual details.

5

4,000원

본 연구는 생성이미지와 실사이미지의 분류를 위한 최적의 이미지 전처리 기법을 탐색한다. Stable diffusion과 크롤링을 통해 구축한 데이터셋에 sharpening, grayscale, canny, diffusion, ELA(Error Level Analysis) 총 5가지 전처리 기법을 적용하여 모델별 성능을 비교하고자 한다.

6

실험실에서 발생하는 화재 사고는 인명과 재산 피해 위험이 커 신속한 탐지와 체계적인 대응이 요구된다. 본 논문에서는 실험실에서 발생할 수 있는 다양한 화재 상황에 효율적으로 대처하기 위하여 4종류(일반, 유류, 전기, 금속) 화재 시나리오를 포함하는 이미지 데이터 셋을 구축하고, 객체 검출(YOLOv5) 기반 화재 특징 추출과 트리 기반 의사결정(Random Forest) 모형을 결합한 이미지 기반 실험실 화재 유형 분류 모델을 제시하고 평가한다.

7

4,200원

최근 인테리어 디자인 플랫폼들이 AR 및 3D 렌더링 기술을 빠르게 도입하고 있으나, 기 존 온라인 시뮬레이터와 AR 서비스가 제공하는 조명 미리보기 기능은 광학 효과가 반영되 지 않아 실제 공간의 분위기를 체감하는 데 한계가 있다. 본 논문에서는 이러한 현실감의 격차를 해소하기 위해 AI 기반 2D 실내 리라이팅 시스템을 제안한다. 제안한 시스템은 이 미지 전체의 스타일을 변환하는 대신 실제 조명 기구 위치에 광원을 적용하는 조명 기구 기반 리라이팅을 수행하며, 사용자 경험 향상을 위해 다중 색상 조명 변경을 지원한다. Latent-Intrinsics와 확산 모델을 기반으로 램프 객체에 특화되도록 U-Net의 Cross-attention을 미세 조정하고, 안정적인 색상 주입을 위해 학습된 ControlNet과 Color Adaptor를 적용하였다. 실험 결과, 기존 베이스라인 모델 대비 우수한 색상재현력과 RMSE와 LPIPS 성능이 약 2.5배, SSIM은 약 1.2배 향상되었다. 연구 결과는 AR 기반 인테리어 서비 스를 위한 정교한 조명 미리보기 기술로 활용되기를 기대한다.

While interior design platforms are rapidly adopting AR and 3D rendering technologies, the lighting preview features in existing online simulators and AR services often fail to incorporate accurate optical effects, limiting users' ability to experience the actual atmosphere of a space. To bridge this realism gap, this paper proposes an AI-based 2D indoor relighting system. The proposed system performs luminaire-grounded relighting—applying the light source directly to the actual fixture location—and supports multi-color lighting changes to enhance the user experience. Built upon Latent-Intrinsics and diffusion models, we fine-tuned the U-Net cross-attention to specialize in lamp objects and employed a trained ControlNet and Color Adaptor for stable color injection. Experimental results demonstrate superior color reproduction compared to baseline models, with performance improvements of approximately 2.5× in RMSE and LPIPS, and 1.2× in SSIM. We anticipate that these findings will serve as a sophisticated lighting preview solution for future AR-based interior services.

8

4,000원

최근 기술의 발전으로 인해 게임 산업은 급격한 성장을 보이고 있습니다. 게임 엔진으로는 유니티, 언리얼 등이 적극적으로 활용되고 있다. 이러한 게임 엔진에서 다양한 플러그인을 제공함으로써 엔진 기능을 자연스럽게 확 장하고 있다. 이 중에서도 3D 그래픽 기술의 혁신은 게임 캐릭터의 표현력을 더욱 높여 사용자의 경험을 향상시 키는 역할을 하고 있다. 그러나 기존의 3D 모델링 접근법은 때로 제약과 어려움을 가져올 수 있다. 특히 AI 스 캔을 통해 일반 사진을 활용할 때 미세한 세부 요소의 표현 문제가 발생하는 경우가 있다. 이에 본 논문에서는 Stable Diffusion을 활용한 게임 3D 캐릭터 생성 Framework를 설계하고 구현하여 이러한 어려움을 극복하고자 한 다. Hair Style가 없는 Concept Art 이미지를 활용하여 제작 프로세스를 개선함으로써, 빠르고 효율적인 캐릭터 제 작을 가능하게 한다. 이 Framework의 적용은 제작 프로세스 단계를 감소시키면서 더욱 신속하고 일관된 캐릭터 제작을 허용하며, 게임 개발자들이 엔진에 캐릭터를 보다 쉽고 빠르게 적용할 수 있게 도와준다. 이를 통해 게임 개발자들이 보다 빠르고 효율적으로 다양한 캐릭터를 엔진에 적용할 수 있게 했다. 또한, 게임엔진에서 애니메이 션 동작을 확인하여 Stable Diffusion으로 생성한 3D 게임 캐릭터의 데이터의 실효성과 확장성을 확인했다.

The gaming industry is experiencing rapid growth due to recent advances in technology. Game engines such as Unity and Unreal are being actively utilized. These game engines offer a variety of plugins to naturally extend the functionality of the engine. Among them, innovations in 3D graphics technology have made game characters more expressive and enhanced the user experience. However, traditional 3D modeling approaches can sometimes bring limitations and challenges. In particular there are problems with the representation of fine details when utilizing general photos through AI scanning. This paper aims to overcome these difficulties by designing and implementing a game 3D character generation framework using Stable Diffusion. By improving the creation process by utilizing concept art images without hairstyles, it enables fast and efficient character creation. The application of this framework allows for faster and more consistent character creation with fewer steps in the creation process and helps game developers to adapt characters to the engine more easily and quickly. In addition, we verified the effectiveness and scalability of the 3D game character data generated by Stable Diffusion by checking the animation behavior in the game engine.

9

4,000원

본 논문은 언리얼 엔진(Unreal Engine) 기반의 가상환경에서 인공지능 시각화의 대표적인 시스템인 Stable Diffusion을 이용하여 다양한 형태의 인터렉티브 레벨 디자인 구현 방법을 분석하고 제안한다. Stable Diffusion 은 Midjourney와 함께 대표적인 인공지능 시각화 시스템이며 다양한 프로그램 및 툴에 응용되고 있으며 프롬 프트 입력 기반으로 Text to Image, Image to Image는 물론이며 간단한 추가 작업으로 바디 모션 변화까지 자 연스럽게 적용되는 인공지능 시각화 시스템이다. Unreal engine은 현재 5.2까지 업데이트 된 대표적인 게임 엔 진이며 본 연구에서는 5.1을 사용하였다. Unreal Engine은 게임은 물론 애니메이션, 영화CG, 건축, 디지털 트윈 까지 광범위하게 사용되고 있으며 게임 엔진이지만 게임 이외의 콘텐츠 저작 툴로써 점점 더 많이 이용되고 있는 프로그램이다. 본 논문은 이 Stable Diffusion과 Unreal engine을 접목시켜 효과적인 인터랙티브 레벨 구현 을 할 수 있는 프로세스에 관해 분석하고 제안한다. 이미 인공지능은 다양한 2D, 3D콘텐츠에 접목이 시도되 어 실제 퍼블리싱가지 이어지고 있는 시대라고 할 수 있다. 본 연구는 현재 Unreal Engine과 인공지능을 효과 적으로 융합하여 시각적 결과물을 구현하는 방법론을 분석하였으며 효과적으로 구축할 수 있는 토대를 확인 하였다.

This study analyzes and proposes various types of interactive level design implementation methods using Stable Diffusion, a representative system of artificial intelligence visualization in a virtual environment based on Unreal Engine. Stable Diffusion is a representative artificial intelligence visualization system along with Midjourney and is applied to various programs and tools. It is an artificial intelligence visualization system that naturally applies text to image and image to image based on prompt input, as well as body motion changes with simple additional work. . Unreal engine is a representative game engine that has been updated up to 5.2, and 5.1 was used in this study. Unreal Engine is widely used not only for games but also for animation, film CG, architecture, and digital twins. Although it is a game engine, it is a program that is increasingly being used as a content authoring tool other than games. This paper analyzes and proposes a process that can implement an effective interactive level by combining this Stable Diffusion with the Unreal engine. It can be said that artificial intelligence is already being applied to various 2D and 3D contents, leading to actual publishing. This study analyzed the methodology of implementing visual results by effectively integrating the current Unreal Engine and artificial intelligence, and confirmed the foundation for effective construction.

10

4,600원

경태람은 중국의 무형문화유산으로서 풍부한 역사적·문화적 함의와 독창적인 예술적 가치를 지니고 있다. 그러나 현대화의 급속한 진전에 따라 전통 공예의 단절 현상이 심화되고 있으며, 경태람 예술 역시 전승과 혁신이라는 이중적 과제에 직면하고 있다. 본 연구는 Stable Diffusion과 Midjourney 기술을 중심으로 생성형 인공지능(AI)을 활용하여 시각 이미지 생성의 혁신적인 방법을 탐구하고, 경태람의 문화상품 이미지를 설계하고 생성하는 데 목적을 두었다. 문헌 연구, 디자인 실험과 전문가 평가의 방법론을 적용하였으며, 연구 범위는 생성형 AI 도구, 무형문화유산과 문화상품의 개념, 무형문화유산 디지털화의 새로운 경로 탐색 및 디자인 실천을 포함한다. 실험 단계에서는 경태람 관련 지식 그래프를 구축하여 이론적 토대를 마련하고, LoRA 모델을 구성하고 학습시킨 후, 이를 기반으로 Stable Diffusion을 제어하여 경태람 문화상품의 시각 이미지를 생성하였다. 생성된 이미지는 Midjourney를 활용하여 최적화하였다. 실험 결과, Stable Diffusion과 Midjourney는 창의성과 예술성을 갖춘 문화상품 시각 이미지를 효과적으로 생성하였으며, 이는 경태람에 새로운 시대적 활력과 창조적 가능성을 부여하고, 무형문화유산의 디지털 전승 및 혁신에 기여할 수 있음을 입증하였다.

As an intangible cultural heritage of China, Cloisonne embodies rich historical and cultural significance alongside unique artistic value. However, rapid modernization has led to a growing disconnection in traditional crafts, posing challenges for both preservation and innovation. This study explores innovative visual image generation methods using generative AI technologies—specifically Stable Diffusion and Midjourney—to design and create cultural product images inspired by Cloisonne. Employing literature review, design practice and expert evaluation, the research covers generative AI tools, concepts of intangible cultural heritage and cultural goods, and new approaches to digitalization and design practice. Experimentally, a Cloisonne knowledge graph was constructed as a theoretical basis, followed by building and training a LoRA model. Stable Diffusion was then used to generate visual images of Cloisonne cultural products, which were further refined with Midjourney. Results demonstrate that these AI tools effectively produce creative and artistic cultural product visuals, revitalizing Cloisonne with fresh vitality and innovation, while advancing the digital preservation and modernization of intangible cultural heritage.

12

A Study on AI Softwear [Stable Diffusion] ControlNet plug-in Usabilities

Chenghao Wang, Jeanhun Chung

국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.15 No.4 2023.12 pp.166-171

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

With significant advancements in the field of artificial intelligence, many novel algorithms and technologies have emerged. Currently, AI painting can generate high-quality images based on textual descriptions. However, it is often challenging to control details when generating images, even with complex textual inputs. Therefore, there is a need to implement additional control mechanisms beyond textual descriptions. Based on ControlNet, this passage describes a combined utilization of various local controls (such as edge maps and depth maps) and global control within a single model. It provides a comprehensive exposition of the fundamental concepts of ControlNet, elucidating its theoretical foundation and relevant technological features. Furthermore, combining methods and applications, understanding the technical characteristics involves analyzing distinct advantages and image differences. This further explores insights into the development of image generation patterns.

13

Comparative Analysis of AI Painting Using [Midjourney] and [Stable Diffusion] - A Case Study on Character Drawing - KCI 등재

Pingjian Jie, Xinyi Shan, Jeanhun Chung

국제문화기술진흥원 International Journal of Advanced Culture Technology(IJACT) Volume 11 Number 2 2023.06 pp.403-408

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The widespread discussion of AI-generated content, fueled by the emergence of consumer applications like ChatGPT and Midjourney, has attracted significant attention. Among various AI applications, AI painting has gained popularity due to its mature technology, user-friendly nature, and excellent output quality, resulting in a rapid growth in user numbers. Midjourney and Stable Diffusion are two of the most widely used AI painting tools by users. In this study, the author adopts a perspective that represents the general public and utilizes case studies and comparative analysis to summarize the distinctive features and differences between Midjourney and Stable Diffusion in the context of AI character illustration. The aim is to provide informative material for those interested in AI painting and lay a solid foundation for further in-depth research on AI-generated content. The research findings indicate that both software can generate excellent character images but with distinct features.

14

Research on AI Painting Generation Technology Based on the [Stable Diffusion]

Chenghao Wang, Jeanhun Chung

국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 12 Number 2 2023.06 pp.90-95

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

With the rapid development of deep learning and artificial intelligence, generative models have achieved remarkable success in the field of image generation. By combining the stable diffusion method with Web UI technology, a novel solution is provided for the application of AI painting generation. The application prospects of this technology are very broad and can be applied to multiple fields, such as digital art, concept design, game development, and more. Furthermore, the platform based on Web UI facilitates user operations, making the technology more easily applicable to practical scenarios. This paper introduces the basic principles of Stable Diffusion Web UI technology. This technique utilizes the stability of diffusion processes to improve the output quality of generative models. By gradually introducing noise during the generation process, the model can generate smoother and more coherent images. Additionally, the analysis of different model types and applications within Stable Diffusion Web UI provides creators with a more comprehensive understanding, offering valuable insights for fields such as artistic creation and design.

15

A Study on the Generative AI Reproduction of Traditional Folding Screens : Focusing on Digital Folding Screen

Young-Ho Kim, Jung-Hyun Han, Ye-Won Kwon, Yeon-Woo Kim, Yu-Da Hee

국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 14 Number 4 2025.12 pp.305-312

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

This paper proposes and validates a reproducible generative AI framework for re-presenting and artistically extending Korean traditional folding screens in digital media. A curated set of 70 high-resolution minhwa landscape images is processed with CLIP-based near-duplicate pruning at a cosine threshold of 0.985 and ViT-based tagging with manual refinement. The model is fine-tuned with LoRA on Stable Diffusion v1.5, and a parameter search identifies an effective regime of 25 sampling steps and a CFG scale of 7.0 to 7.5. Subtle image-to-video motion with a fixed camera preserves contemplative aesthetics while enhancing immersion, yielding three exhibited artifacts: mountain, water, and an integrated scene. The study evidences a shift in digital heritage from preservation to generative extension and provides empirical support for LoRA-based style transfer in media art, while noting limitations in dataset size and genre scope and outlining future work in real-time interactivity and VR/AR deployment.

16

Deaf and Hard of Hearing (DHH) children face challenges in language acquisition due to hearing loss, which is one of the most crucial factors in language development. As a result, most DHH individuals communicate using ASL gloss instead of conventional sentences used by hearing individuals. However, because they are more accustomed to ASL gloss, they often struggle with constructing and understanding standard sentences, making communication with hearing individuals difficult. To address this issue, this study proposes DHHStoryGen, an automated storybook creation system that supports the language enhancement of DHH children. DHHStoryGen is built upon a fine-tuned LLaMA model for text generation and a fine-tuned Stable Diffusion model for image generation, both trained on custom datasets specifically designed for this purpose. The text dataset is created by translating ASL gloss into natural English sentences, which is used to fine-tune LLaMA to generate storybook content. Meanwhile, the image dataset is used to fine-tune the Stable Diffusion model to generate illustrations that match with each chapter's narrative. By converting ASL gloss into structured text and providing visually engaging images, DHHStoryGen fosters a participatory language learning environment, serving as an assistant tool to enhance the literacy skills of DHH children.

17

Research on Character's Consistency in AI-Generated Paintings

Chenghao Wang, Jeanhun Chung

국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.16 No.3 2024.08 pp.199-204

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

This study aims to explore the issue of character consistency in AI-generated artwork. First, the concept of character consistency is explained, including the consistency of appearance, actions, and lighting, and its importance in continuous creation and storytelling is analyzed. Next, the study examines current mainstream AI drawing tools such as MidJourney and Stable Diffusion-based WebUI and ComfyUI, evaluating their strengths and limitations in maintaining character consistency. Finally, methods to improve AI drawing technology were proposed to enhance character consistency, aiming to achieve a higher level of consistency in AI art creation.

18

시각예술 창작과 인공지능 협업의 상호작용에 관한 실증연구 KCI 등재

김현진, 김영조, 윤동현, 이한진

국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.10 No.1 2024.01 pp.517-524

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

ChatGPT와 같은 생성형 AI는 21세기의 인간과 기계 간 상호작용에 새로운 패러다임을 제시했다. 이러한 기술의 발전이 다양한 분야에 빠르게 퍼져나가면서, AI와 꽤 멀리 떨어져 있다고 생각되었던 예술 분야에서도 AI가 어떤 역할을 할 수 있는지에 대한 연구가 활발히 진행되고 있다. 이에 본 연구는 제 4차 산업혁명의 시대에 시각예술 교 육에서 생성형 AI의 활용 가능성을 탐구하고자 한다. 경북에 위치한 4년제 대학에서 진행된 실증연구는 창의적 융합 모듈 수업에 참여한 70명의 학생들을 중심으로, AI와 시각예술 분야에서 협업의 영향, 그 중에서도 전공, 학년, 성별에 따른 차이점을 분석했다. 결과적으로, AI와 함께하는 시각예술 창작 활동이 학생들의 창의성과 디지털 미디어 리터러시 에 긍정적인 영향을 미치는 것을 확인하며, 이를 기반으로 더욱 효과적인 교육 전략과 방향 모색에 관해 제언한다.

Generative AI, exemplified by models like ChatGPT, has revolutionized human-machine interactions in the 21st century. As these advancements permeate various sectors, their intersection with the arts is both promising and challenging. Despite the arts' historical resistance to AI replacement, recent developments have sparked active research in AI's role in artistry. This study delves into the potential of AI in visual arts education, highlighting the necessity of swift adaptation amidst the Fourth Industrial Revolution. This research, conducted at a 4-year global higher education institution located in Gyeongbuk, involved 70 participants who took part in a creative convergence module course project. The study aimed to examine the influence of AI collaboration in visual arts, analyzing distinctions across majors, grades, and genders. The results indicate that creative activities with AI positively influence students' creativity and digital media literacy. Based on these findings, there is a need to further develop effective educational strategies and directions that incorporate AI.

19

This paper introduces a real-time face swapping system that enables video influencers to swap their faces with arbitrary generated face images of their choice. The system is implemented as a Django-based server that uses a REST request to communicate with the generative model, specifically the pretrained stable diffusion model. Once generated, the generated image is displayed on the front page so that the influencer can decide whether to use the generated face or not, by clicking on the accept button on the front page. If they choose to use it, both their face and the generated face are sent to the landmark extraction module to extract the landmarks, which are then used to swap the faces. To minimize the fluctuation of landmarks over time that can cause instability or jitter in the output, a temporal filtering step is added. Furthermore, to increase the processing speed the system works on a reduced set of the extracted landmarks.

20

The Stable diffsuion AI tool is popular among designers because of its flexible and powerful image generation capabilities. However, due to the diversity of its AI models, it needs to spend a lot of time testing different AI models in the face of different design plans, so choosing a suitable general AI model has become a big problem at present. In this paper, by comparing the AI images generated by two different Stable diffsuion models, the advantages and disadvantages of each model are analyzed from the aspects of the matching degree of the AI image and the prompt, the color composition and light composition of the image, and the general AI model that the generated AI image has an aesthetic sense is analyzed, and the designer does not need to take cumbersome steps. A satisfactory AI image can be obtained. The results show that Playground V2.5 model can be used as a general AI model, which has both aesthetic and design sense in various style design requirements. As a result, content designers can focus more on creative content development, and expect more groundbreaking technologies to merge generative AI with content design.

 
1 2
페이지 저장