년 - 년
AI 모델 검증을 위한 기독교 벤치마크 개발을 위한 연구 KCI 등재
한국실천신학회 신학과 실천 제98호 2026.02 pp.571-595
※ 기관로그인 시 무료 이용이 가능합니다.
6,300원
인공지능 기술이 급속도로 발전함에 따라 인공모델에 대한 검증도 활발히 이루어지고 있다. 인공지능 모델 검증은 주로 벤치마크(Benchmark) 방법으로 이루어지는데, 표준 화된 테스트 세트와 평가기준을 적용한다. 대표적인 인공지능 벤치마크로는 MMLU, GPQA 등이 있는데, 새로운 인공지능 모델이 나올 때마다 벤치마크 검증 결과와 분 석이 함께 제시된다. 2025년 AI안전센터와 스케일AI는 기존의 벤치마크의 한계를 극 복하기 위해 전문가들로부터 질문을 수집한 인류 마지막 시험(Humanity’s Last Exam) 프로젝트를 실시하고, 연구한 결과를 보고하였다. 일반 분야와 달리 기독교 분 야에서의 벤치마크 연구는 제대로 이루어지지 않고 있다. 따라서 본 연구에서는 정통 기독교 성경과 교리에 기반한 벤치마크의 사례를 탐색하고 인공지능 모델의 정확성을 검증하기 위해 성경고사 문항을 토대로 기독교 벤치마크 기초 실험을 실시하였다. 실 험 결과로 세 가지 모델인 챗GPT, 제미나이, 클로드의 정확도가 94% 이상을 나타났 으며, 보정 에러는 7% 이하로 나타났다. 또한 거대언어모델별로 정확도와 보정 에러 의 근소한 차이가 나타났다. 전체 정확도는 제미나이, 클로드가 96%로 다소 높았고, 챗GPT는 94%로 상대적으로 낮았다. 주관식 문항은 세 모델 모두 100%의 정답률을 보였고 평균 확신도는 97% 이상으로 보고되었다. 보정 에러는 클로드, 제미나이, 챗GPT 순서로 나타났다. 세 모델 모두 보정 에러가 낮지만 높은 정답률과 확신도에 비 해 틀린 문항이 일부 나타난 것으로 보아 여전히 환각이 존재한다는 것을 알 수 있 다. 문항별 오답과 확신도를 분석한 결과 거대언어모델이 환각을 일으키는 현상이 일 부 분석되었다. 주요 사례는 제시한 성경구절에 명확하게 답이 나타나 있지 않아서 전체 맥락을 살펴보고 추론이 필요한 경우, 유사한 단어에서 정확한 단어를 선택해야 하는 경우, 성경구절을 명확히 참조하지 않고 문장의 앞부분만 계산해서 답을 선택한 경우, 한국어를 제대로 인식하지 못해 오답을 표기한 경우로 나타났다. 현재 개발된 인공지능 모델은 기독교 분야에서 잠재적 활용 가능성이 있으나 실제 활용에서는 주 의가 필요하다는 결론이 도출되었다. 또한 기독교 벤치마크를 개발할 때는 성경 구절 에 나타난 단순 지식을 묻는 문항 보다는 고차원적 추론이나 앞뒤 문맥의 맥락을 이 해해야 풀 수 있는 문항으로 개발할 필요가 있다는 점을 시사한다. 본 연구의 결과는 기독교 분야의 벤치마크 모델 개발의 기초자료가 된다.
As artificial intelligence technology advances rapidly, validation of AI models is also being actively conducted. AI model validation is primarily carried out through benchmarking methods, which apply standardized test sets and evaluation criteria. Representative AI benchmarks include MMLU and GPQA, and whenever a new AI model is released, benchmark validation results and analyses are presented together. In 2025, the Center for AI Safety and Scale AI conducted the Humanity’s Last Exam project, collecting questions from experts to overcome the limitations of existing benchmarks, and reported their findings. Unlike in general domains, benchmark research in the Christian domain has not been adequately conducted. Accordingly, this study explores benchmark cases grounded in orthodox Christian Bible and doctrine and conducted a foundational Christian benchmark experiment based on Bible examination items to verify the accuracy of AI models. The experimental results showed that the accuracies of the three models— ChatGPT, Gemini, and Claude—were at least 94%, and the calibration error was at most 7%. Small differences in accuracy and calibration error were observed across large language models. Overall accuracy was somewhat higher for Gemini and Claude (96%), whereas ChatGPT was relatively lower (94%). For the constructed-response items, all three models achieved a 100% correct answer rate, and the reported an average confidence was at least 97%. Calibration error was observed in the order of Claude, Gemini, and ChatGPT. Although all three models exhibited low calibration error, the presence of some incorrect items despite high accuracy and confidence indicates that hallucinations still exist. An analysis of item-level incorrect responses and confidence revealed some instances of hallucination in large language models. Major cases included situations in which the correct answer was not explicitly stated in the presented biblical passage and required inference from broader context, cases requiring selection of an exact term among similar words, cases in which the answer was chosen by considering only the beginning of a sentence without explicit reference to the biblical passage, and cases in which incorrect answers were produced due to inadequate recognition of Korean. It was concluded that, although currently developed AI models have potential applicability in the Christian domain, caution is required in their practical use. In addition, the findings suggest that, when developing a Christian benchmark, it is necessary to design items that require higher-order reasoning or understanding of contextual coherence rather than items that merely ask for simple knowledge stated in biblical passages. The results of this study provide foundational data for developing benchmark models in the Christian domain.
국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 14 Number 4 2025.12 pp.114-121
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
This study aims to develop an AI-integrated educational model that incorporates generative AI into the UX design curriculum and to validate its effectiveness through Design-Based Research (DBR). The research followed the ADDIE instructional design framework—Analysis, Design, Development, Implementation, and Evaluation—using a university-level UX design course as the experimental setting. In Phase 1, learner needs and the existing curriculum were analyzed to design a project-based learning model centered on AI-supported design processes. In Phase 2, project-based classes were conducted using generative AI tools such as ChatGPT, Gemini, and Midjourney. Learning outcomes were evaluated using a structured rubric focusing on creativity, problem-solving ability, and collaboration. In Phase 3, the curriculum was refined through instructor interviews and learner feedback analysis. The findings indicate that integrating generative AI into design education significantly improved students’ problem-definition accuracy, visualization speed, and collaborative engagement. It also increased their awareness and literacy in AI utilization. This study proposes practical instructional design guidelines for incorporating AI into UX design education and provides foundational insights for developing qualitative evaluation models to support the digital transformation of design pedagogy. We aim to impact the future of design education by establishing a practical and sustainable framework that bridges generative AI technology with creative design learning.
기독교 대안학교 교육 맥락에서의 AI 도구 활용 수업 모형 개발 및 타당도 검증
한국기독교교육학회 기독교교육논총 제86집 2026.06 pp.165-191
※ 원문이용 방식은 연계기관[NRF]의 정책을 따르고 있습니다.
연구 목적 : 본 연구는 기독교 대안학교의 교육 맥락 속에서 AI 도구를 활용하는 수업 모형을 개발하고 이에 대한 타당성을 검증하는 것을 목적으로 하고 있다. 이를 통해 AI 교육에 대한 철학적 논의를 수업 현장으로 연결하고 신앙과 기술의 융합에 기독교교육 현장에 적용하는 방향성을 탐구하고자 한다. 연구 내용 및 방법 : 연구의 목적을 성취하기 위해 48개의 기독교 대안학교 웹페이지에 제공되는 교육 철학과 내용을 분석하고, 기독교 대안학교 교사들과의 인터뷰를 통해 해당 내용을 확인하였다. 또한 국내 주요 교육연구기관의 연구보고서 14건과 52건의 AI 도구 활용 수업 사례 관련 선행 연구를 문헌 분석하여 AI 도구 활용에 관한 설계 원리를 도출하였다. 이를 기반으로 2차례의 델파이 조사를 거쳐 학습자 진단(D)–성경적 접근(A)–맞춤형 내용 제시(C)–성경적 학습 전이(T)–통합적 평가(E)의 5단계 12세부 단계로 구성된 DACTE 수업 모형을 개발하고, 전문가 델파이를 통해 타당성을 검증하였다. 결론 및 제언 : 본 연구는 공교육 중심의 AI 정책에서 사각지대에 놓인 기독교 대안학교를 위하여 AI 도구를 활용하는 수업 모형을 개발하고, 전문가 델파이를 통해 그 타당성을 확인한 것이다. 기독교 대안학교의 교육철학과 AI 활용 교육의 원리를 통합한 DACTE 모형을 현장에 제공한 것은 매우 시의적절한 연구의 결과로 판단된다. 향후 다양한 교과와 학교급에서 DACTE 모형을 적용한 실증 연구와 함께, 기독교 대안학교를 대상으로 한 AI 수업 사례를 지속적으로 축적하는 후속 연구를 제안한다.
Purpose of the Study : This study aims to develop an instructional model utilizing AI tools within the educational context of Christian alternative schools and to verify its validity. Through this process, the study seeks to connect philosophical discussions on AI education to actual classroom practice and to explore directions for applying the integration of faith and technology in Christian educational settings. Research Methods : To achieve the research objectives, the educational philosophies and content of 48 Christian alternative school websites were analyzed, supplemented by teacher interviews. Literature analysis of research reports and prior studies on AI tool utilization was conducted to derive instructional design principles. Based on these findings, the DACTE instructional model ? consisting of five stages and twelve sub-stages, namely Diagnosis (D), Biblical Approach (A), Adaptive & Biblical Content Presentation (C), Biblical Transfer (T), and Holistic Evaluation (E) ? was developed through two rounds of Delphi surveys and validated through expert consultation. Conclusions and Suggestions : This study developed an instructional model utilizing AI tools for Christian alternative schools, which have been marginalized in public education-centered AI policy, and confirmed its validity through expert Delphi consultation. Providing the DACTE model ? which integrates the educational philosophy of Christian alternative schools with the principles of AI-based education ? to the field is deemed a timely and significant research outcome. Future studies are encouraged to conduct empirical research applying the DACTE model across various subjects and school levels, as well as to continuously accumulate AI instructional cases targeting Christian alternative schools.
Validation of RSDA Model in Moral Decision-Making of Artificial Moral Agent or AI Robots
[NRF 연계] 한국지식정보기술학회 (사)한국지식정보기술학회논문지 Vol.18 No.6 2023.12 pp.1535-1545
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The objective of this research is to validate the RSDA model for forecasting decision-making processes in Artificial Moral Agents (AMAs) or AI Robots. The RSDA model delineates the Rubric evaluation, Scenario development, and Data collection phases, ultimately culminating in the creation of an Algorithm capable of predicting robot or AI responses that closely align with human decision-making. This investigation demonstrates the feasibility of a hybrid AMA model that can emulate human ethical judgment by discerning ethical scores through the analysis of human decision-making patterns in real-life scenarios. Data for this study were gathered from a cohort of elementary and college students who responded to four ethical dilemma scenarios involving domestic, medical, and educational robots, as well as autonomous vehicles. The Univariate Dynamic Encoding Algorithm for Searches (uDEAS) was subsequently employed to construct a statistical model that conforms to the decision-making patterns observed in human groups under consideration. According to the results of the RSDA model, the absolute mean ethics score for ethical principle 1 is 0.49 for elementary school students and 1.53 for university students, indicating that ethical awareness of human rights develops as students grow older. In addition, the average standard deviation of the ethics scores of the five principles is 1.02 for elementary school students and 0.67 for college students, indicating that ethical judgment narrows with age and ethical consensus is formed. The findings of this study affirm that the RSDA model, while intuitive, systematically elucidates each step and holds substantial promise for deployment in scenarios wherein intelligent agents necessitate human-like decision-making capabilities. Moreover, the RSDA model is anticipated to enhance the credibility of AMAs by augmenting the transparency and explainability of decisions made by social robots or AI.
이미지 기반 AI 모델의 품질 고도화를 위한 이미지 증강 및 테스트 방법의 융복합 연구
[NRF 연계] 한국전시산업융합연구원 한국과학예술융합학회 Vol.41 No.4 2023.09 pp.205-215
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 이미지 기반 인공지능 모델의 품질을 검증하기 위해 이미지 특징별로 테스트 셋을 구성하기 어려운 점을 해소하기 위해서 시작되었다. 선행 연구에서는 인공지능 모델을 학습하는데 많은 양의 라벨링 데이터가 필요하여 활용에 어려움이 있었다. 본 연구의 목적은 수작업 분류나 라벨링을 하지 않고도 데이터 셋을 편향없이 구성하여 군집별로 테스트하고 인공지능 모델의 성능을 향상시키는 것이다. 따라서 본 논문은 가상 착의 인공지능 모델DC-VTON을 대상으로 두 가지 실험을 수행하고 결과를 비교하였다. 연구 결과 및 내용은 다음과 같다. 첫째, 두 개의 군집화 방법으로 테스트 셋을 각 5개의 군집으로 나누고 모델을 테스트하였다. 두 방법에서 군집들의 모델 품질 지표 SSIM 값의 편차를 비교하였다. 합성곱신경망 기반 군집화 방법의 편차가 픽셀 값 기반 군집화 방법의 편차보다 68% 높아서 군집별 성능 차이가더 발생하는 결과를 보였다. 둘째, 모델의 품질 성능 지수 EV가 목표 값에 도달하도록 재학습 데이터를 군집별로 추가하는 두 가지 실험을 하였다. 성능이 낮았던 군집들의 SSIM 평균을 재학습 이후와 비교한 결과 군집별 차등 추가 방법은 3.2% 높아졌고 군집별 균등 추가 방법은 0.64% 높아졌다. 품질이 낮은 군집은 학습데이터를 보완하면 성능이 향상되는 결과를 보였다. 이러한 연구 결과를 바탕으로 이미지 특징별 군집화를 통해 테스트 셋을 정교화하고, 군집별 테스트 결과에 따라 재학습 데이터를 보강함으로써 이미지 기반의 인공지능 모델의 품질을 높일 수 있을 것으로기대한다.
This study was initiated to solve the difficulty of constructing a test set without bias for each image feature to verify the quality of an image-based artificial intelligence model. In previous studies, it was difficult to use because a large amount of labeling data was required for model learning. The purpose of this study is to improve the performance of artificial intelligence models by organizing learning and test datasets without bias and testing them by cluster without manual classification or labeling. Two experiments were conducted on the artificial intelligence model DC-VTON of the virtual try-on and the results were compared. First, the test set was divided into five clusters each using two clustering methods and the model was tested. In the two methods, the deviation of the model quality indicator SSIM values of the clusters was compared. The deviation of convolutional neural network-based clustering was 0.022 larger than that of pixel value-based clustering, resulting in more differences in performance by cluster. Second, two experiments were conducted to add re-learning data for each cluster to reach the target value of the model's quality performance index. Comparing the SSIM average of clusters that had low performance with after re-learning, the differential addition method by cluster increased by 3.2% and the equal addition method by cluster increased by 0.64%. In clusters with low quality, complementing the learning data resulted in improved performance improvement. Based on the results of these studies, it is expected that the quality of image-based artificial intelligence models can be improved by clustering images, which are unstructured data, according to characteristics, to refine test sets and enhance cluster-specific test results.
CT 영상 기반 근감소증 진단을 위한 AI 영상분할 모델 개발 및 검증
[Kisti 연계] 한국정보처리학회 정보처리학회논문지/컴퓨터 및 통신 시스템 Vol.12 No.3 2023 pp.119-126
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
근감소증은 국내는 2021년 질병으로 분류되었을 만큼 잘 알려져 있지 않지만 고령화사회에 진입한 선진국에서는 사회적 문제로 인식하고 있다. 근감소증 진단은 유럽노인근감소증 진단그룹(EWGSOP)과 아시아근감소증진단그룹(AWGS)에서 제시하는 국제표준지침을 따른다. 최근 진단방법으로 절대적 근육량 이외에 신체수행평가로 보행속도 측정과 일어서기 검사 등을 통하여 근육 기능을 함께 측정할 것을 권고하고 있다. 근육량을 측정하기 위한 대표적인 방법으로 DEXA를 이용한 체성분 분석 방법이 임상에서 정식으로 실시하고 있다. 또한 MRI 또는 CT의 복부 영상을 이용하여 근육량을 측정하는 다양한 연구가 활발하게 진행되고 있다. 따라서 본 논문에서는 근감소증 진단을 위해서 비교적 짧은 촬영시간을 갖는 CT의 복부영상기반으로 AI 영상 분할 모델을 개발하고 다기관 검증한 내용을 기술한다. 우리는 CT 영상 중에 요추의 L3 영역을 분류하여 피하지방, 내장지방, 근육을 자동으로 분할할 수 있는 인공지능 모델을 U-Net 모델을 사용하여 개발하였다. 또한 모델의 성능평가를 위해서 분할영역의 IOU(Intersection over Union)를 계산하여 내부검증을 진행했으며, 타 병원의 데이터를 활용하여 동일한 IOU 방법으로 외부검증을 진행한 결과를 보인다. 검증 결과를 토대로 문제점과 해결방안에 대해서 검증하고 보완하고자 했다.
Sarcopenia is not well known enough to be classified as a disease in 2021 in Korea, but it is recognized as a social problem in developed countries that have entered an aging society. The diagnosis of sarcopenia follows the international standard guidelines presented by the European Working Group for Sarcopenia in Older People (EWGSOP) and the d Asian Working Group for Sarcopenia (AWGS). Recently, it is recommended to evaluate muscle function by using physical performance evaluation, walking speed measurement, and standing test in addition to absolute muscle mass as a diagnostic method. As a representative method for measuring muscle mass, the body composition analysis method using DEXA has been formally implemented in clinical practice. In addition, various studies for measuring muscle mass using abdominal images of MRI or CT are being actively conducted. In this paper, we develop an AI image segmentation model based on abdominal images of CT with a relatively short imaging time for the diagnosis of sarcopenia and describe the multicenter validation. We developed an artificial intelligence model using U-Net that can automatically segment muscle, subcutaneous fat, and visceral fat by selecting the L3 region from the CT image. Also, to evaluate the performance of the model, internal verification was performed by calculating the intersection over union (IOU) of the partitioned area, and the results of external verification using data from other hospitals are shown. Based on the verification results, we tried to review and supplement the problems and solutions.
CT 영상 기반 근감소증 진단을 위한AI 영상분할 모델 개발 및 검증
[NRF 연계] 한국정보처리학회 KIPS Transactions on Computer and Communication Systems Vol.12 No.3 2005.06 pp.119-126
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
근감소증은 국내는 2021년 질병으로 분류되었을 만큼 잘 알려져 있지 않지만 고령화사회에 진입한 선진국에서는 사회적 문제로 인식하고 있다. 근감소증 진단은 유럽노인근감소증 진단그룹(EWGSOP)과 아시아근감소증진단그룹(AWGS)에서 제시하는 국제표준지침을 따른다. 최근 진단방법으로 절대적 근육량 이외에 신체수행평가로 보행속도 측정과 일어서기 검사 등을 통하여 근육 기능을 함께 측정할 것을 권고하고 있다. 근육량을측정하기 위한 대표적인 방법으로 DEXA를 이용한 체성분 분석 방법이 임상에서 정식으로 실시하고 있다. 또한 MRI 또는 CT의 복부 영상을 이용하여근육량을 측정하는 다양한 연구가 활발하게 진행되고 있다. 따라서 본 논문에서는 근감소증 진단을 위해서 비교적 짧은 촬영시간을 갖는 CT의복부영상기반으로 AI 영상 분할 모델을 개발하고 다기관 검증한 내용을 기술한다. 우리는 CT 영상 중에 요추의 L3 영역을 분류하여 피하지방,내장지방, 근육을 자동으로 분할할 수 있는 인공지능 모델을 U-Net 모델을 사용하여 개발하였다. 또한 모델의 성능평가를 위해서 분할영역의IOU(Intersection over Union)를 계산하여 내부검증을 진행했으며, 타 병원의 데이터를 활용하여 동일한 IOU 방법으로 외부검증을 진행한 결과를보인다. 검증 결과를 토대로 문제점과 해결방안에 대해서 검증하고 보완하고자 했다.
Sarcopenia is not well known enough to be classified as a disease in 2021 in Korea, but it is recognized as a social problem in developedcountries that have entered an aging society. The diagnosis of sarcopenia follows the international standard guidelines presented by theEuropean Working Group for Sarcopenia in Older People (EWGSOP) and the d Asian Working Group for Sarcopenia (AWGS). Recently,it is recommended to evaluate muscle function by using physical performance evaluation, walking speed measurement, and standing testin addition to absolute muscle mass as a diagnostic method. As a representative method for measuring muscle mass, the body compositionanalysis method using DEXA has been formally implemented in clinical practice. In addition, various studies for measuring muscle massusing abdominal images of MRI or CT are being actively conducted. In this paper, we develop an AI image segmentation model basedon abdominal images of CT with a relatively short imaging time for the diagnosis of sarcopenia and describe the multicenter validation. We developed an artificial intelligence model using U-Net that can automatically segment muscle, subcutaneous fat, and visceral fat byselecting the L3 region from the CT image. Also, to evaluate the performance of the model, internal verification was performed bycalculating the intersection over union (IOU) of the partitioned area, and the results of external verification using data from other hospitalsare shown. Based on the verification results, we tried to review and supplement the problems and solutions.
학생부종합전형 서류평가의 AI 예비평가 모델 설계 및 탐색적 다중 모델 검증
[NRF 연계] 한국교육과정평가원 교육과정평가연구 Vol.29 No.2 2026.05 pp.275-303
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
학생부종합전형 서류평가는 평가자의 정성적 판단에 기반하므로, 평가자의 경험과 대학별 평가 환경에 따른 판단 편차는 평가의 일관성 및 공정성과 밀접하게 관련된다. 이를 완화하려면 대학 고유의 평가 기준을 일관되게 적용할 보조 도구가 필요하다. 이에 본 연구는 입학사정관의 평가를 보조하는 AI 도구가 대학 고유의 평가 기준을 충실히 구현하여 평가의 구성 타당성을 확보하고, 아울러 평가자 간 편차까지 완화할 수 있는지를 탐색적 다중 모델 비교로 검증하였다. 5개 대학 공동연구(임진택 외, 2022)의 평가항목을 바탕으로 AI 예비평가 모델(A: LLM 자체 채점 기준, B: 사실 추출 레이어, C: 루브릭 없는 통제군, D: 전문가 기준 부호화)을 구성하고, 30건의 모의 학생기록을 네 모델과 서로 다른 대학 전문가 2인이 독립 평가하였다. 분석 결과, 범용 모델은 순위 합의(Kendall W=.859)는 양호하였으나 등급 분류(Fleiss κ=.165)에서는 낮은 일치도를 보였다. 서로 다른 대학 전문가 간 일치도는 κ=.397이었으며, 불일치는 모두 인접 등급 차이에 머물러 평가 기준 차이를 반영하는 교차 기관 평가 기준선임을 보여준다. 모델D는 전문가A의 판단을 κ=.648 수준으로 재현하였고, 전문가B와의 일치도는 κ=.261이었으나, 두 전문가 판정이 일치한 18건에서는 77.8%의 일치율로 네 모델 중 가장 높았다. 평가 기준의 부호화는 대학별 기준 적용을 전제로 유효하며, 사실 추출 레이어는 AI 워딩 과다 기록의 사실성 저하를 효과적으로 감지하였다. AI 기반 평가 보조 도구는 수험생을 평가자·연도 편차로부터 보호하고, 실제 노력과 성장을 보인 학생이 가려지지 않도록 하는 공정성 강화에 기여할 수 있다.
Document review in holistic admissions in Korean universities is based on evaluators’ qualitative judgment; therefore, differences in evaluators’ experience and institutional evaluation environments may directly affect fairness among applicants. To mitigate such variation, a support tool is needed to apply each institution’s evaluation criteria consistently. This study examined, through an exploratory multi-model comparison, whether an AI-based evaluation support tool can implement institution-specific evaluation criteria, secure construct validity, and contribute to reducing inter-evaluator variance. Based on the evaluation items from a joint study by five Korean universities (Lim et al., 2022), four AI preliminary evaluation models were constructed: Model A (LLM self-generated criteria), Model B (fact extraction layer), Model C (no-rubric control), and Model D (expert criteria encoding). Thirty mock student records were independently evaluated by the four models and by two expert evaluators from different universities. The results showed that generic models achieved good rank agreement (Kendall W=.859) but low categorical agreement (Fleiss κ=.165). Agreement between the two cross-institutional experts was κ=.397, and all disagreements were confined to adjacent grades, indicating differences in evaluation criteria across institutions. Model D reproduced Expert A’s judgment at κ=.648, while its agreement with Expert B was κ=.261. However, in the 18 cases where the two experts agreed, Model D showed the highest hit rate among the four models at 77.8%. The findings suggest that encoding evaluation criteria is effective when applied to institution-specific criteria, and that the fact extraction layer effectively detects factuality deficits in records with excessive AI-generated wording. AI-based evaluation support tools may contribute to fairness by protecting applicants from evaluator- and year-based variance and by preventing students’ genuine effort and growth from being obscured.
도자 공예 기반 HCD-AI 융합 ‘공예컨텐츠에듀’ 교과목 설계 기초연구 - 공예심리조향 역량을 반영한 P2IM 교수설계 모델 개발 및 전문가 타당도 검증 -
[NRF 연계] 한국전시산업융합연구원 한국과학예술융합학회 Vol.44 No.3 2026.06 pp.307-320
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구는 디지털 과노출로 심화되는 현대인의 인지 피로에 대응하고, 단순 기능 전수를 넘어 생성형 AI 리터러시를 포괄하는 다감각적 공예교육의 필요성에서 출발하였다. 이에 따라 도자 공예의 물성적 특성과 감각 기반 교육 경험을 토대로, 공예심리조향 역량과 HCD-AI 융합 요소를 통합한 교과 설계 가능성을 탐색하였다. 본 연구의 목적은 선행연구에서 구축한 ‘HCD-AI (Human-Centered Design and AI, 인간중심디자인과인공지능) 융합 기반 IM(Iceberg & Making) 교육모델’을 실천적으로 확장·구체화하여, D대학 리빙세라믹디자인학과 리빙공예디자인트랙 3학년의 기존 ‘응용미술교육’ 교과목을 ‘공예컨텐츠에듀’ 교과목으로 개편하기 위한 P2IM (Problem & Project-based IM) 교수설계 모델을 개발하는 데 있다. 이를 위해 디자인·개발연구(DDR) Type 2 방법론을 적용하여, 천연향료의 아로마콜로지 지식을 도자공예의 물성과 결합한 골든써클 기반 15주 실라버스를 설계하고, 로쉬(Lawshe, 1975)의 내용타당도비율(CVR) 공식을 활용해 전문가타당도 검증을 수행하였다. 연구 결과는 다음과 같다. 첫째, IM 모델에 문제·프로젝트 기반 학습(PBL·PjBL)을 결합한 P2IM 교수법을 구축하여, 생성형 AI를 인지적 파트너(30%)로, 도자 공예 수작업(70%)을 치유적 제작 활동으로 활용하는 동적 순환적 학습체계를 구체화하였다. 둘째, 도자의 다공성 및 열전도성을 활용하여 천연향료 성분(1,8-Cineole, Linalool)의 발향을 물리적으로 조절함으로써, 인지·정서 지원과 연계 가능한 웰니스 교육 요소로서의 정서 지원 교수 체계를 제안하였다. 셋째, 내용타당도 검증 결과, 제안된 설계안은 융합형 공예 에듀케이터의 핵심 직무역량과 높은 부합성을 보이는 것으로 확인되었다. 본 연구는 P2IM 교과목 모델을 통해 K-Craft 기반 웰니스 공예교육의 실천 가능성을 제시하였으며, 향후 2027년 해당 전공의 교육과정 개편 적용 시 커크패트릭(Kirkpatrick) 4단계 평가모형에 근거한 학습자 성취도 추적과 후속 실증 연구를 위한 학술적 기초를 마련하였다는 점에서 의의를 지닌다.
This study addresses the need for multisensory craft education that responds to cognitive fatigue associated with digital overexposure while extending conventional skill-based instruction to include generative AI literacy. Based on the material properties of ceramic craft and sensory educational experience, it explores the design of a curriculum integrating craft psycho-aromatherapy competencies with HCD-AI convergence elements. The purpose of this study is to develop a P2IM (Problem- and Project-based IM) instructional design model by extending the previously established HCD-AI-based IM (Iceberg & Making) educational model and to restructure the existing “Applied Art Education” course into the “Craft Contents Edu” curriculum for the third-year Living Craft Design Track in the Department of Living Ceramic Design at D University. Using Design-Based Research (DDR) Type 2 methodology, this study developed a 15-week Golden Circle-based syllabus integrating aromachology with the material characteristics of ceramic craft and conducted expert validation using Lawshe’s (1975) Content Validity Ratio (CVR). The findings are as follows. First, by integrating Problem- and Project-Based Learning (PBL/PjBL) into the IM model, the study established the P2IM pedagogy as a dynamic cyclical learning system in which generative AI serves as a cognitive partner (30%) and ceramic handcraft practice as the primary medium of embodied learning (70%). Second, by utilizing the porosity and thermal conductivity of ceramics to regulate the diffusion of fragrance compounds such as 1,8-Cineole and Linalool, the study proposed an emotional-support instructional framework applicable as a wellness-oriented educational component linked to cognitive and affective support. Third, the CVR results indicated that the proposed design showed strong alignment with the core competencies required of a convergent Craft Educator. This study suggests the practical applicability of K-Craft-based wellness craft education through the P2IM model and provides a basis for future follow-up studies, including learner achievement tracking based on Kirkpatrick’s four-level evaluation model in the context of the 2027 curriculum reform.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.