Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 28
No
1

6,300원

최근 대형 언어 모델(LLM)의 성능이 전문가 수준으로 발전하고 있다. 본 연구는 고객서비스 도메인에서 기존 인간 중심 평가 방식의 한계를 극복하고자 Expert AI(평가 전문가로서의 역할을 하는 AI)를 활용한 품질 모니터링 시스템 구축 방안을 새롭게 제시하고자 한다. 본 연구에서는 고객 질문에 대해 “고객서비스 도메인 특화된 Local LLM 기반 AI”로부터 생성된 응답 데이터 결과물의 품질을 인간 평가자와 Expert AI(AI 평가자)가 평가하도록 하였다. 단, Local LLM의 품질 평가 시, 다른 LLM AI(ChatGPT와 Perplexity)로부터 답변된 응답과 비교하여 평가되도록 하여, 도메인 특화로 구축된 Local LLM의 현재 품질 상황과 개선할 점을 파악할 수 있도록 하였다. 평가 수행 이후, 인간 평가자와 Expert AI의 품질 평가에 대한 일치도를 분석한 결과, 92.4%에 달하는 높은 일치율을 확인하였다. 특히 Local LLM의 문제점 탐지(응답 품질이 나쁜 경우)에서는 98.9%의 일치율을 보여, Expert AI가 핵심 품질 관리 역할을 수행할 수 있음을 입증하였다. 본 연구는, 지속적이면서 실시간 품질 관리가 필요한 고객서비스 분야에서 Expert AI의 적용 가능성을 실증적으로 검증함으로써, 실무 환경에서 효율성을 극대화할 수 있는 AI 기반 품질 모니터링 시스템 구축 방안을 제안한다.

Recently, the performance of large language models (LLMs) has improved to expert levels. This study proposes a novel approach to building a quality monitoring system utilizing Expert AI (AI acting as an evaluation expert) to overcome the limitations of existing human-centered methods in the customer service (CS) domain. In this study, the quality of responses to customer inquiries generated by a “CS domain-specific Local LLM-based AI” was assessed by human evaluator and Expert AI. For the quality assessment of the Local LLM, its responses were compared with those generated by other LLM-based AIs (ChatGPT and Perplexity) to evaluate the current quality and identify areas for improvement in the Local LLM. After the evaluation, we analyzed the agreement between human and AI evaluators in quality assessment, confirming a high agreement rate of 92.4%. In particular, the agreement rate for detecting Local LLM issues (e.g., poor response quality) was 98.9%, demonstrating that Expert AI can play a key quality control role. This study empirically verified the applicability of Expert AI in the CS field, which requires continuous and real-time quality management. It proposes a method for building an AI-based quality monitoring system that can maximize efficiency in corporate environments.

2

5,100원

이 연구는 ChatGPT 도입이 한국 내 온라인 번역기 사용 패턴에 미친 영향을 정량적으로 분석하였다. 대규모 국내 활동 패널 데이터를 바탕으로 이벤트 타임 사양을 갖춘 매칭 기반 차이의 차이 모형을 적용한 결과, 채택 이후 주간 번역기 사용 시간은 증가하며 도입 후 약 6주에 정점을 보인 뒤 평균적으로 음수로 전환하지 않은 채 0에 가까운 정상상태로 점차 수렴하였다. 평균 효과 뒤에는 뚜렷한 이질성이 존재한다. 보완 효과는 더 젊은 이용자에서 더 크게 나타났고, 대체 패턴은 대학생 및 대학원생 집단에서만 관찰되었다. 이론적 측면에서, 일반 목적 LLM이 유사한 기능을 가지는 단일 도구를 평균적으로 보완한다는 사용자 수준의 인과적 근거를 제시하고, 이질성을 역량에 기반해 설명하며, 이벤트 타임 동학과 사전 추세 공동 검정을 결합해 초기 탐색과 정상상태를 구분하는 관찰자료 분석 템플릿을 제공한다. 실무적 측면에서는 LLM 출력의 번역기 검증으로의 원클릭 연동, 신뢰도 표시를 위한 경량 품질 배지, 전환 비용을 낮추는 명시적 전환 흐름 설계, 보완성이 큰 집단을 겨냥한 세그먼트별 온보딩 등 공용 설계 지침을 제시한다.

We examine how first use of a general purpose LLM, represented by ChatGPT, reallocates user time within the ecosystem of online translators. Using a matched difference-in-differences design with an event-time specification on a large Korean activity panel, we find that weekly translator use increases after adoption, peaks during the first six weeks, and then attenuates toward a near-zero steady state without turning negative on average. The average effect masks systematic heterogeneity: complementarity is larger for younger users, while a substitution pattern appears only among university and graduate students. Theoretically, we provide causal user-level evidence that a general-purpose LLM complements an online translation on average, link heterogeneity to capability, and offer a transparent template that combines event-time dynamics with a joint pre-tend test to separate initial exploration from steady state behavior. Practically, we translate the mechanisms into co-use design guidance, including one-click-pass through from LLM outputs to translator verification, and segment-specific onboarding that targets groups where complementarity is strongest.

3

Predicting Corporate Culture for M&A: Leveraging Fine-Tuning and Chain-of-Thought Strategies with LLMs

Ivan Ivanov, Hee Seok Song

한국정보기술응용학회 JITAM Vol.32 No.4 2025.08 pp.51-73

※ 기관로그인 시 무료 이용이 가능합니다.

6,000원

This study explores a novel approach to assessing cultural fit during the early stages of mergers and acquisitions (M&A) by leveraging publicly available employee review data and large language models (LLMs). Recognizing the limitations of traditional due diligence in accessing internal cultural data, the proposed framework utilizes fine-tuned models and chain-of-thought (CoT) reasoning strategies to infer corporate cultural characteristics based on the Denison and Ko [2016] framework. The model evaluates four key traits-Mission, Consistency, Involvement, and Adaptability-across twelve dimensions, using Low-Rank Adaptation (LoRA) for efficient fine-tuning. Experimental results demonstrate that LoRA-tuned models consistently outperform few-shot prompting across both proprietary (e.g., GPT-4o) and open-source (e.g., Llama 3.2-3B) models, with significant improvements in both text summarization and numerical prediction accuracy. Additionally, CoT reasoning-particularly Multi-step and Hybrid strategies-yields substantial performance gains, especially in smaller models, enabling them to approximate the results of large-scale systems at reduced cost. These findings highlight the practical utility of combining PEFT and CoT methods for scalable, objective, and early-stage cultural assessments in M&A decision-making.

4

HyDE 기반 멀티 홉 검색 기법을 활용한 검색 성능 향상 방안 KCI 등재

김예은, 이재홍, 원상혁, 정우혁, 우지환

한국경영정보학회 경영정보학연구 제27권 제2호 2025.05 pp.127-148

※ 기관로그인 시 무료 이용이 가능합니다.

5,800원

생성형 인공지능과 대형 언어 모델(LLM)의 발전은 기업의 업무 프로세스에 혁신을 가져오며 도메인 특화된 업무를 처리하기 위한 맞춤형 LLM 도입을 촉진하고 있다. 그러나 LLM의 환각(hallucination) 문제와 도메인 적합성 부족을 해결하기 위한 Retrieval-Augmented Generation(RAG) 시스템은 검색 단계에서 복잡한 멀티 홉 질의 처리 시 오류 누적 문제로 인해 성능 저하를 겪는다. 본 연구는 이러한 문제를 해결하기 위해 Hypothetical Document Embedding(HyDE) 기법을 멀티 홉 검색에 통합하여 RAG 성능을 개선하는 프레임워크를 제안한다. HyDE 기법은 질의의 의미를 반영한 가상 문서를 생성하여 검색 정확성을 향상시키며, 본 연구에서는 복잡한 질의를 단계적으로 단일 홉 질의로 분해하고 각 단계에서 HyDE를 적용하는 방식을 채택하였다. 실험은 검색 정확도를 측정하기 위해 precision@k, recall@k, F1 score, MAP, MRR, hit rate와 같은 지표를 사용하여 진행되었다. 실험 결과 HyDE 기반 멀티 홉 검색은 모든 지표에서 기존 대비 향상된 성능을 보였으며, 특히 recall이 약 19.53%, hit rate가 21.21% 증가하였다. 이는 HyDE가 멀티 홉 검색에서 검색의 정확성을 높이는 데 효과적임을 보여준다. 향후 연구에서는 멀티 홉 검색에 최적화된 데이터셋 개발, 가상 문서 생성 전략의 개선 등을 시도할 수 있을 것으로 기대된다.

The development of generative AI and large language models (LLMs) is revolutionizing business processes and fostering the adoption of customized LLMs for domain-specific tasks. However, Retrieval-Augmented Generation (RAG) systems designed to address hallucination and domain relevance issues face performance degradation due to error accumulation in handling complex multi-hop queries. This study proposes a framework that integrates the Hypothetical Document Embedding (HyDE) technique into multi-hop retrieval to enhance RAG performance. HyDE generates virtual documents reflecting query intent, improving retrieval accuracy by decomposing complex queries into single-hop queries and applying HyDE iteratively at each step. Experiments were conducted using precision@k, recall@k, F1 score, MAP, MRR, and hit rate as evaluation metrics. The results demonstrate that HyDE-based multi-hop retrieval improves performance across all metrics, with recall increasing by approximately 19.53% and hit rate by 21.21%. These findings confirm the effectiveness of HyDE in enhancing retrieval accuracy for multi-hop search. Future research directions include the development of optimized datasets for multi-hop retrieval and further refinement of virtual document generation strategies.

5

실시간 문맥 인식 감성 분석을 위한 모듈형 아키텍처 설계 KCI 등재후보

안은희, 안정국

한국융합학회 미래기술융합논문지 제4권 제2호 2025.04 pp.225-231

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

최근 인공지능 기술과 대형 언어모델(Large Language Models, LLM)의 발전으로 감성 분석의 정확도는 꾸준히 향상되 고 있으나, 실시간 문맥 인식, 감성 추론, 사용자 반응 분석 등 실용적 적용을 위한 구조적 기반은 여전히 미흡한 실정이다. 본 연구에서는 기존 감성 데이터와 온톨로지 기반 감성 분석 모델을 바탕으로, 실시간 감성 흐름을 인식하고 반응형 서비스를 지원 할 수 있는 모듈형 아키텍처를 설계하였다. 제안 아키텍처는 LSTM 기반 Attention 모델을 중심으로, 실시간 데이터 스트림 환경 에서 감성 변화 추론과 시각화 기능을 통합할 수 있도록 구조화되었다. 또한, 대형 언어모델의 문맥 이해 능력과 감성 온톨로지의 의미 구조를 결합하여 감성 해석의 신뢰도와 도메인 적응력을 향상시켰다. 실험을 통해 제안된 구조는 기존 감성 분석 시스템 대비 향상된 실시간 처리 성능과 감성 예측 정확도를 입증하였으며, 사회적 이슈 대응, 고객 피드백 분석, 여론 흐름 모니터링 등 다양한 응용 환경에서의 활용 가능성을 확인하였다. 본 연구는 감성 분석 기술의 실시간 응용을 위한 구조적 기반을 제시함으 로써, 사용자 중심의 지능형 감성 서비스 설계에 기여할 수 있다.

Recent advancements in artificial intelligence and large language models (LLMs) have significantly improved the accuracy of sentiment analysis. However, the structural foundation for practical applications— such as real-time context awareness, sentiment inference, and responsive user interaction—remains insufficiently addressed. This study proposes a modular architecture designed to support real-time sentiment flow recognition and reactive services, building upon existing sentiment datasets and ontology-based analysis models. The proposed architecture centers around an LSTM-based attention mechanism and is structured to integrate sentiment variation inference and visualization within real-time data stream environments. Furthermore, by combining the contextual understanding capabilities of large language models with the semantic structure of sentiment ontologies, the system enhances both the reliability of sentiment interpretation and domain adaptability. Experimental results demonstrate that the proposed architecture outperforms conventional sentiment analysis systems in terms of real-time processing and predictive accuracy. The system also shows strong potential for application in various real-world scenarios, including social issue detection, customer feedback analysis, and public opinion monitoring. This research contributes a structural foundation for the real-time deployment of sentiment analysis technologies and supports the development of user-centric, intelligent sentiment-aware services.

6

4,000원

공공 데이터가 폭발적으로 증가함에 따라 데이터 기반 의사결정의 중요성이 더욱 부각되고 있다. 본 연구에 서는 대기오염, 교통사고, 백신 접종률이라는 주요 공공 데이터를 활용하여, 생성형 AI와 대형 언어 모델(LLM)을 결합한 적응형·확장형 공공 데이터 분석 프레임워크를 제안한다. 검증 결과, 교통사고 위험 예측 정확도가 최대 10%p 향상되었고, 대기오염 예측 MAPE가 0.125까지 낮아졌으며, 자동 보고서 생성 및 시각화를 통해 비전문가의 이해도가 40% 이상 향상되었다. 데이터 부족, 모델 편향, ‘블랙박스’ 문제 등 윤리·법적 이슈가 여전히 존재하지만, 본 연구는 공공 정책 의사결정에서 생성형 AI와 LLM이 지닐 수 있는 혁신적 잠재력을 확인하였다.

With the explosive growth of public data, particularly in the domains of air pollution, traffic accidents, and vaccination rates, data-driven decision-making has become increasingly critical. We propose an adaptive, scalable public data analysis framework that integrates Generative AI and Large Language Models (LLMs). Validated on three major datasets—air pollution (WAQI), traffic accidents (US DOT), and vaccination rates (WHO)—our method improves traffic accident risk prediction accuracy by up to 10 percentage points, reduces air pollution MAPE to 0.125, and enhances non-experts’ understanding by over 40%. Despite persistent challenges such as data scarcity, model bias, and the black-box problem, our findings underscore the transformative potential of Generative AI and LLMs in public policy decision-making.

7

The Potential of Large Language Models(LLMs) for Career Transition Education : A Systematic Review Using the PRISMA Method KCI 등재

Seon-Joo Kim, Yun-Kyung Kang, Jeong-Hyun Park, Sung-Tae Lee

한국성인계속교육학회 성인계속교육연구 제15권 제4호 2024.12 pp.65-93

※ 기관로그인 시 무료 이용이 가능합니다.

6,900원

대형언어모델(LLM)은 AI 기술의 일종으로, 인간의 언어를 이해하고 생성하는 능력을 갖추고 있으며, 이를 통해 학습자에게 맞춤형 학습 지원을 제공할 수 있다. 본 연구에서는 경력전환교육 분야에서 LLM의 정의를 명확히 하고, 이를 활용한 교육적 가능성과 효과를 체계적으로 분석하고자 한다. 이를 위해 PRISMA(Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 체계적 문헌고찰 방법론을 활용하여 2015년부터 2024년까지 Google Scholar, Scopus, JSTOR, ProQuest, DOAJ와 같은 주요 데이터베이스에서 수집된 영 어권 문헌과 국내 연구를 검토하였다. 연구는 LLM의 교육적 활용 가능성과 경력전환교육 분야에서의 역할 및 실증적 효과를 중심으로 분석하였다. 연구 결과는 다음 세 가지로 요약된다. 첫째, LLM은 개인화된 학습 환경 설계와 실시간 피드백 제공을 통해 학습 동기를 촉진하고, 학습자 중심의 맞춤형 학습 지원을 가능하게 하는 등 교육적 활용 가능성을 보 여준다. 둘째, LLM의 교육적 활용에 대한 실증적 연구를 통해, 창업 아이디어 개발, 시장 분석, 복잡한 문제 해결 등에서 효과적인 학습 결과가 도출되었으며, 경력전환교육에 적용 할 가능성을 확인하였다. 셋째, 경력전환교육 분야에서 LLM은 창업 지원, 팀워크 훈련, 갈 등 관리와 같은 직무 역량 개발과 학습자 중심의 교육을 지원하는 효과적인 도구로 활용 될 수 있음을 보여준다. 본 연구는 LLM의 교육적 활용 가능성과 실질적 효과를 탐구하여, 경력전환교육의 학습 환경을 혁신하고 창업 및 직업 역량 개발을 체계적으로 지원하는 데 기여한다. 이는 디지털 전환 시대의 요구에 부응하는 경력전환교육 패러다임 구축의 기반 을 마련할 것으로 기대된다.

This study systematically analyzes how Large Language Models(LLMs) can be utilized in career transition education using the PRISMA(Preferred Reporting Items for Systematic Reviews and Meta-Analyses) methodology. The research examines English-language literature collected from major databases such as Google Scholar, Scopus, JSTOR, ProQuest, and DOAJ, as well as Korean literature from domestic databases like RISS, spanning the years 2015 to 2024. The findings of the study are summarized as follows: First, LLMs demonstrate potential for educational applications by fostering learning motivation and enabling personalized, learner-centered education through the design of tailored learning environments and real-time feedback. Second, in the field of career transition education, LLMs can serve as critical tools for practical learning and the development of job competencies, including entrepreneurship support, teamwork training, and conflict management. Third, an empirical analysis of the effects of LLMs reveals positive learning outcomes in areas such as entrepreneurial idea development, market exploration, and complex problem-solving. By concurrently analyzing global and domestic literature, this study validates the educational potential and practical impact of LLMs, contributing to the establishment of a new paradigm for career transition education in the era of digital transformation.

8

4,000원

본 연구는 거대 언어 모델(LLM; Large Language Model) 기반의 메타인지 진단평가 시스템 을 설계하고 구현한 내용을 다룬다. 현재 교육 환경에서 자기 주도적 학습의 중요성이 증가하고 있 는 가운데, 본 시스템은 개인별 메타인지 수준을 효과적으로 평가하고 맞춤형 피드백을 제공할 수 있는 기능을 가진다. 연구 방법으로는 51개의 문항으로 구성된 리커트 5점 척도를 활용하여 사용자 의 메타인지 요소를 분석하였다. 그 결과, LLM을 통해 각 개인의 메타인지 강점과 약점을 구체적으 로 진단하고, 이를 바탕으로 사용자에게 실질적인 피드백을 제공하는 시스템을 개발하였다. 향후 이 시스템은 교육 및 심리 평가 분야에서의 혁신적 접근으로 자리잡을 것으로 기대되며, 추가 연구를 통해 다양한 환경 속에서의 효과성을 검증할 필요가 있다.

This study addresses the design and implementation of a metacognitive diagnostic assessment system based on large language models (LLMs). With the increasing importance of self-directed learning in current educational environments, this system possesses the capability to effectively evaluate individual metacognitive levels and provide personalized feedback. The methodology employed a Likert scale consisting of 51 items to analyze users' metacognitive elements. The findings indicated that through LLMs, strengths and weaknesses in metacognition could be accurately diagnosed, resulting in a system that offers practical feedback to users based on their assessments. Future developments are expected to position this system as an innovative approach in the fields of education and psychological assessment, necessitating further research to validate its effectiveness across diverse educational contexts.

9

4,000원

본 연구는 대규모 언어 모델(LLM)의 기술적 특성과 GDPR 원칙 간 충돌에 관한 개인정보 리터러시 교육 효과를 실증적으로 분석하였다. 이를 위해 인공지능 및 정보보호 전공 대학원생 31명을 대상으로 단일 집단 사전-사후 검사 설계를 적용하여 5가지 기술적 요인(대규모 데이터 사용, 모델 경직성, 데이터 편향성, 블랙박스, 새로운 보안 위협)에 따른 인식 변화를 측정하였다. 분석 결과, 교육 후 ‘모델 경직성’과 ‘데이터 편향성’ 요인에서는 위험 인식이 유의미하게 강화된 반면, ‘블랙박스’와 ‘대규모 데이터 사용’ 관련 항목에서는 기술적 이해도 향상에 따라 막연한 불안감이 합리적으로 조정되는 경향이 나타났다.

This study provides an empirical analysis of the impact of privacy literacy education on the tension between the technical features of Large Language Models (LLMs) and GDPR principles. Utilizing a single-group pre- and post-test design with 31 graduate students in AI and information security, we evaluated perceptual changes based on five technical domains: large-scale data use, model rigidity, data bias, the black box phenomenon, and new security threats. The results indicate that while risk perception was significantly heightened regarding ‘model rigidity’ and ‘data bias,’ arbitrary anxiety surrounding ‘black box’ issues and ‘large-scale data use’ was rationally mitigated through enhanced technical comprehension.

10

5,500원

This study employs Artificial Intelligence to develop experimental future scenarios for North Korean regional development. Focusing on government-proposed areas of inter-Korean economic cooperation - energy, transportation, and agriculture - the research analyzes current challenges including power supply instability, logistics inefficiency, and low agricultural productivity. The study proposes solutions utilizing Fourth Industrial Revolution technologies, such as smart grid systems with Northeast Asian connections, autonomous logistics networks, and advanced agricultural technologies. The use of Large Language Models (LLMs), with their extensive data processing capabilities, enabled generation of creative future scenarios beyond conventional research methods. Policy simulations using virtual personas demonstrated the potential for developing tailored policies considering North Korea’s characteristics. This research holds methodological significance through its experimental use of LLMs and offers valuable reference material for future inter-Korean economic cooperation policies by exploring new possibilities for North Korean regional development.

11

6,000원

2020년 OpenAI가 1,750억 파라미터 규모의 GPT-3를 공개한 이후 간단한 작업부터 복잡한 작업에 이르기까지 다양한 다운스트림 작업에 대응하는 대규모 언어 모델(LLM)의 개발이 가속화되고 있다. LLM이 개발되고 고도화됨에 따라 LLM의 성능을 객관적으로 평가할 수 있는 평가 지표와 데이터셋이 개발되어 활용되고 있다. 이러한 데이터셋은 다양한 분야에 대해 LLM을 객관적으로 평가함에 있어 좋은 성과를 거두었으나, 규모 측면에서 개인이나 소규모 기관에서 활용하기 어렵고 실용적 측면에서 실제 사용자가 체감하는 바와 다소의 괴리를 가지고 있다. 이에 본 연구에서는 사용자의 활용 패턴을 반영하여 비교적 작은 양의 데이터를 활용해 LLM을 평가할 수 있는 평가 지표 및 데이터셋을 제시한다. 더 나아가, 가중치를 일반에 공개하는 공개형 대규모 언어 모델의 개발이 가속화되고 고성능의 공개형 LLM이 출시되고 있음에 따라 연구 수행 시점인 2024년 4월 기준 최신의 폐쇄형 LLM 4종과 공개형 LLM 6종에 대한 평가를 시행하고 폐쇄형 LLM과 공개형 LLM의 비교 평가 결과에 대해 논의한다. 연구 결과 새롭게 개발한 데이터셋이 작은 규모에도 불구하고 기존 데이터셋과 유사한 경향성을 보이는 것으로 나타났다. 상식 추론 및 글 스타일 변환과 같은 간단한 작업에서는 공개형 LLM이 폐쇄형 LLM과 대등하거나 우세한 성능을 보였으나 수학, 코딩, 이미지 질의응답 등의 복잡한 작업에서는 큰 성능 격차를 보임을 확인하였으며, 더 나아가 비교적 작은 규모의 LLM이 규모 대비 좋은 성능을 보임을 확인하였다.

The development of large language models (LLMs) has accelerated since OpenAI released GPT-3, which demonstrated generalizability and capability for various downstream tasks, thanks to its 175 billion parameters. Various metrics and datasets for LLM evaluation have been developed to objectively assess LLMs’ performance. Although existing evaluation metrics and datasets have widely been used across various fields, their large scale hinders their use in small organizations or by individuals. Furthermore, there is degree of discrepancy between evaluation results and actual user experiences. The study proposes evaluation metrics and datasets with relatively small amounts of data while reflecting real-world user experiences. In the process of testing the proposed metrics and datasets, the research evaluates and compares four closed-LLMs and six open-LLMs, which are latest as of April 2024. The results show that proposing datasets exhibited trends similar to existing datasets despite its smaller size, and furthermore, well reflected actual user experiences. Moreover, open-LLMs performed similar, or indeed, better than closed-LLMs in simple tasks while closed-LLMs performed significantly better in complex tasks such as mathematics, coding, and vision question-answering.

12

In this study we explore the integration of Large Language Models (LLMs), specifically GPT-4o-mini, with traditional stock market technical indicators—MACD, RSI, and Bollinger Bands—to enhance stock market prediction and decision-making frameworks. By combining quantitative analysis with qualitative insights generated by LLMs, this research demonstrates how artificial intelligence can improve the interpretability and predictive accuracy of technical trading strategies. Historical data for Tesla and Palantir, spanning six months, was analyzed with indicators calculated to identify market trends and anomalies. LLM integration provided contextual narratives that complemented technical signals, enhancing interpretability for investors. The performance evaluation revealed significant improvements in risk-adjusted returns, alpha generation, and prediction accuracy when LLM insights were incorporated into trading strategies. Key contributions include a novel methodology for structuring indicator outputs for LLM analysis, scalability across diverse stocks, and the potential for democratizing access to advanced financial analytics. Challenges such as computational complexity, data sensitivity, and the dynamic nature of financial markets are discussed, alongside opportunities for real-time adaptive models and expanded indicator integration. This research highlights the transformative potential of LLM-augmented technical analysis and offers a foundation for future innovations in AI-driven financial decision-making.

13

Analysis of Small Large Language Models(LLMs)

Yo-Seob Lee

국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 13 Number 4 2024.12 pp.155-160

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The trend of Small Large Language Models (LLMs) has been developing significantly in recent years. Lightweight LLMs are designed to operate efficiently on mobile devices or edge computing environments, and perform well even with limited resources. These models are optimized for specific domains and provide results that meet the needs of specific industries. In addition, they are easily accessible to non-developers due to their user-friendly interfaces, and are utilized in various fields. The purpose of this paper is to analyze the performance, functionality, and usability of Small Large Language Models (LLMs) to understand how they can be effectively used in various natural language processing (NLP) tasks. In particular, the key goal is to evaluate what advantages and disadvantages small models have compared to large models, and whether they can be optimized for specific tasks. Through this analysis, we aim to provide useful insights for developers and researchers in selecting and utilizing LLMs.

14

LLM의 추론 성과를 향상시키기 위한 영향요인에 관한 연구 KCI 등재

문홍기, 김영준

국제문화기술진흥원 The Journal of the Convergence on Culture Technology (JCCT) Vol.11 No.2 2025.02 pp.131-144

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

최근 기업은 생성형 AI 기술을 기반으로 부가가치 높은 새로운 사업모델을 창출하거나, 제품과 서비스의 지능 화 자동화를 통해 차별화된 경쟁력을 제고하고 기업의 내부 업무 운영 효율화를 위해 많은 노력을 기술이고 있다. 생 성형 AI 접목을 통하여 기업의 의사결정의 정확성과 스피드 향상을 통해 경영 성과 향상에 크게 기여할 것이라고 기 대하고 있다. 본 연구는 LLM을 활용하여 기업의 ERP가 제공하는 정형 데이터를 기반으로 생성형 AI의 추론 성과 에 미치는 영향 요인을 확인하기 위하여 학습방식, 학습 데이터의 품질, LLM모델 아키텍처와 하이퍼 파라미터의 최 적화 변동에 따른 실험가설을 검증하였다. 실험 연구 결과 학습방식에 있어서는 Prompting 방식 보다는 LLM Fine-Tuning 방식이 정형 데이터 학습에 좋은 성능을 보임을 확인하였다. 또한 Fine-Tuning 수행 시, 성능에 영향 을 미치는 주요 요소로서 학습 데이터의 품질, 학습 데이터 구성, 학습주기 (Epoch), LLM 파라미터 (Temperature, Top-k, Top-p)등으로 확인되었다. 본 연구는 정형데이터셋 기반의 검색 기술의 적용과 발전에 기여함과 동시에 LLM 활용 방향성을 제공함으로써 기업 운영에 LLM의 활용을 제고하는데 의미와 가치가 있다.

The recent trend among companies is to leverage generative AI technologies to create high value-added new business models. This involves enhancing differentiated competitiveness through the intelligent automation of products and services, as well as improving the efficiency of business operations. The integration of generative AI is anticipated to significantly contribute to improved management performance by enhancing the accuracy and speed of corporate decision-making. This study investigates the factors influencing the inference performance of generative AI using LLM based on structured data provided by a company's ERP system. The study tested experimental hypotheses concerning the learning method, the quality of training data, and the optimization variations of LLM model architecture and hyperparameters. The experimental results indicated that, in terms of learning methods, LLM Fine-Tuning outperformed Prompting for structured data learning. Additionally, during Fine-Tuning, key factors affecting performance were identified as the quality of training data, the composition of training data, training cycles, and LLM parameters. This research contributes to the application and development of search technologies based on structured datasets and provides direction for the utilization of LLMs, thereby enhancing the use of LLMs in corporate operations.

15

Verification of the Validity and Reliability of Therapeutic Communication Responses Generated by Large Language Models (LLMs): A Comparative Study of Prompt-Engineering Strategies

김성광, 김근면, 차선경, 정미란

[NRF 연계] 한국정신간호학회 정신간호학회지 Vol.34 2025.11 pp.23-35

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Purpose: This study evaluated the validity and reliability of therapeutic communication responses generated by GPT-5 by comparing different prompt-engineering strategies and identifying optimal methods for psychiatric nursing education. Methods: A methodological design compared four strategies?Zero-shot, Few-shot, Chain-of-Thought (CoT), and Role-prompt?applied to ten psychiatric nursing communication questions. Fourteen experts with over ten years of combined clinical and educational experience rated 40 anonymized responses using a 4-point scale. Inter-rater reliability was assessed with Fleiss' κ and Krippendorff's ? from 400 repeated outputs. Semantic consistency was examined using cosine similarity. Analyses were conducted in Python. Results: Role-prompt showed the highest content validity (mean=3.54±0.39, S-CVI/Ave=0.95), followed by CoT and Zero-shot. Inter-rater agreement was slight (κ=0.07; ?=0.15). Semantic consistency was highest for CoT (0.76±0.12) and Role-prompt (0.75±0.11). A Friedman test indicated significant differences among strategies (x2(3)=11.88, p=.008). Conclusion: Role-prompt yielded the most valid therapeutic responses, while CoT produced the most consistent outputs. These findings support the use of Role-prompt and CoT strategies to enhance the validity and reliability of therape.

16

언어, 창조, 하나님 형상 ― 인공지능 대규모 언어 모델(LLM)과 창세기 창조 언어의 학제 간 대화

김윤식

[NRF 연계] 한신대학교 신학사상연구소 신학사상 Vol.211 2025.12 pp.127-160

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 대규모 언어 모델(Large Language Models, LLM)의 등장이 인간 언어와 사고, 그리고 신학적 인간 이해에 제기하는 근본적 질문을 다룬다. 최근 LLM은 심층 신경망과 트랜스포머 아키텍처를 통해 방대한 말뭉치의 통계적 구조를 학습하고, 문맥에 따라 언어를 생성하는 능력을 보여준다. 이러한 기술적 성취는 인간 언어의 창조성, 의미 형성, 사고 과정이 상당 부분 패턴과 맥락적 연관성에 의존한다는 사실을 역설적으로 드러내며, 인간 언어의 본질을 재성찰하도록 요청한다. 이 점에서 LLM은 인간 인지의 구조를 비추는 기술적 ‘거울’로 기능한다. 그러나 동시에 LLM은 세계와의 직접적 관계성, 제도적 효력, 타자와의 상호 인격적 소통이라는 인간 언어의 핵심 요소를 자율적으로 수행하지 못한다는 점에서, 인간 언어와 인공지능 언어의 본질적 차이를 성찰하게 한다. 본 논문은 오스틴(John L. Austin)의 화행 이론과 설(John R. Searle)의 제도적 사실 이론을 이론적 틀로 삼아, LLM의 언어 생성이 수행적 발화(performatives)의 조건을 충족하는지 검토한다. 오스틴에 따르면 수행적 언어는 권위, 절차, 수납의 조건이 충족될 때 현실을 성립시키는 효력을 갖는다. 그러나 LLM은 발화의 사회적·제도적 맥락에 참여하지 못하고, 발화의 결과에 대한 책임을 지지 않으며, 공동체적 규범을 자율적으로 구성할 수 없기 때문에 수반 행위(illocutionary act)의 주체가 될 수 없다. 이에 본 연구는 LLM을 의미 생산의 주체가 아니라, 인간 해석과 책임에 의존하는 매개로 규정하고, 그 한계와 가능성을 인간 언어의 참여성·관계성·책임성이라는 특징과 대비한다. 나아가 본 연구는 이러한 화행적 관점을 창세기 1–4장의 창조 언어에 적용함으로써, 창조가 단순한 명령이나 보고가 아니라 발화–응답–현실화로 구성된 수행적 사건임을 밝힌다. 창세기 1장의 저시브(jussive)와 코호르타티브(cohortative) 발화는 피조 세계를 수동적 객체가 아니라 응답과 참여의 주체로 호명하며, “그대로 되니라”라는 반복 구절은 창조가 하나님의 발화와 피조 세계의 응답 속에서 성취되는 과정을 드러낸다. 또한 창세기 2–4장은 인간이 타자와 관계를 맺고 세계를 명명하며, 창조 행위를 통해 하나님의 형상을 책임적으로 실현해 가는 존재임을 서술한다. 이 과정은 인간이 하나님의 가능성을 현실로 수행하는 참여적·상징적 구조로 이해될 수 있다. 결론적으로 본 연구는 LLM의 언어 구조가 인간 언어와 사고의 일부 형식을 기술적으로 모형화할 수 있음을 인정하면서도, 인간 언어가 본질적으로 발화와 응답 속에서 현실을 구성하는 수행적 사건이며, 참여적·관계적·책임적 성격을 지닌다는 점을 강조한다. 이를 바탕으로 본 논문은 수행적 하나님 형상(Imago Dei) 이해를 기술 시대의 인간 행위를 성찰하기 위한 응답적 책임성에 기초한 신학적 판단 기준으로 제안하며, 인간의 언어와 기술 사용이 창조 세계와 타자를 향한 참여적 응답으로 수행되도록 분별하는 가이드라인의 필요성을 제기한다.

This study addresses the fundamental questions raised by the emergence of Large Language Models (LLMs) concerning human language, thought, and theological understandings of the human being. Recent LLMs, through deep neural networks and transformer architectures, learn the statistical structures of massive corpora and demonstrate the capacity to generate language in context. These technological achievements paradoxically reveal that human linguistic creativity, meaning formation, and cognitive processes rely to a significant extent on patterns and contextual relations, thereby prompting a renewed reflection on the nature of human language. In this sense, LLMs function as a technological “mirror” that reflects the structure of human cognition. At the same time, however, LLMs are unable to autonomously perform core elements of human language?such as direct relational engagement with the world, institutional efficacy, and inter-personal communication with others?which invites critical reflection on the essential differences between human language and artificial language. Employing John L. Austin’s speech act theory and John R. Searle’s theory of institutional facts as its theoretical framework, this article examines whether LLM-generated language fulfills the conditions of performative utterances. According to Austin, performative language brings about reality when conditions of authority, procedure, and uptake are satisfied. LLMs, however, cannot participate in the social and institutional contexts of utterance, bear responsibility for the consequences of speech, or autonomously constitute communal norms; therefore, they cannot serve as subjects of illocutionary acts. Accordingly, this study characterizes LLMs not as agents of meaning production but as mediating instruments that depend on human interpretation and responsibility, contrasting their limits and possibilities with the participatory, relational, and responsible characteristics of human language. Furthermore, this speech-act perspective is applied to the creation language of Genesis 1?4, demonstrating that creation is not merely command or report but a performative event constituted by utterance, response, and realization. The jussive and cohortative forms in Genesis 1 summon the created world not as a passive object but as a subject of response and participation, while the recurring formula “and it was so” reveals creation as a process accomplished through God’s utterance and the created world’s response. Genesis 2?4 further portrays human beings as those who enter into relationships with others, name the world, and responsibly realize the image of God through creative acts. This process can be understood as a participatory and symbolic structure in which humans enact divine possibilities in concrete reality. In conclusion, while acknowledging that the linguistic structures of LLMs can technologically model certain formal aspects of human language and thought, this study emphasizes that human language is fundamentally a performative event that constitutes reality through utterance and response, marked by participatory, relational, and responsible dimensions. On this basis, the article proposes a performative understanding of the Imago Dei as a theological criterion for discerning human action in the technological age, grounded in responsive responsibility, and calls for guidelines to ensure that human language and technological use are enacted as participatory responses toward the created world and toward others.

17

정치학 연구에서의 대형언어모델(LLM): 주제 동향과 사용 현황 분석

이인복

[NRF 연계] 한국국제정치학회 국제정치논총 Vol.65 No.3 2025.09 pp.257-328

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

2022년 11월 챗GPT 출현 이후, 대형언어모델(LLM)에 기반한 인공지능 기술은 지식 생산 방식을 근본적으로 재편할 가능성을 제기하며 활발한 논의가 일어나고 있다. 정치학 분야 역시 예외는 아니다. 한편으로는 연구나 교육의 효과성 증진을 위해 대형 언어모델을 적용하는 방식에 대한 고민이 이루어지고 있으며, 다른 한편으로는 대형언 어모델이 글을 편집하고 작성하는 것에 대한 우려도 존재한다. 본 논문은 대형언어모델의 출현과 조금더 광의의 인공지능(AI)이 정치학 지식 생산에 미치는 영향을 검토하고, 윤리적, 제도적 고려사항을 다룬 논평이다. 먼저 챗GPT 출현 이후(2023-2025) SSCI와 KCI에 색인된 LLM 관련 정치학 논문의 주제 분포를 분석했다. SSCI의 경우, 63편 분석 결과 연구방법론(22.2%), 교육(20.6%), 정치커뮤니케이션(15.9%) 순으로 실용적 적용에 중점을 두는 반면, KCI의 경우는 식별된 논문 수가 9편에 불구하며 아직 관련 연구는 탐색적 단계로 파악되었다. 나아가 논문 초록 정보에 기반한 대형언어모델 사용 탐지 분석을 통해 한국 정치학계에서 LLM 사용이 광범위하게 이루어지고 있음을 보여준다. 키워드 기반 탐지와 언어학적 특징 분석이라는 두 방법론을 통해, 영문 초록 에서는 챗GPT 특징 키워드 사용률이 2022년 대비 2023년에 약 85% 증가했으며, 국문 초록에서는 쉼표 관련 언어학적 지표들이 챗GPT 출시 시점을 기점으로 급격한 변화를 보였다. 마지막으로 LLM 사용에 따른 연구 윤리적 쟁점들을 검토하여 투명성과 가이드 라인의 필요성을 제시한다. 본 연구는 정치학 분야 LLM 연구의 전체 지형을 체계적으로 파악했으며, 이중 탐지 방법론을 통해 한국 정치학계 LLM 사용 현황을 정량적으로 제시했다는 점에서 의의가 있다.

Since the release of ChatGPT in November 2022, AI technologies based on large language models (LLMs) have sparked active debate about their potential to fundamentally reshape the way knowledge is produced. Political science is no exception. On one hand, researchers are exploring how LLMs can enhance the effectiveness of research and education; on the other, concerns have been raised about the ethical implications of their use, especially in research writing and editing. This article provides a commentary on the impact of LLMs on knowledge production in political science, addressing both ethical and institutional considerations. It begins by analyzing the topical distribution of LLM-related political science articles indexed in SSCI and KCI between 2023 and 2025, following the release of ChatGPT. Among 63 SSCI articles, practical applications dominated, with research methods (22.2%), education (20.6%), and political communication (15.9%) as the top categories. In contrast, there were only nine KCI articles, which took more exploratory approaches to the topic. The study further conducts LLM usage detection based on article abstracts, revealing widespread adoption of LLMs in Korean political science. Using keyword-based detection and linguistic feature analysis, it shows that the usage rate of ChatGPT-related keywords in English abstracts rose by approximately 85% from 2022 to 2023, while linguistic markers such as comma usage in Korean abstracts showed a sharp shift after ChatGPT’s release. Finally, the paper reviews key ethical challenges raised by LLM usage and calls for greater transparency and clearer guidelines. This research offers a comprehensive mapping of LLM-related studies in political science and provides quantitative insights into LLM usage in Korean political science through a dual-method detection approach.

18

대규모 언어모델의 초등학교 수학 문제 풀이 성능 분석: OpenAI o1, Claude 3.5 Sonnet, Gemini 1.5 Pro를 중심으로

김수진, 김용성

[NRF 연계] 한국교육정보미디어학회 교육정보미디어연구 Vol.31 No.3 2025.06 pp.823-845

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 초등학교 수학 문제 해결에서 대규모 언어모델(LLM)의 성능을 실증적으로 분석하고, 그 교육적 활용 가능성과 한계를 탐색하기 위해 OpenAI o1, Claude 3.5 Sonnet, Gemini 1.5 Pro를 대상으로 정답률, 정답 일관성, 오류 유형을 비교․분석하였다. 연구 대상은 2022 개정 교육과정의 초등학교 4학년 수학 ‘수와 연산’ 영역 문제이다. 분석 결과 OpenAI o1이 0.93으로 가장 높은 정답률을 보였고, Gemini 1.5 Pro(0.88), Claude 3.5 Sonnet(0.54) 순으로 나타났다. 표나 이미지가 결정적 정보를 포함한 문제에서는 세 모델 모두 낮은 정답률을 기록하며 시각적 정보 처리의 한계를 보였다. '곱셈과 나눗셈' 단원에서는 높은 정답률을, '큰 수' 단원에서는 낮은 정답률을 보이는 등 단원별 차이도 관찰되었다. 일관성 측면에서는 OpenAI o1이 가장 안정적이었으며, Claude 3.5 Sonnet은 동일 문제 반복 해결 시 정답이 일관되지 않는 경향을 보였다. 오류 유형 분석에서는 정보 인식 오류, 구조적 관계 해석 오류, 문맥 파악 실패 등이 나타났다. 이러한 결과는 LLM이 초등 수학 학습에서 활용 가능성을 보이지만, 문제 유형에 따른 성능 차이와 시각적 정보 해석 및 논리적 사고 과정의 한계로 인해 신중한 적용이필요함을 시사한다. 따라서 교사는 LLM의 문제 유형별 성능 차이를 인식하고 학습 지원 전략을 설계해야 한다. 또한 LLM의 다양한 풀이 표현 양식을 전략적으로 활용하고, 오류 특성을 교육적 기회로전환하는 교수․학습 전략 개발이 필요하다. 향후 연구에서는 다양한 학년과 수학 영역으로 대상을확장하고, 오류 분석을 바탕으로 LLM의 한계를 보완할 수 있는 프롬프트 설계 전략 및 수업 적용 방안을 실증적으로 검토할 필요가 있다.

This study investigates the performance of large language models (LLMs) in solving elementary mathematics problems by comparing accuracy, consistency, and error types across OpenAI o1, Claude 3.5 Sonnet, and Gemini 1.5 Pro. The dataset consists of problems from the “Number and Operations” domain within 4th-grade mathematics of the 2022 revised Korean national curriculum. Results show OpenAI o1 achieved the highest accuracy (0.93), followed by Gemini 1.5 Pro (0.88) and Claude 3.5 Sonnet (0.54). All models demonstrated lower performance on problems containing critical information in tables or images, revealing limitations in visual information processing. Performance was high in “Multiplication and Division” but low in “Large Numbers,” indicating topic-based differences. OpenAI o1 was most consistent, while Claude 3.5 Sonnet showed inconsistent answers when solving identical problems repeatedly. Error analysis revealed information recognition errors, structural relationship misinterpretation, and contextual understanding failures. These findings suggest that while LLMs show potential for elementary math learning, cautious application is needed due to performance differences across problem types and limitations in visual information interpretation and logical reasoning. Teachers should recognize LLMs' performance differences and design appropriate learning support strategies. Future research should expand to various grade levels and mathematical domains, and empirically examine prompt design strategies based on error analysis.

19

대규모 언어 모델(LLM)의 인공소통과 세미오시스- 에스포지토의 소통 이론과 퍼스 기호학의 관점에서

강병창

[NRF 연계] 한국기호학회 기호학 연구 Vol.82 2026.04 pp.7-37

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문은 대규모 언어 모델(LLM)을 ‘인공지능’이 아닌 ‘인공소통’ 체계로 재정의하고, 퍼스(C. S. Peirce) 기호학의 세미오시스 개념을 통해 그 작동 방식과 한계를 분석한다. 2022년 ChatGPT의 등장 이후 생성형 AI는 언어와 사고, 창의성의 경계를 허물며 인문학적 성찰을 요구하고 있다. 그러나 현행 AI 담론은 ‘인공지능’이라는 의인화된 명명법에 의해 인간과 기계를 동일 선상의 경쟁 관계로 설정하는 범주 오류에 빠져 있다. 본고는 이러한 대체-방어 이분법을 넘어서기 위해, 에스포지토(E. Esposito)의 ‘인공소통’ 개념과 루만(N. Luhmann)의 소통 이론을 경유하여 LLM의 매체적 본질을 규명한다. 루만적 관점에서 소통은 발신자의 의도 전달이 아니라 수신자의 관찰과 해석을 통해 완성되는 사회적 작동이다. LLM은 자신이 생성한 텍스트의 의미를 이해하지 못하지만, 인간 사용자가 해석 가능한 출력을 생성함으로써 소통에 참여한다. 이때 LLM은 인간이 남긴 방대한 데이터에 내재된 지향성을 기생적으로 활용하여 ‘가상적 이중 우발성’을 생성하며, 단순한 패턴 반복을 넘어선 기호적 행위자로 기능한다. 퍼스 기호학의 틀에서 LLM은 의식 없이도 세미오시스를 수행하는 ‘유사 마음(quasi-mind)’으로 위치 지어진다. LLM은 사용자의 프롬프트(기호)를 받아 충족 조건(대상)을 매개하는 응답(해석체)을 산출하며, 이 과정에서 훈련 데이터에 각인된 인간의 지향성에 기생하는 ‘기생적 지향성’을 드러낸다. 그러나 LLM의 세미오시스에는 결정적 한계가 존재한다. LLM은 실재 세계와의 직접적 접촉(2차성)이 결여된 채 기호와 규칙의 조작(3차성)에만 머물며, 퍼스가 강조한 비판적 자기 제어 능력―자신의 추론을 도덕적 이상에 비추어 검토하고 수정하는 능력―에는 도달하지 못한다. 본 논문은 이러한 분석을 통해 LLM과의 소통이 ‘인간 대 기계’의 대체가 아니라 세미오시스의 재배치임을 논증하며, 인공소통의 윤리와 지성이 기계의 성능이 아닌 인간의 비판적 실천과 제도적 설계에 달려 있음을 제시한다.

This paper redefines large language models (LLMs) not as “artificial intelligence” but as systems of “artificial communication,” and analyzes their operative mechanisms and inherent limitations through the concept of semiosis drawn from the semiotics of Charles Sanders Peirce. Since the advent of ChatGPT in 2022, generative AI has blurred the boundaries of language, thought, and creativity, prompting urgent humanistic reflection. Yet the prevailing AI discourse remains trapped in a category error that, through the anthropomorphizing label “artificial intelligence,” positions humans and machines as competitors on the same plane. To move beyond this replacement-versus-defense dichotomy, this study draws on Elena Esposito’s notion of “artificial communication” and Niklas Luhmann’s communication theory to clarify the medial nature of LLMs. In Luhmann’s framework, communication is not the transmission of a sender’s intention but a social operation completed through the observer’s interpretation. Although LLMs do not understand the meaning of the texts they generate, they participate in communication by producing outputs that human users can interpret. In doing so, LLMs parasitically exploit the intentionality embedded in vast human data to generate a “virtual double contingency,” functioning as semiotic agents beyond mere pattern repetition. Within the Peircean semiotic framework, LLMs are positioned as “quasi-minds” capable of performing semiosis without consciousness.

20

거대언어모델(LLM)을 활용한 한국어 학습자 쓰기 평가 가능성 탐색 연구 - GPT-4o와 GPT-4.5를 활용하여 -

이진

[NRF 연계] 한국문법교육학회 문법 교육 Vol.53 2025.04 pp.281-314

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This study explores the possibility of using Large Language Models (LLMs) as tools for evaluating Korean writing by examining the agreement between human raters and GPT models (GPT-4o, GPT-4.5). Intraclass correlation coefficients (ICCs) were computed to assess the level of agreement between scores assigned by human raters and the GPT models on essays written by Korean learners. The ICC for total scores between human raters and GPT-4o was .664, while the ICC with GPT-4.5 was .794, indicating good and excellent agreement, respectively. However, both models showed poor agreement with human raters in “sociolinguistic competence” and “linguistic diversity.” Kruskal-Wallis H tests with post-hoc analyses also revealed significant differences in scores among the three raters in most evaluation categories. Although GPT models demonstrated human-like tendencies in some areas, clear discrepancies were observed in others. These findings suggest that GPT models have considerable potential as supplementary tools in automated writing evaluation, though careful attention to evaluation domains and model behavior is required.

 
1 2
페이지 저장