Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 82
No
1

Artificial intelligence (AI) is a significant tool in modern military operations in that it helps to analyze a large volume of strategic, tactical, and operational data. On the other hand, current large language models (LLMs) like GPT-4 or Falcon have difficulty resolving problems in defense-specific contexts because of issues related to security, data confidentiality, and the lack of explainability. This document presents MilGPT, a secure and explainable LLM structure that aims at solving military problems only. To the model, fine-tuned open-source architectures with domain-specific defense datasets are integrated to elevate intelligence synthesis, decision-making, and threat prediction. On the benchmark, performance evaluation tasks show that MilGPT accounts for a 27% increase in contextual accuracy, an 18% reduction in hallucination rate, and an 33% improvement in explainability as measured by gradient-based feature attribution. In the proposed framework, military intelligence systems are not only secured but also made adaptive and humaninterpretable, thus, setting up a basis for the coming generation of AI models capable of defense-grade tasks.

2

대규모 언어 모델의 교육적 활용 방안 탐색

이영호

한국정보교육학회 정보교육연구 제1권 제3호 2023.10 pp.29-34

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 연구에서는 대규모 언어 모델의 발전 및 교육적 활용 동향을 살펴보고, 이를 바탕으로 교육적 활용 방안을 제시 하였다. 현재 대규모 언어 모델은 GPT 등의 모델이 상용화되어 사용되고 있으며, 특정 데이터에 최적화된 sLLM도 지속적으로 연구 개발되고 있다. 이러한 인공지능 기술을 교육적으로 활용하는 연구에서 그 가능성과 한계점 분석이 이루어지고 있다. 이에 본 연구에서는 대규모 언어 모델을 교육적으로 활용하기 위한 방안으로 교육용 LLM 가이드 라인 개발과 sLLM 최적화 방안을 제시하였다. 이를 통해 교육에 보다 효과적으로 LLM을 적용할 수 있을 것으로 기대된다.

This study examined the advancements and educational utilization trends of large language models (LLM), and based on this, proposed educational application strategies. Currently, LLMs, such as GPT, are being commercialized and used, and sLLM optimized for specific data is also continuously being researched and developed. The potential and limitations of using these artificial intelligence technologies in education are being analyzed in related studies. In this research, we proposed the development of educational LLM guidelines and optimization strategies for sLLM. Through these efforts, we anticipate that LLM can be applied more effectively in education.

3

4,200원

본 연구는 2022 개정 교육과정 초등학교 5~6학년군 5개 교과(국어, 수학, 실과, 체육, 영어)를 대상으로 생성형 AI 기반 초거대언어모델 프롬프트의 교육적 활용 가능 사례를 탐구하여 향후 생성형 AI 활용 교육의 방향성과 시 사점을 도출하고자 하였다. 분석 결과를 통해 교과별 목표와 성격에 따라 초거대언어모델의 교육적 활용 분야가 다르게 적용될 수 있다는 점을 알 수 있었다. 교육 현장에서 적절하게 활용되는 생성형 AI는 학습자 중심 역량 강 화를 지원할 수 있는 혁신적 교육 방안으로서의 잠재력을 가지고 있다. 생성형 AI를 올바르게 사용하기 위해서는 이러한 긍정적인 가능성과 더불어 생성형 AI의 기술적 한계와 윤리적 사용에 대해서도 충분한 이해가 필요할 것 이다. 또한 AI교육에 대한 사회적 수요와 관심이 꾸준하게 늘어나는 만큼 관련한 후속 연구 및 논의가 필요한 상황이다.

This study aimed to explore the potential educational applications of a generative AI-based large language model prompts for the revised 2022 elementary school curriculum, specifically targeting 5th and 6th grade in five subjects: Korean, Mathematics, Physical arts, Physical Education, and English. Through our analysis, we found that the educational uses of the large language model can vary depending on the goals and characteristics of each subject. Appropriately utilized, generative AI has the potential to serve as an innovative educational approach that supports learner-centered competency enhancement. However, it is necessary to have a sufficient understanding of both its positive possibilities and the technical limitations and ethical considerations associated with its use. Furthermore, given the growing social demand and interest in AI education, further research and discussions on this topic are necessary.

4

4,000원

생활SOC 복합시설의 공간 구성에서 다양한 이해관계자의 의견수렴은 효과적인 우선순위 도출의 핵심 요 소이다. 그러나 기존 주민참여 방식은 시간과 비용 소요, 참여율 저조 등 현실적 제약을 가지고 있다. 이에 본 연 구는 대규모 언어모델(Large Language Model, LLM) 기반 다중에이전트 시스템과 델파이 방법론을 결합한 생활 SOC 복합시설 공간 우선순위 도출 프레임워크를 제시하였다. 연령별 인구구조를 반영한 가상 주민 에이전트를 모델링하고, 3단계 델파이 과정을 통한 체계적 의견수렴 메커니즘을 구현하였다. 프레임워크 검증을 위해 반복 시 뮬레이션을 수행한 결과, 주민 에이전트들이 선정한 상위 핵심공간에 대해 높은 일관성을 확인하였다. 연령대별 공간 중요도 평가 분석을 통해 모든 공간에서 연령대 간 통계적으로 유의한 차이가 확인되어, LLM 기반 에이전 트가 연령별 특성을 적절히 반영한 차별화된 의견을 생성할 수 있음을 확인하였다. 본 연구는 기존 주민참여 방식 의 제약을 보완하는 새로운 방법론적 대안을 제시하고 그 기초적 유효성을 검증하였다.

Opinion convergence among diverse stakeholders is critical for Community-based SOC complex facility spatial configuration, but existing resident participation methods face time, cost, and participation constraints. This study presents a spatial priority determination framework combining Large Language Model (LLM)-based multi-agent systems with Delphi methodology. Virtual resident agents reflecting age-specific population structures were modeled and implemented through a three-stage Delphi process. Repetitive simulations confirmed high consistency for selected core spaces. Age-specific spatial importance evaluations showed statistically significant differences among age groups, verifying that LLM-based agents generate differentiated opinions reflecting age characteristics. This study presents a methodological alternative complementing existing participation limitations and verifies its fundamental validity.

5

4,200원

효과적인 사이버 위협 대응 전략을 수립하기 위해서는 다단계로 수행되는 사이버공격의 단계별 정보를 정확하게 식별할 수 있어야 한다. 그러나 사이버보안의 핵심인 침입탐지시스템은 다단계로 구성된 사이버공격을 하나의 연관된 공격으로 탐지 하지 못하고 개별적인 다수의 공격이 발생한 것처럼 경보를 생성한다. 따라서 대응은 개별 공격에 대한 단편화된 대응으로 이 어질 수밖에 없으며, 동일한 침해사고의 재발 방지를 위한 효과적인 대책을 수립하는 것도 어려워진다. 이러한 문제를 해결하 기 위해 본 연구에서는 거대 언어 모델을 기반으로 별도의 미세조정 없이 프롬프트 엔지니어링만을 활용하여 침입탐지 경보 로부터 다단계 공격을 시나리오 형태로 식별하는 기법을 제안한다. 프롬프트 엔지니어링 과정에서는 모델의 사이버 위협 분석 과 관련된 역량이 활성화되도록 페르소나를 부여하였다. 또한 논리적인 연쇄적 사고 기법을 적용하고, 환각 현상을 억제하기 위해 제약조건을 명시하였다. 제안된 기법의 효과성은 5종의 거대 언어 모델을 이용하여 검증하였으며, 웹 기반 모델인 Gemini 2.5 Pro 06-05가 가장 정확한 공격 시나리오를 식별하였다.

Developing an effective cyber threat response strategy requires accurately identifying information at each stage of a multi-stage cyber attack. However, intrusion detection system, a core element of cybersecurity, fail to detect multi-stage cyber attack as a single, correlated attack and instead generates alerts as if multiple individual attacks had occurred. This inevitably leads to fragmented responses to individual attacks. Developing effective countermeasures to prevent the recurrence of the same cyebr incident also becomes difficult. To address this problem, in this paper, we propose a novel approach for identifying multi-stage cyber attack in attack scenario from intrusion detection alerts using prompt engineering alone, without any additional fine-tuning, based on a Large Language Model(LLM). During the prompt engineering process, we assigned personas to the LLM to activate its cyber threat analysis capability. We also applied Chain-of-Thought technique to facilitate logical reasoning and specified constraints to suppress hallucinations. The effectiveness of the proposed approach was verified using five LLMs. As a result, the web-based LLM, Gemini 2.5 Pro 06-05, identified the most accurate attack scenario.

6

4,800원

연구목적: 본 연구는 뉴스 빅데이터를 활용해 한파로 인한 사회적 피해와 위험요소를 체계적으로 분류 하는 것을 목표로 한다. 또한 한파 위험이 시대별로 어떻게 보도되는지 분석하고자 한다. 연구방법: 2010년부터 2025년까지 수집된 약 1억 건의 네이버 뉴스를 대상으로 BERT 기반 모델과 거대언어모델 (GPT-4o-mini)을 활용한 4단계 정제 프로세스를 설계하였다. 단순 날씨 예보와 비유적 표현을 제거하 여 최종 1,973건의 한파 관련 뉴스를 확정하고, 최종 선별된 뉴스에서 2,829개의 위험 문장을 추출하여 WHO 기준을 바탕으로 구축한 5개 대분류와 24개 세분류 체계에 따라 위험요소를 분류·분석하였다. 연구결과: 한파 위험요소 분류 결과 취약계층(42.0%), 취약시설(33.2%) 관련 보도가 가장 높은 비중을 차지하여, 뉴스에서는 한파가 사회적 약자에게 상대적으로 큰 영향을 미치는 재난으로 다루어지고 있다. 시대별로는 2010년대에 인프라 손상 등 물리적 피해에 집중되었으나, 2020년대에는 코로나19와 결합된 복합재난 및 취약계층 보호 이슈 가 부각되었다. 결론: 한파는 단순한 기상 현상을 넘어 취약계층과 취약시설에 피해가 집중되는 복합적 위험으로 인식되고 있으며, 감염병 등 복합재난 상황에서의 특수성 또한 부각되고 있는 것으로 확인되었다. 이에 한파 피해가 특정 계층에 집중되는 특성을 고려하여, 취약계층의 수요와 상황을 반영한 ‘수요자 중심’의 정책과 복합재난 관점에서의 한파 대응 체계 마련이 필요하다.

Purpose: This study aims to systematically classify social risks associated with cold waves using news big data and to analyze how media coverage of cold wave risks has changed over time. Method: A four-stage refinement process was designed using the BERT based model and a Large Language Model (GPT-4o-mini) on approximately 100 million Naver news articles collected from 2010 to 2025. After removing simple weather forecasts and figurative expressions, 1,973 news articles specifically related to cold waves were finalized. From these, 2,829 risk-related sentences were extracted and analyzed according to a classification system of five major and twenty-four subcategories established based on WHO standards. Result: The classification of cold wave risk factors revealed that the number of new articles related to vulnerable groups (42.0%) and vulnerable facilities (33.2%) accounted for the highest proportions, indicating that news media perceives cold waves as disasters with a disproportionately large impact on the socially disadvantaged. While reports in the 2010s focused on physical damage such as infrastructure failure, the 2020s saw a shift toward highlighting complex disasters associated with COVID-19 and issues regarding the protection of vulnerable populations. Conclusion: The study confirms that cold waves are recognized as complex risks where damage is concentrated on vulnerable groups and facilities, especially under special circumstances like pandemic-driven complex disasters. Considering that cold wave damage is concentrated on specific classes, it is necessary to establish “user-centered” policies that reflect the needs and situations of vulnerable groups, alongside a cold wave response system from a complex disaster perspective.

7

4,000원

연구목적: 119 신고접수체계에 적용된 인공지능(AI) 기술의 현황과 효과를 분석하고, 실제 운영 환경에서 AI 기반 신고접수 지원 시스템이 재난 대응 효율성에 미치는 영향을 실증적으로 검토하는 데 있다. 이를 위 해 음성인식(STT), 자연어처리(NLP), 기계학습, 대규모 언어모델(LLM) 등 주요 AI 기술의 적용 동향을 고 찰하고, 서울종합방재센터를 중심으로 AI 기반 119 신고접수 시스템의 도입 사례를 분석하였다. 연구방 법: AI 콜봇 실증 운영 데이터를 활용하여 신고접수 대기시간 변화, 긴급도 판단 정확도, 비긴급 신고 선별효과를 정량적으로 분석하였으며, 동시에 AI 도입 이후 종합상황실의 업무 구조와 운영 방식 변화를 정성적으로 검토하였다. 분석결과: AI 시스 템 도입 이후 평균 신고접수 대기시간이 약 7.33% 단축되었으며, AI 콜봇을 활용한 비긴급 신고 자동 선별을 통해 상담요원의 업무 부담이 완화 되는 효과가 확인되었다. 특히 신고 폭주 상황에서도 AI 기반 초기 대응을 통해 긴급 신고 처리 집중도가 향상되었으며, 긴급도 판단 정확도 또한 초기 모델 기준 약 70% 이상의 수준을 보였다. 결론: 본 연구는 119 신고접수체계에서 AI가 단순한 자동화 수단을 넘어 접수요원의 판단을 보조 하는 협업형 의사결정 지원 기술로 기능할 수 있음을 실증적으로 제시한다. 아울러 향후 AI 기반 신고접수체계의 안정적 확산을 위해 학습데이 터 고도화, 설명가능한 AI 도입, 제도적·거버넌스 기반 마련이 필요함을 정책적·기술적 과제로 제안한다.

Purpose: This study is to analyze the current status and effects of artificial intelligence (AI) applications in the 119 emergency call-taking system and to empirically examine how an AI-based call-taking support system influences the efficiency of disaster response in real operational environments. To this end, this study reviews recent trends in key AI technologies—including speech-to-text (STT), natural language processing (NLP), machine learning, and large language models (LLMs)—and analyzes a practical application case of an AI-based 119 emergency call-taking system implemented at the Seoul Emergency Operations Center. Method: Consists of both quantitative and qualitative analyses. Quantitatively, operational data from an AI callbot pilot were analyzed to assess changes in call waiting time, the accuracy of emergency priority classification, and the effectiveness of non-emergency call triage. Qualitatively, organizational and operational changes in the emergency operations center following the introduction of AI were examined. Results: Indicate that the introduction of the AI-based system reduced the average call waiting time by approximately 7.33%. In addition, the automated pre-screening of non-emergency calls through the AI callbot significantly alleviated call-takers’ workloads. Even under conditions of heavy call surges, the AI-supported initial response improved the concentration of resources on high-priority emergency calls, while the emergency classification accuracy achieved a level exceeding 70% in the initial model implementation. Conclusion: This study demonstrates that AI in the 119 emergency call-taking system can function not merely as an automation tool but as a collaborative decision-support technology that assists human call-takers. Furthermore, the findings suggest that future advancements in AI-based emergency call-taking systems should focus on enhancing training data quality, adopting explainable AI, and establishing institutional and governance frameworks to ensure sustainable and reliable deployment in public safety services.

8

폐쇄망 기반 로컬 RAG에서 RAG 오염 공격 위협 실증 및 프롬프트 기반 방어 효과 분석 KCI 등재

김동훈, 송정환, 이수진

한국융합보안학회 융합보안논문지 제26권 제2호 2026.03 pp.105-113

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

검색 증강 생성(Retrieval-Augmented Generation, RAG)은 거대 언어 모델(Large Language Model, LLM)이 외부 데이터 베이스를 참조하여 신뢰할 수 있는 응답을 생성하도록 돕는 기술로서, 다양한 분야에서 광범위하게 활용되고 있다. 그리고 국 방과 같이 데이터 보안과 도메인 지식이 중요한 폐쇄망 환경에서는 로컬 LLM과 RAG 기술을 결합한 내부 시스템 구축이 중 요한 대안으로 부상하고 있다. 그러나 RAG는 공격자가 소수의 악성 문서를 주입함으로써 특정 질의에 대해 공격자가 의도한 오답을 생성하도록 유도하는 RAG 오염 공격에 취약하다. 이에 본 연구는 2종의 로컬 LLM과 한국어 위키피디아 데이터세트 를 기반으로 폐쇄망 환경의 RAG를 구현하고, 3가지 유형의 RAG 오염 공격 기법을 설계하여 그 위협을 실증적으로 분석하였 다. 나아가 생성 단계에서 데이터베이스의 허위 가능성을 경고하고 비판적 검증을 유도하는 프롬프트를 적용하여 공격 성공률 을 20% 수준으로 감소시켰다.

Retrieval-Augmented Generation (RAG) is a technique that enables Large Language Models (LLMs) to generate reliable responses by referencing external databases, and is widely used in various fields. In closed-network environments, such as defense area where data security and domain knowledge are crucial, building internal systems that combine local LLM and RAG techniques is emerging as a viable alternative. However, RAG systems are vulnerable to poisoning attacks, wherein an adversary injects a small fraction of malicious documents to manipulate the model into generating intended incorrect responses for specific queries. Consequently, this study implemented a closed-network RAG environment utilizing two local LLMs and Korean Wikipedia datasets, and empirically analyzed the associated threats by designing three distinct RAG poisoning attack techniques. Furthermore, by applying prompts that alert the model to potential inaccuracies within the database and encourage critical verification during the generation phase, the attack success rate was successfully reduced to 20%.

9

대규모 언어 모델의 발전은 외부 도구 그리고 API와의 상호작용 방식을 크게 변화시켰다. 대표적인 접근법으로는 Tool-Calling, MCP(Model Context Protocol), RAG-MCP 세 가지 접근법으로 구분된다. 그러나 도구 수가 증가할수록 프롬프트 팽창(prompt bloat)이나 선택 정확도 저하와 같은 구조적 한계가 나타난다. 본 연구에서는 BFCL(Berkeley Function Calling Leaderboard) 데이터셋을 기반으로 구축된 데이터셋에 LLaMA3-8b- Instruct 모델을 적용하여 세 접근법을 비교하였다. 특히 도구 집합의 규모를 50, 100, 150개로 변화시키며 호출 정확도, 파라미터 삽입 정확도, 토큰 소비량, 지연 시간을 측정하였다. 이러한 측정 결과를 바탕으로 각 방식의 성능 및 비용상의 특성을 규명하고, LLM(Large Language Model) 기반 시스템 설계 시 도구 호출 방식의 합리적인 선택을 위한 정량적 기준을 제시하고자 한다.

The development of large language models has significantly changed the way they interact with external tools and APIs. Representative approaches can be classified into three categories: Tool-Calling, Model Context Protocol (MCP), and RAG-MCP. However, as the number of tools increases, structural limitations such as prompt bloat and a decline in selection accuracy emerge. This study compares the three approaches by applying the LLaMA3-8b-Instruct model on the Berkeley Function Calling Leaderboard (BFCL) dataset. In particular, We varied the toolset size to 50, 100, and 150 tools, and measured tool invocation accuracy, parameter filling accuracy, token consumption, and latency. Through this analysis, we identify the performance and cost characteristics of each method and provide quantitative criteria for rationally selecting tool integration strategies in LLM-based system design.

10

OpenAlex를 이용한 모빌리티 분야 기술 트렌드 분석 KCI 등재

이호, 박준우, 정시교, 윤일수

한국ITS학회 한국ITS학회논문지 제25권 제1호 통권123호 2026.02 pp.75-98

※ 기관로그인 시 무료 이용이 가능합니다.

6,100원

급격한 인공지능의 부상 등으로 급변하고 혼란스러운 상황에서 모빌리티 산업 분야에서 미 래에 사용될 기술에 대한 예측의 중요성이 더욱 부각되고 있으나, 소수 전문가의 주관적 판단 에 의존하는 기존 분석 방식에는 한계가 존재한다. 본 연구는 이러한 한계를 극복하기 위해, 대규모 학술 데이터베이스인 OpenAlex를 활용하여 방대한 학술 문헌 데이터를 자동으로 채굴, 클러스터링, 평가하는 정량적·재현가능한 기술 트렌드 분석 방법론을 제안한다. 제안된 방법론 은 거대 언어 모델 기반 데이터 추출, BERTopic을 활용한 계층적 토픽 모델링을 통한 기술 구 조 구축, 그리고 Prophet 모델과 boston consulting group(BCG) 매트릭스 변용을 통한 미래 성장 잠재력 예측 및 기술 유형을 자동 분류(부상, 성장, 성숙, 쇠퇴)하는 절차로 구성된다. 이 연구 는 데이터 기반의 객관적인 기술 탐색과 미래 예측을 통한 선제적 R&D 전략 수립의 기틀을 마련하고 급변하는 기술 환경에서 효과적인 R&D 전략 수립을 위한 이론적 토대를 제공할 수 있을 것으로 기대된다.

Future technology prediction is critical in the rapidly changing mobility industry. Existing analysis methods relying on subjective judgment have limitations. This study proposes a quantitative, reproducible technology trend analysis methodology using vast academic literature data from OpenAlex. The methodology involves LLM-based data extraction, hierarchical topic modeling via BERTopic to construct technology structures, and adapting the Prophet model and boston consulting group(BCG) matrix to classify technology types (emerging, growing, mature, declining). This research provides a foundation for data-driven, objective technology exploration and preemptive R&D strategy establishment in a dynamic environment.

11

4,800원

This study investigates the performance of a state-of-the-art Large Language Model(LLM) in classifying Korean neutral sentiment without additional fine-tuning and proposes an effective prompt design to improve neutral sentiment classification. To enhance the accuracy of neutral sentiment detection, we introduce a prompt based on Aspect-Based Sentiment Analysis(ABSA) and conduct a comparative evaluation using the proposed approach. The proposed prompt consists of a five-step procedure, including aspect identification and sentiment ratio–based classification, enabling more fine-grained sentiment reasoning. Experimental results demonstrate that the proposed prompt significantly improves classification accuracy when applied to the GPT-4-turbo model, thereby validating the effectiveness of prompt-based control for neutral sentiment analysis in Korean.

12

4,000원

13

6,100원

본 연구는 외국인 한국어 학습자의 작문에서 나타나는 오류를 인간 교 정자와 거대언어모델의 교정 양상을 비교ㆍ분석하였다. 이를 통해 현 재 LLM의 한국어 자동 교정 도구로서의 활용 가능성을 평가하고자 하였다. 연구 자료는 국립국어원의 <한국어학습자말뭉치>에서 수집 된 고급 학습자 작문 100개였으며, 교정 결과는 문장 단위로 비교하여 코사인 유사도를 산출하였다. 분석 결과, LLM의 교정 결과와 인간 교 정자의 교정 결과는 평균 0.8의 유사도를 보였으며, 전체 문장의 약 40%는 완전히 일치하였다. 그러나 문체와 화계, 피동 구문, 문법 구조, 연어 생성 등에서 차이가 나타났다. LLM은 문체를 일관되게 하지 못 하거나, 한국어의 특정 문법 구조를 올바르게 처리하지 못하는 경향을 보였으나, 연어 생성에서는 상대적으로 더 자연스러운 표현을 제공하 는 경우도 있었다. 본 연구는 LLM의 한국어 교정 능력에 대한 연구로 서 의의가 있으며 향후 연구에서는 다양한 학습자 변인과 데이터의 확 장을 통한 심층 분석을 제언한다.

This study compares and analyzes the error correction patterns in the writing of Korean language learners between human correctors and large language models (LLMs). The primary goal is to evaluate the feasibility of utilizing LLMs as automatic Korean error correction tools. The research data comprises 100 writing samples from advanced learners, sourced from the National Institute of Korean Language's Korean Learners Corpus. The correction results were compared sentence by sentence, and cosine similarity was calculated to measure the alignment. The findings indicate an average similarity score of 0.8 between LLM corrections and human corrections, with around 40% of the sentences showing complete agreement. However, notable differences were observed in terms of stylistic consistency, speech levels, passive constructions, grammatical structures, and collocation generation. LLMs exhibited inconsistencies in maintaining style and handling specific grammatical structures in Korean, yet in some cases, they provided more natural collocations. This research holds significance as an evaluation of the current capabilities of LLMs in Korean error correction. Future studies are suggested to expand on learner variables and data to provide deeper insights.

14

6,100원

[연구목적] 2022년 연말 인공지능 언어모델인 OpenAI의 ChatGPT가 처음 출시된 이후 새로운 버전의 모델 이 계속 출시되고 있다. 또한 범용 인공지능 언어모델을 특수한 전문분야에 활용하려는 시도도 이루어지고 있다. 본 연구는 최신 인공지능 언어모델을 회계분야에 적용하고, ChatGPT 3.5 버전과 비교하여 개선된 최근 모델의 문제해결 성능을 평가하였다. [연구방법] ChatGPT 3.5 버전을 회계분야 문제해결에 적용한 선행연구에서 보고한 한계점을 바탕으로 2024 년 8월 현재 최신 모델인 ChatGPT 4o의 성능 개선여부를 평가하였다. 특히 회계분야 문제해결 측면에서 최신 인공지능 언어모델의 성능을 평가하였다. [연구결과] 최신 모델인 OpenAI의 ChatGPT 4o는 기존의 한계사항으로 지적되었던 후입선출방법을 적용한 재고자산 원가 배분, 채권투자시의 유효이자율법을 적용한 회계처리, 간접법 방식의 영업활동현금흐름 도출, 간접비 배부에서의 단계배부법 및 상호배부법 등의 과제를 모두 올바르게 해결하는 것으로 나타났다. 다만, 인터넷 실시간 검색기능이 가능함에도 개정된 최신 세법을 적용하거나, 특정 기업의 최근 재무정보를 조회하는 경우 불안정하거나 부정확한 답변을 제시하는 한계가 존재하였다. [연구의 시사점] 범용형 인공지능 모델의 성능이 개선됨에 따라, 회계분야의 문제 해결 성능도 함께 향상된 것으로 나타났다. 회계 분야에 인공지능 모델을 활용하려는 다양한 시도와 함께 새로운 활용 가능성도 지속적으로 확대될 것으로 예상된다.

[Purpose] OpenAI's ChatGPT was first introduced in late 2022. The company has been released new versions of the model with improvements. In addition, a wide range of attempts has been made to take advantage of generative AI models for specialized fields. This study uses the latest ChatGPT 4o and evaluates its performance, which has been improved since ChatGPT 3.5. [Methodology] Based on the limitations reported in previous studies using ChatGPT 3.5, this study evaluated whether the performance was improved by utilizing ChatGPT 4o, the latest model as of August 2024. In particular, the performance was evaluated with respect to the accounting area. [Findings] The OpenAI's ChatGPT 4o correctly solved issues such as cost allocation of inventory assets using the last-in, first-out method, accounting for bond investments using the effective interest rate method, determining cash flows from operating activities using the indirect method, and allocation of overhead expenses using the step method and reciprocal method. However, there were limitations in applying the latest revised tax laws or providing financial information of a specific company with the Internet search. [Implications] As the performance of general-purpose AI models has improved, the model will be applied to the accounting. It is expected that the application of AI model to various accounting fields will continue, which will open up new possibilities.

15

인간과 유사한 AI 모델의 언어처리 - 심리언어학적 관점에서 바라보기 - KCI 등재

윤홍옥, 이은경, 송상헌

제주대학교 인문과학연구소 인문학연구 제37집 2024.07 pp.1-28

※ 기관로그인 시 무료 이용이 가능합니다.

6,700원

이 논문은 거대언어모델(Large Language Model) 문장처리와 인간 문장처리의 양상을 비교하여, 이 두 유형 간 처리결과의 유사점과 차이점을 심리언어학적 관점에서 고찰해 보는 것을 목표로 한다. 기대 기반의 문장처리 관점에 따르면, 인간 처리자는 주어진 맥락 정보에서 언어적 단서를 적극적으로 추출하여 앞으로 등장할 단어를 확률적으로 예측하고, 이를 문장처리에 즉시 반영한다고 한다. 이런 인간 처리자의 예측 기반의 확률적 행동은 거대언어모델의 작동 원리와 기본적인 측면에서 공유하는 점이 많다. 이 논문에서는 기대 기반의 인간 문장처리와 거대언어모델의 시뮬레이션을 다면적으로 비교하고자 하였다. 이를 위해, 영어 도구 논항을 연구 대상으로 하였고, 도구 명사가 논항(argument)의 위상을 가졌을 때와 부가어(adjunct)로 사용될 때 관찰된 인간 행동의 시뮬레이션을 처리 (processing)와 생성(generation)의 영역에서 시도했다. 거대언어모델은 도구 논항을 결정하 는 동사유형 효과를 성공적으로 모사하였다. 도구 명사를 생성하는 과제에서는 인간보다 도구명사 생성률이 낮아서 소극적 생성이라고 할 수 있었으나, 거대언어모델이 생성한 도구 명사와 인간이 생성한 도구 명사의 의미적 유사성은 낮지 않았다. 이 논문에서 소개된 일련의 언어모델 시뮬레이션은 인간의 행동과 대체로 유사하다고 판단되며, 이 결과는 거대언어모델이 인간 언어처리 행동을 예측하는 도구로 사용될 수 있다는 견해를 지지한다.

This paper examines the similarities and differences between Large Language Models (LLMs) and humans in sentence comprehension and generation. In psycholinguistics, expectation-based approaches suggest human processors use linguistic cues from contexts to predict upcoming words before encountering them. Our study compares the predictive sentence processing of humans and LLMs across several domains. Using English instrument arguments and adjuncts in prepositional phrases within declarative sentences, we present findings indicating that LLMs effectively simulate human behaviors in processing instrument nouns: The LLMs successfully demonstrated the significant role of verb type in distinguishing instrument arguments from adjuncts. Although LLMs were more passive than humans in generating instruments from sentence fragments, the semantic features of LLM-generated instruments closely matched those generated by humans. Our results suggest that LLMs are fairly successful in simulating human sentence processing, supporting their use as a pratical tool for predicting human processing behaviors.

16

대형 언어 모델의 문화적 편향 측정

정가연, 강채안, 김민선, 노소현, 최혜지, 김한샘

한국코퍼스언어학회 Corpus Linguistics Research Vol. 9 No. 1 2024.06 pp.111-137

※ 기관로그인 시 무료 이용이 가능합니다.

6,600원

This study analyzes cultural biases in major large language models from the United States, South Korea, and China (GPT-4, CLOVA X, and Qwen1.5) through story generation tasks using culture-specific names. Morphological analysis of the generated stories revealed that all models exhibited certain cultural biases. GPT-4 did not show negative biases toward Korean and Chinese cultures but tended to prefer traditional and rural settings when describing these cultures. In contrast, CLOVA X and Qwen1.5, which are specialized for their respective national languages, portrayed their own cultures in modern and positive terms while using a relatively higher proportion of negative adjectives and unrealistic settings when describing Western contexts. These findings are significant because they go beyond the conventional focus on biases in Western-centric models toward non-Western contexts. They newly reveal that East Asian-based models can also exhibit similar biases when representing Western cultures. This research suggests that current language model has fundamental limitations in achieving cultural neutrality and highlights the importance of balanced learning and reflection of diverse cultural contexts as a crucial challenge in language model development.

17

4,000원

18

7,500원

최근의 인공지능 기술 발전, 특히 생성형 인공지능 기술의 성장은 문화예술경영 분야에서도 혁신적인 변화를 가 져올 것으로 예상된다. 본 연구는 인공지능 거대언어모형, 즉 Large Language Model (LLM)이 문화예술경영에 어떻 게 활용될 수 있는지 고찰하고 이를 통해 문화예술기관들이 LLM이 가져올 변화에 효과적으로 대응할 수 있도록 돕 고자 한다. 본 논문은 먼저 LLM 기술의 본질을 소개하고 그 주요 장점인 자연어의 이해 및 생성, 프로그래밍 지원 기능 등을 살펴본다. 그 다음 이를 문화예술경영의 여러 분야, 즉 관객 개발, 기금 모금, 관객 체험 향상, 창작 프로세 스 등에 활용할 수 있는 가능성을 탐구하며 그에 따른 예시들을 제시한다. 아울러 LLM 기술의 가능성과 한계, 발전 방향에 대한 토론도 제시한다. 결론적으로 이 연구는 문화예술경영 분야에서도 LLM과 같은 인공지능 기술의 활용이 필수적이며, 그러한 기술이 가져오는 변화에 적극 대응함으로써 문화예술기관들이 업무 효율성 및 창의적 역량을 크 게 향상시킬 수 있음을 보인다.

The recent advancements in artificial intelligence technology, particularly the growth of generative AI, are expected to bring about innovative changes in the field of arts and cultural management. This study examines how Large Language Models (LLMs) can be used in arts and cultural management, aiming to help cultural institutions proactively respond to the changes brought about by LLM technology. This paper first seeks to understand LLM technology and explores its main advantages, such as natural language understanding and generation, and coding support. Next, it investigates the potential for applying these advantages to various areas of arts and cultural management, including audience development, fundraising, audience experience enhancement, and the creative process, presenting concrete examples accordingly. Furthermore, the paper provides a discussion on the possibilities, limitations, and future directions of LLM technology based on the analysis. In conclusion, this study demonstrates that the utilization of AI technologies like LLMs is essential in arts and cultural management, and that cultural institutions can significantly improve their operational efficiency and creative capabilities by proactively adapting to the changes these technologies bring.

19

Fine-Tuning the Llama2 Large Language Model Using Books on the Diagnosis and Treatment of Musculoskeletal System in Physical Therapy

김준희

[NRF 연계] KEMA학회 Journal of Musculoskeletal Science and Technology Vol.8 No.2 2024.12 pp.65-73

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Background In physical therapy, accurate information on musculoskeletal diagnosis and treatment is essential for patient care. Generative language models use machine learning to perform tasks like text generation and question answering, mimicking human language understanding. Purpose This study aimed to fine-tune the Llama2 language model using text data from books on the diagnosis and treatment of the musculoskeletal system in physical therapy, and to compare it to the base model to evaluate its usability in medical fields. Study design Technical evaluation study Methods The Llama2-13B Chat model was fine-tuned using text from books on musculoskeletal diagnosis and treatment in physical therapy and trained on 1.2 million tokens. Responses were compared to the base model based on relevance and specialization to clinical questions. Results Compared to the base model, the fine-tuned model consistently generated answers specific to the musculoskeletal system diagnosis and treatment, demonstrating improved understanding of the specialized domain. Conclusions The model fine-tuned for musculoskeletal diagnosis and treatment books provided more detailed information related to musculoskeletal topics, and the use of this fine-tuned model could be helpful in medical education and the acquisition of specialized knowledge.

20

LLM(Large Language Model) 속성과 성능 연관성 연구 KCI 등재

임철홍

한국EA학회 정보화연구 제20권 3호 2023.09 pp.257-266

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

OpenAI의 ChatGPT는 기존 기술보다 놀랍게 향상된 기능을 보여주었고 생성형 AI에 대한 활 발한 연구개발과 활용이 시작되었다. 컴퓨터 성능의 비약적인 발전으로 거대한 크기의 인공지능 언어 모델(LLM)을 활용하게 되었다. 본 연구는 LLM을 구축할 때 성능을 평가할 방법과 성능을 높이기 위 해 LLM의 구축 모든 과정에 걸쳐 고려해야 할 요인을 제시한다. 원시 데이터로부터 학습하여 backbone 모델을 만드는 것부터 만들어진 모델에 fine-tuning을 통해서 최종 모델을 구축하는 과정을 포함한다. 학습 데이터의 양과 품질, 모델 및 파라미터크기, 학습횟수 및 fine-tuning에 관련된 방안을 제시한다. 영문모델은 크기를 줄이면서 성능을 높이는 성과가 있는 반면에 한국어 모델은 작고 성능이 우수하면 서 연구개발과 서비스에 제한 없이 활용 가능한 오픈소스 모델의 구축과 보급이 필수적이다.

OpenAI's ChatGPT showed remarkable improvements over existing technologies, and active research and development and utilization of generative AI began. Advances in computing power have led to the use of large-scale artificial intelligence language models (LLMs). This study presents methods to evaluate performance when building an LLM and factors to be considered throughout the entire process of building an LLM to increase performance. It includes the entire process of building a model by learning from raw data and fine-tuning the created backbone model. We present how to improve performance from considering the quantity and quality of learning data, model and parameter size, learning frequency, and fine tuning. When it comes to the English dataset, there have been achievements in improving performance while reducing the model size, while in the case of Korean, it is essential to build and disseminate an open source model that is small but has excellent performance and can be used without restrictions for research and development and services.

 
1 2 3 4 5
페이지 저장