Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 494
No
1

대화시스템에서 자연어 생성은 대화관리 단계에서 결정한 시스템 발화의 의미표현을 사람이 이해할 수 있는 자연어 로 생성하는 것이다. 기존의 자연어 생성 연구는 의미표현에 대하여 매우 제한된 종류의 발화만을 생성하거나 문법 적으로 불완전한 발화를 생성한다는 문제점이 있다. 그래서 본 논문에서는 문제점들을 동시에 처리하기 위하여 Long Short Term Memory 기반의 언어모델을 이용한 한국어 자연어 생성 모델을 제안한다. 특히 우리는 시스템 발화의 다양성과 문법적 정확성을 높이기 위하여 빔서치 디코딩을 적용한다. 실험은 어절, 형태소, 음절단위에 따라 개별적으로 진행하였으며, 생성한 문장들은 정량적, 정성적 평가를 모두 진행하였다. 그 결과 형태소 단위로 학습한 제안모델에 빔서치 디코딩을 적용한 방법은 가장 좋은 성능을 보였다. 실제로 해당 생성 문장은 정량평가 결과에서 BLEU 지표는 0.86, Slot Error Rate 지표는 0.03을 기록하였으며 정성평가 역시 문법적으로 정확하고 문맥적 으로 충분히 자연스러운 결과임을 확인하였다.

Natural language generation in the dialogue system is a task that transforms the semantic frame of the system utterance determined in the dialogue management phase into a natural language that can be understood by humans. Existing studies have still faced some obstacles in that only very limited types of utterances or grammatically incomplete ones are generated from the semantic frames. In order to address these issues simultaneously, we propose a Korean natural language generation model using a long short term memory based language model. In particular, we exploit the beam search decoding method to obtain system utterances with diverse structures and grammatical correctness. The experiments were conducted individually with respect to the word, morpheme, and syllable units, and the generated utterances were evaluated in both quantitative and qualitative ways. As a result, the morpheme-based model with the beam search decoding has achieved the most robust result of all. In fact, in the quantitative evaluation result of the generated sentence, the BLEU-4 score was 0.86 and the SER was 0.03, and the qualitative evaluation was also confirmed to be grammatically correct and contextually natural.

2

언어모델 기반 계층화 분석법을 활용한 산업제어시스템 에이전트 도입 타당성 분석 연구 KCI 등재

김가경, 김준석, 엄익채

한국EA학회 정보화연구 제23권 1호 2026.03 pp.15-24

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

산업제어시스템은 안전과 직결되어 인간 운전자의 지속적인 감시가 요구되는 운영 구조를 가짐 에 따라, 인력에 대한 높은 의존도, 인적 오류에 따른 안전사고, 경제적 비효율성 문제 등을 초래할 수 있다. 이를 완화하기 위한 대안으로 인공지능 기반 지능형 에이전트의 도입이 주목받고 있지만, 인공지 능 고유의 불확실성과 복잡성으로 인한 위험이 우선적으로 해결되어야 할 필요성이 존재한다. 이와 같 이 산업제어시스템 에이전트 도입에 대한 타당성 분석이 요구됨에 따라, 본 연구에서는 언어모델을 활 용한 계층화 분석 방안을 제시하고자 한다. 제안하는 언어모델 기반 계층화 분석 방안은 기존의 타당성 분석 방법론과 달리, 학습된 언어모델을 전문가 개체로 활용하였다는 점에서 차별점이 존재한다. 학습 된 언어모델을 제안하는 방법론에 활용한 결과, 지능형 에이전트의 자율성 수준에 따라 의미적 유사도 가 상이한 결과를 식별하였다. 이는 언어모델이 타당성 분석을 위한 초기 또는 보완적 자료로써 작용할 수 있음을 의미한다.

To ensure human safety, industrial control systems are structured in a way that requires continuous monitoring by human operators. This operational model leads to high dependence on specialized personnel and results in inefficiencies in labor costs. As an alternative, the introduction of intelligent agent–based operational technologies has gained attention; however, the inherent uncertainty and complexity of artificial intelligence pose significant risks. In this context, this study proposes a hierarchical analysis approach using language models to evaluate the feasibility of deploying intelligent agents for industrial control system operations. The key contribution of the proposed approach lies in leveraging a domain-specialized, pre-trained language model as an expert entity. Applying the proposed methodology to trained language models demonstrated that they can produce differing semantic similarity scores depending on the autonomy level of intelligent agents. This experiment confirmed that language models can serve as one clue for feasibility analysis.

3

GRU 언어 모델을 이용한 Fuzzy-AHP 기반 영화 추천 시스템 KCI 등재

오재택, 이상용

한국디지털정책학회 디지털융복합연구 제19권 제8호 2021.08 pp.319-325

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

무선 기술의 고도화 및 이동통신 기술의 인프라가 빠르게 성장함에 따라 AI 기반 플랫폼을 적용한 시스템이 사용자의 주목을 받고 있다. 특히 사용자의 취향이나 관심사 등을 이해하고, 선호하는 아이템을 추천해주는 시스템은 고도화된 전자상거래 맞춤형 서비스 및 스마트 홈 등에 적용되고 있다. 그러나 이러한 추천 시스템은 다양한 사용자들의 취향이나 관심사 등에 대한 선호도를 실시간으로 반영하기 어렵다는 문제가 있다. 본 연구에서는 이러한 문제를 해소하 기 위해 GRU(Gated Recurrent Unit) 언어 모델을 이용한 Fuzzy-AHP 기반 영화 추천 시스템을 제안하였다. 본 시스 템에서는 사용자의 취향이나 관심사를 실시간으로 반영하기 위해 Fuzzy-AHP를 적용하였다. 또한 대중들의 관심사 및 해당 영화의 내용을 분석하여 사용자가 선호하는 요인과 유사한 영화를 추천하기 위해 GRU 언어 모델 기반의 모델을 적용하였다. 본 추천 시스템의 성능을 검증하기 위해 학습 모듈에서 사용된 스크래핑 데이터를 이용하여 학습 모델의 적합성을 측정하였으며, LSTM(Long Short-Term Memory) 언어 모델과 Epoch 당 학습 시간을 비교하여 학습 수행 속도를 측정하였다. 그 결과 본 연구의 학습 모델의 평균 교차 검증 지수가 94.8%로 적합하다는 것을 알 수 있었으며, 학습 수행 속도가 LSTM 언어 모델보다 우수함을 확인할 수 있었다.

With the advancement of wireless technology and the rapid growth of the infrastructure of mobile communication technology, systems applying AI-based platforms are drawing attention from users. In particular, the system that understands users' tastes and interests and recommends preferred items is applied to advanced e-commerce customized services and smart homes. However, there is a problem that these recommendation systems are difficult to reflect in real time the preferences of various users for tastes and interests. In this research, we propose a Fuzzy-AHP-based movies recommendation system using the Gated Recurrent Unit (GRU) language model to address a problem. In this system, we apply Fuzzy-AHP to reflect users' tastes or interests in real time. We also apply GRU language model-based models to analyze the public interest and the content of the film to recommend movies similar to the user's preferred factors. To validate the performance of this recommendation system, we measured the suitability of the learning model using scraping data used in the learning module, and measured the rate of learning performance by comparing the Long Short-Term Memory (LSTM) language model with the learning time per epoch. The results show that the average cross-validation index of the learning model in this work is suitable at 94.8% and that the learning performance rate outperforms the LSTM language model.

4

4,000원

본 연구의 목적은 영어 말하기 교육에 GPT-3.5 기반 웹 어플리케이션, GPTalk의 적용 가능성 살펴보기 위한 사례연구를 수행하는 것이다. 연구 대상은 B 초등학교 3학년에서 6학년까지의 학생 86명과 영어 교사 3명이다. 연구 수행을 위해 GPTalk를 활용한 수업 내용의 관찰, 성찰일지 분석, 그리고 반구조화된 인터뷰를 실시하였다. 연구 결과, GPTalk를 활용하여 수업에 참여한 학생들은 다양한 대화 상황과 주제에 대한 말하기 능력이 향상되 었다고 보았다. 그러나 문법적 피드백이 부족하고, 사용자경험 측면에서의 한계도 확인하였다. 교사들은 GPTalk 의 활용 방안에 대해 개별 학생의 필요에 따라 조절 가능하다는 의견을 공유했으며, 데이터 분석을 통한 학습 지원의 필요성을 지적하였다. 본 연구를 통해 초등학교 수준에서의 GPT-3.5 기반 언어모델의 교육적 적용에 대 한 가능성과 한계를 밝힘과 함께 GPT-3.5 기반의 기술적 접근을 통한 영어 말하기 교육의 방향을 탐색하였다.

The purpose of this study is to conduct a case study examining the feasibility of implementing GPT-3.5-based web application, GPTalk, in English speaking education. The subjects of this research consist of 86 students ranging from the 3rd to 6th grades at B Elementary School and English teachers. To carry out the study, various methodologies including classroom observations using GPTalk, reflective journal analyses, and semi-structured interviews were employed. The findings suggest that students who engaged in lessons using GPTalk displayed improved speaking abilities across diverse conversational situations and topics. However, the study also identified limitations, such as the lack of grammatical feedback and issues related to the user experience. Teachers expressed the opinion that the application of GPTalk could be tailored according to the individual needs of students, and underscored the necessity for data-driven educational support. Through this research, we illuminate the potential and limitations of applying GPT-3.5-based language models in elementary school English education, aiming to explore future directions for English speaking education through technological approaches grounded in GPT-3.5.

5

4,000원

작성된 문서의 문법적 오류 교정을 할 때 맞춤법 검사기를 사용하는 것이 일반적이다. 그러나 초등학생들이 작성한 글 중에는 문법적으로는 옳더라도 자연스럽지 않은 문장이 있을 수 있다. 본 논문에서는 동일한 의미를 가진 2개의 문장이 주어졌을 때, 어떤 것이 더 자연스러운 문장인지 자동 판별할 수 있는 방법을 소개한다. 이 방법은 순환 신경망(recurrent neural network)을 이용하여 장기 의존성(long-term dependencies) 문제를 해결하 고, 보조 단어(subword)를 사용하여 희소 단어(rare word) 문제를 해결한다. 약 200만 문장의 단일어 코퍼스를 통해 순환 신경망 기반 언어 모델을 학습하였다. 그 결과, 초등학생들이 주로 틀리는 표현들과 그에 대응하는 올 바른 표현을 입력으로 주었을 때, 모든 경우에 대해 자연스러운 표현을 자동으로 선별할 수 있었다. 본 소프트웨 어가 스마트 기기에 사용될 수 있는 형태로 구현된다면 실제 초등학교 현장에서 활용 가능할 것으로 기대된다.

We often use spellcheckers in order to correct the syntactic errors in our documents. However, these computer programs are not enough for elementary school students, because their sentences are not smooth even after correcting the syntactic errors in many cases. In this paper, we introduce an automated method for evaluating the smoothness of two synonymous sentences. This method uses a recurrent neural network to solve the problem of long-term dependencies and exploits subwords to cope with the rare word problem. We trained the recurrent neural network language model based on a monolingual corpus of about two million English sentences. In our experiments, the trained model successfully selected the more smooth sentences for all of nine types of test set. We expect that our approach will help in elementary school writing after being implemented as an application for smart devices.

6

양방향 순환 신경망 언어 모델을 이용한 Fuzzy-AHP 기반 영화 추천 시스템 KCI 등재

오재택, 이상용

한국디지털정책학회 디지털융복합연구 제18권 제12호 2020.12 pp.525-531

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

다양한 정보가 대량으로 유통되는 IT 환경에서 사용자의 요구를 빠르게 파악하여 의사결정을 도와줄 수 있는 추천 시스템이 각광을 받고 있다. 그러나 현재 추천 시스템은 사용자의 취향이나 관심사가 바뀌었을 때 선호도가 즉시 시스템에 반영이 되지 않을 수가 있으며, 광고 유도로 인하여 사용자의 선호도와 무관한 아이템이 추천될 수가 있다는 문제점이 있다. 본 연구에서는 이러한 문제점을 해결하기 위해 양방향 순환 신경망 언어 모델을 이용한 Fuzzy-AHP 기반 영화 추천 시스템을 제안하였다. 본 시스템은 사용자의 취향이나 관심사를 명확하고 객관적으로 반영하기 위해 Fuzzy-AHP를 적용하였다. 그리고 사용자가 선호하는 영화를 예측하기 위해 양방향 순환 신경망 언어 모델을 이용하여 실시간으로 수집되는 영화 관련 데이터를 분석하였다. 본 시스템의 성능을 평가하기 위해 그리드 서치를 이용하여 전체 단어 집합의 크기에 대한 학습 모델의 적합성을 확인하였다. 그 결과 본 시스템의 학습 모델은 전체 단어 집합의 크기에 따른 평균 교차 검증 지수가 97.9%로 적합하다는 것을 확인할 수 있었다. 그리고 본 모델은 네이버의 영화 평점 대비 평균 제곱근 오차가 0.66, LSTM 언어 모델은 평균 제곱근 오차가 0.805으로, 본 시스템의 영화 평점 예측성이 더 우수 함을 알 수 있었다.

In today’s IT environment where various pieces of information are distributed in large volumes, recommendation systems are in the spotlight capable of figuring out users’ needs fast and helping them with their decisions. The current recommendation systems, however, have a couple of problems including that user preference may not be reflected on the systems right away according to their changing tastes or interests and that items with no relations to users’ preference may be recommended, being induced by advertising. In an effort to solve these problems, this study set out to propose a Fuzzy-AHP-based movie recommendation system by applying the BRNN(Bidirectional Recurrent Neural Network) language model. Applied to this system was Fuzzy-AHP to reflect users’ tastes or interests in clear and objective ways. In addition, the BRNN language model was adopted to analyze movie-related data collected in real time and predict movies preferred by users. The system was assessed for its performance with grid searches to examine the fitness of the learning model for the entire size of word sets. The results show that the learning model of the system recorded a mean cross-validation index of 97.9% according to the entire size of word sets, thus proving its fitness. The model recorded a RMSE of 0.66 and 0.805 against the movie ratings on Naver and LSTM model language model, respectively, demonstrating the system’s superior performance in predicting movie ratings.

7

생성형 AI 기반 언어모델 성능 최적화를 위한 학습데이터 구축 전략 KCI 등재

주영진, 엄정호

한국융합보안학회 융합보안논문지 제25권 제1호 2025.03 pp.217-224

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

생성형 언어모델은 최근 AI 기술의 발전과 함께 다양한 산업 분야에서 혁신을 주도하고 있다. 모델이 최적의 성능을 발휘하기 위해서는 충분한 양의 고품질 학습데이터 확보가 필수적이며, 이는 모델의 일반화 성능을 향상시키고 신뢰성 높은 결과를 제공하는 데 중요한 역할을 한다. 이에 본 연구에서는 학습데이터의 양과 품질이 생성형 AI 기반 언어모델 의 성능에 미치는 영향을 분석하고, 기존의 데이터 양(Quantity) 중심 접근 방식에서 벗어나 데이터 품질(Quality)을 반 영한 손실함수(loss function) 확장 모델을 제안하였다. 제안된 모델의 유효성을 정량․실증적으로 검증 후 데이터의 양 과 질을 기반으로 한 2×2 모델을 적용하여 충분한 고품질 학습데이터의 구축이 AI 성능 최적화에 필수적인 요소임을 입증하고 모델의 성능을 극대화하기 위한 최적의 학습데이터 구축 전략을 제시한다.

Generative language models have been driving innovation across various industries with the rapid advancement of AI technology. To achieve optimal performance, these models require a sufficient amount of high-quality training data, which plays a crucial role in enhancing generalization capabilities and ensuring reliable outputs. This study analyzes the impact of training data quantity and quality on the performance of generative AI-based language models and proposes an extended loss function model that incorporates data quality, moving beyond the traditional quantity-centric approach. The proposed model is quantitatively and empirically validated, followed by the application of a 2×2 model based on data quantity and quality. Through this approach, the research demonstrates that constructing high-quality training data is essential for optimizing AI performance and presents a strategic framework for maximizing language model performance.

8

딥러닝 기반 의료 영상 판독 소견서 생성 연구는 최근 몇 년간 의료 영상 분석 분야에서 중요한 발전을 이루었으며, 주로 비전 모델과 언어 모델의 융합에 초점을 맞춰왔다. 이러한 연구는 의료 전문가들의 업무 효율성을 높이고, 진 단의 일관성을 향상시키며 오류를 줄이는 데 기여할 것으로 기대된다. 본 논문에서는 자연어 처리 기술과 의료 영상 분석 기술의 융합을 통해 흉부 방사선 영상에 대한 판독 소견서를 자동으로 생성하는 최근 연구들을 소개하고, 이를 위한 주요 딥러닝 모델과 공용 데이터셋인 MIMIC-CXR 및 IU-Xray를 활용하여 성능 비교를 수행하였다. 성능 분석 결과, MIMIC-CXR 데이터셋에서는 RGRG 모델이 BLEU-4에서 0.126의 점수로 우수한 성능을 보였으며, IU-Xray 데이터셋에서는 COMG 모델이 BLEU-4에서 0.206을 기록하였다. 연구 결과, 각 모델이 특정 지표에서 강점을 보이는 반면, 데이터 불균형 및 개인정보 보호 문제와 같은 한계가 존재함을 확인하였다. 이를 바탕으로 향 후 연구의 발전 방향을 제시하며, 이러한 기술이 임상 현장에서 실질적으로 적용될 수 있는 가능성을 논의한다.

Research on deep learning-based automatic generation of radiology reports has seen significant advancements in recent years, with a primary focus on the integration of vision and language models. These advancements are expected to enhance the efficiency of medical professionals, improve diagnostic consistency, and reduce errors. In this paper, we introduce recent studies that combine natural language processing and medical image analysis techniques to automatically generate radiology reports for chest X-rays. Using key deep learning models and public datasets, including MIMIC-CXR and IU-Xray, we conducted a comparative performance evaluation. The analysis shows that on the MIMIC-CXR dataset, the RGRG model achieved superior performance with a BLEU-4 score of 0.126, while on the IU-Xray dataset, the COMG model recorded a BLEU-4 score of 0.206. Our findings reveal that while each model excels in specific evaluation metrics, limitations such as data imbalance and privacy concerns persist. Based on these findings, we propose future research directions and discuss the potential for these technologies to be practically applied in clinical settings.

9

딥러닝 기반 언어모델을 이용한 한국어 학습자 쓰기 평가의 자동 점수 구간 분류 - KoBERT와 KoGPT2를 중심으로 - KCI 등재

조희련, 이유미, 임현열, 차준우, 이찬규

국제한국언어문화학회 한국언어문화학 제18권 제1호 2021.04 pp.217-241

※ 기관로그인 시 무료 이용이 가능합니다.

6,300원

이 연구에서는 '한국어 딥러닝 모델'이 '한국어 학습자의 쓰기 자료에 대한 한국어 교사의 평가 점수'를 어느 정도 유사하게 예측할 수 있는지 살펴보았다. 구체적으로 이 연구에서는 304편의 한국어 쓰기 자료와 각각에 대한 평가 점수를 KoBERT와 KoGPT2로 학습시킨 후 그것이 인간 채점자(한국어 교사)의 평가 점수를 어느 정도 유사하게 예측하는지 실험하였다. 학습 데이터는 주제에 따라 '직업'과 '행복'으로 구분하였고, 점수에 따라 4종 레이블을 부착하였다. 7겹 교차 검증을 통한 실험 결과, KoBERT에서는 '직업' 데이터에서 48.8%, '행복' 데이터에서 65.2%의 분류 정확도를 나타냈다. KoGPT2에서는 같은 데이터에 대해 각각 50.6%와 58.9%의 분류 정확도를 나타냈다. 더불어, 모든 주제를 통합한 데이터에서는 KoBERT와 KoGPT2에 대해 각각 54.5%와 46.5%의 분류 정확도를 확인할 수 있었다. 이 연구를 통해 한국어 쓰기 자료에 대한 자동 채점 시스템의 가능성을 확인할 수 있었다. 향후 GPT-3의 한국어 모델이 개발되는 등의 기술 발전이 이루어진다면, 이 연구에서 시도한 한국어 자동 채점 시스템도 충분히 가능할 것으로 기대한다.

Automatic Score Range Classification of Korean Essays Using Deep Learning-based Korean Language Models-The Case of KoBERT & KoGPT2-. We investigate the performance of deep learning-based Korean language models on a task of automatically classifying Korean essays written by foreign students. We construct an experimental data set containing a total of 304 essays, which include essays discussing the criteria for choosing a job (‘job’), conditions of a happy life (‘happiness’), relationship between money and happiness, and definition of success. These essays were divided into four scoring levels, and using this 4-class data set, we fine-tuned two Korean deep learning-based language models, namely, KoBERT and KoGPT2, to use them in the automatic essay classification experiment. The 7-fold cross validation classification accuracies of ‘job’ and ‘happiness’ essays were 48.8% and 65.2% respectively for KoBERT, and 50.6% and 58.9% respectively for KoGPT2. Furthermore, the 7-fold cross validation classification accuracies of the integrated dataset that combined all essays were 54.5% and 46.5% for KoBERT and KoGPT2 respectively.

10

Fine-Tuning the Llama2 Large Language Model Using Books on the Diagnosis and Treatment of Musculoskeletal System in Physical Therapy

김준희

[NRF 연계] KEMA학회 Journal of Musculoskeletal Science and Technology Vol.8 No.2 2024.12 pp.65-73

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Background In physical therapy, accurate information on musculoskeletal diagnosis and treatment is essential for patient care. Generative language models use machine learning to perform tasks like text generation and question answering, mimicking human language understanding. Purpose This study aimed to fine-tune the Llama2 language model using text data from books on the diagnosis and treatment of the musculoskeletal system in physical therapy, and to compare it to the base model to evaluate its usability in medical fields. Study design Technical evaluation study Methods The Llama2-13B Chat model was fine-tuned using text from books on musculoskeletal diagnosis and treatment in physical therapy and trained on 1.2 million tokens. Responses were compared to the base model based on relevance and specialization to clinical questions. Results Compared to the base model, the fine-tuned model consistently generated answers specific to the musculoskeletal system diagnosis and treatment, demonstrating improved understanding of the specialized domain. Conclusions The model fine-tuned for musculoskeletal diagnosis and treatment books provided more detailed information related to musculoskeletal topics, and the use of this fine-tuned model could be helpful in medical education and the acquisition of specialized knowledge.

11

LLM(Large Language Model) 속성과 성능 연관성 연구 KCI 등재

임철홍

한국EA학회 정보화연구 제20권 3호 2023.09 pp.257-266

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

OpenAI의 ChatGPT는 기존 기술보다 놀랍게 향상된 기능을 보여주었고 생성형 AI에 대한 활 발한 연구개발과 활용이 시작되었다. 컴퓨터 성능의 비약적인 발전으로 거대한 크기의 인공지능 언어 모델(LLM)을 활용하게 되었다. 본 연구는 LLM을 구축할 때 성능을 평가할 방법과 성능을 높이기 위 해 LLM의 구축 모든 과정에 걸쳐 고려해야 할 요인을 제시한다. 원시 데이터로부터 학습하여 backbone 모델을 만드는 것부터 만들어진 모델에 fine-tuning을 통해서 최종 모델을 구축하는 과정을 포함한다. 학습 데이터의 양과 품질, 모델 및 파라미터크기, 학습횟수 및 fine-tuning에 관련된 방안을 제시한다. 영문모델은 크기를 줄이면서 성능을 높이는 성과가 있는 반면에 한국어 모델은 작고 성능이 우수하면 서 연구개발과 서비스에 제한 없이 활용 가능한 오픈소스 모델의 구축과 보급이 필수적이다.

OpenAI's ChatGPT showed remarkable improvements over existing technologies, and active research and development and utilization of generative AI began. Advances in computing power have led to the use of large-scale artificial intelligence language models (LLMs). This study presents methods to evaluate performance when building an LLM and factors to be considered throughout the entire process of building an LLM to increase performance. It includes the entire process of building a model by learning from raw data and fine-tuning the created backbone model. We present how to improve performance from considering the quantity and quality of learning data, model and parameter size, learning frequency, and fine tuning. When it comes to the English dataset, there have been achievements in improving performance while reducing the model size, while in the case of Korean, it is essential to build and disseminate an open source model that is small but has excellent performance and can be used without restrictions for research and development and services.

12

4,000원

본 연구는 기업의 AI Transformation(AX)을 지원하기 위해, Vision Language Model(VLM) 기반 지능 형 문서처리 플랫폼을 설계하고, Qwen2.5VL-7B를 활용한 영수증 처리 프로토타입을 구현하였다. 제안된 플랫폼 은 3-Tier 마이크로서비스 아키텍처를 기반으로, 프롬프트 관리 체계와 기능별 모듈화를 통해 유연하고 확장 가능 한 구조를 구현하였다. 실험 결과, 평균 91.7%의 정보 추출 정확도를 달성하였으며, 사전 템플릿 없이 다양한 문서 형식에 대응 가능한 처리 유연성을 바탕으로 실무 적용 가능성을 입증하였다. 본 연구는 OCR 중심 기술의 한계를 보완하는 프롬프트 기반 VLM 아키텍처를 실증적으로 제시하고, 금융·물류·의료 등 산업 전반에서 적용 가능한 문 서 자동화 기반을 제공하였다는 점에서 학문적·실무적 의의를 갖는다.

This study supports corporate AI Transformation (AX) by designing a document processing platform based on a Vision Language Model (VLM) and implementing a prototype using Qwen2.5VL-7B. The platform employs a three-tier microservice architecture with prompt management and modular components to ensure flexibility and scalability. Experiments showed an average information extraction accuracy of 91.7%, and the system demonstrated practical applicability by handling diverse document formats without predefined templates. This research provides an empirical implementation of a prompt-based VLM architecture that overcomes limitations of OCR technologies, offering academic and practical value as a foundation for document automation across sectors such as finance, logistics, and healthcare.

13

Artificial intelligence (AI) is a significant tool in modern military operations in that it helps to analyze a large volume of strategic, tactical, and operational data. On the other hand, current large language models (LLMs) like GPT-4 or Falcon have difficulty resolving problems in defense-specific contexts because of issues related to security, data confidentiality, and the lack of explainability. This document presents MilGPT, a secure and explainable LLM structure that aims at solving military problems only. To the model, fine-tuned open-source architectures with domain-specific defense datasets are integrated to elevate intelligence synthesis, decision-making, and threat prediction. On the benchmark, performance evaluation tasks show that MilGPT accounts for a 27% increase in contextual accuracy, an 18% reduction in hallucination rate, and an 33% improvement in explainability as measured by gradient-based feature attribution. In the proposed framework, military intelligence systems are not only secured but also made adaptive and humaninterpretable, thus, setting up a basis for the coming generation of AI models capable of defense-grade tasks.

14

The Topic Detection and Tracking with Topic Sensitive Language Model

Ruifang He, Bing Qin, Ting Liu, Sheng Li

한국어정보학회 한국어정보학 제8권 1호 2006.06 pp.64-68

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

In this paper, we explore the language model with topic sensitive features for the topic detection and tracking, formulate the relationship among the Chinese internet new words, language model with topic sensitive feature and the scheduling logic and the interval temporal reasoning and the key techniques. we use the Chinese internet new words to strengthen the detection and tracking of the topic and try to employ the scheduling logic and interval temporal reasoning to educe the reciprocal influences of events. At last we summarize the potential issues and the future work.

15

4,000원

16

Small Language Models are competitive without large-scale infrastructure, their performance is highly contingent on prompt design. This study analyzes the sensitivity of BitNet b1.58-2B-4T to label exposure and fewshot exemplar composition on a 36-class medical query classification task. We generated 504 items consisting of 6 direct and 8 indirect questions for each disease and after removing cross-exemplar leakage the final evaluation set contained 494 items. With no parameter updates, 0/1/2/5/10- shot prompting was evaluated using Accuracy. Under the nolabel- exposure setting accuracy increased as more exemplars were provided. However, these gains were accompanied by growing prediction concentration on exemplar labels. In contrast with label-exposure, zero-shot achieved the highest accuracy, while the inclusion of exemplars reduced accuracy and amplified label bias. These results show that the structure of the prompt tends to shift few-shot effects from beneficial to detrimental. This highlights the importance of controlled prompt design and domain-adaptive training to ensure trustworthy performance.

17

4,000원

본 논문에서는 한국어 쓰기 교육을 위한 교육용 게임으로서 text-to-scene 모델을 제안한다. 사용자 입력 문장에 대해 즉각적인 피드백을 제공하기 위해서, 제안하는 text-to-scene 모델은 사용자 입력 문장의 의미구조를 분석하여 장면을 자동으로 생성한다. 사용자가 문장을 교정하면서 자연스럽게 한국어 쓰기 능력이 향상될 수 있도록, 제안하는 text-to-scene 모델은 정답 장면과 사용자 입력 문 장에 대한 장면을 함께 제공하므로, 사용자는 입력 문장의 어느 부분이 잘못되었는지 쉽게 파악하 고 문장을 수정할 수 있다. 기존의 text-to-scene 연구에서 정교하게 장면을 생성하기 위해서는 필 요한 자원 구축이 어렵고 처리 과정이 매우 복잡하다는 점을 고려하여, 제안하는 text-to-scene 모 델은 외국인이 쉽게 이해할 수 있는 수준으로 장면을 단순하게 표현한다. 실험결과 격조사나 문장 요소 생략 등을 고려하여 15개의 단어를 임의로 조합하여 5개 어절 이내의 문장 1,048,576개 문장 을 만들어 실험한 결과 제안하는 모델은 한 문장당 평균 14.94ms가 소요하면서 안정적으로 장면을 생성하였다. 외국인을 대상으로 한 작문 시험 결과, 종이를 이용한 작문 결과에서는 10개 문항 중 오답 문장이 평균 4.7개 나왔는데, 제안하는 text-to-scene 모델을 이용한 작문 결과에서는 한 문장 당 평균 1.56회의 미리보기를 하고 문장을 수정하여 0.87개가 나왔다.

In this paper, we propose a text-to-scene model as one of educational games for Korean language writing education as a foreign language. In order to immediately provide feedback on user input sentences, the proposed text-to-scene model automatically generates the scene corresponding to the user input sentence by analyzing its semantic structure. For the purpose of improving the user's writing ability by correcting errors in the user input sentence, the proposed text-to-scene model provides both the scenes of the correct sentence and the scenes of the user input sentence; so that the user can easily recognize the errors in the user input sentence and correct them. Previous text-to-scene approaches have a weak point that it is difficult to construct the resources and to generate the elaborate scene. Considering that the construction cost and the processing time, the proposed text-to-scene model generates the simple scene that foreigner can easily understand. Experimental results show that the proposed text-to-scene model takes 14.94ms per sentence and produces every scene for 1,048,576 sentences, which is made by arbitrarily combining 15 words. On the paper writing test, a foreigner writes 4.7 incorrect sentences among 10 sentences; while the foreigner writes 0.87 incorrect sentences with the proposed text-to-scene model on the writing test.

18

Dual Language Education/Instruction Model-based Teaching Method for Korean Students’ English Communication Skills KCI 등재

Hyangki Jung, Inn-Chull Choi

한국외국어교육학회 외국어교육 제25권 제3호 2018.09 pp.53-78

※ 기관로그인 시 무료 이용이 가능합니다.

6,400원

The main objective of this research was to probe the DLM-based teaching method for Korean students’ English communication skills. For this study, the corpus was made of Korean reading materials with 307 frequently used sentence patterns extracted. The present research employing a quasi-experimental design and a comparative analysis of Korean and English corpora revealed the followings: first, the difference between the DLM (experimental group) and the CTR (control group) was examined in terms of the pretest score in order to identify the students’ level of English productive skills. After the pretest, meaningful translation focused DLM was employed in the instruction of the experimental group but not in the instruction of the control group. After one semester of teaching, on the posttest, the students in the DLM group outperformed the students in the CTR group. In conclusion, it can be inferred that the DLM teaching method is effective for English productive skills and can be a good solution to our English education environments in which both teachers and students use Korean as a medium of teaching English.

19

The problem of discrimination due to information alienation among Korean sign language users is continuously mentioned, and sign language translation research is actively being conducted to solve this problem. However, due to technological limits in translating text into Korean Sign Language, the need for specialist equipment causes annoyance and spatial constraints. Furthermore, it isn't easy to replicate the vocabulary and grammatical structure of the Korean language in Korean Sign Language. Furthermore, the service's commercialization is complicated by the need for more nonmanual signal (NMS) identification technology.

20

The CIPP Model : Applications in Language Program Evaluation SCOPUS KCI 등재

Satja Sopha, Alexander Nanni

아시아영어교육학회 The Journal of AsiaTEFL Vol.16 No.4 2019.12 pp.1360-1367

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

This article introduces the CIPP Model (Context, Input, Process, and Product) and explains its application in the evaluation of language programs. The model has long been used in various fields to evaluate programs both before they begin (e.g., by assessing the alignment of the contexts and input) and after they are complete (e.g., by evaluating how well the process has been implemented and whether the product is up to standard). The flexibility of the model is a major strength. In the field of language teaching, this model is highly relevant to curriculum development and can be applied at both the course level and the program level. This paper first introduces competing models of language program evaluation. It then introduces the CIPP Model and explains its applicability in the field of language education, providing suggestions for the application of the model. The CIPP Model has the potential to assist TESOL professionals — teachers and administrators — in improving their professional practice, curriculum design, and program evaluation. Educators in a variety of contexts would benefit from furthering their knowledge of this widely applied evaluation model.

 
1 2 3 4 5
페이지 저장