Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 18
No
1

엔트리 텍스트 모델 학습을 활용한 초등 인공지능 교육 내용 개발 KCI 등재

김병조, 김현배

한국정보교육학회 정보교육학회논문지 제26권 제1호 2022.02 pp.65-73

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 연구에서는 엔트리 텍스트 모델 학습을 활용해서 초등학교 인공지능 교육 내용을 개발하고 이를 실제 수 업에 적용한다. 초·중등 인공지능 내용 체계표를 바탕으로 실과 소프트웨어 교육과 인공지능 교육의 성취 기준 을 재구성한다. 기계학습이 가능한 텍스트, 이미지, 소리 중에서 다양한 플랫폼에서 지원하고 초등학생의 데이터 준비 시간을 줄일 수 있으면서 손쉽게 이해가 가능한 ‘텍스트 모델 학습을 활용한 감정 인식 프로그램 제작’을 교육 내용으로 선정한다. 엔트리 인공지능을 교육 플랫폼으로 선정해서 텍스트 모델 학습을 활용한 감정인식 프 로그램을 만드는 인공지능 교육 내용을 개발하고 실제 초등학교 수업에 적용한다. 수업 적용 결과 엔트리 인공 지능 수업에 긍정적인 반응과 흥미를 보였다. 본 연구 내용을 기반으로 초등학생을 대상으로 한 수업의 효과성 에 대한 양적 연구가 후속 연구로 필요함을 제언한다.

In this study, by using Entry text model learning, educational contents for artificial intelligence education of elementary school students are developed and applied to actual classes. Based on the elementary and secondary artificial intelligence content table, the achievement standards of practical software education and artificial intelligence education will be reconstructed.. Among text, images, and sounds capable of machine learning, "production of emotion recognition programs using text model learning" will be selected as the educational content, which can be easily understood while reducing data preparation time for elementary school students. Entry artificial intelligence is selected as an education platform to develop artificial intelligence education contents that create emotion recognition programs using text model learning and apply them to actual elementary school classes. Based on the contents of this study, As a result of class application, students showed positive responses and interest in the entry AI class. it is suggested that quantitative research on the effectiveness of classes for elementary school students is necessary as a follow-up study.

2

텍스트 마이닝과 딥러닝 알고리즘을 이용한 가짜 뉴스 탐지 모델 개발 KCI 등재

임동훈, 김건우, 최근호

한국경영정보학회 경영정보학연구 제23권 제4호 2021.11 pp.127-146

※ 기관로그인 시 무료 이용이 가능합니다.

5,500원

가짜 뉴스는 정보화 시대라는 현대사회의 특성에 의해 진위 여부의 검증과는 상관없이 빠른 속도로 확대, 재생산되어 퍼진다. 전체 뉴스의 1%를 가짜라고 가정했을 경우 우리사회에 미치는 경제적 비용이 30조 원에 달한다고 하니 가짜 뉴스는 사회적, 경제적으로 매우 중요한 문제라고 할 수 있다. 이에 본 연구는 뉴스의 진위 여부를 신속하고 정확하게 확인하고자 자동화된 가짜 뉴스 탐지 모델을 개발하는데 목적을 두고 있다. 이를 위해 본 연구에서는 크롤링(crawling)을 통해 진위 여부가 밝혀진 뉴스 기사를 수집하였고, 워드 임베딩(Word2Vec, Fasttext)과 딥러닝 기법(LSTM, BiLSTM)을 이용하여 가짜 뉴스 예측 모델을 개발하였다. 실험 결과, Word2Vec과 BiLSTM의 조합이 가장 높은 84%의 정확도를 보였다.

Fake news isexpanded and reproduced rapidly regardless of their authenticity by the characteristics of modern society, called the information age. Assuming that 1% of all news are fake news, the amount of economic costs is reported to about 30 trillion Korean won. This shows that the fake news isvery important social and economic issue. Therefore, this study aims to develop an automated detection model to quickly and accurately verify the authenticity of the news. To this end, this study crawled the news data whose authenticity is verified, and developed fake news prediction models using word embedding (Word2Vec, Fasttext) and deep learning algorithms (LSTM, BiLSTM). Experimental results show that the prediction model using BiLSTM with Word2Vec achieved the best accuracy of 84%.

3

A Lightweight Deep Learning Model for Text Detection in Fashion Design Sketch Images for Digital Transformation

Ju-Seok Shin, Hyun-Woo Kang

[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.28 No.10 2023 pp.17-25

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 의류 디자인 도면 이미지의 글자 검출을 위한 경량화된 딥러닝 네트워크를 제안하였다. 최근 의류 디자인 산업에서 Digital Transformation의 중요성이 대두되면서, 디지털 도구를 활용한 의류 디자인 도면 작성이 강조되고 있으며, 디지털화된 의류 디자인 도면의 활용 가능성을 고려할 때, 도면에서 글자 검출과 인식이 중요한 첫 단계로 간주된다. 이 연구에서는 기존의 글자 검출 딥러닝 모델을 기반으로 의류 도면 이미지의 특수성을 고려하여 경량화된 네트워크를 설계하였으며, 별도로 수집한 의류 도면 데이터 셋을 추가하여 딥러닝 모델을 학습시켰다. 실험 결과, 제안한 딥러닝 모델은 의류 도면 이미지에서 기존 글자 검출 모델보다 약 20% 높은 성능을 보였다. 따라서 이 논문은 딥러닝 모델의 최적화와 특수한 글자 정보 검출 등의 연구를 통해 의류 디자인 분야에서의 Digital Transformation에 기여할 것으로 기대한다.

In this paper, we propose a lightweight deep learning architecture tailored for efficient text detection in fashion design sketch images. Given the increasing prominence of Digital Transformation in the fashion industry, there is a growing emphasis on harnessing digital tools for creating fashion design sketches. As digitization becomes more pervasive in the fashion design process, the initial stages of text detection and recognition take on pivotal roles. In this study, a lightweight network was designed by building upon existing text detection deep learning models, taking into consideration the unique characteristics of apparel design drawings. Additionally, a separately collected dataset of apparel design drawings was added to train the deep learning model. Experimental results underscore the superior performance of our proposed deep learning model, outperforming existing text detection models by approximately 20% when applied to fashion design sketch images. As a result, this paper is expected to contribute to the Digital Transformation in the field of clothing design by means of research on optimizing deep learning models and detecting specialized text information.

4

Machine Learning Based Automatic Categorization Model for Text Lines in Invoice Documents

Shin, Hyun-Kyung

[Kisti 연계] 한국멀티미디어학회 멀티미디어학회논문지 Vol.13 No.12 2010 pp.1786-1797

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Automatic understanding of contents in document image is a very hard problem due to involvement with mathematically challenging problems originated mainly from the over-determined system induced by document segmentation process. In both academic and industrial areas, there have been incessant and various efforts to improve core parts of content retrieval technologies by the means of separating out segmentation related issues using semi-structured document, e.g., invoice,. In this paper we proposed classification models for text lines on invoice document in which text lines were clustered into the five categories in accordance with their contents: purchase order header, invoice header, summary header, surcharge header, purchase items. Our investigation was concentrated on the performance of machine learning based models in aspect of linear-discriminant-analysis (LDA) and non-LDA (logic based). In the group of LDA, na$\"{\i}$ve baysian, k-nearest neighbor, and SVM were used, in the group of non LDA, decision tree, random forest, and boost were used. We described the details of feature vector construction and the selection processes of the model and the parameter including training and validation. We also presented the experimental results of comparison on training/classification error levels for the models employed.

5

An Optimal Weighting Method in Supervised Learning of Linguistic Model for Text Classification

Mikawa, Kenta, Ishida, Takashi, Goto, Masayuki

[Kisti 연계] 대한산업공학회 Industrial engineering & management systems Vol.11 No.1 2012 pp.87-93

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper discusses a new weighting method for text analyzing from the view point of supervised learning. The term frequency and inverse term frequency measure (tf-idf measure) is famous weighting method for information retrieval, and this method can be used for text analyzing either. However, it is an experimental weighting method for information retrieval whose effectiveness is not clarified from the theoretical viewpoints. Therefore, other effective weighting measure may be obtained for document classification problems. In this study, we propose the optimal weighting method for document classification problems from the view point of supervised learning. The proposed measure is more suitable for the text classification problem as used training data than the tf-idf measure. The effectiveness of our proposal is clarified by simulation experiments for the text classification problems of newspaper article and the customer review which is posted on the web site.

6

Machine learning-based item difficulty prediction model using item text features

최윤하

[NRF 연계] 한국중등영어교육학회 중등영어교육 Vol.17 No.1 2024.02 pp.46-66

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This study investigates the difficulty of English reading passages in the College Scholastic Ability Test (CSAT) in South Korea, analyzing corpus linguistic features from 2015 to 2023 using Coh-Metrix software. A notable innovation of this research is the application of the Random Forest algorithm, which overcomes the limitations of traditional regression analysis, such as multicollinearity and overfitting. Focusing on textual complexity, the analysis utilizes 108 explanatory variables to predict the difficulty of CSAT English items. Rigorous methods, including grid search and cross-validation, were employed to ensure the reliability of the results. The Random Forest regression model demonstrated high predictive accuracy. Crucial variables influencing the difficulty of items were identified, providing valuable insights for educators and policymakers to improve CSAT preparation strategies. This research marks a significant advancement in language assessment by combining corpus analysis with machine learning algorithms. It not only enhances the prediction of test difficulty but also supports students and teachers with effective, tailored preparation strategies to address the varied challenges encountered in English public education.

7

Text-to-Image를 위한 아동 손그림 학습 모델 생성 연구

이은채, 문미경

[Kisti 연계] 한국컴퓨터정보학회 한국컴퓨터정보학회 학술대회논문집 2022 pp.505-506

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

인공지능 기술은 점차 빠른 속도로 발전되며 응용 분야가 확대되어 창작 산업에서의 역할도 커져 예술, 영화 및 기타 창조적인 산업에도 영향을 주고 있다. 이러한 인공지능 기술을 이용하여 텍스트로 설명하면 다양한 스타일의 이미지를 생성해내는 기술이 있지만 아동이 직접 그린 손그림 스타일의 그림을 생성하지는 못한다. 본 논문에서는 아동 손그림 데이터를 통해 Text-to-Image를 학습시켜 새로운 학습 모델을 생성하는 과정에 대해서 기술한다. 이 연구를 통해 생성된 픽셀을 결합하여 텍스트를 기반으로 하나의 아동 손그림을 만들 수 있을 것으로 기대한다.

8

자연어 처리에 기반한 사상체질 치험례의 텍스트 마이닝 분석과 체질 진단을 위한 머신러닝 모델 선정

김진석, 박소현, 정로아, 이은수, 김윤서, 성현동, 유준상

[Kisti 연계] 대한한의학회 Journal of Korean Medicine Vol.45 No.3 2024 pp.193-210

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Objectives: We analyzed Sasang constitution case reports using text mining to derive network analysis results and designed a classification algorithm using machine learning to select a model suitable for classifying Sasang constitution based on text data. Methods: Case reports on Sasang constitution published from January 1, 2000, to December 31, 2022, were searched. As a result, 343 papers were selected, yielding 454 cases. Extracted texts were pretreated and tokenized with the Python-based KoNLPy package. Each morpheme was vectorized using TF-IDF values. Word cloud visualization and centrality analysis identified keywords mainly used for classifying Sasang constitution in clinical practice. To select the most suitable classification model for diagnosing Sasang constitution, the performance of five models-XGBoost, LightGBM, SVC, Logistic Regression, and Random Forest Classifier-was evaluated using accuracy and F1-Score. Results: Through word cloud visualization and centrality analysis, specific keywords for each constitution were identified. Logistic regression showed the highest accuracy (0.839416), while random forest classifier showed the lowest (0.773723). Based on F1-Score, XGBoost scored the highest (0.739811), and random forest classifier scored the lowest (0.643421). Conclusions: This is the first study to analyze constitution classification by applying text mining and machine learning to case reports, providing a concrete research model for follow-up research. The keywords selected through text mining were confirmed to effectively reflect the characteristics of each Sasang constitution type. Based on text data from case reports, the most suitable machine learning models for diagnosing Sasang constitution are logistic regression and XGBoost.

9

Naive Bayes 문서 분류기를 위한 점진적 학습 모델 연구

김제욱, 김한준, 이상구

[Kisti 연계] 한국데이타베이스학회 한국데이타베이스학회 학술대회논문집 2001 pp.331-341

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 Naive Bayes 문서 분류기를 위한 새로운 학습모델을 제안한다. 이 모델에서는 라벨이 없는 문서들의 집합으로부터 선택한 적은 수의 학습 문서들을 이용하여 문서 분류기를 재학습한다. 본 논문에서는 이러한 학습 방법을 따를 경우 작은 비용으로도 문서 분류기의 정확도가 크게 향상될 수 있다는 사실을 보인다. 이와 같이, 알고리즘을 통해 라벨이 없는 문서들의 집합으로부터 정보량이 큰 문서를 선택한 후, 전문가가 이 문서에 라벨을 부여하는 방식으로 학습문서를 결정하는 것을 selective sampling이라 한다. 본 논문에서는 이러한 selective sampling 문제를 Naive Bayes 문서 분류기에 적용한다. 제안한 학습 방법에서는 라벨이 없는 문서들의 집합으로부터 재학습 문서를 선택하는 기준 측정치로서 평균절대편차(Mean Absolute Deviation), 엔트로피 측정치를 사용한다. 실험을 통해서 제안한 학습 방법이 기존의 방법인 신뢰도(Confidence measure)를 이용한 학습 방법보다 Naive Bayes 문서 분류기의 성능을 더 많이 향상시킨다는 사실을 보인다.

10

딥러닝 모델을 활용한 실시간 인쇄물 문자 탐지 시스템

최예준, 김송원, 문미경

[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.19 No.3 2024 pp.523-530

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

웹페이지나 디지털 문서 등과 같은 온라인에서는 사용자가 검색하고 싶은 특정 단어나 특정 문구를 실시간으로 검색하는 기능이 있다. 인쇄된 도서나 참고서 등과 같은 인쇄물에는 실시간으로 특정 단어나 특정 문구를 찾는 기능이 없어 어려움을 겪는 경우가 많다. 본 논문에서는 텍스트를 탐지(Detection)하는 딥러닝 모델과 텍스트를 인식(Recognition)하는 OCR을 활용한 실시간 문자 탐지 시스템의 개발내용에 관해 기술한다. 본 연구에서는 EAST 모델을 사용하여 텍스트를 탐지하는 방법, 탐지한 텍스트를 EasyOCR을 사용하여 인식하는 방법, 인식한 텍스트를 사용자가 검색하고 싶은 특정 단어나 특정 문구를 비교하여 bounding box로 나타내는 방법을 제안한다. 이 시스템을 통해 사용자는 도서나 참고서 등과 같은 인쇄물에서 실시간으로 검색하고 싶은 특정 단어나 특정 문구를 찾아 필요한 정보를 쉽고 빠르게 찾는 것에 효과적일 것을 기대한다.

Online, such as web pages and digital documents, have the ability to search for specific words or specific phrases that users want to search in real time. Printed materials such as printed books and reference books often have difficulty finding specific words or specific phrases in real time. This paper describes the development of a deep learning model for detecting text and a real-time character detection system using OCR for recognizing text. This study proposes a method of detecting text using the EAST model, a method of recognizing the detected text using EasyOCR, and a method of expressing the recognized text as a bounding box by comparing a specific word or specific phrase that the user wants to search for. Through this system, users expect to find specific words or phrases they want to search in real time in print, such as books and reference books, and find necessary information easily and quickly.

11

예체능계 대학생의 학술적인 글쓰기 교수-학습 모형 개발을 위한 텍스트 분석

양태영

[NRF 연계] 학습자중심교과교육학회 학습자중심교과교육연구 Vol.16 No.7 2016.07 pp.541-562

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

예체능계열 대학생은 전공의 특성상 작품이나 실기를 통하여 자기를 직접 표현하는 것에 익숙하고, 문자를 통한 표현에는 익숙하지 않은 편이다. 이로 인해 쓰기 학습에 대한 동기 부여가 쉽지 않으며, 쓰기에 대한 관심과 흥미에 개인차가 크다. 그러나 이 학생들도 대학 과정을 성공적으로 수행하고, 자신이 속한 전공 분야의 전문 담화 집단에서 인정받을 수 있는 내용과 형식을 갖춘 학술적인 텍스트를 쓸 수 있어야 전문가가 될 수 있다. 본 연구는 예체능계열의 대학생과 전문가의 텍스트를 분석하여 계열별 특징을 밝히고, 쓰기 교수-학습 모형에 정보를 제공하고자 하였다. 연구 대상으로 삼은 자료는 예체능계열 대학생 보고서 25편, 27,468어절, 예체능계열 전문 필자의 글 12편, 11,815어절 총 3,9283어절이다. 텍스트 분석을 위해서 미시구조, 거시구조, 최상위구조의 세 층위에서 나타나는 여러 특징을 분석하였다. 그 결과, 예체능계열의 텍스트는 작품을 매개로 대중과의 소통을 목적으로 하고 있었다. 이를 위해 작품에 대한 객관적인 정보가 중심이 되는 설명텍스트, 작품에 대한 감상과 이해를 함께 전달하는 감상텍스트의 두 종류가 나타났다. 설명텍스트와 감상텍스트라는 텍스트 유형에 따른 차이와 집필 관점의 차이, 화제 구성의 차이를 학습하도록 쓰기 교수-학습 모형을 설계할 수 있다면, 예체능계열의 특징을 반영한 효과적인 쓰기 교육이 가능할 것이다. 또한 연구 결과를 활용한다면 예체능 계열 보고서의 객관적인 평가 기준으로 활용될 수도 있을 것이다. 본 연구의 성과가 예체능계열 대학생의 쓰기 교육뿐만 아니라 다양한 계열의 특징을 밝히는 연구의 기초자료로 활용되어 학술적인 쓰기 능력 신장에 기여하기를 기대한다.

The purpose of this study is to provide basic materials for designing writing instruction models that are needed by Arts and Physical Education majors. In order to understand the qualities and characteristics of the Arts and Physical Education academic text, I analyzed students' reports. There were 25 pieces with 27,468 words and 12 pieces of experts' writing with 11,815 words. I built these written samples into a corpus. In the course of this study, four steps were carried out in the following order: Corpus building, analysis, comparing the results of the analysis, and information construction. In order to identify the features with three-dimensional understanding, I used a combination of quantitative and qualitative methods. The materials were micro-structure sentences, which were analyzed by vocabulary and sentence units and macro-structure of the statements and paragraph unit. Furthermore, a top-level structure was used to analyze the overall structure of an article. I built various information needed for teaching and learning through a multi-layer, three-dimensional analysis. To determine the quality of text, we analyzed the distribution of topic, which consists of the smallest units in each paragraph to form meaning. The topic of reports has a parallel structure and is not complicated. However, the experts’ writings demonstrate complex contents. This complexity is evident in the depth of topic, which is more profound due to the casual relationship or sequential progress with similar topics. By comparing the students' writing and the experts' writing, I developed the practical information that could be used for developing Arts and Physical Education majors' teaching and learning models.

12

설명 가능한 개인화 영화 추천 서비스를 위한 딥러닝 기반 텍스트 요약 모델

진요요, 강경모, 김재경

[Kisti 연계] 한국IT서비스학회 한국IT서비스학회지 Vol.21 No.2 2022 pp.109-126

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The number and variety of products and services offered by companies have increased dramatically, providing customers with more choices to meet their needs. As a solution to this information overload problem, the provision of tailored services to individuals has become increasingly important, and the personalized recommender systems have been widely studied and used in both academia and industry. Existing recommender systems face important problems in practical applications. The most important problem is that it cannot clearly explain why it recommends these products. In recent years, some researchers have found that the explanation of recommender systems may be very useful. As a result, users are generally increasing conversion rates, satisfaction, and trust in the recommender system if it is explained why those particular items are recommended. Therefore, this study presents a methodology of providing an explanatory function of a recommender system using a review text left by a user. The basic idea is not to use all of the user's reviews, but to provide them in a summarized form using only reviews left by similar users or neighbors involved in recommending the item as an explanation when providing the recommended item to the user. To achieve this research goal, this study aims to provide a product recommendation list using user-based collaborative filtering techniques, combine reviews left by neighboring users with each product to build a model that combines text summary methods among deep learning-based natural language processing methods. Using the IMDb movie database, text reviews of all target user neighbors' movies are collected and summarized to present descriptions of recommended movies. There are several text summary methods, but this study aims to evaluate whether the review summary is well performed by training the Sequence-to-sequence+attention model, which is a representative generation summary method, and the BertSum model, which is an extraction summary model.

13

프로토콜 분석을 활용한 쓰기 과정 지도 및 평가 - 논술 텍스트 생산 과정 모형을 중심으로

김평원

[NRF 연계] 한국국어교육학회 새국어교육 Vol.87 2011.04 pp.5-36

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

쓰기 교육의 관심이 결과 중심(product-oriented approach)에서 글을 완성하기까지의 과정을 강조하는 과정 중심(process-oriented approach)으로 변화하고 있음에도 현장에서는 과정 중심 글쓰기 학습 모형과 같은 거시적인 전략을 구체화 한 미시적인 방법론이 부족한 실정이다. 기존의 쓰기 지도는 학생이 작성한 텍스트를 첨삭하는 방식으로 피드백을 제공하기 때문에, 실제 쓰기 상황에서 겪는 문제점을 진단하고 이를 교정하기 위한 정보를 제공하는 데 한계가 있었다. 이 연구의 목적은 프로토콜 분석(protocol analysis)을 통해 학생이 생산한 논술 텍스트가 어떠한 과정을 거쳐 나왔는지를 파악하고, 이를 토대로 피드백을 제공하는 구체적인 방법을 제안하는 것이다. 프로토콜 분석은 필자의 의미 구성 과정을 추론한 것으로 한 마디로 글을 쓰는 동안 각 학생들이 어떤 생각을 하고 있는지를 파악하는 진단 방법이다. 프로토콜 분석의 토대가 되는 ‘논술 텍스트 생산 과정 모형’은 전형적인 논술 과정을 정보처리 모형으로 정리한 것으로 학생의 논술 과정을 분석하는 유용한 교육 모형임은 물론 논술 과정을 명시적으로 정리한 교육 자료로 활용할 수 있다. 이 모형을 토대로 피드백을 제공받은 학생은 기존의 결과 중심의 지면 첨삭 지도가 놓치고 있었던 논술 과정에 대한 구체적인 전략을 점검하고 수정할 수 있게 된다.

This study shows the method of providing feedback to enhance argumentative essay writing skills, diagnosing the information processing in the brain through the protocol analysis in the course of writing argumentative essays. The objective of this study is to collect the data of thinking process during the argumentative essay writing and to analyze and theorize problem solving process. Indicating the problems and providing information to correct them in real essay writing have been limiting, because in the conventional argumentative essay writing education and assessment, the feedback is provided by checking the details of the exams and calculating the scores. This study developed the exemplary argumentative essay writing process into the information processing model, and suggested the analysis and feedback model of students' protocol process. The model of argumentative essay writing process through the protocol analysis can be used not only for measuring the argumentative essay writing education but also for training the test takers.

14

EfficientNet 모델과 전이학습을 이용한 상품 이미지와 텍스트 데이터의 결합

임수빈, 김범윤, 김선재, 한정우, 유동영

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2023 pp.334-335

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 논문에서는 이미지 데이터와 각종 텍스트 기반의 데이터를 적절히 결합하여 유용한 데이터를 만들어 내는 방법을 제안한다. 그 사례로 편의점 상품 이미지와 편의점 프로모션 데이터, 사용자 위치정보 데이터를 적절히 결합하여 사용자가 편의점 상품 전면 이미지를 제공했을 때, 해당 상품이 어떤 편의점 브랜드에서 어떤 프로모션을 진행하고 있는지, 그리고 현재 위치에서 가까운 점포가 어디인지를 사용자에게 제공하는 시스템을 구현한다. 이미지를 어떤 데이터와 결합하는지에 따라 다양한 요구사항에 대응할 수 있다.

15

딥러닝 기반 분류 모델의 성능 분석을 통한 건설 재해사례 텍스트 데이터의 효율적 관리방향 제안

김하영, 장예은, 강현빈, 손정욱, 이준성

[Kisti 연계] 한국건설관리학회 건설관리 Vol.22 No.5 2021 pp.73-85

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 딥러닝 기반의 텍스트 데이터 분류 모델의 성능 고찰을 통해 한국어 건설 재해사례의 효율적 관리방향을 제안한다. 이를 위해 비정형 텍스트 문서인 건설 재해 보고서를 활용해 건설 사고의 대표적 유형인 추락, 감전, 낙하, 붕괴, 협착의 5개 범주로 분류하는 딥러닝 모델을 구현하였다. 초기 모델 테스트 결과, 추락 재해의 분류 정확도가 상대적으로 높게 도출되며 타 유형을 추락 재해로 분류하는 경우가 많이 발생한다는 특징이 나타났다. 원인 분석 결과, 1) 구체적인 사고 유발 행동, 2) 유사한 문장 구조, 3) 여러 유형에 해당되는 복합사고가 위의 특징에 영향을 미치는 것으로 분석되었으며, 이 중 추가 실험을 통해 검증이 가능한 복합사고에 대한 두 가지 정확도 개선 실험을 진행하였다: 1) 재분류, 2) 제외. 실험 결과, 복합사고 제외 시 분류 성능이 185.7% 향상되었으며, 이를 통해 여러 사고 유형에 대한 내용을 동시에 포함하는 복합사고의 다중공선성(multicollinearity)이 해소되었음을 알 수 있다. 결론적으로 본 연구에서는 향후 사고에 대한 상황을 상세히 서술하는 체계를 마련함과 동시에 복합사고를 독립적으로 관리할 필요성을 시사한다.

This study proposes an efficient management direction for Korean construction accident cases through a deep learning-based text data classification model. A deep learning model was developed, which categorizes five categories of construction accidents: fall, electric shock, flying object, collapse, and narrowness, which are representative accident types of KOSHA. After initial model tests, the classification accuracy of fall disasters was relatively high, while other types were classified as fall disasters. Through these results, it was analyzed that 1) specific accident-causing behavior, 2) similar sentence structure, and 3) complex accidents corresponding to multiple types affect the results. Two accuracy improvement experiments were then conducted: 1) reclassification, 2) elimination. As a result, the classification performance improved with 185.7% when eliminating complex accidents. Through this, the multicollinearity of complex accidents, including the contents of multiple accident types, was resolved. In conclusion, this study suggests the necessity to independently manage complex accidents while preparing a system to describe the situation of future accidents in detail.

16

중국어 미디어텍스트 번역수업 설계 모형 ― Blended Learning 기법을 중심으로

최병학

[NRF 연계] 중국어문학연구회 중국어문학논집 Vol.114 2019.02 pp.203-226

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The usefulness of self-directed learning, which emphasizes the active education of learners, is emerging, and blended learning is being presented as a new paradigm of future education. Blended Learning-based translation education offers some advantages, including: First, it supports powerful multimedia tools. Second, it offers a variety of interaction opportunities. Third, it enables smooth contact with Chinese culture. Fourth, it enables level-specific learning. Fifth, it enhances self-directed learning ability. In foreign language education, differentiated learning activities and learning materials are required according to individual language level. Therefore, it is a very appropriate teaching strategy to create a self-directed learning environment in which learners choose learning activities that meet their own level and lead learning activities at their own pace.

17

한글 텍스트 감정 이진 분류 모델 생성을 위한 미세 조정과 전이학습에 관한 연구

김종수

[Kisti 연계] 한국산업정보학회 한국산업정보학회논문지 Vol.28 No.5 2023 pp.15-30

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

근래에 트랜스포머(Transformer) 구조를 기초로 하는 ChatGPT와 같은 생성모델이 크게 주목받고 있다. 트랜스포머는 다양한 신경망 모델에 응용되는데, 구글의 BERT(bidirectional encoder representations from Transformers) 문장생성 모델에도 사용된다. 본 논문에서는, 한글로 작성된 영화 리뷰에 대한 댓글이 긍정적인지 부정적인지를 판단하는 텍스트 이진 분류모델을 생성하기 위해서, 사전 학습되어 공개된 BERT 다국어 문장생성 모델을 미세조정(fine tuning)한 후, 새로운 한국어 학습 데이터셋을 사용하여 전이학습(transfer learning) 시키는 방법을 제안한다. 이를 위해서 104 개 언어, 12개 레이어, 768개 hidden과 12개의 집중(attention) 헤드 수, 110M 개의 파라미터를 사용하여 사전 학습된 BERT-Base 다국어 문장생성 모델을 사용했다. 영화 댓글을 긍정 또는 부정 분류하는 모델로 변경하기 위해, 사전 학습된 BERT-Base 모델의 입력 레이어와 출력 레이어를 미세 조정한 결과, 178M개의 파라미터를 가지는 새로운 모델이 생성되었다. 미세 조정된 모델에 입력되는 단어의 최대 개수 128, batch_size 16, 학습 횟수 5회로 설정하고, 10,000건의 학습 데이터셋과 5,000건의 테스트 데이터셋을 사용하여 전이 학습시킨 결과, 정확도 0.9582, 손실 0.1177, F1 점수 0.81인 문장 감정 이진 분류모델이 생성되었다. 데이터셋을 5배 늘려서 전이 학습시킨 결과, 정확도 0.9562, 손실 0.1202, F1 점수 0.86인 모델을 얻었다.

Recently, generative models based on the Transformer architecture, such as ChatGPT, have been gaining significant attention. The Transformer architecture has been applied to various neural network models, including Google's BERT(Bidirectional Encoder Representations from Transformers) sentence generation model. In this paper, a method is proposed to create a text binary classification model for determining whether a comment on Korean movie review is positive or negative. To accomplish this, a pre-trained multilingual BERT sentence generation model is fine-tuned and transfer learned using a new Korean training dataset. To achieve this, a pre-trained BERT-Base model for multilingual sentence generation with 104 languages, 12 layers, 768 hidden, 12 attention heads, and 110M parameters is used. To change the pre-trained BERT-Base model into a text classification model, the input and output layers were fine-tuned, resulting in the creation of a new model with 178 million parameters. Using the fine-tuned model, with a maximum word count of 128, a batch size of 16, and 5 epochs, transfer learning is conducted with 10,000 training data and 5,000 testing data. A text sentiment binary classification model for Korean movie review with an accuracy of 0.9582, a loss of 0.1177, and an F1 score of 0.81 has been created. As a result of performing transfer learning with a dataset five times larger, a model with an accuracy of 0.9562, a loss of 0.1202, and an F1 score of 0.86 has been generated.

18

음악가사와 오디오 특성 변수의 통합 분석 모델 구축에 관한 연구: 텍스트 마이닝과 기계학습을 중심으로

김성범

[NRF 연계] 한국엔터테인먼트산업학회 한국엔터테인먼트산업학회논문지 Vol.16 No.8 2022.12 pp.271-283

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구의 목표는 음악의 가사(lyrics)와 오디오(a특성을 변수화하여 통합한 후 음악의 특성을 예측하는 모델을 구축하는 것이다. 이 연구는 텍스트 마이닝을 이용한 가사 분석과 스포티파이가 제공하는 오디오 특성치 변수를 통합하여 두 변수 간의 통합 모델을 제시한다. 연구1에서는 텍스트 마이닝의 군집분석과 감성분석을 사용하여 음악 가사로부터 변수를 도출하고, 연구2에서는 상관분석을 사용한다. 마지막으로 연구3에서는 연구1,2에서 도출한 입력변수와 기계학습 방법론을 사용하여 음악의 장르를 추정한다. 가사 군집변수와 감성변수가 포함되면 오디오 특성만을 사용한 장르 예측에 비하여 정확도를 높일 수 있음을 확인하였고, 이 연구에서 RnB 힙합 장르가 고유한 가사 군집을 구성하는 것처럼, 특정 장르가 가사 내용과 단어 선택의 차별화를 가지고 있을 때, 이 해당 가사의 특성이 오디오 특성보다 장르 추정에 기여도가 높다는 것이 확인되었다. 이 연구는 광고를 포함한 특정 용도에 필요한 음악을 예측하는 작업으로 확장할 수 있다. 실무적인 활용성을 높이기 위하여 다양한 음악을 수집하고 그에 적합한 입력변수 도출 작업의 노력이 병행되어야 한다.

The objective of this study is to build a model that could predict the characteristics of music by extracting and integrating the lyrics and audio characteristics. The study presents an integrated model between two variables by implementing lyrics analysis using text mining and audio features provided by Spotify. In Study 1, variables were derived from lyrics using cluster analysis and sentiment analysis. In Study 2, correlation analysis was used to explore the relationship between lyrics and audio features. And finally, in Study 3, musical genres are predicted using the input variables derived from Studies 1 and 2 as well as machine learning methodology. We confirmed that lyrical themes and sentiment variables could increase the accuracy of the estimation model in comparison to genre prediction that relied only on audio characteristics. Furthermore, we found that lyrical characteristics contribute more to genre estimation than audio characteristics when a specific genre has differentiation in its lyrical content. We believe that this study can be extended to predict music needed for specific applications, including in advertising. To increase the necessary practical utilization, efforts should be made in parallel to collect various music type and derive appropriate input variables.

 
페이지 저장