년 - 년
자동 분류 기술을 활용한 온라인 강의 평가 방법 KCI 등재
한국정보교육학회 정보교육학회논문지 제24권 제4호 2020.08 pp.291-300
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
국내외 온라인 강의에 대한 학습자와 프로그램 수요는 증가하고 있지만 이에 대한 평가 방법은 설문지에 의한 정량적인 수치에 의존하고 있으며 객관적인 학습 만족도에 대한 평가 방법은 마련돼 있지 않다는 것이 문 제점으로 드러나고 있다. 본 연구에서는 온라인 학습 시스템의 게시판에 있는 빅 데이터 메시지를 분석하여 온 라인 강의를 평가하는 방법을 제안하려고 한다. 실제로 빅 데이터 분석기법 중 중요한 기술로 인식되는 자동분 류 기법을 적용하여 온라인 강의 평가에 시범 적용해 보았으며 델파이 분석 결과에서도 평가 항목과 분류 결과 등이 온라인 강의 평가에 적합하고 학교나 기관에서 적용해볼 만하다는 결론을 얻었다. 본 연구는 빠르게 축적 되고 있는 빅 데이터 분석기술을 가장 변화가 늦은 교육 분야에 적용해 보고 확장 가능성을 진단해보는데 의의 가 있다.
Although the need for international online courses and the number of online learners has been rapidly increasing, the online class evaluation has been mostly relying on the quantitative survey analysis. So a more objective evaluation method has to be developed to more accurately assess online course satisfaction. This study highlights the benefits of using big data analysis from the bulletin board messages of online learning system as a method to evaluate the online courses. In fact, automatic classification technology is recognized as an important technology among big data analysis techniques. Our team applied this technique to evaluate the online courses. From the delphi analysis results, suggested method was concluded that the evaluation items and classification results are suitable for online course evaluation and applicable in schools or institutions. This study has confirmed that the rapidly accumulating big data analysis technology can be successfully applied to the education sector with the least change. It also diagnosed a meaningful possibility to expand the big data analysis for further application.
CEFRに対応した日本語例文自動分類システムのBERT適用による精度改善の試み KCI 등재
한국일본학회 일본학보 제137권 2023.11 pp.43-62
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
최근 CanDo에 의한 언어 능력 척도의 하나로 CEFR(유럽 언어 공통 기준)가 관심을 받고 있으며, 외국어 교육에 분야에서도 전세계적으로 도입되고 있는 것은 잘 알려진 사 실이다. 한편, 일본어 교육을 위한 CEFR의 연구 사례는 많지 않으며, 일본어 CEFR 준거 텍스트 코퍼스에 관한 연구도 거의 이루어지고 있지 않다. 본 연구에서는 코퍼스를 작성 할 때, 예문에 CEFR의 독해력을 나타내는 CDS(능력 기술문)에 관한 정보를 부여하는 것 이 상정되었을 때 그 노력을 경감하기 위한 자동 분류의 구현을 지속적으로 연구하고 있 다. 분류를 위한 특징량으로 현재 문서 타입, 전문성, 문장, 한자율을 채용하고 있으며, 그 속의 문서 타입이나 전문성 분류를 위해 fastText가 이용되고 있다. 이에 대해 본 논문 에서는 지금까지 많은 자연언어 처리 태스크에서 유효하게 동작하는 것이 확인되어 정평 이 있는 BERT 알고리즘(Bidirectional Encoder Representations from Transformers)을 새롭게 적용하여 시도한 결과, 예측 정밀도를 향상시키는 것에 성공하였으며, 예측한 CDS의 수 를 보다 적절하게 억제할 수 있었다. 향후 예측할 수 있는 수를 줄임으로써 적합률을 높 이고, 보다 효과적인 특징량의 작성을 검토하여 CDS 정보가 더 많은 데이터를 수집하고 자 한다. 또한 이번 연구 대상은 CEFR가 초기에 설정한 6단계의 언어 능력 레벨(초급 레 벨의 A1부터 A2, B1, B2, C1와 최상급 레벨의 C2까지) 중, 모국어화자라도 난이도가 높 은 C1과 C2를 뺀 A1, B2을 더하여 2017년 CEFR을 보완한 것으로 공개된 CEFR Companion Volume에서 새롭게 추가된 PreA1 레벨도 포함하였다. 또한 기술 항목은 Reading, Writing, Speaking, Listening, Interaction 중 Reading에 초점을 맞추고, 그에 대응하 는 CDS는 34으로 세었다. 즉, 본 연구는 입력 예문을 34개의 라벨로 분류하는 것을 중심 으로 하고 있다는 것을 의미한다.
The Common European Framework of Reference for Languages (CEFR) is an example of a Can-Do language proficiency scale that has attracted large attention in recent years and has been introduced in foreign language education worldwide. On the other hand, there are a few examples of CEFR research related to Japanese language education. As far as the present authors investigated, there is no Japanese CEFR-compliant text corpus. In the current research, to create a corpus, we focused on the implementation of automatic classification in order to reduce the effort of adding Can-Do Statement (CDS), thereby enhancing the reading comprehension of CEFR in the example sentences. Document type, specialty, sentence length, and Kanji rate are commonly used as the parameters for classification; however, the current version of the aforementioned implementation uses fastText to identify document type and specialty. The present study attempted to apply the Bidirectional Encoder Representations from Transformers (BERT) algorithm, which has been confirmed to work effectively in many natural language processing tasks. The findings of the research showed that prediction accuracy was improved and it was possible to suppress the number of CDSs for more appropriate prediction. Prospects include improving the precision by further narrowing down the number of predictions, creating more compelling features, and collecting more data using CDS information. The targets in this study were CEFR proficiency levels A1, A2, B1, B2, and PreA1 for Reading skill items (34 CDSs correspond to those levels).
6,000원
최근 학습중인 언어를 사용하여 구체적으로 무엇을 할 수 있는지를 나타내는 범용 체계에 큰 관심이 모아지고 있다. 그 중에서도 2001년에 유럽위원회가 발표한 Common European Framework of Reference for Languages (CEFR)는 언어능력의 국제표 준으로 세계적으로 평가가 높다. 2017년에는 그것을 보완하는 CEFR Companion Volume이 공개되어, PreA1 레벨이 추가되는 등 더욱 더 레벨이 세분화되었다. CEFR 를 사용한 연구와 실천 예는 영어를 비롯한 많은 언어에서 이루어지고 있는 반면, 일 본어 교육을 염두에 둔 CEFR연구는 수적으로도 여전히 적으며, 일본어 CEFR준수 텍 스트 코퍼스도 현재까지의 연구 결과 존재하지 않는다. 본 연구에서는, 코퍼스를 작성 할 때 발생하는, 예문에 CEFR의 독해력을 반영하는 Can-Do Statements (CDS)를 부여 하는 노력을 경감하기 위해 자동분류 실장(実装)에 대해 지속적으로 연구하고 있다. 분류 방법에는 Support Vector Machine과 랜덤 포레스트에 의한 지도 학습을 적용하 고, 기계 학습을 위한 예문의 특징량으로써 문서 유형, 전문성, 문장 길이, 한자 비율 4 개를 사용한다. Pre-A1 레벨은 종전 레벨 군과 난이도와 구성언어요소에 큰 차이가 있기 때문에, 모든 레벨의 CDS를 한 번에 분류하는 과거 방식에 비해 2 단계에 따른 CDS 분류에 따른 정확도 향상을 목표로 하였다. 또한, Web 어플리케이션의 개발을 실시하여, 자동 분류 알고리즘을 내부 구현함으로써, 주어진 예문에 대하여 이에 대응 하는CDS를 자동 부여하는 기능과, 특정 CDS를 선택함으로써 이에 해당하는 예문 리 스트를 그 확실성 순으로 제공하는 기능을 제공하고 있다.
In recent years, a lot of attention is being paid to general-purpose frameworks that show what can be concretely done in using a target language for learners. In 2017, the Common European Framework of Reference for Languages (CEFR) Companion Volume was released. This volume complements the CEFR initially published in 2001, which is widely considered as an international standard for language ability, and introduces a Pre-A1 level. Conversely, there are few studies on CEFR for Japanese language education, and from the past studies, it was noted that there are no Japanese CEFR compliant text corpora. Thus, the present study aims to classify example sentences according to their corresponding Can-Do Statements (CDSs) to reduce efforts in creating a corpus. Support Vector Machine and Random Forest were applied to the classification approach where document types, specialty, sentence length, and kanji ratio have been given as the features of example sentences. The Pre-A1 level has a great difference in difficulty level and constituent language elements from the previous level groups. Therefore, our study seeks to improve the accuracy through binary classification combined with incorporation of the past method of classifying all levels of CDSs at once. Moreover, we also developed a web application that would help attach CDSs efficiently to example sentences and provide example sentence collections corresponding to specific CDSs.
4,000원
일반적으로 자동분류는 학습문서의 개수에 영향을 받는다고 알려져 있지만 실제로 학습문서의 수가 텍스트 자동분류에 어떻게 영향을 주는지 입증한 연구는 거의 없었다. 본 연구에서는 학습문서 수가 자동분류에 어떤 영향을 주는지 알아보기 위해 최근에 개발된 편차기반 분류방법을 중심으로 다른 분류 알고리즘과 비교하는데 초점을 두었다. 실험결과, 편차기반 분류모델은 학습문서의 수가 총 21개(7개 장르)인 상황에서 정확도가 0.8로 베이지안이나 지지벡터기계보다 우수하게 나타났다. 이것은 편차기반 분류모델이 장르내의 주제정보를 이용하여 학습하기 때문에 학습문서의 수가 적더라도 다른 학습방법보다 좋은 자질 선택 능력을 갖는다는 것을 입증한 것이다.
It is generally accepted that classification accuracy is affected by the number of learning documents, but there are few studies that show how this influences automatic text classification. This study is focused on evaluating the deviation-based classification model which is developed recently for genre-based classification and comparing it to other classification algorithms with the changing number of training documents. Experiment results show that the deviation-based classification model performs with a superior accuracy of 0.8 from categorizing 7 genres with only 21 training documents. This exceeds the accuracy of Bayesian and SVM. The Deviation-based classification model obtains strong feature selection capability even with small number of training documents because it learns subject information within genre while other methods use different learning process.
[NRF 연계] 한국통신학회 ICT Express Vol.9 No.5 2023.10 pp.834-840
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Due to the influence of noise in the received signal in non-cooperative communication, it is difficult for existing Automatic modulation classification methods to balance classification accuracy and model complexity. This paper proposes a novel Convolutional Adaptive Noise Reduction (CANR) network, which consists of an Adaptive Noise Reduction (ANR) module and a Convolutional Feature Extraction (CFE) module. The ANR and CFE modules denoise the combined input and capture the spatiotemporal features in the time series. Experiments on benchmark datasets show that the proposed network has the fewest training parameters and state-of-the-art recognition accuracy under the same conditions.
자동분류 알고리즘을 이용한 지능형 정보검색시스템 구축에 관한 연구
[NRF 연계] 한국도서관·정보학회 한국도서관·정보학회지 Vol.39 No.4 2008.12 pp.283-304
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구의 목적은 이용자의 탐색 행태, 시스템의 정보 구축 행태를 기반으로 초기 질의어의 범주에 해당하는 연관 용어들(해당 용어의 지식구조와 관련된 연관 용어들)을 학습기능을 통해 자동으로 제시해 줄 수 있는 지능형 검색 시스템을 구현하는 것이다. 이를 위해 학습을 통해 전문가 수준의 색인어를 추출할 수 있는 지능형자동색인 알고리즘, 자동분류에 관련한 클러스터링 알고리즘과 문서 범주화 알고리즘 그리고 범주 표현 알고리즘에 대한 이론적 연구를 수행하였으며, 이들 이론적 연구를 근거로 비용과 시간적인 측면에서 그리고 재현율과 정도율이란 측면에서 우수한 성능을 발휘할 수 있는 지능형검색시스템을 구현하였다.
This is to develop Intelligent Retrieval System which can automatically present early query's category terms(association terms connected with knowledge structure of relevant terminology) through learning function and it changes searching form automatically and runs it with association terms. For the reason, this theoretical study of Intelligent Automatic Indexing System abstracts expert's index term through learning and clustering algorism about automatic classification, text mining(categorization), and document category representation. It also demonstrates a good capacity in the aspects of expense, time, recall ratio, and precision ratio.
자동분류기반 성격 유형별 도서추천시스템 개발을 위한 실험적 연구
[NRF 연계] 한국도서관·정보학회 한국도서관·정보학회지 Vol.48 No.2 2017.06 pp.215-236
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
이 연구의 목적은 개인별 성향이나 성격 유형에 따라 선호하는 도서에 차이가 있음을 전제로, 어린이·청소년을 위한 추천도서의 책소개 정보를 활용하여 개인별 성격유형에 적합한 도서를 합리적으로 추천할 수 있는 서평 자동분류시스템을 개발하는 것이다. 연구에서 사용한 데이터는 국립어린이청소년도서관에서 제공하는 501권의 유아 및 아동도서를 대상으로 하였다. 실험에 활용된 2가지 기계학습 모델(비선형 커널 및 선형 커널) 각각에 대해서 총 6가지의 색인어 가중치 계산 방법과 자질 선택 방법, 그리고 10가지의 자질 선정 임계치 조합으로 구성된 360 개의 분류 모델들을 구성하고 각각의 성능을 측정하였다. 전체적으로는 선형 커널을 이용한 SVM 기반 학습 방법(LIBLINEAR)이 비선형 분류를 지원하는 LibSVM(RBF 커널) 모델보다 더 나은 성능을 보이는 것으로 나타났다. 다만 성능 측정 결과는 뉴스 기사나 논문을 대상으로 한 문헌 분류 성능에 비해서 낮은 것으로 나타났으나, 합리적인 분류 기준이 존재하는 뉴스기사나 주제 분류에 비해서 성격 유형 기반 분류는 그 난이도가 높다는 것을 감안할 때, 초기 실험 결과로서의 의미는 있다.
The purpose of this study is to develop an automatic classification system for recommending appropriate books of 9 enneagram personality types, using book information data reviewed by librarians. Data used for this study are book review of 501 recommended titles for children and young adults from National Library for Children and Young Adults. This study is implemented on the assumption that most people prefer different types of books, depending on their preference or personality type. Performance test for two different types of machine learning models, nonlinear kernel and linear kernel, composed of 360 clustering models with 6 different types of index term weighting and feature selections, and 10 feature selection critical mass were experimented. It is appeared that LIBLINEAR has better performance than that of LibSVM(RBF kernel). Although the performance of the developed system in this study is relatively below expectations, and the high level of difficulty in personality type base classification take into consideration, it is meaningful as a result of early stage of the experiment.
AttentionMesh를 활용한 국가과학기술표준분류체계 소분류 키워드 자동추천에 관한 연구
[NRF 연계] 한국도서관·정보학회 한국도서관·정보학회지 Vol.53 No.2 2022.06 pp.95-115
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
이 연구의 목적은 국가과학기술표준분류체계의 소분류 용어를 기계학습 알고리즘을 적용하여 기술키워드 변환하는 것이 목적이다. 이를 위해 본 연구에서는 주제어 추천에 적합한 학습 알고리즘으로 AttentionMeSH를 활용했다. 원천데이터는 한국과학기술기획평가원이 정제한 2017년부터 2020년까지 4개년 연구현황 파일을 사용하였다. 학습은 과제명, 연구목표, 연구내용, 기대효과와 같이 연구내용을 잘 표현하고 있는 4개 속성을 사용했다. 그 결과 임계치(threshold)가 0.5일 때 MiF 0.6377이라는 결과가 도출됨을 확인하였다. 향후 실제 업무에 기계학습을 활용하고, 기술키워드 확보를 위해서는 용어관리체계 구축과 다양한 속성들의 데이터 확보가 필요할 것으로 보인다.
The purpose of this study is to transform the sub-categorization terms of the National Science and Technology Standards Classification System into technical keywords by applying a machine learning algorithm. For this purpose, AttentionMeSH was used as a learning algorithm suitable for topic word recommendation. For source data, four-year research status files from 2017 to 2020, refined by the Korea Institute of Science and Technology Planning and Evaluation, were used. For learning, four attributes that well express the research content were used: task name, research goal, research abstract, and expected effect. As a result, it was confirmed that the result of MiF 0.6377 was derived when the threshold was 0.5. In order to utilize machine learning in actual work in the future and to secure technical keywords, it is expected that it will be necessary to establish a term management system and secure data of various attributes.
광고 문구의 수사학적 설득 요소 자동 분류를 위한 인공지능 기반 접근 : Zero-shot Classification의 적용 KCI 등재후보
한국융합학회 미래기술융합논문지 제4권 제6호 2025.12 pp.59-66
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 광고 문구의 수사학적 설득 요소(Ethos, Pathos, Logos)를 인공지능 기반 Zero-shot Classification 기법으로 자동 분류하는 방법을 제안한다. 총 2,870개의 광고 문구를 수집하고, 이 중 30%를 전문 가 합의 라벨로 구축하여 성능을 검증하였다. facebook/bart-large-mnli 모델과 정의문 포함 프롬프트를 적용한 결과, 정확도 87.4%, F1 점수 85.8%, 카파 계수 0.81을 기록하였다. 제안 기법은 라벨링 없이도 높은 일관성과 확장성을 보였으며, 광고 분석 및 전략 설계에 활용 가능함을 확인하였다.
This study proposes an AI-based method to automatically classify rhetorical persuasive strategies—Ethos, Pathos, and Logos—in advertising copy using Zero-shot Classification. A dataset of 2,870 advertisements was collected, with 30% manually labeled for evaluation. Using the facebook/bart-large-mnli model with definition-embedded prompts, the approach achieved 87.4% accuracy, an F1-score of 85.8%, and a Cohen’s kappa of 0.81. The results show that the method can deliver consistent, scalable analysis without labeled data, supporting applications in advertising analysis and strategy design.
T1-MRI 뇌영상으로부터 히포캠퍼스 특징 기반 알츠하이머 병 자동 분류기법 KCI 등재
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.15 No.2 2019.04 pp.17-28
알츠하이머 병 (Alzheimer's disease, AD) 에서 해마는 신경 세포 및 신경 섬유 얽힘 스레드 침착으로 인해 영향 을 받는 첫 번째 구조물 중 하나이며 결국 신경 세포의 손실을 초래합니다. 또한 최근에 많은 자기 공명 영상 연구에 서 AD 환자는 건강한 피검자에 비해 해마의 체적이 작다는 의견을 제시했다. 해마 자기 공명 영상 체적은 AD를 위한 잠재적인 바이오 마커이지만 작은 구조로 인해 수동 분할 및 낮은 시각화의 한계가 에 의해 어려움이 있다. 이 러한 문제를 해결하기 위해 대부분의 이미지 분석가는 뇌 구조의 자동 세분화를 위해 Freesurfer와 FSL에 익숙했 습니다. 본 연구에서는 다른 그룹의 AD 환자의 조기 진단을 위해 Freesurfer (v.6.0) 자동화 도구 상자를 사용하여 해마 영역을 추출하는 방법을 이용한다. 38 명의 환자가 AD에 속하며, 46 명의 환자가 MCI에 속한다 (안정-MCI, 18 개월 후에 AD로 전환되지 않음) 및 36 명의 환자가 MCIc에 속한다 (전환-MCI, 18 개월 이내에 AD로 전환), 나머지 38 명의 노인은 정상 대상 (NC)에 속하며 이러한 모든 환자들의 정보는 알츠하이머병 Neuroimaging Initiative database (ADNI)에서 다운로드 되었습니다. AD, MCIs, MCIc 및 노인 대조군 (NC) 환자를 구별하 기 위해 우리는 구입한 해마 체적을 사용했습니다. 여기서 차원적 감소를 위해 우리는 다양한 학습 기반의 ISOMAP 기법을 사용했다. 이 기법은 많은 기능들 중 중요한 특징만을 선택하고 나중에 랜덤 포레스트와 softmax 분류기를 사용하여 이진 분류 문제들을 분류했다.
In Alzheimer’s disease (AD), the hippocampus is among the first structures, which is affected, due to the neuropil and neurofibrillary tangle thread deposition, which eventually results in a neuronal loss. Moreover, recently a large number of magnetic resonance imagining studies have stated that AD patients has get smaller hippocampus volume as compared to healthy subjects. Hippocampal magnetic resonance imaging volumetric is a potential biomarker for an AD but is hindered by the limitation of manual segmentation and low visualization because of its small structure. To solve that problems, most image analyst used to Freesurfer and FSL for automatic segmentation of brain structures. In this paper, we apply to extract hippocampus region using Freesurfer (v.6.0) automated toolbox for early diagnosis of AD subjects with different groups. a total 158 baseline subjects were used for the experiment, from which 38 patients belong to AD, 46 patients belong to MCIs (stable-MCI, not converted to AD after 18th month of periods) and 36 patients belong to MCIc (converted-MCI, converted to AD within 18th month of periods), and remaining 38 patients belong to elderly normal subjects (NC), and all these subjects were downloaded from Alzheimer’s disease Neuroimaging Initiative database (ADNI). We used the gained hippocampal volumes to discriminate between AD, MCIs, MCIc, and elderly controls (NC) patients. Here, for dimensionality reduction purpose we have used manifold learning based ‘isometric feature mapping’ ISOMAP technique, which only selects the important feature from a bunch of features and later random forest and softmax classifier were used to classify the binary classification problems.
TOEFL11 코퍼스에서 주제 프롬프터의 자동 분류 KCI 등재
국제언어인문학회 인문언어 제20권 1호 2018.06 pp.157-177
※ 기관로그인 시 무료 이용이 가능합니다.
5,700원
The aim of the paper is to develop classification models of the prompts of the TOEFL essays in the TOEFL11 corpus. The corpus is a collection of TOEFL essays written in response to one of 8 prompts of various topics and by test-takers of different proficiency levels who are from 11 different countries. The number of essays is 11,000 for each language (that is, 121,000 in total). The paper aims at developing prompt classification models using an automatic method of Support Vector Machine (SVM), to which a number of different features are fed: The input features to the model include high frequency words which are observed in the raw essay texts, and high frequency nouns which are extracted from a POS-tagged essay texts. High frequency nouns among three different proficiency levels are also used as input features. The results indicated that even though high frequency words taken from raw textual materials performed quite well with an accuracy of 90.4%, the words tagged as nouns did even better with an accuracy of 97.3%. The inspecting of high frequency nouns revealed that the words were independently distributed among prompts with nearly no overlapping across different prompts. The classification test of essay samples of different proficiency levels confirmed that the accuracy rate of automatically classifying prompts by observing the frequency occurrence of nouns in texts increased in general as the proficiency levels of the essay samples increase. The paper serves as a foundation for further details studies on topic modeling used by learners of English.
SLAM을 위한 MDLP 기반의 수신 데이터 자동 분류 기법
한국ITS학회 한국ITS학회 학술대회 2017년 한국ITS학회 춘계학술대회 2017.04 pp.549-551
※ 기관로그인 시 무료 이용이 가능합니다.
3,000원
제목의 단어 가중치를 이용한 중등학교 공문서 자동분류시 스템 KCI 등재후보
한국정보교육학회 정보교육학회논문지 제7권 제2호 2003.08 pp.219-226
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
자율주행 데이터셋 자동 분류 및 딥러닝 알고리즘 통합 학습 환경 구성에 관한 연구
한국ITS학회 한국ITS학회 학술대회 Net-Zero Mobility 2023.04 pp.269-272
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
4,000원
UML(Unified Modeling Language) 클래스 다이어그램은 시스템의 정적인 측면을 표현하며 분석 및 설계 부터 문서화, 테스팅까지 사용된다. 클래스 다이어그램을 이용한 모델링이 소프트웨어 개발에 있어 필수적이지만, 경험이 많지 않은 모델러에게 쉽지 않은 작업이다. 도메인 카테고리별로 분류된 클래스 다이어그램 데이터 세트가 제공된다면, 모델링 작업의 생산성을 높일 수 있을 것이다. 본 논문은 클래스 다이어그램 이미지 데이터를 구축하기 위한 자동 분류 기술을 제공한다. 추가 정보 없이 단지 UML 클래스 다이어그램 이미지를 식별하고 도메인 카테 고리에 따라 자동 분류한다. 먼저, 웹상에서 수집된 이미지들이 UML 클래스 다이어그램 이미지인지 여부를 판단한다. 그리고, 식별된 클래스 다이어그램 이미지에서 클래스 이름을 추출하여 도메인 카테고리에 따라 분류한다. 제안된 분류 모델은 정밀도, 재현율, F1점수, 정확도에서 각각 100.00%, 95.59%, 97.74%, 97.77%를 달성했으며, 카테 고리별 분류에 대한 정확도는 81.1%와 95.2% 사이에 분포한다. 해당 실험에 사용된 클래스 다이어그램 이미지 개수가 충분히 크지 않지만, 도출된 실험 결과는 제안된 자동 분류 방식이 고려할 만한 가치가 있음을 나타낸다.
UML class diagrams are used to visualize the static aspects of a software system and are involved from analysis and design to documentation and testing. Software modeling using class diagrams is essential for software development, but it may be not an easy activity for inexperienced modelers. The modeling productivity could be improved with a dataset of class diagrams which are classified by domain categories. To this end, this paper provides a classification method for a dataset of class diagram images. First, real class diagrams are selected from collected images. Then, class names are extracted from the real class diagram images and the class diagram images are classified according to domain categories. The proposed classification model has achieved 100.00%, 95.59%, 97.74%, and 97.77% in precision, recall, F1-score, and accuracy, respectively. The accuracy scores for the domain categorization are distributed between 81.1% and 95.2%. Although the number of class diagram images in the experiment is not large enough, the experimental results indicate that it is worth considering the proposed approach to class diagram image classification.
전통문화 콘텐츠 표준체계를 활용한 자동 텍스트 분류 시스템 KCI 등재
한국융합학회 한국융합학회논문지 제8권 제12호 2017.12 pp.39-47
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
한국 문화의 역사, 전통과 관련된 디지털 웹 문서가 증가하게 되었다. 하지만 창작자 또는 전통 문화와 관련된 소재를 찾는 사용자들은 정보를 검색해도 결과가 충분하지 않았으며 원하는 정보를 얻지 못하는 경우가 나타나고 있다. 이런 효과적인 정보를 접하기 위해서는 문서 분류가 필요하다. 과거에 문서 분류는 작업자가 수작업으로 문서 분류하여 시간과 비용이 많이 소비하는 어려움이 있었지만, 최근 기계학습 기반으로 한 자동 문서 분류를 통해 효율적인 문서 분류가 이루어진다. 이에 본 논문은 전통문화 콘텐츠를 체계적인 분류체계로 구성한 한민족정보문화마당 데이터를 기반으로 전통문화 콘텐츠 자동 텍스트 분류 모델을 개발한다. 본 연구는 한민족정보문화마당 텍스트 데이터에 대해 단어 빈도수를 추출하기 위해 TF-IDF모델, Bag-of-Words 모델, TF-IDF/Bag-of-Words를 결합한 모델을 적용하여 각각 SVM 분류 알고리즘을 사용하여 전통문화 콘텐츠 자동 텍스트 분류 모델을 개발하여 성능평가를 확인하였다.
The Internet have increased the number of digital web documents related to the history and traditions of Korean Culture. However, users who search for creators or materials related to traditional cultures are not able to get the information they want and the results are not enough. Document classification is required to access this effective information. In the past, document classification has been difficult to manually and manually classify documents, but it has recently been difficult to spend a lot of time and money. Therefore, this paper develops an automatic text classification model of traditional cultural contents based on the data of the Korean information culture field composed of systematic classifications of traditional cultural contents. This study applied TF-IDF model, Bag-of-Words model, and TF-IDF/Bag-of-Words combined model to extract word frequencies for 'Korea Traditional Culture' data. And we developed the automatic text classification model of traditional cultural contents using Support Vector Machine classification algorithm.
AI 학습모델 및 AI모델 서빙 서버 개발을 통한 생활안전 예방 서비스 신고 이미지 자동분류 시스템 개발에 대한 연구 KCI 등재
한국재난정보학회 한국재난정보학회논문집 제19권 2호 통권60호 2023.06 pp.432-438
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
연구목적: 생활안전 예방서비스 앱에서 신고되는 이미지를 AI를 사용하여 실시간으로 위험 카테고리를 자동으로 분류하여 사용자에게 편리한 위험신고를 가능하게 하는 것을 목적으로 한다. 연구방법: 인터넷 으로 상호연결되는 생활안전 예방서비스 플랫폼, 생활안전 예방서비스 앱, AI 모델 서빙 서버와 sftp 서버 로 구성되는 시스템을 통하여 신고된 생활안전 이미지를 실시간으로 자동분류하며, 이때 사용되는 AI모 델 생성을 위한 AI 학습 알고리즘도 개발하였다. 연구결과: 이미지를 실시간으로 AI 처리하여 자동으로 분류할 수 있게 되어, 신고자가 생활안전 관련 사항을 보다 편리하게 신고할 수 있게 되었다. 결론: 본 논문 에서 제시하는 AI 이미지 자동분류 시스템은 90% 이상의 분류 정확도로 신고 이미지를 실시간으로 자동 분류하여 신고자가 간편하게 생활안전 관련 이미지를 신고할 수 있게 되었으며 향후 생활안전 예방서비 스 앱의 사용자의 증가에 따라 더욱 빠르고 정확한 AI 모델 개발 및 시스템 처리용량 향상이 필요하다.
Purpose: The purpose of this study is to enable users to conveniently report risks by automatically classifying risk categories in real time using AI for images reported in the life safety prevention service app. Method: Through a system consisting of a life safety prevention service platform, life safety prevention service app, AI model serving server and sftp server interconnected through the Internet, the reported life safety images are automatically classified in real time, and the AI model used at this time An AI learning algorithm for generation was also developed. Result: Images can be automatically classified by AI processing in real time, making it easier for reporters to report matters related to life safety.Conclusion: The AI image automatic classification system presented in this paper automatically classifies reported images in real time with a classification accuracy of over 90%, enabling reporters to easily report images related to life safety. It is necessary to develop faster and more accurate AI models and improve system processing capacity.
AI 어시스턴트의 요청문 인식 및 분류 주석을 위한 음악청취 관련 문형 패턴 연구 KCI 등재
한국언어과학회 언어과학 제27권 1호 2020.02 pp.95-132
※ 기관로그인 시 무료 이용이 가능합니다.
8,200원
This study aims to describe music-related request patterns observed in AI assistant platforms and formalize them by using the local grammar graph(LGG) methodology. These patterns, transformed into Finite-State Transducers in the Unitex platform, allow us to automatically classify and annotate various sentences used for request, which is crucial for diverse machine learning approaches. In this study, we examined a corpus of tweet texts crawled by Deco-T-Crawler as seed data and classified music-related requests into 3 categories: music on/off, volume up/down and music-play control. Then each category was divided into sub-categories which were represented by several LGGs. By applying these LGGs and the DECO Korean dictionary on test data, we obtained the information retrieval performance. In this regard, the linguistic resources proposed in this study are proved concrete and meaningful.
Word2Vec과 2계층 양방향 장단기 기억 네트워크를 이용한 특허 문서의 자동 IPC 분류 KCI 등재
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.15 No.2 2019.04 pp.50-60
자연어 처리를 이용한 문서 분류 분야에서도 전통적인 방법에서 벗어나 단어 임베딩을 활용한 합성곱 신경망과 순환 신경망 등 심층 신경망을 이용한 다양한 연구가 진행되고 있다. 본 논문에서는 Word2Vec과 두 개의 계층으로 구성 된 양방향 장단기 기억 네트워크를 이용한 특허 문서의 IPC(International Patents Classification) 자동 분류 모델을 제안한다. IPC는 세계지식재산권기구에서 제정한 국제적으로 통일된 특허 분류 기준이며, 각 국가의 공인된 기관에서 수작업으로 분류하고 있다. IPC 자동 분류를 위하여 입력 시퀀스에 Word2Vec을 이용한 단어 임베딩가 중치를 사용한다. 그리고 가중치가 부여된 시퀀스를 두 개의 계층을 갖는 깊은 구조의 양방향 장단기 기억 네트워크 신경망에 입력하여 IPC를 분류한다. 실험 결과 특허 문서의 분류 정확도가 합성곱 신경망 보다는 약 7% 향상되었 으며, 순환 신경망을 단일로 이용하는 것 보다는 약 5% 향상된 것을 확인할 수 있었다. 또한 전통적인 방법인 나이 브 베이시안, 로지스틱 분류 및 서포트 벡터 머신보다는 5~12% 이상 우수한 성능을 나타내었다.
There are various studies using Deep Neural Network such as CNN(Convolutional Neural Network) and RNN(Recurrent Neural Network) that utilize word embedding in document classification using natural language processing out of traditional methods. In this paper, we propose the IPC(International Patents Classification) automatic classification model of patent documents using two layers BLSTM (Bidirectional Long Short Term memory) network. The IPC is an internationally uniform standard for patent classification established by the World Intellectual Property Organization and is categorized by hand in authorized agencies in each country. For the IPC automatic classification, we use word embedding weight with Word2Vec in the input sequences. And they are classified by entering a weighted sequences into a deep neural network with two layers BLSTM. The experimental results showed that the accuracy of classification is improved by about 7% than that of CNN, and about 5% than that of single layer LSTM that is a field of RNN. Also it showed more than 5~12% higher performance than traditional methods such as Naive Bayes, Logistic and Support Vector Machine classification.
딥 러닝 기반 이미지 자동 분류 및 랭킹 시스템을 이용한 사용자 편의 중심의 유실물 등록 및 조회 관리 시스템 KCI 등재
한국융합보안학회 융합보안논문지 제18권 제4호 2018.10 pp.19-25
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 논문은 딥 러닝(Deep-Learning) 기반의 계층형 이미지 분류 체계와 가중치 기반의 랭킹 시스템을 이용한 사용자 편의 중심의 유실물 등록 및 조회 관리 시스템을 제안한다. 제안된 시스템은 딥 러닝을 통해 이미지를 자동으로 분류하는 계층형 이미지 분류 시스템과 조회 과정의 편의를 위해 시스템상의 등록된 유실물 정보를 고려해 가중치 순으로 정렬하는 랭킹 시스 템 모듈로 구성된다. 등록 과정에서 한 장의 사진만으로 카테고리 분류와 브랜드, 연관 태그 등 여러 정보가 자동으로 인식되 어 사용자의 번거로움을 최소화하였다. 그리고 랭킹 시스템을 통해 사용자들이 자주 찾는 유실물을 상위에 노출함으로써 유실 물 검색의 효율성을 높였다. 실험 결과, 제안된 시스템은 사용자가 쉽고 편리하게 시스템을 이용할 수 있음을 확인하였다.
In this paper, we propose an user-centered integrated lost-goods management system through a ranking system based on weight and a hierarchical image classification system based on Deep Learning. The proposed system consists of a hierarchical image classification system that automatically classifies images through deep learning, and a ranking system modules that listing the registered lost property information on the system in order of weight for the convenience of the query process.In the process of registration, various information such as category classification, brand, and related tags are automatically recognized by only one photograph, thereby minimizing the hassle of users in the registration process. And through the ranking systems, it has increased the efficiency of searching for lost items by exposing users frequently visited lost items on top. As a result of the experiment, the proposed system allows users to use the system easily and conveniently
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.