년 - 년
Semantic Feature Analysis for Multi-Label Text Classification on Topics of the Al-Quran Verses
[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.20 No.1 2024 pp.1-12
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Nowadays, Islamic content is widely used in research, including Hadith and the Al-Quran. Both are mostly used in the field of natural language processing, especially in text classification research. One of the difficulties in learning the Al-Quran is ambiguity, while the Al-Quran is used as the main source of Islamic law and the life guidance of a Muslim in the world. This research was proposed to relieve people in learning the Al-Quran. We proposed a word embedding feature-based on Tensor Space Model as feature extraction, which is used to reduce the ambiguity. Based on the experiment results and the analysis, we prove that the proposed method yields the best performance with the Hamming loss 0.10317.
머신러닝 기반의 기업 리뷰 다중 분류 : 부분 문법 적용을 중심으로 KCI 등재
한국경영정보학회 경영정보학연구 제25권 제3호 2023.08 pp.27-41
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
최근 많은 분야에서 기계학습에 대한 연구가 활발히 진행되고 있는데, 상당수의 연구들이 학습 모델의 성능을 개선하는 최신 방법론을 제시하고 있다. 본 연구에서는 방법론의 개발 못지않게 기계학습에 투입되는 훈련용 데이터의 ‘품질’을 개선하는 것 역시 중요하다는 점에 착안하여, 코퍼스 분석에서 자주 사용되는 ‘부분 문법’ 처리 프로세스를 통해 훈련 데이터의 품질을 향상시키는 방법을 제시한다. 우리나라 100대 기업에 근무하는 재직자들이 채용플랫폼에 게시하는 방대한 양의 비정형 기업 리뷰 텍스트 데이터를 수집하고, 데이터 품질을 부분 문법 프로세스로 개선한 후, 부분 문법이 적용된 분류 모델이 적용되지 않은 모델보다 분류 성능이 우수함을 확인하였다. 분류 카테고리는 직원 몰입의 5가지 요인으로 상정하였는데, 국내 직장인들이 기업 리뷰가 각 유형별로 빈도에 차이가 있는지를 분석하였다. 추가로 리뷰 양상이 코로나 팬데믹 전후로 어떠한 변화가 있었는지도 분석하였다. 본 연구를 통해 국내 직장인들의 생생한 일터 경험들을 자동적으로 식별하고 분류하여, 이직을 포함한 주요한 조직문화 현상의 행태와 유발 원인 등을 유추해 볼 수 있는 근거를 제공한다.
Unlike the previous works focusing on the state-of-the-art methodologies to improve the performance of machine learning models, this study improves the ’quality' of training data used in machine learning. We propose a method to enhance the quality of training data through the processing of 'local grammar,' frequently used in corpus analysis. We collected a vast amount of unstructured corporate review text data posted by employees working in the top 100 companies in Korea. After improving the data quality using the local grammar process, we confirmed that the classification model with local grammar outperformed the model without it in terms of classification performance. We defined five factors of work engagement as classification categories, and analyzed how the pattern of reviews changed before and after the COVID-19 pandemic. Through this study, we provide evidence that shows the value of the local grammar-based automatic identification and classification of employee experiences, and offer some clues for significant organizational cultural phenomena.
효과적인 애스팩트 마이닝을 위한 다중 레이블 분류접근법 KCI 등재
한국경영정보학회 경영정보학연구 제22권 제3호 2020.08 pp.81-97
※ 기관로그인 시 무료 이용이 가능합니다.
5,100원
최근의 감성분류 연구는 출력변수가 하나인 단일레이블 분류방법을 사용한 연구가 많다. 특히, 이러한 연구는 하나의 극성 값(긍정, 부정)만을 찾는 연구가 많다. 그러나 한 문장 안에는 다중적인 의미가 내포되어 있다. 그 중에서도 감정과 오피니언이 이러한 특징을 갖는다. 본 논문은 두 가지 연구목적을 제시한다. 첫째, 한 문장 안에 다양한 토픽(주제 또는 애스팩트)이 있다는 사실을 기반으로, 해당 문장을 각 애스팩트 별로 감성을 분류하는 애스팩트 마이닝을 수행한다. 둘째, 두개 이상의 종속변수(출력 값)를 한 번에 분석하는 다중레이블 분류방법을 적용한다. 이에 본 연구는 감성분류의 연구가 단일분류기에 의해서만 이루어진 연구를 개선하고자 다중레이블 분류방법에 의한 애스팩트 마이닝을 수행하고자 한다. 이와 같은 연구목적을 달성하기 위해 국내 뮤지컬 데이터를 수집하였다. 분석결과 문장 안에 있는 다양한 애스팩트별 감성을 추출하였고, 유의한 결과를 얻었다.
Recent trends in sentiment analysis have been focused on applying single label classification approaches. However, when considering the fact that a review comment by one person is usually composed of several topics or aspects, it would be better to classify sentiments for those aspects respectively. This paper has two purposes. First, based on the fact that there are various aspects in one sentence, aspect mining is performed to classify the emotions by each aspect. Second, we apply the multiple label classification method to analyze two or more dependent variables (output values) at once. To prove our proposed approach’s validity, online review comments about musical performances were garnered from domestic online platform, and the multi-label classification approach was applied to the dataset. Results were promising, and potentials of our proposed approach were discussed.
기계학습 기반 다중 레이블 분류를 이용한 실시간 전략 게임에서의 상대 행동 예측 KCI 등재
한국융합학회 한국융합학회논문지 제11권 제10호 2020.10 pp.45-51
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
최근 많은 게임이 사용자의 게임 플레이와 관련된 데이터를 제공하고 있고, 이에 기계학습 기법을 결합하여 상대의 행동을 예측하는 연구들이 있다. 본 연구는 실시간 전략 게임(클래시로얄)의 경기 데이터와 기계학습 기반의 다 중 레이블 분류를 사용하여 상대 플레이어의 행동을 예측한다. 초기 실험은 이진 형태의 카드 특성과 카드 배치 좌표 그리고 정규화된 시간 정보를 입력받아 카드 타입, 카드 배치 좌표를 랜덤포레스트와 다층 퍼셉트론을 이용하여 예측한 다. 이후, 순차적으로 3 가지 전처리 방식을 사용하여 실험을 진행했다. 먼저 입력 데이터의 특성 정보 일부를 변환 시켜 예측했다. 다음으로 입력 데이터를 연속된 카드 입력 방식까지 고려한 중첩 형태로 변환 시켜 예측했다. 마지막으 로 모든 이전 단계의 데이터들을 정규화된 시간 기준에 따라 초반, 후반으로 분할하여 예측했다. 그 결과 가장 개선을 보인 전처리 방식은 중첩 형태의 데이터를 초반으로 분할하였을 경우로 카드 타입이 약 2.6%, 카드 배치 좌표가 약 1.8% 개선을 보였다.
Recently, many games provide data related to the users' game play, and there have been a few studies that predict opponent move by combining machine learning methods. This study predicts opponent move using match data of a real-time strategy game named ClashRoyale and a multi-label classification based on machine learning. In the initial experiment, binary card properties, binary card coordinates, and normalized time information are input, and card type and card coordinates are predicted using random forest and multi-layer perceptron. Subsequently, experiments were conducted sequentially using the next three data preprocessing methods. First, some property information of the input data were transformed. Next, input data were converted to nested form considering the consecutive card input system. Finally, input data were predicted by dividing into the early and the latter according to the normalized time information. As a result, the best preprocessing step was shown about 2.6% improvement in card type and about 1.8% improvement in card coordinates when nested data divided into the early.
ReliefF-based Multi-label Feature Selection SCOPUS
보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.8 No.4 2015.08 pp.307-318
※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.
In recent years, multi-label learning has been used to deal with data attributed to multiple labels simultaneously and has been increasingly applied to various applications. As many other machine learning tasks, multi-label learning also suffers from the curse of dimensionality; so extracting good features using multiple labels of the datasets becomes an important step prior to classification. In this paper, we study the problem of multi-label feature selection for classification and have proposed a method based on single label feature selection ReliefF, termed ML-ReliefF, to select discriminant features in order to boost multi-label classification accuracy. Compared to other multi-label feature selection methods that only consider the relationship between pairwise classes, the proposed method introduces the concept of label set to further consider the relationship among more than two labels, modifies the regulation of the nearest neighbors computation reflecting the influence between samples and multiple labels, and considers and adds the similarity between samples to reinforce the effect. With the classifier, ML-kNN, experiments on five different datasets show that the proposed method is effective in removing irrelevant or redundant features and the selected features are more discriminant for classification.
Multi-Label Classification Approach to Location Prediction
[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.22 No.10 2017 pp.121-128
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this paper, we propose a multi-label classification method in which multi-label classification estimation techniques are applied to resolving location prediction problem. Most of previous studies related to location prediction have focused on the use of single-label classification by using contextual information such as user's movement paths, demographic information, etc. However, in this paper, we focused on the case where users are free to visit multiple locations, forcing decision-makers to use multi-labeled dataset. By using 2373 contextual dataset which was compiled from college students, we have obtained the best results with classifiers such as bagging, random subspace, and decision tree with the multi-label classification estimation methods like binary relevance(BR), binary pairwise classification (PW).
HANDWRITTEN HANGUL RECOGNITION MODEL USING MULTI-LABEL CLASSIFICATION
[Kisti 연계] 한국산업응용수학회 Journal of the Korean society for industrial and applied mathematics Vol.27 No.2 2023 pp.135-145
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Recently, as deep learning technology has developed, various deep learning technologies have been introduced in handwritten recognition, greatly contributing to performance improvement. The recognition accuracy of handwritten Hangeul recognition has also improved significantly, but prior research has focused on recognizing 520 Hangul characters or 2,350 Hangul characters using SERI95 data or PE92 data. In the past, most of the expressions were possible with 2,350 Hangul characters, but as globalization progresses and information and communication technology develops, there are many cases where various foreign words need to be expressed in Hangul. In this paper, we propose a model that recognizes and combines the consonants, medial vowels, and final consonants of a Korean syllable using a multi-label classification model, and achieves a high recognition accuracy of 98.38% as a result of learning with the public data of Korean handwritten characters, PE92. In addition, this model learned only 2,350 Hangul characters, but can recognize the characters which is not included in the 2,350 Hangul characters
A Novel Posterior Probability Estimation Method for Multi-label Naive Bayes Classification
[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.23 No.6 2018 pp.1-7
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
A multi-label classification is to find multiple labels associated with the input pattern. Multi-label classification can be achieved by extending conventional single-label classification. Common extension techniques are known as Binary relevance, Label powerset, and Classifier chains. However, most of the extended multi-label naive bayes classifier has not been able to accurately estimate posterior probabilities because it does not reflect the label dependency. And the remaining extended multi-label naive bayes classifier has a problem that it is unstable to estimate posterior probability according to the label selection order. To estimate posterior probability well, we propose a new posterior probability estimation method that reflects the probability between all labels and labels efficiently. The proposed method reflects the correlation between labels. And we have confirmed through experiments that the extended multi-label naive bayes classifier using the proposed method has higher accuracy then the existing multi-label naive bayes classifiers.
협동 학습 방법을 이용한 문서의 멀티 라벨 분류의 개선
[Kisti 연계] 한국정보과학회 정보과학회논문지:데이타베이스 Vol.40 No.6 2013 pp.406-410
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 연구에서는 멀티 라벨 문서의 분류학습을 위한 다항 나이브 베이스 알고리즘의 성능을 향상시키는 새로운 방법을 제안한다. 멀티 라벨의 분류학습에서는 클래스 값들 간의 종속성을 효과적으로 활용하는 것이 학습의 성능에 많은 영향을 미치며 따라서 본 연구에서는 클래스 값들 사이의 종속성을 협동 학습의 기법을 이용하여 향상시키는 새로운 방법을 제시한다. 다수의 멀티 라벨 문서를 이용한 실험에서 본 연구의 방법은 다른 알고리즘들에 비하여 좋은 성능을 보임을 알 수 있었다.
In this paper, we propose a new method for improving the performance of multi-label classification learning using the multinomial naive Bayesian model. Incorporating the dependencies among class values is an important issue in the multi-label classification, and we employ a co-training method for this purpose. A number of multi-label document data sets are selected for testing the performance of the method. Our experiments show that the proposed method outperforms the state-of-the-art methods.
NSGA-II 알고리즘을 이용한 다중 레이블 분류 문제에서 특징 선별
[Kisti 연계] 한국정보과학회 정보과학회논문지 : 소프트웨어 및 응용 Vol.40 No.3 2013 pp.133-140
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 하나 이상의 클래스 레이블을 가지는 데이터에 대한 클래스 분류 기법들이 연구되고 있다. 그 중 몇몇 연구들에서 다중 레이블 분류의 성능을 높이기 위해 특징 선별 기법을 사용했다. 그러나 다중 레이블 데이터의 복잡성으로 인해 기존의 특징 선별 기법을 적용하는 것은 성능 향상에 한계가 있다. 본 논문은 다목적 최적화 알고리즘인 NSGA-II를 다중 레이블 특징 선별 문제에 활용했다. 또한 본 논문에서는 특징 선별 문제에 적합한 유전자 조작 기법을 제안하여 NSGA-II에 적용했다. 그리고 실제 다중 레이블 데이터 셋에서 실험을 통해 제안하는 알고리즘이 기존 유전자 알고리즘을 사용한 특징 선별 기법보다 더 좋은 성능을 가지는 것을 보인다.
Recently, a lot of researchers are interested in multi-label classification. Some of the researchers use feature subset selection to improve performance in multi-label classification. However, because multi-label problem is more complex than single-label, single-label feature selection algorithms have limitation to be applied to multi-label data. In this paper, NSGA-II, which is multi-objective optimize algorithm, is used for multi-label feature selection. In addition, this study proposes a novel genetic operator for feature selection problem. Finally, experiments on real world data set show that proposed algorithm achieves better performance than traditional genetic algorithm.
단백질의 세포내 위치 예측을 위한 다중레이블 분류 방법의 성능 비교
[Kisti 연계] 한국정보통신학회 한국정보통신학회논문지 Vol.18 No.4 2014 pp.992-999
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
단백질이 존재하는 세포내의 다중 위치를 정확하게 예측하기 위하여 다중레이블 학습 방법을 광범위하게 비교한다. 이를 위하여 다중레이블 분류의 접근 방법인 알고리즘 적응, 문제 변환, 메타 학습의 여러 방법을 비교 평가한다. 다양한 관점에서 다중레이블 분류 방법의 특성을 평가하기 위하여 12가지 평가 척도를 사용하였고, 최적의 성능을 보이는 방법을 찾기 위하여 새로운 요약 척도를 사용하였다. 비교 실험 결과, 흔하지 않은 다중레이블 집합을 가지치기 하는 멱집합 방법과, 관련 레이블들을 추가된 특징으로 나타내는 분류기-체인 방법의 성능이 높았다. 또한, 이들 방법들로 구성된 여러 개의 분류기를 조합하면 더욱 성능이 향상되었다. 즉, 세포내 위치간의 연관관계를 사용하는 것이 예측에 효과적인데, 특정 생물학적 기능을 수행하는 단백질의 세포내 위치들의 관계는 독립적이지 않고 서로 관련되어 있기 때문이라 판단된다.
This paper presents an extensive experimental comparison of a variety of multi-label learning methods for the accurate prediction of subcellular localization of proteins which simultaneously exist at multiple subcellular locations. We compared several methods from three categories of multi-label classification algorithms: algorithm adaptation, problem transformation, and meta learning. Experimental results are analyzed using 12 multi-label evaluation measures to assess the behavior of the methods from a variety of view-points. We also use a new summarization measure to find the best performing method. Experimental results show that the best performing methods are power-set method pruning a infrequently occurring subsets of labels and classifier chains modeling relevant labels with an additional feature. futhermore, ensembles of many classifiers of these methods enhance the performance further. The recommendation from this study is that the correlation of subcellular locations is an effective clue for classification, this is because the subcellular locations of proteins performing certain biological function are not independent but correlated.
다중 레이블 분류를 활용한 안면 피부 질환 인식에 관한 연구
[Kisti 연계] 한국정보처리학회 정보처리학회논문지/소프트웨어 및 데이터 공학 Vol.10 No.12 2021 pp.555-560
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
최근 안면 피부 미용에 대한 사람들의 관심이 높아짐에 따라 딥 러닝을 활용한 안면 피부 미용을 위한 피부 질환 인식 연구가 진행되고 있다. 이러한 연구들은 여드름을 비롯한 다양한 피부 질환을 인식한다. 기존의 연구들은 단일 피부 질환만을 인식하지만, 안면에 발생하는 피부 질환은 더 다양하고 복합적으로 발생할 수 있다. 따라서 본 논문에서는 Inception-ResNet V2 모델을 활용하여 다중 레이블 분류 방법으로 여드름, 블랙헤드, 주근깨, 검버섯, 일반 피부, 화이트헤드에 관한 복합적인 피부 질환을 인식한다. 사용한 평가 지표 중 정확도는 98.8%, 해밍 손실은 0.003을 달성하였고, 단일 클래스별 정밀도, 재현율, F1-점수는 모두 96.6% 이상을 달성하였다.
Recently, as people's interest in facial skin beauty has increased, research on skin disease recognition for facial skin beauty is being conducted by using deep learning. These studies recognized a variety of skin diseases, including acne. Existing studies can recognize only the single skin diseases, but skin diseases that occur on the face can enact in a more diverse and complex manner. Therefore, in this paper, complex skin diseases such as acne, blackheads, freckles, age spots, normal skin, and whiteheads are identified using the Inception-ResNet V2 deep learning mode with multi-label classification. The accuracy was 98.8%, hamming loss was 0.003, and precision, recall, F1-Score achieved 96.6% or more for each single class.
다중 레이블 분류 작업에서의 Coarse-to-Fine Curriculum Learning 메카니즘 적용 방안
[Kisti 연계] 한국컴퓨터정보학회 한국컴퓨터정보학회 학술대회논문집 2022 pp.29-30
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Curriculum learning은 딥러닝의 성능을 향상시키기 위해 사람의 학습 과정과 유사하게 일종의 'curriculum'을 도입해 모델을 학습시키는 방법이다. 대부분의 연구는 학습 데이터 중 개별 샘플의 난이도를 기반으로 점진적으로 모델을 학습시키는 방안에 중점을 두고 있다. 그러나, coarse-to-fine 메카니즘은 데이터의 난이도보다 학습에 사용되는 class의 유사도가 더욱 중요하다고 주장하며, 여러 난이도의 auxiliary task를 차례로 학습하는 방법을 제안했다. 그러나, 이 방법은 혼동행렬 기반으로 class의 유사성을 판단해 auxiliary task를 생성함으로 다중 레이블 분류에는 적용하기 어렵다는 한계점이 있다. 따라서, 본 논문에서는 multi-label 환경에서 multi-class와 binary task를 생성하는 방법을 제안해 coarse-to-fine 메카니즘 적용을 위한 방안을 제시하고, 그 결과를 분석한다.
다중 감정 분류를 위한 실시간 음성 기반 감정 인식 시스템 구현 및 평가
[Kisti 연계] 한국전자통신학회 The Journal of the Korean institute of electronic communication sciences Vol.20 No.3 2025 pp.611-620
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
본 논문은 KEMDy19 한국어 감정 음성 데이터셋과 AI-Hub에서 자체 구축한 커스텀 음성 데이터셋을 기반으로 MFCC 특징을 추출하고, CNN-BiLSTM 모델을 활용하여 실시간 다중 감정 인식 시스템을 구축하였다. 데이터 증강, 정규화, 감정별 오버샘플링을 최적화하여 모델의 일반화 성능과 감정 간 데이터 불균형 문제를 개선하였다. 감정별 최적 임계값(threshold)을 적용하여 예측 성능을 향상시킨 결과, 최종적으로 완전 일치 정확도 72%와 부분 감정 일치 정확도 93%를 달성하였으며, 본 시스템은 실시간 감정 인식 기반 스마트홈, AI 인터페이스, IoT 서비스 등 다양한 분야에 활용될 수 있다.
This study constructed a real-time multi-label emotion recognition system by extracting MFCC features from the KEMDy19 Korean emotional speech dataset and a custom speech dataset built using AI-Hub. By optimizing data augmentation, normalization, and emotion-specific oversampling, the model's generalization performance was improved and the issue of emotion class imbalance was addressed. Applying optimal threshold values for each emotion further enhanced prediction performance, resulting in an exact match accuracy of 72% and a partial match accuracy of 93%. The proposed system can be applied to various fields such as real-time emotion-aware smart homes, AI interfaces, and IoT services.
레이블 멱집합 분류와 다중클래스 확률추정을 사용한 단백질 세포내 위치 예측
[Kisti 연계] 한국정보통신학회 한국정보통신학회논문지 Vol.18 No.10 2014 pp.2562-2570
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
단백질의 기능을 유추할 수 있는 중요한 정보중의 하나는 단백질이 존재하는 세포내 위치이다. 최근에는 하나의 단백질이 동시에 존재하는 여러 세포내 위치를 예측하는 연구가 활발하다. 본 논문에서는 단백질이 존재하는 세포내의 다중위치를 예측하기 위해서 레이블 멱집합 방법을 개선한다. 레이블 멱집합 방법으로 분류한 다중위치들을 예측 확률에 따라 결합하여 최종적인 다중레이블로 분류한다. 각 다중위치에 대한 정확한 확률적 기여를 구하기 위하여 쌍별 비교와 오류정정 출력코드를 사용한 다중클래스 확률추정 방법을 적용하였다. 단백질 세포내 위치 예측 실험에 제안한 방법을 적용하여 성능이 향상됨을 보였다.
One of the important hints for inferring the function of unknown proteins is the knowledge about protein subcellular localization. Recently, there are considerable researches on the prediction of subcellular localization of proteins which simultaneously exist at multiple subcellular localization. In this paper, label power-set classification is improved for the accurate prediction of multiple subcellular localization. The predicted multi-labels from the label power-set classifier are combined with their prediction probability to give the final result. To find the accurate probability estimates of multi-classes, this paper employs pair-wise comparison and error-correcting output codes frameworks. Prediction experiments on protein subcellular localization show significant performance improvement.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.