Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 60
No
1

준지도학습 기반의 P2P 대출 부도 위험 예측에 대한 연구 KCI 등재

김현정

한국디지털정책학회 디지털융복합연구 제20권 제4호 2022.04 pp.185-192

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 연구는 P2P(Peer-to-Peer) 대출의 부도위험 예측을 위하여 준지도학습(SSL) 기반의 모델을 개발하고자 한 다. 검증된 성능에도 불구하고 지도학습(SL) 방법은 완전 지불 또는 채무불이행과 같이 레이블이 결정된 다수의 데이터 가 필요한데 충분한 수의 레이블 데이터를 수집하려면 많은 자원과 시간이 필요하다. P2P 플랫폼이 급성장하면서 대출 건수도 매해 급증하였고, 레이블이 없는 데이터도 지속적으로 증가하고 있다. 본 연구는 P2P 대출 플랫폼인 LendingClub에서 수집한 데이터를 사용하였다. P2P 대출 중 레이블이 결정된 대출에서 추출한 정보뿐만 아니라 레이 블이 결정되지 않은 대출에서 추출한 정보도 사용하여 부도 위험을 예측하는 SSL 모델을 개발하여 연구를 수행한 결과, 적은 수의 레이블이 결정된 데이터를 사용함에도 불구하고 SSL 방법으로 구축된 모델이 많은 수의 레이블이 결정된 데이터를 사용하여 학습시킨 SL 방법으로 구축된 모델보다 부도 위험 예측성과가 향상되었다.

This study investigates the effect of the semi-supervised learning(SSL) method on predicting default risk of peer-to-peer(P2P) loans. Despite its proven performance, the supervised learning(SL) method requires labeled data, which may require a lot of effort and resources to collect. With the rapid growth of P2P platforms, the number of loans issued annually that have no clear final resolution is continuously increasing leading to abundance in unlabeled data. The research data of P2P loans used in this study were collected on the LendingClub platform. This is why an SSL model is needed to predict the default risk by using not only information from labeled loans(fully paid or defaulted) but also information from unlabeled loans. The results showed that in terms of default risk prediction and despite the use of a small number of labeled data, the SSL method achieved a much better default risk prediction performance than the SL method trained using a much larger set of labeled data.

3

4,000원

This study examines factors influencing occupational injuries among plant and machine operators using the Semi-supervised MarginBoost algorithm. Data from the 2007-2009 Korean National Health and Nutrition Examination Survey (KNHANES) were analyzed, covering 4,062 employed participants. The MarginBoost model achieved 84.3% accuracy, outperforming other models. Key factors identified included exposure to hazardous substances, ergonomic conditions, and psychosocial stress. The findings emphasize the need for targeted interventions to enhance workplace safety and offer a robust predictive tool for the effective management of occupational health.

4

4,600원

오늘날 국제 항공산업은 1980년 이후 여객 및 화물 운송에서 지속적인 성장세를 보이고 있다. 그러나 항공규제완화 정책과 저가 항공사(LCC: low cost carrier)의 등장으로 성장세에도 불구하고 항공업계의 경쟁 환경은 더욱 심화되고 있는 실정이다. 이런 경쟁의 심화는 항공사들로 하여금 서비스 품질을 파악하고 향상시켜야 한다는 당위성을 제공하였다. 현재 항공사들은 사용자가 집적 작성한 콘텐츠를 이용하여 품질을 측정하고 그 결과에 따라 항공사 서비스 질을 향상시키려는 움직임이 전개하고 있다. 때문에 본 연구에서는 설문지 데이터 대신 항공사 리뷰 텍스트를 항공사 서비스 품질 모델(AIRQUAL: airline service quality) 차원으로 분류하여 각 차원 별 감성분석을 통해 감성 지표 도출한다. 또한 각 리뷰의 제목을 감성 분석하여 감성 지표를 도출한 후 두 가지 감성 지표를 변수로 활용하여 항공사 평가에 유의한지를 평가하였다. 연구결과에 따르면 AIRQUAL의 5가지 차원의 감성 지표가 항공사 서비스 품질을 나타내는 지표로서 유의하다는 것을 발견하였다. 또한, 리뷰의 제목을 감성 분석하여 작성된 감성 지표를 추가 지표로 사용하여 진행한 분석에서는 리뷰 텍스트의 감성 지표 보다 제목의 감성 지표를 추가한 것이 더 유의한 것으로 나타났다. 또한, 항공사의 서비스 품질 차원 중 ‘항공사 유형’이 중요한 지표로 나타났으며, 리뷰 제목의 감성 지표를 추가 변수로 사용한 모델에서는 ‘항공사 유형’ 보다 ‘리뷰 제목’의 감성 지표에 더 중요한 지표로 나타났다.

5

Anomaly recognition in visual and audio data has gained increasing significance in computer vision, as it plays a crucial role in protecting human lives and property. In this work, we developed a semi-supervised multimodal framework for anomaly recognition that combines audio and visual data for better performance. The proposed framework employs a hybrid network consisting of a convolutional neural network, Bi-Directional Long Short-Term Memory, a multi-head attention module, and a fully connected layer for anomalous pattern recognition. We created a novel real-time visual-audio anomaly recognition dataset and evaluated our framework on it, achieving promising results.

6

Although living organisms differ in shape and size, all are fundamentally structured by genetic sequences. Interpreting these sequences helps explain how organisms function. With the advancement of AI, significant breakthroughs have been made in protein sequencing and understanding protein function. However, there is still room for improvement, as data-intensive models require a substantial amount of protein sequences, many of which are not publicly available or lack quality. In this paper, we present a semi-supervised learning scheme to address the shortage of high-quality training data necessary for training protein language models. We demonstrate that this approach enhances the model's capability to classify toxic fungi protein sequences.

7

Recently, due to the recent significant advances in machine learning and deep learning, it is being utilized in many fields. However, real-world data in the medical field significantly degrades the performance of machine learning algorithms due to problems that are heavily skewed to specific states or that the distribution of data is unbalanced. Therefore, this study solves the problem of not being learned by converting the dependent variable into a regression problem that predicts using a new dependent variable by pseudo labeling. Also, this study present ensemble methods to improve the performance of the model and prevent overfitting.

9

교차 도메인 혼합 샘플링 기법을 활용한 준 지도 학습 기반 도메인 적응 기법 KCI 등재

김대한, 이상우, 최동걸, 서민석, 이재민, 장래영, 이건우, 최명석, 이용

한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 논문지 Vol.17 No.6 2021.12 pp.28-37

의미론적 영상 분할을 위한 컨볼루션 신경망 기반 접근 방식은 픽셀 단위 레이블을 통한 지도 학습에 의존한다. 하 지만 접근 불가능한 도메인으로 일반화되지 않을 수 있는 문제점이 존재한다. 의미론적 영상 분할에서 모든 데이터 에 사람이 직접 레이블을 지정하는 작업은 노동 집약적이기 때문에 소량의 라벨링 데이터를 이용해 네트워크의 일반 화 성능을 향상시키는 것은 매우 중요하다. 본 논문에서는 가상 도메인과 현실 도메인 사이의 교차 도메인 혼합 샘 플링을 활용한 도메인 적응 학습 방법을 제안한다. 이를 위하여 혼합 샘플링 방법의 대표적인 기법을 분석하고 활용 하여 도메인 적응 문제에 적용하고 일반화 성능을 비교한다. 제안하는 학습 방법을 통해 학습된 네트워크는 도메인 적응 문제에서 정량적 수치와 정성적 결과 모두 기준 모델을 능가하는 성능을 보인다.

In semantic segmentation, a convolutional neural network-based approach relies on supervised learning through pixel-level labels. However, there is a problem that may not be generalized to other inaccessible domains. Therefore, it is very important to improve the generalization performance of the network using a small amount of labeling data because it is labor intensive to manually label all data. In this paper, we propose a domain adaptive learning method using cross-domain mixed sampling between the synthetic domain and the real-world domain. To this end, we analyze and utilize the representative techniques of the mixed sampling method, apply them to the domain adaptation problem, and compare the generalization performance. The network trained through the proposed learning method outperforms the baseline model in both quantitative and qualitative results in domain adaptation problems.

10

수도 레이블을 활용한 준지도 학습 기반의 도로노면 파손 탐지 KCI 등재

전찬준, 류승기

한국ITS학회 한국ITS학회논문지 제18권 제4호 통권84호 2019.08 pp.71-79

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

의미론적 분할 형태로 합성곱 신경망을 구성하여 도로노면의 파손을 탐지하는 연구가 진행 되고 있다. 이러한 합성곱 신경망 형태의 모델을 생성하기 위해서는 입력 이미지와 이에 상응 한 레이블된 이미지 데이터셋으로 수집해야 하고, 이러한 과정에서는 굉장히 많은 시간과 비 용이 발생하게 된다. 본 논문에서는 이러한 작업을 완화하기 위하여 수도 레이블링을 활용한 준지도 학습 기반의 도로노면 파손 탐지 기술을 제안하고자 한다. 레이블된 데이터셋과 레이 블되지 않은 데이터셋을 적절하게 혼합하여 도로노면 파손을 탐지하는 모델을 업데이트하고, 이를 레이블된 데이터셋만을 활용한 기존 모델과 성능을 비교한다. 주관적인 성능결과, 민감도 부분에서는 조금 저하된 성능을 보였지만, 정밀도 부분에서는 대폭 성능 향상이 있었으며, 최 종적으로 F1-score 또한 높은 수치로 평가되었다

By using convolutional neural networks (CNNs) based on semantic segmentation, road surface damage detection has being studied. In order to generate the CNN model, it is essential to collect the input and the corresponding labeled images. Unfortunately, such collecting pairs of the dataset requires a great deal of time and costs. In this paper, we proposed a road surface damage detection technique based on semi-supervised learning using pseudo labels to mitigate such problem. The model is updated by properly mixing labeled and unlabeled datasets, and compares the performance against existing model using only labeled dataset. As a subjective result, it was confirmed that the recall was slightly degraded, but the precision was considerably improved. In addition, the F1-score was also evaluated as a high value.

11

Semi-supervised learning (SSL) has shown great promise in utilizing both labeled and unlabeled data, especially using consistency regularization To enhance the model's learning efficiency, we proposed a dualbranch co-training approach. After dividing the training data into two subsets, each branch utilizes a teacherstudent model pair, where the teacher's weights are updated via an exponential moving average (EMA). The framework combines supervised loss and unsupervised loss to optimize the model. Upon label expansion (e.g pseudo labeling), an additional otherside view is introduced, promoting agreement between the branches' predictions on shared data. This loss mitigates errors arising from incorrect pseudo-labels and enhances the overall robustness of the training process. By dynamically adjusting pseudo-label inclusion based on confidence thresholds, our methodology reduces the impact of noisy data and prevents overfitting. As a result, we could demonstrate the effectiveness of the proposed method in leveraging unlabeled data while maintaining high performance.

12

Semi-Supervised Learning Based Anomaly Detection for License Plate OCR in Real Time Video KCI 등재

Bada Kim, Junyoung Heo

국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 9 Number 1 2020.03 pp.113-120

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Recently, the license plate OCR system has been commercialized in a variety of fields and preferred utilizing low-cost embedded systems using only cameras. This system has a high recognition rate of about 98% or more for the environments such as parking lots where non-vehicle is restricted; however, the environments where non-vehicle objects are not restricted, the recognition rate is about 50% to 70%. This low performance is due to the changes in the environment by non-vehicle objects in real-time situations that occur anomaly data which is similar to the license plates. In this paper, we implement the appropriate anomaly detection based on semisupervised learning for the license plate OCR system in the real-time environment where the appearance of non-vehicle objects is not restricted. In the experiment, we compare systems which anomaly detection is not implemented in the preceding research with the proposed system in this paper. As a result, the systems which anomaly detection is not implemented had a recognition rate of 77%; however, the systems with the semisupervised learning based on anomaly detection had 88% of recognition rate. Using the techniques of anomaly detection based on the semi-supervised learning was effective in detecting anomaly data and it was helpful to improve the recognition rate of real-time situations.

13

Semi-supervised Learning for Automatic Image Annotation Based on Bayesian Framework SCOPUS

Dongping Tian

보안공학연구지원센터(IJCA) International Journal of Control and Automation Vol.7 No.6 2014.06 pp.213-222

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

In this paper, we present a new method for automatic image annotation by applying semi- supervised learning based on the Bayesian framework. On the one hand, we employ the semi- supervised learning, i.e., transductive support vector machine (TSVM) to enhance the quality of training image data, which is a promising way to find out the underlying relevant data from the unlabeled ones. On the other hand, a simple yet very efficient Bayesian model is built to implement image annotation by the maximum a posteriori (MAP) criterion. The novelty of our method mainly lies in two aspects: exploiting TSVM to improve the quality of training image dataset and utilizing the Bayesian model to predict the candidate annotations for the unseen images. Experimental results on the standard Corel dataset demonstrate that the proposed method is superior or highly competitive to several state-of-the-art approaches.

14

Data Ranking in Semi-Supervised Learning

Amin Allahyar, Hadi Sadoghi Yazdi

보안공학연구지원센터(IJAST) International Journal of Advanced Science and Technology Vol.53 2013.04 pp.1-10

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

The real challenge in pattern recognition tasks and machine learning processes is to train a discriminator using labeled data and use it to distinguish between future data points as accurate as possible. However, most of the problems in the real world have numerous data. Therefore assigning labels to every data points in these problems are a cumbersome or even impossible matter. Semi-supervised learning is one approach to overcome these types of problems. It uses only a small set of labeled with the company of huge remain and unlabeled data to train the discriminator. In semi-supervised learning, it is very essential that which data is labeled and depend on position of data it effectiveness changes. In this paper, we proposed an evolutionary approach called Artificial Immune System (AIS) to determine which data is better to be labeled to get the high quality data. The experimental results represent the effectiveness of this algorithm in finding these data points.

15

Short Text Classification Algorithm Based on Semi-Supervised Learning and SVM SCOPUS

Chunyong Yin, Jun Xiang, Hui Zhang, Zhichao Yin, Jin Wang

보안공학연구지원센터(IJMUE) International Journal of Multimedia and Ubiquitous Engineering Vol.10 No.12 2015.12 pp.195-206

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Short text is a popular text form, which is widely used in real-time network news, short commentary, micro-blog and many other fields. With the development of the application such as QQ, mobile phone text messages and movie websites, the size of data is also becoming larger and larger. Most data is useless for us while other data is significant for us. Therefore, it is necessary for us to extract the useful short text from the big data. However, there are many problems with the short text classification, such as fewer features, irregularity and so on. To solve these problems, we should pretreat the short text set first, and then choose the significant features. This paper use semi-supervised learning method and SVM classifier to improve the traditional methods and it can classify a large number of short texts to mining the useful massage from the short text. The experimental results in this paper also show a good promotion.

16

세미감독형 학습 기법을 사용한 소프트웨어 결함 예측 KCI 등재

홍의석

국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제19권 제3호 2019.06 pp.127-133

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

소프트웨어 결함 예측 연구들의 대부분은 라벨 데이터를 훈련 데이터로 사용하는 감독형 모델에 관한 연구들이 다. 감독형 모델은 높은 예측 성능을 지니지만 대부분 개발 집단들은 충분한 라벨 데이터를 보유하고 있지 않다. 언라벨 데이터만 훈련에 사용하는 비감독형 모델은 모델 구축이 어렵고 성능이 떨어진다. 훈련 데이터로 라벨 데이터와 언라벨 데이터를 모두 사용하는 세미 감독형 모델은 이들의 문제점을 해결한다. Self-training은 세미 감독형 기법들 중 여러 가정과 제약조건들이 가장 적은 기법이다. 본 논문은 Self-training 알고리즘들을 이용해 여러 모델들을 구현하였으며, Accuracy와 AUC를 이용하여 그들을 평가한 결과 YATSI 모델이 가장 좋은 성능을 보였다.

Most studies of software fault prediction have been about supervised learning models that use only labeled training data. Although supervised learning usually shows high prediction performance, most development groups do not have sufficient labeled data. Unsupervised learning models that use only unlabeled data for training are difficult to build and show poor performance. Semi-supervised learning models that use both labeled data and unlabeled data can solve these problems. Self-training technique requires the fewest assumptions and constraints among semi-supervised techniques. In this paper, we implemented several models using self-training algorithms and evaluated them using Accuracy and AUC. As a result, YATSI showed the best performance.

17

점진적 능동준지도 학습 기반 고효율 적응적 얼굴 표정 인식 KCI 등재

김진우, 이필규

국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제17권 제2호 2017.04 pp.165-171

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

사람의 얼굴 표정을 실제 환경에서 인식하는 데에는 여러 가지 난이한 점이 존재한다. 그래서 학습에 사용된 데이터베이스와 실험 데이터가 여러 가지 조건이 비슷할 때에만 그 성능이 높게 나온다. 이러한 문제점을 해결하려면 수많은 얼굴 표정 데이터가 필요하다. 본 논문에서는 능동준지도 학습을 통해 다양한 조건의 얼굴 표정 데이터를 쉽게 모으고 보다 빠르게 성능을 확보할 수 있는 방법을 제안한다. 제안하는 알고리즘은 딥러닝 네트워크와 능동 학습 (Active Learning)을 통해 초기 모델을 학습하고, 이후로는 준지도 학습(Semi-Supervised Learning)을 통해 라벨이 없는 추가 데이터를 확보하며, 성능이 확보될 때까지 이러한 과정을 반복한다. 위와 같은 능동준지도 학습(Active Semi-Supervised Learning)을 통해서 보다 적은 노동력으로 다양한 환경에 적합한 데이터를 확보하여 성능을 확보할 수 있다.

It is difficult to recognize Human’s facial expression in the real-world. For these reason, when database and test data have similar condition, we can accomplish high accuracy. Solving these problem, we need to many facial expression data. In this paper, we propose the algorithm for gathering many facial expression data within various environment and gaining high accuracy quickly. This algorithm is training initial model with the ASSL (Active Semi-Supervised Learning) using deep learning network, thereafter gathering unlabeled facial expression data and repeating this process. Through using the ASSL, we gain proper data and high accuracy with less labor force.

18

Semi-supervised learning 기법을 활용한 병리학 이미지 분석

이유진, 박지영, 이상민

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2023 pp.675-677

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 병리학 이미지 분석에서 자주 발생하는 문제 중 하나인 레이블링 불일치 문제를 해결하고자 준지도학습(semi-supervised learning) 기법을 적용하였다. 기존의 병리 진단 과정은 정확한 판정 및 치료를 위해 전문가의 판단을 필요로 한다. 이로 인해, 시간이 매우 많이 소모되며 전문가의 피로도가 증가한다. 최근 이를 해결하고자 지도학습(supervised learning) 기법을 사용하여 업무의 피로도를 감소시키고자 하는 연구가 진행되고 있다. 하지만 병리 이미지 데이터에 대한 접근이 어렵고, 병변의 위치를 레이블링 하는 부분에서 많은 비용이 발생한다. 또한 암 병변의 스펙트럼적 특성으로 인해 레이블링 과정 속에서 레이블링 불일치 문제가 발생할 가능성이 높다. 이러한 문제를 극복하기 위해, 우리는 제한된 레이블 된 데이터와 많은 양의 레이블 되지 않은 데이터를 활용하는 준지도학습 방법론을 제안한다. 이 제안하는 방법은 필요한 수동 레이블링 작업량을 줄여, 병리학자들에게 보다 효과적인 진단 도구를 제공할 것으로 예상된다.

19

GAN기반의 Semi Supervised Learning을 활용한 이미지 생성 및 분류

정도윤, 최광미, 김남호

[Kisti 연계] 한국스마트미디어학회 스마트미디어저널 Vol.13 No.3 2024 pp.27-35

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 GAN(Generative Adversarial Network)을 기반으로 한 Semi Supervised Learning을 활용하여 이미지 생성과 ResNet50을 이용한 이미지 분류를 결합하는 방법에 대해 다루고 있다. 이를 통해 새로운 접근법을 제시하여 이미지 생성과 분류를 통합함으로써 더 정확하고 다양한 결과를 얻을 수 있도록 하였다. 생성자와 판별자를 학습시켜 생성된 이미지와 실제 이미지를 구별하고, ResNet50을 활용하여 이미지 분류를 수행한다. 실험 결과에서는 생성된 이미지의 품질이 epoch에 따라 변화함을 확인할 수 있었으며, 이를 통해 산업재해 예측 정확성을 향상하고자 한다. 또한, GAN과 ResNet50의 결합을 통해 이미지 생성의 품질을 향상시키고 이미지 분류의 정확도를 높이는 효율적인 방법을 제시하고자 한다.

This study deals with a method of combining image generation using Semi Supervised Learning based on GAN (Generative Adversarial Network) and image classification using ResNet50. Through this, a new approach was proposed to obtain more accurate and diverse results by integrating image generation and classification. The generator and discriminator are trained to distinguish generated images from actual images, and image classification is performed using ResNet50. In the experimental results, it was confirmed that the quality of the generated images changes depending on the epoch, and through this, we aim to improve the accuracy of industrial accident prediction. In addition, we would like to present an efficient method to improve the quality of image generation and increase the accuracy of image classification through the combination of GAN and ResNet50.

20

게임 환경을 통제할 수 있는 규칙 기반 Semi-Supervised Learning 오목 인공지능 프레임 워크

김선민, 구본우

[Kisti 연계] 한국정보처리학회 한국정보처리학회 학술대회논문집 2022 pp.618-620

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

게임은 수많은 NPC 와 규칙에 의해 작동되는 가상 공간을 의미한다. 이런 가상 공간에서는 규칙을 엄격히 지키면서 수행되는 AI 를 필수로 요구하게 된다. 하지만 강화 학습 기반의 AI 는 복잡한 게임의 규칙을 온전히 지키지 못하고 예상 밖의 행동을 돌출하면서 이를 해결하기 위한 많은 연구도 수행되고 있다. 본 논문에서는 규칙 기반으로 획득한 오목판의 확률 맵과 학습을 통해 획득한 확률맵 데이터를 병합하여 가장 높은 Value 를 가지는 위치를 다음 수로 반환하는 방법을 사용하였다. 향후 연구에서는 ANN(Approximate Nearest Neighbor)알고리즘을 적극 활용하여, 커널의 State 와 보드의 State 비교를 확률적으로 개선할 예정이다. 본 논문에서 제안된 프레임 워크는 게임 AI 연구에 기여할 수 있길 바란다.

 
1 2 3
페이지 저장