년 - 년
인공지능 학습데이터 라벨링 정확도에 따른 인공지능 성능 KCI 등재
대한산업경영학회 산업융합연구(구 대한산업경영학회지) 제22권 제1호 2024.01 pp.177-183
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
본 연구는 데이터의 품질이 인공지능(AI) 성능에 미치는 영향을 검토한다. 이를 위해, 데이터 특성변수(Feature)의 유사도와 클래스(Class) 구성의 불균형을 고려한 모의실험(Simulation)을 통해 라벨링 오류 수준이 인공지능의 성능에 미치 는 영향을 비교 분석하였다. 그 결과, 특성변수 간 유사성이 높은 데이터에서는 특성 변수 간 유사성이 낮은 데이터에 비해 라 벨링 정확도에 더 민감하게 반응하였으며, 클래스 불균형이 증가함에 따라 인공지능 정확도가 급격히 감소되는 경향을 관찰 하였다. 이는 인공지능 학습데이터의 품질평가 기준 및 관련 연구를 위한 기초자료가 될 것이다.
The study investigates the impact of data quality on the performance of artificial intelligence (AI). To this end, the impact of labeling error levels on the performance of artificial intelligence was compared and analyzed through simulation, taking into account the similarity of data features and the imbalance of class composition. As a result, data with high similarity between characteristic variables were found to be more sensitive to labeling accuracy than data with low similarity between characteristic variables. It was observed that artificial intelligence accuracy tended to decrease rapidly as class imbalance increased. This will serve as the fundamental data for evaluating the quality criteria and conducting related research on artificial intelligence learning data.
Cross-domain autonomous driving visual segmentation based on enhanced target data learning
[NRF 연계] 한국통신학회 ICT Express Vol.11 No.1 2025.02 pp.53-58
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Within the broader context of Information and Communications Technology (ICT), the quest for reliable and scalable visual segmentation methods poses significant challenges, particularly in autonomous driving, where real-world scene complexity requires advanced solutions. To address data scarcity and improve segmentation performance, we propose a novel unsupervised domain adaptation (UDA) approach that enhances target domain learning. Our method introduces multiple perturbations consistency, leveraging spatial context within the target domain to improve recognition. By applying perturbations at input and feature levels and using a consistency loss, we enhance contextual learning. Additionally, a weight mapping technique reduces the impact of detrimental source domain information. Experimental results demonstrate that our approach outperforms baseline methods on the GTAV->Cityscapes and SYNTHIA->Cityscapes datasets.
Data Preprocessing for Machine-Learning-Based Adaptive Data Center Transmission
[NRF 연계] 한국통신학회 ICT Express Vol.8 No.1 2022.03 pp.37-43
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
To enable optical interconnect fluidity in next-generation data centers, we propose adaptive transmission based on machine learning in a wavelength-routing network. We consider programmable transmitters that can apply possible code rates to connections based on predicted bit error rate (BER) values. To classify the BER, we employ a preprocessing algorithm to feed the traffic data to a neural network classifier. We demonstrate the significance of our proposed preprocessing algorithm and the classifier performance for different values of and switch port count.
[NRF 연계] 한국통신학회 ICT Express Vol.6 No.3 2020.09 pp.220-228
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Environmental monitoring has become an active research area due to the current rise in the global climate change crises. Current environmental monitoring solutions, however, are characterized by high cost of acquisition and complexity of installation; often requiring extensive resources, infrastructure and expertise. It is infeasible to achieve with these solutions, high density in-situ networks such as are required to build refined scale models to facilitate robust monitoring, thus, leaving large gaps within the collected dataset. Low-Cost Sensors (LCS) can offer high-resolution spatiotemporal measurements which could be used to supplement existing dataset from current environmental monitoring solutions. LCS however, require frequent calibration in order to provide accurate and reliable data as they are often affected by environmental conditions when deployed on the field. Calibrating LCS can help to improve their data quality and ensure they are collecting accurate data. Achieving effective calibration, however, requires identifying factors that affect sensor’s data quality for a given measurement. This study evaluates the performance of three Feature Selection (FS) algorithms including Forward Feature Selection (FFS), Backward Elimination (BE) and Exhaustive Feature Selection (EFS) in identifying factors that affect data quality of low-cost IoT sensors in environmental monitoring networks. Applying the concept of data fusion, sensors data were merged with environmental factors and integrated into a single calibration equation to calibrate cairclipO3/NO2 and cairclipNO2 sensors using Linear Regression (LR) and Artificial Neural Networks (ANN). The study showed the effectiveness of calibration in improving low-cost IoT sensor data quality and also demonstrated the convenience of feature selection and the ability of data fusion to provide more consistent, accurate and reliable information for calibration models. The analysis showed that the cairclipO3/NO2 sensor provided measurements that have good correlation with reference measurements whereas the cairclipNO2 sensor showed no reasonable correlation with the reference data. Calibrating the cairclipO3/NO2 yielded good improvement in its measurement outputs when compared to reference measurements (R2=0.83). However, calibrating the cairclipNO2 sensor data yielded no significant improvement in its data quality.
Toward storage-aware learning with compressed data an empirical exploratory study on JPEG
[NRF 연계] 한국통신학회 ICT Express Vol.12 No.1 2026.02 pp.50-54
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
On-device machine learning is fundamentally constrained by limited storage, especially in continuous data collection scenarios where sensor or vision streams accumulate rapidly. This paper empirically investigates storage-aware learning, characterizing the trade-off between data quantity and data quality under lossy compression. Using the CIFAR-10 dataset as a controlled benchmark, we systematically vary both the amount and the fidelity of training data to understand their joint impact on model performance. Our results reveal that (1) neither maximizing quantity nor quality alone yields optimal accuracy, emphasizing that the optimal trade-off between them depends nonlinearly on the available storage budget, and (2) data samples exhibit differential sensitivity to compression, motivating a sample-wise adaptive compression policy. These findings challenge uniform data-retention strategies such as naive data dropping or fixed-rate compression, and establish a foundation for adaptive, storage-efficient learning systems on resource-limited devices. This work opens new directions toward generalizable, storage-aware on-device intelligence.
Selective imitation for efficient online reinforcement learning with pre-collected data
[NRF 연계] 한국통신학회 ICT Express Vol.10 No.6 2024.12 pp.1308-1314
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Deep reinforcement learning (RL) has emerged as a promising solution for autonomous devices requiring sequential decision-making. In the online RL framework, the agent must interact with the environment to collect data, making sample efficiency the most challenging aspect. While the off-policy method in online RL partially addresses this issue by employing a replay buffer, learning speed remains slow, particularly at the beginning of training, due to the low quality of data collected with the initial policy. To overcome this challenge, we propose Reward-Adaptive Pre-collected Data RL (RAPD-RL), which leverages pre-collected data in addition to online RL. We employ two buffers: one for pre-collected data and another for online collected data. The policy is trained using both buffers to increase the objective and imitate the actions in the dataset. To maintain resistance to poor-quality (i.e., low-reward) data, our method selectively imitates data based on reward information, thereby enhancing sample efficiency and learning speed. Simulation results demonstrate that the proposed solution converges rapidly and achieves high performance across various dataset qualities.
[NRF 연계] 한국통신학회 ICT Express Vol.10 No.4 2024.08 pp.863-870
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
To address the accuracy degradation as well as prolonged convergence time due to the inherent data heterogeneity among end-devices in federated learning (FL), we introduce the joint batch size and weighted aggregation adjustment problem, which is non-convex problem. To adjust optimal hyperparameters, we develop deep reinforcement learning (DRL) to empower a mechanism known as Batch size and Weighted aggregation Adjustment (BWA). Experimental evaluation demonstrates that BWA not only outperforms methods optimized solely from either a local training or server perspective but also achieves higher accuracy, with an increase of up to 5.53% compared to FedAvg, and additionally accelerates convergence speeds.
[NRF 연계] 한국기초간호학회 Journal of korean biological nursing science Vol.27 No.4 2025.11 pp.586-597
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
PurposeThis study aimed to classify self-care patterns among Korean adults with prediabetes using an unsupervised machine learning approach. The classification was grounded in Orem’s Self-Care Theory, focusing on self-care demands, self-care agencies, and self-care behaviors. MethodsA secondary data analysis was conducted using the 2023 Korea National Health and Nutrition Examination Survey. Variables were selected and categorized according to the theoretical components of Orem’s model. Principal component analysis was applied for dimensionality reduction, followed by K-means clustering to identify distinct self-care pattern groups. All variables were standardized using min-max normalization. Group differences were examined using analysis of variance and the chi-square test. ResultsThree self-care pattern groups were identified: the high self-care performance group, the latent self-care risk group, and the self-care vulnerable group. These groups exhibited distinct profiles across self-care demands, agencies, and behaviors. Significant intergroup differences were also observed in education level, income, health literacy, fasting blood glucose, and hemoglobin A1c levels. ConclusionSelf-care patterns among adults with prediabetes can be effectively classified through unsupervised learning techniques. The findings highlight the importance of developing tailored nursing interventions that consider multidimensional self-care profiles. This study underscores the applicability of Orem’s Self-Care Theory and demonstrates the potential of machine learning in identifying at-risk subgroups for early intervention.
A Deep Learning based HTTP Slow DoS Classification Approach Using Flow Data
[NRF 연계] 한국통신학회 ICT Express Vol.7 No.2 2021.06 pp.210-214
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
The popularity of the Internet introduces many network-enabled services that can be accessed by the user. But the adversaries are trying to deny these critical services to the user through Denial of Service (DoS) attacks. Presently, dealing with DoS attack which targets the application layer using slow traffic rate is one of the key challenges faced by the service providers. In this paper, a deep classification model using flow data is proposed to detect slow DoS attack on HTTP. The classifier is evaluated using CICIDS2017 dataset. The results obtained show that the classifier can obtain 99.61% accuracy.
[NRF 연계] KEMA학회 Journal of Musculoskeletal Science and Technology Vol.7 No.2 2023.12 pp.71-79
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Background Pressure pain hypersensitivity (PPH) is used to measure pain sensitivity in deep tissues, but factors contributing to PPH remain unclear. Abnormal neck and scapula posture are thought to play a role in shoulder pain. Traditional statistical methods like logistic regression have limitations in capturing complex relationships, while machine learning (ML) can model nonlinear relationships effectively. Purpose The purpose of the present study was to develop, evaluate, and compare the predictive performance of ML models and logistic regression for classifying food service workers (FWs) with and without PPH based on postural analysis data. Study design Cross sectional study. Methods FWs (n=150) meeting specific criteria were assessed for PPH and underwent postural analysis. ML algorithms (logistic regression, neural network, random forest, gradient boosting, decision tree, and support vector machine) were used for classification. Model performance was evaluated using the area under the curve (AUC), accuracy, recall, precision, and F1 score. Feature importance was assessed. Results Gradient boosting exhibited the best performance (AUC: 0.867) in classifying PPH, followed by random forest (AUC: 0.822) in the test dataset. Logistic regression performed less effectively (AUC: 0.613). For feature importance analysis, scapular downward rotation ratio, forward head posture, BMI and rounded shoulder angle were the top four important predictors of PPH in gradient boosting model. Conclusions Gradient boosting, along with identified predictors, offers promise for early intervention and risk assessment tools in addressing musculoskeletal pain in food service workers.
[NRF 연계] 한국기초간호학회 Journal of korean biological nursing science Vol.26 No.4 2024.11 pp.300-310
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Purpose: This study aimed to develop and compare machine learning models for predicting prediabetes in young adults in Korea using dietary intake data and to identify the most effective model. Methods: Data from the ninth Korea National Health and Nutrition Examination Survey were used, with 823 participants aged 19-35 years selected after excluding those with missing data. Logistic regression, k-nearest neighbors, and random forest models were applied to predict prediabetes, and the analysis was conducted using the Orange 3.5 program. Five-fold cross-validation was performed to reduce performance variability, and test data were used for final model validation. Results: In the dataset, 14%-15% of participants were classified as having prediabetes. The random forest model showed the highest performance in terms of classification accuracy, harmonic mean of precision and recall, and precision. Logistic regression had the highest performance regarding the model’s ability to distinguish between individuals with and without prediabetes. Age, thiamine intake, and water intake emerged as the most important predictors. Conclusion: This study demonstrated the utility of using dietary intake data to predict prediabetes in young adults. The random forest model provided the highest prediction accuracy, supporting early detection and intervention, which could help to reduce unnecessary treatment. This highlights nurses’ important role in educating patients about lifestyle changes and implementing preventive care. Future studies should incorporate additional factors, such as psychological and lifestyle variables, to improve the model's performance.
위기관리 이론과 실천 한국위기관리논집 제19권 제12호 2023.12 pp.13-28
※ 기관로그인 시 무료 이용이 가능합니다.
4,900원
최근 발생하고 있는 기후변화는 전세계적으로 많은 피해를 발생시키고 있다. 한반도의 기존 강우형태는 장시간에 걸친 강우와 발생기간을 예측할 수 있었으나, 기후 변화로 인한 강우 형태의 변화는 기존에 마련한 위기대응방안이 무용한 상황을 만들고 있다. 2022년 8월 집중호우로 인해 서울·경기·강원·충남 등 10개 지자체가 특별재난지역으로 우선 선포되었다. 도시화에 따른 불투수면적의 증가와 건물의 지하 공간의 활용 증가하고 있는 가운데 기후 변화로 인한 집중호우는 도심지 내 하수관로 설계용량을 초과하 는 경우가 빈번해지고 있으며, 범위는 국부적이지만 피해의 규모는 커지는 사례들이 증가하고 있다. 도시침수는 다양한 조건에 의해 침수범위가 결정되며 불시에 일어나는 침수의 형태로 인해 신속한 침수상황 인지 및 위기대응이 필요하다. 이를 위하여 AI기반 영상분석 기술을 통해 전국에 분포되어 있는 CCTV를 활용하여 침수상황을 상시 모니터링 할 수 있는 시스템 구축이 필요하며, 정형화된 IoT기반 실시간 계측센서 데이터와 비정형 데이터인 CCTV영상을 분석하고 연계한 새로운 형태의 도시침수 모니터링 기술이 필요하다.
Recently, climate change has caused a lot of damage worldwide. Existing rainfall patterns on the Korean Peninsula could predict long-term rainfall patterns and periods, but changes in rainfall patterns due to climate change made crisis response measures useless. Due to heavy rain in August 2022, 10 local governments, including Seoul, Gyeonggi, Gangwon, and Chungnam, were declared as special disaster areas for the first time. As the impermeable area increases due to urbanization and the use of the underground space of buildings increases, the design capacity of sewage pipes is often exceeded, and the scope is local but the damage scale is increasing. Urban floods require rapid situational awareness and crisis response. To this end, it is necessary to establish a system that can monitor the flood situation at all times using CCTV distributed across the country through AI-based image analysis technology, and a new type of urban flood monitoring technology that analyzes and links it based on standardized IoT unstructured data such as real-time measurement sensor data and CCTV images is required.
An Efficient Learning Methodology Using Curriculum Learning and Data Reduction
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 10th International Conference on Next Generation Computing 2024 2024.11 pp.285-286
The advancement of deep learning has significantly improved image classification performance. However, the complexity of large-scale datasets continues to present challenges in terms of training time and resource consumption. One approach to address these issues is Self- Paced Curriculum Learning. Self-Paced Curriculum Learning begins by training on relatively easy data, allowing the model to autonomously select data and adjust difficulty levels as training progresses, gradually incorporating more complex data. This method improves the efficiency of the training process while minimizing performance degradation. In this study, we propose an approach that combines Self- Paced Curriculum Learning with the exclusion of low-noise data from the training process to further enhance training speed. The experimental results show that reducing the amount of data improves training speed. However, accuracy tends to decrease as the extent of data reduction increases.
Data-Driven Learning (DDL) Approach: Is It Applicable to Korean EFL Classrooms? KCI 등재
한국외국어교육학회 외국어교육 제20권 제3호 2013.09 pp.45-86
※ 기관로그인 시 무료 이용이 가능합니다.
8,800원
한국경영정보학회 Asia Pacific Journal of Information Systems 제34권 제2호 2024.06 pp.400-420
※ 기관로그인 시 무료 이용이 가능합니다.
5,700원
Nowadays, social media has evolved into a powerful networked ecosystem in which governments and citizens publicly debate economic and political issues. This holds true for the pros and cons of Indonesia’s ore nickel export restriction to Europe, which we aim to investigate further in this paper. Using Twitter as a dependable channel for conducting sentiment analysis, we have gathered 7070 tweets data for further processing using two sentiment analysis approaches, namely Support Vector Machine (SVM) and Long Short Term Memory (LSTM). Model construction stage has shown that Bidirectional LSTM performed better than LSTM and SVM kernels, with accuracy of 91%. The LSTM comes second and The SVM Radial Basis Function comes third in terms of best model, with 88% and 83% accuracies, respectively. In terms of sentiments, most Indonesians believe that the nickel ore provision will have a positive impact on the mining industry in Indonesia. However, a small number of Indonesian citizens contradict this policy due to fears of a trade dispute that could potentially harm Indonesia’s bilateral relations with the EU. Hence, this study contributes to the advancement of measuring public opinions through big data tools by identifying Bidirectional LSTM as the optimal model for the dataset.
Data Sharing for Learning Analytics – Exploring Risks and Benefits through Questioning
에듀테크학회(구 이러닝학회) 이러닝학회 논문지 제1권 제1호 2016.12 pp.1-13
※ 기관로그인 시 무료 이용이 가능합니다.
4,500원
Moving learning analytics from the research labs to the classrooms, lecture halls, and digital learning spaces requires data sharing. Routine production and consumption of data is no limited to controlled settings as data can be sourced from an increasing diversity of data points and combined or aggregated from these sources within and often beyond the institution. This scaling up of learning analytics raises a host of questions on behalf of the data subjects and therefore informs requirements for design of new solutions and practices. This paper analyses a corpus of more than 250 questions gathered by a European project that organized international workshops and facilitated community exchange. It explores how these questions could inform both the problem space of data sharing and the solution space yet to be fully scoped by research and development within this emerging field. The analysis shows that the discourse on data sharing and big data for education is still at an early stage. Conceptual issues dominate this discourse; however, the elicited questions also hold numerous challenges for technical development and implementation. In concluding we propose a short inventory of what the emerging solution space may entail and suggest a path for further work.
한국정보기술응용학회 JITAM Vol.1 1999.03 pp.173-211
※ 기관로그인 시 무료 이용이 가능합니다.
8,400원
The task of classification permeates all walks of life, from business and economics to science and public policy. In this context, nonlinear techniques from artificial intelligence have often proven to be more effective than the methods of classical statistics. The objective of knowledge discovery and data mining is to support decision making through the effective use of information. The automated approach to knowledge discovery is especially useful when dealing with large data sets or complex relationships. For many applications, automated software may find subtle patterns which escape the notice of manual analysis, or whose complexity exceeds the cognitive capabilities of humans. This paper explores the utility of a collaborative learning approach involving integrated models in the preprocessing and postprocessing stages. For instance, a genetic algorithm effects feature-weight optimization in a preprocessing module. Moreover, an inductive tree, artificial neural network (ANN), and k-nearest neighbor (kNN) techniques serve as postprocessing modules. More specifically, the postprocessors act as second0order classifiers which determine the best first-order classifier on a case-by-case basis. In addition to the second-order models, a voting scheme is investigated as a simple, but efficient, postprocessing model. The first-order models consist of statistical and machine learning models such as logistic regression (logit), multivariate discriminant analysis (MDA), ANN, and kNN. The genetic algorithm, inductive decision tree, and voting scheme act as kernel modules for collaborative learning. These ideas are explored against the background of a practical application relating to financial fraud management which exemplifies a binary classification problem.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.