년 - 년
Learning Agent를 고려한 연합학습 시뮬레이터 고안 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제6권 12호 2022.12 pp.2225-2239
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
실제 환경에서 연합학습을 수행하기 위해 많은 기기가 준비되어야 하고 다양한 상황들에 대한 고려가 필 요하다. 이는 높은 비용 소모를 야기하므로 실제 환경에서 연합학습을 구축하기 전에 시뮬레이션을 통해 성능 측 정을 해보는 과정이 필요하다. 또한 기존의 연합학습 방법은 개인이 소유하고 있는 다양한 기기들을 고려하지 않 는다는 한계를 가지고 있다. 이와 같은 한계를 극복하기 위해 본 논문에서는 연합학습에 사용되는 다양한 인자들을 파라미터화한 시뮬레이터를 구현하였고 개인이 소유한 다양한 기기가 활용된 학습이 고려될 수 있도록 Learning Agent 개념을 추가하였다. 제안한 시뮬레이터는 Windows 11, Python3.9의 환경에서 구현되었으며, 본 논문에서 가정한 상황에서 시뮬레이션 한 결과 Learning Agent를 사용한 경우, 평균적으로 전체 학습 시간은 1/2 가량 줄 어들었으며 상한 정확도에 수렴하기까지의 소요 시간은 1/3 가량 줄어듦을 확인하였다. 따라서 제안된 시뮬레이터 는 사용자의 다양한 기기를 활용한 연구에 활용될 수 있을 것이며 실제 환경에 연합학습을 구축하려는 상황에서 성능 측정을 위해 활용할 수 있는 도구가 될 것으로 기대한다.
To perform federated learning (FL) in real environments, many devices are required and various situations should be considered. This leads to high-cost consumption, so we need a process to measure performance through simulation before conducting FL in a real environment. In addition, the existing FL has a limitation in that it does not consider users’ devices which can be utilized for training. To overcome such limitation, in this paper, we implemented the FL simulator that parameterizes various factors in FL to consider diverse devices and situations. Also, the concept of Learning Agent was considered so that the simulator includes training by using users’ idle devices. The simulator was implemented on Windows 11 using Python 3.9. Using the simulator, we performed the various simulations in the assumed situations. As a result, on average, the total training time was halved and the time required to reach the saturation point of accuracy was reduced by one-third. These results showed that the proposed simulator can be used for researches using various users’ idle devices and measure performance before constructing FL systems in real environments.
Learning Agent 활용 및 다중 경로 기반 모델 분할 전송을 통한 연합학습 개선 KCI 등재
국제차세대융합기술학회 차세대융합기술학회논문지 제8권 3호 2024.03 pp.607-620
※ 기관로그인 시 무료 이용이 가능합니다.
4,600원
연합학습은 데이터를 중앙 서버에 모아서 학습하는 중앙집중식 학습과 달리 개별 기기에서 모델 학습을 진행한 후 중앙 서버로 가중치들을 전송하는 분산식 학습 방식을 가진다. 따라서 개별 기기의 연산 성능과 통신 성 능 및 상황이 모델 학습에 큰 영향을 끼칠 수 있으므로 적절한 기기와 통신 방식을 선택해야 한다. 하지만 기존 연 합학습은 클라이언트 주위의 신뢰할 수 있는 가용 유휴 장치와 다양한 통신 및 다중 경로를 사용하지 않는 한계가 존재한다. 이를 극복하기 위해, 본 논문에서는 Learning agent 활용 및 다중 경로 기반 모델 분할 전송을 통한 연 합학습 개선 연구를 수행하였다. 시뮬레이션과 실제 장치를 활용한 실험을 통해 제안한 기법을 평가하였으며, 기존 연합학습과 비교하여 제안한 기법을 활용한 경우 학습 시간의 감소 및 학습 성능의 향상을 확인할 수 있었다.
Unlike centralized learning, federated learning has a distributed learning method that individual devices perform model training locally and then transmit weights to the central server. Therefore, the appropriate devices and communication methods should be selected as the computational and communication performance, and situation of individual devices can have a great impact on learning. However, the existing federated learning has limitations in not using various communications through multiple paths simultaneously and reliable and available idle devices around clients. To overcome this, in this paper, we conduct a study on improvement of federated learning using learning agent and multipath-based model split transmission. We evaluated the proposed technique through simulations and experiments using the real devices. The evaluation results show that the training time was reduced and the learning performance improved when the proposed technique was used compared to the existing method.
Learning graph based individual intrinsic reward for multi-agent reinforcement learning
[NRF 연계] 한국통신학회 ICT Express Vol.12 No.2 2026.04 pp.301-305
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Designing a reward function is a critical challenge in reinforcement learning. However, as environments become more complex and tasks grow more difficult, designing a reward function that drives optimal behavior becomes increasingly challenging. To overcome these issues, Preference based reinforcement learning has proposed methods that learn reward functions based on the preference between two trajectories, thereby eliminating the need for handcrafted reward function. In multi-agent reinforcement learning, the challenge is even greater due to the complex interactions among agents, which makes designing a single global reward function even more difficult. In this paper, we show that when a single global reward function is learned via preference-based reinforcement learning in multi-agent setting, it often fails to capture sufficient information for optimal policy learning. Instead, we propose a method for learning individual reward functions that provide additional guidance for each agent’s optimal policy. Our approach, which leverages graph structures and preference-based reinforcement learning, outperforms the method based on learning a single, global reward function.
Multi-agent reinforcement learning for a distributed multi-channel access game
[NRF 연계] 한국통신학회 ICT Express Vol.11 No.5 2025.10 pp.863-869
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this work, we model multi-user distributed channel access as a game with channels and users, and propose the Multi-Agent Thompson Sampling (MA-TS) algorithm. It uses Bayes’ theorem to dynamically optimize action selection. This optimization aims to maximize throughput. We derive the algorithm’s computational complexity as . Simulations show that MA-TS converges to a pure strategy Nash equilibrium (PNE) and outperforms existing methods in average throughput.
[NRF 연계] 한국통신학회 ICT Express Vol.10 No.3 2024.06 pp.472-478
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In CRNs, it is crucial to develop an efficient and reliable spectrum detector that consistently provides accurate information about the channel state. In this work, we investigate a CSS in a fully-distributed environment where all secondary users (SUs) are equipped with directional antennas and make decisions based solely on their local knowledge without information sharing between SUs. First, we establish a stochastic sequential optimization problem, which is an NP-hard, that maximizes the SU’s detection accuracy by the dynamic and optimal control of the energy sensing/detection threshold. It can enable SUs to select an available channel and sector without causing interference to the primary network. To address it in a distributed environment, the problem is transformed into a decentralized partially observed Markov decision process (Dec-POMDP) problem. Second, in order to determine the best control for the Dec-POMDP in a practical environment without any prior knowledge of state?action transition probabilities, we develop a multi-agent deep deterministic policy gradient (MADDPG)-based algorithm, which is referred to as MA-DCSS. This algorithm adopts the centralized training and decentralized execution (CTDE) architecture. Third, we analyzed its computational complexity and showed the proposed approach’s scalability by the polynomial computational complexity, in terms of the number of channels, sectors, and SUs. Lastly, the simulation confirms that the proposed scheme provides enhanced performance in terms of convergence speed, accurate detection, and false alarm probabilities when it is compared to baseline algorithms.
[NRF 연계] 한국통신학회 ICT Express Vol.8 No.4 2022.12 pp.525-529
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In B5G heterogeneous cellular networks, a rapid increase in the number of small cell base stations (SBSs) to support a massive number of devices tends to waste a considerable amount of energy. Therefore, intelligent management of SBSs’ power consumption is one of the most important research issues. We herein propose quasi-distributed Q-learning-based cell breathing (QD-QCB) considering full and partial SBS collaborations for maximizing network energy efficiency. Also, the concept of an aggregated active SBS set based on regional user distributions is proposed for computing- and energy-efficient operation. Through intensive simulations, we show that the proposed QD-QCB algorithm can achieve optimal energy efficiency, and improve the network energy efficiency significantly compared with conventional algorithms such as no transmit power control, random cell breathing, and greedy cell breathing algorithms.
[NRF 연계] 한국통신학회 ICT Express Vol.9 No.5 2023.10 pp.776-782
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper investigates the Contention Window (CW) optimization problem in multi-agent scenarios, where the fully cooperative among mobile stations is considered. A partially observable environment is employed to model and analyze the CW optimization problem, and Smart Exponential-Threshold-Linear with Deep Q-learning Network (SETL-DQN) Multi-Agent (MA) algorithm is proposed to obtain the optimal system throughput through the CW Threshold optimization. In the determined scenarios, SETL-DQN(MA) can effectively cope with the mutual interaction among mobile stations. The simulation results show that our proposed method is superior from both static and dynamic scenarios and has the highest optimum packet transmission efficiency.
[NRF 연계] 한국통신학회 ICT Express Vol.11 No.3 2025.06 pp.473-480
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This study introduces the Multi-Agent, Multi-Parameter, Interaction-Driven Contention Window Optimization (M2I-CWO) algorithm, a novel Multi-Agent Deep Reinforcement Learning (MADRL) framework designed to optimize multiple CW parameters in IEEE 802.11 Wireless LANs. Unlike single-parameter or specialized multi-agent methods, M2I-CWO employs a Dueling-DQN architecture and an Adaptive Interaction Reward Function?spanning independent, cooperative, competitive, and mixed modes?and accommodates Hierarchical Multi-Agent System (HMAS) or Federated RL (FRL) for further scalability. First, multiple CW parameters are simultaneously adjusted to enhance collision management. Second, M2I-CWO consistently achieves throughput improvements in both static and dynamic scenarios. Extensive results confirm M2I-CWO's superiority in efficiency and adaptability.
A Research on Machine Learning Agent in Rogue-like game KCI 등재
한국컴퓨터게임학회 컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) 제37권 제1호 2024.04 pp.33-39
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
실제세계에서 데이터 수집의 비용과 한계를 고려할 때, 시뮬레이션 생성 환경은 데이터 생성 과 다양한 시도에 있어 효율적인 대안이다. 이 연구에서는 Unity ML Agent를 로그라이크 장 르에 적합한 강화학습 모델로 구현하였다. 간단한 게임에Agent를 이식하고, 이 Agent가 적을 인식하고 대응하는 과정을 코드로 작성하였다. 초기 모델은 조준사격의 한계를 보였으나 RayPerceptionSensor-Component2D를 통해 Agent의 센서 정보를 직접 제공함으로써, Agent가 적을 감지하고 조준 사격을 하는 능력을 관찰할 수 있었다. 결과적으로, 개선된 모델 은 평균3.81배 향상된 성능을 보여주었으며, 이는 Unity ML Agent가 로그라이크 장르에서 강화학습을 통한 데이터 수집이 가능함을 입증한다.
Collecting large amounts of data in the real world is expensive and has clear limitations. Simulation-generated environments, on the other hand, offer the opportunity to efficiently generate the necessary data and to try different things easily and quickly. In this research, we utilized one of the tools that addresses these challenges, by the Unity Machine Learning tool, to study an efficient automation model that responds to the characteristics of the rogue-like genre. For testing purposes, we implemented a simple game, implanted an agent into the main character of the game, and fed the agent with code to shoot and avoid hostile. The implemented ML Agent successfully recognized the hostile targets and responded by shooting and dodging them. However, instead of learning to prioritize the hostile targets over time by reinforcing itself and shooting the high-risk targets first, it consistently fired in only one of the 360-degree directions given to it at the beginning, which we didn’t expected, so we improved the code. By utilizing the RayPerceptionSensor-Component2D element to directly feed the agent's sensors with information about hostile targets, we found that the agent was able to utilize its ray sensor to detect them and make much more precise aimed shots. In fact, it outperformed the original model by an average of 3.81x, proving that Unity ML Agentcan collect data through reinforcement learning in the roguelike genre.
A Multi-agent Reinforcement Learning Approach of Dynamic Traffic Assignment
한국ITS학회 한국ITS학회 학술대회 Net-Zero Mobility 2023.04 pp.731-737
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Emergent Emotional Dynamics and Intrinsic Motivation in Multi-Agent Reinforcement Learning Systems.
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 ICNGC 2025 The 11th International Conference on Next Generation Computing 2025 2025.12 pp.124-128
This study investigates the emergence of emotional dynamics and intrinsic motivation within multi-agent reinforcement learning (MARL) systems. Traditional MARL frameworks rely solely on extrinsic task rewards, which often limit exploration and adaptability. To address this, we propose an Affective-Motivated MARL (AM-MARL) framework where agents integrate curiosity-based intrinsic rewards and emotionmodulated affective feedback alongside extrinsic reinforcement. Agents operate in a continuous multiagent environment, learning through Q-learning, Actor- Critic, or Advantage Actor-Critic (A2C) methods depending on their action space. The intrinsic reward is defined as the state-prediction error between observed and expected future states, while the affective reward arises from temporal changes in emotional state and the social influence among peers. Experimental results show that incorporating intrinsic and affective rewards enhances exploration coverage, stabilizes emotional trajectories, and improves coordination efficiency compared to extrinsic-only baselines. These findings suggest that emotional feedback, when coupled with curiosity-driven intrinsic signals, fosters more humanlike adaptability, cooperative intelligence, and stable affect regulation in MARL environments.
한국ITS학회 한국ITS학회 학술대회 Inclusive ITS Technologies 2024.04 pp.321-328
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
한국컴퓨터게임학회 컴퓨터게임및콘텐츠논문지(구 한국컴퓨터게임학회논문지) 제36권 제4호 2023.12 pp.23-29
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
다중 에이전트 강화학습의 발전과 함께 게임 분야에서 강화학습을 레벨 디자인에 적용하려는 연구가 계속되 고 있다. 플랫폼의 형태가 레벨 디자인의 중요한 요소임에도 불구하고 지금까지의 연구들은 플레이어의 스킬 수준이나, 스킬 구성 등 플레이어의 매트릭에 초첨을 맞춰 강화학습을 활용하였다. 따라서 본 논문에서는 레 벨 디자인에 플랫폼의 형태가 사용될 수 있도록 시각 센서의 가시성과 구조물의 복잡성을 고려하여 플랫폼 이 플레이 경험에 미치는 영향을 연구한다. 이를 위해Unity ML-Agents Toolkit과MA-POCA 알고리즘, Self-play 방식을 기반으로2vs2 대전 슈팅 게임 환경을 개발하였으며 다양한 플랫폼의 형태를 구성하였다. 분석을 통해 플랫폼의 형태에 따른 가시성과 복잡성의 차이가 승률 밸런스에는 크게 영향을 미치지 않으나 전체 에피소 드 수, 무승부 비율, Elo의 증가폭에 유의미한 영향을 미치는 것을 확인했다.
As multi-agent reinforcement learning advance, it has been utilized for level design in game with multiple players and complex interactive environments. Despite the significant impact that the platform's design has on game level design, player-related metrics including behavior, skill level, and skill composition have been the primary focus of most reinforcement learning studies. In this paper, we study the impact of platform design on playing experience by utilizing multi-agent reinforcement learning. we use the Unity Engine to create various platform designs based on a 2vs2 shooting game and visual sensor visibility and structural complexity are taken into consideration for analysis. With Unity ML-Agents Toolkit, we construct a reinforcement learning environment based on the MA-POCA algorithm and the self-play method. Though the agent, we analyze changes in play experience with win rate, the number of episode, draw rate, and the Elo rating value. With the study, we figure out that visibility and complexity of platform design exhibit minimal impact on win rate balance in self-play method, but does influence the total number of episodes, draw rate, and the increase in Elo rating value.
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 10th International Conference on Next Generation Computing 2024 2024.11 pp.319-322
In this paper, we investigate the application of Nash equilibrium strategy to enhance coordination and decision-making in multi-agent systems, specifically focusing on Automated Guided Vehicles (AGVs) systems. Traditional reinforcement learning methods often face challenges in multiagent environments due to the non-stationarity introduced by multiple learning agents and the complexity of coordinating actions among them. To address these challenges, we propose an approach that integrates game-theoretic principles of Nash equilibrium into existing multi-agent reinforcement learning frameworks. By incorporating Nash equilibrium considerations into the policy update mechanisms, agents can anticipate and respond to the strategies of other agents proactively. This integration reduces conflicts and improves cooperation without relying solely on reward shaping or penalization for undesirable behaviors, such as collisions. Additionally, we introduce a collaboration cost into the reward function to further incentivize cooperative behavior among agents. We validate the effectiveness of our approach in a flexible manufacturing system simulated using PyBullet, utilizing the default URDF models to create a realistic and standardized environment. Multiple AGVs operate as autonomous agents tasked with collaboratively optimizing production tasks. Experimental results demonstrate that our Nash equilibrium-based method significantly outperforms traditional algorithms—including MADDPG, NDQN, CQL, COMA, IQL, PPO, SAC, and DQN—in terms of cumulative reward, policy convergence speed, and overall system throughput.
Aspect-based Sentiment Analysis of Product Reviews using Multi-agent Deep Reinforcement Learning KCI 등재 SCOPUS
한국경영정보학회 Asia Pacific Journal of Information Systems 제32권 제2호 2022.06 pp.226-248
※ 기관로그인 시 무료 이용이 가능합니다.
6,000원
The existing model for sentiment analysis of product reviews learned from past data and new data was labeled based on training. But new data was never used by the existing system for making a decision. The proposed Aspect-based multi-agent Deep Reinforcement learning Sentiment Analysis (ADRSA) model learned from its very first data without the help of any training dataset and labeled a sentence with aspect category and sentiment polarity. It keeps on learning from the new data and updates its knowledge for improving its intelligence. The decision of the proposed system changed over time based on the new data. So, the accuracy of the sentiment analysis using deep reinforcement learning was improved over supervised learning and unsupervised learning methods. Hence, the sentiments of premium customers on a particular site can be explored to other customers effectively. A dynamic environment with a strong knowledge base can help the system to remember the sentences and usage State Action Reward State Action (SARSA) algorithm with Bidirectional Encoder Representations from Transformers (BERT) model improved the performance of the proposed system in terms of accuracy when compared to the state of art methods.
Animated Pedagogical Agent: New Possibility for Computer-Assisted Language Learning
한국어정보학회 한국어정보학 제7권 1호 2005.06 pp.37-41
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
An animated pedagogical agent is a computerized lifelike character that cohabits an interactive computer-based learning environment with a learner. It is a product of recent technological advances in user interface and autonomous software agents. The animated pedagogical agents have great potential in promoting foreign language learning and instruction. Animated pedagogical agents can be used to implement the interactionist approach and task-based approach to second language acquisition. Moreover, an agent-embedded L2 learning environment can lower learner’affective filter by providing interesting and entertaining learning experiences, that is, by generating the Persona Effect. Animated edagogical agents can offer great possibilities for designing and developing computer-based foreign language learning environment.
Animated Pedagogical Agent : New Possibility for Computer-Assisted Language Learning
한국어정보학회 한국어정보학회 국제학술대회 종전과 광복 60주년 및 안의사 서거 95주년 기념 2005.08 pp.115-118
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Generation of Hypotheses on the Evolution of Agent-Based Business using Inductive Learning
한국경영정보학회 한국경영정보학회 정기 학술대회 2002년 춘계학술대회 2002.06 pp.293-302
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
매치 3 게임 플레이를 위한 PPO 알고리즘을 이용한 강화학습 에이전트의 설계 및 구현 KCI 등재
중소기업융합학회 융합정보논문지(구 중소기업융합학회논문지) 제11권 제3호 2021.03 pp.1-6
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
매치 3 퍼즐 게임들은 주로 MCTS(Monte Carlo Tree Search) 알고리즘을 사용하여 자동 플레이를 구현 하였지만 MCTS의 느린 탐색 속도로 인해 MCTS와 DNN(Deep Neural Network)을 함께 적용하거나 강화학습 으로 인공지능을 구현하는 것이 일반적인 경향이다. 본 연구에서는 매치 3 게임 개발에 주로 사용되는 유니티3D 엔진과 유니티 개발사에서 제공해주는 머신러닝 SDK를 이용하여 PPO(Proximal Policy Optimization) 알고리 즘을 적용한 강화학습 에이전트를 설계 및 구현하여, 그 성능을 확인해본 결과, 44% 정도 성능이 향상되었음을 확인하였다. 실험 결과 에이전트가 게임 규칙을 배우고 실험이 진행됨에 따라 더 나은 전략적 결정을 도출 해 낼 수 있는 것을 확인할 수 있었으며 보통 사람들보다 퍼즐 게임을 더 잘 수행하는 결과를 확인하였다. 본 연구에서 설계 및 구현한 에이전트가 일반 사람들보다 더 잘 플레이하는 만큼, 기계와 인간 플레이 수준 사이의 간극을 조절 하여 게임의 레벨 디지인에 적용된다면 향후 빠른 스테이지 개발에 도움이 될 것으로 기대된다.
Most of the match-3 puzzle games supports automatic play using the MCTS algorithm. However, implementing reinforcement learning agents is not an easy job because it requires both the knowledge of machine learning and the way of complex interactions within the development environment. This study proposes a method in which we can easily design reinforcement learning agents and implement game play agents by applying PPO(Proximal Policy Optimization) algorithms. And we could identify the performance was increased about 44% than the conventional method. The tools we used are the Unity 3D game engine and Unity ML SDK. The experimental result shows that agents became to learn game rules and make better strategic decisions as experiments go on. On average, the puzzle gameplay agents implemented in this study played puzzle games better than normal people. It is expected that the designed agent could be used to speed up the game level design process.
의사결정 트리를 이용한 학습 에이전트 단기주가예측 시스템 개발 KCI 등재후보
대한안전경영과학회 대한안전경영과학회지 제6권 제2호 2004.06 pp.211-229
※ 기관로그인 시 무료 이용이 가능합니다.
5,400원
The basis of cyber trading has been sufficiently developed with innovative advancement of Internet Technology and the tendency of stock market investment has changed from long-term investment, which estimates the value of enterprises, to short-term investment, which focuses on getting short-term stock trading margin. Hence, this research shows a Short-term Stock Price Forecasting System on Learning Agent System using DTA(Decision Tree Algorithm) ; it collects real-time information of interest and favorite issues using Agent Technology through the Internet, and forms a decision tree, and creates a Rule-Base Database. Through this procedure the Short-term Stock Price Forecasting System provides customers with the prediction of the fluctuation of stock prices for each issue in near future and a point of sales and purchases. A Human being has the limitation of analytic ability and so through taking a look into and analyzing the fluctuation of stock prices, the Agent enables man to trace out the external factors of fluctuation of stock market on real-time. Therefore, we can check out the ups and downs of several issues at the same time and figure out the relationship and interrelation among many issues using the Agent. The SPFA (Stock Price Forecasting System) has such basic four phases as Data Collection, Data Processing, Learning, and Forecasting and Feedback.
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.