Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

-

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 46
No
1

This study investigates the emergence of emotional dynamics and intrinsic motivation within multi-agent reinforcement learning (MARL) systems. Traditional MARL frameworks rely solely on extrinsic task rewards, which often limit exploration and adaptability. To address this, we propose an Affective-Motivated MARL (AM-MARL) framework where agents integrate curiosity-based intrinsic rewards and emotionmodulated affective feedback alongside extrinsic reinforcement. Agents operate in a continuous multiagent environment, learning through Q-learning, Actor- Critic, or Advantage Actor-Critic (A2C) methods depending on their action space. The intrinsic reward is defined as the state-prediction error between observed and expected future states, while the affective reward arises from temporal changes in emotional state and the social influence among peers. Experimental results show that incorporating intrinsic and affective rewards enhances exploration coverage, stabilizes emotional trajectories, and improves coordination efficiency compared to extrinsic-only baselines. These findings suggest that emotional feedback, when coupled with curiosity-driven intrinsic signals, fosters more humanlike adaptability, cooperative intelligence, and stable affect regulation in MARL environments.

2

In this paper, we investigate the application of Nash equilibrium strategy to enhance coordination and decision-making in multi-agent systems, specifically focusing on Automated Guided Vehicles (AGVs) systems. Traditional reinforcement learning methods often face challenges in multiagent environments due to the non-stationarity introduced by multiple learning agents and the complexity of coordinating actions among them. To address these challenges, we propose an approach that integrates game-theoretic principles of Nash equilibrium into existing multi-agent reinforcement learning frameworks. By incorporating Nash equilibrium considerations into the policy update mechanisms, agents can anticipate and respond to the strategies of other agents proactively. This integration reduces conflicts and improves cooperation without relying solely on reward shaping or penalization for undesirable behaviors, such as collisions. Additionally, we introduce a collaboration cost into the reward function to further incentivize cooperative behavior among agents. We validate the effectiveness of our approach in a flexible manufacturing system simulated using PyBullet, utilizing the default URDF models to create a realistic and standardized environment. Multiple AGVs operate as autonomous agents tasked with collaboratively optimizing production tasks. Experimental results demonstrate that our Nash equilibrium-based method significantly outperforms traditional algorithms—including MADDPG, NDQN, CQL, COMA, IQL, PPO, SAC, and DQN—in terms of cumulative reward, policy convergence speed, and overall system throughput.

3

다중 에이전트 강화 학습(Multi-Agent Reinforcement Learning, MARL)은 에이전트가 시행착오를 통해 최적 의 정책을 상호 작용하고 학습하는 기계학습의 한 분야로, 여러 에이전트가 동시에 동일한 환경에서 상호 작용하고 학습하는 복잡한 시나리오를 처리한다. 이러한 복잡한 상호작용을 분석하고 이해하기는 어려운 일이며, 기존의 분석 방법으로는 이러한 복잡성을 충분히 분석에 반영하고 해석하는 데 한계가 있다. 이러한 문제를 해결하기 위해 우리 는 MARL 환경에서 에이전트의 정책과 상호작용을 시각화하고 분석할 수 있는 시각화 시스템인 MARLViz를 제공 한다. 이 시스템은 다양한 환경에서의 에이전트를 산점도로 나타내어 행동 차이를 요약적으로 보여주고 에이전트의 행동 패턴을 공간상에 시각화하는 열지도를 통해 사용자가 에이전트 간 복잡한 상호작용 패턴을 이해하고 비교할 수 있도록 설계되었다. 또한, MARLViz를 통해 에이전트들의 행동을 이해하기 위한 과업 수행을 효과적으로 수행할 수 있음을 보이기 위해 대표적인 분석 시나리오를 사례 연구의 형식으로 제시한다.

Multi-Agent Reinforcement Learning (MARL) is a subfield of machine learning in which agents interact and learn optimal policies through trial and error in a shared environment, handling complex scenarios where multiple agents interact simultaneously. Analyzing and understanding these intricate interactions is a challenging task, and existing methods often fall short of adequately capturing and interpreting this complexity. To address this issue, we present MARLViz, a visualization system designed to visualize and analyze agent policies and interactions in MARL environments. MARLViz represents agents in various environments using scatterplots to provide a summary of behavioral differences and visualizes agents' behavioral patterns on a spatial heatmap, enabling users to understand and compare complex interaction patterns among agents. Furthermore, we demonstrate the effectiveness of MARLViz in comprehending agents' behavior through a case study, showcasing its capability in facilitating the understanding of task performance in MARL environments.

4

Multi-agent reinforcement learning for a distributed multi-channel access game

Zhongyang Li, Yu Zhao, 이주현

[NRF 연계] 한국통신학회 ICT Express Vol.11 No.5 2025.10 pp.863-869

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this work, we model multi-user distributed channel access as a game with channels and users, and propose the Multi-Agent Thompson Sampling (MA-TS) algorithm. It uses Bayes’ theorem to dynamically optimize action selection. This optimization aims to maximize throughput. We derive the algorithm’s computational complexity as . Simulations show that MA-TS converges to a pure strategy Nash equilibrium (PNE) and outperforms existing methods in average throughput.

5

Learning graph based individual intrinsic reward for multi-agent reinforcement learning

주석훈, 한승엽, 조태현, 이정우, Lee Taeyoung, Kim Minkyoung, An Jinho

[NRF 연계] 한국통신학회 ICT Express Vol.12 No.2 2026.04 pp.301-305

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Designing a reward function is a critical challenge in reinforcement learning. However, as environments become more complex and tasks grow more difficult, designing a reward function that drives optimal behavior becomes increasingly challenging. To overcome these issues, Preference based reinforcement learning has proposed methods that learn reward functions based on the preference between two trajectories, thereby eliminating the need for handcrafted reward function. In multi-agent reinforcement learning, the challenge is even greater due to the complex interactions among agents, which makes designing a single global reward function even more difficult. In this paper, we show that when a single global reward function is learned via preference-based reinforcement learning in multi-agent setting, it often fails to capture sufficient information for optimal policy learning. Instead, we propose a method for learning individual reward functions that provide additional guidance for each agent’s optimal policy. Our approach, which leverages graph structures and preference-based reinforcement learning, outperforms the method based on learning a single, global reward function.

6

Multi-agent reinforcement learning based optimal energy sensing threshold control in distributed cognitive radio networks with directional antenna

PHAMTHITHUHIEN, 노원종, Cho Sungrae

[NRF 연계] 한국통신학회 ICT Express Vol.10 No.3 2024.06 pp.472-478

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In CRNs, it is crucial to develop an efficient and reliable spectrum detector that consistently provides accurate information about the channel state. In this work, we investigate a CSS in a fully-distributed environment where all secondary users (SUs) are equipped with directional antennas and make decisions based solely on their local knowledge without information sharing between SUs. First, we establish a stochastic sequential optimization problem, which is an NP-hard, that maximizes the SU’s detection accuracy by the dynamic and optimal control of the energy sensing/detection threshold. It can enable SUs to select an available channel and sector without causing interference to the primary network. To address it in a distributed environment, the problem is transformed into a decentralized partially observed Markov decision process (Dec-POMDP) problem. Second, in order to determine the best control for the Dec-POMDP in a practical environment without any prior knowledge of state?action transition probabilities, we develop a multi-agent deep deterministic policy gradient (MADDPG)-based algorithm, which is referred to as MA-DCSS. This algorithm adopts the centralized training and decentralized execution (CTDE) architecture. Third, we analyzed its computational complexity and showed the proposed approach’s scalability by the polynomial computational complexity, in terms of the number of channels, sectors, and SUs. Lastly, the simulation confirms that the proposed scheme provides enhanced performance in terms of convergence speed, accurate detection, and false alarm probabilities when it is compared to baseline algorithms.

7

Applying multi-agent deep reinforcement learning for contention window optimization to enhance wireless network performance

Ke Chih-Heng, Astuti Lia

[NRF 연계] 한국통신학회 ICT Express Vol.9 No.5 2023.10 pp.776-782

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This paper investigates the Contention Window (CW) optimization problem in multi-agent scenarios, where the fully cooperative among mobile stations is considered. A partially observable environment is employed to model and analyze the CW optimization problem, and Smart Exponential-Threshold-Linear with Deep Q-learning Network (SETL-DQN) Multi-Agent (MA) algorithm is proposed to obtain the optimal system throughput through the CW Threshold optimization. In the determined scenarios, SETL-DQN(MA) can effectively cope with the mutual interaction among mobile stations. The simulation results show that our proposed method is superior from both static and dynamic scenarios and has the highest optimum packet transmission efficiency.

8

A comprehensive multi-agent deep reinforcement learning framework with adaptive interaction strategies for contention window optimization in IEEE 802.11 Wireless LANs

Yi-Hao Tu, Yi-Wei Ma

[NRF 연계] 한국통신학회 ICT Express Vol.11 No.3 2025.06 pp.473-480

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

This study introduces the Multi-Agent, Multi-Parameter, Interaction-Driven Contention Window Optimization (M2I-CWO) algorithm, a novel Multi-Agent Deep Reinforcement Learning (MADRL) framework designed to optimize multiple CW parameters in IEEE 802.11 Wireless LANs. Unlike single-parameter or specialized multi-agent methods, M2I-CWO employs a Dueling-DQN architecture and an Adaptive Interaction Reward Function?spanning independent, cooperative, competitive, and mixed modes?and accommodates Hierarchical Multi-Agent System (HMAS) or Federated RL (FRL) for further scalability. First, multiple CW parameters are simultaneously adjusted to enhance collision management. Second, M2I-CWO consistently achieves throughput improvements in both static and dynamic scenarios. Extensive results confirm M2I-CWO's superiority in efficiency and adaptability.

9

A Multi-agent Reinforcement Learning Approach of Dynamic Traffic Assignment

Hyunsoo Yun, Dong-Kyu Kim

한국ITS학회 한국ITS학회 학술대회 Net-Zero Mobility 2023.04 pp.731-737

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

11

4,000원

다중 에이전트 강화학습의 발전과 함께 게임 분야에서 강화학습을 레벨 디자인에 적용하려는 연구가 계속되 고 있다. 플랫폼의 형태가 레벨 디자인의 중요한 요소임에도 불구하고 지금까지의 연구들은 플레이어의 스킬 수준이나, 스킬 구성 등 플레이어의 매트릭에 초첨을 맞춰 강화학습을 활용하였다. 따라서 본 논문에서는 레 벨 디자인에 플랫폼의 형태가 사용될 수 있도록 시각 센서의 가시성과 구조물의 복잡성을 고려하여 플랫폼 이 플레이 경험에 미치는 영향을 연구한다. 이를 위해Unity ML-Agents Toolkit과MA-POCA 알고리즘, Self-play 방식을 기반으로2vs2 대전 슈팅 게임 환경을 개발하였으며 다양한 플랫폼의 형태를 구성하였다. 분석을 통해 플랫폼의 형태에 따른 가시성과 복잡성의 차이가 승률 밸런스에는 크게 영향을 미치지 않으나 전체 에피소 드 수, 무승부 비율, Elo의 증가폭에 유의미한 영향을 미치는 것을 확인했다.

As multi-agent reinforcement learning advance, it has been utilized for level design in game with multiple players and complex interactive environments. Despite the significant impact that the platform's design has on game level design, player-related metrics including behavior, skill level, and skill composition have been the primary focus of most reinforcement learning studies. In this paper, we study the impact of platform design on playing experience by utilizing multi-agent reinforcement learning. we use the Unity Engine to create various platform designs based on a 2vs2 shooting game and visual sensor visibility and structural complexity are taken into consideration for analysis. With Unity ML-Agents Toolkit, we construct a reinforcement learning environment based on the MA-POCA algorithm and the self-play method. Though the agent, we analyze changes in play experience with win rate, the number of episode, draw rate, and the Elo rating value. With the study, we figure out that visibility and complexity of platform design exhibit minimal impact on win rate balance in self-play method, but does influence the total number of episodes, draw rate, and the increase in Elo rating value.

12

Aspect-based Sentiment Analysis of Product Reviews using Multi-agent Deep Reinforcement Learning KCI 등재 SCOPUS

M. Sivakumar, Srinivasulu Reddy Uyyala

한국경영정보학회 Asia Pacific Journal of Information Systems 제32권 제2호 2022.06 pp.226-248

※ 기관로그인 시 무료 이용이 가능합니다.

6,000원

The existing model for sentiment analysis of product reviews learned from past data and new data was labeled based on training. But new data was never used by the existing system for making a decision. The proposed Aspect-based multi-agent Deep Reinforcement learning Sentiment Analysis (ADRSA) model learned from its very first data without the help of any training dataset and labeled a sentence with aspect category and sentiment polarity. It keeps on learning from the new data and updates its knowledge for improving its intelligence. The decision of the proposed system changed over time based on the new data. So, the accuracy of the sentiment analysis using deep reinforcement learning was improved over supervised learning and unsupervised learning methods. Hence, the sentiments of premium customers on a particular site can be explored to other customers effectively. A dynamic environment with a strong knowledge base can help the system to remember the sentences and usage State Action Reward State Action (SARSA) algorithm with Bidirectional Encoder Representations from Transformers (BERT) model improved the performance of the proposed system in terms of accuracy when compared to the state of art methods.

13

다수의 자동 운반 차량(AGV)을 운용하는 시스템에서는 AGV 간의 충돌, 교착 상태 발생, 그리 고 비효율적인 대기 시간 증가로 인해 전체 시스템의 효율성이 저하되는 문제가 빈번하게 발생한 다. 본 논문은 이러한 문제를 해결하기 위해, 연속적인 공간으로 모델링된 공장 환경에서 100대의 AGV가 상호 협력을 통해 최적의 경로를 계획하도록 하는 PPO(Proximal Policy Optimization) 기 반 분산 강화학습 모델의 실험 설계를 제안한다. 제안하는 설계는 시뮬레이션 환경 구성, 개별 AGV 에이전트 모델링, 보상 함수 체계, 그리고 학습 구조를 체계적으로 정의함으로써, AGV 간의 충돌을 회피하고 교착 상태를 방지하며 이동 경로상의 비효율을 최소화하는 것을 목표로 한다. 이 설계를 통해 대규모 다중 AGV 시스템의 운영 안전성과 효율성을 향상시킬 수 있을 것으로 기대한다.

14

다중 에이전트 강화학습 기반 특징 선택에 대한 연구 KCI 등재

김민우, 배진희, 왕보현, 임준식

한국디지털정책학회 디지털융복합연구 제19권 제12호 2021.12 pp.347-352

※ 기관로그인 시 무료 이용이 가능합니다.

4,000원

본 논문은 다중 에이전트 강화학습 방식을 사용하여 입력 데이터로부터 분류에 효과적인 특징 집합을 찾아내는 방식을 제안한다. 기계 학습 분야에 있어서 분류에 적합한 특징들을 찾아내는 것은 매우 중요하다. 데이터에는 수많은 특징들이 존재할 수 있으며, 여러 특징들 중 일부는 분류나 예측에 효과적일 수 있지만 다른 특징들은 잡음 역할을 함으 로써 올바른 결과를 생성하는 데에 오히려 악영향을 줄 수 있다. 기계 학습 문제에서 분류나 예측 정확도를 높이기 위한 특징 선택은 매우 중요한 문제 중 하나이다. 이러한 문제를 해결하기 위해 강화학습을 통한 특징 선택 방법을 제시한다. 각각의 특징들은 하나의 에이전트를 가지게 되며, 이 에이전트들은 특징을 선택할 것인지 말 것인지에 대한 여부를 결정 한다. 에이전트들에 의해 선택된 특징들과 선택되지 않은 특징들에 대해서 각각 보상을 구한 뒤, 보상에 대한 비교를 통해 에이전트의 Q-value 값을 업데이트 한다. 두 하위 집합에 대한 보상 비교는 에이전트로 하여금 자신의 행동이 옳은지에 대한 판단을 내릴 수 있도록 도와준다. 이러한 과정들을 에피소드 수만큼 반복한 뒤, 최종적으로 특징들을 선별한다. 이 방법을 통해 Wisconsin Breast Cancer, Spambase, Musk, Colon Cancer 데이터 세트에 적용한 결과, 각각 0.0385, 0.0904, 0.1252, 0.2055의 정확도 향상을 보여주었으며, 최종적으로 0.9789, 0.9311, 0.9691, 0.9474 의 분류 정확도를 보여주었다. 이는 우리가 제안한 방법이 분류에 효과적인 특징들을 잘 선별하고 분류에 대한 정확도를 높일 수 있음을 보여준다.

In this paper, we propose a method for finding feature subsets that are effective for classification in an input dataset by using a multi-agent reinforcement learning method. In the field of machine learning, it is crucial to find features suitable for classification. A dataset may have numerous features; while some features may be effective for classification or prediction, others may have little or rather negative effects on results. In machine learning problems, feature selection for increasing classification or prediction accuracy is a critical problem. To solve this problem, we proposed a feature selection method based on reinforced learning. Each feature has one agent, which determines whether the feature is selected . After obtaining corresponding rewards for each feature that is selected, but not by the agents, the Q-value of each agent is updated by comparing the rewards. The reward comparison of the two subsets helps agents determine whether their actions were right. These processes are performed as many times as the number of episodes, and finally, features are selected. As a result of applying this method to the Wisconsin Breast Cancer, Spambase, Musk, and Colon Cancer datasets, accuracy improvements of 0.0385, 0.0904, 0.1252 and 0.2055 were shown, respectively, and finally, classification accuracies of 0.9789, 0.9311, 0.9691 and 0.9474 were achieved, respectively. It was proved that our proposed method could properly select features that were effective for classification and increase classification accuracy.

15

4,000원

투자 포트폴리오는 현대 포트폴리오 이론에 근거하여 해리 마코위츠가 제안한 계량 모델로서 비체계적 위험을 감소시키는 것을 목표로 삼는 이론 체계이다. 현대 사회를 급진전시킨 컴퓨팅 환경하에서 최근 통계나 기계학습 기법을 기반으로 한 포트폴리오 전략 수립이 지배적 현상으 로 나타나고 있다. 특히, 기계학습 기법 중 강화학습을 적용한 투자 포트폴리오 구축이 주목받으면서 이에 관한 국내 연구도 수행된 바 있다. 하 지만 투자 포트폴리오에 강화학습을 적용한 연구는 활성화되지 못한 채 여전히 미흡한 수준에 머물러 있는 것으로 보인다. 본 연구는 자산 선택 과 강화학습이 통합된 포트폴리오를 위험조정 수익 극대화 측면에서 벤치마크 지수와 비교·분석하였다. 주요 실증분석 결과는 다음과 같다. 첫 째, 제안한 모든 포트폴리오들은 위험 대비 높은 투자 성능을 비교해 보면 벤치마크 지수에 비해 상대적으로 높은 누적 수익 성과를 보여 주었 다. 더불어 포트폴리오 모형 1은 모든 기간에서 정(+)의 샤프비율을 보였는데, 이러한 결과는 동작 확률 분포에 입각해 동작 선택 후 보상의 획 득을 가능케 하고, 이것을 상태 가치와 비교해 이익 계산을 실행함으로써 최적 정책의 학습 확률이 증진되었다고 판단된다. 둘째, 제안한 포트 폴리오들의 평균 최대손실폭도 벤치마크 지수보다 낮거나 비슷하게 나타났다. 이로 인해 수익 기회의 온전성 유지 및 변동성 감소를 동시에 추 구할 수 있는 것이다. 이를 통해 제안한 통합된 포트폴리오가 적정히 잘 작동되고 있음이 실증되었다. 이런 맥락에서 자산 선택과 강화학습이 통합된 포트폴리오에 관한 연구는 학문적 측면에서 핵심 연구 주제가 될 뿐만 아니라 현업의 다이렉트 인덱싱 고도화에 있어서도 근원적 탐구 대상이 된다.

Over 70 years ago, Harry Markowitz published his portfolio optimization works. Nevertheless, research is currently underway to devise more effective strategies with the aim of attaining improved returns. Recent advances in reinforcement learning have enabled the exploration of new methods to create portfolios that maximize risk-adjusted returns. This study unleashed a new approach that integrates asset selection with reinforcement learning and benchmark index. The main empirical analysis results are as follows. First, all of the proposed portfolios showed relatively high cumulative return performance compared to the benchmark index. Especially, portfolio model 1 showed a positive sharp ratio for all periods. One reason for this is that we determine an advantage based on the state value by comparing it to an action probability, which requires understanding action probabilities. Second, the average maximum drawdown of the proposed investment portfolios was also lower or similar to the benchmark index. The proposed integrated portfolio was demonstrated to result in reduced volatility while preserving profit opportunities. In this context, the propped portfolio presents an opportunity to take direct indexing investement strategy to the next level.

16

This paper investigates the application of multi-agent deep reinforcement learning in the fighting game Samurai Shodown using Proximal Policy Optimization (PPO) and Advantage Actor-Critic (A2C) algorithms. Initially, agents are trained separately for 200,000 timesteps using Convolutional Neural Network (CNN) and Multi-Layer Perceptron (MLP) with LSTM networks. PPO demonstrates superior performance early on with stable policy updates, while A2C shows better adaptation and higher rewards over extended training periods, culminating in A2C outperforming PPO after 1,000,000 timesteps. These findings highlight PPO's effectiveness for short-term training and A2C's advantages in long-term learning scenarios, emphasizing the importance of algorithm selection based on training duration and task complexity. The code can be found in this link https://github.com/Lexer04/Samurai-Shodown-with-Reinforcement-Learning-PPO.

17

Intelligent Warehousing: Comparing Cooperative MARL Strategies

Yosua Setyawan Soekamto, Dae-Ki Kang

국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.16 No.3 2024.08 pp.205-211

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Effective warehouse management requires advanced resource planning to optimize profits and space. Robots offer a promising solution, but their effectiveness relies on embedded artificial intelligence. Multi-agent reinforcement learning (MARL) enhances robot intelligence in these environments. This study explores various MARL algorithms using the Multi-Robot Warehouse Environment (RWARE) to determine their suitability for warehouse resource planning. Our findings show that cooperative MARL is essential for effective warehouse management. IA2C outperforms MAA2C and VDA2C on smaller maps, while VDA2C excels on larger maps. IA2C’s decentralized approach, focusing on cooperation over collaboration, allows for higher reward collection in smaller environments. However, as map size increases, reward collection decreases due to the need for extensive exploration. This study highlights the importance of selecting the appropriate MARL algorithm based on the specific warehouse environment's requirements and scale.

18

This study presents a reinforcement learning-based simulation of agent behaviors within a synthetic ecosystem environment. Using Unity ML-Agents and the Proximal Policy Optimization (PPO) reinforcement learning algorithm, we designed three distinct agents—a chicken, a dog, and a tiger—each with unique survival objectives such as foraging or hunting. These agents autonomously learned their behaviors through interaction and reward feedback, without predefined rules. We investigated how variations in reward structure and environmental complexity influenced policy convergence and strategic behavior formation. Experimental results showed that each agent developed distinct behavioral strategies aligned with its reward design. Furthermore, despite the high complexity of the simulated environment—including uneven terrain and multiagent interactions—all agents achieved stable learning convergence when reward signals were properly calibrated. This work contributes to the modeling of reward-driven adaptive strategies in multi-agent ecological simulations and offers a foundation for future studies involving cooperation, emergent behavior, or survival-based competition.

19

Reinforcement learning multi-agent using unsupervised learning in a distributed cloud environment

Seo-Yeon Gu, Seok-Jae Moon, Byung-Joon Park

국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.14 No.2 2022.05 pp.192-198

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Companies are building and utilizing their own data analysis systems according to business characteristics in the distributed cloud. However, as businesses and data types become more complex and diverse, the demand for more efficient analytics has increased. In response to these demands, in this paper, we propose an unsupervised learning-based data analysis agent to which reinforcement learning is applied for effective data analysis. The proposal agent consists of reinforcement learning processing manager and unsupervised learning manager modules. These two modules configure an agent with k-means clustering on multiple nodes and then perform distributed training on multiple data sets. This enables data analysis in a relatively short time compared to conventional systems that perform analysis of large-scale data in one batch.

20

A Multi-agent Cooperation Model using Reinforcement Learning for Planning Multiple Goals

Mal Rey LEE

보안공학연구지원센터(JSE) 보안공학연구논문지 Vol.2 No.3 2005.11 pp.238-243

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

 
1 2 3
페이지 저장