년 - 년
Multi-agent reinforcement learning for a distributed multi-channel access game
[NRF 연계] 한국통신학회 ICT Express Vol.11 No.5 2025.10 pp.863-869
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In this work, we model multi-user distributed channel access as a game with channels and users, and propose the Multi-Agent Thompson Sampling (MA-TS) algorithm. It uses Bayes’ theorem to dynamically optimize action selection. This optimization aims to maximize throughput. We derive the algorithm’s computational complexity as . Simulations show that MA-TS converges to a pure strategy Nash equilibrium (PNE) and outperforms existing methods in average throughput.
[NRF 연계] 한국통신학회 ICT Express Vol.10 No.3 2024.06 pp.472-478
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In CRNs, it is crucial to develop an efficient and reliable spectrum detector that consistently provides accurate information about the channel state. In this work, we investigate a CSS in a fully-distributed environment where all secondary users (SUs) are equipped with directional antennas and make decisions based solely on their local knowledge without information sharing between SUs. First, we establish a stochastic sequential optimization problem, which is an NP-hard, that maximizes the SU’s detection accuracy by the dynamic and optimal control of the energy sensing/detection threshold. It can enable SUs to select an available channel and sector without causing interference to the primary network. To address it in a distributed environment, the problem is transformed into a decentralized partially observed Markov decision process (Dec-POMDP) problem. Second, in order to determine the best control for the Dec-POMDP in a practical environment without any prior knowledge of state?action transition probabilities, we develop a multi-agent deep deterministic policy gradient (MADDPG)-based algorithm, which is referred to as MA-DCSS. This algorithm adopts the centralized training and decentralized execution (CTDE) architecture. Third, we analyzed its computational complexity and showed the proposed approach’s scalability by the polynomial computational complexity, in terms of the number of channels, sectors, and SUs. Lastly, the simulation confirms that the proposed scheme provides enhanced performance in terms of convergence speed, accurate detection, and false alarm probabilities when it is compared to baseline algorithms.
[NRF 연계] 한국통신학회 ICT Express Vol.8 No.4 2022.12 pp.525-529
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
In B5G heterogeneous cellular networks, a rapid increase in the number of small cell base stations (SBSs) to support a massive number of devices tends to waste a considerable amount of energy. Therefore, intelligent management of SBSs’ power consumption is one of the most important research issues. We herein propose quasi-distributed Q-learning-based cell breathing (QD-QCB) considering full and partial SBS collaborations for maximizing network energy efficiency. Also, the concept of an aggregated active SBS set based on regional user distributions is proposed for computing- and energy-efficient operation. Through intensive simulations, we show that the proposed QD-QCB algorithm can achieve optimal energy efficiency, and improve the network energy efficiency significantly compared with conventional algorithms such as no transmit power control, random cell breathing, and greedy cell breathing algorithms.
[NRF 연계] 한국통신학회 ICT Express Vol.9 No.5 2023.10 pp.776-782
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This paper investigates the Contention Window (CW) optimization problem in multi-agent scenarios, where the fully cooperative among mobile stations is considered. A partially observable environment is employed to model and analyze the CW optimization problem, and Smart Exponential-Threshold-Linear with Deep Q-learning Network (SETL-DQN) Multi-Agent (MA) algorithm is proposed to obtain the optimal system throughput through the CW Threshold optimization. In the determined scenarios, SETL-DQN(MA) can effectively cope with the mutual interaction among mobile stations. The simulation results show that our proposed method is superior from both static and dynamic scenarios and has the highest optimum packet transmission efficiency.
[NRF 연계] 한국통신학회 ICT Express Vol.11 No.3 2025.06 pp.473-480
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
This study introduces the Multi-Agent, Multi-Parameter, Interaction-Driven Contention Window Optimization (M2I-CWO) algorithm, a novel Multi-Agent Deep Reinforcement Learning (MADRL) framework designed to optimize multiple CW parameters in IEEE 802.11 Wireless LANs. Unlike single-parameter or specialized multi-agent methods, M2I-CWO employs a Dueling-DQN architecture and an Adaptive Interaction Reward Function?spanning independent, cooperative, competitive, and mixed modes?and accommodates Hierarchical Multi-Agent System (HMAS) or Federated RL (FRL) for further scalability. First, multiple CW parameters are simultaneously adjusted to enhance collision management. Second, M2I-CWO consistently achieves throughput improvements in both static and dynamic scenarios. Extensive results confirm M2I-CWO's superiority in efficiency and adaptability.
Learning graph based individual intrinsic reward for multi-agent reinforcement learning
[NRF 연계] 한국통신학회 ICT Express Vol.12 No.2 2026.04 pp.301-305
※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.
Designing a reward function is a critical challenge in reinforcement learning. However, as environments become more complex and tasks grow more difficult, designing a reward function that drives optimal behavior becomes increasingly challenging. To overcome these issues, Preference based reinforcement learning has proposed methods that learn reward functions based on the preference between two trajectories, thereby eliminating the need for handcrafted reward function. In multi-agent reinforcement learning, the challenge is even greater due to the complex interactions among agents, which makes designing a single global reward function even more difficult. In this paper, we show that when a single global reward function is learned via preference-based reinforcement learning in multi-agent setting, it often fails to capture sufficient information for optimal policy learning. Instead, we propose a method for learning individual reward functions that provide additional guidance for each agent’s optimal policy. Our approach, which leverages graph structures and preference-based reinforcement learning, outperforms the method based on learning a single, global reward function.
Multi-Agent를 이용한 강화학습 기반 실시간 신호제어
한국ITS학회 한국ITS학회 학술대회 SMART CITY 새롭게 펼쳐지는 교통 시스템 2018.04 pp.279-283
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
A Multi-Agent Collaborative Framework for StakehoIder-Aware Recommendation Governance
한국경영정보학회 한국경영정보학회 정기 학술대회 AX 시대 데이터 경제와 비즈니스 혁신: 가치창출 경영과 융합 생태계 2026.06 pp.951-959
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
A Multi-agent Reinforcement Learning Approach of Dynamic Traffic Assignment
한국ITS학회 한국ITS학회 학술대회 Net-Zero Mobility 2023.04 pp.731-737
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
한국정보기술응용학회 JITAM Vol.16 No.1 2009.03 pp.1-20
※ 기관로그인 시 무료 이용이 가능합니다.
5,500원
Negotiation is a process of reaching an agreement on the terms of a transaction, such as price, quantity, for two or more parties. Negotiation tries to maximize the benefits for all parties concerned. Instead of using human-based negotiation, the e-commerce environment provides such an environment as adopting automated negotiation. Thus, choosing agent technology is appropriate for an automatic electronic negotiation platform, since autonomous software agents strive for the best deal on behalf of the human participants. Negotiation agents need a clear-cut definition of negotiation models or strategies. In reality, most bargaining systems embody nearly one negotiation model. In this article, we present a mobile agent negotiation system with reusable negotiation strategies that allows agents to dynamically embody a user’s favorite negotiation strategy which can be preinstalled as a component in the system. We develop a prototype system, which is fully implemented in compliance with FIPA specifications, and then, describe the benefits of using the system.
Optimal Multi-Agent Performance Measures for Team Contracts
한국재무학회 한국재무학회 학술대회 2007년 5개 학회 공동학술연구발표회 2007.04 pp.896-917
※ 기관로그인 시 무료 이용이 가능합니다.
5,800원
We present a continuous-time contracting model under moral hazard with many agents. The principal contracts many agents as a team, and they jointly produce correlated outcomes. We show the optimal contract for each agent is linear in outcomes of all other agents as well as his/her own. The structure of the optimal contract strikingly reveals that the optimal aggregate performance measure in general can be orthogonally decomposed into two statistics: one is a sucient statistic, and the other a non-sucient statistic. As a consequence, the optimal aggregate performance measure in general is not a sucient statistic, unless the principal is risk neutral. We further discuss agents' optimal eort choices using a \quadratic-cost" example, which also strikingly suggests that team contracts sometimes provide lower-powered eort incentives than individually separate contracts do.
A Multi-agent Architecture for Coordination of Supply Chains with Local Information Sharing
한국경영정보학회 한국경영정보학회 정기 학술대회 2003년 춘계학술대회 2003.06 pp.659-667
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
4,000원
본 도시 교통 혼잡은 간선도로망 전반에서 장시간의 지연, 배출 증가, 그리고 안전성 저하를 초래한다. 본 논문 은 스마트 폴 엣지 컴퓨팅과 연합형 다중 에이전트 강화학습을 결합한 교차로 제어용 엣지-네이티브 프레임워크를 제안 한다. 각 교차로의 경량 에이전트는 폴 탑재 센서로부터 로컬 학습을 수행하고, 프라이버시를 보존하는 연합 절차가 네 트워크 전역에서 주기적으로 모델 업데이트를 집계한다. 통신 계층은 저지연 로컬 V2X 교환을 위한 IEEE 802.11p와 모니터링 및 모델 조정을 위한 MQTT 5.0을 통합한다. 평가는 802.11p와 NR-V2X 링크 동작을 반영하기 위해 VEINS(OMNeT++–SUMO)와 Simu5G를 결합한 디지털 트윈-인-더-루프 환경에서 수행된다. 배출량은 SUMO의 HBEFA 기반 모형으로 추정한다. 고정주기, 감응식, 비연합형 MARL 기준선과 비교할 때, 제안 방법은 네트워크 전역 평균 지연과 대기행렬 길이를 줄이는 동시에 CO₂와 NOx를 낮출 잠재력을 보인다. 본 논문은 아키텍처, 학습 및 통신 설계, 그리고 시나리오 구성·평가지표·통계 검정을 위한 재현 가능한 프로토콜을 상세히 기술한다.
Urban traffic congestion leads to prolonged delays, elevated emissions, and diminished safety across arterial networks. This paper proposes an edge‑native framework for intersection control that combines smart‑pole edge computing with federated multi‑agent reinforcement learning. Lightweight agents at each intersection learn locally from on‑pole sensors while a privacy‑preserving federated procedure periodically aggregates model updates across the network. The communication layer integrates IEEE 802.11p for low‑latency local V2X exchange and MQTT 5.0 for monitoring and model coordination. Evaluation is conducted in a digital‑twin‑in‑the‑loop environment that couples VEINS (OMNeT++–SUMO) with Simu5G to account for 802.11p and NR‑V2X link behavior. Emissions are estimated using SUMO’s HBEFA‑based model. Compared with fixed‑time, actuated, and non‑federated MARL baselines, the proposed approach shows the potential to reduce network‑wide average delay and queue length while simultaneously lowering CO₂ and NOx. The paper details the architecture, the learning and communication design, and a reproducible protocol for scenario construction, metrics, and statistical testing.
도로선형에 따른 Multi-agent 주행 시뮬레이터 기반 자율차-비자율차 차량추종 이벤트 주행안전성 평가
한국ITS학회 한국ITS학회 학술대회 ITS와 함께하는 미래 스마트 시티 2022.06 pp.465-489
※ 기관로그인 시 무료 이용이 가능합니다.
6,300원
Beyond the Numbers : How Multi-Agent Compassionate AI Can Foster Fairer Financial Inclusion
한국경영정보학회 한국경영정보학회 정기 학술대회 Beyond AI: Building an Inclusive and Ethical Digital Economy with Web3 2025.10 pp.101-107
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
Integrating Large Language Models (LLMs) into high-stakes evaluations like hiring and loans poses dual challenges: reducing algorithmic bias and embedding human compassion. We propose an AI evaluation framework grounded in the four-factor model of organizational justice, restructured into two dimensions: Structural Justice (procedural and distributive fairness) and Interactional Justice (interpersonal and informational compassion). Our modular, multi-agent system includes a Criteria Generator for fair rubric design and an Application Evaluator with two LLM agents—a “just bureaucrat” scoring structural fairness and a “compassionate communicator” scoring interactional fairness. These qualitative scores integrate with quantitative predictions through a Budget-Aware Fair Ranker to produce optimized outcomes. This framework offers a blueprint for AI systems that balance fairness with empathy, advancing beyond bias mitigation to foster just and compassionate decision-making in automated evaluations.
Route Selection in a Dynamic Multi-Agent Multilayer Electronic Supply Network KCI 등재
한국정보기술응용학회 JITAM Vol.17 No.1 2010.03 pp.141-155
※ 기관로그인 시 무료 이용이 가능합니다.
4,800원
We develop an intelligent information system in a multilayer electronic supply chain network. Using the internet for supply chain management (SCM) is a key interest for contemporary managers and researchers. It has been realized that the internet can facilitate SCM by making real time information available and enabling collaboration between trading partners. Here, we propose a multi-agent system to analyze the performance of the elements of a supply network based on the attributes of the information flow. Each layer consists of elements which are differentiated by their performance throughout the supply network. The proposed agents measure and record the performance flow of elements considering their web interactions for a dynamic route selection. A dynamic programming approach is applied to determine the optimal route for a customer in the end-user layer.
Emergent Emotional Dynamics and Intrinsic Motivation in Multi-Agent Reinforcement Learning Systems.
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 ICNGC 2025 The 11th International Conference on Next Generation Computing 2025 2025.12 pp.124-128
This study investigates the emergence of emotional dynamics and intrinsic motivation within multi-agent reinforcement learning (MARL) systems. Traditional MARL frameworks rely solely on extrinsic task rewards, which often limit exploration and adaptability. To address this, we propose an Affective-Motivated MARL (AM-MARL) framework where agents integrate curiosity-based intrinsic rewards and emotionmodulated affective feedback alongside extrinsic reinforcement. Agents operate in a continuous multiagent environment, learning through Q-learning, Actor- Critic, or Advantage Actor-Critic (A2C) methods depending on their action space. The intrinsic reward is defined as the state-prediction error between observed and expected future states, while the affective reward arises from temporal changes in emotional state and the social influence among peers. Experimental results show that incorporating intrinsic and affective rewards enhances exploration coverage, stabilizes emotional trajectories, and improves coordination efficiency compared to extrinsic-only baselines. These findings suggest that emotional feedback, when coupled with curiosity-driven intrinsic signals, fosters more humanlike adaptability, cooperative intelligence, and stable affect regulation in MARL environments.
한국차세대컴퓨팅학회 한국차세대컴퓨팅학회 학술대회 The 10th International Conference on Next Generation Computing 2024 2024.11 pp.319-322
In this paper, we investigate the application of Nash equilibrium strategy to enhance coordination and decision-making in multi-agent systems, specifically focusing on Automated Guided Vehicles (AGVs) systems. Traditional reinforcement learning methods often face challenges in multiagent environments due to the non-stationarity introduced by multiple learning agents and the complexity of coordinating actions among them. To address these challenges, we propose an approach that integrates game-theoretic principles of Nash equilibrium into existing multi-agent reinforcement learning frameworks. By incorporating Nash equilibrium considerations into the policy update mechanisms, agents can anticipate and respond to the strategies of other agents proactively. This integration reduces conflicts and improves cooperation without relying solely on reward shaping or penalization for undesirable behaviors, such as collisions. Additionally, we introduce a collaboration cost into the reward function to further incentivize cooperative behavior among agents. We validate the effectiveness of our approach in a flexible manufacturing system simulated using PyBullet, utilizing the default URDF models to create a realistic and standardized environment. Multiple AGVs operate as autonomous agents tasked with collaboratively optimizing production tasks. Experimental results demonstrate that our Nash equilibrium-based method significantly outperforms traditional algorithms—including MADDPG, NDQN, CQL, COMA, IQL, PPO, SAC, and DQN—in terms of cumulative reward, policy convergence speed, and overall system throughput.
Aspect-based Sentiment Analysis of Product Reviews using Multi-agent Deep Reinforcement Learning KCI 등재 SCOPUS
한국경영정보학회 Asia Pacific Journal of Information Systems 제32권 제2호 2022.06 pp.226-248
※ 기관로그인 시 무료 이용이 가능합니다.
6,000원
The existing model for sentiment analysis of product reviews learned from past data and new data was labeled based on training. But new data was never used by the existing system for making a decision. The proposed Aspect-based multi-agent Deep Reinforcement learning Sentiment Analysis (ADRSA) model learned from its very first data without the help of any training dataset and labeled a sentence with aspect category and sentiment polarity. It keeps on learning from the new data and updates its knowledge for improving its intelligence. The decision of the proposed system changed over time based on the new data. So, the accuracy of the sentiment analysis using deep reinforcement learning was improved over supervised learning and unsupervised learning methods. Hence, the sentiments of premium customers on a particular site can be explored to other customers effectively. A dynamic environment with a strong knowledge base can help the system to remember the sentences and usage State Action Reward State Action (SARSA) algorithm with Bidirectional Encoder Representations from Transformers (BERT) model improved the performance of the proposed system in terms of accuracy when compared to the state of art methods.
한국ITS학회 한국ITS학회 학술대회 Inclusive ITS Technologies 2024.04 pp.321-328
※ 기관로그인 시 무료 이용이 가능합니다.
4,000원
0개의 논문이 장바구니에 담겼습니다.
선택하신 파일을 압축중입니다.
잠시만 기다려 주십시오.