Earticle

현재 위치 Home 검색결과

결과 내 검색

발행연도

년 - 년

학문분야

자료유형

간행물

검색결과

검색조건
검색결과 : 15건
No
1

Distributed Deep Reinforcement Learning for Autonomous Aerial eVTOL Mobility in Drone Taxi Applications

윤원준, 정소이, 김중헌, 김재현

[NRF 연계] 한국통신학회 ICT Express Vol.7 No.1 2021.03 pp.1-4

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

The urban aerial mobility (UAM) system, such as drone taxi or air taxi, is one of future on-demand transportation networks. Among them, electric vertical takeoff and landing (eVTOL) is one of UAM systems that is for identifying the locations of passengers, flying to the positions where the passengers are located, loading the passengers, and delivering the passengers to their destinations. In this paper, we propose a distributed deep reinforcement learning where the agents are formulated as eVTOL vehicles that can compute the optimal passenger transportation routes under the consideration of passenger behaviors, collisions among eVTOL, and eVTOL battery status.

2

Multi-agent reinforcement learning for a distributed multi-channel access game

Zhongyang Li, Yu Zhao, 이주현

[NRF 연계] 한국통신학회 ICT Express Vol.11 No.5 2025.10 pp.863-869

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this work, we model multi-user distributed channel access as a game with channels and users, and propose the Multi-Agent Thompson Sampling (MA-TS) algorithm. It uses Bayes’ theorem to dynamically optimize action selection. This optimization aims to maximize throughput. We derive the algorithm’s computational complexity as . Simulations show that MA-TS converges to a pure strategy Nash equilibrium (PNE) and outperforms existing methods in average throughput.

3

Deep-reinforcement-learning-based range-adaptive distributed power control for cellular-V2X

Wooyeol Yang, Han-Shin Jo

[NRF 연계] 한국통신학회 ICT Express Vol.9 No.4 2023.08 pp.648-655

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

A distributed congestion control must be adaptable to varying target communication ranges as cellular V2X (C-V2X) is evolving to support flexible coverage suitable for various service scenarios. This study proposes range-adaptive distributed power control (Ra-DPC) based on deep reinforcement learning (DRL) with the Monte Carlo policy gradient algorithm. A key finding is that the agents learn Ra-DPC more effectively when the cumulative interference power of the subchannels is adopted as the state of the DRL model, rather than the channel busy ratio. The proposed Ra-DPC algorithm performs better in energy efficiency and packet delivery ratio than the existing technologies.

4

Multi-agent reinforcement learning based optimal energy sensing threshold control in distributed cognitive radio networks with directional antenna

PHAMTHITHUHIEN, 노원종, Cho Sungrae

[NRF 연계] 한국통신학회 ICT Express Vol.10 No.3 2024.06 pp.472-478

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In CRNs, it is crucial to develop an efficient and reliable spectrum detector that consistently provides accurate information about the channel state. In this work, we investigate a CSS in a fully-distributed environment where all secondary users (SUs) are equipped with directional antennas and make decisions based solely on their local knowledge without information sharing between SUs. First, we establish a stochastic sequential optimization problem, which is an NP-hard, that maximizes the SU’s detection accuracy by the dynamic and optimal control of the energy sensing/detection threshold. It can enable SUs to select an available channel and sector without causing interference to the primary network. To address it in a distributed environment, the problem is transformed into a decentralized partially observed Markov decision process (Dec-POMDP) problem. Second, in order to determine the best control for the Dec-POMDP in a practical environment without any prior knowledge of state?action transition probabilities, we develop a multi-agent deep deterministic policy gradient (MADDPG)-based algorithm, which is referred to as MA-DCSS. This algorithm adopts the centralized training and decentralized execution (CTDE) architecture. Third, we analyzed its computational complexity and showed the proposed approach’s scalability by the polynomial computational complexity, in terms of the number of channels, sectors, and SUs. Lastly, the simulation confirms that the proposed scheme provides enhanced performance in terms of convergence speed, accurate detection, and false alarm probabilities when it is compared to baseline algorithms.

5

Distributed Resource Allocation in IoT Networks Based on Deep Reinforcement Learning

Xueping Han

[Kisti 연계] 한국정보처리학회 Journal of information processing systems Vol.22 No.1 2026 pp.88-99

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

Currently, most Internet of Things (IoT) resource-allocation solutions are based on centralized management, rendering it difficult to successfully establish diverse and dynamic IoT networks. Consequently, the fields of reinforcement learning and distributed computing require additional technological advancements. We propose an innovative approach for slot scheduling in IoT networks. This approach focuses on the utilization of distributed resource blocks. The purpose of this work is to demonstrate, via simulations, the influence of distributed slot assignment on the signal-to-interference ratio (SIR) and the probability of accidents occurring. The results of this research suggest that the proposed approach, in which each device in the IoT network is provided with an appropriate slot that possesses acceptable SIR levels, was successful. As the process of distributed slot allocation progresses, it is beneficial to build network convergence by utilizing the learning capabilities of each device. This was accomplished using a distributed slot-allocation process. Therefore, the existing bandwidth can be utilized more efficiently.

6

This paper introduces a decision-making framework for offloading tasks in home network environments, utilizing Distributed Reinforcement Learning (DRL). The proposed scheme optimizes energy efficiency while maintaining system reliability within a lightweight edge computing setup. Effective resource management has become crucial with the increasing prevalence of intelligent devices. Conventional methods, including on-device processing and offloading to edge or cloud systems, need help to balance energy conservation, response time, and dependability. To tackle these issues, we propose a DRL-based scheme that allows flexible and enhanced decision-making regarding offloading. Simulation results demonstrate that the proposed method outperforms the baseline approaches in reducing energy consumption and latency while maintaining a higher success rate. These findings highlight the potential of the proposed scheme for efficient resource management in home networks and broader IoT environments.

7

Reinforcement learning multi-agent using unsupervised learning in a distributed cloud environment

Seo-Yeon Gu, Seok-Jae Moon, Byung-Joon Park

국제인공지능학회(구 한국인터넷방송통신학회) International Journal of Internet, Broadcasting and Communication Vol.14 No.2 2022.05 pp.192-198

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

Companies are building and utilizing their own data analysis systems according to business characteristics in the distributed cloud. However, as businesses and data types become more complex and diverse, the demand for more efficient analytics has increased. In response to these demands, in this paper, we propose an unsupervised learning-based data analysis agent to which reinforcement learning is applied for effective data analysis. The proposal agent consists of reinforcement learning processing manager and unsupervised learning manager modules. These two modules configure an agent with k-means clustering on multiple nodes and then perform distributed training on multiple data sets. This enables data analysis in a relatively short time compared to conventional systems that perform analysis of large-scale data in one batch.

8

This study proposes a distributed reinforcement learning system that incorporates the Safe Proper Time (SPT) protocol to address latency issues in cloud-based environments. The system is architected to operate efficiently under limited computational resources, making it suitable for small-scale enterprises. By combining physicallevel technologies such as InfiniBand and TOE with software-level optimization, the SPT protocol enables low-latency, high-throughput data transmission across distributed nodes. Experimental results show that the proposed system reduces response failure rates and achieves faster processing times compared to centralized models. Furthermore, a comparative analysis demonstrates that the system offers competitive advantages over existing machine learning platforms in terms of deployment flexibility and initial cost efficiency. This research contributes to the field by presenting a scalable and resource-efficient approach to distributed reinforcement learning. Future work will focus on enhancing the security and stability of data transmission in SPT-based systems.

9

심층강화학습 기반 분산형 전력 시스템에서의 수요와 공급 예측을 통한 전력 거래시스템 KCI 등재

이승우, 선준호, 김수현, 김진영

국제인공지능학회(구 한국인터넷방송통신학회) 한국인터넷방송통신학회 논문지 제21권 제6호 2021.12 pp.163-171

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

본 논문은 분산형 전력 시스템에서 심층강화학습 기반의 전력 생산 환경 및 수요와 공급을 예측하며 자원 할당 알고리즘을 적용해 전력거래 시스템 연구의 최적화된 결과를 보여준다. 전력 거래시스템에 있어서 기존의 중앙집중식 전력 시스템에서 분산형 전력 시스템으로의 패러다임 변화에 맞추어 전력거래에 있어서 공동의 이익을 추구하며 장기적 인 거래의 효율을 증가시키는 전력 거래시스템의 구축을 목표로 한다. 심층강화학습의 현실적인 에너지 모델과 환경을 만들고 학습을 시키기 위해 날씨와 매달의 패턴을 분석하여 데이터를 생성하며 시뮬레이션을 진행하는 데 있어서 가우시 안 잡음을 추가해 에너지 시장 모델을 구축하였다. 모의실험 결과 제안된 전력 거래시스템은 서로 협조적이며 공동의 이익을 추구하며 장기적으로 이익을 증가시킨 것을 확인하였다.

In this paper, the energy transaction system was optimized by applying a resource allocation algorithm and deep reinforcement learning in the distributed power system. The power demand and supply environment were predicted by deep reinforcement learning. We propose a system that pursues common interests in power trading and increases the efficiency of long-term power transactions in the paradigm shift from conventional centralized to distributed power systems in the power trading system. For a realistic energy simulation model and environment, we construct the energy market by learning weather and monthly patterns adding Gaussian noise. In simulation results, we confirm that the proposed power trading systems are cooperative with each other, seek common interests, and increase profits in the prolonged energy transaction.

10

A Study of Collaborative and Distributed Multi-agent Path-planning using Reinforcement Learning

Kim, Min-Suk

[Kisti 연계] 한국컴퓨터정보학회 Journal of the Korea society of computer and information Vol.26 No.3 2021 pp.9-17

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

동적 시스템 환경에서 지능형 협업 자율 시스템을 위한 기계학습 기반의 다양한 방법들이 연구 및 개발되고 있다. 본 연구에서는 분산 노드 기반 컴퓨팅 방식의 자율형 다중 에이전트 경로 탐색 방법을 제안하고 있으며, 지능형 학습을 통한 시스템 최적화를 위해 강화학습 방법을 적용하여 다양한 실험을 진행하였다. 강화학습 기반의 다중 에이전트 시스템은 에이전트의 연속된 행동에 따른 누적 보상을 평가하고 이를 학습하여 정책을 개선하는 지능형 최적화 기계학습 방법이다. 본 연구에서 제안한 방법은 강화학습 기반 다중 에이전트 최적화 경로 탐색 성능을 높이기 위해 학습 초기 경로 탐색 방법을 개선한 최적화 방법을 제안하고 있다. 또한, 분산된 다중 목표를 구성하여 에이전트간 정보 공유를 이용한 학습 최적화를 시도하였으며, 비동기식 에이전트 경로 탐색 기능을 추가하여 실제 분산 환경 시스템에서 일어날 수 있는 다양한 문제점 및 한계점에 대한 솔루션을 제안하고자 한다.

In this paper, an autonomous multi-agent path planning using reinforcement learning for monitoring of infrastructures and resources in a computationally distributed system was proposed. Reinforcement-learning-based multi-agent exploratory system in a distributed node enable to evaluate a cumulative reward every action and to provide the optimized knowledge for next available action repeatedly by learning process according to a learning policy. Here, the proposed methods were presented by (a) approach of dynamics-based motion constraints multi-agent path-planning to reduce smaller agent steps toward the given destination(goal), where these agents are able to geographically explore on the environment with initial random-trials versus optimal-trials, (b) approach using agent sub-goal selection to provide more efficient agent exploration(path-planning) to reach the final destination(goal), and (c) approach of reinforcement learning schemes by using the proposed autonomous and asynchronous triggering of agent exploratory phases.

11

멀티 엣지 네트워크에서 협업 엣지컴퓨팅을 위한 심층강화학습 기반 분산 오프로딩 정책 연구

정준호, 윤주상

[Kisti 연계] 한국산업정보학회 한국산업정보학회논문지 Vol.29 No.5 2024 pp.11-19

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

유저 디바이스의 태스크 오프로딩을 처리하는 위치가 클라우드에서 엣지로 이동함에 따라, 이를 효과적으로 처리하기 위한 자원 관리 기술의 필요성이 대두되고 있다. 많은 연구에서 강화 학습을 통해 이 문제를 해결하고자 하였으나, 실제 오프로딩 태스크에서 발생하는 오버헤드를 충분히 반영하지 못하였다. 본 논문에서는 태스크의 오버헤드를 고려한 강화학습 기반 분산 오프로딩 정책 생성 기법을 제안하고, 이를 검증하기 위한 시뮬레이션 환경을 구축하였다. 실험을 통해 해당 기법이 엣지의 큐 대기시간을 감소시켜 기존 기법 대비 최대 46.3%의 성능 향상이 있음을 보였다.

As task offloading from user devices transitions from the cloud to the edge, the demand for efficient resource management techniques has emerged. While numerous studies have employed reinforcement learning to address this challenge, many fail to adequately consider the overhead associated with real-world offloading tasks. This paper proposes a reinforcement learning-based distributed offloading policy generation method that incorporates task overhead. A simulation environment is constructed to validate the proposed approach. Experimental results demonstrate that the proposed method reduces edge queueing time, achieving up to 46.3% performance improvement over existing approaches.

12

MANET에서 종단간 통신지연 최소화를 위한 심층 강화학습 기반 분산 라우팅 알고리즘

Choi, Yeong-Jun, Seo, Ju-Sung, Hong, Jun-Pyo

[Kisti 연계] 한국정보통신학회 한국정보통신학회논문지 Vol.25 No.9 2021 pp.1267-1270

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In this paper, we propose a distributed routing algorithm for mobile ad hoc networks (MANET) where mobile devices can be utilized as relays for communication between remote source-destination nodes. The objective of the proposed algorithm is to minimize the end-to-end communication delay caused by transmission failure with deep channel fading. In each hop, the node needs to select the next relaying node by considering a tradeoff relationship between the link stability and forward link distance. Based on such feature, we formulate the problem with partially observable Markov decision process (MDP) and apply deep reinforcement learning to derive effective routing strategy for the formulated MDP. Simulation results show that the proposed algorithm outperforms other baseline schemes in terms of the average end-to-end delay.

13

강화학습과 분산유전알고리즘을 이용한 자율이동로봇군의 행동학습 및 진화

이동욱, 심귀보

[Kisti 연계] 대한전자공학회 電子工學會論文誌. Journal of the Korean Institute of Telematics and Electronics S. S Vol.s34 No.8 1997 pp.56-64

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

In distributed autonomous robotic systems, each robot must behaves by itself according to the its states and environements, and if necessary, must cooperates with other orbots in order to carray out a given task. Therefore it is essential that each robot has both learning and evolution ability to adapt the dynamic environments. In this paper, the new learning and evolution method based on reinforement learning having delayed reward ability and distributed genectic algorithms is proposed for behavior learning and evolution of collective autonomous mobile robots. Reinforement learning having delayed reward is still useful even though when there is no immediate reward. And by distributed genetic algorithm exchanging the chromosome acquired under different environments by communication each robot can improve its behavior ability. Specially, in order to improve the perfodrmance of evolution, selective crossover using the characteristic of reinforcement learning is adopted in this paper, we verify the effectiveness of the proposed method by applying it to cooperative search problem.

14

실시간 차량 밀도에 대응하는 심층강화학습 기반 C-V2X 분산혼잡제어

전병철, 양우열, 조한신

[Kisti 연계] 한국전기전자학회 Journal of IKEEE Vol.27 No.4 2023 pp.379-385

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

분산혼잡제어는 높은 밀도의 차량 네트워크에서 채널 혼잡을 완화하고, 통신 성능을 개선하는 기술이다. 기존 분산혼잡제어 기술은 quality of service(QoS) 요구사항을 고려하지 않은 채 채널 혼잡을 줄이는 방향으로 동작한다. 이러한 분산혼잡제어 알고리즘 설계는 과도한 DCC 동작으로 인하여 다른 QoS를 저하시킬 수 있다. 이와 같은 문제를 해결하기 위해 심층강화학습 기반 QoS 적응형 DCC 알고리즘을 제안한다. 시뮬레이션은 준 실환경 시뮬레이터를 기반으로 동적인 차량 밀도를 생성하여 평가하였으며, 시뮬레이션 결과 기존 DCC 알고리즘 보다 목표 QoS에 더 근접한 결과를 확인하였다.

Distributed congestion control (DCC) is a technology that mitigates channel congestion and improves communication performance in high-density vehicular networks. Traditional DCC techniques operate to reduce channel congestion without considering quality of service (QoS) requirements. Such design of DCC algorithms can lead to excessive DCC actions, potentially degrading other aspects of QoS. To address this issue, we propose a deep reinforcement learning-based QoS-adaptive DCC algorithm. The simulation was conducted using a quasi-real environment simulator, generating dynamic vehicular densities for evaluation. The simulation results indicate that our proposed DCC algorithm achieves results closer to the targeted QoS compared to existing DCC algorithms.

15

에이전트 기반 분산형 제조실행 환경에서의 강화학습 기반 동적 디스패칭

김란, 신문수

[Kisti 연계] 한국시뮬레이션학회 한국시뮬레이션학회논문지 Vol.35 No.2 2026 pp.41-51

※ 협약을 통해 무료로 제공되는 자료로, 원문이용 방식은 연계기관의 정책을 따르고 있습니다.

원문보기

본 연구는 분산형 제조실행 환경에서의 강화학습 기반 동적 디스패칭 메커니즘을 제안한다. 이는 다품종 혼류 생산 환경의 불확실성에 실시간으로 대응하며, 중앙집중형 제어 방식의 높은 연산 부담 해소와 분산형 접근에서 나타나는 협상 과정의 복잡성 및 행동 공간의 기하급수적 증가 문제를 해결하는 것에 목적을 둔다. 이를 위해 중앙 시스템의 개입을 배제하고, 개별 공정 설비 에이전트가 병렬적으로 의사결정을 수행하는 자율적 제어 구조를 개발한다. 특히, 에이전트가 처리할 작업을 직접 할당하는 대신 사전 정의된 휴리스틱 디스패칭 규칙 중 하나를 선택하도록 문제를 재정식화함으로써, 행동 공간을 제한하고 학습의 안정성을 높인다. 또한 연속형 상태변수를 효과적으로 처리하기 위해 심층 Q-네트워크(deep Q-network, DQN)를 도입하며, 셋업 시간 및 납기 지연 최소화라는 다목적 최적화를 위한 보상 함수를 적용한다. 시뮬레이션 결과, 제안 모델은 관측되지 않은 생산 시나리오 하에서도 기존의 정적 디스패칭 규칙 및 이산형 Q-learning 모델 대비 통계적으로 유의미한 성능 향상을 보임으로써 강건한 일반화 성능을 확인할 수 있다.

This study proposes a reinforcement learning-based dynamic dispatching mechanism within a distributed manufacturing execution environment. The objective is to respond in real time to the uncertainties of high-mix, multi-product production environments, while simultaneously alleviating the heavy computational burden of centralized control and resolving the negotiation complexity and exponential action-space expansion found in decentralized approaches. To achieve this, we develop an autonomous control architecture where individual processing equipment agents make parallel decisions without the intervention of a centralized system. In particular, rather than directly assigning jobs to be processed, the problem is reformulated to let agents select one from predefined heuristic dispatching rules, thereby constraining the action space and enhancing training stability. Furthermore, a deep Q-network (DQN) is introduced to effectively handle continuous state variables, and a multi-objective reward function is applied to minimize both setup time and total tardiness. Simulation results demonstrate that the proposed model achieves statistically significant improvements in performance compared to conventional static dispatching rules and discrete Q-learning models, even under unobserved production scenarios, thereby validating its robust generalization capability.

 
페이지 저장