This study investigates the applicability and effectiveness of a PPO(Proximal Policy Optimization)- based idle vehicle relocation strategy for the Demand Responsive Transport (DRT) service “Tabara” in Gijang-gun, Busan, using a SUMO-based simulation environment calibrated with actual operational data. Based on approximately three months of service records, three representative demand scenarios were constructed: weekday peak, weekday offpeak, and weekend/holiday. The PPO-based policy was evaluated against a no-relocation baseline, a rule-based heuristic, and a DQN-based policy under identical episode conditions. The results indicate that the PPO-based relocation policy reduced average pickup waiting time by approximately 2.9–6.2% and cumulative unserved calls by approximately 2.5–5.3% relative to the baseline, with the largest improvement observed in the weekday offpeak scenario. However, its additional advantage over the heuristic and DQN policies was limited. These findings suggest that, under the operating conditions considered in this study, the main contribution of PPO lies not in clear absolute superiority, but in its potential as a learning-based adaptive relocation strategy for addressing dynamic demand–supply imbalances in DRT operations.
한국어
본 연구는 부산광역시 기장군 수요응답형 교통(Demand Responsive Transport, DRT) ‘타바라’ 를 대상으로 실제 운행자료를 반영한 SUMO(Simulation of Urban MObility) 기반 시뮬레이션 환 경을 구축하고, 근접 정책 최적화(Proximal Policy Optimization, PPO) 기반 유휴차량 재배치 전 략의 적용 가능성과 효과를 분석하였다. 실제 3개월 운행자료를 바탕으로 평일 첨두, 평일 비 첨두, 주말·공휴일 시나리오를 구성하였으며, PPO 정책을 재배치 미적용 정책, 규칙기반 휴리 스틱, DQN 기반 정책과 비교하였다. 분석 결과, PPO 기반 재배치 정책은 재배치 미적용 대비 평균 픽업 대기시간을 약 2.9~6.2%, 누적 미처리 호출 수를 약 2.5~5.3% 감소시키는 경향을 보 였으며, 특히 평일 비첨두 시나리오에서 상대적으로 큰 개선 효과가 나타났다. 반면 휴리스틱 및 DQN과 비교한 추가적 우월성은 제한적으로 나타나, 본 연구 환경에서는 PPO의 절대적 성 능 우위보다 학습기반 적응형 재배치 정책으로서의 적용 가능성이 더 크게 확인되었다.
목차
요약 ABSTRACT Ⅰ. 서론 Ⅱ. 선행연구 고찰 1. 국내 DRT 연구 동향과 운영전략 연구의 확장 2. 유휴차량 재배치 및 운영 최적화 전략 관련 선행연구 Ⅲ. 자료 수집 및 시뮬레이션 환경 구축 1. 분석 대상 2. 자료 수집 및 전처리 3. 수요 및 교통 특성 분석 4. 시뮬레이션 환경 구축 5. 강화학습 기반 재배치 모형 설계 Ⅳ. 분석 결과 1. 평가 지표 및 비교 방법 2. PPO 학습 결과 3. 시나리오별 성능 비교 및 결과 분석 Ⅴ. 결론 REFERENCES