Abstract
<title>Abstract</title> <p>Vehicular networks must support stringent reliability and latency requirements for high mobility, rapidly varying channels, and heterogeneous communication options. In this study, we investigated dynamic mode selection in a clustered vehicular network where each vehicle (or cluster entity) can transmit using either direct sidelink / device-to-device (D2D) or a cellular-assisted Cellular Vehicle to Everything (C-V2X) mode. Mode selection is formulated as a multi-agent sequential decision problem in which agents observe local network conditions (e.g., topology/ relative distances, channel/ interference indicators, and traffic or queue states). The agents then choose transmission mode and additional control variables such as power, to optimize vehicular Quality of Service (QoS). A multi- agent parameter sharing Proximal Policy Optimization (PPO) approach is proposed. This policy enables cluster-head agents to adapt to time varying interference and mobility while satisfying reliability requirements. The reward design explicitly reflects system-level Key Performance Indicators (KPIs) latency, reliability, spectral efficiency, and transmit power to discourage solutions that achieve reliability through high power. We evaluate the proposed method in a highway simulation setting with realistic channel and interference modeling and compare against intuitive baselines such as distance-threshold mode selection and fixed-mode policies. Results show that PPO improves reliability by about 4 to 12 percentage points compared with the strongest non-learning baseline, reduces latency by at least 69.4\% compared with the random policy, and lowers transmit power by up to 8.39 dB relative to the 18 dBm fixed-power baselines. This demonstrates the benefit of learning-based adaptive mode selection in clustered vehicular environments.</p>