TRUST-AWARE ENVIRONMENT-SENSITIVE DEEP REINFORCEMENT LEARNING FOR ADAPTIVE CLUSTER HEAD SELECTION IN MOBILE AD HOC NETWORKS
DOI:
https://doi.org/10.29121/shodhai.v3.i2.2026.109Keywords:
MANET, Clustering, Deep Reinforcement Learning, Cluster Head Selection, Trust Management, Network Security, Energy EfficiencyAbstract
Mobile Ad hoc Networks (MANETs) suffer from performance degradation as node count and mobility increase, since every node must independently manage routing without fixed infrastructure. Clustering is a well-established remedy: nodes are grouped under an elected cluster head (CH) that coordinates intra- and inter-cluster communication, which reduces routing-table size and control overhead. Most existing CH-selection schemes score candidate nodes using a fixed combination of residual energy, node degree, mobility and distance, but they treat trust as a secondary or static attribute and re-evaluate it too infrequently to catch adaptive insider attacks such as wormhole or grey-hole routing. This paper extends a reward-optimized Deep Q-Learning (RoDQL) clustering formulation by folding a continuously-updated, feedback-driven trust score into the state representation and reward function of the learning agent, alongside energy, mobility, node-degree and distance. A guard-node subsystem audits packet-forwarding behaviour, updates each node's trust value on every routing cycle, and feeds this into the agent so that a node's suitability as CH is re-assessed as its behaviour changes, not only at cluster-formation time. The resulting framework, referred to here as Trust-Augmented RoDQL (TA-RoDQL), is presented together with its Markov Decision Process formulation, state/action/reward design, and a two-stage algorithm for cluster-head election and trust-triggered re-election. The expected benefits — more balanced cluster sizes, lower energy depletion at CHs, and faster detection and isolation of malicious nodes — are discussed against the underlying reward-shaping mechanics, and a simulation protocol is laid out so the approach can be validated empirically against classical Q-learning, DQN-only and static-weight clustering baselines.
References
Alowish, M., Shiraishi, Y., Takano, Y., Mohri, M., and Morii, M. (2020). Stabilized clustering enabled V2V communication in an NDN-SDVN environment for content retrieval. IEEE Access, 8, 135138–135151.
Desai, A. M., and Jhaveri, R. H. (2018). Secure routing in mobile ad hoc networks: A predictive approach. International Journal of Information Technology.
Feng, Y., Teng, G. F., Wang, A. X., and Yao, Y. M. (2007). Chaotic inertia weight in particle swarm optimization. In 2007 Second International Conference on Innovative Computing, Information and Control (ICICIC 2007) (p. 475). IEEE.
Haridas, S. (2026). Environment-aware rewards optimized deep-Q-learning for cluster head selection in MANET. International Journal of Advanced Research in Computer and Communication Engineering (IJARCCE), 15(8), 363–368.
Haridas, S., and Rama Prasath, A. (2020). Enhancement of network lifetime in MANET: Improved particle swarm optimization for delayless and secured geographic routing. Journal of Advanced Research in Dynamical and Control Systems, 12(07-Special Issue).
Janani, V. S., and Manikandan, M. S. (2020). Hexagonal clustered trust based distributed group key agreement scheme in mobile ad hoc networks. Wireless Personal Communications, 1–20.
Janani, V. S., and Manikandan, M. S. K. (2017). Enhanced security using cluster based certificate management and ECC-CRT key agreement schemes in mobile ad hoc networks. Wireless Personal Communications, 97(4), 6131–6150.
Kanagasundaram, H., and A, K. (2018). EIMO-ESOLSR: Energy efficient and security-based model for OLSR routing protocol in mobile ad-hoc network. IET Communications.
Kavitha, G. (2018). QoS improvement based security enhancement for link activity monitoring service in mobile ad hoc network. Cluster Computing.
Kavitha, T., and Muthaiah, R. (2018). A lightweight FFT based enciphering system for extending the lifetime of mobile ad hoc networks. Cluster Computing.
Khan, B. U. I., Anwar, F., Olanrewaju, R. F., Pampori, B. R., and Mir, R. N. (2020). A game theory-based strategic approach to ensure reliable data transmission with optimized network operations in futuristic mobile adhoc networks. IEEE Access, 1–1.
Krishnan, R. S., Julie, E. G., Robinson, Y. H., Kumar, R., Son, L., Tuan, T. A., and Long, H. V. (2020). Modified zone based intrusion detection system for security enhancement in mobile ad hoc networks. Wireless Networks, 26, 1275–1289.
Kumar, K. V., Jayasankar, T., Eswaramoorthy, V., and Nivedhitha, V. (2020). SDARP: Security based data aware routing protocol for ad hoc sensor networks. International Journal of Intelligent Networks, 1, 36–42.
Mehrkanoon, S., Alzate, C., Mall, R., Langone, R., and Suykens, J. A. (2014). Multiclass semisupervised learning based upon kernel spectral clustering. IEEE Transactions on Neural Networks and Learning Systems, 26(4), 720–733.
Ochola, E. O., Mejaele, L., Eloff, M. M., and Poll, J. (2017). MANET reactive routing protocols node mobility variation effect in analysing the impact of black hole attack.
Robinson, Y. H., and Julie, E. G. (2019). MTPKM: Multipart trust based public key management technique to reduce security vulnerability in mobile ad-hoc networks. Wireless Personal Communications.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Dr. Haridas S.

This work is licensed under a Creative Commons Attribution 4.0 International License.
With the licence CC-BY, authors retain the copyright, allowing anyone to download, reuse, re-print, modify, distribute, and/or copy their contribution. The work must be properly attributed to its author.
It is not necessary to ask for further permission from the author or journal board.
This journal provides immediate open access to its content on the principle that making research freely available to the public supports a greater global exchange of knowledge.



















