Cyber-physical power systems face coordinated false data injection attack (FDI) attacks that exploit the cyber measurement pipeline and knowledge of the power-flow Jacobian to bypass bad-data detection. Existing moving target defense (MTD) approaches address each layer in isolation, leaving a critical cross-layer seam that adaptive adversaries can exploit. This paper presents a unified, agentic artificial intelligence (AI)-guided reinforcement learning (RL) framework that jointly optimizes physical MTD actions such as topology switching, flexible alternating current transmission system (FACTS) configuration, and distributed energy resource (DER) control with measurementcentric cyber-agility actions, including dynamic watermarking, adaptive scheduling, and estimator reconfiguration. A large language model (LLM) orchestrator, grounded via retrieval augmented generation (RAG) over domain threat knowledge, guides a hierarchical multi-agent reinforcement learning (MARL) network operating over a high-fidelity digital twin with crosslayer interdependency modeling. Preliminary results on the IEEE 14-bus system demonstrate improved FDI detection over physical-only and cyber-only baselines while maintaining frequency deviation within $\pm 0.05 \text{Hz}$, with full evaluation planned across IEEE 14, 39, 57, and 118-bus benchmarks.