A learned control policy is only useful in the field if it never violates transformer ratings, voltage limits, or battery state-of-charge bounds. I develop deep reinforcement learning methods that embed such hard constraints directly into training, together with techniques for reducing very large action spaces, coordinating multiple agents, and transferring policies across sites. We are also beginning to explore what large language models can contribute to user behavior modeling and load forecasting.