AI-driven operations fail when the estate has no reliable source of truth. Models cannot reason over missing CMDB data, unnamed subscriptions, or networks that exist only in a Visio from 2019. Preparation is inventory, telemetry and access design.

Inventory: what runs, where, and who owns it. Even a maintained spreadsheet plus tags is better than five conflicting CMDBs. Telemetry: metrics and logs that arrive, with resource IDs that match the inventory. Identity: workload identities for automation, not shared admin passwords in a vault nobody rotates.

On one estate, an ops assistant asked what changed before an outage found nothing, because change tickets were free text, subscription names were inconsistent and half the resources had no owner tag. Before any model could help, the estate needed an inventory with owners and a change log a machine could read.

Change records matter. If production changes happen outside a recorded path, an agent will propose work that collides with a human who already did it. Put a thin, simple change process in place before you add a planner.

Environment access should be least privilege and break-glassed. An operations agent in production with Owner is a future incident. Start in read-only investigation, then gated write actions, then a small set of well-tested remediations.

This work looks like infrastructure hygiene because it is. Organisations that skip it buy a demo. Organisations that do it can later use models on the messy edges without gambling the estate. Do the inventory and access work first, before buying an “AI operations” product.

Want AI running on infrastructure you can trust? AI Infrastructure & Automation

All resources