Nutanix has introduced new oversight tools for enterprise AI agents after witnessing customers exhaust monthly AI budgets in a single week, according to the company's latest platform update. The release of Nutanix Enterprise AI 2.8 addresses growing governance and cost-management challenges as businesses deploy AI agents closer to production systems and data. The update is part of Nutanix's broader strategy to enable organizations to run AI workloads alongside existing virtual machines and containers within a unified management environment.

The new platform delivers granular controls over agent permissions, spending, and infrastructure access. An MCP gateway allows companies to define which applications, tools, and data each agent can reach, while a separate MCP Server enables agents to interact with Nutanix infrastructure without unrestricted access. Organizations can now monitor token consumption and establish spending caps at the team, user, or individual agent level. The release also expands Private Inference capabilities with LoRA fine-tuning for models under 8 billion parameters, multi-GPU inference, batch inference, and speculative decoding, which the company says can boost token-generation speed by up to 2.5 times depending on model and infrastructure configuration. The forthcoming NKP 2.19 extends Kubernetes management across both virtualized and bare-metal environments, with NKP Metal automating operating system, firmware, and container deployment, while an AI applications catalog includes curated deployments of Kubeflow, Milvus, and Slurm.

Greg Cornely from Nutanix noted the company has observed instances where customers allocated a budget meant to last a month that was consumed within a week. The platform now enables Private Inference to route appropriate workloads to privately deployed open-weight models, where customers pay for underlying infrastructure rather than per-token charges. Nutanix is also making Service Provider Central generally available, offering a multitenant control plane for infrastructure, application, cloud-native, and AI services, while the new Powered by Nutanix: Verified Services program provides partners with onboarding resources, service delivery materials, and badging for building services around hybrid cloud infrastructure, Kubernetes, and VM migration.

The update responds to organizations moving AI agents into production environments where uncontrolled access and spending pose operational risks. Token-based pricing models can produce unpredictable costs when agents generate more queries than anticipated, making budget controls essential as deployments scale. The shift toward private inference reflects growing interest in cost structures that trade per-use fees for infrastructure investment, particularly as customers seek more predictable expenses. Cornely also emphasized that companies migrating away from VMware have an opportunity to modernize infrastructure rather than simply replicating existing setups. "You don't just replace the VMware with the same old thing," he said, suggesting that the migration process itself creates a window for architectural change rather than a like-for-like swap.

The platform positions governance and cost transparency as prerequisites for production AI deployments, with token-level tracking and permission boundaries designed to prevent runaway spending and unauthorized data access. By integrating agent controls with existing VM and container management, Nutanix aims to help organizations extend current operational practices to AI workloads without building separate governance frameworks. The agent spending limits and access restrictions provide a foundation for scaling deployments while maintaining financial and security guardrails. The real test will be whether enterprises treat agent governance as a strategic discipline or a box-checking exercise, and whether cost controls prompt smarter workload design or simply shift budget conversations without changing underlying efficiency.