Splunk’s May 2026 Observability Cloud release pushes AI observability into a more practical phase. The headline addition is AI Infrastructure Monitoring, with built-in dashboards for components including agentgateway LLM proxy, Amazon Bedrock AgentCore Gateway, Kong AI Gateway Proxy, llm-d, NVIDIA Dynamo and vLLM. Splunk also introduced OpenTelemetry Fleet Management, a phased-rollout feature that provides APIs for inventorying and configuring agents and collectors at scale.
That matters because enterprise AI stacks are becoming production infrastructure, not experiments. LLM gateways, model-serving runtimes, vector databases, GPU layers and agent frameworks now sit directly in the path of customer service, developer workflows, search, analytics and internal automation. Splunk says AI Infrastructure Monitoring supports telemetry across LLM services, model-serving platforms, language frameworks, vector databases, models, infrastructure services and microservices, using the Splunk Distribution of the OpenTelemetry Collector.
This extends the argument we raised in our earlier piece on Splunk’s agentic observability push: the more observability data flows into AI assistants, IDEs, chatbots and internal LLMs, the more enterprises must think about access, identity, auditability and governance.
What IT teams should instrument first
For platform teams, the immediate priority is not to monitor “AI” as a vague category. It is to instrument the choke points first: the LLM gateway, the model-serving runtime, GPU utilisation, request latency, error rates, token consumption, queue depth, cost signals and downstream dependencies. Those are the places where small failures become expensive outages.
OpenTelemetry Fleet Management is equally important.
AI observability will not scale if every collector, agent and integration is configured manually. Centralised fleet visibility gives teams a better chance of knowing what is deployed, what is misconfigured and where telemetry coverage is missing.
There is also a cautionary note.
Splunk’s own May release notes warn that enabling GenAI content capture can cause performance issues when captured input and output attributes exceed backend limits.
That is the lesson for enterprises racing into AI infrastructure: visibility is essential, but indiscriminate capture is not observability. It is risk with a dashboard.