The platform team's quarterly review has a new slide. Eighteen months ago there was one AI agent in production. Now there are twenty-three. The customer service team built theirs in LangGraph. The data team uses CrewAI. A group that works closely with a cloud provider chose Google ADK; another went with AWS Strands. The innovation lab has something in the OpenAI Agents SDK, and a .NET-heavy business unit is on Microsoft Agent Framework.
Each team has also, separately, solved the same production problems: retries, checkpointing, crash recovery, secrets, tracing, cost tracking. Some solved them well. Some didn't solve them at all. And the security team has just asked the platform team a question nobody can answer: which agents can call which internal APIs, and how would we know if one misbehaved?
If you run a platform engineering, developer productivity or AI platform team at a mid-sized or large organization, this article is about how to bring order to that sprawl without becoming the team that forbids everyone's favourite framework.
Why "pick one framework" doesn't work
The instinctive response is standardization by decree: choose a framework, write a guide, and require everyone to migrate. It rarely works, for three reasons.
The frameworks are genuinely different. Graph-based orchestration, role-based crews and model-provider SDKs suit different problems. Forcing them into one shape costs teams productivity.
The landscape is moving fast. The framework that looks like the obvious winner today may not be in eighteen months. Betting the company's agent strategy on one is a big risk.
Migration has no business value. Rewriting a working agent in a different framework delivers nothing to customers. Product teams will push back, and they'll be right.
The more useful question isn't "which framework?" It's "which layer should the platform own?"
Separate the agent logic from the agent runtime
Platform teams solved a similar problem with microservices a decade ago. They didn't standardize on one web framework. They standardized the runtime: containers, Kubernetes, service mesh, observability, CI/CD. Teams kept their languages and frameworks; the platform made everything run, scale and fail the same way.
The same split works for AI agents:
- Agent logic (prompts, tools, graphs, crews, reasoning patterns) belongs to product teams and their chosen frameworks.
- Agent runtime (durability, recovery, identity, communication, observability, deployment, audit) belongs to the platform.
The goal is a runtime layer that any supported framework can plug into with minimal code changes, so every agent inherits the same production guarantees regardless of how it was written.
What the platform's agent runtime should provide
Based on what tends to go wrong in production, a shared agent runtime should cover six areas.
1. Durable execution. Every model call and tool call is recorded, and crashed runs resume automatically from the last completed step. This is the single biggest reliability win, and it's the one most teams implement inconsistently on their own.
2. Workload identity and zero-trust access. Each agent gets its own cryptographic identity, and all calls to tools, MCP servers and other agents go over mutual TLS. Access policies define which agents can call which services. This answers the security team's question directly.
3. Standard communication. Agents that collaborate do so through a common mechanism, such as event-driven pub/sub or direct invocation through the platform, rather than ad-hoc HTTP calls with hard-coded URLs.
4. Unified observability. Every agent emits the same traces and metrics in the same format, so the platform team can answer "what is each agent doing, how often does it fail, and what does it cost?" from one dashboard.
5. Human-in-the-loop primitives. Approvals and escalations work the same way across agents, with durable waits that survive restarts.
6. Consistent deployment and audit. Agents deploy through the same pipelines, into the same environments, and produce the same kind of execution record for audits.
The open-source foundation
Building this runtime from scratch is a large project. Fortunately, much of it maps onto existing open-source infrastructure. Dapr, a CNCF graduated project, is a popular choice because it already provides most of the six areas as runtime building blocks: a durable workflow engine, SPIFFE-based workload identity with mTLS, pub/sub messaging, service invocation, and OpenTelemetry-based tracing, all delivered through a sidecar that works the same way regardless of language.
The key enabler for agent standardization is that durable workflow integrations now exist for most popular agent frameworks. Instead of rewriting agents, teams wrap them: a LangGraph graph is handed to a workflow-backed runner, a CrewAI agent's tool calls become workflow activities, and so on. The agent code stays the same; the runtime underneath changes.
A paved road, not a mandate
The most successful platform teams present the agent runtime as a "paved road": the easiest way to get an agent into production, rather than a compliance hurdle. That usually means:
- A project template per supported framework, preconfigured with durable execution, identity and tracing.
- Clear documentation of the few constraints the runtime imposes, such as keeping I/O inside tool calls so recovery works correctly.
- Self-service onboarding, so a team can register a new agent, get an identity and deploy it without filing a ticket.
- Visible benefits. Show teams their agents' recovery events, latency and failure rates on day one. Nothing sells a platform like seeing a crash that didn't matter.
Teams that stay off the paved road can still ship, but they take on the reliability, security and audit requirements themselves, and they find out quickly that the road is the cheaper option.
Build, buy or both
Platform teams with deep Kubernetes and distributed-systems expertise can assemble this runtime from open-source Dapr. Many organizations, though, would rather spend their platform engineers' time on internal developer experience than on running workflow infrastructure at scale.
Commercial options sit on top of the same foundation. Diagrid, for example, provides integrations for a wide range of agent frameworks (including LangGraph, CrewAI, Google ADK, AWS Strands, the OpenAI Agents SDK and Microsoft Agent Framework) through its Catalyst platform, along with hosted or self-hosted deployment, agent identity and signed execution history. A reasonable evaluation approach is to take your three most-used frameworks, wrap one real agent from each, and compare the effort and results against your build estimate.
Measuring success
How do you know the standardization is working? Track a few metrics over two or three quarters:
- Coverage: the share of production agents running on the shared runtime
- Recovery: how many runs recovered automatically from crashes or restarts, which shows value the teams would otherwise not see
- Time to production for a new agent, from first commit to live traffic
- Security posture: agents using shared credentials versus agents with their own identity
- Audit readiness: how long it takes to reconstruct a specific agent run on request
Conclusion
AI agent sprawl is the microservices sprawl of this decade, and the solution looks familiar: let product teams choose their tools, and give them a platform that makes everything they build reliable, secure and observable by default. Standardizing the runtime instead of the framework lets the organization move as fast as the agent ecosystem does, without twenty-three teams each solving production reliability on their own. For platform teams weighing how to provide that runtime, a durable execution layer for AI agents built on open standards is a solid place to start.
