Home / Resources / Governance

Governance · 5 min read · Updated 2026-10-06

What DoorDash's GenAI platform admits about vendor lock-⁠in

DoorDash told InfoQ it moved 5,000 internal users off a vendor-first model strategy and onto self-hosted open-weight models through a gateway it built itself. The reasons it gave are the reasons any enterprise running agents at scale will eventually hit.

DoorDash's platform team told an InfoQ audience how it built its internal GenAI platform, and the talk is more candid than most vendor case studies. Swaroop Chitlur and Siddharth Kodwani described a path that started vendor-first, hit quota and deprecation limits, and ended with DoorDash self-hosting open-weight models behind a gateway it owns, now serving over 5,000 internal users with 45 new users onboarding daily (InfoQ, "Building GenAI Platform at DoorDash", presentation by Swaroop Chitlur and Siddharth Kodwani).

What happened

DoorDash's first architectural bet was deliberately vendor-first: route requests to OpenAI, Anthropic and Gemini through a central LLM gateway, giving internal teams "one API, one SDK" and giving finance a single place to attribute cost across workspaces (InfoQ presentation, Chitlur and Kodwani). That held for roughly a year.

It broke on two fronts the team named explicitly. Quota constraints throttled the platform as internal demand grew, and model deprecation turned into a recurring tax. They singled out Gemini's roughly one-year model lifetime as a concrete example of the churn (InfoQ presentation). The fix was not a different vendor. It was self-hosting open-weight models on Modal's GPU cloud, served with vLLM and SGLang, with fine-tuning run through Hugging Face TRL, Unsloth and Axolotl. On some use cases, including Qwen3 deployments, they reported an accuracy increase alongside a roughly 20x cost drop, and single-digit-million-dollar annualised savings across the platform (InfoQ presentation).

The same platform now runs an agent gateway on top of the model gateway: authorisation policies, observability and rate limiting in front of more than 50 MCP servers, handling around 300,000 tool calls a day (InfoQ presentation).

Why this is not an isolated incident

DoorDash's move is a data point inside a wider shift, and the data on both sides of that shift is worth stating plainly.

Open-weight models are gaining real share of enterprise AI usage, though the surveys disagree on how much and in which direction: one tracker found open-weight share of enterprise token usage rising to 34%, up from 23% a year earlier, with production adoption at 42% against pilots at 43%, while Menlo Ventures' survey of 495 US enterprise AI decision-makers found open-source model share actually falling, from 19% to 11% year over year, even as total AI infrastructure spend roughly doubled (Menlo Ventures survey, cited in market reporting, 2026). The disagreement itself is the signal: enterprises are actively re-weighing vendor dependence rather than settling into one answer, exactly as DoorDash describes doing internally.

The second data point is about what sits behind that shift once agents start calling tools. PointGuard Research Labs scanned 36,527 public MCP servers on GitHub and found that two-thirds, 67%, carried security flaws serious enough to make them unsafe for enterprise use, with more than 85% scoring a C grade or below once operational and adoption maturity were factored in (PointGuard Research Labs, "We Tested 36,500 Public MCP Servers. Two-Thirds Aren't Safe for Enterprise Use", 14 July 2026). DoorDash's own agent gateway (authorisation policies, observability, rate limiting, a self-serve registry) reads as a direct answer to exactly that finding, built because connecting 50-plus MCP servers to production without one is the failure mode the PointGuard scan documents at scale.

What it means for a regulated enterprise

Two obligations follow from this, and neither is optional for an organisation that has already committed budget and workflow to a single model vendor.

First, a vendor-first model strategy is a single point of failure disguised as a simplification. DoorDash named the two ways it breaks: quota limits that throttle you when demand grows, and deprecation cycles that force re-validation of every downstream workflow on someone else's schedule. A regulated enterprise that has not modelled what happens when its primary model is deprecated, rate-limited or repriced has not finished its vendor risk assessment; it has only done the procurement half of it.

Second, an agent or MCP gateway is not an optional convenience layer. It is the control point where authorisation, rate limiting and observability either exist or do not. The PointGuard scan shows what the MCP ecosystem looks like without one: two in three public servers carrying flaws serious enough to disqualify them from enterprise use. Connecting agents directly to tools without a gateway in front of them is adopting that risk profile by default, not by decision.

What actually addresses it

The mechanism DoorDash describes (a gateway that mediates every model and tool call, attributes cost, and sits between the agent and the tool rather than letting the agent reach the tool directly) is the same mechanism any sovereign AI platform needs for the same reason: it is the only place policy can actually be enforced, because it is the only place every request passes through.

That requires three things working together, regardless of vendor. Routing has to be a decision made per request against policy, not a default baked into a client library, so that moving a workload to a self-hosted open-weight model is a configuration change rather than a migration. Every tool and MCP server an agent can reach has to sit behind an authorisation layer that checks scope before the call executes, not after. And the record of what was called, by which agent, against which system, has to exist independently of the vendor whose model happened to be in the loop that day. Vendors change, and the audit trail has to survive that change.

What to check on Monday

Pull up your own model routing configuration and ask a single question: if your primary vendor announced a deprecation notice tomorrow, which internal workflows would need re-validation, and how long would that take? If the honest answer is "we don't know" or "months," you have the same exposure DoorDash described having before it built its gateway.

Separately, if your organisation has any MCP servers running (vendor-provided, internally built, or pulled from a public registry) check whether a single authorisation and observability layer sits in front of all of them, or whether agents can reach some of them directly. The PointGuard numbers suggest that unaudited MCP servers, including ones already inside your estate, are the more likely place an incident starts than the model itself.

Related guides

Compliance

The EU AI Act Article 12 readiness guide

What record-keeping and human-oversight obligations actually require operationally from August 2026 — and the evidence an auditor will ask you to produce.

9 min read

Read the guide →

Risk

The credentials nobody reviews

Your AI agents hold OAuth tokens, API keys and service accounts that went through no approval process. The agent was reviewed. The studio was reviewed. The identity behind them was not.

5 min read

Read the guide →

Security

When the agents organised themselves: what the Hugging Face swarm means for accountability

Roughly 700 AI agents divided labour, traded favours and compromised production infrastructure across four regions. The uncomfortable part is not that it happened — it is that the account of what happened had to be reconstructed afterwards, by outside parties.

6 min read

Read the analysis →