---
title: "&quot;Platform team topologies: organizing for leverage, not tickets&quot; | Dhruv Doshi"
description: "Every platform team starts with the same noble intent: build the paved road so product teams can move fast. And nearly every platform team ends up in the same place: a ticket queue.…"
canonical: https://doshidhruv.com/notes/platform-team-topologies-organizing-for-leverage-not-tickets/
author: Dhruv Doshi
---

[Skip to content](#main-content)

[Dhruv Doshi](/)

[Work](/projects)[Experience](/resume)[Notes](/notes)[Guides](/guides)[Research](/research)[Investing](/investing)[About](/about)[Contact](/contact)[Search /](/search)dark mode

[All notes](/notes)

[Engineering leadership](/topics/engineering-leadership)

# "Platform team topologies: organizing for leverage, not tickets"

2026-10-01 · 13 minute read

Every platform team starts with the same noble intent: build the paved road so product teams can move fast. And nearly every platform team ends up in the same place: a ticket queue. "Please provision a database." "Please add us to the deploy pipeline." "Please debug our Helm chart." The team becomes internal IT with a Kubernetes cluster — reactive, overloaded, resented. The difference between platform teams that create leverage and ones that drown in tickets is almost never technology. It's topology: how the team is organized, what it owns, and — most importantly — what it refuses to do.

## The ticket trap: how good platform teams die

**The trap starts with being helpful.** A new platform team wants adoption, so it says yes to everything: custom Terraform modules per team, bespoke pipeline configurations, one-off infrastructure requests. Each yes feels like service. In aggregate, the yeses become a support burden that consumes the team's entire capacity — and because every request was "reasonable," nobody notices the team has stopped building the platform and started operating everyone's infrastructure by hand.

**Tickets scale linearly; platforms must scale sublinearly.** A ticket-driven platform team grows its backlog in direct proportion to the number of product teams. Ten teams means ten times the tickets. The math never works, because the platform team can't hire proportionally — and shouldn't. A platform's value proposition is that N product teams get leverage from one platform team's work. The moment the platform team's effort scales with N, the leverage is gone and you just have a centralized ops team with better branding.

**The symptom is a roadmap of other people's requests.** Look at the platform team's backlog: if it's dominated by tickets filed by product teams, the topology is wrong. A healthy platform backlog is dominated by the platform team's own initiatives — capabilities they're building because their product research says product teams need them, not because someone filed a JIRA ticket. The ticket queue should be the exception path, not the planning input.

Verdict: if your platform team measures itself by ticket throughput, it has already lost. Throughput of tickets is a measure of demand for manual work — the thing the platform exists to eliminate.

## The four topologies, honestly assessed

Team Topologies gives us the vocabulary — stream-aligned, enabling, complicated-subsystem, platform — but the labels matter less than the interaction modes. Here's how platform organizations actually arrange themselves, with the trade-offs nobody puts in the conference talk.

### The centralized platform team

One team owns the platform: CI/CD, Kubernetes clusters, observability, cloud accounts, developer tooling. Product teams consume it.

**When it works:** small to mid-size engineering organizations (roughly under 150 engineers), where one team can genuinely know all its consumers. The platform is coherent — one way to deploy, one observability stack, one set of conventions — and coherence is the whole point. Centralization also concentrates scarce expertise: you don't need every team to understand cluster autoscaling.

**When it breaks:** the team becomes the bottleneck for everything. Every new need waits in the platform queue. Product teams route around it — shadow infrastructure in personal cloud accounts, CI pipelines built on side projects — and the platform team discovers it's governing an infrastructure landscape that no longer matches reality. The breaking point is usually around the time the platform team needs a ticket triage process: that's the organization telling you the topology no longer fits.

**Rule of thumb:** centralize while the platform team can name every consuming team and their primary contact. When they can't, the topology is already failing.

### The federated platform model

A small central team owns the platform's core — the control plane, the contracts, the paved road's foundation — while platform *capabilities* are built by federated teams embedded in or aligned to product areas. Think: central team owns the Kubernetes platform and the deployment API; the data platform team (aligned to data-heavy products) builds the streaming and warehouse tooling on top of it.

**When it works:** larger organizations where one team can't hold the whole platform in its head. Federation keeps the core coherent while letting domain-specific platform needs (ML infrastructure, data pipelines, edge computing) evolve at their own pace with people who understand the domain. The central team's job shifts from building everything to maintaining the contracts everything builds on.

**When it breaks:** the federation becomes fragmentation. Each federated team builds its own deployment story, its own observability conventions, its own IAM model — and the "platform" is five platforms wearing a trench coat. This happens when the central team's contracts are weak: if the core doesn't provide genuinely useful primitives with real enforcement, federated teams will build around it rather than on top of it.

**The contract is the product.** In a federated model, the central team's primary output isn't infrastructure — it's the interfaces: the deployment API, the service scaffolding contract, the observability schema, the security baseline. These contracts must be *better* than what teams would build themselves, or federation decays into everyone building their own. This is the hardest version of platform-as-product, because your "customers" are other platform builders with strong opinions.

### The embedded platform engineer

Platform engineers sit inside product teams — not as visitors, but as members — building platform capabilities from within and feeding patterns back to the central platform.

**When it works:** as a transitional topology. Embedding is how you discover what the platform should be: the embedded engineer feels the product team's pain directly, builds the tooling that relieves it, and generalizes the pattern back into the platform. It's also how you seed platform thinking in teams that don't have it. Time-box it — six to twelve months — with an explicit goal of extracting a reusable capability.

**When it breaks:** the embedded engineer becomes the team's DevOps person. They spend their days on the team's tickets, their platform work evaporates, and the "embedding" is just staff augmentation with extra steps. This is the default outcome unless the embedding has a written charter: what capability they're building, what they're explicitly *not* doing (the team's on-call, the team's tickets), and when the embedding ends.

**Rule of thumb:** embedding without a charter and an end date is just reassignment. Write both down before the engineer moves desks.

### Platform as a product team (the meta-topology)

This isn't a separate topology — it's the operating model that makes any of the above work. The platform team treats product teams as customers: it does user research (actually talks to engineers about their pain), maintains a roadmap driven by that research, measures adoption and satisfaction, and markets its capabilities. The platform has a product manager — or the tech lead does the PM work explicitly — because "build it and they will come" has never worked for internal tools.

**The product discipline is what kills the ticket queue.** A product-minded platform team doesn't take feature requests as tickets; it takes them as *input to roadmapping*. "Team X needs a database" becomes a research question: how many teams need databases provisioned? What's the common pattern? Can we build self-service provisioning with guardrails instead of doing it by hand? The ticket is the symptom; the product work is the cure.

## Interaction modes: the actual lever

Topology diagrams are static; what matters is how teams interact day to day. Three modes, in increasing order of leverage:

**Ticket/queue mode (lowest leverage).** Product team files a request; platform team does the work. Necessary for genuinely exceptional cases — a novel compliance requirement, a production incident in shared infrastructure — but every ticket in this mode is a failure of self-service. Track the ratio: if more than 20% of platform team effort goes here, the platform isn't a platform, it's a service desk.

**Self-service mode (the goal).** Product teams do it themselves through platform-provided interfaces: scaffold a service from a template, provision a database through an API, deploy through the standard pipeline. The platform team's work is building and maintaining these interfaces. This is where leverage lives — one platform engineer's work on the provisioning API serves every team that provisions.

**Collaboration mode (for the hard stuff).** For complex, novel work — a new service mesh migration, a multi-region architecture — platform and product engineers work together temporarily, with the explicit goal of the product team being self-sufficient afterward. Collaboration without the self-sufficiency goal becomes permanent embedding (see above).

```
  leverage ▲
           │   collaboration ──► self-service
           │   (temporary,          (the goal:
           │    goal: independence)  teams serve themselves)
           │
           │   ticket/queue
           │   (exception path only)
           └─────────────────────────────────► platform maturity

```

**The migration path is always queue → collaboration → self-service.** When a new need appears, it starts as a ticket (someone needs something the platform doesn't provide). If it recurs, it becomes a collaboration (build it together once). If it's a pattern, it becomes self-service (productize it). Platform teams that skip straight from ticket to self-service build abstractions for needs they don't understand; teams that never progress past tickets drown.

Verdict: audit your platform team's week by interaction mode. The percentages tell you the topology's health more honestly than any architecture diagram.

## What the platform team owns (and refuses)

**Own the paved road, not every road.** The platform provides one blessed path for each common need: one way to deploy a service, one way to provision a database, one way to observe an application. The paved road is opinionated — that's the point. Teams can leave the paved road, but they own everything about the wilderness: their own pipelines, their own runbooks, their own 3 AM pages. The platform team's support boundary ends at the pavement's edge, and that boundary must be explicit and enforced.

**Own the contracts, not the implementations.** The platform defines what a service *must* provide — health endpoints, structured logging, OpenTelemetry instrumentation, resource declarations — and provides libraries and templates that make compliance easy. It does not write every team's Helm charts. Contracts scale; hand-holding doesn't.

**Refuse to be the only people who understand production.** If the platform team is the only team that can debug a deployment failure, the topology has failed regardless of what the diagram says. Every self-service interface must come with documentation, runbooks, and error messages that let product engineers resolve their own issues. "Contact the platform team" in an error message is a design defect.

**Refuse custom snowflakes.** The request will come: "we need a slightly different pipeline because our service is special." Sometimes it's true — genuinely novel requirements exist. Usually it's preference disguised as requirement. The platform team's job is to distinguish them, and the default answer to customization requests is no, with a clear appeals path (the collaboration mode) for the ones that survive scrutiny. Every snowflake the platform accepts becomes permanent maintenance burden.

Rule of thumb: the platform team should be able to describe its scope in one paragraph. If the description needs bullet points for all the exceptions, the scope has already dissolved.

## Staffing the platform team

**Platform engineering is a senior discipline.** Building interfaces that hundreds of engineers depend on — getting the abstractions right, handling the edge cases, maintaining backward compatibility — requires engineers who've seen what goes wrong. A platform team staffed mostly with junior engineers will build junior abstractions: leaky, undocumented, and abandoned when the author moves on. This isn't about prestige; it's about the blast radius of bad platform decisions being organization-wide.

**Don't staff the ticket queue; staff its elimination.** When leadership asks "the platform team needs more people to handle tickets," the right answer is usually "the platform team needs people to eliminate the tickets." Headcount that goes to queue throughput entrenches the queue. Headcount that goes to self-service shrinks it. This is a hard conversation because the tickets are *urgent* and the self-service work is *important* — but that's exactly why it needs leadership air cover.

## Measuring platform teams (without ticket metrics)

If ticket throughput is out, what goes in? Measure the things that indicate leverage:

**Adoption of the paved road.** What percentage of services deploy through the standard pipeline? What percentage of new services start from the template? Adoption is the platform's revenue — it measures whether the product is good. Declining adoption means teams are routing around you; find out why before it becomes structural.

**Time-to-first-deploy for a new service.** From "I want a service" to "it's running in production" — this is the platform's core value proposition, and it should be measured continuously with a canary service. If it takes two weeks and three tickets, the platform is failing at its one job regardless of how sophisticated the underlying infrastructure is.

**Self-service ratio.** What percentage of platform interactions happen without human involvement from the platform team? Provisioning, deployments, rollbacks, scaling — each should trend toward self-service over time. This is the metric that directly measures progress out of the ticket trap.

**Product team DORA metrics, attributed.** The platform exists to make product teams faster and more reliable. Track deployment frequency, lead time, change failure rate, and recovery time across consuming teams — and be honest about attribution. The platform team doesn't own these numbers, but if they're not moving, the platform isn't delivering its value proposition.

**Platform reliability as a product SLA.** The platform is production infrastructure for every team. It gets SLOs like any production system: pipeline availability, provisioning latency, control-plane error rates. A platform team that doesn't hold itself to SLOs is asking product teams to build on a foundation it won't warrant.

Verdict: the platform team's dashboard should look like a product dashboard — adoption, time-to-value, reliability — not an IT service desk report.

## The reorganization test

Topologies aren't permanent. Every 12–18 months, run the test:

1. **Can the platform team name its top 5 initiatives, and are they platform-initiated?** If the roadmap is ticket-driven, the topology or the operating model needs fixing.
2. **What's the self-service ratio trend?** Flat or declining means the team is treading water.
3. **Are product teams routing around the platform?** Shadow infrastructure is a vote of no confidence — investigate the cause, not the symptom.
4. **Does the topology match the org size?** The centralized team that worked at 80 engineers is drowning at 300\. Topologies have a valid size range; outgrowing it without reorganizing is how you get the worst of both worlds.
5. **Is anyone's job "ticket triage"?** If yes, you don't have a platform team — you have a queue with engineers attached.

## Anti-patterns

**The platform team that doesn't run production.** A platform team that builds deployment tooling but has never been on-call for a real production incident builds tools for a fantasy. Platform engineers should carry operational responsibility for the platform itself — its control plane, its pipelines — so they feel the consequences of their abstractions.

**Conway's law in reverse.** Organizing the platform team around technologies (the "Kubernetes team," the "CI team," the "observability team") instead of around product-team workflows produces fragmented tools that don't compose. Organize around the developer journey: from code to production, as the consumer experiences it. "All teams must use the platform" works only if the platform is genuinely the best option. Mandating a bad platform breeds resentment and shadow infrastructure. Earn the mandate through quality first; the mandate then just accelerates what's already happening.

**Vanity platforms.** Building a service mesh, a developer portal, or an AI assistant because it's interesting — while CI takes 45 minutes and nobody can provision a database without a ticket. Platform roadmaps must be ordered by product-team pain, not by technology fashion.

## Closing

Platform team topology is an organizational design problem wearing an infrastructure costume. The centralized team, the federated model, the embedding — each is right for some size and wrong for others, and each decays into a ticket queue without the product discipline to prevent it. The test is always the same: is the team's effort scaling sublinearly with the number of teams it serves? If yes, you have a platform. If no, you have a queue — and no amount of Kubernetes will fix an organizational design problem.

## Continue reading

- ["Hiring bars for senior engineers: what the bar is really made of"](/notes/hiring-bars-for-senior-engineers-what-the-bar-is-really-made-of) · Note
- ["RFC culture that actually decides"](/notes/rfc-culture-that-actually-decides) · Note
- ["Staff-plus archetypes: the four paths that actually exist"](/notes/staff-plus-archetypes-the-four-paths-that-actually-exist) · Note

**Dhruv Doshi** · Toronto, Canada · [work@doshidhruv.com](mailto:work@doshidhruv.com)

[Resume](/resume)[Notes](/notes)[Guides](/guides)[Topics](/topics)[Search](/search)[Research](/research)[Investing](/investing)[LinkedIn](https://www.linkedin.com/in/dhruvdoshi25071999)[GitHub](https://github.com/DhruvDoshi)

[Sitemap](/sitemap.xml)[RSS](/feed.xml)[LLMs](/llms.txt)

© 2026 Dhruv Doshi
