---
title: "Platform engineering for AI workloads | Dhruv Doshi"
description: "AI workloads are becoming ordinary production workloads, but platforms built for stateless microservices handle them poorly. GPU scheduling, model serving, evaluation infrastructure,…"
canonical: https://doshidhruv.com/notes/platform-engineering-for-ai-workloads/
author: Dhruv Doshi
---

[Skip to content](#main-content)

[Dhruv Doshi](/)

[Work](/projects)[Experience](/resume)[Notes](/notes)[Guides](/guides)[Research](/research)[Investing](/investing)[About](/about)[Contact](/contact)[Search /](/search)dark mode

[All notes](/notes)

[Platform architecture](/topics/platform-architecture)

# Platform engineering for AI workloads

2026-10-01 · 4 minute read

AI workloads are becoming ordinary production workloads, but platforms built for stateless microservices handle them poorly. GPU scheduling, model serving, evaluation infrastructure, and experiment tracking each need platform capabilities that most internal platforms do not yet offer. Platform teams that treat AI as someone else's problem will watch shadow infrastructure grow.

## Serve models as a platform capability

The most repeated undifferentiated work in AI adoption is every team standing up its own model serving: choosing inference runtimes, configuring autoscaling for bursty token generation, managing API keys, and implementing fallbacks. A platform model gateway consolidates this: approved model endpoints, authentication, rate limiting, request logging, cost attribution, and provider failover in one place.

The gateway is also a governance enforcement point. Allowed models per risk tier, data-handling rules, prompt and completion logging policy, and emergency disablement all attach naturally here. Teams get a simple interface; the organization gets consistent control.

Support both hosted APIs and self-hosted models. Some workloads need data to stay in specific regions or networks, and some need latency or cost profiles that public APIs cannot meet. The platform should offer model serving as an internal service with the same operational standards as any other — versioned, monitored, and boring.

## Make evaluation infrastructure shared

Every AI team needs the same evaluation machinery: dataset management with versioning, rubric-based graders with calibration tracking, regression suites tied to releases, and production sampling joined back to traces. Built per-team, this is duplicated effort with inconsistent quality. Built once as a platform capability, it raises the bar for everyone.

The evaluation platform should integrate with delivery: evaluation gates in CI, quality dashboards per use case, and automatic regression runs on model, prompt, or data changes. When evaluation is easy, teams do it. When it requires building infrastructure first, they skip it and hope.

## Treat data pipelines as AI inputs

AI systems are only as good as the data pipelines feeding them. Retrieval corpora need ingestion with metadata preservation, freshness guarantees, and deletion propagation. Fine-tuning needs governed datasets with provenance. Evaluation needs production-derived data with consent handling.

Platform data capabilities — ingestion frameworks, metadata standards, lineage tracking, access policy enforcement — are AI capabilities. The teams building retrieval systems should not be reinventing document parsing and permission filtering; the platform should provide them as composable building blocks.

## Plan for accelerators explicitly

GPU capacity planning differs from CPU planning: it is expensive, lumpy, and contention-sensitive. Platform teams need to offer GPU access as a managed resource with fair scheduling, quota management, and cost visibility — whether on cloud instances, reserved capacity, or on-premises clusters.

For most organizations, the right answer is a mix: burst to cloud for training and experiments, reserved or committed capacity for steady-state inference. The platform abstracts this choice where possible and makes the cost visible where it matters. What it must not do is let every team negotiate its own GPU access.

## Operate AI systems with the same rigor

AI services need SLOs, on-call rotations, incident processes, and capacity planning like any production system — plus the AI-specific additions: quality signals alongside availability signals, probe sets for silent change, and rollback procedures that cover prompts, indexes, and model versions as well as code.

The platform team's job is to make the safe path the easy path: golden paths for model serving, evaluation, and deployment that embed the controls by default. AI workloads are not exotic anymore. They are production workloads with additional dimensions, and the platform should treat them that way.

## Do not build a second platform

The failure mode to avoid is the AI platform team building a parallel universe: its own CI, its own deployment system, its own monitoring, its own identity — duplicating everything the existing platform provides, slightly worse and permanently understaffed. AI workloads need new capabilities, not a new platform.

Integrate instead. Model serving rides on the existing deployment and networking foundations. Evaluation gates plug into the existing CI. Cost attribution extends the existing FinOps model. GPU scheduling extends the existing compute abstraction. Each integration point is a place where AI-specific needs meet platform-wide standards — and where the AI team benefits from years of operational hardening it did not have to build.

This requires the existing platform team to treat AI as a first-class workload rather than an exception, which means real roadmap commitment, not just tolerance. The organizational design that works is usually one platform team with AI workload expertise embedded, not two teams negotiating a boundary. The boundary will move constantly as AI capabilities mature; a single team moves with it, while two teams file tickets at each other.

The test is simple: can a product team ship an AI feature using the same golden paths, the same deployment process, and the same on-call model as everything else? If yes, the platform has absorbed AI. If the AI work happens in a separate system with separate rituals, you have built the shadow infrastructure you were trying to prevent.

## Continue reading

- [API versioning strategy that survives contact with reality](/notes/api-versioning-strategy-that-survives-contact-with-reality) · Note
- ["Feature flag governance: flags as inventory, not litter"](/notes/feature-flag-governance-flags-as-inventory-not-litter) · Note
- [GitOps at scale: what works past the demo](/notes/gitops-at-scale-what-works-past-the-demo) · Note
- [The internal developer platform adoption playbook](/notes/the-internal-developer-platform-adoption-playbook) · Note

**Dhruv Doshi** · Toronto, Canada · [work@doshidhruv.com](mailto:work@doshidhruv.com)

[Resume](/resume)[Notes](/notes)[Guides](/guides)[Topics](/topics)[Search](/search)[Research](/research)[Investing](/investing)[LinkedIn](https://www.linkedin.com/in/dhruvdoshi25071999)[GitHub](https://github.com/DhruvDoshi)

[Sitemap](/sitemap.xml)[RSS](/feed.xml)[LLMs](/llms.txt)

© 2026 Dhruv Doshi
