"Web performance budgets: make speed a release gate"
· 10 minute read
Every web team agrees performance matters and almost none treat it as a requirement. Features ship, bundles grow, images accumulate, third-party scripts multiply, and performance degrades one merged PR at a time — each individually reasonable, collectively fatal. The decay is never anyone's decision, which is exactly why it needs a mechanism: performance budgets that turn "we care about speed" into "this release does not ship if it exceeds the budget." A budget is a contract between the team and its future self, enforced by tooling, with a defined process for spending more. This note covers what to budget, how to enforce it, and how to keep the system honest over time.
Why budgets instead of goals
Goals ("make the site fast") fail because they're evaluated never and enforced never. Budgets work for the same reason financial budgets work: they convert a vague aspiration into a number, attach the number to a decision point (the release), and force an explicit trade when the number is exceeded. "We want LCP under 2.5s" is a wish. "PRs that push the product page's JS over 180KB fail CI" is governance.
Budgets are SLOs for the frontend. The SRE framing is deliberate and useful: define the metric, set the target, measure in production, alert on violation, and treat sustained violation as an incident-class problem. Frontend teams that already think in SLOs for backends should extend the same discipline to the user-facing surface — it's the same reliability story, measured where the user experiences it.
The budget conversation forces prioritization. Without budgets, every feature's performance cost is invisible, so nothing is ever too expensive. With budgets, adding a 90KB chat widget means finding 90KB elsewhere or getting explicit approval to raise the budget. That conversation — "is this widget worth 200ms of LCP?" — is the entire point. Most performance waste isn't malicious; it's unpriced.
Verdict: if performance isn't gating a release somewhere in your pipeline, you don't have a performance strategy — you have a performance hope.
What to budget: the metrics that matter
Core Web Vitals as the headline budget. Largest Contentful Paint (loading), Interaction to Next Paint (responsiveness), Cumulative Layout Shift (visual stability). These are what search ranking uses and what users feel. Set budgets at the page-template level — the product page, the checkout flow, the landing page — because a site-wide average hides the pages that matter. The checkout's budget should be tighter than the blog's.
JavaScript weight: the most predictive budget. JS bytes correlate brutally with interactivity on real devices. Budget by route: total JS transferred (compressed) and, more importantly, JS executed on load. A useful starting shape: under 170KB compressed JS for critical flows on mobile, with per-route budgets scaled by the page's job. Track both the bundle and the third-party share separately — third-party scripts are the fastest-growing line item in most budgets and need their own accounting.
Image and media weight. Images are typically the largest bytes on the page. Budget total image weight per page template and enforce format discipline (modern formats, responsive sizes, lazy loading below the fold). A single unoptimized hero image can blow an otherwise disciplined budget — the budget is what catches it before the marketing team ships a 4MB PNG.
Request counts and third-party limits. Cap the number of third-party scripts and the domains they load from. Each third party is a reliability and privacy dependency as well as a performance cost: their outage becomes your outage, their latency becomes your latency. Budget them like vendor dependencies, because that's what they are — with a review process for adding new ones.
Custom metrics for your product. Beyond the standard vitals: time to first product interaction, search-results render time, checkout step transition time. The vitals measure the web; custom metrics measure your business. Instrument the two or three user journeys that drive revenue and budget those specifically.
Enforcement: budgets that bite
CI gates on bundle size. Every PR gets a bundle-size report with the delta against the budget; PRs exceeding the budget fail the check. Tools like bundlesize, size-limit, or Lighthouse CI in the pipeline make this mechanical. The key design decision: fail vs warn. Warn-only budgets decay within a quarter — everyone clicks through. Fail-by-default with an explicit override process (a label, an approver, a recorded justification) keeps the gate real while allowing genuine exceptions.
# Conceptual CI budget gate (Lighthouse CI style)
ci:
collect:
url:
- https://staging.example.com/product/demo
- https://staging.example.com/checkout
assert:
assertions:
largest-contentful-paint: ["error", { maxNumericValue: 2500 }]
interactive: ["error", { maxNumericValue: 3500 }]
cumulative-layout-shift: ["error", { maxNumericValue: 0.1 }]
resource-summary:script:size: ["error", { maxNumericValue: 180000 }]
Lighthouse CI on critical paths. Run lab tests against staging for the budgeted routes on every PR. Lab data is noisy — run multiple passes, compare against baselines, fail on regression beyond a threshold rather than absolute numbers. Lab tests catch regressions before merge; they're the unit tests of performance.
RUM as the source of truth. Lab tests gate PRs; real-user monitoring judges reality. Instrument Core Web Vitals and custom metrics in production, segmented by device class, network, and geography — the p75 on a mid-tier Android in a secondary market is the number that matters, not the p50 on office wifi. Alert on budget violation in RUM the way you'd alert on an SLO breach: page the owner, not a mailing list.
Budgets in code review culture. The CI gate catches the numbers; review culture catches the judgment calls. "This adds 40KB — what's it buying us?" should be a normal review comment. Maintain a short performance review checklist for PRs touching critical paths: new dependencies justified, images optimized, no synchronous third-party scripts, lazy loading where applicable. Culture is the enforcement layer tooling can't reach.
Setting the numbers: baselines, not aspirations
Start from measured reality. Run the vitals and bundle analysis on production today, per page template. The initial budget is current performance plus a small improvement margin — not an aspirational number pulled from a conference talk. An unattainable budget gets ignored on day one; a budget set at "current minus 10%" creates a ratchet that compounds.
Differentiate by page value and device. The checkout flow gets the tightest budget; the rarely-visited settings page gets a looser one. Mobile budgets are tighter than desktop — the devices are slower and the networks worse, and mobile is usually the majority of traffic. One global budget is simple and wrong.
Revisit on a schedule. Budgets set once rot: the product changes, traffic shifts to new pages, frameworks evolve. Review budgets quarterly alongside the performance data. Tighten where there's headroom, loosen where the budget is blocking value with no user-visible gain — a budget that's never adjusted is either ignored or cargo-culted into irrelevance.
The "performance bank" pattern. Give teams a small discretionary allowance — they can spend up to X over budget without approval, drawn from a shared pool that refills quarterly. This acknowledges reality (sometimes the feature is worth the bytes) while keeping the spending visible and bounded. What gets measured gets managed; what gets a shared pool gets debated, which is the point.
Third parties: the budget's biggest leak
Inventory everything. Most teams can't list their third-party scripts from memory — tag managers silently accumulate pixels, A/B tools, chat widgets, survey popups. Audit quarterly: every script needs a named owner, a measured cost (bytes, execution time, impact on vitals), and a business justification. Scripts without owners get removed. This audit alone typically recovers more performance than months of code optimization.
Load third parties like untrusted code, because they are. Async or deferred loading, never render-blocking; resource hints (preconnect) only for the critical ones; facade patterns for heavy embeds (load the chat widget's facade, fetch the real thing on interaction). A third-party script that blocks rendering is a vendor with deploy rights to your user experience — treat the integration accordingly.
Contract with vendors. Where third parties are business-necessary (payments, analytics), put performance terms in the relationship: maximum script weight, execution time budgets, no document.write-era patterns, notification before major changes. You can't CI-gate a vendor's deploy, but you can monitor their cost in RUM and escalate with data.
The image and font pipeline: budgets need infrastructure
Budgets fail when compliance depends on individual diligence. Images are the proof: telling every engineer to "optimize images" produces nothing; building an image pipeline that optimizes automatically produces compliance. The pipeline — enforced at build or at the CDN layer — should resize to breakpoints, convert to modern formats, strip metadata, and lazy-load below the fold by default. Then the budget measures the pipeline's output, not human discipline. Any image bypassing the pipeline (CMS uploads, marketing embeds) gets flagged in CI. The same applies to fonts: subset to the character sets you need, self-host instead of render-blocking third-party font CSS, use font-display: swap, and budget the number of font families and weights. A page loading six weights of two families has a typography problem wearing a performance costume.
Video is the new image-weight crisis. Autoplay hero videos and inline product clips dwarf every other asset. The pipeline rules: never autoplay above the fold without a poster fallback and user-gesture gating, stream adaptive bitrates rather than shipping one MP4, and count video bytes in the media budget with a separate, explicit line item — because a single unreviewed video can exceed the entire page's JS budget tenfold. If marketing needs video, the pipeline makes it cheap; without the pipeline, video is where budgets go to die.
Keeping it honest over time
Performance reviews as a ritual. Monthly or quarterly: RUM trends per budgeted route, budget adherence, the biggest regressions and their causes, third-party cost changes. Present it like any other reliability review — to engineering leadership, with the same seriousness as uptime. Performance that isn't reviewed regularly isn't managed.
Regressions get root-caused. When RUM shows a vitals regression, treat it like an incident: what changed, when, which deploy or vendor update, and what prevents recurrence. "The new hero video" is a finding; "hero videos bypass the image pipeline" is the systemic fix. The goal is fewer regressions over time, which requires learning from each one.
Celebrate the wins visibly. Performance work is invisible when it succeeds — the page just feels fine. Publish the improvements: "checkout LCP down 400ms, conversion up" connects the discipline to business outcomes and buys goodwill for the next budget fight. The teams that sustain performance culture are the ones that can show it pays, quarter after quarter.
Anti-patterns
Budgets nobody enforces. The wiki page with numbers nobody checks. Worse than no budget — it teaches the team that performance commitments are decorative.
Lab-only measurement. Gating on Lighthouse while RUM shows users suffering. Lab tests are the smoke alarm; RUM is checking whether the house is actually on fire. You need both.
Optimizing the average. p50 improvements that leave p75/p90 untouched. The users having the worst experience are the ones most affected by your performance — and often the ones you most need to convert.
One budget for the whole site. The marketing blog and the checkout flow have nothing in common. Undifferentiated budgets are either too lax to matter or too strict to follow.
Performance as a project. The quarterly "performance sprint" that recovers six months of decay, followed by six more months of decay. Performance is a property maintained by gates, not a project completed by sprints.
Closing
Performance budgets work because they move speed from aspiration to mechanism: numbers attached to release gates, enforced in CI, judged by real-user data, reviewed like reliability. Budget the vitals, the JavaScript, the images, and the third parties — per route, per device class, starting from measured reality. Make the gate fail by default with an explicit override path, audit the third-party inventory ruthlessly, build pipelines that make compliance automatic rather than heroic, and review the trends like an SLO. Speed held by a budget compounds; speed held by good intentions decays. Pick the mechanism.