पाठशाला Pathshala · उत्पाद Utpād, The product · Lesson 28 · Scale
Reliability, SLAs and the price of downtime
Every nine of uptime costs more engineering than the last. Choose targets by customer tier, measure them the way customers feel them, and budget the work that keeps the promise.
Pathshala, The Founder Library · 11 October 2026 · 7 min read

An enterprise contract arrives with a clause promising 99.99 per cent uptime, and the founder signs it because the competitor did. That clause allows about four minutes of downtime a month, and the company’s own payment gateway, SMS provider and cloud database cannot deliver it together.
This lesson sets out what the three reliability terms mean, how to choose a target for each customer tier, the arithmetic of nines and dependencies with a figure to test a promise before signing it, how to measure honestly, and how to budget the engineering that keeps the number.
Uptime is a product decision
More reliability is not always better. Google’s site reliability engineers put it bluntly in the chapter Embracing Risk: “100% is probably never the right reliability target.” Each additional nine costs more than the last, in redundant infrastructure, in on-call engineers and in slower releases, and past some point customers cannot tell. The same chapter notes that “a user on a 99% reliable smartphone cannot tell the difference between 99.99% and 99.999% service reliability.” For a product used over Indian mobile networks the device and the connection often fail before the service does.
So reliability belongs on the product roadmap beside features. The question is not how reliable the system can be made but how reliable each customer needs it to be, what that customer pays, and what each extra nine costs to deliver.
Three terms, three different promises
The Google SRE book’s chapter on Service Level Objectives gives the definitions most teams use. A service level indicator is “a carefully defined quantitative measure of some aspect of the level of service”, such as the share of payment requests that succeed within two seconds. A service level objective is a target for that indicator, set internally. A service level agreement is a contract with customers “that includes consequences of meeting (or missing) the SLOs they contain.”
The consequences are usually service credits. Amazon’s EC2 SLA is a common template: a region-level commitment of 99.99 per cent monthly uptime, with a credit of 10 per cent of the bill below that, 30 per cent below 99.0 per cent and 100 per cent below 95.0 per cent. Note the order. The internal objective should be tighter than the contract, so the team gets a warning before the company owes money. In India some downtime already has a regulated price: the RBI’s circular on failed transactions of 20 September 2019 requires banks to compensate customers ₹100 a day when a failed UPI or IMPS transfer is not reversed within the set time. A product inside a payment flow inherits that pressure from its bank partners.
Targets by customer tier
Few companies need one target for everyone. A workable pattern has three tiers. Self-serve customers on monthly plans get a published objective, often 99.5 per cent, and no credits. Mid-market customers on annual contracts get 99.9 per cent on the journeys they depend on, with modest credits. Enterprise customers may get 99.9 or 99.95 per cent with credits, named support response times and a status page, and pay for it in the contract price. Set the tier by what downtime costs the customer, not by what the sales team can be talked into: an hour down on payroll day is a crisis for an HR team and an annoyance on any other day.
Define availability per journey, not per server. Login, the core transaction and data export are separate indicators, and the contract should name which ones it covers. A promise of 99.9 per cent for the whole product is both harder to keep and less useful to the customer than 99.9 per cent for payroll processing and 99.5 per cent for reports.
The arithmetic of nines and dependencies
In a 30-day month of 43,200 minutes, 99 per cent allows 432 minutes of downtime, 99.9 per cent allows 43, and 99.99 per cent allows about four. A service that calls other services in series can be no more available than the product of all of them. Five components at 99.95 per cent each multiply to about 99.75 per cent, which is roughly 108 minutes a month. No amount of care in the company’s own code reaches 99.9 per cent on top of that chain without redundancy: a second payment gateway, a fallback SMS route, a replica database in another zone.

At the defaults, a 99.9 per cent target allows 43 minutes a month, but a service at 99.95 per cent calling four dependencies at 99.95 per cent should expect about 108. On a contract worth ₹15 lakh a month, a typical month would cost a 10 per cent credit. Remove two dependencies from the critical path, or raise each to 99.99 per cent with a fallback, and the chain meets the target. Choose 99.99 per cent and almost no chain of rented services gets there. Read the figure before the contract, not after the first outage.
Promise customers less than engineering aims for, and never promise more than your dependencies can deliver together.
Measuring honestly
Measure from where the customer sits. Server uptime says the machine was on; it does not say a payroll run completed. Use the share of real requests that succeed within a latency threshold, plus synthetic probes that run the critical journey every minute from outside the cloud region. Count partial outages: if 20 per cent of users cannot log in for an hour, that is downtime for them. Publish planned maintenance windows in the contract or count them.
Watch for the opposite failure: being too reliable for too long. Google’s SLO chapter describes how teams came to depend on its Chubby lock service because it almost never failed, so in any quarter where real failures had not already pulled availability below the target, SRE took it down on purpose, exposing dependencies that assumed it never would. A startup need not go that far. It should make sure its own internal services publish an objective, so other teams design for failure.
Communication during an outage is part of reliability as customers experience it. Keep a status page that is hosted outside the product’s own infrastructure, update it within minutes of an incident being declared, and say what is affected in the customer’s language: payroll runs are delayed, not database latency is elevated. After any incident that breached an objective, send the affected customers a short written account within a week: what happened, how long, who was affected, what has changed so it does not recur. Enterprise buyers in India often judge a vendor more by that note than by the outage itself, and a frank account is the cheapest renewal insurance there is.
Budgeting the engineering
The SRE book’s error budget turns the target into a rule. If the objective is 99.9 per cent, the team may spend 0.1 per cent of the month on failure. While budget remains, releases continue. When it runs out, feature releases pause and the team works on reliability until the budget recovers. This ends the argument between product and engineering, because the number decides.
Price each nine before selling it. A Mumbai HR software company with ₹1.8 crore a month in revenue found that moving its payroll journey from 99.5 to 99.9 per cent needed a second payment gateway, a replica database in a second zone and a paid on-call rota: roughly ₹9 lakh a month in infrastructure and people. Its enterprise accounts, about ₹70 lakh of that monthly revenue, asked for it. It sold 99.9 per cent on payroll processing to enterprise customers only, priced the tier to cover the cost, and kept everyone else on a published 99.5 per cent objective with no credits.
Count the people as well as the servers. A credible 99.9 per cent promise needs someone able to respond within minutes at three in the morning, which means an on-call rota of at least four or five engineers so nobody carries it more than one week a month, paid for the burden and given the next morning off after a bad night. A rota of two burns out within a quarter, and the promise fails with it.
The monthly reliability review
Once a month, engineering brings one page. For each critical journey: the indicator, the objective, the actual and the error budget left. Every incident with its minutes, its customer impact and the credits paid. The dependency chain with each provider’s measured availability. Contracts due for signature with any reliability clause and whether the chain can meet it. The founder or the head of sales signs no SLA that this page has not checked.
The figure is a simplified model and the Mumbai company is illustrative. Contract terms and the RBI circular are as published, checked 11 October 2026. Nothing here is legal advice.
Sources
- Google, Site Reliability Engineering, chapter 3: Embracing Risk — 100 per cent is the wrong target; error budgets; the 99 per cent smartphone.
- Google, Site Reliability Engineering, chapter 4: Service Level Objectives — Definitions of SLI, SLO and SLA; the Chubby planned outage.
- Amazon Web Services, Amazon Compute Service Level Agreement (checked 11 October 2026) — 99.99 per cent region-level commitment; credits of 10, 30 and 100 per cent.
- Reserve Bank of India, Harmonisation of TAT and customer compensation for failed transactions, RBI/2019-20/67, 20 September 2019 — ₹100 a day compensation beyond the turnaround time for failed UPI and IMPS transfers.