An uptime SLA measures one service's availability, not the reliability your customers experience. A host can meet a 99.99% SLA while DNS, gateway, CDN, plugin, or support failures stop sales.
A 99.9% uptime SLA allows up to 43.8 minutes of downtime monthly. It does not promise fast fixes or a working customer journey.
For SMBs, check incident history, MTTR, dependencies, and recovery readiness. The headline percentage alone cannot protect revenue.
A 99.99% SLA can still leave you offline
A Service Level Agreement, or SLA, is a contract for one named service. It defines availability during a set measurement period.
An SLA is useful, but it does not promise a full customer journey. Visitors may still fail to find your site, log in, order, get email, or reach staff.
Monthly uptime usually subtracts covered downtime from total minutes in a billing month. Providers may exclude planned work, customer mistakes, third-party networks, DDoS events, and unnamed services.
The contract may protect the server while customers still cannot buy.
| Monthly availability | Downtime allowed per month | Business meaning |
|---|
| 99% | About 7 hours 18 minutes | Too loose for most revenue-producing sites |
| 99.5% | About 3 hours 39 minutes | May suit a low-impact internal tool |
| 99.9% | About 43.8 minutes | Often acceptable only with recovery controls |
| 99.99% | About 4.4 minutes | Strong target, but only for covered components |
For an SMB, the best reliability test is simple: Can a customer reach the site, complete the key action, and receive confirmation during a failure? If services outside the SLA decide that result, the percentage cannot protect the business.
Your SLA misses customer-facing dependencies
Real-world reliability depends on the weakest link in the customer journey. It does not depend on the best SLA in your vendor list.
A visitor needs DNS to find your site. They also need a CDN or network path, working code, and data.
Many customers also need payment, email, or booking APIs to finish. Every link can stop the sale.
DNS can take down a healthy server
DNS means Domain Name System. It turns a domain like yourstore.com into a server address.
Think of DNS as a business directory. Your building can be open, but visitors cannot come if the listing gives the wrong street.
A healthy server is useless when DNS sends customers elsewhere.
Fast pages can still have failed checkout
Map the full transaction from start to finish. Include the browser, CDN, web app, database, inventory system, fraud tool, payment gateway, email, and support inbox.
Give each link an owner, a status page, and a fallback choice. This shows who acts when a service fails.
The most common mistake is testing page loads but not checkout completion. A fast home page does not mean customers can pay.
Availability is not fault tolerance
High availability means a service keeps working when one part fails. It often uses duplicate parts and automatic switching.
Fault tolerance goes further and allows failure with little or no break. It usually needs redundancy, load balancing, and automatic failover.
These terms sound alike, but they protect different risks.
Choose providers by recovery, not just uptime
A provider with clear incident reports, short recovery times, and reachable support may be safer. That can matter more than a higher advertised SLA.
Mean time to recovery, or MTTR, measures restoration time after an incident. For small businesses, it often predicts lost revenue better than annual availability claims.
Choose the provider that shows how it recovers, not only what it promises.
Audit the SLA before comparing prices
Read the SLA as a contract you might enforce during an outage. Confirm the covered service, uptime percentage, measurement method, and calculation window.
Check whether the provider measures availability from its own network or outside locations. Then list exclusions for planned work, customer settings, DDoS events, third-party networks, and beta features.
Check the maintenance notice period and the deadline for credit claims. Also check the maximum credit you can receive.
Read the evidence behind the promise
Compare the contract with the provider's status page and incident history. Read root-cause reports, postmortems, and support records where available.
Contract wording matters less when recovery history is slow or unclear. Prompt and open recovery gives the promise more weight.
SLA credits may work in theory, but they rarely fix a lost peak-hour sale. Recovery evidence is what protects the next incident.
Compare the hosting model to the workload
| Workload | Reasonable setup | Reliability proof to request |
|---|
| Brochure site | Managed hosting with tested backups | External monitoring and restore test |
| Ecommerce store | Managed cloud or VPS with managed database | Checkout monitoring, MTTR history, payment fallback |
| Bookings or payments | Redundant application path where justified | Failover test and 24/7 escalation path |
| Internal SaaS tool | Cloud service with backup and access plan | Recovery time target and support response evidence |
An SLA is the vendor's contract minimum and often triggers service credits. An SLO is the reliability target that a team tries to meet.
An error budget is the allowed unreliability before teams pause new changes. Your business may set a stricter SLO for checkout, login, or bookings.
Track incident frequency, failed requests, transaction errors, MTTR, support response time, and repeat failures. These measures show whether customers can finish their task.
A short DNS or CDN outage can use little monthly downtime. It can still ruin a peak sales period.
Match resilience spending to outage cost
The right resilience level depends on what stops when your service stops. A company site, store, payment flow, booking system, and team tool need different protection.
For an SMB, outage cost is more useful than a generic uptime target. It ties technical spending to actual business loss.
Build an affordable recovery baseline
Small and medium businesses often gain more from a tested baseline than complex design. Start with monitoring that runs outside your host.
An internal monitor can fail with the server it watches. Outside monitoring can alert you when customers cannot reach the service.
Use tested backups and confirm that you can restore them. A backup you never restore is like a spare key you never checked.
Know when more redundancy is justified
Use multi-zone or multi-region hosting when one offline hour costs more than yearly added infrastructure and staff work. This can fit payments, time-sensitive bookings, and SaaS tools used all day.
Estimate outage cost by adding lost gross profit, idle payroll, support time, refunds, recovery discounts, and lost leads. Do not count only the missed order value.
For example, a store may earn $1,200 in gross profit during a peak hour. A checkout failure may cost more through abandoned carts and support work.
Extra redundancy should pay for itself through avoided loss.
An enterprise SLA is not a priority for a personal project without revenue, users, or business processes tied to availability. In that case, price, ease of use, and basic backups may matter more than costly high-availability design.
This calculation makes availability a business choice. A low-traffic brochure site may accept more downtime.
A booking or payment service may need stronger recovery controls. The next questions cover the terms buyers ask most often.
Common questions
Is 99.99% uptime good for an SMB?
99.99% uptime is good only when it covers services customers actually use. It allows about 4.4 minutes of covered downtime each month.
DNS, checkout, and app failures may sit outside that promise. Check each dependency before trusting the number.
What does a 99.9% uptime SLA mean?
A 99.9% SLA allows about 43.2 minutes of covered downtime in a 30-day month. It allows about 43.8 minutes in an average 30.4-day month.
Read exclusions and credit rules. Providers often exclude planned work and third-party failures.
Are SLA credits worth relying on?
SLA credits rarely cover an outage's real cost for a revenue-producing SMB. Credits often equal a small part of the monthly bill.
Providers may require a claim within 30 to 60 days. Keep incident records and claim details.
Is cloud hosting more reliable than a VPS?
Cloud hosting can be more resilient than a VPS when it has managed redundancy across zones. A basic cloud server can still have one database, region, or DNS dependency.
Architecture matters more than the hosting label. Ask how each part fails and recovers.
Which reliability metric matters most after uptime?
MTTR matters most after uptime because it measures recovery time. Review incident frequency, failed-request rate, and support response time with MTTR.
That combination gives a fuller picture. One number cannot show every customer risk.
Make the next hosting decision on evidence
Choose the provider and design that protect the customer action that earns revenue. A 99.99% SLA is a useful filter, but it cannot replace dependency maps and tested recovery plans.
Ask each provider for the covered service, past incidents, escalation path, and recovery proof in your US region. Vague answers mean the advertised availability should not decide the purchase.
The best host is the one that proves recovery under pressure.
What matters most:- An SLA covers a defined service, not the full customer journey.
- MTTR, incident patterns, and support quality show more than a headline availability percentage.
- Protect DNS, payments, applications, and backups because any one can cause a business outage.
- Match redundancy spending to the financial cost of being unavailable.