Most teams only read their SLA properly during an outage. That is the worst possible moment: the incident channel is on fire, your users are emailing, and you are squinting at a PDF trying to work out whether a four-hour response means four hours to a human or four hours to an automated acknowledgement.
A custom software sla review fixes that asymmetry. Instead of discovering gaps under pressure, you audit the agreement on a quiet afternoon, score each clause against how your business actually runs, and walk into the renewal conversation with a specific list of changes.
This guide gives you a clause-by-clause method you can complete in a couple of hours, plus the negotiation points that matter most for custom-built systems rather than off-the-shelf SaaS.
Key takeaways
- Audit the SLA against your own incident history, not in the abstract, the gaps that matter are the ones you have already hit.
- Check the definitions before the numbers: uptime percentage, severity tiers, and what counts as "response" decide almost every dispute.
- Credits are usually a weak remedy. Fix response and restoration commitments first, then negotiate credits as a backstop.
- Start the review at least one full quarter before auto-renewal so you have leverage and time to compare alternatives.
- Document every ambiguity you find in writing and get the vendor to confirm the interpretation, that email is worth more than the clause.
Why custom software SLAs drift out of alignment
Off-the-shelf vendors can offer one SLA to thousands of customers because the product is identical for everyone. Custom software is the opposite: your system has bespoke integrations, a specific data model, and dependencies on third-party APIs that the vendor may or may not control.
That mismatch is where drift creeps in. The SLA was likely written during the build phase, when everyone was optimistic and the priority was shipping. Two years later, the system carries more traffic, more integrations, and more business criticality, but the agreement still says what it said on day one.
Common symptoms we see when teams finally sit down to review:
- The uptime target is measured monthly, so one bad week is diluted by three good ones.
- Severity tiers are defined by the vendor's judgment rather than objective criteria.
- Response time commitments exclude "waiting on customer" time with no cap on how long that pause can run.
- Maintenance windows are open-ended and can be invoked with short notice.
- Exclusions cover so much ground that the headline uptime number is effectively decorative.
Build your evidence base before you read the contract
An SLA review without data becomes an argument about feelings. Spend the first thirty minutes assembling facts.
Pull your incident history
Export every ticket, outage, and escalation from the last twelve months. For each one, note when it started, when the vendor acknowledged it, when a human actually engaged, and when it was resolved. If you do not track this today, start now, even a simple spreadsheet gives you negotiating power next cycle.
Map business impact to each incident
Grade incidents by what they cost you: revenue blocked, users unable to complete a core task, compliance exposure, or internal workaround time. This becomes your severity model, and it is the single most persuasive artifact you can bring to a renewal meeting.
Identify your true critical paths
Not every part of a custom system deserves the same protection. Payment processing, authentication, and core data writes usually need tighter commitments than an internal reporting screen. A good SLA review often ends with tiered coverage rather than one blanket target.
A clause-by-clause custom software SLA review
Work through the agreement in this order. Definitions first, because they silently govern everything downstream.
1. Uptime definition and measurement
Find the exact wording of the availability commitment. Ask:
- Is it measured monthly, quarterly, or annually? Shorter windows are friendlier to you.
- Is it measured per service, per component, or across the whole platform? Platform-wide averaging hides single-component failures.
- How is availability probed, synthetic checks, real user monitoring, or vendor-reported status? Who owns the measurement?
- What is the measurement interval? A check every five minutes will miss short but painful outages.
If the vendor controls both the system and the measurement, ask for read access to the monitoring data. That single request resolves more disputes than any clause rewrite.
2. Exclusions and maintenance windows
Exclusions are where headline uptime numbers quietly evaporate. Read them carefully and push back on anything that is not genuinely outside the vendor's control.
- "Scheduled maintenance" should have defined notice periods, defined maximum durations, and a defined window outside your peak hours.
- "Third-party service failure" should not excuse a vendor who chose a single upstream provider with no fallback for a critical path.
- "Customer-caused" incidents should require the vendor to demonstrate causation, not just assert it.
- Force majeure clauses should be narrow enough that routine scaling failures do not qualify.
3. Severity tiers and classification
This is the most negotiated and most abused part of any agreement. Vague tiers let a vendor downgrade your emergency to fit their staffing.
Push for objective, testable criteria. For example: a severity-one incident is one where a defined critical user journey is unavailable to all users, or where data integrity is at risk. Severity two is a critical journey degraded or unavailable to a subset. Severity three is a non-critical function impaired with a workaround available.
Then settle two process questions: who has authority to declare severity, and how fast can you escalate if you disagree? A good answer is a joint declaration within a short window and an automatic escalation if the customer disputes the classification and the vendor does not respond.
4. Vendor SLA response times and the clocks that matter
"Response time" is the most misleading phrase in the document. It usually means acknowledgement, not engagement. Build a table of commitments and check each one:
- Acknowledgement time: an automated ticket reply often satisfies this. Nearly worthless on its own.
- Human engagement time: when a qualified engineer starts working the issue. This is the number that actually matters.
- Workaround or mitigation time: when service is restored to a usable state, even if the root cause is unresolved.
- Resolution time: full fix, which is often deliberately left uncommitted.
Also check the clock's business hours. A four-hour response during business hours is a very different promise from a four-hour response around the clock. For anything customer-facing, insist on coverage that matches your users' time zones, not your vendor's office hours.
Finally, look for pause conditions. If the clock stops whenever the vendor is "awaiting customer information," ask for a cap on total paused time per incident.
5. Escalation and communication
Response times are meaningless if you cannot reach a decision-maker. Confirm that the agreement names escalation contacts by role, defines trigger points for each level, and requires a status cadence during severe incidents, for example, proactive updates at fixed intervals until mitigation.
Ask for a post-incident review commitment too: a written root cause analysis within a defined period for severity-one events, delivered to a named stakeholder.
6. Credits, remedies, and the honest math
Service credits feel like protection but rarely compensate real losses. A credit is typically a percentage of the monthly fee, capped, and available only if you claim it within a window. If your custom system processes revenue, the credit will not come close to the business impact.
Treat credits as a backstop, not the remedy. The stronger remedies are structural:
- A right to escalate and require a remediation plan after repeated misses.
- A termination-for-convenience or termination-for-cause trigger tied to sustained SLA failure.
- Source code escrow or a documented handover obligation, so you are never trapped.
- Priority access to engineering capacity during severe incidents.
On sla uptime credits negotiation, the useful move is to widen the trigger and shorten the claim process rather than chase a bigger percentage. Automatic credits remove the awkwardness of asking.
7. Custom software maintenance agreement terms
The SLA covers incidents; the maintenance agreement covers everything else, and the two are often in separate documents. Review them together.
- What is included in the support retainer versus billed as change requests?
- How are security patches and dependency upgrades handled, and on what timeline?
- What is the response commitment for non-incident work, feature requests, small fixes, environment changes?
- How are rates defined, and how much notice is required before they change?
- What happens at the end of the term: knowledge transfer, documentation, access to repositories and infrastructure?
That last point is the one most teams regret skipping. Exit terms are cheap to negotiate now and expensive to negotiate later.
A software support SLA checklist you can run today
Use this as a working software support sla checklist during your review. Score each item as adequate, ambiguous, or missing.
- Availability target, measurement window, and measurement owner are explicit.
- Monitoring data is accessible to you, not just reported by the vendor.
- Exclusions are narrow, enumerated, and require evidence.
- Maintenance windows have notice periods and duration caps.
- Severity tiers use objective, testable criteria.
- You can dispute a severity classification and get a response.
- Response times distinguish acknowledgement from human engagement.
- Coverage hours match your users, not the vendor's office hours.
- Pause conditions are capped per incident.
- Escalation contacts, triggers, and update cadence are named.
- Post-incident reviews are required for severe events.
- Credits are automatic, with a clear claim path.
- Repeated failure triggers a remediation plan or exit right.
- Maintenance scope, security patching, and change-request boundaries are defined.
- Exit and handover obligations are documented.
How to renegotiate before renewal
Timing is most of the leverage. Auto-renewal clauses often require notice well in advance, so diarize the review a full quarter before the term ends.
Bring three things to the conversation: your incident data, your business-impact grading, and a short prioritized list of changes. Vendors respond far better to "here are the four clauses that caused us pain this year, with evidence" than to a general request for a better SLA.
Prioritize ruthlessly. If you can only win two changes, choose human engagement times for severity-one incidents and objective severity criteria. Those two determine how an outage actually feels.
If the vendor will not move on substance, ask for the ambiguities to be resolved in writing instead. A written clarification of how they interpret a clause is enforceable in practice and much easier to obtain than a redline.
And if you are weighing whether to stay or rebuild, it helps to talk to a team that has taken over systems mid-life. Avaton builds and maintains custom software, including taking over codebases from other vendors, so we have seen these agreements from both sides. You can review our engineering and support services, look at past work we have delivered, or simply start a conversation about your setup. We also publish ongoing notes on maintaining and modernizing custom software if you want more context before you decide.
Frequently Asked Questions
How often should you review a custom software SLA?
At minimum once a year, and always at least one full quarter before any auto-renewal date. Review it sooner if your system's criticality has changed, if you have onboarded significantly more users, or if you have experienced an incident where the vendor's obligations felt unclear. A short review after every severity-one incident is a good habit because the details are still fresh.
What is the difference between response time and resolution time in an SLA?
Response time is typically the period before the vendor acknowledges or engages with your issue, while resolution time is the period before the problem is fully fixed. Many agreements commit to response times but leave resolution uncommitted. For custom software, the most valuable commitment is mitigation time, meaning when service is restored to a usable state even if the underlying root cause is still being investigated.
Are SLA uptime credits worth negotiating?
They are worth having, but they rarely compensate real business losses because they are usually a capped percentage of the monthly fee. Treat credits as a backstop and focus your negotiation on structural remedies instead: faster human engagement, objective severity criteria, a remediation plan requirement after repeated failures, and a termination right tied to sustained SLA misses. Automatic credits with no claim process are more useful than a larger percentage you have to fight for.
What should a software support SLA checklist include?
At a minimum: an explicit availability target with a defined measurement window and owner, narrow and evidence-based exclusions, capped maintenance windows, objective severity tier definitions, separate acknowledgement and human engagement times, coverage hours matching your users, capped pause conditions, named escalation contacts and update cadence, required post-incident reviews, automatic credits, a repeated-failure remediation trigger, and documented exit and handover obligations.
Can you renegotiate an SLA mid-contract?
Sometimes, especially if you can show documented failures against the current terms. Even when the vendor will not amend the contract text, you can often obtain a written clarification of how ambiguous clauses are interpreted, which is practically enforceable. If the vendor refuses both, that is useful information: it tells you how the next outage will go and gives you a clear reason to evaluate alternatives before renewal.
Cover: Photo by https://kaboompics.com/ on Pexels
