A structured diagnostic of your full incident lifecycle.
No disruption to your team. A clear picture of where reliability is breaking down — and what to fix first.
An assessment of the full incident lifecycle.#
We assess all three phases — prepare, respond, and learn — using a combination of structured interviews, incident history review, and, where appropriate, simulated response drills. The goal isn't to grade you across fixed dimensions. It's to find where your reliability is actually breaking down.
Schedule a free 30-min callPrepare — before things break
- Change validation
Are deploys validated against API contracts and load baselines before they reach production?
- Blast radius awareness
Can teams assess the impact of a change before it ships? Are service dependencies mapped and current?
- Alert signal quality
Are alerts actionable, or is the team desensitised to noise? What failure modes aren't being measured?
Respond — when things are breaking
- Ownership clarity
Every service has a clear owner. Every incident has an accountable responder within minutes.
- Runbook quality
Runbooks exist, are current, and are actually followable under pressure.
- Escalation paths
The right people get looped in at the right time. Escalation is documented and tested.
- On-call coverage
No timezone gaps. Coverage doesn't silently disappear on bank holidays or during busy periods.
Learn — after things are fixed
- Post-mortem quality
Do post-mortems surface real systemic causes, or produce a polished write-up that nobody acts on?
- Action item follow-through
Do remediation items get completed, or sink into backlogs against the product roadmap?
- Learning propagation
Do learnings reach adjacent teams with the same latent risk — or stay local to the team involved?
Three things you walk away with.#
Not a generic template. Not a PDF you'll file and forget. Three specific deliverables built around your actual incidents, your actual team, your actual stack.
A written findings report, structured by phase
For each of the three phases — prepare, respond, learn — a clear breakdown of what's working, what's failing, and what's quietly building risk. Written in plain language, structured so you can share it with your engineering leadership without translation.
Example
"Prepare: no contract validation in place — three of your last five major incidents traced back to an API change that wasn't caught before shipping. Respond: ownership is solid. Runbooks are the weak link — most haven't been updated in over a year. Learn: post-mortems are being written but action items are consistently deprioritised. The same database connection exhaustion pattern appears in four separate incidents over the last 12 months."
A ranked fix list with rationale
The 3–5 highest-leverage things to address, ranked by impact not severity. Each one comes with a clear "why this first" — grounded in what we actually found, not generic best practice. Something you can bring to a planning meeting and defend.
A live readout with your team
Findings are never just emailed. We present on a call, walk through what we found, take questions, and push back where needed. The conversation is part of the deliverable — it's where the real priorities get set.
Three weeks. No disruption to your team.#
The assessment works around your team's actual rhythm — not pull people out of it.
Scoping and interviews
We start by understanding your stack, team structure, and incident history. Structured interviews with engineers, on-call leads, and engineering managers — looking for the pattern of where things consistently stall.
Validation and drills
Simulated incident scenarios to verify ownership, runbook usability, escalation paths, and on-call gaps under realistic pressure. We also review your last 6–12 months of incidents and post-mortems directly.
Findings and next steps
A clear diagnostic report — where reliability is breaking down, what the highest-leverage fixes are, and what we'd recommend working on first. Presented to the team, not just emailed as a PDF.
Built for teams with real incident history.#
Most valuable for teams that have been running long enough to have patterns — and are tired of addressing the same symptoms without finding the cause.
VP Engineering
You keep explaining the same incidents to the CTO. Recoverable gives you a clear, honest picture of where your reliability investment is actually going — and what's blocking improvement.
Engineering Manager
You know the debt is there but can't make the case against the product roadmap. Recoverable surfaces the data that makes the argument concrete.
SRE / Platform Engineer
You've seen the patterns for years. Recoverable gives you an external, structured view to confirm what you already suspect — and the report to act on it.
Good fit if...
20+ engineers, a real incident cadence, and the same problems keep coming back. You want to understand the system, not just fix the latest symptom.
Something specific brought you here.#
The assessment is most useful at inflection points — moments where reliability stops being abstract and starts being urgent.
After a major incident
One to four weeks post-outage, the post-mortem keeps surfacing patterns the team already knew about.
Before a board review or fundraise
The VP Engineering wants an independent read before reliability becomes a diligence topic.
New VP Eng or Head of SRE in seat
Fresh leadership wanting an independent baseline before committing to a roadmap.
Same incidents keep recurring
Different services, same failure modes — and nobody can name the underlying pattern.
Q4 budget planning
Reliability investment needs a defensible diagnostic, not a vendor pitch.
Post-acquisition integration
Two platforms merging, different reliability practices — need a single picture before you consolidate.
Find out where your reliability breaks down.
No commitment, no pitch. A 30-minute conversation to understand whether the problem you're facing is one we can help with.