Find where your reliability
breaks down.
You're investing in reliability. Incidents are still recurring. We help you find exactly where the gap is — before the next one happens.
You're investing in reliability.
But where is that investment actually going?#
Most reliability work focuses on what happens during an incident. But the causes are usually upstream and the consequences downstream — in the way changes are validated before they ship, and in whether learnings actually stick after things are fixed. Recoverable looks across the full lifecycle to find where your reliability investment is actually going — and where it isn't.
Most teams only see part of the picture.#
Reliability problems span three phases. Most tooling and attention focuses on the middle one — what happens once something is already broken.
A deploy breaks something it wasn't supposed to touch.
An API contract changed silently. Latency regressed by 60% under load. Nobody caught it before it shipped because the pipeline only checks whether tests pass — not whether the system holds up in reality.
Twenty minutes pass before the right person is even paged.
The service has no clear owner. The runbook exists but hasn't been touched in a year. Escalation depends on someone knowing who to Slack. Every minute of confusion is a minute of downtime.
The post-mortem is written. The same incident recurs six months later.
Action items get created and deprioritised against the product roadmap. Learnings stay with the team that was involved. Adjacent teams with the same latent risk never hear about it.
Where we start.#
We assess all three phases — prepare, respond, and learn — using structured interviews, incident history review, and simulated response drills. The goal is to find where your reliability is actually breaking down.
Most engagements start here. What happens after depends on what we find.
See the full assessment frameworkReasons to choose Recoverable.#
Evidence-based, not opinionated
Every finding is grounded in your incident data, team interviews, and observed processes — not generic best-practice checklists.
Actionable, not academic
You get a ranked list of fixes with effort estimates and expected impact — not a 40-page report that sits in a drawer.
Full lifecycle, not just response
Most assessments focus on incident response. We look at what happens before, during, and after — because the causes are usually upstream.
No disruption to your team
The assessment runs alongside your normal work. We need 2–3 hours of interview time total, spread across the people who matter.
Built by someone who's been on every side of an incident.#
Recoverable is built by Hugo van der Horst. Seven years at Adyen — spanning Technical Support, Platform Reliability, Site Reliability Engineering, and Platform Observability — including leading the EMEA 24/7 on-call support team for one of Europe's largest payment platforms. He's led the team responding to incidents, owned the incident management programme, and built the observability tooling that detects them. The assessment draws directly from that operational range.
Find out where your reliability breaks down.
No commitment, no pitch. A 30-minute conversation to understand whether the problem you're facing is one we can help with.