Incident Management and SRE
Reduce blast radius, shorten recovery, clarify ownership, and strengthen the operating model around production systems.
I help organisations reduce incident impact, improve SRE maturity, control cloud cost, and remove the conditions that make human mistakes expensive. Restoring service is the first job. Changing the system so the next failure hurts less is the work that follows.
The name comes from years spent handling production incidents: overnight pages, unclear ownership, noisy alerts, brittle releases, and postmortems that stopped too early. People started calling me when systems were down and the path back was unclear.
Over time the work widened. Incident response led into SRE practices, observability, architecture reviews, FinOps, DevOps, DevSecOps, and careful use of AIOps. The through-line stayed the same: keep customers safer when systems fail, and make failure cheaper to recover from.
Reduce blast radius, shorten recovery, clarify ownership, and strengthen the operating model around production systems.
Make cloud spend visible, assign cost ownership, remove waste, and keep cost decisions from quietly eroding reliability.
Tighten delivery paths, put safety checks where they belong, and keep security inside the same workflow that ships changes.
Use automation and signal correlation to cut noise and speed triage. Treat AIOps as decision support, not autonomous incident resolution.
These sites publish the operating lessons that do not fit inside a single consulting engagement.
Built from lessons learned during incident response. Still improving, with a long road ahead.
Independent OceanBase guides, migration notes, runbooks, and production lessons.
Practical cloud architecture, FinOps, migration, reliability, and APAC infrastructure guidance.
After an incident, stopping at human error leaves the next person exposed. Treat people as part of the system and change the conditions around them.
More paging rarely shortens recovery. Clear ownership, better signals, and practiced recovery paths do more for MTTR than another noisy threshold.
Cloud cost work belongs in the operating model. Cut waste with ownership and clear trade-offs, not by quietly removing the capacity and observability that keep incidents small.
If your organisation needs help with recurring incidents, SRE maturity, cloud cost, or production architecture, write to me directly.