I have spent years in the room when production breaks.

I help organisations reduce incident impact, improve SRE maturity, control cloud cost, and remove the conditions that make human mistakes expensive. Restoring service is the first job. Changing the system so the next failure hurts less is the work that follows.

Why The Incident Guy

The name comes from years spent handling production incidents: overnight pages, unclear ownership, noisy alerts, brittle releases, and postmortems that stopped too early. People started calling me when systems were down and the path back was unclear.

Over time the work widened. Incident response led into SRE practices, observability, architecture reviews, FinOps, DevOps, DevSecOps, and careful use of AIOps. The through-line stayed the same: keep customers safer when systems fail, and make failure cheaper to recover from.

Where I work with teams

Incident Management and SRE

Reduce blast radius, shorten recovery, clarify ownership, and strengthen the operating model around production systems.

FinOps and Cloud Cost

Make cloud spend visible, assign cost ownership, remove waste, and keep cost decisions from quietly eroding reliability.

DevOps and DevSecOps

Tighten delivery paths, put safety checks where they belong, and keep security inside the same workflow that ships changes.

AIOps and Automation

Use automation and signal correlation to cut noise and speed triage. Treat AIOps as decision support, not autonomous incident resolution.

Projects connected to this work

These sites publish the operating lessons that do not fit inside a single consulting engagement.

Vigiles

Built from lessons learned during incident response. Still improving, with a long road ahead.

OceanDB Pro

Independent OceanBase guides, migration notes, runbooks, and production lessons.

CloudArch Pro

Practical cloud architecture, FinOps, migration, reliability, and APAC infrastructure guidance.

Latest articles

Incident Management7 min read

Human error is not a root cause

After an incident, stopping at human error leaves the next person exposed. Treat people as part of the system and change the conditions around them.

SRE8 min read

Cut MTTR without adding more alerts

More paging rarely shortens recovery. Clear ownership, better signals, and practiced recovery paths do more for MTTR than another noisy threshold.

FinOps7 min read

FinOps decisions that keep reliability intact

Cloud cost work belongs in the operating model. Cut waste with ownership and clear trade-offs, not by quietly removing the capacity and observability that keep incidents small.

Contact

If your organisation needs help with recurring incidents, SRE maturity, cloud cost, or production architecture, write to me directly.

hey@theincidentguy.com