Articles

Field notes on incidents, recovery, cost, and the operating practices that keep production work safer.

Incident Management7 min read

Human error is not a root cause

After an incident, stopping at human error leaves the next person exposed. Treat people as part of the system and change the conditions around them.

SRE8 min read

Cut MTTR without adding more alerts

More paging rarely shortens recovery. Clear ownership, better signals, and practiced recovery paths do more for MTTR than another noisy threshold.

FinOps7 min read

FinOps decisions that keep reliability intact

Cloud cost work belongs in the operating model. Cut waste with ownership and clear trade-offs, not by quietly removing the capacity and observability that keep incidents small.

AIOps6 min read

What AIOps can and cannot do in incident work

AIOps can correlate signals and reduce noise. It does not replace ownership, runbooks, or judgment about customer impact. Set expectations before you buy a platform story.