Reliability & incident audit · Kubernetes teams

Find out what an agent fleet would have caught in your last 90 days of incidents.

A fixed-scope, fixed-price audit of your incident history, alert noise, on-call load, and observability posture. You get an incident-by-incident map of where automation pays off first, and a plan to get there.

Three weeks. Read-only access. Fee credited against remediation if you continue.

audit-summary · 90-day window Illustrative
Incidents
47
Median MTTR
74min
Pages / on-call wk
31
Alerts w/ no action
68%
Incident disposition, if agents were on-calln=47
Auto-remediable · 22 Triage & hand-off · 15 Needs a human · 10
  • AUTO11 incidents were pod restarts after OOMKill on the same 3 deployments. Memory limits + a rollback runbook close them without a page.
  • NOISE14 alert rules paged 210 times and never led to a change. Retire 9, reroute 5 to a ticket queue.
  • SPENDCustom metrics cardinality drives 38% of the observability bill. Two label drops recover most of it.
What the audit covers

One question, asked of every incident: what should have handled this?

The audit looks at the last 90 days of your production reality, not a maturity model. Every incident is classified by what an agent could have done with it. Everything else in the report exists to support that classification.

01 · Incident history

Every incident, scored for automation

Detection lag, time to diagnosis, time to fix, and who did what. Each one gets a disposition: could an agent have detected it earlier, triaged it, or remediated it outright, and what would it have needed to do so.

Output: an incident-by-incident automation map, grouped by root-cause pattern.

02 · Alert noise

Which pages lead to a change, and which never do

Alert rules cross-referenced against incident and ticket history. Rules that fire and never produce an action get a cut list. Rules that fire late or never for real incidents get a coverage note.

Output: a retire / reroute / rewrite list per rule, with expected page reduction.

03 · On-call load

What your rotation actually costs

Pages per engineer per week, off-hours share, escalation depth, and the handful of services responsible for most of it. This is the number your engineers already feel and your board has not seen.

Output: an on-call load profile and the services to fix first.

04 · Observability posture

Can an agent see enough to act?

Trace, metric, and log coverage across the services that appear in incidents. OpenTelemetry readiness. Gaps that would leave an agent blind. Tooling spend is reviewed here too, and the findings usually cover the cost of the audit.

Output: coverage gaps ranked by incident impact, plus a spend findings appendix.

How it works

Three weeks, read-only access, one readout.

You give access to the data once. I do the analysis. Your team spends about four hours total across intake and the readout. Nothing changes in production during the audit.

Week 1 · Intake

Pull the data

A 60-minute kickoff, then read-only exports from the systems you already run.

  • Incident and postmortem history
  • PagerDuty / Opsgenie page and escalation log
  • Alert rule definitions and firing history
  • Observability config and last 3 invoices
  • Cluster topology and deployment inventory
Week 2 · Analysis

Score and classify

Every incident gets a disposition. Every alert rule gets a verdict. The analysis is scripted, so the same questions get asked of every incident the same way.

  • Incident-by-incident automation map
  • Alert noise cut list with expected page reduction
  • On-call load model
  • Coverage gaps and spend findings
  • Midpoint check-in to validate early findings
Week 3 · Readout

Report and roadmap

A written report and a 90-minute readout with your engineering leadership, with the automation roadmap ordered by incidents avoided per week of effort.

  • Written report, yours to keep and circulate
  • Prioritized 90-day automation roadmap
  • Executive summary for the board or CEO
  • Optional retainer proposal, scoped to the findings only
What you get

A diagnosis your team can act on, with or without me.

The report is written to be handed to an engineer on Monday. No slideware, no framework names, no vendor pitch.

  • Incident automation mapEach incident from the window, with disposition, pattern group, and the prerequisites to automate it.
  • Alert noise cut listPer-rule verdicts with projected page reduction, ready to apply.
  • On-call load profilePages per engineer, off-hours share, and the services driving it.
  • Observability coverage gapsRanked by how many incidents each gap touched.
  • Tooling spend findingsConcrete reductions in your observability bill, as an appendix.
  • 90-day automation roadmapOrdered by incidents avoided per week of effort, with effort estimates.
Pricing
Fixed fee. Fixed scope. Quoted before we start.
  • One price for the three-week engagement, agreed on the scoping call.
  • Credited in full against your first month if you continue into remediation.
  • Read-only access only. Nothing changes in production during the audit.
  • The spend findings alone typically cover the fee within a quarter.
Book a scoping call
After the audit

Remediation as a managed service, with agents doing the work.

If the findings justify it, the audit rolls into an ongoing reliability retainer. I implement the roadmap: alert cleanup, runbook automation, OpenTelemetry instrumentation, and the agent workflows that handle the incidents the audit flagged as automatable.

The agents are how the work gets delivered, not a separate product you have to buy. The measure of success is simple: the same reliability outcomes with fewer human hours each month.

Retainers are only proposed against findings the audit actually supports. No retainer, no problem. The report stands on its own.

Delivery hours per client per month · illustrative
Same reliability outcomes, fewer hours. As automation takes over the incidents the audit identified, human delivery time falls quarter over quarter. That curve is the retainer's report card.
Who it's for

Cloud-native teams who are paying for incidents in engineer time.

The audit is built for a specific shape of company. If that's you, the findings will be sharp. If it isn't, I'll say so on the scoping call.

Good fit

  • Kubernetes in production, 50 to 500 people, typically Series A to C SaaS.
  • Hiring for SRE, DevOps, or platform and want the reliability work moving before the hire lands.
  • A recent public incident, or a status page that's had a rough quarter.
  • Datadog, Grafana, or New Relic plus PagerDuty or Opsgenie already in the stack.
  • A VP of Engineering, CTO, or Head of Platform who owns the decision.

Probably not

  • Fewer than a dozen incidents in the last 90 days. There isn't enough signal to audit yet.
  • No container orchestration, or a mostly serverless or managed-PaaS footprint.
  • Looking for a cloud cost audit. Spend findings are included, but they aren't the point.
  • Looking for staff augmentation by the hour.
Who's doing the work

One operator. Kubernetes, observability, and incident automation.

Cactus Matrix is run by Mike, a Director of SRE who has spent his career on the receiving end of the pager: running Kubernetes platforms at scale, standardizing observability on OpenTelemetry, and building the incident automation that turned recurring pages into closed tickets.

The audit is the same analysis he runs on his own platforms, packaged so another team can get it in three weeks. The agent workflows behind the retainer are the ones he built to stop doing the same remediation twice, including the pipeline that runs this business.

You work with him directly. No account managers, no bench.

Kubernetes OpenTelemetry Datadog · Grafana PagerDuty · Opsgenie Incident automation Agentic SRE workflows
Questions

The things people ask on the scoping call.

What access do you need?
Read-only, and only to the systems the audit uses: your incident tool, paging platform, alert definitions, observability billing, and cluster inventory. Exports are fine if you'd rather not grant accounts. Nothing is changed in production during the audit, and access is revoked at the readout.
How is this different from the AI features Datadog, PagerDuty, or incident.io already sell?
Those tools answer "what can our AI do?" The audit answers "what would have happened in your last 90 days?" It's evidence from your own incidents, before you commit to a platform or a retainer. If the answer is that a vendor feature covers most of it, the report will say that.
Is this a cost-cutting engagement?
No. The audit is about incidents and what handles them. Observability spend is reviewed because an agent needs good telemetry to act, and because the findings are usually large enough to pay for the audit. They sit in an appendix, not the headline.
Do we have to continue into the retainer?
No. The report is written so your own team can execute the roadmap. If you do continue, the audit fee is credited against your first month, and the retainer scope is limited to what the audit found.
How much of my team's time does it take?
About four hours across the three weeks: a kickoff, the data pull, a midpoint check-in, and the readout. Engineers are welcome at the readout but not required before it.
Get started

Bring your last 90 days. Leave with a plan.

A 30-minute scoping call covers your stack, your incident volume, and whether the audit will find enough to be worth your money. If it won't, I'll tell you on the call.

Book a scoping call hello@cactusmatrix.com Replies within one business day.