Find out what an agent fleet would have caught in your last 90 days of incidents.
A fixed-scope, fixed-price audit of your incident history, alert noise, on-call load, and observability posture. You get an incident-by-incident map of where automation pays off first, and a plan to get there.
Three weeks. Read-only access. Fee credited against remediation if you continue.
- AUTO11 incidents were pod restarts after OOMKill on the same 3 deployments. Memory limits + a rollback runbook close them without a page.
- NOISE14 alert rules paged 210 times and never led to a change. Retire 9, reroute 5 to a ticket queue.
- SPENDCustom metrics cardinality drives 38% of the observability bill. Two label drops recover most of it.
One question, asked of every incident: what should have handled this?
The audit looks at the last 90 days of your production reality, not a maturity model. Every incident is classified by what an agent could have done with it. Everything else in the report exists to support that classification.
Every incident, scored for automation
Detection lag, time to diagnosis, time to fix, and who did what. Each one gets a disposition: could an agent have detected it earlier, triaged it, or remediated it outright, and what would it have needed to do so.
Output: an incident-by-incident automation map, grouped by root-cause pattern.
Which pages lead to a change, and which never do
Alert rules cross-referenced against incident and ticket history. Rules that fire and never produce an action get a cut list. Rules that fire late or never for real incidents get a coverage note.
Output: a retire / reroute / rewrite list per rule, with expected page reduction.
What your rotation actually costs
Pages per engineer per week, off-hours share, escalation depth, and the handful of services responsible for most of it. This is the number your engineers already feel and your board has not seen.
Output: an on-call load profile and the services to fix first.
Can an agent see enough to act?
Trace, metric, and log coverage across the services that appear in incidents. OpenTelemetry readiness. Gaps that would leave an agent blind. Tooling spend is reviewed here too, and the findings usually cover the cost of the audit.
Output: coverage gaps ranked by incident impact, plus a spend findings appendix.
Three weeks, read-only access, one readout.
You give access to the data once. I do the analysis. Your team spends about four hours total across intake and the readout. Nothing changes in production during the audit.
Pull the data
A 60-minute kickoff, then read-only exports from the systems you already run.
- Incident and postmortem history
- PagerDuty / Opsgenie page and escalation log
- Alert rule definitions and firing history
- Observability config and last 3 invoices
- Cluster topology and deployment inventory
Score and classify
Every incident gets a disposition. Every alert rule gets a verdict. The analysis is scripted, so the same questions get asked of every incident the same way.
- Incident-by-incident automation map
- Alert noise cut list with expected page reduction
- On-call load model
- Coverage gaps and spend findings
- Midpoint check-in to validate early findings
Report and roadmap
A written report and a 90-minute readout with your engineering leadership, with the automation roadmap ordered by incidents avoided per week of effort.
- Written report, yours to keep and circulate
- Prioritized 90-day automation roadmap
- Executive summary for the board or CEO
- Optional retainer proposal, scoped to the findings only
A diagnosis your team can act on, with or without me.
The report is written to be handed to an engineer on Monday. No slideware, no framework names, no vendor pitch.
- Incident automation mapEach incident from the window, with disposition, pattern group, and the prerequisites to automate it.
- Alert noise cut listPer-rule verdicts with projected page reduction, ready to apply.
- On-call load profilePages per engineer, off-hours share, and the services driving it.
- Observability coverage gapsRanked by how many incidents each gap touched.
- Tooling spend findingsConcrete reductions in your observability bill, as an appendix.
- 90-day automation roadmapOrdered by incidents avoided per week of effort, with effort estimates.
- One price for the three-week engagement, agreed on the scoping call.
- Credited in full against your first month if you continue into remediation.
- Read-only access only. Nothing changes in production during the audit.
- The spend findings alone typically cover the fee within a quarter.
Remediation as a managed service, with agents doing the work.
If the findings justify it, the audit rolls into an ongoing reliability retainer. I implement the roadmap: alert cleanup, runbook automation, OpenTelemetry instrumentation, and the agent workflows that handle the incidents the audit flagged as automatable.
The agents are how the work gets delivered, not a separate product you have to buy. The measure of success is simple: the same reliability outcomes with fewer human hours each month.
Retainers are only proposed against findings the audit actually supports. No retainer, no problem. The report stands on its own.
Cloud-native teams who are paying for incidents in engineer time.
The audit is built for a specific shape of company. If that's you, the findings will be sharp. If it isn't, I'll say so on the scoping call.
Good fit
- Kubernetes in production, 50 to 500 people, typically Series A to C SaaS.
- Hiring for SRE, DevOps, or platform and want the reliability work moving before the hire lands.
- A recent public incident, or a status page that's had a rough quarter.
- Datadog, Grafana, or New Relic plus PagerDuty or Opsgenie already in the stack.
- A VP of Engineering, CTO, or Head of Platform who owns the decision.
Probably not
- Fewer than a dozen incidents in the last 90 days. There isn't enough signal to audit yet.
- No container orchestration, or a mostly serverless or managed-PaaS footprint.
- Looking for a cloud cost audit. Spend findings are included, but they aren't the point.
- Looking for staff augmentation by the hour.
One operator. Kubernetes, observability, and incident automation.
Cactus Matrix is run by Mike, a Director of SRE who has spent his career on the receiving end of the pager: running Kubernetes platforms at scale, standardizing observability on OpenTelemetry, and building the incident automation that turned recurring pages into closed tickets.
The audit is the same analysis he runs on his own platforms, packaged so another team can get it in three weeks. The agent workflows behind the retainer are the ones he built to stop doing the same remediation twice, including the pipeline that runs this business.
You work with him directly. No account managers, no bench.
The things people ask on the scoping call.
What access do you need?
How is this different from the AI features Datadog, PagerDuty, or incident.io already sell?
Is this a cost-cutting engagement?
Do we have to continue into the retainer?
How much of my team's time does it take?
Bring your last 90 days. Leave with a plan.
A 30-minute scoping call covers your stack, your incident volume, and whether the audit will find enough to be worth your money. If it won't, I'll tell you on the call.