Skip to content

Anomaly detection

Plutus compares each day’s spend against a baseline and flags a day that’s both unusually high and large enough to matter. There’s no machine learning model behind it — every flag is a specific comparison, and the alert text always states which comparison fired.

The detection runs once a day per (cost source, service) group, using the trailing 29 days of cost data:

  1. A group needs at least 3 prior days of history before its latest day can be judged. A connection you just added doesn’t have a baseline yet, so its first few days never trigger a flag.
  2. The baseline is day-of-week aware. If there are at least 4 prior occurrences of the same weekday (4 previous Tuesdays for a Tuesday being evaluated), the baseline is the average of just those days. That keeps a quiet weekend from dragging down the average and making an ordinary Monday look like a spike. Without enough same-weekday history yet, it falls back to a flat average of the prior 7 days.
  3. Two conditions both have to be true. The day has to be at least 50% above its baseline, and the dollar increase has to clear a noise floor (the greater of $10 or 2% of the account’s average daily spend, capped at $500). The floor exists so a swing on a near-zero line — $0.10 going to $0.20 — doesn’t fire just because the percentage looks dramatic.

A flagged day gets a severity of high, medium, or low, based on how far past the noise floor it is — the same severity labels mean the same thing whether the account spends $300 a day or $50,000 a day.

Example: an account’s average Tuesday spend on Amazon EC2 has been $2,100 over the last 4 Tuesdays. Today (a Tuesday) EC2 comes in at $3,400. That’s 62% above baseline and well past the noise floor, so it’s flagged. The alert reads:

Tuesday’s spend of $3,400.00 is 62% above the average Tuesday spend of $2,100.00 over the last 4 weeks.

A flagged day appears two places:

  • On the Cost Explorer chart, as an annotation on the day it happened — the same overlay used for deploys, incidents, and other timeline events.
  • On the Alerts page, in a list sorted worst-delta-first, so the account’s biggest mover is always at the top rather than buried under smaller flags.
  • As a notification, sent to whichever alert channels are subscribed to account-wide alerts — the same delivery path budget alerts and event-overage alerts use.

The same detection also runs over any tag you’ve defined rules for (see Virtual tagging & cost allocation) — one series per tag value, so “the platform team’s spend doubled” can be flagged and routed to that team, not just to whoever watches the account-wide feed. This can catch things a service-level check can’t: if a brand-new service appears inside a team’s spend, that service alone has no baseline and won’t flag on its own, but the team’s total jumping is still visible.

Turning on shared-cost redistribution or re-weighting a split rule never creates a false anomaly by itself — the same rule is applied to the current day and its whole baseline, so the ratio between them doesn’t move just because an allocation setting changed.

Because one real incident often shows up in two places — a service spike and the team spike it feeds — a tag-value alert that’s driven by a specific service names that service and, if the service was flagged too, marks it also flagged, so the two don’t read as unrelated notifications about the same money.

Some flagged spend is real and recurring — a monthly batch job, a predictable seasonal spike — and re-flagging it every time it happens is noise. Marking an anomaly as expected creates a suppression that matches on what the anomaly is about, not on the specific day it happened:

Anomaly type Suppression matches on
Cost source + service account, cost source, service
Tag value account, tag key, tag value

This is the important part: suppressing an anomaly doesn’t just dismiss the alert you’re looking at. It stops that same slice of spend from being flagged or notified again in the future, for as long as the suppression exists. A suppressed day still doesn’t get a chart annotation and still doesn’t send a notification — the suppression applies everywhere the anomaly would have shown up, not just the alert inbox.

The suppression is deliberately scoped to the dimensions above and nothing narrower — not the day, not the percentage, not the dollar amount. If it matched on those too, it would only ever apply to the exact anomaly you dismissed and would silently stop working the next time the same spike recurred. It’s also deliberately not scoped any broader — suppressing one service on one cost source doesn’t suppress anything else on that cost source, so a genuinely new problem elsewhere isn’t swallowed along with it.

A suppression can optionally carry an expiration date, but it doesn’t expire by default — an anomaly you’ve marked expected stays suppressed until you remove the suppression yourself. Removing it un-suppresses future occurrences; it doesn’t retroactively restore any alerts that were withheld while it was active.

Each suppression tracks how many anomalies it’s matched, shown as “muted N anomalies” next to the rule, so you can see whether it’s still doing anything.

Marking an anomaly as expected also closes any Jira issue or PagerDuty incident its alert opened. Nothing closes them automatically when spend returns to normal, because detection is a per-day check and doesn’t record when a spike ends.

Suppression covers the spend you’ve already judged. Auto-triage is the other end: an anomaly that Plutus can both explain and show a history for is filed away instead of notifying you, so the triage list is the anomalies that actually need a decision.

It’s off by default, on every account. Without the opt-in, every cost anomaly notifies every subscribed channel exactly as before — nothing about your alerts changes until you turn this on.

There’s no toggle in the app yet. It’s a per-account setting, anomaly_auto_triage_enabled, set through the account settings API by an account admin:

PUT /api/accounts/:accountId/settings

That endpoint replaces the whole settings object rather than merging into it, so read your current settings first (GET /api/accounts/:accountId/settings) and send them back with "anomaly_auto_triage_enabled": true added — otherwise you’ll clear other settings, like the account’s default display currency. Setting it back to false (or removing it) returns every anomaly to the ordinary list and to normal notification.

Every anomaly is scored 0–5 at detection time, whether or not the account has opted in, from three parts:

Part Range What it measures
Explainability 0–2 Whether something specific accounts for the spike. A single driver at 50% or more of it scores 2, a partial one scores 1, nothing named 0.
Recurrence 0–2 How often this exact signature has fired before. 3 or more previous times scores 2, one or two scores 1, first time 0.
Magnitude safety 0–1 The inverse of severity. A low-severity anomaly scores 1; medium and high score 0.

The tiers:

  • High (4–5, and explainability above 0) — eligible for auto-triage.
  • Medium (2–3) — shown and notified normally.
  • Low (0–1) — shown and notified normally.

Only the High tier changes anything. Medium and Low behave exactly as anomalies did before this existed.

Neither magnitude safety nor recurrence can carry an anomaly to High on its own: High requires a named driver. That’s the gate that keeps the promise honest — nothing is filed away without a stated reason, no matter how familiar or how small it is. Each anomaly carries its reasons in plain text, one per part, shown on the row.

Tag anomalies and service anomalies read different signals, and the two aren’t mixed.

  • Tag anomalies use the driver breakdown: which service on which cost source accounts for the tag’s increase. A top driver at 50% or more of the move scores 2. A smaller top driver scores 1. No named driver scores 0.
  • Service anomalies use event correlation. A service has no finer breakdown to name, so Plutus instead looks at deploy, incident, and release events for the whole account in the spike day and the day before. For each event family (one source plus one event type), it compares the number of events in that two-day window with the family’s daily rate over the days the cost baseline was taken from. Only the excess over that rate counts. A family that accounts for 50% or more of the window’s total excess scores 2, a smaller one scores 1, and nothing unusual scores 0.

Example: EC2 spikes on a Tuesday. The account normally has about 1 deploy a day from one repository, and there were 6 in the Monday–Tuesday window. That is 4 more than the baseline predicts, and nothing else in the window is unusual. The deploy family accounts for all of the excess, so explainability scores 2. The anomaly’s reasons show the counts: 6 events, against a baseline of 1 a day.

Correlation is temporal only. It shows that an unusual burst of events lines up with the spike, not that the events caused it, and the reason text says only what was counted. Two further limits:

  • A family that lines up with more than 3 of the same run’s service anomalies is capped at 1. An event that explains everything explains none of them in particular.
  • Only operational change events count, such as deploys, releases, and incidents. Business events (customer lifecycle, CRM) and news events don’t.

Correlation is scored for service anomalies when they’re detected. An anomaly recorded before this shipped shows “cross-event correlation was not measured” and scores 0.

A service anomaly can therefore reach High and be auto-triaged, but only when one unusual event family dominates its window. Most service anomalies score 0 or 1 here and stay in the normal list.

Only two things: whether the anomaly notifies, and where it’s listed.

  • Detection and recording are unchanged. The anomaly is still detected, still recorded, and still annotates the Cost Explorer chart on the day it happened.
  • No notification is sent for a High-confidence anomaly on an opted-in account.
  • It’s filed under a collapsed “Auto-triaged” section on the Alerts page, with a count badge, instead of the main triage list — the same pattern as muted signatures.

Suppression still comes first. A suppressed signature never reaches scoring at all, exactly as it never reaches recording. Auto-triage changes nothing about what suppression already does.

Each auto-triaged row has a Flag for review button, which moves that one anomaly back into the ordinary list. It applies to that single occurrence, not the signature: the next time the same slice spikes, it’s scored again from scratch.

Two things it deliberately doesn’t do. It doesn’t re-send the notification that was withheld — the moment for that alert has passed. And it doesn’t feed back into future scoring, so your triage clicks never quietly change how the next spike is scored. If a signature keeps auto-triaging and you don’t want it to, turn the setting off; there’s no per-signature override. Flagging for review needs the member role and is recorded in the account’s activity log.