7. Require guardrails
About 1400 wordsAbout 5 min
2026-09-30
At the end of step 5, sigil check warned that payments invokes the platform's page rules under a when. Any team can do that, including one that writes when false { routing() } and never pages anyone again. In this step the platform moves the rules every team must get into a policy of their own, and your program refuses to load a team policy that doesn't invoke it unconditionally.
Split out the page rules
Pages are the rules that must always apply: a critical production alert pages the on-call, and so does a warning that won't go away. Create platform/paging.sigil with both, and the threshold as a param with bounds on both sides:
policy platform.paging: AlertRouting@1
use platform.alerts.{pre_production}
param page_after: duration = 30m, min: 5m, max: 1h
when not pre_production and alert.severity == critical {
page(reason: critical_alert, target: team.oncall)
}
when not pre_production and alert.severity == warning and alert.firing_for >= page_after {
page(reason: sustained, target: team.oncall)
}The max matters as much as the min: without it, paging(page_after: 1000h) would switch sustained paging off as surely as a when false.
What's left in platform/routing.sigil is the part teams may shape freely:
policy platform.routing: AlertRouting@1
use platform.alerts.{pre_production}
param muted: list<string> = []
when not pre_production and alert.severity == warning {
notify(reason: routine, channel: team.channel)
}
when pre_production {
drop(reason: not_production)
}
when alert.name in muted {
drop(reason: muted)
}Require it
Tell sigil check that every team policy must invoke platform.paging:
$ sigil check --require platform.paging
checkout/alerts.sigil:5:9: error: policy platform.routing has no param `page_after`
|
5 | routing(page_after: 10m, muted: ["CheckoutCanaryLatency"])
| ^^^^^^^^^^
= help: platform.routing declares: muted
payments/alerts.sigil:7:11: error: policy platform.routing has no param `page_after`
|
7 | routing(page_after: 5m)
| ^^^^^^^^^^
= help: platform.routing declares: muted
✗ checked 6 files, 2 errorsThe threshold moved, so both team policies need updating. Give checkout/alerts.sigil both calls:
policy checkout.alerts: AlertRouting@1
use platform.paging
use platform.routing
paging(page_after: 10m)
routing(muted: ["CheckoutCanaryLatency"])And payments/alerts.sigil:
policy payments.alerts: AlertRouting@1
use platform.alerts.{pre_production}
use platform.paging
use platform.routing
paging(page_after: 5m)
routing()
when not pre_production and alert.severity == info and alert.labels["component"] == "ledger" {
notify(reason: routine, channel: "#payments-ledger")
}$ sigil check --require platform.paging
✓ checked 6 files, no problems foundThe gated-deny warnings are gone too: routing no longer holds page rules, so gating it can't silence a page.
Payments lost something in the move: its ledger-only threshold. A required policy is invoked once, at the top level, so each team gets one page_after. That's the trade: the platform can now promise that every team pages for sustained warnings within an hour.
Try to get around it
Put the payments call back under a condition:
when alert.labels["component"] == "ledger" {
paging(page_after: 5m)
}$ sigil check --require platform.paging
payments/alerts.sigil:8:3: error: platform.paging must be invoked unconditionally
|
8 | paging(page_after: 5m)
| ^^^^^^^^^^^^^^^^^^^^^^
= help: the host requires platform.paging for every AlertRouting policy; move the call to the top level
✗ checked 6 files, 1 errorLeave it out entirely, and the error names the policy that's missing:
$ sigil check --require platform.paging
checkout/alerts.sigil:1:1: error: checkout.alerts doesn't invoke platform.paging
|
1 | policy checkout.alerts: AlertRouting@1
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
= help: the host requires platform.paging for every AlertRouting policy; import it with `use platform.paging` and invoke it at the top levelAnd a threshold outside the bounds:
payments/alerts.sigil:8:22: error: page_after: 5h is above the maximum 1h
|
8 | paging(page_after: 5h)
| ^^
= help: platform.paging declares `param page_after: duration = 30m, min: 5m, max: 1h`Put everything back the way it was before moving on.
Make it stick
Typing --require on every run is easy to forget. Put the requirement in policies/sigil.yaml, which check, eval, explain and test read from the directory they run in:
require:
- policy: platform.paging
trusted: [platform]
roots: ["*.alerts"]trusted says platform.paging must come from platform/, so a team can't satisfy the requirement with a platform.paging of its own. roots names the policies it applies to.
The check that counts is the one in your program, because that's what serves alerts. Change the Load call in cmd/route/main.go:
p, err := alerting.Kind.Load(alerting.Policies, "checkout.alerts",
policy.Require("platform.paging"))Now a team policy that skips the call doesn't load:
$ go run ./cmd/route
2026/09/30 17:31:40 policies/checkout/alerts.sigil:1:1: error: checkout.alerts doesn't invoke platform.paging
|
1 | policy checkout.alerts: AlertRouting@1
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
= help: the host requires platform.paging for every AlertRouting policy; import it with `use platform.paging` and invoke it at the top level
exit status 1With the call back in place:
$ go run ./cmd/route
CheckoutErrorRate critical production 2m0s → page checkout-primary (critical_alert)
CheckoutLatencyHigh warning production 12m0s → page checkout-primary (sustained)
CheckoutLatencyHigh warning production 45m0s → page checkout-primary (sustained)
CheckoutErrorRate critical staging 2m0s → drop (not_production)
CheckoutQueueStuck critical - 3m0s → page checkout-primary (critical_alert)
CheckoutCanaryLatency warning production 5m0s → drop (muted)The 12-minute warning pages now, because checkout's page_after is 10 minutes. A page outranks every other decision, so nothing a team adds can beat a page from platform.paging.
What a guardrail can't stop
A team can't remove the platform's page, but it has two other ways to keep it from being the answer.
It can outrank it. Reasons are ranked within a decision, critical_alert above sustained, so a team page with the higher reason beats the platform's sustained page without any conflict. Add this to checkout/alerts.sigil:
when alert.severity == warning {
page(reason: critical_alert, target: "nobody")
}sigil check --require platform.paging still passes, and the 45-minute warning from step 4 now pages nobody:
$ sigil eval --policy checkout.alerts --input checkout/testdata/latency.json
checkout.alerts: page(reason: critical_alert)
target = "nobody"
trace: 3 candidates
* page(reason: critical_alert) checkout/alerts.sigil:10:3
when alert.severity == warning
target = "nobody"
page(reason: sustained) checkout/alerts.sigil:6:1 → platform/paging.sigil:12:3
when not pre_production and alert.severity == warning and alert.firing_for >= 10m
target = "checkout-primary"
notify(reason: routine) checkout/alerts.sigil:7:1 → platform/routing.sigil:8:3
channel = "#checkout-alerts"It can also make the evaluation fail. Change the rule to page for critical alerts instead, with the same reason as the platform's:
when alert.severity == critical {
page(reason: critical_alert, target: "nobody")
}Save a critical production alert as critical.json, the way you saved latency.json in step 4, and evaluate it:
$ sigil eval --policy checkout.alerts --input critical.json
checkout.alerts: the candidates conflict, the host falls back to notify(reason: unrouted), the kind's default
conflict: collect one: 2 candidates at the top rank
page(reason: critical_alert) checkout/alerts.sigil:6:1 → platform/paging.sigil:8:3
page(reason: critical_alert) checkout/alerts.sigil:10:3
= help: a conflict is a defect in the policy: rank the reasons with precedence, or keep the exclusive outcomes' conditions apartTwo pages with the same reason and different targets can't both win, so the evaluation fails and Eval returns an error along with the kind's default. In this kind the default is a post to #alerts, not a page. A failing assert or a timeout has the same effect.
Neither case is visible to sigil check, because both depend on the input. A service that must page whatever teams write evaluates platform.paging on its own as well, and uses its page when the team's result doesn't match it or fails. Handle failed evaluations covers the error, and What the guarantee doesn't cover explains why composition can't prevent either case.
Remove the rule again before moving on.
What you built
A Go service that routes alerts through policies it checks against a typed contract; a kind file that lets anyone check, evaluate and test those policies with the CLI; a shared library that teams use with their own values; and guardrails that no team can leave out.
The same router, grown into a service, is examples/alert-routing. It runs the policies you wrote here in a TypeScript app on Sigil's WebAssembly build, with an HTTP API, an operator console, hot reload, metrics, traces, a Grafana dashboard and load tests. From here:
- Per-team policies and Policies in a ConfigMap take the same ideas into production, including loading team policies from a directory the platform doesn't control, where
policy.Frommakes the guardrails come from the platform's own copy. - Composition without templating explains what guarantees composition gives and which it doesn't.
- The language reference is the precise version of everything you used.
