Pester, Drift Detection, Monitoring, and Alerts
📦 Run it yourself — routing-drift examples are in
companion/tests/Drift/RoutingDrift.Tests.ps1.
Pester can run on a schedule.
That does not automatically make it a monitoring platform.
This distinction matters because scheduled tests often begin innocently:
We will run these ten checks every morning.
Then someone wants email. Then history. Then suppression. Then dashboards. Then escalation. Soon one PowerShell script is responsible for test execution, state storage, deduplication, routing, paging, maintenance windows, and the emotional well-being of everyone receiving its messages.
Pester is excellent at evaluating expectations. Let it do that job well.
What you will learn
- The difference between a test, health check, monitor, and alert
- When Pester is appropriate for drift detection
- How to return structured differences instead of one red result
- How to map different failures to different actions
- How to prevent flaky or unactionable alerts
- When results should flow into another operational system
Four related jobs
Test
Evaluates a claim under defined conditions.
Given this approved routing configuration and this production configuration, do the required properties match?
Health check
Evaluates whether a system can perform a bounded operation now.
Can the synthetic partner batch complete and reconcile?
Monitoring
Collects and retains observations over time.
How often do partner batches fail, how long do they take, and is the trend worsening?
Alerting
Routes selected conditions to people or systems based on severity, ownership, suppression, and escalation rules.
The NORTHWIND route now points to a different destination; suspend intake and notify the partner platform owner.
Pester can produce the first two and supply data to the last two. It should not have to become all four.
The routing-drift example
The approved configuration is version-controlled:
$approved = @(
[pscustomobject]@{
PartnerId = 'NORTHWIND'
SchemaVersion = '2.1'
Destination = 'sales-west'
TransformRuleSet = 'northwind-v4'
QuarantinePolicy = 'hold-30-days'
AcknowledgementMethod = 'api'
MaximumBatchSize = 5000
}
)Production is queried separately:
$actual = Get-PfgProductionRoutingConfigurationThe comparison function returns differences, not just $true or $false:
$differences = Compare-PfgRoutingConfiguration `
-Approved $approved `
-Actual $actualA useful difference contains:
PartnerId : NORTHWIND
Difference: Destination
Expected : sales-west
Actual : sales-east
That is actionable evidence.
Test no-drift and focused-drift behavior
Describe 'Partner routing drift' `
-Tag 'Environment.Production', 'Risk.ReadOnly', 'Gate.Drift' {
It 'returns no differences for matching configuration' {
$actual = $approved | ForEach-Object { $_.PSObject.Copy() }
$result = Compare-PfgRoutingConfiguration `
-Approved $approved `
-Actual $actual
$result | Should-BeCollection @()
}
It 'returns the changed route destination' {
$actual = $approved | ForEach-Object { $_.PSObject.Copy() }
$actual[0].Destination = 'sales-east'
$result = Compare-PfgRoutingConfiguration `
-Approved $approved `
-Actual $actual
$result.Count | Should-Be 1
$result[0] | Should-BeEquivalent ([pscustomobject]@{
PartnerId = 'NORTHWIND'
Difference = 'Destination'
Expected = 'sales-west'
Actual = 'sales-east'
})
}
}These are tests of the comparison logic. A scheduled production run then uses the real approved and actual sources.
Not all drift has the same severity
A changed destination may risk data disclosure or loss. A lower batch-size limit may slow processing without misrouting data. An acknowledgement-method change may be expected during a rollout.
Map the difference to policy:
| Difference | Suggested response |
|---|---|
| Destination | Suspend that partner and escalate immediately |
| Transform rule set | Block the next batch; urgent ticket |
| Schema version | Block incompatible batches; notify owner |
| Maximum batch size | Create ticket; warn before the next oversized batch |
| Unexpected partner | Investigate before accepting any data |
| Missing disabled partner | Record evidence; normal-priority review |
The Pester test identifies the difference. A policy layer determines the action.
Export structured results
Do not make downstream systems parse console symbols.
Use the Pester result object and a separate drift payload:
$pesterResult = Invoke-Pester -Configuration $config
$driftRecord = [pscustomobject]@{
CheckedAtUtc = [datetime]::UtcNow
Environment = 'Production'
CodeVersion = $releaseVersion
PesterVersion = $pesterResult.Version
Differences = $differences
TestSummary = [pscustomobject]@{
Passed = $pesterResult.PassedCount
Failed = $pesterResult.FailedCount
Skipped = $pesterResult.SkippedCount
Inconclusive = $pesterResult.InconclusiveCount
}
}
$driftRecord | ConvertTo-Json -Depth 8 |
Set-Content -Path $outputPathAnother system can retain, route, suppress, correlate, or visualize the result.
Pester should not own long-term history
A test-result XML file is useful evidence. A directory containing 18 months of XML files is not automatically an operational data platform.
Use an appropriate store for:
- Trend queries
- Retention policy
- Deduplication
- Dashboards
- Alert state
- Maintenance suppression
- Acknowledgement and escalation
Pester’s job is to evaluate the expectation consistently.
Make every alert actionable
Before routing a failure, answer:
Who owns it?
What should they do?
How urgent is it?
What evidence is included?
What condition clears it?
What maintenance or rollout can suppress it?
What happens if it repeats?
Avoid alerts such as:
Pester test failed.
Prefer:
NORTHWIND production routing destination differs from the approved configuration.
Expected: sales-west
Actual: sales-east
Automatic action: NORTHWIND intake suspended
Owner: Partner Platform
Correlation: DRIFT-20260717-0830
Keep paging rare
A page should represent something that requires prompt human action.
Do not page because:
- A noncritical diagnostic suite could not run
- A test was flaky once
- An advisory configuration field changed
- A maintenance window intentionally changed state
- The same known failure repeated every five minutes
Scheduled Pester results can create tickets, warnings, or evidence without paging.
Treat flakiness as a defect
A test that fails randomly teaches people to ignore it.
Common causes include:
- Eventual consistency
- Timing assumptions
- Shared test state
- Provider throttling
- Network instability
- Uncontrolled dates
- Parallel execution collisions
Do not leave a flaky production check permanently enabled with the explanation “just rerun it.”
A responsible quarantine process includes:
- Named owner
- Defect or issue
- Date quarantined
- Risk assessment
- Replacement evidence, if needed
- Repair deadline
A required safety check cannot simply disappear because it is inconvenient.
Distinguish system failure from test failure
A drift suite can fail because:
- Production differs from approval
- The production query failed
- The approved configuration could not be loaded
- The test file did not discover
- The comparison code threw
These need different routing.
Use the Pester result object’s failed tests, failed blocks, and failed containers to classify the run. A failed BeforeAll is not the same as an assertion that found drift.
Schedule at the speed of the risk
Not every configuration needs hourly checks.
Choose frequency based on:
- How quickly the value can change
- How much harm delay creates
- Whether changes are already event-driven
- Provider cost and rate limits
- Noise and ownership capacity
For routing that can change only through a controlled deployment, post-deployment verification plus a daily drift check may be enough. For a critical dynamic route, faster detection may be justified.
Retire tests that no longer protect a risk
Scheduled suites accumulate historical checks long after the system changes.
Review regularly:
- Does this condition still exist?
- Does the failure still require action?
- Is another control now authoritative?
- Is the owner still correct?
- Has the test become a duplicate of monitoring?
More tests are not always more assurance. Unowned and outdated tests dilute trust.
What the drift suite proves
It can prove that, at the time of the query, the selected production configuration matched or differed from the approved configuration according to the comparison rules.
What it does not prove
It does not provide historical trends, guarantee the configuration did not change immediately afterward, or manage the full alert lifecycle. Those responsibilities belong to the systems consuming the result.
Try it yourself
Choose one version-controlled production configuration and build:
- A function that returns structured differences
- Unit tests for matching, missing, unexpected, and changed items
- A read-only production query
- A severity map for each difference type
- A structured result export
- An owner and failure action
Then decide whether the result belongs in a ticket, dashboard, alert, or page.
Common mistakes
Calling every scheduled Pester suite monitoring. Pester evaluates expectations; monitoring manages observations over time.
Sending console output as an alert. Export structured, contextual evidence.
Paging on every failure. Match the response to urgency and actionability.
Ignoring flaky tests. Flakiness is an operational defect in the evidence system.
Never retiring old checks. Tests that no longer protect a real risk reduce confidence in the suite.
Recap
Pester is a strong engine for drift checks and bounded operational validation. It is not automatically a history store, dashboard, alert router, or incident-management platform.
Keep responsibilities clear: Pester evaluates; operational systems retain, route, suppress, escalate, and visualize.
Next up: Part 10 — Rollback, Recovery, and Incident-Generated Tests. We will prove that the system can return to a known state and turn a real production failure into a permanent test.