Rollback, Recovery, and Incident-Generated Pester Tests
📦 Run it yourself — gate-policy examples are in
companion/tests/Recovery/GatePolicy.Tests.ps1; the complete capstone adds rollback verification and an incident regression.
Most test strategies spend far more time proving installation than proving recovery.
That is understandable. Successful deployment is the plan. Rollback is the conversation nobody wants to have while the release meeting is going well.
Production does not care which conversation was more comfortable.
A recovery procedure that has never been executed and verified is not a control. It is a document describing something everyone hopes will work.
What you will learn
- What to test before rollback is needed
- How rollback differs from recovery
- How to verify restored code and restored state
- How to preserve incident evidence
- How to turn a production failure into a regression test
- When a rollback is unsafe because the new version changed data
Rollback and recovery are not identical
Rollback
Returns code, configuration, or deployment artifacts to a prior version.
Recovery
Returns the service or workflow to an acceptable operating state.
Copying version 0.3.2 over version 0.4.0 is rollback.
Proving that version 0.3.2 loaded, the configuration is compatible, unfinished batches are reconciled, and processing can resume is recovery.
You need both.
The incident
A new Partner Feed Guardian release mistakenly accepts schema 2.1-preview as if it were the approved 2.1 schema.
One batch reaches production before the problem is detected.
The response must answer:
- How do we stop additional processing?
- What happened to the affected batch?
- Can the prior module safely read the current state?
- Can we restore the prior version?
- How do we prove the restored system is working?
- How do we prevent this exact defect from returning?
Test rollback prerequisites before deployment
Before releasing 0.4.0, verify:
- The prior
0.3.2package is available - Its hash or signature is approved
- The current configuration is exported
- The prior version supports the current state-store schema
- The rollback command parses and runs in a controlled environment
- Required credentials remain available
Example:
Describe 'Rollback prerequisites' -Tag 'Gate.Recovery', 'Risk.ReadOnly' {
It 'has the approved prior package' {
$package = Get-PfgReleaseArtifact -Version '0.3.2'
$package | Should-NotBeNull
$package.Hash | Should-BeString $approvedPriorHash
}
It 'confirms the prior version can read the current state schema' {
Test-PfgStateCompatibility `
-ModuleVersion '0.3.2' `
-StateSchemaVersion $currentStateSchema |
Should-BeTrue
}
}If the prior version cannot understand state written by the new version, copying old code may make recovery worse.
Stop new work before changing recovery state
The first recovery operation may be to disable intake for the affected partner or rollout wave.
That action should be explicit, scoped, and verified:
Suspend-PfgPartnerIntake `
-PartnerId 'NORTHWIND' `
-Reason 'INC-2417 unsupported schema acceptance' `
-Confirm:$false
Get-PfgPartnerIntakeState -PartnerId 'NORTHWIND' |
Should-BeString 'Suspended'Do not confuse stopping new work with resolving work already in flight.
Preserve evidence before cleanup
Keep:
- Batch ID
- Correlation ID
- Original manifest
- Payload hashes
- Module version
- Configuration version
- Execution identity
- Pester result files
- Durable operation state
- Provider responses
- Relevant timestamps
Do not let an eager cleanup task erase the only evidence showing whether publication occurred.
Synthetic data and test records can have expiration policies, but incident preservation should override normal cleanup when required.
Verify the restored artifact
After deploying the prior package:
Describe 'Post-rollback module verification' -Tag 'Gate.Recovery' {
It 'loads the restored module version' {
Remove-Module PartnerFeedGuardian -ErrorAction SilentlyContinue
Import-Module PartnerFeedGuardian -RequiredVersion 0.3.2
(Get-Module PartnerFeedGuardian).Version.ToString() |
Should-BeString '0.3.2'
}
It 'loads from the approved module path' {
(Get-Module PartnerFeedGuardian).ModuleBase |
Should-BeString $approvedRollbackPath
}
}Then execute the prior version’s known-good controlled workload. File presence is not recovery evidence.
Verify state, not only code
The affected batch may be in one of several states:
Rejected before publication
PublishedPendingAcknowledgement
Completed
Unknown after timeout
Partially published by a defective adapter
Recovery must reconcile the specific batch:
$remote = Get-PfgPublicationStatus `
-BatchId $batchId `
-CorrelationId $correlationId
$local = Get-PfgOperationState -BatchId $batchId
Compare-PfgOperationState -Local $local -Remote $remoteThe response may be:
- Quarantine and do not reprocess
- Complete acknowledgement only
- Reverse or compensate the publication
- Correct and reprocess under a new batch ID
- Escalate because the outcome remains unknown
Pester can verify the chosen recovery rule against controlled cases.
Write the incident regression first
The defect was wildcard schema matching.
Capture it directly:
It 'rejects schema 2.1-preview when only 2.1 is approved' `
-Tag 'Regression', 'INC-2417', 'Gate.PullRequest' {
$batch = New-PfgBatch `
-PartnerId 'NORTHWIND' `
-BatchId 'INC-2417-REGRESSION' `
-SchemaVersion '2.1-preview' `
-EffectiveDate ([datetime]'2026-07-17') `
-Payloads @('orders.json') `
-ManifestPayloads @('orders.json')
$result = Test-PfgBatchEligibility `
-Batch $batch `
-SupportedSchemas @('2.1') `
-Now ([datetime]'2026-07-17T12:00:00Z')
$result.Status | Should-Be 'Rejected'
$result.Reasons |
Should-ContainCollection @('UnsupportedSchema')
}Run it against the defective version and watch it fail. That confirms the test reproduces the bug rather than simply agreeing with the fixed code.
Then fix the implementation and watch it pass.
Add the test at the cheapest effective layer
The incident happened in production, but the permanent regression test belongs in the fast pull-request gate because the defect is pure decision logic.
Do not require a production synthetic batch to rediscover a string-comparison bug on every release.
The incident may also reveal missing integration or deployment tests. Add those only for claims that truly require those layers.
Test the recovery procedure itself
A recovery exercise can run in a disposable or canary environment:
- Deploy current version.
- Create controlled in-flight state.
- Execute rollback.
- Load prior version.
- Reconcile the state.
- Complete or quarantine the controlled workload.
- Run the post-recovery suite.
Use Pester to evaluate each stage, while an orchestration script performs the sequence.
Do not hide the entire exercise inside one It block. Separate stages make failure location and evidence clearer.
Treat recovery tests as required evidence
A gate helper can reject incomplete recovery suites:
$result = Invoke-Pester -Configuration $recoveryConfig
Assert-PfgGateResult -Result $result -GateName 'Recovery exercise'The companion helper rejects failed, skipped, inconclusive, not-run, failed-block, and failed-container counts.
A skipped rollback verification is not a successfully tested rollback.
Know when rollback is the wrong response
Rollback can be unsafe when the new version:
- Migrated data to a format the old version cannot read
- Changed an external contract
- Rotated credentials
- Published irreversible messages
- Updated provider-side configuration
- Removed information needed by the prior version
In those cases, roll-forward or a dedicated recovery tool may be safer.
The test plan should identify compatibility before deployment, not discover it during the incident.
Record recovery time honestly
A successful recovery exercise should capture:
- Time to detect
- Time to stop new work
- Time to restore code
- Time to reconcile state
- Time to verify service
- Manual decisions required
Pester durations can contribute, but they are not the entire recovery time. Human approval, artifact retrieval, and provider reconciliation also count.
Review the test after the incident
Ask:
- Which earlier test could have caught this?
- Did a test exist but fail to block the release?
- Was required evidence skipped?
- Did monitoring detect the issue promptly?
- Did the recovery instructions match reality?
- Which test created noise without helping?
An incident should improve the testing model, not merely add one more It block.
What recovery tests prove
They can prove that the selected rollback and recovery procedures work in the controlled environment and that the restored system satisfies defined expectations.
What they do not prove
They do not guarantee every future incident is reversible or that production conditions will match the exercise. They convert recovery from an untested theory into practiced evidence.
Try it yourself
Choose one deployed PowerShell automation and document:
Prior version artifact:
Compatibility requirement:
Configuration backup:
Stop-new-work command:
Rollback command:
Post-rollback verification:
State reconciliation command:
Controlled recovery exercise:
Incident regression location:
Owner:
Then run the recovery exercise somewhere safe and record what did not work as written.
Common mistakes
Calling file restoration recovery. Verify the loaded code and the operating state.
Writing the regression test only after the fix. First prove the test fails against the defect.
Putting every incident regression in a production suite. Use the cheapest layer that can prove the behavior.
Rolling back across an incompatible data change. Test backward compatibility before release.
Cleaning up before preserving evidence. Recovery without facts can create a second incident.
Recap
Rollback restores an artifact. Recovery restores an acceptable operating state.
Test prerequisites before deployment, preserve evidence during incidents, verify both code and state after rollback, and turn each reproducible defect into a permanent test at the least expensive effective layer.
Next up: Part 11 — The Complete Partner Feed Guardian. We will assemble every suite, configuration, decision, owner, and failure action into one production testing strategy.