Rollback, Recovery, and Incident-Generated Pester Tests

PesterForge · July 2026 · 7 min read

📦 Run it yourself — gate-policy examples are in companion/tests/Recovery/GatePolicy.Tests.ps1; the complete capstone adds rollback verification and an incident regression.

Most test strategies spend far more time proving installation than proving recovery.

That is understandable. Successful deployment is the plan. Rollback is the conversation nobody wants to have while the release meeting is going well.

Production does not care which conversation was more comfortable.

A recovery procedure that has never been executed and verified is not a control. It is a document describing something everyone hopes will work.

What you will learn

Rollback and recovery are not identical

Rollback

Returns code, configuration, or deployment artifacts to a prior version.

Recovery

Returns the service or workflow to an acceptable operating state.

Copying version 0.3.2 over version 0.4.0 is rollback.

Proving that version 0.3.2 loaded, the configuration is compatible, unfinished batches are reconciled, and processing can resume is recovery.

You need both.

The incident

A new Partner Feed Guardian release mistakenly accepts schema 2.1-preview as if it were the approved 2.1 schema.

One batch reaches production before the problem is detected.

The response must answer:

  1. How do we stop additional processing?
  2. What happened to the affected batch?
  3. Can the prior module safely read the current state?
  4. Can we restore the prior version?
  5. How do we prove the restored system is working?
  6. How do we prevent this exact defect from returning?

Test rollback prerequisites before deployment

Before releasing 0.4.0, verify:

Example:

Describe 'Rollback prerequisites' -Tag 'Gate.Recovery', 'Risk.ReadOnly' {
    It 'has the approved prior package' {
        $package = Get-PfgReleaseArtifact -Version '0.3.2'

        $package | Should-NotBeNull
        $package.Hash | Should-BeString $approvedPriorHash
    }

    It 'confirms the prior version can read the current state schema' {
        Test-PfgStateCompatibility `
            -ModuleVersion '0.3.2' `
            -StateSchemaVersion $currentStateSchema |
            Should-BeTrue
    }
}

If the prior version cannot understand state written by the new version, copying old code may make recovery worse.

Stop new work before changing recovery state

The first recovery operation may be to disable intake for the affected partner or rollout wave.

That action should be explicit, scoped, and verified:

Suspend-PfgPartnerIntake `
    -PartnerId 'NORTHWIND' `
    -Reason 'INC-2417 unsupported schema acceptance' `
    -Confirm:$false

Get-PfgPartnerIntakeState -PartnerId 'NORTHWIND' |
    Should-BeString 'Suspended'

Do not confuse stopping new work with resolving work already in flight.

Preserve evidence before cleanup

Keep:

Do not let an eager cleanup task erase the only evidence showing whether publication occurred.

Synthetic data and test records can have expiration policies, but incident preservation should override normal cleanup when required.

Verify the restored artifact

After deploying the prior package:

Describe 'Post-rollback module verification' -Tag 'Gate.Recovery' {
    It 'loads the restored module version' {
        Remove-Module PartnerFeedGuardian -ErrorAction SilentlyContinue
        Import-Module PartnerFeedGuardian -RequiredVersion 0.3.2

        (Get-Module PartnerFeedGuardian).Version.ToString() |
            Should-BeString '0.3.2'
    }

    It 'loads from the approved module path' {
        (Get-Module PartnerFeedGuardian).ModuleBase |
            Should-BeString $approvedRollbackPath
    }
}

Then execute the prior version’s known-good controlled workload. File presence is not recovery evidence.

Verify state, not only code

The affected batch may be in one of several states:

Rejected before publication
PublishedPendingAcknowledgement
Completed
Unknown after timeout
Partially published by a defective adapter

Recovery must reconcile the specific batch:

$remote = Get-PfgPublicationStatus `
    -BatchId $batchId `
    -CorrelationId $correlationId

$local = Get-PfgOperationState -BatchId $batchId

Compare-PfgOperationState -Local $local -Remote $remote

The response may be:

Pester can verify the chosen recovery rule against controlled cases.

Write the incident regression first

The defect was wildcard schema matching.

Capture it directly:

It 'rejects schema 2.1-preview when only 2.1 is approved' `
    -Tag 'Regression', 'INC-2417', 'Gate.PullRequest' {

    $batch = New-PfgBatch `
        -PartnerId 'NORTHWIND' `
        -BatchId 'INC-2417-REGRESSION' `
        -SchemaVersion '2.1-preview' `
        -EffectiveDate ([datetime]'2026-07-17') `
        -Payloads @('orders.json') `
        -ManifestPayloads @('orders.json')

    $result = Test-PfgBatchEligibility `
        -Batch $batch `
        -SupportedSchemas @('2.1') `
        -Now ([datetime]'2026-07-17T12:00:00Z')

    $result.Status | Should-Be 'Rejected'
    $result.Reasons |
        Should-ContainCollection @('UnsupportedSchema')
}

Run it against the defective version and watch it fail. That confirms the test reproduces the bug rather than simply agreeing with the fixed code.

Then fix the implementation and watch it pass.

Add the test at the cheapest effective layer

The incident happened in production, but the permanent regression test belongs in the fast pull-request gate because the defect is pure decision logic.

Do not require a production synthetic batch to rediscover a string-comparison bug on every release.

The incident may also reveal missing integration or deployment tests. Add those only for claims that truly require those layers.

Test the recovery procedure itself

A recovery exercise can run in a disposable or canary environment:

  1. Deploy current version.
  2. Create controlled in-flight state.
  3. Execute rollback.
  4. Load prior version.
  5. Reconcile the state.
  6. Complete or quarantine the controlled workload.
  7. Run the post-recovery suite.

Use Pester to evaluate each stage, while an orchestration script performs the sequence.

Do not hide the entire exercise inside one It block. Separate stages make failure location and evidence clearer.

Treat recovery tests as required evidence

A gate helper can reject incomplete recovery suites:

$result = Invoke-Pester -Configuration $recoveryConfig
Assert-PfgGateResult -Result $result -GateName 'Recovery exercise'

The companion helper rejects failed, skipped, inconclusive, not-run, failed-block, and failed-container counts.

A skipped rollback verification is not a successfully tested rollback.

Know when rollback is the wrong response

Rollback can be unsafe when the new version:

In those cases, roll-forward or a dedicated recovery tool may be safer.

The test plan should identify compatibility before deployment, not discover it during the incident.

Record recovery time honestly

A successful recovery exercise should capture:

Pester durations can contribute, but they are not the entire recovery time. Human approval, artifact retrieval, and provider reconciliation also count.

Review the test after the incident

Ask:

An incident should improve the testing model, not merely add one more It block.

What recovery tests prove

They can prove that the selected rollback and recovery procedures work in the controlled environment and that the restored system satisfies defined expectations.

What they do not prove

They do not guarantee every future incident is reversible or that production conditions will match the exercise. They convert recovery from an untested theory into practiced evidence.

Try it yourself

Choose one deployed PowerShell automation and document:

Prior version artifact:
Compatibility requirement:
Configuration backup:
Stop-new-work command:
Rollback command:
Post-rollback verification:
State reconciliation command:
Controlled recovery exercise:
Incident regression location:
Owner:

Then run the recovery exercise somewhere safe and record what did not work as written.

Common mistakes

Calling file restoration recovery. Verify the loaded code and the operating state.

Writing the regression test only after the fix. First prove the test fails against the defect.

Putting every incident regression in a production suite. Use the cheapest layer that can prove the behavior.

Rolling back across an incompatible data change. Test backward compatibility before release.

Cleaning up before preserving evidence. Recovery without facts can create a second incident.

Recap

Rollback restores an artifact. Recovery restores an acceptable operating state.

Test prerequisites before deployment, preserve evidence during incidents, verify both code and state after rollback, and turn each reproducible defect into a permanent test at the least expensive effective layer.

Next up: Part 11 — The Complete Partner Feed Guardian. We will assemble every suite, configuration, decision, owner, and failure action into one production testing strategy.


← All posts