State Changes, Partial Failure, and Safe Repetition
📦 Run it yourself — this post’s examples are in
companion/tests/Unit/Processing.Tests.ps1.
The cleanest test scenario is also one of the least realistic:
- The operation begins.
- Every dependency succeeds.
- The operation ends.
Production prefers more interesting stories.
Publication succeeds, but acknowledgement fails. A timeout occurs after the remote system accepted the request. The state record updates, but the evidence write does not. Four items succeed and the fifth does not. The job is restarted before anyone knows where it stopped.
The important question is not simply:
Did the function throw?
It is:
What state did the function leave behind, and what happens when it runs again?
What you will learn
- Why retries are dangerous after state changes
- The difference between idempotence and “run it twice and hope”
- How to preserve partial progress explicitly
- How to test unknown and partial outcomes
- Why correlation IDs and durable state matter
- How to prove that acknowledgement retries do not republish data
The failure that matters
The Partner Feed Guardian performs two major external actions:
- Publish the validated batch.
- Send the partner acknowledgement.
Suppose publication succeeds and acknowledgement fails.
The worst possible retry logic is:
try {
Publish-PfgBatch -Batch $Batch
Send-PfgAcknowledgement -Batch $Batch
}
catch {
Start-Sleep -Seconds 10
Publish-PfgBatch -Batch $Batch
Send-PfgAcknowledgement -Batch $Batch
}This treats the pair as if neither action occurred. The second publication may duplicate data.
Record progress at meaningful boundaries
The operation state should distinguish:
NotStarted
Publishing
PublishedPendingAcknowledgement
CompletedPendingEvidence
Completed
Failed
PublishedPendingAcknowledgement is not an error-message detail. It changes the safe next action.
When the operation reruns:
NotStartedmeans publish, then acknowledge.PublishedPendingAcknowledgementmeans do not publish again; retry acknowledgement.CompletedPendingEvidencemeans publication and acknowledgement succeeded; repair evidence without repeating either external action.Completedmeans return the prior success state.Publishingmay represent an unknown outcome requiring reconciliation.Failedrequires examining the recorded failure and recovery policy.
The processing function
The companion module uses the durable state to decide where to resume:
$state = Get-PfgOperationState -BatchId $Batch.BatchId
if ($state -eq 'Completed') {
return [pscustomobject]@{
BatchId = $Batch.BatchId
Status = 'AlreadyCompleted'
Published = $true
Acknowledged = $true
}
}
$published = $state -in @(
'PublishedPendingAcknowledgement',
'CompletedPendingEvidence',
'Completed'
)
$acknowledged = $state -in @(
'CompletedPendingEvidence',
'Completed'
)
if (-not $published) {
Set-PfgOperationState -BatchId $Batch.BatchId -State 'Publishing'
Publish-PfgBatch -Batch $Batch -CorrelationId $CorrelationId
Set-PfgOperationState `
-BatchId $Batch.BatchId `
-State 'PublishedPendingAcknowledgement'
$published = $true
}
if (-not $acknowledged) {
Send-PfgAcknowledgement -Batch $Batch -CorrelationId $CorrelationId
Set-PfgOperationState `
-BatchId $Batch.BatchId `
-State 'CompletedPendingEvidence'
$acknowledged = $true
}
Write-PfgEvidence -Record $evidence
Set-PfgOperationState -BatchId $Batch.BatchId -State 'Completed'The state records the last known safe boundary.
Test that the retry skips publication
It 'does not publish a second time when publication already succeeded' {
Mock -ModuleName PartnerFeedGuardian Get-PfgOperationState {
'PublishedPendingAcknowledgement'
}
Mock -ModuleName PartnerFeedGuardian Set-PfgOperationState { }
Mock -ModuleName PartnerFeedGuardian Publish-PfgBatch { }
Mock -ModuleName PartnerFeedGuardian Send-PfgAcknowledgement { }
Mock -ModuleName PartnerFeedGuardian Write-PfgEvidence { }
$result = Invoke-PfgBatch `
-Batch $batch `
-Eligibility $approved `
-Confirm:$false
$result.Status | Should-Be 'Completed'
Should-NotInvoke -ModuleName PartnerFeedGuardian Publish-PfgBatch
Should-Invoke `
-ModuleName PartnerFeedGuardian `
Send-PfgAcknowledgement `
-Times 1 `
-Exactly
}This test protects a production guarantee:
A retry after confirmed publication does not publish the batch again.
Test the partial failure result
It 'reports publication as complete when acknowledgement fails' {
Mock -ModuleName PartnerFeedGuardian Get-PfgOperationState { 'NotStarted' }
Mock -ModuleName PartnerFeedGuardian Set-PfgOperationState { }
Mock -ModuleName PartnerFeedGuardian Publish-PfgBatch { }
Mock -ModuleName PartnerFeedGuardian Send-PfgAcknowledgement {
throw 'Partner acknowledgement API unavailable'
}
Mock -ModuleName PartnerFeedGuardian Write-PfgEvidence { }
$result = Invoke-PfgBatch `
-Batch $batch `
-Eligibility $approved `
-Confirm:$false
$result.Status | Should-Be 'Failed'
$result.Published | Should-BeTrue
$result.Acknowledged | Should-BeFalse
$result.Error |
Should-BeString 'Partner acknowledgement API unavailable'
Should-Invoke `
-ModuleName PartnerFeedGuardian `
Set-PfgOperationState `
-ParameterFilter {
$State -eq 'PublishedPendingAcknowledgement'
}
}The returned result and durable state must agree. A generic Failed state would discard the most important recovery fact: publication already happened.
Idempotence is specific to an operation
“Make it idempotent” is good advice that can become vague very quickly.
An operation is idempotent when repeating it produces the same intended final state without compounding the effect.
Examples:
- Setting a route destination to
sales-westcan be idempotent. - Incrementing a retry counter is not.
- Sending an email is not naturally idempotent.
- Publishing a batch may be idempotent only if the receiver honors a unique batch key.
PowerShell cannot declare an arbitrary API idempotent by believing in it very hard.
The system needs one or more mechanisms such as:
- A provider-supported idempotency key
- A unique batch identifier enforced by the receiver
- A durable local operation record
- A reconciliation query before retry
- A compensating action
Tests should prove the mechanism, not merely the word.
Use correlation IDs everywhere
A correlation ID should connect:
- The PowerShell execution
- The durable state record
- The publication request
- The acknowledgement request
- The evidence record
- The incident or ticket
$correlationId = [guid]::NewGuid().Guid
Publish-PfgBatch `
-Batch $Batch `
-CorrelationId $correlationIdA test can assert that the same identifier reaches every dependency:
It 'uses one correlation id across publication and acknowledgement' {
$correlationId = 'TEST-CORRELATION-1042'
Invoke-PfgBatch `
-Batch $batch `
-Eligibility $approved `
-CorrelationId $correlationId `
-Confirm:$false
Should-Invoke -ModuleName PartnerFeedGuardian Publish-PfgBatch `
-ParameterFilter { $CorrelationId -eq 'TEST-CORRELATION-1042' }
Should-Invoke -ModuleName PartnerFeedGuardian Send-PfgAcknowledgement `
-ParameterFilter { $CorrelationId -eq 'TEST-CORRELATION-1042' }
}Without correlation, partial failures become a guessing exercise across several systems.
Test an unknown outcome separately
A timeout does not always mean failure.
Suppose Publish-PfgBatch sends the request, the receiver commits it, and the connection drops before the response reaches PowerShell.
Locally, the last state is Publishing. The actual remote state is unknown.
Do not automatically republish.
The recovery path should query by batch ID or correlation ID:
$remote = Get-PfgPublicationStatus `
-BatchId $Batch.BatchId `
-CorrelationId $CorrelationId
switch ($remote.Status) {
'Published' {
Set-PfgOperationState `
-BatchId $Batch.BatchId `
-State 'PublishedPendingAcknowledgement'
}
'NotFound' {
# Safe to retry publication under the documented provider rules.
}
default {
# Quarantine for investigation. Do not guess.
}
}Create tests for all three outcomes. “Unknown” is a legitimate state and often safer than inventing certainty.
Test failure at every boundary
For a state-changing workflow, inject failure before and after each durable boundary:
| Failure point | Expected safe result |
|---|---|
| Before publication | No publication; state remains retryable |
| During publication with known rejection | State records failure; no acknowledgement |
| During publication with unknown result | Reconcile before retry |
| After publication, before state update | Reconcile by correlation ID |
| During acknowledgement | Do not republish; retry acknowledgement |
| During evidence write | Preserve completed external state; repair evidence separately |
This table is more useful than a single test named handles errors.
Cleanup failures are first-class failures
If a synthetic or integration operation creates test state and cleanup fails, the test should report that clearly.
Do not hide it in AfterAll -ErrorAction SilentlyContinue.
A cleanup failure may not mean the product is broken, but it does mean the test changed the environment and did not restore it.
Return or record:
- Resource identifier
- Correlation ID
- Cleanup command
- Owner
- Expiration time
Retrying the test is not always retrying the operation
Pester rerun behavior should not accidentally rerun unsafe setup.
Keep state-changing calls inside clearly tagged tests and require deliberate configuration. A test adapter that creates synthetic state should use unique IDs and recovery logic.
For production-facing tests, avoid automatic “retry the whole failed test three times” wrappers. Retry a known transient dependency at the smallest safe boundary. Repeating the whole It block may repeat successful state changes.
What these tests prove
They can prove that, for known recorded states and simulated failures:
- Publication happens before acknowledgement
- Acknowledgement does not occur after a failed publication
- A confirmed publication is not repeated
- Partial progress is returned and recorded
- Correlation IDs remain consistent
- Completed external work is not repeated when only evidence remains
- Completed work is recognized on rerun
What they do not prove
They do not prove the real receiver enforces uniqueness, the remote status query is accurate, or every failure mode has been modeled. Integration and production tests must verify the provider behavior.
Try it yourself
For one state-changing automation, draw its state machine.
At minimum, answer:
What state exists before the first change?
What durable boundary exists after each irreversible step?
What does a timeout mean at each boundary?
Which steps are safe to repeat?
Which steps require reconciliation?
What identifier connects retries to the original attempt?
Then write one Pester test proving that a retry skips an already-completed irreversible step.
Common mistakes
Retrying the whole workflow after any exception. Some earlier steps may already have succeeded.
Recording only
SucceededorFailed. Partial states often determine the only safe recovery action.
Calling an operation idempotent without a mechanism. Prove uniqueness, reconciliation, or provider idempotency.
Treating timeout as rejection. The remote result may be unknown, not failed.
Hiding cleanup failure. Test-created state must remain identifiable and recoverable.
Recap
Production testing must examine the state left behind after failure.
Safe repetition requires durable progress, correlation IDs, explicit partial states, and provider-aware reconciliation. A retry should resume from the last confirmed boundary—not replay the entire story.
Next up: Part 6 — Preflight: Stop Before the First Change. We will build a required Pester gate that refuses to begin when evidence is missing, skipped, inconclusive, or unsafe.