I measured my own test generator against 195 real functions. It was humbling.

PesterForge · July 2026 · 5 min read

This morning, the Pester test files my generator produces failed an average of 2.99 assertions each, straight out of the box. By the end of the day it was 1.01. Nothing about the generator’s ambitions changed in between. What changed is that I finally measured it.

The setup

I build PesterForge, a PowerShell module that reads a function and writes a runnable Pester test file for it — deterministically by default, no AI calls unless you explicitly opt in. It had a test suite of its own, green. It had passed code review many times. I would have told you it was in good shape, and I would have half-known that “good shape” was a feeling, not a number.

So I built the number. Invoke-PfEvalCorpus.ps1 (it ships in the repo’s Tools\ folder) runs the generator against a corpus of 195 real-world functions — 150 from dbatools, which is about as gnarly as production PowerShell gets, plus 42 from my own GitEasy module and a few fixtures — and scores every output: does it parse, does it run under Pester, how much of it is real assertions versus placeholders, what fails.

The good news the corpus confirmed

Three numbers came back clean, and they’re the hard part of the problem:

If you’ve ever generated code programmatically, you know those three lines don’t happen by accident. I’ll take the win.

The humbling part

Then the other columns. On average, 69.5% of every generated test body was a placeholder stub, not an assertion. And out of 195 outputs, the number that were genuinely ready to commit as-is was zero.

The most useful finding was the one that looked backwards at first: functions my profiler rated most testable produced the most stubs — 72% for High-testability functions versus 57% for Medium. The explanation fell out of the data. Stub ratio tracks the number of functions the target calls:

Called functions Mean stub ratio Mean failures
0 42% 1.0
1–3 60% 2.1
4–8 71% 3.0
9+ 80% 4.4

For every collaborator, the generator emits a mock-setup stub, because it can’t infer what the mock should do. The “highly testable” functions are the big orchestrators with the most collaborators, so they stub the most. That table is the roadmap.

What the 2.99 failures actually were

Reading the worst outputs by hand turned up three concrete bugs, none of which my own test suite could have caught, because all three only appear against someone else’s code:

  1. I was asserting my conventions against their source. The generated static test checked that the target file carries my change-log header style. dbatools uses a different convention, so the assertion failed on every single file. The fix: check whether the source actually uses the pattern at generation time, and emit an honest skipped stub when it doesn’t.
  2. A null-array crash on unresolvable parameter types. If any parameter in a function’s signature references a type that isn’t loaded (dbatools’ [DbaInstance]), Get-Command hands back a null parameter collection, and my default-value extraction indexed into it. Every parameter assertion in the file failed with Cannot index into a null array.
  3. Unguarded execution against missing dependencies. Generated tests dot-sourced and invoked targets that need modules the test machine doesn’t have, producing a wall of red that told the user nothing.

The fixes share one philosophy, which is also the module’s house rule: when the generator doesn’t know, it says so — it never fakes green. Placeholder assertions skip explicitly:

It 'Credential is an accepted parameter -- FILL IN' { Set-ItResult -Skipped -Because 'stub' }

And a missing dependency now surfaces as Inconclusive with the dependency named, instead of seven failures that all mean the same unstated thing. After those fixes, the corpus re-run came back at 1.01 mean failures per file — and the remaining failures are real findings about the target code, which is what failures are for.

The bug I’m least proud of

The corpus also caught my generator baking my own machine’s absolute paths into generated files — twice, by two different mechanisms. Once in the line that locates the function under test (a string-replace “relativization” that silently did nothing when the source lived outside the project root), and once in a comment that echoed a parameter’s literal default value. Both are fixed, and both are now behind permanent Pester gates that scan full generated content for any absolute drive path, so they can’t come back quietly.

A tool that writes files for other people’s repos has no business knowing my directory layout. Finding that in my own output was uncomfortable in exactly the way measurement is supposed to be.

The takeaway

The corpus run takes about six minutes, and re-running it is now part of how I cut a release — each of today’s fixes shipped with a fresh 195-function measurement, so the numbers get re-earned instead of quoted from the day they were good. (The whole surface is also exposed as an MCP server now, so editors and AI assistants can call the generator through a versioned JSON contract; that’s another post.)

Usual scope caveat: this is a one-person suite, Windows, and your mileage may vary. But the pattern travels: if your tool generates code, its own unit tests only tell you the generator runs — they don’t tell you the output is any good against code you didn’t write.

Do you measure the output quality of your code generators against a real corpus — or, like me until yesterday, do you have a green test suite and a feeling?

← All posts