How we test DetentShell's UI in Windows Sandbox
Our unit suite once told us the debugger was fine while F5 was dead at every breakpoint. Nine hundred tests, all green. Every command and view-model behaved correctly in isolation; the bug lived in the seam between WPF’s keyboard routing and an async command’s state, a place unit tests don’t reach. Any user pressing the actual key on the actual window would have found it in seconds.
So we built a user. A test rig drives the real compiled app: real window, real keystrokes, real mouse coordinates, through more than forty scenarios. It presses F5 and asserts execution stopped at the breakpoint. It right-clicks the editor and works the context menu. It types into the search panel, drives the Save As dialog, generates Pester tests from the Tools menu, runs Quick Fix through its preview and verifies the buffer changed. Every scenario records a video, so a failure arrives with footage.
The rig runs inside Windows Sandbox for two reasons. Synthesized input is indiscriminate. A test that sends F5 will send it to whatever window has focus, including your email, so the input needs a desktop nobody is using. And the disposable VM doubles as an honesty check: it’s a factory-fresh Windows with no dev-machine leftovers, so if the app works there, it works because of what we ship, not because of what happened to be installed on the build machine.
The bug list this rig has produced is the kind unit tests structurally miss. The dead F5, where the toolbar button worked and the key didn’t. File Open leaving the caret past the last line, which quietly broke the debugging step that followed it. Generated tests checking a temp path that no longer existed. A single unnamed toolbar control that a screen reader would announce as just “button,” caught because the accessibility scenario counts names and got sixteen where it expected seventeen.
The rig has also had to police itself. Early screenshots showed the editor area as a black void, which looked like a serious rendering bug in the app. It wasn’t. The sandbox window had opened tiny, the virtual display matched it, and the app simply had no pixels to draw the editor in. The harness now refuses to start until the display is a usable size. A test system you can’t audit will lie to you with complete confidence.
The thousand-test unit suite still runs first; it’s fast and catches most regressions cheaply. The sandbox rig is for the bugs that live where the user lives, at the keyboard, in the window. Its verdict gates every release, and when someone asks whether the debugger works, there’s a video.