Scenario 8 names a device that doesn't exist · Brandon Miller
·6 min read· agents· mobile· tooling· automation

Scenario 8 names a device that doesn't exist

The API 28 export bug had a written QA scenario four months before it reached the field. The scenario specifies an emulator image that has never existed on this machine, and nothing in the loop treats 'could not run' as different from 'passed'.

The export bug I wrote about last week had a test scenario. It has been sitting in the repo since 6 May, four months before the first failure arrived from the field.

Scenario 8, in .claude/agents/qa.md:

### 8. SAF + scoped storage on API 26–28
- On `Medium_Phone_API_28` AVD, verify export prompts for `WRITE_EXTERNAL_STORAGE`
- After grant, file appears in `Movies/Pulse/`

That is the bug. Not a scenario that might have tripped over it sideways. It names the API range, the storage mechanism, the permission and the destination directory. If it had run, it would have failed.

There is no Medium_Phone_API_28 AVD. There never has been. This machine has two Android emulator images, Medium_Phone and Pulse_API_34, and every artifact in test_evidence/, from 7 May through 1 August, was captured on a third: Medium_Phone_API_36.

What ran instead

test_evidence/ is the good part of this setup. Screenshots named by release and step, a full logcat and a filtered one, a version dump per run. Two months of it. The loop demonstrably ran.

It ran on API 36 every time.

The QA agent also carries a five-device priority matrix, each entry with a stated reason. Galaxy A54 is listed as "midrange Samsung, MediaMuxer bug risk." Pixel 4a is the Baseline-tier timing check. None of the five has ever been attached. SHIP_CHECKLIST.md gets to the same place by a different route, and its wording is the tell:

Test on real hardware (PRD §1.2 priority devices). Emulator validated SHIP, but on-device verification on at least one Full-tier and one Baseline-tier device is recommended.

Recommended. Every other item on that checklist is an imperative with a command under it.

Why nothing noticed

Here is the mechanism, and it is duller than a missing device.

An agent told to run scenario 8 would try to boot Medium_Phone_API_28, get nothing, and move on to scenario 9. Nothing downstream distinguishes that from a pass. The PostToolUse hooks in settings.json only echo a reminder string and exit 0. There are no PreToolUse hooks. No CI exists in either Pulse repo. The release record for versionCode 9 lists its gates as a flat list of PASS, because PASS and FAIL are the only two things it knows how to write down.

So the scenario existed, the scenario could not run, and the absence of a result was indistinguishable from a good result. Four months later the bug arrived on four devices across three brands and both chip vendors, every one of them API 28, every failure landing 6 to 30 ms into open_output.

A gate that cannot run is not a gate that passed. If the difference is not written down, it does not exist, and the default reading is always the optimistic one.

The other repo already fixed this

The part I find hard to sit with is that I solved this months ago, in a different codebase, and never carried it back.

Quartra's release workflow runs seven gates in parallel, and every gate prompt ends with the same sentence: a gate you could not run is NOT_RUN, never PASS. The gates return a three-state schema rather than a boolean. A separate release manager reads them and is told, in its own prompt, a NOT_RUN gate is not a pass, say what it would take to run it, do not soften a failure. Its QA agent has the same rule stated for scenarios instead of gates.

And where Quartra could not trust an agent to remember a rule, it made the environment refuse instead. The benchmark harness reads ro.kernel.qemu before it writes a result. If it finds an emulator, it rewrites its own CSV role column to reference and redirects the output into a separate directory, so an emulator number is structurally incapable of producing a ship verdict. The comparison tool refuses to read device rows from the other direction.

That is the difference between a protocol and a gate. One asks the agent to remember something at the moment it is least likely to. The other removes the capability.

Pulse has none of it. Same author, same month, adjacent directories.

What the scenario was guarding

minSdk is 26. Export output selection is four lines:

private fun openOutput(context: Context, filename: String): OutputHandle =
    if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.Q) {
        openMediaStoreOutput(context, filename)
    } else {
        openDirectFileOutput(context, filename)
    }

Three shipped API levels take the second branch. Zero of them have ever been booted. The branch is not obscure, it is not deep in a dependency, and it is not hard to reach. It is the first thing that happens on the export path for a quarter of the supported range.

The instrumentation that eventually caught it tells the same story from the other side. The doc comment on ExportFailure reads:

The field instant-abort investigation burned a day of session archaeology for lack of exactly this.

That comment is about the stage property, added after a day was spent reconstructing which part of the export died. It was written by someone who had just learned what missing instrumentation costs. The fix was an inline staged(stageName) { } wrapper, about six lines. The same person had already written scenario 8 and did not notice it had never run.

What I'm changing

Three things, in the order they matter.

NOT_RUN becomes a status. Every scenario and gate reports one of three values, and a run that produces no result reports the third. This is one word in a prompt and a column in the release record. It is also the entire fix for this class of failure, because everything downstream stops being able to read silence as success.

The device matrix gets materialized or deleted. Five priority devices and an API 28 image, created as AVDs and checked in as a list something can iterate, or struck from the agent files. A device matrix that exists only as prose is worse than no matrix, because it reads as coverage to anyone auditing the file, including me on the day I wrote scenario 8.

Environments refuse rather than remind. The API 28 run either happens on an API 28 surface or reports NOT_RUN. No path where a scenario names one device, runs on another, and files the result under the first.

None of this is about the model. Every instruction here was followed correctly. The agent was told to run a scenario, the scenario named a device, the device was not there, and nothing in the system had an opinion about what that meant.

The honest version of the last four months: I did not have a gap in coverage. I had a gap between coverage and the belief that I had it, and the belief was documented in more detail than the coverage was.