The scheduler opened the pull request
An export bug on four devices across three brands and both chip vendors, every one API 28. The scheduled run read the closed bucket, diffed it against the stored baseline, wrote the diagnosis, and opened the PR. I reviewed it over coffee.
The Pulse export funnel had been quietly broken on one class of device for weeks. Not broken enough to show up as a crash, not broken enough to move the blended success rate much, and invisible on any chart, because the failures all carried the same event name as two entirely unrelated bugs.
The run that found it read a closed daily bucket, diffed it against a stored baseline, checked its own hypothesis against a rule it had been given for exactly this shape of mistake, and opened a pull request. I read the PR before I had read anything else that morning.
Here is what it actually did, and the two guardrails that turned out to be the whole ballgame.
The digest
The run reports on closed buckets only. Nothing rolling, ever, because a rolling counter can't be diffed against yesterday's rolling counter without lying to you about which direction things moved.
What it had to work with was export_fail_by_error, a stat registered a month earlier during a different investigation, dimensioned by a schema-normalised error token, stage, and app_version. The stage dimension is the entire reason any of this was possible. Before it existed, export.fail was one line.
The finding, once it had resolved the dimensions:
| stage | exception | duration | API | devices |
|---|---|---|---|---|
open_output | FileNotFoundException | 6–30 ms | 28 | Sony, vivo 1902, vivo Y85, OPPO CPH2015 |
copy | ErrnoException | ~36 s | 35 | 1 |
cpu | IllegalStateException | ~76 s | 36 | 1 |
Three bugs under one event name. Three different owners. The run separated them by stage and then, crucially, stopped treating them as one number.
The open_output cluster is the one that mattered: four devices, three brands, both Qualcomm and MediaTek, every single one on API 28, every single failure landing 6 to 30 ms in. That timing is the diagnosis by itself. Thirty milliseconds is not a failed encode. It is not a disk-space problem and it is not a codec problem. It is a file that could not be opened for writing before any encoding was attempted, on the one API level where the app still took the legacy external-storage write path.
The rule that stopped it being wrong
Days earlier the same hypothesis had been narrowed, badly, on a contradicting device. A Galaxy S8 on API 28 had processed a video without incident, and that looked like a clean refutation of "this is an API 28 problem." The theory got narrowed to one manufacturer.
That was wrong, and it was wrong in a way that is extremely easy for an agent to be wrong. The S8 never attempted an export at all. It was not evidence about exports in either direction.
An agent sweeping a dataset will always find some case that appears to contradict the pattern. Before a contradicting case is allowed to move a hypothesis, it has to be shown to have exercised the code path in question.
That rule was in the run's memory as a written lesson, attached to this specific investigation, because it had cost something the first time. So when the sweep surfaced API 28 devices with no export failures, it checked whether those devices had any export.start at all before letting them weaken the finding. They didn't. The pattern held, and it held across three more brands.
This is the part I want to be precise about, because it is the difference between a scheduled agent that is useful and one that is a liability. The value wasn't the sweep. The sweep is easy. The value was a written rule that survived from a previous mistake into a later run, and fired at the exact moment it was relevant.
The second guardrail: a streak is not a denominator
A month earlier, an Android build went thirteen-for-thirteen on exports and got called fixed. The next day was three successes and fourteen failures.
That is now encoded as a claim rule. A fix claim needs a version, a success count and a failure count, and it has to survive the next closed bucket before it can be stated as fact anywhere. The run applies it to its own output, which means it is structurally incapable of declaring victory on the strength of one good day, and it is also why the PR description says "candidate fix" rather than "fixes."
The verification came later and separately, on its own closed buckets: Android has been clean on open_output since 15 September.
The pull request
The diagnosis was already a task description. It named the failing stage, the API level, the code path and the remedy, which is more than most bug reports I get from humans contain.
What landed: a branch, a change swapping the legacy external-storage write for a MediaStore-backed, app-scoped write on API ≤ 28, and a test exercising the export path with the storage API pinned to 28. The description carried the stage table above, the four device models with their brands and chip vendors, the 6-to-30 ms timing with the argument for why it rules out an encode failure, and a link to the stat series the claim came from, so the number was reproducible rather than asserted.
It did not carry a victory lap, because of the rule above.
What I changed in review
Two things, and neither was the diagnosis.
The write path it chose was correct but the failure handling wasn't. The original change let a MediaStore insert failure propagate as the same generic export error, which would have put the new code's failures straight back into the bucket it was trying to empty and made the next run's diff meaningless. That's a subtle thing to get wrong and it is exactly the class of mistake I'd expect: it optimised the fix and not the observability of the fix.
The second was scope. It had also tidied an unrelated path in the same file. Correct tidying, wrong PR.
Neither of those is an argument against the loop. They are an argument for reviewing the pull request, which I would be doing anyway.
What this actually changes
Detection latency stops being a function of my attention. That is the whole of it, and it is bigger than it sounds.
The old shape was that findings waited for me to ask. Over the preceding two months the analysis had been genuinely good and the scheduling had been me typing "what's the latest data?" on the days I happened to think of it. Every insight in that period, including the ones that changed product decisions, sat in a transcript until I next opened it.
Three things had to be true before any of this could run unattended, and only one of them was about analytics.
The catalog. Definitions registered over months, each carrying its caveats in its own description field, read over closed buckets. That is what makes a number reproducible between two runs a day apart. A raw query cannot give you that, and an agent that hand-aggregates raw rows produces a different number every time it is asked.
Absence detectors. A hard crash cannot emit a failure event. Two orphan detectors, one for a video.pick with no pipeline.start inside 60 seconds, one for a pipeline.start with no terminal event, mean a run that dies silently still surfaces as a number. Those are precisely the things nobody checks by hand.
Memory written at discovery time. This was the blocker, and it is the one that took longest. The S8 rule only worked because it had been written down when it was learned rather than at the end of the day. The transcript is not memory. Sessions compact, and a finding survives compaction only if it happens to have been restated recently enough to land in the summary, which is a coin flip dressed up as a process.
What I am still watching for
One instance is not a mechanism, and a scheduled agent with a budget to open pull requests is exactly the thing that would ship a fix for a coincidence. There is a case in the data right now of a user hitting an export failure, subscribing three and a half minutes later, and then running eight successful exports in three minutes. The tempting story is a free-tier restriction surfacing as a generic error and driving a conversion. It is one user. I want it to stay one user in the record until it is more than that.
And the green-check problem, which is the failure I trust least. The most transferable mistake this loop has produced had nothing to do with data: an overflow check went green because the fix had silenced the metric that detected the overflow, while leaving a five-thousand-pixel element on the page. It shipped.
Anything that acts without a human in the loop needs at least one verification that is not the thing it is optimising. For the export fix that verification is a test that opens a real file on a real API 28 surface, and it is not negotiable, because the number in the dashboard is the thing the agent is trying to move.