The scheduler is a human, and that's the bug
Seven weeks of agentic analytics on Pulse found real bugs, some of them diagnosed down to the failing code path. Every one of them waited for me to ask. The next version of this loop runs on a cron and opens the pull request itself.
The Pulse analytics loop has been running since 16 July. It has found things a dashboard would not have: one event name resolved into three unrelated bugs, an ad thesis inverted by a metric invented halfway through answering the question, a paywall recommendation reversed on its own evidence.
It has one structural flaw, and it is not in any of the analysis. Every cycle begins with me typing "what's the latest data?"
Across seven weeks, roughly half my turns in that session have been bare polls, on most of the days the app saw any traffic at all. That is not a workflow. That is me working as a cron job with worse uptime.
What a human scheduler costs
Three costs, all visible in the record.
Detection latency is a function of my attention. Aug 13, Android 1.6.0 went 13-for-13 on exports and got called fixed. The refutation arrived the next day: 3 ok, 14 failed. That data existed at the close of the Aug 14 bucket. It got read when I next happened to ask.
Nothing is checked on the days you don't ask. The gap between days with activity and days with a poll is small, and it is still the wrong shape, because the days I don't ask are uncorrelated with the days nothing happened.
Session state evaporates. More than half my queries needed the project re-pinned before they could run, because a stateful MCP session against a human who shows up once a day keeps timing out. Pure tax, and it exists entirely because a person at a keyboard is driving.
What a scheduled agent would actually do
Not "summarise yesterday's numbers". That is a dashboard with extra latency and a worse interface.
The useful version is narrow: a scheduled run, on closed buckets only, that diffs today's registered stats against a remembered baseline and a remembered set of already-falsified hypotheses, and speaks only when something crosses a line it can name.
Three preconditions make that worth building, and all three now exist.
- The catalog. Thirty-one definitions, each carrying its caveats in its own description. A scheduled run reading
export_success_rate_by_platformover closed buckets gets the same number every time it asks. That is the reproducibility a cron job needs and a raw query cannot give. - Absence detectors.
orphaned_video_pickandorphaned_pipeline_runmean a crash that cannot report itself still surfaces as a number. Those are precisely the things nobody checks by hand. - Long retention where it matters.
purchase_ledgerandmonetization_funnel_by_platformcarry 3650-day retention, so a scheduled run has something to diff against that outlives the 21-day raw window.
The part that isn't analytics
Here is where I want this to go, and the reason it is worth building rather than just wiring up a nicer alert.
export.fail used to be one line on a chart. Adding a stage prop resolved it into three clusters that have nothing to do with each other: a failure at open_output in the first few milliseconds, one at copy after half a minute, one at cpu after more than a minute. Different stages, different exceptions, different API levels, three different owners.
That is not a hunch, and the first of those three is close to a task description already: it names the failing stage, the timing that rules out an encode problem, and the write path implicated. What it doesn't have is anyone acting on it, because the only thing between a diagnosis and a code change is me reading it and opening an editor.
The target shape: a scheduled run detects the regression against a baseline it remembers, writes the finding to durable memory, and opens a pull request with the diagnosis in the description and a test that reproduces it. I review the PR. I don't discover the bug.
Three things that have to be true first, and one is currently false
Durable memory, written at discovery time. This is the false one, and it is the blocker.
Everything the loop has found lives in a single chat transcript and four untracked HTML files. Almost nothing has been written to durable memory. The one analytics memory that does exist was written on day one, describes the SDK plumbing, and still records two infrastructure bugs as open long after they were fixed. It lists the original 11 stat definitions and none of the 20 added since. A fresh session reading it would be actively misled.
That is specifically fatal for a scheduled agent, because surveillance is a diffing job. Today's numbers only mean anything against a remembered baseline and a remembered list of hypotheses you have already killed. Without written memory the agent re-derives, and re-derivation is not free. The hypothesis that a highlight's hype score predicts whether it gets exported or deleted has been proposed and falsified more than once, on the same data, because nothing recorded that it was already dead. Hype 1 through 4 have all been both exported and deleted.
And the compaction boundary is the real enemy. That session has compacted twice, and each time a finding survived only if it happened to have been restated recently enough to land in the summary. Memory has to be written at the moment of discovery, not at the end of the day.
One instance is not a mechanism. An agent with a budget to open pull requests and a taste for narrative is exactly the thing that would ship a fix for a coincidence it saw once. Every interesting anomaly in this dataset so far has arrived as a single device doing a single strange thing, and most of them stayed that way.
Render it and look at it. The most transferable failure so far had nothing to do with data. An oversized SVG caused horizontal overflow on a published page, and the fix applied silenced the overflow metric being used to detect the problem while leaving a 5,030 px element on the page. It shipped. The agent had a numeric check, the check went green, and the check going green was the bug.
Where it stands
Not built. The catalog is there, the absence detectors are there, the long-retention stats are there. The memory discipline is the gap, and it has to close first, because a scheduled agent without memory doesn't compound. It just repeats, once a day, forever.
The honest summary of seven weeks: the analysis was good and the scheduling was me. Only one of those scales.