A question asked three times becomes a definition · Brandon Miller
·7 min read· analytics· distill· agents· mcp

A question asked three times becomes a definition

Five weeks of running Pulse's production analytics through an MCP loop with no dashboard, no analyst and no scheduled jobs. The mechanic that makes it compound is promoting a repeated question into a stored definition.

Pulse has been instrumented on Distill since 16 July. No dashboard, no analyst, no scheduled jobs. An agent with MCP access to the event store, and me pinging it when I care.

Five weeks in, here is what the loop actually looks like, and the three things it has done that a chart structurally cannot.

The loop

  ping ("what's the latest data?")
        ↓
  summary counters → has anything moved?
        ↓
  read registered stats over closed buckets
        ↓
  ── anomaly? ──→ raw-event drilldown on the specific device / session
        ↓
  narrate: what changed, what it means, what is still unknown
        ↓
  ── recurring question? ──→ register a new stat def

The second-to-last edge is the entire thing. A question asked once gets answered. A question asked three times becomes a definition, and from then on it answers itself over closed daily buckets, indefinitely, with its caveats attached.

That is the ratchet that turns an ad-hoc chat into an asset.

The catalog, and what triggered each wave

Twenty-odd definitions now, registered in waves. Every wave has a trigger, and the trigger column is the interesting one:

DateRegisteredTriggered by
Jul 16dau wau mau d1_retention d7_retention pipeline_success_rate activation_rate paywall_conversion_rate export_complete_count pipeline_p95_processing_ms pipeline_realtime_ratio_avginitial instrumentation (templates)
Jul 27the four _by_platform variants"1 stray Android": blended metrics describe neither platform
Aug 5orphaned_video_pick orphaned_pipeline_runa crash stack pasted in
Aug 11daily_event_volume daily_pipeline_flowclock-skewed device fleet, plus wanting charts over time
Aug 15export_success_rate_by_platform export_fail_by_error export_engine_splitthe instant-abort export investigation
Aug 16wait_time_minutes_by_phase processing_attention_split processing_wait_buckets empty_completion_by_signaturethe in-app-ad thesis
Aug 21purchase_ledger monetization_funnel_by_platform"is there an index on the purchases to track forever?"

Only the first row is a template. Everything after it exists because a specific question got asked, and then asked again.

Three things a chart can't do

Split one event name into three unrelated bugs.

export.fail is a single line on any chart. Adding a stage prop in Android 1.6.1 resolved it into this:

stageexceptiondurationAPIdiagnosis
open_outputFileNotFoundException6–30 ms28legacy external-storage write path
copyErrnoException~36 s35likely disk full
cpuIllegalStateException~76 s36GPU fallback (gpu_error: 3002), then CPU failure

Three fixes, three owners, one event name. The live one is API 28 open_output, confirmed across four devices and three brands (Sony, vivo 1902, vivo Y85, OPPO CPH2015), both Qualcomm and MediaTek, every failure 6 to 30 ms in, which is to say before any encoding begins. Nobody was going to see that in a bar chart.

Answer a business question with a metric that didn't exist when it was asked.

Can we sell ad inventory against processing wait time? The surface numbers say obviously yes: Android median processing wait is 63 s, 70% of runs are at least 30 s, 51% are at least 60 s.

Then processing_attention_split got written specifically to ask whether anybody is looking at the screen during that wait. 45% of waits are backgrounded, and the backgrounded ones are the long ones: median 125 s backgrounded against 31 s in the foreground. Only about 28% of the wall clock is addressable at all.

The long waits that make the inventory look attractive are precisely the waits where the user has left the app. That inverts the thesis, and it required inventing a metric in the middle of answering the question.

Measure the absence of an event.

A hard crash cannot emit a failure event. pipeline_success_rate is structurally blind to it, because the run appears in neither the numerator nor the denominator. Hence two orphan detectors. orphaned_video_pick is a video.pick with no pipeline.start within 60 s in the same session, which is the signature of the app dying at service start. orphaned_pipeline_run is a pipeline.start with no terminal event before the next one in that session. Baseline: 3 of 43 starts, 7% vanishing silently.

A dashboard charts the events that happened. The interesting ones are the events that should have happened and didn't, and those can only be expressed as a definition.

The definition work is the analysis

empty_completion_by_signature is the cleanest case. "Pipeline completed with zero highlights" looks like one number. It is two completely different things, separated by whether the content_type prop is present. Present means the classifier ran and the video was genuinely too short, which is correct behaviour. Absent means the classifier never ran, which is the bug. Baseline over three weeks: 17 missing against 9 ran.

A blended "empty completions" count averages a defect together with a correct result and tells you nothing at all. The insight was not in a chart. It was in choosing the split.

The same shape shows up in export_fail_by_error, which has to normalise across a schema change: before 1.6.0 the error prop held an exception class name, and after it holds a stable token with the class moved to cause. Legacy rows get prefixed legacy: so two vocabularies never mix inside one bucket.

Put the caveat inside the definition

Every stat carries its own epistemics in its description field. Verbatim from the catalog:

wait_time_minutes_by_phase NOTE: counts successful runs only, so crashed or cancelled pipelines and failed exports contribute nothing, and true wall-clock wait is higher.

This is the highest-leverage habit of the whole engagement. A number that travels without its caveat gets misused three weeks later by somebody, very much including the agent that produced it. Putting the caveat inside the definition makes it impossible to read the stat without reading the warning.

purchase_ledger goes further and writes its own birth defect into its description: it was registered too late to capture the first sandbox purchase, whose raw rows had already expired. That fact is now permanent, attached to the series it damages.

Two ways this has already gone wrong

Declaring a fix on one good day. Aug 13: Android 1.6.0 went 13-for-13 on exports, and I reported it fixed. Aug 14: 3 ok, 14 failed. A streak is not a denominator. A fix claim needs a version, a success count and a failure count, and it needs to survive the next daily bucket before it goes anywhere anyone reads.

Reading a rolling window as a fixed one. The project summary returns 7-day top-event counts. They got reported as 21-day figures, and that made it onto a published page before being caught. The same mistake produced a "highest export.start yet" that was an artifact of the window sliding rather than a real peak. Closed daily buckets for anything that is a claim; rolling counters only for "has anything moved since yesterday".

And one that is about the loop itself

I asked whether the agent was using aggregates to query. It wasn't. It had been pulling raw events and aggregating them by hand, which is why several queries timed out and two spilled to files of 90 KB and 266 KB.

There were already 22 registered stats at that point, including three created six days earlier during the export investigation and never read once. The agent built the tool and then didn't use it.

Raw querying feels more responsive turn to turn. It answers the exact question asked with no schema negotiation. The cost is invisible in any single turn and enormous in aggregate: non-reproducible numbers, inconsistent windows between answers, and nothing that survives the 21-day raw window.

That is the obvious objection to a no-dashboard design, and it landed anyway, on a catalog that already existed and was already good. Building the ratchet is not the same as pulling it.

Where this goes next

The loop works, and it compounds, and the compounding is entirely down to that one edge in the diagram.

But look at the first box again. It says ping. Every cycle begins with me asking. Across five weeks, roughly half my turns in that session have been bare polls: "what's the latest data?", "anything new?", "check again".

That is a scheduler made out of a human, and it is the next thing to fix.