Writing
Scenario 8 names a device that doesn't exist
The API 28 export bug had a written QA scenario four months before it reached the field. The scenario specifies an emulator image that has never existed on this machine, and nothing in the loop treats 'could not run' as different from 'passed'.
The store tax on shipping on-device ML
Quartra carries a 158 MB neural network inside the app and runs it on the phone. Four things about that turned out to be store problems rather than engineering problems, including a Play declaration that cannot be filed until after you've already shipped.
The scheduler opened the pull request
An export bug on four devices across three brands and both chip vendors, every one API 28. The scheduled run read the closed bucket, diffed it against the stored baseline, wrote the diagnosis, and opened the PR. I reviewed it over coffee.
A reviewer that forgets is a linter
Every review agent on Quartra keeps its own memory directory: indexed, one file per lesson, written at the moment it learned the thing. Four entries in, it had already caught a test assertion that always passes.
You can't parallelize what you can't verify
Quartra ran seven agent tracks in parallel against a frozen C ABI. One property made that possible, and it wasn't the model or the context window. It was an 8-second WAV file with a PyTorch tensor sitting next to it.
The scheduler is a human, and that's the bug
Seven weeks of agentic analytics on Pulse found real bugs, some of them diagnosed down to the failing code path. Every one of them waited for me to ask. The next version of this loop runs on a cron and opens the pull request itself.
A question asked three times becomes a definition
Five weeks of running Pulse's production analytics through an MCP loop with no dashboard, no analyst and no scheduled jobs. The mechanic that makes it compound is promoting a repeated question into a stored definition.
The dashboard is the wrong artifact
Distill ships three SDKs and no UI. The query surface is MCP, which means the thing that accumulates over time is a catalog of definitions instead of a wall of charts.
The win grows with conversation depth
Same prompt. Same K=128 budget. Same hardware (a Pixel-class on-device runtime). Three policies. SlidingWindow forgot subprime mortgages. Plain TemporalKV forgot 2008. The hybrid kept both. Here's the trace, and why.
The ReLU is doing all the work
A linear scorer on the same features scores AUC 0.859. Add one hidden layer of 8 ReLU units (49 parameters total) and AUC jumps to 0.900. We opened up the trained weights to see what the nonlinearity actually bought us. It was not what we expected.
When StreamingLLM beats us, and why
At K/T = 1/16, the dumbest policy in the comparison ('always keep the last 64 tokens, no exceptions') outperforms our learned policy on perplexity by 17 points. We dug in to figure out where it was spending its budget.
It's all about K/T
If you plot a learned eviction policy's win over heuristics against cache size, you get a confusing picture. Against context length, also confusing. Against their ratio, you get a wall.
Sparse eviction is the right llama.cpp primitive
We took our eviction policy off the PyTorch research stack and onto llama.cpp on a real device. Decoding got 5x slower. The fix was four lines of code, plus a re-read of the cache's data model.
Where AI coding agents go blind on mobile
Three structural blind spots that limit what AI coding agents can do on iOS and Android, and what it takes to fix them.
Why I'm rebuilding how I ship mobile
Mobile has been slower to absorb AI coding agents than web. The interesting engineering is in the scaffolding around the agent, not the models.
MCP at 97 million
Model Context Protocol hit 97M monthly SDK downloads in March and now sits under the Linux Foundation. What that means for what a mobile engineer should build.
Xcode 26.3: Apple, late but serious
Xcode 26.3 ships with agentic coding, Claude and Codex integrations, and MCP support. The MCP part is the real news.
Switching the default to Sonnet 4.6
Sonnet 4.6 landed as the new default in Claude Code. Practical notes on when it replaces Opus for mobile work and when it doesn't.
One million tokens and the legacy codebase problem
Claude Opus 4.6 ships with a 1M context window. Necessary but not sufficient. A real Android codebase still needs retrieval, not just volume.
Cowork and the question of surface
Claude Cowork is a clean read of where desktop agents belong. Mobile agents need a different surface: one that can see the device, not just the files.
Vibe coding won't work on a device
Stack Overflow's latest survey shows 72% of pros refuse to ship AI-generated code without review. On mobile, that review gate isn't sentiment. It's load-bearing.
The year agents stop being a demo
The 2026 AI pragmatism narrative makes sense on the web. Mobile is still a cycle behind: the demos look good because the feedback loops are still carrying the weight.