You can't parallelize what you can't verify
Quartra ran seven agent tracks in parallel against a frozen C ABI. One property made that possible, and it wasn't the model or the context window. It was an 8-second WAV file with a PyTorch tensor sitting next to it.
Quartra splits a song into vocals, drums, bass and other, entirely on the phone. HT-Demucs under ONNX Runtime, one shared C++ engine reached over JNI on Android and a bridging header on iOS, native UI on both sides. It was built as seven parallel tracks (model export, core engine, Android, iOS, billing, store, bench), one git worktree each, against a C ABI that no track but its owner is allowed to edit.
The question anyone asks about that shape is how you stop seven agents from producing seven locally-correct, mutually incompatible pieces of work.
The repo answers it under a heading called The property that makes this parallelizable:
There is a numerical oracle. Track 1's golden fixtures mean any agent on any track can prove its work is correct without asking a human whether the output "sounds right."
Everything else in the process exists to protect that one property.
What the oracle actually is
fixtures/golden-001/. Eight seconds of structured stereo audio, 352,800 samples at 44.1 kHz, seed 42. The length is not arbitrary: the exported graph takes a fixed 343,980-sample segment, so eight seconds is the shortest input that forces two segments and therefore exercises overlap-add.
Next to the input sit expected_vocals_ft_sub3.npy and expected_htdemucs_4stem.npy, generated by Demucs' own apply_model. A manifest.json carries the recipe (shifts=0, split=True, overlap=0.25, transition_power=1), pins torch 2.13.0+cpu / demucs 4.1.0 / numpy 2.5.2, and records the sha256 of every file.
Audio separation is close to the worst case for machine verification. The output is a waveform. "Did the vocals come out clean" is a listening test, and if that's your acceptance criterion then every agent's work funnels through one person's ears and your seven parallel tracks are a queue with extra steps.
The oracle turns that into ctest.
The first gate was wrong, and the fixture proved it
The original bound was max-abs ≤ 1e-4, set against a measurement taken on a single random-noise segment. Real structured audio broke it, in the most useful possible order. It caught two real bugs first, then it caught the metric.
| Bug | Symptom | Fix |
|---|---|---|
| Short final chunk left-aligned, zero-filled on the right | four-stem max error 3e-2 – 9.5e-2 | Centered padding, matching Demucs' TensorChunk.padded + center_trim |
| Outer whole-track normalisation (channel-mean std) | ~5e-4 spikes in chunk 1 | Removed. HT-Demucs normalises inside forward() and the ONNX export carries that graph. Raw in, raw out |
The first one is a genuine modelling mistake rather than an off-by-one: the model reaches backward for left context, so a short tail has to be padded on both sides and trimmed, not left-aligned and stuffed with zeros on the right.
Then the interesting part. The residual that survived both fixes is not the engine's. A pure single forward pass of chunk 1, with no engine involved at all, reproduces it to three significant figures: fp16 vocals 5.63e-4 either way, fp32 1.107e-4 against 1.11e-4. ORT graph optimization level made no difference anywhere from DISABLE_ALL to ENABLE_ALL. So fp16-versus-fp32 is weight quantisation, and fp32-versus-torch is ORT-versus-torch kernel numerics. Neither is reachable from the engine.
Which means max-abs was the wrong shape of bound. It tolerated neither real bug by mean, and it is dominated by isolated single-sample fp16 excursions that no engine change can ever remove.
The rejected option is the part I keep coming back to:
Keep 1e-4 max-abs. Then no fp16 export can ever pass, nor this fp32 one. The oracle becomes a permanent red and gets ignored. The worst outcome.
An oracle that is always red is not a strict oracle. It is no oracle. The failure mode isn't a bad build shipping. It's everyone learning to scroll past a red check.
The replacement gates on mean-abs ≤ 3e-5 and per-stem SDR-versus-oracle ≥ 55 dB, and still reports max-abs on every run without gating it. What makes that credible rather than convenient is the sanity check attached to it: both bugs this fixture found would have tripped the new gate. Their means were 6.2e-5 and 4.8e-5, and the padding bug alone cost about 25 dB of SDR.
Current margins are roughly 2× on mean and 6+ dB on the worst stem, at 61.3 / 74.1 / 90.2 / 76.7 / 68.1 dB per stem against the reference implementation. For scale, the product's own definition of parity is 0.2 dB of end SDR at around 9 dB absolute. Sixty-one dB of agreement is more than fifty dB below anything audible.
The general claim
What limited agent throughput on this project was never the model and never the context window. It was how much of "correct" could be asserted without me.
Where the oracle reaches, agents run flat out. The core track can rewrite overlap-add windowing overnight and know by morning whether it's right. Android and iOS can build against a stub engine and swap in the real one, because parity is a number rather than an opinion.
Where it doesn't reach, everything serialises on one person. Whether the paywall appears at the right moment. Whether the Stems screen feels right on the first reveal. Whether four stem colors stay distinct on an OLED panel in sunlight. All of that came through me, one item at a time, and all of it was the slow part of the project.
That's an uncomfortable read for most mobile codebases, because very little in a typical app is machine-checkable at that grain. "The list scrolls smoothly." "Login works." A UI test asserts a button exists, not that the screen is right.
So the question I'd now ask before starting any parallel-agent project isn't which model or how large the context window. It's: what fraction of correct can be asserted, and what is the cheapest thing I can build that moves that fraction up.
For Quartra the answer was an eight-second WAV and a PyTorch script. It cost one track about a day. It is the reason the other six could run at all.