The store tax on shipping on-device ML
Quartra carries a 158 MB neural network inside the app and runs it on the phone. Four things about that turned out to be store problems rather than engineering problems, including a Play declaration that cannot be filed until after you've already shipped.
Quartra separates a song into stems on the device. HT-Demucs under ONNX Runtime, no server, no upload. The model ships inside the app rather than downloading on first run.
The engineering case for that is easy and mostly already written. The store case is where it costs you, and that part is much less documented, because the parts of Google Play and App Store Connect you end up in are the ones almost nobody visits.
Four things, in the order they bit.
1. The listing is around 300 MB, and that was the cheap option
The original plan was on-demand download. A day of design went into it: two delivery backends, an iOS availability matrix weighing AssetPackManager against BADownloadManager, a deployment-target decision, a consent flow for guideline 4.2.3(ii), and a CDN that was never built.
Two measurements deleted all of it. fp16 weight storage is identical in throughput to fp32 at half the size: 316 MB down to 158 MB, free. And both models embedded measure 295.8 MB compressed in a real release AAB, against Play's 500 MB base-module limit. The premise the whole design rested on, a payload north of a gigabyte, had never been measured.
So the download path bought nothing except a funnel for users to drop out of. Embedding didn't mitigate that risk, it removed it. What's left is a ~300 MB listing, taken deliberately and with eyes open.
One number worth carrying into your own budgeting: dense neural network weights barely compress. gzip -6 takes the 158 MB model to 139 MB. If you're sizing a model against a store limit, size the raw file.
2. Play blocks 4 KB-aligned native libraries
Play won't accept new apps targeting Android 15+ whose native libraries use 4 KB page alignment. ONNX Runtime was already 16 KB-aligned. Microsoft got there early. libstemcore_jni.so, our own JNI glue, was not, because NDK r27 still defaults to 4 KB.
The fix is one cmake argument:
-DANDROID_SUPPORT_FLEXIBLE_PAGE_SIZES=ON
And the check:
llvm-objdump -p libstemcore_jni.so | grep LOAD
Every segment should report 2**14.
This is trivial once you know it and completely opaque when you don't, and there's a specific trap in it: it's easy to assume you're covered because your heavyweight dependency is compliant. The 200 MB dependency was. The 40 KB of glue wasn't.
The bundle had already gone up before that check, so the fixed one is the next number. And because this project runs a single version number across both platforms, the iOS CURRENT_PROJECT_VERSION moved with it, which stranded an iOS build that was already sitting in Apple's review queue.
3. mediaProcessing is an API 35 foreground service type
A separation run is minutes of continuous inference. The user starts it and then leaves the app, or locks the screen, and waits for a notification. That's a foreground service, and Android 15 has a type that describes it precisely: FOREGROUND_SERVICE_TYPE_MEDIA_PROCESSING.
It does not exist below API 35. A build shipped with it crashed every separation start on Android 14 (InvalidForegroundServiceTypeException: Starting FGS with type unknown), and for targetSdk 34+ there is no untyped foreground service to fall back to.
The fix is a split: dataSync on API 29–34, mediaProcessing on 35+, with the manifest declaring mediaProcessing|dataSync. Because mediaProcessing covers everything from 35 up, dataSync's 35+ time limit never actually applies.
There was an earlier trap in the same few lines, and it's the more insidious of the two. ServiceCompat.startForeground silently masked the type to zero: 8192 is newer than core-ktx 1.15's table of known types, so the compat layer dropped it. The platform then rejected the call with "Starting FGS with type none" while the manifest and the constant were both perfectly correct. The compat wrapper, whose entire job is to handle version skew, was the thing introducing it. Fix: call the platform API directly.
4. The Play declaration for that permission cannot be filed in advance
This is the one I'd never have predicted, and the reason I'm writing any of this down.
Play generates the Foreground service permissions form from the permissions it finds in released bundles. Checked in the console: "App bundles and APKs using sensitive permissions" listed only the one production versionCode, and the form offered a mediaProcessing section and nothing else. A bundle sitting as an unreleased draft is never scanned. There is no DATA_SYNC section to fill in, because as far as Play is concerned that permission does not exist yet.
So the order is forced:
- Release a build carrying the new permission to a track.
- Let Play rescan. The declaration appears under App content → Need attention.
- Fill it in, which in this case includes a video demonstrating the service.
- Then promote to production.
You cannot prepare step 3 ahead of time. Which means adding a foreground service type costs a full release cycle plus a review, quoted at up to seven days, and no amount of preparation compresses it.
And the API and the console disagree about step 1. edits:commit on a track release returns 403 ("You must let us know whether your app uses any Foreground Service permissions") for the same bundle that the console's own release flow reports as "Ready to release" and then happily publishes. So the ship script can upload the bundle but cannot roll it out. The first build carrying a new FGS type has to go out from the console by hand.
That last one is worth generalising: if an automated path fails on a Play policy check, try the console before you debug the API. They are not the same code path and they do not enforce the same gates.
What this actually trades
None of these are hard problems. Each is a day, or a release cycle, spent on something that is not the app. And they cluster around on-device inference specifically: the size is the model, the page alignment is the native runtime, the foreground service exists because inference takes minutes instead of a round trip.
The cloud version of this app has none of them. It uploads a file and polls an endpoint. That's the real trade, and it isn't "on-device is harder to build". That's roughly untrue now that ONNX Runtime ships usable mobile binaries and the export paths are public. It's that on-device puts you in the corners of both stores that get the least traffic, and therefore the least documentation, the least Stack Overflow, and the least sympathy from the reviewer.
Worth it here, because zero marginal cost per separation is the entire product. It's what lets the free tier carry no track-length limit at all, which is the one thing a competitor paying for GPU time per minute of audio cannot copy at any price.
But it is a bill, and it's better to know the amount before you commit.