In this guide
Short answer
ML Drift deserves attention from developers maintaining on-device AI, particularly those planning a LiteRT GPU migration. It is infrastructure for running models, not a new chatbot. The buying or engineering decision should rest on device-specific correctness, responsiveness and resource use.
What was released
Google's October 8 announcement describes ML Drift as an Apache 2.0 GPU compute engine supporting OpenGL ES, OpenCL, Metal and WebGPU. It powers LiteRT acceleration and is also available as a standalone library. Google highlights unified shaders, custom operators, five-dimensional tensors and inference-stage optimizations.
The company says the legacy TFLite GPU delegate will stop receiving new features and encourages migration. Desktop expansion is presented as previews. Google supplies performance examples, but we have not reproduced them and they do not establish a universal improvement across devices or workloads.
Why users may eventually notice
Our interpretation is that the practical benefit could be fewer delays in a local workflow: applying an effect, summarizing a short passage or processing a camera frame. The relevant unit is the complete interaction, including initialization and post-processing. Fast model execution can coexist with a slow-feeling app.
A local feature must also have a clear fallback when the device cannot support it. Decide whether the app should use a smaller model, show a limitation or request permission for remote processing. Avoid silently changing the data path when performance is disappointing.
Proposed migration checklist
Keep the existing runtime as a baseline. Freeze the model, input examples and expected outputs before changing the accelerator. Include malformed input, long sessions and background interruptions as well as a short ideal example.
Build a device matrix around what your users actually own. Test an older phone, a low-memory device and representative desktop hardware where relevant. Record cold-start time, steady response time, peak memory, battery behavior and output differences. Run sustained sessions so thermal effects have a chance to appear.
Count crashes and fallbacks alongside completed tasks. Do not report only the best device or fastest trial. If quality changes, investigate precision, preprocessing and unsupported operations before attributing the difference to the model. These are proposed engineering checks, not a DiscoverAI benchmark.
Local inference and privacy are separate claims
An accelerator running locally tells you where that computation happens. It does not tell you whether the application uploads inputs, synchronizes generated content or records identifying telemetry.
Map the complete path from capture to deletion. Inspect network behavior with authorized test material. Check what happens when a feature falls back, when a user signs in and when diagnostic reporting is enabled. Document any remote step in the user-facing explanation.
A practical decision
Proceed when representative devices preserve expected output and the full interaction improves within your resource budget. Use a reversible rollout and keep enough diagnostics to compare versions. A migration that only improves the benchmark device may not improve the product.
For a related local-search development, read our [EmbeddingGemma 2 analysis](/articles/embeddinggemma-2-local-multimodal-search-2026). For the broader question of accepted outcomes, see the [Haiku task-cost guide](/articles/claude-haiku-5-5-task-cost-2026).
Transparency
How this guide was checked
Editorial accountability
Who checked this guide
- Evaluation type
- Research-based verification
- Last materially checked
- Evidence
- 1 listed sources
Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.
Editorial basis
What this guidance is based on
- Editorial basis
- Source-led analysis
- Primary references
- 1
- Products covered
- 0
- Last checked
- 2026-10-09
Important limits
- • DiscoverAI has not independently tested the product or reproduced vendor benchmarks.
- • Availability, prices and policies may change. Evaluation exercises are proposed reader-run tests.
Sources and verification
Product details and claims were checked against the following primary sources.
Frequently asked questions
Is ML Drift an AI model?
No. It is an on-device GPU compute engine for AI and machine-learning inference.
What license does Google describe?
The announcement describes an Apache 2.0 open-source release.
Does local inference make the whole app private?
No. Uploads, telemetry, storage and connected services need separate checks.
Should teams assume desktop support is equally mature?
No. Google's announcement labels the desktop expansion as previews; evaluate the specific runtime and platform.
Read next

