🔍 Read the full analysis: Can LFM2.5-VL-DSpark Speed Up Vision-Language Models? on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Liquid AI released LFM2.5-VL-DSpark, an experimental 280M-parameter speculative-decoding drafter for its LFM2.5-VL-3B vision-language model. According to the company, it delivers decoding speedups of up to 3.13x on Apple silicon and 2.66x on H100 without changing outputs, with day-one llama.cpp, MLX-VLM, and SGLang support.
Liquid AI has released LFM2.5-VL-DSpark, an experimental draft model that accelerates inference of its open-weight LFM2.5-VL-3B vision-language model through speculative decoding. According to the company, the 280M-parameter drafter adds roughly 8.9% to the target model’s parameter count while delivering decoding speedups of up to 3.13x on Apple silicon and up to 2.66x on an NVIDIA H100 — without changing output quality under greedy decoding. The model is available now on Hugging Face in Safetensors and GGUF formats.
The drafter extends Liquid AI’s DSpark recipe — previously applied to its text-only LFM2.5 models — to a multimodal target. It captures the target model’s hidden states at a fixed set of tapped layers and drafts blocks of candidate tokens conditioned on those states. Because image patches and text tokens are projected into a shared representation before the tapped layers, the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality, and the inference algorithm is unchanged from the text models.
Trained on a mixture of vision-language supervised fine-tuning data weighted toward expected serving workloads, the final drafter is a simplified attention-only model with 4 layers, selected via ablations across 3, 4, and 5 layers, with a block size of 9. Liquid AI reports that acceptance improved over 10 training epochs before reaching diminishing returns. The drafter comprises a 193.0M-parameter decoder stack, a 21.0M hidden-state projection, a 65.5M Markov head, and roughly 6.4k parameters in norms and a confidence head — about 279.5M in total.
Benchmarks follow the MMSpec protocol across six vision tasks: general VQA, text VQA, image captioning, chart VQA, complex reasoning, and multi-turn conversation. On-device, with MLX on an M5 Max, decoding runs 2.30x to 3.13x faster and end-to-end latency improves 1.56x to 2.62x. With llama.cpp on an M3 Ultra, decoding improves 1.57x to 2.14x, end-to-end 1.30x to 1.77x. On H100, the company reports decoding speedups ranging up to 2.66x and end-to-end gains of 1.64x to 2.27x, using a DSpark block size of 8.
Faster Vision AI on Consumer Hardware
The release targets a persistent bottleneck for local and edge AI: vision-language models are slower than text models because images must pass through a vision encoder and then be processed as hundreds of visual tokens alongside the text prompt. A drafter that roughly doubles or triples decode speed — for under 9% more parameters — could make strong>3B-class multimodal models practical on laptops and phones, where users feel latency directly.
Day-one integration matters as much as the speedup itself. DSpark support in llama.cpp, MLX-VLM, and SGLang means the acceleration works in the toolchains hobbyists and deployers already use, rather than requiring custom inference code. Combined with Liquid AI’s open-weight licensing — download, fine-tune, and deploy without restrictions, per the company — the release positions the LFM2.5 family for on-device multimodal applications.
Speculative decoding is exact: the target model verifies every proposed token, so greedy output matches the target alone. The company is candid about the ceiling — speculative decoding accelerates only the decode phase, not vision encoding or prefill, so end-to-end gains trail decode gains on edge devices.
From Text Drafters to Multimodal
Speculative decoding is an established technique in which a small, fast “draft” model proposes candidate tokens that the larger target model verifies in batch, accepting matching tokens and discarding the rest. Because verification is cheaper than sequential generation, accepted drafts translate into net speedup with identical outputs under greedy decoding.
Liquid AI released its first LFM2.5-DSpark drafter models for text-only LFM2.5 targets earlier in 2026. The new vision drafter reuses the same architecture and inference algorithm, differing mainly in training data and the shared multimodal representation. The LFM2.5 family spans base models, audio, and vision variants, with the 3B vision model positioned as an edge-capable multimodal option.
“It adds a speculative decoding path that trades a minimal increase in memory footprint for a larger speedup without changing output quality.”
— Liquid AI, announcement post
Experimental Status and Benchmark Gaps
The release is explicitly labeled experimental, and the company has not indicated when or whether the drafter will be promoted to a stable release. The reported speedups are Liquid AI’s own measurements, not independently verified third-party benchmarks, and results will vary with hardware, prompt composition, and image resolution.
One figure in the company’s GPU results appears internally inconsistent — a lower bound described as “20.4x” in a range stated as “20.4x to 2.66x” — which reads as a typo, likely for 2.04x; Liquid AI has not clarified the figure. It is also unclear how acceptance rates behave on out-of-distribution vision tasks, how the drafter affects sampling-based (non-greedy) generation quality, and what the memory footprint increase is in runtime terms beyond parameter count. Details of the training data mixture have not been published.
Community Uptake and Clarifications
The model and required integration patches are public: SGLang support requires a build with DSpark for LFM2.5 targets (PR #40651), llama.cpp requires PR #29339, and MLX-VLM requires PR #2280. Early adopters will likely publish independent benchmarks, which will test the company’s reported speedup ranges on varied hardware and workloads. Whether Liquid AI promotes the drafter from experimental to stable status, and whether it clarifies the inconsistent GPU figure or publishes training-data details, remains to be seen.
Key Questions
Does DSpark change the model’s outputs?
No, according to Liquid AI. Speculative decoding is exact: the target model verifies every proposed token, so greedy output matches the target model alone. Effects on sampling-based (non-greedy) generation are not yet documented.
How much faster is it, and on what hardware?
Per the company’s own benchmarks: decoding runs 2.30x–3.13x faster on an M5 Max with MLX, 1.57x–2.14x on an M3 Ultra with llama.cpp, and up to 2.66x on an H100. End-to-end latency gains are lower, ranging 1.30x–2.62x across setups.
Is the model free to use?
Yes. The drafter is available on Hugging Face in Safetensors and GGUF formats, and Liquid AI says its open-weight licensing permits download, fine-tune, and deployment without restrictions.
Why are end-to-end gains smaller than decoding gains?
Speculative decoding accelerates only the decode phase, not vision encoding or prefill. On edge devices those stages take a larger share of total time — an Amdahl’s-law limitation that Liquid AI itself highlights.
How do I enable DSpark in my inference stack?
Integration requires pending patches: PR #29339 for llama.cpp, PR #2280 for MLX-VLM, and a build with DSpark for LFM2 targets (PR #40651) for SGLang.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
