Liquid AI releases LFM2.5-VL-DSpark for faster vision-language model inference
Liquid AI announced LFM2.5-VL-DSpark, a draft model for its LFM2.5-VL-3B vision-language model. The model uses speculative decoding to speed up inference by up to 3.13x on-device and 2.66x on H100 GPUs, with end-to-end gains of 2.62x and 2.27x, respectively. It adds 280M parameters (8.9% increase) but maintains output quality, according to the company’s benchmarks across tasks like VQA,…
Key points
- LFM2.5-VL-DSpark speeds up vision-language model inference by up to 3.13x on-device and 2.66x on H100 GPUs
- Adds 280M parameters (8.9% increase) to LFM2.5-VL-3B with no output quality trade-off, per Liquid AI
- Supports day-one integration with llama.cpp, MLX-VLM, and SGLang for edge and GPU deployment
The model supports day-one integration with llama.cpp, MLX-VLM, and SGLang, targeting edge and GPU deployments. Liquid AI emphasizes its open-weight approach, allowing unrestricted fine-tuning and deployment. The release aligns with the lab’s goal of AI running across devices, from base models to specialized variants like audio and vision.
Model page: LFM2.5-VL-DSpark →
The story so far
3 episodes →- Liquid AI releases LFM2.5-VL-DSpark for faster vision-language model inferencethis story
Accelerating vision-language models with LFM2.5-VL-DSpark
Hugging Face · 24 September 2026
Loading the full article…
This text was published by Hugging Face. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- Anthropic releases Claude Opus 5.5; OpenAI halves prices for GPT-6 Sol and Luna · 15 src
- OpenAI releases ChatGPT Images 2.5 with sketch, point, and template tools · 4 src
- FreedomIntelligence releases HuatuoGPT-3 multimodal model on Hugging Face · 1 src
- PrismML launches 1-bit Bonsai LLM for Qualcomm smart glasses · 1 src
- BottleCap AI cuts Qwen3.8-27B’s reasoning tokens by 37.2% with ThinkingCap-Qwen3.8-27B · 1 src
Comments
via GitHub Discussions