Liquid AI Launches Zero-Token d1 Decision Models
Liquid AI has released its open-weight d1-3B and d1-omni-600M models, introducing a zero-token decision architecture designed to run real-time classification and routing at the edge.

Liquid AI has launched two open-weight multimodal decision models, d1-3B and d1-omni-600M, designed for real-time decisions on NVIDIA DGX, RTX, and Jetson edge boards. Unlike generative models that write answers token by token, these decision models evaluate a state and questions to return calibrated probability distributions in a single forward pass with zero output tokens. Practitioners can use them for routing, moderation, and agent guardrails. Both models are on Hugging Face with llama.cpp support under the LFM Open License v1.0, which is free for commercial use below 10 million dollars in annual revenue.
The d1-3B model features 3.12 billion parameters, created by averaging the weights of LFM2.5-2.6B with the text backbone of LFM2.5-VL-3B. It integrates a 400 million parameter SigLIP2 NaFlex vision encoder and supports a 32,768-token context window. On the Decision Index v0.2.1, d1-3B scored 48.57, outperforming Decider 35B-A3B at 47.11, while trailing Winnow-12B at 50.02. It led the Tools category at 74.5 and Arts at 36.3, but trailed on Knowledge at 23.8. It averaged 82.9 across seven text benchmarks, ahead of Decider 4B at 81.1, and scored 74.1 across 11 image benchmarks. End-to-end latency for one question is 8 milliseconds on an RTX 4090 and 9 milliseconds on an AMD MI325X. On Jetson AGX Thor, it takes 16 milliseconds for one question and 20 milliseconds for three, while AGX Orin takes 26 milliseconds and Orin Nano takes 50 milliseconds.
The smaller d1-omni-600M model contains 587 million parameters and is built on the LFM2.5-Encoder-350M. It features a 381 million parameter shared trunk, a 94 million parameter vision encoder, and a 112 million parameter audio encoder utilizing a 17-layer FastConformer. With a 16,384-token context window, it processes text alongside either an image or an audio clip of up to 30 seconds. It scored 15.95 on the Decision Index v0.2.1 and averaged 78.4 across seven text benchmarks, achieving scores of 95.8 on Civil Comments and 79.5 on PAWS-X.
This is our own summary of reporting by MarkTechPost



