d1-3B and d1-omni from LiquidAI
preview.redd.it/owhrvbbiu2uh1.png d1-omni-600M d1-omni-600M is a 600M parameter decision model built on LFM2.5-Encoder-350M . You give it a state (text or JSON, with images or a voice clip) and a set of named questions. It returns typed answers with zero output tokens : every answer is read directly from the model's distribution over the options, with no generation and no parsing. Vision-language : text and images (tiled for large frames, several images per state) in a single forward pass. Audio-language : text and up to 30 s of speech in a single forward pass. Edge-sized : 587M parameters: a 381M shared trunk and decision head, a 94M vision encoder and a 112M audio encoder. Every modality runs the same trunk weights. preview.redd.it/f2wkfqcku2uh1.png d1-3B d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B . You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens . Best decision model under 10B on the Decision Index 0.2.1 : 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11). Multimodal : images and text in the same state. It scores 74.1 on 11 public image benchmarks (LFM2.5-VL-3B: 73.9). Fast : 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro. huggingface.co/LiquidAI/d1-3B-GGUF huggingface.co/LiquidAI/d1-3B huggingface.co/LiquidAI/d1-omni-600M-GGUF huggingface.co/LiquidAI/d1-omni-600M preview.redd.it/e11c4qlut2uh1.png