IT Home On October 9, Liquid AI posted a blog post on October 7, announcing the launch of two open-weight decision-making models in the d1 series: d1-3B supports text and image classification with 3.12 billion parameters, and d1-omni-600M supports text, image, and voice input with 587 million parameters.
According to a blog post on IT Home, after the popular launch of the Jev model in September, companies such as OpenAI also introduced similar decision-making models. Liquid AI is a new member joining this battle. The newly released d1-3B and d1-omni-600M have been made public on Hugging Face, allowing users to download, fine-tune, and deploy them.
d1-3B has 3.12 billion parameters and is trained based on the visual language model LFM2.5-VL-3B. It supports text and image input. The model scored 48.57 points in the public test set of Decision Index v0.2.1, and Liquid AI claims it leads all models below 10B parameters.

d1-omni-600M has 587 million parameters and was trained using the bidirectional encoder LFM2.5-Encoder-350M. This model supports combined input of text and images, or text and speech, and is the first experimental multimodal decision-making model checkpoint of Liquid AI.

In terms of inference speed, d1-3B takes 8 milliseconds to answer a single question on the NVIDIA RTX 4090, and 102 milliseconds to process a 384-pixel image. On the Jetson AGX Thor, it takes 16 milliseconds for a single question; on the Jetson Orin Nano, it takes 50 milliseconds.

User tests show that d1-3B has a single judgment time of 25 milliseconds on a GeForce RTX 3060 with 12GB of VRAM, and takes 146 milliseconds to process a 640×480 pixel image, with a peak VRAM usage of approximately 6GB. Other users reported that the model has a single judgment time of 8 milliseconds on the RTX 3090.

