The DecoderOriginal · English

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

Reka AI's Rho-1 is a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single neural network. Trained on 320 H100 GPUs in about three months, it uses a…

Matthias Bastian
Image source · The Decoder

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

Matthias Bastian Matthias Bastian View the LinkedIn Profile of Matthias Bastian Oct 5, 2026

Reka AI has released a research preview of Rho-1. The 19-billion-parameter omni-model processes and generates text, images, video, and robot control actions in a single neural network. Unlike most AI systems that route tasks to specialized models, Rho-1 runs all modalities as tokens in one shared context window with no tool calls or external models. The model generates continuous video in real time and responds to new instructions on the fly without restarting.

The same weights that predict camera images also drive robot movements. To work around scarce robot training data, Reka AI built an inverse dynamics model that pulls control signals from ordinary internet videos. Rho-1 trained on 320 H100 GPUs over about three months.

Reka AI isn't new to multimodal AI. In April 2024, the company shipped Reka Core, a multimodal language model that competed with GPT-4, Claude 3, and Gemini Ultra on benchmarks. The release fits a broader push in AI research toward so-called world models.Ad

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Source: Reka AI
Original source

The Decoder

Content notes

Original publication and rights belong to the source.