← Back to live feed · 1 stories across 1 day

Wednesday, Sep 23, 2026

1 story
1
Rigel 2.3B Model Rivals Llama-3.2-3B With <1% Pretraining Compute
topics 🤖 AI tags AIAI ModelsAI Research keywords Mayank Mishkin

A new machine learning architecture called Rigel has produced performance results comparable to Meta's Llama-3.2-3B using a reduced compute footprint. Developed by Mayank Mishkin and a research team, the 2.3B parameter Mixture-of-Experts (MoE) model employs a Hybrid Mamba-2 design with 360M active parameters to score within a few points of the Meta model while using less than 1% of its pretraining floating-point operations (FLOPs).

The model was developed without a dedicated compute cluster, instead utilizing a "guerrilla style" strategy that shifted a single codebase across fragmented resources. This training run moved between Nvidia H100, A100, and V100 GPUs as well as TPU v5p and v6e processors, demonstrating a method for conducting large-scale academic pretraining without unified supercomputing hardware.

Image via @mayankmish98 on X
Earlier version from Tuesday, Sep 22
Rigel AI Hits Llama 3.2 Performance Using Under 1% of FLOPs
4 tweets • 4 sources
See all 6 tweets →