Research

Ai2 unveils Olmo Hybrid model and AstaBrief at COLM 2026

The Allen Institute for AI showcased its Olmo Hybrid architecture at COLM 2026, demonstrating that blending attention with recurrence can halve model training costs without losing performance.

AI22 days agoResearch
Image: AI2

At the COLM 2026 conference in San Francisco, the Allen Institute for AI (Ai2) showcased breakthroughs in its foundational model architectures and scientific tooling. Among the key developments presented was the paper "Olmo Hybrid: From Theory to Practice and Back," which details an architecture combining transformer attention mechanisms with linear recurrent layers. While attention retrieves specific details across context, recurrence maintains a compact memory state that updates as tokens process. In empirical tests, this hybrid design reached identical accuracy to the Olmo 3 7B baseline on the MMLU multi-subject benchmark while consuming 49% fewer training tokens.

Ai2 is leveraging these findings to build its next generation of models using a hybrid mixture-of-experts system that routes tokens through specific sub-components for greater efficiency. To support community research, the lab open-sourced Olmo-core 3, an infrastructure codebase redesigned to facilitate training large mixture-of-experts systems alongside open checkpoints and technical reports. Highlighting the benefits of open scientific artifacts, Google recently validated Ai2's methodology by reproducing the Olmo 3 7B training run inside MaxText on Google Cloud TPUs.

Expanding beyond foundation models, Ai2 highlighted its Asta agentic platform created to assist domain experts. The initiative includes AstaBrief, an open-weights model capable of writing fully cited reports by querying literature bases. Scientists can run AstaBrief locally or utilize it through Asta's Fast mode interface. Additionally, the team celebrated the publication of "Retrofitting language models to operate over bytes" in Nature. The work outlines Bolmo, a model operating directly on raw bytes rather than traditional subword tokenization.

Senior director of NLP research Noah A. Smith and communications lead Kyle Wiggers discussed these developments during the event, emphasizing how byte-level processing avoids tokenization errors. Senior research scientist Bodhisattwa Prasad Majumder also led discussions with academic communities to identify current limitations in general-purpose models, reinforcing Ai2's non-profit mandate to share data, weights, and training code publicly.

This is our own summary of reporting by AI2

More in Research