An instruction-tuned, hybrid-reasoning Mixture-of-Experts model built on Llama-4-Scout-17B-16E. Cogito v2 can answer directly or engage an extended “thinking” phase, with alignment guided by Iterated Distillation & Amplification (IDA). It targets coding, STEM, instruction following, and general helpfulness, with stronger multilingual, tool-calling, and reasoning performance than size-equivalent baselines. The model supports long-context use (up to 10M tokens) and standard Transformers workflows. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs(opens in new tab)
Modalities
Context
131K
Released
Sep 2, 2025
Knowledge Cutoff
Aug 2024
An instruction-tuned, hybrid-reasoning Mixture-of-Experts model built on Llama-4-Scout-17B-16E. Cogito v2 can answer directly or engage an extended “thinking” phase, with alignment guided by Iterated Distillation & Amplification (IDA).
Cogito V2 Preview Llama 109B has a 131,072 token context window.
Cogito V2 Preview Llama 109B accepts images and text as input and returns text.
Cogito v2.1 671B is another text model from the same author.
Cogito V2 Preview Llama 109B was released on September 2, 2025. Its knowledge cutoff is August 31, 2024.
Token volume and request traffic to this model over time.