
Kimi-VL is a lightweight Mixture-of-Experts vision-language model that activates only 2.8B parameters per step while delivering strong performance on multimodal reasoning and long-context tasks. The Kimi-VL-A3B-Thinking variant, fine-tuned with chain-of-thought and reinforcement learning, excels in math and visual reasoning benchmarks like MathVision, MMMU, and MathVista, rivaling much larger models such as Qwen2.5-VL-7B and Gemma-3-12B. It supports 128K context and high-resolution input via its MoonViT encoder.
Modalities
Context
131K
Released
Apr 10, 2025
Knowledge Cutoff
Dec 2024
Kimi-VL is a lightweight Mixture-of-Experts vision-language model that activates only 2.8B parameters per step while delivering strong performance on multimodal reasoning and long-context tasks. The Kimi-VL-A3B-Thinking variant, fine-tuned with chain-of-thought and reinforcement learning, excels in math and visual reasoning benchmarks like MathVision, MMMU, and MathVista, rivaling much larger...
Kimi VL A3B Thinking has a 131,072 token context window.
Kimi VL A3B Thinking accepts images and text as input and returns text.
Kimi K3, Kimi K2.7 Code, Kimi K2.6 and 4 more are other text models from MoonshotAI.
Kimi VL A3B Thinking was released on April 10, 2025. Its knowledge cutoff is December 31, 2024.
Token volume and request traffic to this model over time.