Skip to content

Kimi-VL-A3B-Thinking-2506, Explained Simply

What Kimi-VL-A3B-Thinking-2506 is, what it does well, and what it costs — released 2025-06-21 by Moonshot AI.

TLDR
  • It is a free, open-source AI model made by Moonshot AI that can read text, look at pictures, and watch videos to solve tricky problems.
  • It is an upgraded version of their older "Thinking" model, but it uses 20% fewer words to "think" through answers while getting better grades on tests.
  • It can now see super high-resolution images and is great at helping navigate computer screens.

Want a heads-up when Kimi updates?

We email you the moment a new Kimi version drops — plus plain-English release notes like these. No spam.

Choose update emails Kimi · Vibe Mastermind Updates
Choose which update emails you want

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Kimi-VL-A3B-Thinking-2506 is a newly updated AI model created by Moonshot AI. It is a "vision-language" model, which means it can understand both written text and visual things like photos or videos. It is built to be highly efficient, meaning it runs on less computing power than many larger models. This makes it a great tool for students, developers, or curious users who want a smart assistant that can reason through tough math and science problems without needing a massive, expensive computer to run.

What it does well

  • Multimodal reasoning: It can look at a picture or video and use it to solve logic or math problems. It scored huge improvements on visual math tests compared to its older version.
  • Screen navigation: It is really good at understanding what is on a computer screen and figuring out where to click or what to do, making it useful as a computer assistant (often called an "agent").
  • Video understanding: It can watch video clips and answer questions about what is happening in them, beating many other open-source models on video reasoning tests.
  • Text-only smarts: Even though it is built for images and video, it is also great at standard text subjects, scoring an 82.0 on a general knowledge test (MMLU) and a 91.8 on a text math test (MATH).

How it compares

  • Versus the previous Kimi model: The older version was focused mostly on "thinking" (reasoning through problems step-by-step), but it struggled a bit with basic image understanding. This new version keeps the deep reasoning but is now just as good at everyday visual tasks as their non-thinking model. It also cut down its "thinking length" by 20%, meaning it gets to the answer faster.
  • Versus rivals: When put head-to-head with other open-source models of similar or even much larger sizes (like Qwen2.5-VL-7B, Gemma3-12B-IT, and Qwen2.5-VL-72B), Kimi often comes out on top, especially in visual math, video reasoning, and finding things on a computer screen. In several tests, it even matches or beats the scores of GPT-4o.

What it costs and what it can handle

  • Resolution: It can handle images up to 3.2 million total pixels. That is four times sharper than the previous version, allowing it to read fine details in big photos.
  • Output length: It can generate up to 32K "tokens" (chunks of words) in a single response, which is very long and allows it to think deeply about complex problems.
  • Price: The model is open-source, meaning developers can download and run it for free on their own hardware.

Worth knowing

  • Hardware requirements: Even though it is considered "efficient" (the "A3B" in its name means it uses about 3 billion active parameters), you still need a decent computer with a good graphics card to run it locally.
  • Setup: The creators recommend using a specific software tool called VLLM to run it, which might be a bit technical for someone without coding experience.

This is our plain-language summary. Read the complete, official notes on the Kimi official changelog ↗.

Vibe Mastermind in 00d 00h 00m 00s
Call LIVE — Join Now