Ling 3.1 Flash: Ant Group's 560B MoE Model Explained
Updated 2026-10-04
Ling-3.1-Flash is a large language model from Ant Group's Ling (百灵) team, the group that publishes open models under the inclusionAI name. It was announced on September 30, 2026 as a mixture-of-experts (MoE) model with about 560B total parameters and about 25B active per token, and a context window of up to 1M tokens. At launch it was offered as a two-week free trial (with a 256K context limit), followed by a planned paid service with the full 1M context; the team said it plans to open-source the weights, but as of October 4, 2026 no weights or licence have been published.
Ling 3.1 Flash key facts
| Ling-3.1-Flash | |
|---|---|
| Developer | Ant Group, Ling (百灵) team / inclusionAI |
| Announced | September 30, 2026 |
| Type | Text LLM, mixture-of-experts |
| Parameters | ~560B total, ~25B active per token (vendor-reported) |
| Context window | Up to 1M tokens (256K during the free trial) |
| Access at launch | Two-week free trial, then paid access |
| Open weights | Planned, not yet released (as of Oct 4, 2026) |
| Licence | Not yet published |
Important: Ling 3.1 Flash is a text model, not an image or video generator. If you were looking for a video tool, you probably mean Kling, Kuaishou's AI video model – a different product from a different company.
What Ant Group says it is built for
According to the launch coverage, the team tuned Ling-3.1-Flash for:
- General-purpose agents – multi-step tasks with tool calls
- Search – answering with retrieved information
- Office work – documents, spreadsheets and writing
- Software development – coding assistance
- Domain research – the announcement specifically names medical, financial and materials-science scenarios
Ant published a benchmark comparison chart with the launch, but these results are vendor-reported. Independent evaluations (for example on Artificial Analysis) had not appeared for 3.1 at the time of writing, so we don't repeat score claims here.
Ling 3.1 Flash vs Ling 3.0 Flash
Ling-3.1-Flash is a much larger model than its predecessor, released only about two months earlier:
| Ling-3.0-Flash | Ling-3.1-Flash | |
|---|---|---|
| Released | July 27, 2026 | September 30, 2026 |
| Total / active parameters | 124B / 5.1B | ~560B / ~25B |
| Context | 256K native, extendable to 1M | Up to 1M |
| Weights | Open on Hugging Face (inclusionAI), MIT licence | Planned, not yet published |
| API | OpenRouter, Vercel AI Gateway and other routers | Trial first; third-party APIs not yet listed |
The two launches follow the same pattern: Ling-3.0-Flash was free on OpenRouter and Vercel AI Gateway until August 3, 2026, and the weights were open-sourced after that free window. If 3.1 follows the same path, open weights would arrive after the trial – but that is a plan, not a release, until Ant publishes the repository.
How to access Ling 3.1 Flash
Step 1: Use the free trial
During the launch window Ant offered free use of Ling-3.1-Flash for two weeks with a 256K context cap. Coverage points to Ant's own Ling channels (including the Ling Studio app) as the place to try it. Check the inclusionAI pages on Hugging Face or ModelScope, which link to the official access points, rather than unofficial mirrors.
Step 2: Watch for paid API access
After the trial, Ant said the model moves to a paid service with the full 1M context. At the time of writing no per-token price had been published and the model was not listed by major API routers. Don't trust a price you see on a reseller page until it matches an official source.
Step 3: Self-host when the weights land
When inclusionAI publishes the weights, they will most likely appear under the inclusionAI organization on Hugging Face and ModelScope, like the 3.0 models (which also had FP8 versions). Keep in mind that a ~560B MoE model needs multi-GPU server hardware even though only ~25B parameters are active per token – all experts still have to be stored in memory. Check the licence in the model card before commercial use; 3.0 used MIT, but 3.1's licence is not yet known.
Need something you can use today?
Ling-3.0-Flash is open (MIT) and available on Hugging Face and through API routers, so it is a practical stand-in while 3.1 is in trial. For hosted alternatives from other vendors, see our explainers on GPT-6.1 Sol, Claude Opus vs Sonnet and Gemini for coding.
Who should care about Ling 3.1 Flash?
- Teams that want open-weight models for agents and long documents – once released, a 1M-context open MoE model is a serious option for self-hosting.
- Developers in China or using Chinese cloud stacks – Ant's Ling models are distributed on ModelScope as well as Hugging Face.
- Cost-sensitive API users – Ling-3.0-Flash was priced aggressively on routers; watch whether 3.1 follows.
FAQ
What is Ling 3.1 Flash?
Ling-3.1-Flash is a mixture-of-experts language model from Ant Group's Ling team (inclusionAI), announced on September 30, 2026, with about 560B total parameters, about 25B active per token, and up to a 1M-token context window.
Is Ling 3.1 Flash an official AI model?
Yes. It was announced by Ant Group's Ling (百灵) team on September 30, 2026. Its specifications are vendor-reported, and its weights and licence had not been published as of October 4, 2026.
Is Ling 3.1 Flash free?
At launch Ant offered a two-week free trial limited to 256K context. After the trial it is planned to become a paid service with the full 1M context. If the weights are open-sourced, self-hosting will be free apart from your hardware costs.
Is Ling 3.1 Flash open source?
Not yet. The team said it plans to open-source the model, but as of October 4, 2026 there is no public repository or licence. Its predecessor, Ling-3.0-Flash, is open under the MIT licence on Hugging Face.
Is Ling 3.1 Flash an image or video model?
No, it is a text model for chat, agents, search, office work and coding. "Kling" is Kuaishou's AI video generator – a separate product that is easy to confuse with Ling.
What is the difference between Ling 3.1 Flash and Ling 3.0 Flash?
Ling-3.1-Flash is roughly 4.5 times larger (about 560B vs 124B total parameters, about 25B vs 5.1B active) and launched with up to 1M context. Ling-3.0-Flash is already open-source and available through APIs, while 3.1 started with a trial.