TMCnet News
Ant Group Unveils Ling-3.0-Flash Delivering Top-Tier Performance at a Fraction of the Parameter ScaleAnt Group today announced the release of Ling-3.0-Flash, a next-generation native hybrid-reasoning foundational model engineered specifically for production-grade AI agent workflows. Designed to deliver rapid response capabilities, it serves as a high-speed execution node that offers a superior balance of intelligence density and cost-efficiency. This press release features multimedia. View the full release here: https://www.businesswire.com/news/home/20260726584441/en/
Ling-3.0-Flash delivers strong performance across multiple core benchmarks. Featuring 124B total parameters with only 5.1B active parameters per token, Ling-3.0-Flash achieves remarkable performance despite its streamlined footprint. It matches or surpasses industry-leading models with two to three times its parameter scale across core benchmarks, including foundational reasoning, instruction following, and long-context processing. Architectural Innovation for Efficiency Ling-3.0-Flash moves away from the traditional approach of simply scaling parameter counts. Instead, it is built from the ground up with a native hybrid-linear attention architecture. By alternating KDA (Kimi Delta Attention) and MLA layers at a 5:1 ratio, the model optimally balances long-context efficiency with robust state memory. Key architectural advancements include:
Rather than aiming to replace ultra-large, general-purpose reasoning models, Ling-3.0-Flash is designed to complete the "planning-execution separation" paradigm in AI workflows. It serves as a cost-controllable, fast, and highly stable execution node, delegating deep planning and high-frequency execution to specialized models. To support this, Ling-3.0-Flash has been deeply refined for real-world agent scenarios, expanding its training to over 10,000 interactive environments. It features enhanced self-correction and long-horizon planning mechanisms, enabling autonomous, end-to-end delivery in complex tasks such as coding, task decomposition, and deep multi-source research. This resolves common issues of deviation or context loss in traditional models during large-scale operations. Engineering for Speed and Stability To ensure fast and reliable agent performance, Ant Group has paired Ling-3.0-Flash with a supporting engineering and collaboration architecture:
Ling-3.0-Flash is now available on OpenRouter and Vercel AI Gateway, offering a free API through August 3, 2026. Following this limited-time free access period, the model weights will be open-sourced to support further development and innovation within the global AI community. Developers are encouraged to integrate Ling-3.0-Flash into their coding, search, research, and tool-use workflows to experience its high-speed execution and stable tool-calling capabilities. About Ant Group Ant Group is a global digital technology provider and the operator of Alipay, a leading internet services platform in China, connecting over one billion users to more than 10,000 types of consumer services from partners. Through innovative products and solutions powered by AI, blockchain and other technologies, Ant Group supports partners across industries to thrive through digital transformation in an ecosystem for inclusive and sustainable development. For more information, visit www.antgroup.com.
View source version on businesswire.com: https://www.businesswire.com/news/home/20260726584441/en/ |

