Ling 3.0 Tiny is now available on AI Gateway
Ling 3.0 Tiny from ANT Group is now available on Vercel AI Gateway. The model is free to use until 8:00am PT on 8/14. It takes the free slot previously held…
By Dillip Chowdary • Aug 06, 2026 • Source: Vercel Blog
Ling 3.0 Tiny from ANT Group is now available on Vercel AI Gateway. The model is free to use until 8:00am PT on 8/14. It takes the free slot previously held by Ling 3.0 Flash.
Ling 3.0 Tiny is a mixture-of-experts (MoE) model with 7.9B total parameters and about 1.3B active per token. It supports a 256K token context window and up to 32K output tokens. The design targets responsive agents and strong instruction following, with a small active parameter budget per token so inference stays light while the larger total parameter pool preserves capacity.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers and builders on AI Gateway, a free MoE with long context and large max output is useful for agent loops, tool-calling chains, and multi-step instruction work without spinning up a separate provider path. The 1.3B active-per-token profile is the practical lever: lower per-token compute than dense models of similar total size, while 256K context covers long transcripts, codebases, or multi-document agent state.
The free-slot rotation from Ling 3.0 Flash to Ling 3.0 Tiny is a product signal: AI Gateway is cycling ANT Group’s Ling 3.0 line rather than leaving one model pinned indefinitely. Tiny’s MoE shape (7.9B total / ~1.3B active) positions it as the lightweight, high-throughput option in that lineup for gateway users who want cost-free experimentation before the free window ends.
Use Ling 3.0 Tiny on AI Gateway while it is free, through 8:00am PT on 8/14. Validate agent latency, instruction adherence, and long-context behavior against your own workloads before the free slot moves again. After that window, check whether Tiny remains on the gateway under paid routing or whether another Ling 3.0 variant takes the free seat.
Advertisement
🔎 More interesting news
- Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
- Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill
- Loop Engineering with native model switching in Codex and Claude
- Show HN: Reduck GEO, open source Skill to measure and optimize Claude citations
- Today's full Tech Pulse briefing →