AI Gateway: GPT-5.6 pricing and speed updates
By Dillip Chowdary • Aug 03, 2026 • Source: Vercel Blog
Vercel’s AI Gateway has updated GPT-5.6 routing: GPT-5.6 Luna and GPT-5.6 Terra are now cheaper, and GPT-5.6 Sol is faster. The changes cover both short-context and long-context token pricing. Because AI Gateway adds no markup on token pricing, the new rates and speed characteristics pass through at the upstream rate.
On the product side, AI Gateway sits between your app and the model provider. Token cost is not inflated by the gateway layer, so a price cut or speed change from the upstream model shows up as the same unit economics in your gateway bill. That applies whether a request is billed under short-context or long-context pricing, so long-running or large-window calls are not left on an older rate structure while short calls move.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers and builders already on AI Gateway, Luna and Terra become cheaper default choices for workloads where cost per token dominates—batch jobs, high-volume chat, or multi-step agent loops. Sol’s faster path is the lever when latency, not cost, is the constraint: interactive UIs, tool-calling chains, and anything where time-to-first-token or total wall time drives user experience. You can rebalance model selection without changing how you authenticate or route through the gateway.
In a market where multi-model gateways compete on routing, observability, and pass-through pricing, zero markup is the practical differentiator. When upstream GPT-5.6 pricing and speed move, customers on AI Gateway get those moves without a second pricing negotiation or a custom discount stack. That keeps Luna, Terra, and Sol comparable to direct upstream use on pure token cost, while still using a single integration for model access.
Watch for how your traffic splits across Luna, Terra, and Sol after the cut and the speed bump: cost-heavy routes should favor the cheaper Luna/Terra paths; latency-sensitive routes should sample Sol and measure end-to-end latency in your stack. Confirm in your AI Gateway usage that short- and long-context lines both reflect the new rates, and treat any remaining variance as application-side (prompt size, retries, parallel calls) rather than gateway markup.
Advertisement
🔎 More interesting news
- AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward…
- Agentic coding goes hands-free as OpenAI brings GPT-Live's full duplex voice control to…
- Gemini 2.5 Pro and Gemini 3 Flash deprecated
- Thinking Machines debuts Inkling Small open source AI model nearing performance of…
- Today's full Tech Pulse briefing →