Laguna S 2.1 is now available on AI Gateway
By Dillip Chowdary • Jul 21, 2026 • Source: Vercel Blog
Poolside’s Laguna S 2.1 is now available on AI Gateway. Builders can call a free route at poolside/laguna-s-2.1-free or a paid route at poolside/laguna-s-2.1 when they need higher throughput and higher rate limits. The model ships as open-weight software with a context window of up to 1M tokens and supports both thinking and no-thinking run modes.
Architecturally, Laguna S 2.1 is a Mixture-of-Experts model: inference routes work across a set of experts rather than always activating a single dense stack. The 1M-token window is the hard capacity limit for long prompts, retrieval-heavy packs, and multi-file or multi-document sessions. Thinking mode keeps intermediate reasoning on the path; no-thinking mode drops that path for lower latency and simpler completion-style calls.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers already on AI Gateway, the split IDs matter more than the headline. Free traffic can land on poolside/laguna-s-2.1-free for prototypes, evals, and low-volume tools. Production paths that need stable rate limits and higher throughput move to poolside/laguna-s-2.1 without changing provider surface. Open weights keep local or private replication as an option if gateway dependency or data residency later becomes a constraint.
On the market side, this is Poolside placing an open-weight MoE behind a gateway product surface with a clear free-vs-paid metering line. Availability through AI Gateway reduces integration cost for teams that already route multiple models through one gateway layer. The paid SKU’s pitch is operational headroom—throughput and rate limits—not a different base architecture from the free route.
Practical next steps: wire both IDs into staging, measure latency and quality under thinking vs no-thinking for the same tasks, and load-test free limits before promoting long-context or high-QPS traffic to poolside/laguna-s-2.1. Watch for how thinking mode trades off cost and latency against no-thinking on your real 1M-window workloads, and whether free-tier rate limits force an early cutover to the paid route.
Advertisement
🔎 More interesting news
- Product Jun 30, 2026 Introducing Claude Sonnet 5 Sonnet 5 delivers frontier performance…
- Inviting hard questions Announcements Jul 9, 2026 We’re asking the public for their…
- Apple teams up with Klarna to launch a lease-to-own program for iPhones, iPads, and Macs
- Show HN: Tokenmaxx – CLI that merges usage across Claude Code and Codex accounts
- Today's full Tech Pulse briefing →