Inkling Small from Thinking Machines is now available on AI Gateway
Inkling Small from Thinking Machines is now available on AI Gateway, according to the Vercel Blog. The release puts a smaller variant of the Inkling line…
By Dillip Chowdary • Aug 04, 2026 • Source: Vercel Blog
Inkling Small from Thinking Machines is now available on AI Gateway, according to the Vercel Blog. The release puts a smaller variant of the Inkling line behind the same gateway path teams already use for model routing and access. The headline claim is direct: Inkling Small reaches performance comparable to the larger Inkling model at about a quarter of the size, while using much less compute per task.
On product mechanics, Inkling Small is positioned as a broad generalist rather than a narrow specialist. It supports native reasoning over audio and images, not text alone. Controllable thinking effort is built in, so callers can trade quality against cost and latency instead of treating every request as a fixed-depth run. The model is also described as holding up well on reasoning, agentic coding, and tool use—the workloads that usually expose weak small models first.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers and builders, the size-to-performance ratio is the practical lever. A model that tracks the larger Inkling at roughly one-quarter the size, with lower compute per task, changes how you budget multi-step agents, tool loops, and multimodal pipelines. Controllable thinking effort maps cleanly onto production patterns where some steps need deep reasoning and others should stay cheap and fast. Native audio and image reasoning also reduces the need to bolt on separate vision or speech models for many agent flows.
In market terms, this sits in the small-model generalist lane: capable enough for reasoning, coding agents, and tools, but priced and sized for higher volume. Shipping through AI Gateway matters because availability on a shared routing layer lowers the friction of trying a new vendor model beside existing ones. The competitive pressure is less about a single benchmark win and more about whether smaller models can stay close to their larger siblings without collapsing on agentic and multimodal work.
What to watch next is how far controllable thinking effort stretches in real agent stacks—whether teams can keep quality high on hard steps while cutting cost and latency on routine ones. Also watch whether the “comparable to larger Inkling at about a quarter the size” claim holds under sustained tool-use and coding workloads once traffic moves beyond demos. If it does, Inkling Small becomes a default candidate for production agents that need multimodality without paying large-model rates on every call.
Advertisement