Inkling Small from Thinking Machines is now available on AI Gateway
Inkling Small from Thinking Machines is now available on AI Gateway, according to the Vercel Blog. The model is positioned as a smaller sibling of the larger…
By Dillip Chowdary • Aug 05, 2026 • Source: Vercel Blog
Inkling Small from Thinking Machines is now available on AI Gateway, according to the Vercel Blog. The model is positioned as a smaller sibling of the larger Inkling model: it reaches performance comparable to that larger model at about a quarter of the size, and it uses much less compute per task.
On the product side, Inkling Small is described as a broad generalist with native reasoning over audio and images. It is also reported to hold up well on reasoning, agentic coding, and tool use. Controllable thinking effort is part of the design, so callers can trade quality against cost and latency rather than running a single fixed compute budget for every request.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For engineers and builders, the main practical shift is access through AI Gateway rather than a one-off integration path. A model that keeps near-Inkling performance at roughly a quarter of the size, with lower compute per task, is relevant for agent loops, coding tools, and multimodal pipelines where token and latency costs compound. Controllable thinking effort gives a direct knob for production tradeoffs: raise effort when quality matters, lower it when cost or latency dominates.
In market terms, this sits in the same lane as other gateway-distributed models that compete on size-to-capability ratio and multimodal + tool-use coverage. The combination of audio and image reasoning plus agentic coding and tool use targets workflows that used to require larger general models or multi-model stacks. Availability on AI Gateway also means teams already on that stack can evaluate Inkling Small without standing up a separate provider integration first.
What to watch next is how Controllable thinking effort behaves under real traffic: where the quality-cost-latency curve breaks for agentic coding and tool-use loads, and whether the quarter-size claim holds when audio and image reasoning are in the path. For early trials, compare Inkling Small against the larger Inkling model on the same gateway routes, with effort fixed high on hard tasks and dialed down on high-volume, lower-stakes ones.
Advertisement
🔎 More interesting news
- Show HN: OldHand A Claude/Codex plugin to verify the development flow end-to-end
- Show HN: Clayrune – Run Claude Code agents in parallel without losing context
- Agent skills that bring team coding standards to Claude Code and Codex
- AI coding agents are blowing through budgets — Replit, Kilo Code, and Symbotic explain…
- Today's full Tech Pulse briefing →