Home / Blog / Inkling Small from Thinking Machines is now available on AI…
Tech News

Inkling Small from Thinking Machines is now available on AI Gateway

Inkling Small from Thinking Machines is now available on AI Gateway, according to the Vercel Blog. The model is positioned as a smaller sibling of the larger…

By Dillip Chowdary • Aug 05, 2026 • Source: Vercel Blog

Inkling Small from Thinking Machines is now available on AI Gateway

Inkling Small from Thinking Machines is now available on AI Gateway, according to the Vercel Blog. The model is positioned as a smaller sibling of the larger Inkling model: it reaches performance comparable to that larger model at about a quarter of the size, and it uses much less compute per task.

On the product side, Inkling Small is described as a broad generalist with native reasoning over audio and images. It is also reported to hold up well on reasoning, agentic coding, and tool use. Controllable thinking effort is part of the design, so callers can trade quality against cost and latency rather than running a single fixed compute budget for every request.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For engineers and builders, the main practical shift is access through AI Gateway rather than a one-off integration path. A model that keeps near-Inkling performance at roughly a quarter of the size, with lower compute per task, is relevant for agent loops, coding tools, and multimodal pipelines where token and latency costs compound. Controllable thinking effort gives a direct knob for production tradeoffs: raise effort when quality matters, lower it when cost or latency dominates.

In market terms, this sits in the same lane as other gateway-distributed models that compete on size-to-capability ratio and multimodal + tool-use coverage. The combination of audio and image reasoning plus agentic coding and tool use targets workflows that used to require larger general models or multi-model stacks. Availability on AI Gateway also means teams already on that stack can evaluate Inkling Small without standing up a separate provider integration first.

What to watch next is how Controllable thinking effort behaves under real traffic: where the quality-cost-latency curve breaks for agentic coding and tool-use loads, and whether the quarter-size claim holds when audio and image reasoning are in the path. For early trials, compare Inkling Small against the larger Inkling model on the same gateway routes, with effort fixed high on hard tasks and dialed down on high-volume, lower-stakes ones.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →