Home / Blog / MiniMax H3 now available on AI Gateway
Tech News

MiniMax H3 now available on AI Gateway

**MiniMax H3** is now available on **AI Gateway**, according to the Vercel Blog. The model produces **2K video** from a text prompt, a starting image, a…

By Dillip Chowdary • Aug 05, 2026 • Source: Vercel Blog

MiniMax H3 now available on AI Gateway

**MiniMax H3** is now available on **AI Gateway**, according to the Vercel Blog. The model produces **2K video** from a text prompt, a starting image, a first-and-last frame pair, or reference material, so teams can call one gateway path instead of wiring a separate MiniMax integration for each input style.

H3 covers **text-to-video** and **first-frame image-to-video**, plus **first-to-last keyframe** transitions that interpolate between two fixed endpoints. It also supports **multimodal reference-to-video**, conditioning a generation on reference images, video, or audio in a single request. That packs prompt, frame, and reference controls into one request surface rather than a chain of single-modality calls.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

For builders, gateway access matters because product flows rarely stay in one mode. A marketing clip may start from a still, a product demo may lock open and close frames, and a brand pass may need image or audio references without leaving the same API path. Putting H3 on AI Gateway keeps those modes behind one auth, routing, and usage model instead of ad hoc SDKs per experiment.

In the market, this is Vercel expanding AI Gateway beyond chat and still-image models into **2K video** with multi-input conditioning. Competing video stacks often split text-to-video, image-to-video, and reference workflows across providers or endpoints; H3 on the gateway compresses that surface for teams already standardized on Vercel’s AI layer.

Practical next step: try the four input paths on a real asset—plain text, single start frame, first/last keyframes, and multimodal references—and measure latency, cost, and how tightly reference or keyframe modes hold identity and motion. Watch whether gateway docs and quotas treat reference and keyframe modes as first-class peers to text-to-video, and whether audio-conditioned runs stay production-stable under concurrent load.

Advertisement

🔎 More interesting news

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →