How to Install / Upgrade: Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock
Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock
By Dillip Chowdary β’ Aug 03, 2026 β’ Source: AWS Machine Learning Blog
Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock
OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock. This release also adds explicit prompt caching, which gives you precise control over which parts of your prompt are cached and reused. Together these changes let you run the new GPT-5.6 models on Bedrock and cut inference cost by caching the stable portions of your prompts.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
To get started, enable access to GPT-5.6 Sol, Terra, and Luna in Amazon Bedrock and point new or existing inference calls at the model you need. Set up explicit prompt caching so you mark which parts of each prompt should be cached and reused across requests. If you already run GPT workloads, migrate them to these models on Bedrock and apply explicit caching on the reusable prompt segments to reduce inference cost.
Watch for incomplete migration: existing GPT traffic will not pick up the new models or caching until you retarget it and define which prompt parts to cache. After you switch, verify that requests use Sol, Terra, or Luna on Bedrock as intended, that explicit caching is applied only to the segments you chose, and that inference cost moves in the direction you expect from reuse of cached prompt content.
Advertisement
π More interesting news
- When Cloud AI Escapes: OpenAI and Anthropic Models Breach Live Networks
- Anthropic's Claude breached 3 orgs, uploaded PyPI malware during tests
- Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI
- Boris Cherny on Trying to Get Claude Code to Rewrite the Claude App
- Today's full Tech Pulse briefing β