AWS’s new access route gives developers two cross-Region model IDs, four reasoning settings, and Bedrock-specific pricing to evaluate.
Listen
AI narration
14:01
0:00 / 14:01
AI SummaryGenerated from this article
xAI's Grok 4.7 is now available through Amazon Bedrock, giving AWS-native developers access to the model via AWS credentials, APIs, and billing. This is a distribution change announced September 28, not a new model release; xAI introduced Grok 4.7 on September 21.
Developers must use cross-Region inference profiles—either us.xai.grok-4.7 or global.xai.grok-4.7—rather than the base model ID. Grok 4.7 supports a 500K-token context window, image input, and four reasoning-effort settings that always remain active. AWS's Standard pricing runs $2.00–$2.20 per million input tokens and $6.00–$6.60 per million output tokens depending on the chosen profile.
Bedrock lets developers use the model through AWS accounts, credentials, APIs, and billing. Its availability does not establish new model capabilities or independently validate xAI’s performance claims.
The integration details are worth checking early. Grok 4.7 uses cross-Region inference profiles rather than in-Region model calls, reasoning is always active, and pricing depends on both the routing profile and service tier. A request sent to the wrong model ID will fail before any of the model’s advertised capabilities come into play.
Use a Cross-Region Model ID, Not the Base ID
Bedrock lists Grok 4.7 with a 500,000-token context window, text and image input, and text output. Its model card gives developers two inference IDs to use on the bedrock-runtime endpoint:
Inference option
Model ID
Where processing can be routed
US geographic cross-Region
us.xai.grok-4.7
Supported Regions within the US geography
Global cross-Region
global.xai.grok-4.7
Supported commercial AWS Regions globally
The base identifier, xai.grok-4.7, appears in AWS’s model information, but the model card says in-Region inference is not supported on bedrock-runtime. Requests must name one of the two cross-Region inference IDs.
The endpoint’s Region still matters. For OpenAI-compatible calls, AWS shows the base URL as https://bedrock-runtime.{region}.amazonaws.com/openai/v1. The Region in that URL is where the client sends its request; the selected inference profile determines the permitted routing beyond it. Before deploying, check that the model and chosen profile are available from the AWS Region your application will use.
Bedrock Supports Both OpenAI-Compatible and AWS APIs
AWS lists four supported invocation options: the Responses API, Chat Completions API, , and Converse. Responses and Chat Completions are available through Bedrock’s OpenAI-compatible path. Converse offers a familiar AWS SDK interface for teams already calling multiple Bedrock models.
Work with Zeniteq
Let’s work together
We’re open to thoughtful collaborations with teams building in AI. Explore the ways we can work together.
A minimal Responses request can use the OpenAI Python SDK, pointed at Bedrock rather than OpenAI:
from openai import OpenAI
client = OpenAI(
api_key="<Bedrock API key or short-term bearer token>",
base_url="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1",
)
response = client.responses.create(
model="us.xai.grok-4.7",
reasoning={"effort": "low"},
input="Explain the difference between Geo and Global inference.",
)
print(response.output_text)
The SDK choice does not change the provider receiving this request: the URL sends it to Amazon Bedrock. AWS says the bearer credential can be a Bedrock API key or a short-term token derived from AWS Identity and Access Management (IAM) credentials. An existing AWS SDK application can instead use Converse with ordinary AWS credentials and the same cross-Region ID in its modelId field.
Check permissions before troubleshooting application code. AWS says bedrock:InvokeModel is evaluated against the account’s default project, the named inference profile, and the underlying foundation model. Bearer-token authentication additionally requires bedrock:CallWithBearerToken. A policy allowing the US profile does not automatically allow the Global profile. AWS recommends treating a long-term Bedrock API key as an exploration credential and using short-term, IAM-derived credentials for production.
The APIs also differ beyond request syntax. AWS lists application inference profiles for InvokeModel and Converse, but not for Responses or Chat Completions. Teams relying on that Bedrock feature should not assume an OpenAI-compatible client is a drop-in replacement for their Converse workflow.
A 500K Window Does Not Make Every Request Cheap
Grok 4.7’s 500K context window gives applications room for long documents, conversation history, or agent state. It is a capacity limit, not a recommendation to fill every prompt. Longer requests consume more input tokens, and sustained reasoning can substantially increase output-token use.
AWS documents four reasoning-effort settings: low, medium, high, and xhigh. Reasoning is always active, with high as the default. On the Responses API, developers set the reasoning.effort field, as in the example above. On Converse, AWS uses additionalModelRequestFields={"reasoning_effort": "low"}.
Test settings against the work the application actually does. A short classification or extraction task may not need xhigh; a difficult coding or multi-step analysis task might justify it. Compare answer quality, latency, and tokens consumed, along with whether the request completes.
AWS’s announcement cites an Artificial Analysis evaluation in which Grok 4.7 used roughly 81,000 output tokens per Intelligence Index task, versus roughly 38,000 for Grok 4.6. Grok 4.7 was measured at xhigh, while the older model’s reported effort varied by measure. Those figures warn of possible task costs, but they are not a like-for-like estimate of what a default Bedrock call will spend.
Response handling differs across APIs, too. AWS says the Responses API can return encrypted reasoning content when requested, so an application can pass it into later turns. Chat Completions does not return reasoning tokens. For Converse, AWS’s example notes that reasoning may occupy a content block before the answer; code should find the text block rather than assume the first block contains the final response.
Routing Determines Both Residency and Price
The US geographic profile keeps processing within the US geography, according to AWS. The Global profile can route to any supported commercial AWS Region, giving Bedrock more capacity to draw on at a lower published Standard rate. Despite that price advantage, Global is a poor default for a workload that requires US-only processing.
Neither option is in-Region inference. If a policy requires processing in one specific AWS Region, the US profile’s geographic boundary is not the same guarantee. Selecting US routing also does not settle every legal or contractual residency requirement; teams need to check the complete data-handling requirements for their workload.
Routing may affect performance as well as compliance. AWS notes that Global’s broader placement can bring more variable latency. Teams with latency-sensitive traffic should test the profile they intend to deploy instead of inferring response time from the model name alone.
Bedrock lists other operational features for Grok 4.7, including response streaming, invocation logging, guardrails, structured outputs, and implicit prompt caching. Caching can reduce charges for eligible repeated prompt prefixes; it does not make a newly submitted 500K-token document free. Logging and guardrails are useful integration options, but teams should configure them deliberately, particularly when requests contain sensitive material.
Check the Bedrock Rate, Not Just xAI’s Launch Price
AWS’s Grok 4.7 pricing table lists these Standard-tier on-demand rates, in US dollars per million tokens:
Profile
Input
Output
Cache read
US geographic
$2.20
$6.60
$0.55
Global
$2.00
$6.00
$0.50
At those published rates, one million uncached input tokens and one million output tokens would cost $8.80 through the US profile or $8.00 through Global. That comparison excludes any eligible cache reads and says nothing about how many tokens a real task will use. The $2 input and $6 output starting prices in xAI’s September launch announcement should not be applied indiscriminately to Bedrock traffic.
Service tier is a second pricing choice. AWS lists Standard as pay-per-token with no commitment, Priority at 1.75 times the Standard rate, and Flex at half the Standard rate. Requests can select priority or flex through service_tier; omitting the field selects the default Standard tier. Priority buys prioritized processing, while Flex is intended for work that can tolerate less predictable timing.
For budgeting, compare cost per completed task, not just price per million tokens. Run the same prompts across the intended profile, service tier, and reasoning effort. Check AWS’s current rates and regional availability before deployment, since a pricing table and a model ID alone cannot predict an agent’s total spend.
Final Thoughts
Bedrock gives AWS-native applications a way to evaluate Grok 4.7 alongside models they already use, within existing authentication, billing, and invocation workflows. September 28 marks that new distribution route, not a new xAI model release.
Start by confirming that the chosen profile meets routing requirements and IAM permits the complete request path. Then measure whether the selected reasoning effort delivers enough value for its token cost. Bedrock makes that comparison possible inside an AWS workflow; it does not decide the result for the team.
Frequently Asked Questions
4 questions
1
What Model ID Should I Use for Grok 4.7 on Bedrock?
Use us.xai.grok-4.7 for US geographic cross-Region inference or global.xai.grok-4.7 for Global cross-Region inference. AWS says Grok 4.7 does not support in-Region inference on bedrock-runtime, so name one of those profiles rather than the base xai.grok-4.7 identifier. Check availability from the AWS Region where your client sends requests.
2
Does Grok 4.7 on Bedrock Support Images and Reasoning Controls?
Yes. AWS lists text and image input, text output, a 500K-token context window, and low, medium, high, and xhigh reasoning effort. Reasoning is always active and defaults to high. The Responses API accepts a reasoning.effort setting; Converse takes reasoning_effort through additionalModelRequestFields.
3
How Much Does Grok 4.7 Cost on Amazon Bedrock?
At AWS’s published Standard rates, Global inference costs $2 per million input tokens and $6 per million output tokens; US geographic inference costs $2.20 and $6.60, respectively. Cache reads have separate lower rates. AWS lists Priority at 1.75 times Standard and Flex at half Standard, so check current pricing and measure token use for your workload.
4
Was Grok 4.7 Released on September 28?
No. xAI announced Grok 4.7 on September 21, 2026; September 28 was AWS’s announcement that the existing model had become available through Amazon Bedrock. The later date marks an additional access, API, and billing route. Bedrock availability alone does not demonstrate a change to the model or establish new benchmark gains.