Azure developers can now deploy Grok 4.7 as a generally available model in Microsoft Foundry and call it through OpenAI-compatible APIs. Microsoft’s Grok deployment documentation lists grok-4.7 as a Foundry Model sold by Azure, with support for Chat Completions, Responses, Microsoft Entra ID authentication, API keys, and function calling.
Grok 4.7 itself launched earlier. SpaceXAI announced the model on September 21, 2026, positioning it for coding and knowledge work. Microsoft’s documentation, updated October 6, now establishes generally available Foundry deployment.
Azure customers can access the model through Azure-managed endpoints, use Microsoft’s identity system for authentication, and bill usage through Azure. How much existing application infrastructure can carry over depends on the API integration and deployment details.
OpenAI Clients Can Call Grok Through Foundry
Microsoft documents both OpenAI-compatible API routes: Chat Completions and Responses. Its examples use the official OpenAI Python and JavaScript clients, configured to send requests to Foundry instead of OpenAI’s service.
At the resource level, Microsoft’s documented base URL follows this pattern:
https://<resource-name>.services.ai.azure.com/openai/v1
Chat Completions requests go to /openai/v1/chat/completions; Responses requests go to /openai/v1/responses.
The model parameter takes the deployment name chosen by the customer, which need not match the underlying model identifier grok-4.7. An application that sends the catalog name without matching its actual deployment could therefore target the wrong identifier.
Teams already using an OpenAI client can retain familiar SDK methods. Microsoft’s documented resource-level authentication pattern can be expressed as follows:
import os
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
from openai import OpenAI
token_provider = get_bearer_token_provider(
DefaultAzureCredential(),
"https://ai.azure.com/.default",
)
client = OpenAI(
base_url=f"{os.environ['AZURE_ENDPOINT'].rstrip('/')}/openai/v1",
api_key=token_provider,
)
response = client.responses.create(
model=os.environ["DEPLOYMENT_NAME"],
input="Summarize the main risks in this deployment plan.",
reasoning={"effort": "high"},
)
print(response.output_text)
This follows Microsoft’s documented integration pattern; it is not a report of hands-on testing. It requires the openai and azure-identity packages, the appropriate Azure permissions, and environment variables pointing to the resource endpoint and deployed model.
Familiar request methods reduce integration work, but API compatibility does not establish complete behavioral equivalence with an OpenAI model. Prompts, tool-selection behavior, response quality, and optional parameters still need application-level evaluation.
Entra ID Separates Model Access From API-Key Management
Microsoft supports both Microsoft Entra ID and API-key authentication for Grok 4.7, recommending Entra ID in its deployment guide.
The Entra ID approach uses a bearer-token provider with the https://ai.azure.com/.default scope. In the SDK example, that provider is passed through the client’s api_key argument, even though authentication uses Azure identity tokens instead of a static model-service key.
Azure teams can use identity-based access controls without distributing another long-lived credential. SpaceXAI’s Foundry guide describes support for managed identities, service principals, and local-development authentication through DefaultAzureCredential.
Deployment and inference have separate permission requirements. Microsoft lists the Cognitive Services Contributor role on the Foundry resource as a deployment prerequisite. SpaceXAI’s integration guide identifies Cognitive Services OpenAI User, or an appropriate equivalent role, for identities that call models.
An API key remains an alternative: Microsoft shows passing the Foundry resource key directly to the OpenAI client. That may suit an initial integration, but teams should choose their production authentication method deliberately instead of carrying a development key forward unchanged.
The 500,000-Token Window Supports Text and Image Input
Microsoft documents a context window of up to 500,000 tokens, alongside text and image input, advanced reasoning, and tool calling. The Foundry model catalog also marks Grok 4.7 generally available and identifies text as its output type.
The large window creates room for substantial source material in one request, such as document collections or code excerpts. Capacity alone does not demonstrate that the model will reliably find every relevant detail. Accuracy on a particular long-context task needs its own evaluation.
Text and image input allows an application to submit visual material for analysis. The listing does not announce image generation through this deployment.
Microsoft identifies coding, data extraction, summarization, and agentic applications as intended uses. Those are reasonable evaluation targets, though a particular enterprise workflow still needs testing against its accuracy requirements.
Function calling is supported through both Chat Completions and Responses. An application can describe a function, provide its argument schema, and allow the model to select that function and generate structured arguments.
SpaceXAI’s guide explicitly describes the application executing tool calls and continuing the conversation with the results. Developers remain responsible for validating arguments, enforcing permissions, handling failures, and deciding which operations require approval.
Agents with write access need particular care. Function calling provides an integration mechanism, but a requested database change, ticket update, or code modification is not necessarily safe to execute automatically.
Reasoning Controls Differ Between the API Examples
Grok 4.7 exposes configurable reasoning effort, with different controls documented for the two API surfaces.
Microsoft’s Chat Completions example uses the top-level reasoning_effort field and lists low, medium, high, and xhigh, with high as the default.
Its Responses example uses a nested reasoning object, such as {"effort": "high"}, and lists low, medium, and high, again with high as the default.
Teams should verify supported values before copying settings between endpoints, especially when carrying a Chat Completions configuration using into a Responses integration.
Sources
- Microsoft’s Grok deployment documentationlearn.microsoft.com
- announced the modelx.ai
- SpaceXAI’s Foundry guidedocs.x.ai
- Foundry model catalogai.azure.com





