A million uncached input tokens cost $1 through OpenRouter’s newly listed Step 5 Preview. Output costs $2.70 per million tokens, while eligible cached input costs $0.05 per million. Developers comparing agent backends now have another access point with explicit AI pricing.
OpenRouter lists October 8, 2026, as its release date for stepfun/step-5-preview. That date refers to the OpenRouter listing, not StepFun’s first announcement of the underlying model. Developers can now try it through OpenRouter’s OpenAI-compatible API at published prices.
The listing describes Step 5 Preview as StepFun’s flagship model for agentic work, using a sparse mixture-of-experts architecture with 600 billion total parameters and 27 billion active parameters. Its more immediately useful specifications are a 1M-token context window, up to 64K output tokens, text/image/video input, and support for tools and structured responses.
The Hacker News discussion adds a chronology caveat: commenter james2doyle said the model had been available on ZenMux for at least two weeks. That earlier availability remains a commenter’s report, not an independently verified launch timeline. Views on speed, price, and benchmark competitiveness were mixed; the thread offered no consensus that the model was an obvious replacement for existing options.
The Published Prices Make Agent Experiments Easier to Budget
OpenRouter publishes three relevant rates:
| Token category | Price per million tokens |
|---|---|
| Uncached input | $1.00 |
| Output | $2.70 |
| Cached input reads | $0.05 |
Workloads that repeatedly send large amounts of context can make these rates especially relevant. A coding agent might carry repository material, instructions, and previous tool results through several iterations. A research agent might repeatedly reference the same collection of documents while adding new evidence.
Consider an illustrative request with 200,000 input tokens and 10,000 billed output tokens. At the listed uncached rates, those quantities cost $0.20 for input and $0.027 for output, or $0.227 combined. If all 200,000 input tokens qualified for the cache-read rate, the input charge would fall to $0.01, bringing the same calculation to $0.037.
These figures are arithmetic from the price sheet. They do not describe a measured workload or promise that an entire prompt will receive cached pricing.
The cache-read rate is 95% below the uncached input rate. Whether an application captures that saving depends on which portions of its requests qualify for caching. Developers should inspect returned usage and billing; repeating a conversation does not guarantee a discounted request.
Low token rates also don't establish a low cost per completed task. An agent that needs additional attempts, produces excessive output, or spends many calls recovering from mistakes can consume the savings.
Call stepfun/step-5-preview Through OpenRouter
StepFun’s own API documentation uses the model identifier step-5-preview; OpenRouter uses the namespaced slug stepfun/step-5-preview.
For an application using the OpenAI Python SDK, a minimal text request looks like this:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
)
response = client.chat.completions.create(
model="stepfun/step-5-preview",
messages=[
{
"role": "user",
"content": (
"Review this proposed code change and identify "
"the tests needed to verify it."
),
}
],
max_tokens=2048,
)
print(response.choices[0].message.content)
This makes a bounded text-generation call. Repository access, search, and code execution require application-provided tools and an execution loop.
Developers already using OpenRouter can evaluate another model without adopting StepFun’s direct API as a separate backend. Existing orchestration still needs testing, particularly for tool calls, multimodal payloads, and provider-specific options.
The listing shows StepFun as the model’s sole provider. OpenRouter supplies the familiar access layer here, but developers should not assume that this listing provides multiple upstream providers or automatic failover between them.
A 1M Context Window Expands the Test, Not the Guarantee
StepFun’s model documentation specifies a 1M-token context window and up to 64K output tokens. It positions that capacity for cross-document analysis, software engineering, and multi-step research.
Coding teams can test whether supplying more relevant code, dependency information, and error logs helps the agent locate a problem or propose a correct change. In research, a larger context can support questions spanning several documents without requiring each source to be reduced to a short summary first.
Capacity alone doesn't demonstrate reliable recall. A model may accept a long request yet miss a constraint, confuse similar passages, or produce a conclusion that the supplied evidence doesn't support. Nothing in the listing establishes how consistently Step 5 Preview handles those cases.
An evaluation should test information spread across the context, including conflicting sources and details far from the final instruction. Confirming that a large request succeeds establishes acceptance, not understanding.
The output limit deserves similar discipline. A 64K ceiling allows lengthy deliverables, but applications should still request an appropriate response budget. A small extraction task doesn't benefit from permission to generate a report-sized answer, and long outputs require their own checks for completeness and accuracy.
The 600B-total/27B-active architecture provides technical context. The mixture-of-experts designation describes selective activation within a larger model; neither that designation nor parameter counts establish coding accuracy, long-context reliability, or agent success rates.
Multimodal Input and Tools Need an Application Around Them
StepFun documents native support for text, images, and video as input, with text as output. This supports a broader set of agent experiments than text-only document processing.
A development assistant could inspect a screenshot alongside an error report. A research workflow could analyze a chart together with its accompanying text. A video-analysis application could request a written summary or extract information into a specified format. These are possible workflows based on the documented modalities, not results from hands-on testing.
The model understands these inputs; it does not generate images or video. Developers also need to check the accepted payload format and media limits when routing requests through OpenRouter. A direct StepFun API example should not be assumed to transfer unchanged.
With tool calling, the model can request a function supplied by the application. The application executes it and returns the result for the next step.
StepFun explicitly states that search, code execution, and access to external services come from the integrating application. The model does not independently gain access to a developer’s local environment. Repository permissions, command execution, approval gates, and stopping conditions remain application responsibilities.
For structured responses, the documentation lists both JSON Mode and JSON Schema support. JSON Mode is useful when an application needs machine-readable JSON. A schema provides a more specific contract for the response’s structure.
Neither establishes that the content is correct. A valid object can still contain an unsupported finding or a mistaken file reference. Applications should validate the structure and separately check consequential values before acting on them.
Sources
- Step 5 Previewopenrouter.ai
- Hacker News discussionnews.ycombinator.com
- StepFun’s model documentationplatform.stepfun.ai





