Perplexity has cut PPLX Decider’s API input price to $0.02 per million tokens, down from $0.04, alongside an updated open-weight checkpoint. The company’s October 6 announcement also claims the new model leads the Decision Index 0.3 benchmark.
The release, pplx-decider-v1.1-27b, keeps the Qwen3.8-27B backbone. Perplexity attributes its reported performance gains mainly to removing the causal attention mask and training on more data, rather than moving to a different underlying model.
For developers building LLM-powered applications, the attraction is specific: a decision model for classification, routing, scoring, and multimodal choices, without using generated prose as the intermediate result. The practical distinction is between buying a decision through the API and operating the checkpoint yourself. Local deployment requires the supplied noncausal inference implementation and roughly 49 GiB for weights, plus additional working memory.
The Backbone Stays, but Attention Changes
According to Perplexity’s v1.1 model card, the updated checkpoint retains Qwen3.8-27B and gains mainly from two changes: lifting the causal mask and expanding the training data.
The attention change is relevant to what this model does. In a conventional causal language model, an input position cannot attend to positions that come after it. That restriction fits next-token generation, where the model must predict a continuation without seeing future text.
A decision task has a different requirement. The application already has the input and wants a classification, score, or choice. Removing the causal restriction allows attention to use information from later positions in that supplied input. For such tasks, access to the complete input can be more appropriate than preserving the directional constraint used for text generation.
That explains the architectural rationale, but it does not establish how much of the benchmark improvement comes from attention alone. More training data changed at the same time. The overall score cannot separate those contributions or show which change matters most for a particular application.
Keeping the backbone also narrows what developers should infer from this model release. Perplexity is reporting a better decision checkpoint, not demonstrating that the underlying Qwen model has become better at every generative task. The relevant question is whether v1.1 makes more accurate or useful decisions on the inputs an application actually receives.
The Reported Benchmark Gain Is 5.16 Points
Perplexity’s model card reports the following Decision Index results:
| Model | Reported Decision Index Score |
|---|---|
| PPLX Decider v1.1 | 61.56 |
| PPLX Decider v1 | 56.4 |
| Jev | 57.9 |
Those figures put v1.1 5.16 points above v1 and 3.66 points above Jev. Perplexity’s announcement points readers to the Jev Decision Index on Hugging Face and says the updated checkpoint scores highest on version 0.3 of the benchmark.
These are Perplexity’s published results, not an independent reproduction presented here. The distinction matters because a benchmark lead is narrower than a claim of general superiority. The scores support the company’s case for evaluating v1.1, but they do not establish that it will outperform alternatives on every classification or multimodal workload.
Nor should 61.56 be rewritten as “61.56% accuracy.” A named index score needs its own metric definition and evaluation context before it can be interpreted that way. Similarly, the 5.16-point increase is not automatically a 5.16% reduction in production errors.
For developers, the most useful follow-up is a comparison on representative application data. A routing system should be evaluated on whether it selects the appropriate destination. A classifier needs examples from its actual categories, including ambiguous inputs. A scoring system needs a test of whether its scores support the decisions made downstream.
The aggregate benchmark result is a reason to run those tests, not a replacement for them.
Cheaper Decisions Are Useful When Prose Isn’t the Product
At the announced rate, 100 million input tokens would cost $2, compared with $4 at the previous $0.04-per-million rate. That is straightforward arithmetic for the stated input charge, not an estimate of an application’s complete operating cost.
The appeal of a decision model is clearest when the output the application needs is a choice rather than an explanation. Possible uses include assigning incoming requests to predefined categories, selecting a processing route, scoring material against a criterion, or making a decision using multimodal input.
For example, a support workflow might need to choose between billing, technical support, and account-access queues. Asking a general-purpose LLM to write an explanation and then extracting the queue name adds a generated-text step that the application may not need. A decision-focused model is a candidate for handling that narrower task directly.
These are application examples, not reported tests of PPLX Decider v1.1. Developers still need to check whether the model supports their decision structure and whether its behavior is reliable enough for the consequences attached to the result.
The release does not make ordinary text-generating AI models obsolete. An application that needs a customer-facing explanation, a written response, or a draft document still needs generation somewhere in its workflow. Decider’s potential role is the decision stage before that work, or a standalone decision where no prose is required.
The pricing change lowers the cost of evaluating that approach. It does not establish lower latency, better throughput, or fewer failures. Those operational properties need measurement separately from both the input rate and the Decision Index score.
Local Use Requires More Than Downloading the Weights
The checkpoint uses an Apache 2.0 license, giving developers an open-weight deployment option alongside Perplexity’s API. But its inference requirements are a material part of the release.
The published local-use instructions call for:
- Python 3.12 or newer.
- Authenticated Hugging Face access.
- The supplied inference implementation, which supports the noncausal attention path.
- GPU capacity for approximately 49 GiB of weights, with additional working memory.
The custom inference requirement follows directly from the attention change. A standard causal generation path should not be assumed to reproduce the model’s intended behavior. Removing the mask is part of how v1.1 operates, so a deployment must preserve that behavior rather than treating the checkpoint as an interchangeable text-generation model.
Sources
- October 6 announcementcommunity.perplexity.ai
- v1.1 model cardhuggingface.co
- Jev Decision Indexhuggingface.co





