Perplexity has cut PPLX Decider’s API input price to $0.02 per million tokens, down from $0.04, alongside an updated open-weight checkpoint. The company’s October 6 announcement also claims the new model leads the Decision Index 0.3 benchmark.
The release, pplx-decider-v1.1-27b, keeps the Qwen3.8-27B backbone. Perplexity attributes its reported performance gains mainly to removing the causal attention mask and training on more data, rather than moving to a different underlying model.
For developers building LLM-powered applications, the attraction is specific: a decision model for classification, routing, scoring, and multimodal choices, without using generated prose as the intermediate result. The practical distinction is between buying a decision through the API and operating the checkpoint yourself. Local deployment requires the supplied noncausal inference implementation and roughly 49 GiB for weights, plus additional working memory.
The Backbone Stays, but Attention Changes
According to Perplexity’s v1.1 model card, the updated checkpoint retains Qwen3.8-27B and gains mainly from two changes: lifting the causal mask and expanding the training data.
The attention change is relevant to what this model does. In a conventional causal language model, an input position cannot attend to positions that come after it. That restriction fits next-token generation, where the model must predict a continuation without seeing future text.
A decision task has a different requirement. The application already has the input and wants a classification, score, or choice. Removing the causal restriction allows attention to use information from later positions in that supplied input. For such tasks, access to the complete input can be more appropriate than preserving the directional constraint used for text generation.
That explains the architectural rationale, but it does not establish how much of the benchmark improvement comes from attention alone. More training data changed at the same time. The overall score cannot separate those contributions or show which change matters most for a particular application.
Keeping the backbone also narrows what developers should infer from this model release. Perplexity is reporting a better decision checkpoint, not demonstrating that the underlying Qwen model has become better at every generative task. The relevant question is whether v1.1 makes more accurate or useful decisions on the inputs an application actually receives.





