LLM Guide Turns Qwen Into a Single-Pass Classifier
Nish Tahir’s tutorial pairs constrained answer-token scoring with CommonsenseQA evaluation, fine-tuning and temperature scaling, backed by a public implementation repository.
Listen
AI narration
11:00
0:00 / 11:00
AI SummaryGenerated from this article
Nish Tahir's tutorial demonstrates using Qwen3-1.7B as a single-pass classifier by restricting outputs to five valid answer tokens and scoring them in one forward pass. Evaluated on CommonsenseQA, the approach achieves a 3.03 percentage point accuracy improvement after fine-tuning. The guide's key finding reveals a calibration gap: predictions in the highest-confidence bin averaged 98.55% probability but were correct only 70.09% of the time. Temperature scaling addresses this overconfidence, and Tahir's public repository provides scripts for dataset preparation, fine-tuning, and calibration, offering developers an inspectable workflow for building constrained-output decision systems.
In Nish Tahir’s Qwen-based classification demonstration, predictions in the highest-confidence bin averaged 98.55% probability but were correct only 70.09% of the time. That gap is the most useful finding in his new decision-model guide: restricting an LLM to valid answers is straightforward; making its probabilities trustworthy requires additional work.
The October 10, 2026 tutorial uses Qwen3-1.7B to score five possible answers from a single forward pass. It evaluates the classifier on CommonsenseQA, reports a modest improvement after fine-tuning, and applies temperature scaling to adjust its probability distribution.
As a third-party implementation guide, it offers a concrete workflow developers can inspect and adapt, including the evaluation steps needed to assess a constrained answer generator as a decision system. It is not a new Qwen or Jev model release.
One Forward Pass, Five Allowed Answers
Tahir’s implementation starts with Qwen3-1.7B, a tokenizer and a multiple-choice prompt. Each answer receives a label, A through E, and the question and answer text appear together in the input.
The prompt uses Qwen’s chat template with enable_thinking=False, so the implementation can score the answer label immediately without first generating a reasoning sequence.
Instead of calling generate and waiting for a textual response, the code runs the model directly. It takes the logits at the final input position, where the model predicts the next token, and keeps only those corresponding to the five answer labels.
A logit is an unnormalized score. Applying softmax to those five selected scores produces a probability distribution over the allowed answers. The highest-scoring label becomes the prediction.
The central operation looks like this adapted fragment, assuming the model, tokenizer and tokenized prompt are already prepared:
options = ["A", "B", "C", "D", "E"]
option_token_ids = []
for option in options:
ids = tokenizer.encode(option, add_special_tokens=False)
assert len(ids) == 1, "Each answer label must be one token"
option_token_ids.append(ids[0])
temperature = 1.0
model.eval()
with torch.no_grad():
logits = model(**model_inputs).logits[0, -1]
scores = logits[option_token_ids].float()
probabilities = torch.softmax(scores / temperature, dim=-1)
prediction = options[probabilities.argmax().item()]
Work with Zeniteq
Let’s work together
We’re open to thoughtful collaborations with teams building in AI. Explore the ways we can work together.
The single-token check is an implementation safeguard. Tahir’s example takes the first token ID returned for each label; developers adapting the method should verify that their labels really are single tokens for the tokenizer they use. Taking only the first token of a multi-token label would score something different from the intended answer.
Here, “single-pass” means the model processes the prompt and supplies all five answer scores without an autoregressive output-generation loop. Computation still depends on prompt length, and the guide does not establish a measured latency advantage on specified hardware.
For a fixed-choice decision, though, the efficiency opportunity is clear: the application need not generate a paragraph or JSON object just to recover one label.
A Valid Answer Is Not a Reliable Confidence Score
Restricting the output vocabulary solves a formatting problem. Correctness remains a separate problem.
The probabilities are normalized over the five selected answer tokens, not over the model’s entire vocabulary. They describe the relative preference among those allowed labels after every other token has been excluded. That distribution always sums to one, so even when none of the answers fits the input well, the classifier must allocate all its probability mass among them.
Tahir illustrates the problem with an ambiguous question: “Where would you most likely find a bat?” The choices include a cave, a baseball game, an attic, a zoo and a sporting goods store. His example assigns 99.78% probability to “Cave,” despite the ambiguity between the animal and sporting equipment.
This is an illustration, not a formal calibration test. The stronger evidence comes from grouping predictions by confidence and comparing each group’s mean probability with its observed accuracy.
For a well-calibrated classifier, predictions assigned roughly 80% confidence should be correct roughly 80% of the time across comparable examples. Calibration is a statistical property of a collection of predictions, not a guarantee about any individual answer.
A sharply peaked next-token distribution does not automatically meet that requirement. Training a model to predict text and restricting its possible outputs do not, by themselves, make the resulting scores empirical probabilities of correctness.
CommonsenseQA Shows a Modest Fine-Tuning Gain
Tahir reports testing the classifier on a random held-out sample of 1,221 CommonsenseQA questions. The tutorial’s evaluation results show:
Configuration
Correct answers
Accuracy
Macro F1
Before fine-tuning
725 / 1,221
59.38%
0.5844
After fine-tuning
762 / 1,221
62.41%
0.6234
Fine-tuning adds 37 correct answers, an improvement of approximately 3.03 percentage points. These are the author’s reported results, not independently reproduced measurements.
The class-level results warrant a closer look. Before fine-tuning, recall for answer positions A and B is much higher than for D and E. After fine-tuning, the recall values are more evenly distributed.
That pattern gives developers a reason to investigate sensitivity to answer position, although it does not establish the cause. A useful follow-up experiment would shuffle the same answer choices and check whether predictions remain stable when the correct answer moves to a different label.
The evaluation demonstrates a particular prompt, answer-label scheme and multiple-choice task. It does not support the broader conclusion that “Qwen becomes a general decision model.” Developers adapting the workflow to another classification problem still need representative evaluation data from that problem.
Accuracy and macro F1 leave the calibration question open, too. A classifier can improve its label predictions while continuing to attach excessive confidence to its mistakes.
Temperature Scaling Adjusts Probabilities, Not Answers
Of the 1,221 predictions in the tutorial’s raw confidence table, 809 fall in the 90–100% confidence bin. Their average confidence is 98.55%, compared with 70.09% accuracy.
The mismatch extends beyond that group: the 80–90% bin averages 85.55% confidence but achieves only 47.11% accuracy.
Tahir addresses this overconfidence with temperature scaling: divide the selected logits by a positive scalar, T, before applying softmax. At T = 1, the distribution is unchanged. A temperature above one flattens it, reducing the dominance of the highest-scoring answer; a temperature below one makes it sharper.
The guide reports a fitted temperature of approximately 3.7973. Its post-scaling table shows closer alignment in several bins. For example, the 50–60% bin has 54.75% mean confidence and 54.82% accuracy.
Positive scalar temperature scaling preserves the ranking of the logits. It therefore leaves the selected answer unchanged and cannot, on its own, improve top-choice accuracy. Its purpose is to adjust how much confidence the system attaches to its existing decisions.
The tables come with a reporting caveat: the pre- and post-scaling calibration tables imply different aggregate accuracies. Because temperature scaling alone cannot cause that change, they should not be read as a controlled before-and-after comparison of an unchanged classifier without clarification of the checkpoints and evaluation splits involved.
Nor is the reported temperature a universal Qwen setting. Developers should fit a temperature for their own trained model, prompt and data, using a calibration set separate from the final evaluation set. A value that works for these multiple-choice questions should not be assumed to transfer to an unrelated application.
The Repository Makes the Workflow Inspectable
The accompanying build-your-own-jev repository provides the scripts Tahir links for dataset preparation, evaluation, fine-tuning and calibration. The supplied repository listing also includes a pyproject.toml and uv.lock, giving developers dependency information alongside the implementation.
Readers can inspect how the stages connect and attempt to reproduce the demonstration before changing the model or task, making the guide more actionable than an isolated code snippet. The public code provides a reproducibility starting point; it does not establish that the reported results have been independently verified.
For an application-specific adaptation, the important sequence is:
Establish a baseline with the exact prompt, tokenizer and allowed labels.
Evaluate both prediction accuracy and confidence reliability.
Fine-tune on training data, then reassess the resulting checkpoint.
Fit temperature on separate calibration data.
Measure accuracy and calibration on an untouched final test set.
Keeping those stages separate prevents a fitted confidence adjustment from looking more reliable simply because it is assessed on the same examples used to choose it.
The implementation can also support a confidence threshold that routes uncertain cases for review. That would be an application policy built around the calibrated scores, not a capability automatically supplied by constrained decoding.
Tahir’s guide is most useful as an inspectable starting point for fixed-choice LLM applications. Extracting a decision takes little code; deciding whether its confidence score is reliable takes separate evaluation. Before those scores govern automated actions, developers need to reproduce the workflow, resolve the evaluation details and test calibration on the data their system will actually encounter.