Moderation teams can download PolicyLM-1.7B, define their own categories in plain language, and score messages against those rules without retraining the model. Rather than generating a written verdict, it returns a score for each category in one pass.
Musubi announced PolicyLM-1.7B on October 6, 2026, releasing the weights under Apache 2.0. The company reports a median inference time of 35 milliseconds for short chat messages with six categories on a 24GB NVIDIA L4 GPU. It also says the model runs on laptop CPUs and Apple silicon, although that GPU result should not be read as a laptop performance claim.
The practical appeal is the combination: locally deployable weights, platform-specific rules, and category-level thresholds. The limits are equally important. PolicyLM-1.7B handles text only, shares a 2,048-token context between policy and message, and provides no explanation for its decisions. Its performance figures are vendor-reported, not independently reproduced.
Download the Weights and Run Locally
The Hugging Face repository contains the weights, model card, and inference helper. Musubi describes the package as self-contained: it requires neither trust_remote_code nor a second model download and can operate offline.
For teams that want control over infrastructure and where message content is processed, that is a meaningful difference from a hosted moderation API. Apache 2.0 permits use under the license’s terms, but downloadable weights do not eliminate compute, integration, or operational costs.
The documented quickstart uses a fresh virtual environment and pins Transformers to version 4.57.6. Musubi says Transformers 5.x is not yet supported by the helper. The supplied instructions download revision v1.2:
pip install "transformers==4.57.6" "torch>=2.5.1" \
"safetensors>=0.4" "huggingface_hub>=0.34,<1.0"
hf download musubilabs/policylm-1.7b \
--revision v1.2 --local-dir policylm-1.7b
python policylm-1.7b/inference/quickstart.py policylm-1.7b
Pinning the revision and dependencies also gives a team a reproducible starting point for evaluation. “Runs locally” is the first deployment milestone, not evidence that the model is ready to enforce a community’s rules.
Write Short Categories, Not a Whole Policy Document
PolicyLM-1.7B’s custom-policy mode accepts categories defined at inference time. Each category can include a violation rule, a non-violation rule, and an exception override. Teams can add or revise those instructions without a training cycle.
For a game community, Musubi’s example distinguishes insults or threats aimed at another player from criticism of gameplay. A separate category flags off-platform trading while excluding scam warnings and questions about trading rules. Those distinctions show the intended use: translating a platform’s enforcement boundaries into compact instructions.
The model supports up to 16 custom categories and 1,662 policy tokens. Policy and message together must fit within the 2,048-token context. The helper’s model.check_policy(policy) reports token counts.
That makes this a policy-authoring task as much as a model-integration task. Teams need concise, decisive rules rather than a pasted handbook containing background, examples, appeals procedures, and internal commentary.
There are two qualifications to the promise of customization:
- Categories interact. Musubi warns that unrelated categories read together can affect one another’s scores and recommends separate calls where appropriate. That trades some one-pass efficiency for cleaner separation.
- Policy interpretation has limits. The model retains meanings learned during training. Instructions can refine those meanings, but Musubi says they are not intended to invert a concept such as abuse into support.
Changes take effect without retraining, but a wording change does not guarantee the expected decision change. Teams should regression-test policy edits against labeled examples before deploying them.
Calibrate Scores Before Turning Them Into Actions
PolicyLM-1.7B returns scores between 0 and 1. Enforcement depends on the cutoff a team chooses, not simply on whether the model produces a nonzero score.
Musubi provides two cutoff presets:
| Preset | Custom-policy cutoff | Built-in Aegis cutoff | Intended starting point |
|---|---|---|---|
| Precision, the default | 0.335 | 0.69 | Traffic where violations are rare, such as live chat |
| Balanced | 0.275 | 0.45 | Queues with more violations or a higher cost of misses |
The built-in Aegis mode uses NVIDIA’s 23-category taxonomy and a separate “Prompt harmful” score for its overall decision. Its thresholds should therefore not be confused with the custom-category thresholds.
Lowering a cutoff generally catches more violations while risking more false flags. The helper also supports category-specific cutoffs, so a team need not use the same sensitivity for credential sharing and routine profanity.
A sensible rollout would separate four tasks:
- Label representative examples from the actual community, including permitted content that resembles violations.
- Measure false positives and missed violations separately for each category and relevant language.
- Choose thresholds for the intended action, distinguishing automatic blocking from human-review triage.
- Recheck results after policy edits or changes to the serving setup.
These are deployment recommendations, not results from a Zeniteq test. Musubi itself says scores can shift with policy wording, language, device, and numeric format.
The distinction between blocking and triage matters. A threshold that is useful for sending suspicious messages to moderators may be unacceptable for automatically removing benign messages. The preset names do not establish a particular precision or recall level on a team’s own traffic.
The 35-Millisecond Result Has Specific Conditions
Musubi’s headline latency measurement is a 35-millisecond median per short chat message, using up to six categories on one 24GB NVIDIA L4 in bfloat16. The model card reports a 22-millisecond median on an NVIDIA H100 PCIe under the corresponding short-message setup.
These figures support investigating PolicyLM-1.7B for real-time screening. They do not establish end-to-end latency under every workload. A production request also encounters preprocessing, scheduling, queueing, and the application’s enforcement logic. Longer text and different category configurations require their own measurements.
Throughput is another separate question. Musubi reports 39.3 messages per second on the L4 in its batched built-in-taxonomy test. In a different test with randomly arriving chat messages and one five-category policy, it reports sustaining 34.4 messages per second while keeping p95 latency at or below 150 milliseconds.
Those tests have different configurations and should not be combined into a universal capacity estimate.
The underlying approach helps explain the speed. As TechCrunch’s reporting describes, decision models restrict their outputs rather than writing an answer token by token. PolicyLM-1.7B produces category scores, avoiding the generated verdict that an LLM-based moderation workflow would otherwise need to parse.
Musubi also reports that a custom fine-tuned PolicyLM version is deployed on a platform handling more than one million messages daily. The platform is unnamed, and the claim concerns a customized version, not demonstrated production performance of the downloadable release.
The Benchmark Supports a Tradeoff, Not a Universal Win
Sources
- Musubi announced PolicyLM-1.7Bmusubilabs.ai
- Hugging Face repositoryhuggingface.co
- TechCrunch’s reportingtechcrunch.com





