Baseten customers can now use Goodfire’s AI-agent monitors to check signals inside a running model and send suspicious activity for closer inspection, according to TechCrunch’s October 8 report. The approach avoids requiring another large model to examine every step.
Goodfire’s pitch is economic: run inexpensive detectors continuously, then spend more on review when they flag a potential problem. For teams building AI products around long-running agents, that could make monitoring substantially cheaper than repeatedly feeding an expanding transcript to a second model.
The company’s reported Kimi K3 tests support that possibility. They do not yet establish how well the approach works across other models, real customer workloads, or every risk category offered through the integration.
Small Probes Decide What Needs a Closer Look
The architecture has two stages. Small classifiers called probes inspect the model’s internal activations, the intermediate numerical representations it produces while calculating its response. A flagged step can then receive more expensive scrutiny from another AI model.
This is the “inside-out” distinction: a transcript monitor evaluates text or actions exposed by the agent, while an activation monitor examines signals within the computation that produced them.
Goodfire says the probes reuse calculations the model already performs during its forward pass. They still require additional computation, but they do not need to generate a separate large-model response for every monitored step. That is the basis of the claimed cost advantage.
According to TechCrunch, Baseten customers can choose risk categories including offensive hacking, chemical and biological weapons misuse, and reward hacking. They can also configure responses such as logging, human review, or refusal. Finding a suspicious signal does not, by itself, determine what happens next.
That separation between detection and intervention gives inference providers room to set different policies. Some applications may need a record for later investigation; others may require review before allowing a sensitive action. Those choices have different operational consequences even when they use the same detector.
Deployment also requires access to internal activations. A developer working only with an ordinary text API cannot perform this inspection unless the model host exposes or processes the necessary signals. The Baseten integration puts the monitoring where the model runs.
The Cost Estimates Are Not a Customer Price List
TechCrunch reports that, in Goodfire’s Kimi K3 tests, monitoring approximately one million exchanges would cost roughly:
- $185 using the activation probes.
- $5,420 using a cheaper AI model to check every step.
- $200,000 using a top-tier AI model for that checking.
These vendor-reported estimates cover monitoring. They should not be interpreted as the total cost of operating an agent or as a published Baseten customer tariff. The report does not establish what a particular deployment will pay after adding inference, escalated reviews, human handling, and other infrastructure.
The architectural comparison is nevertheless meaningful. Each check by a second-model monitor requires another model invocation and the processing of relevant context. A probe extracts a signal from computation already happening inside the monitored model.
Adding second-stage review changes the cost. The bill depends on how often the probes escalate, how much context the reviewer receives, and which reviewer the operator chooses. Frequent alerts bring back more of the expensive checking.
Goodfire also says running four probes together added less than 2% to the time before the model began responding. This specific latency claim does not establish that every measure of runtime performance changes by less than 2%, or show the overhead for an entire agent session or a step that triggers additional review.
Hosts therefore need to compare the cost and latency of the complete monitoring pipeline at an acceptable detection rate, beyond the price of the first-stage classifier.
A Flag Is Not a Confirmed Safety Violation
In the Kimi K3 results reported by TechCrunch, Goodfire says its probes caught 93% of malicious hacking sessions and sent 5.5% of harmless sessions for a second look.
A harmless session escalated for review is not necessarily incorrectly blocked. A second-stage model or human reviewer might decide that the activity is legitimate and allow it to continue. Escalation still consumes reviewer capacity and may introduce delay, depending on how the operator handles alerts. In a deployment dominated by benign traffic, a relatively small escalation percentage can produce a substantial review queue.
The detection figure also leaves missed malicious sessions. In the basic two-stage architecture, behavior that the probes fail to flag will not automatically receive the additional review triggered by those probes. Other controls may catch it, but the reported result does not demonstrate that.
Neither figure should be confused with precision, which measures how many flagged cases are genuinely problematic. Precision depends partly on how common problematic behavior is in the evaluated workload. A test containing many malicious sessions can produce a different operational picture from a production service where those sessions are rare.
The published summary does not supply enough detail to reconstruct the test, including the session mix and uncertainty around the estimates. The numbers, attributed to Goodfire, offer evidence for a promising monitoring design. They are not an independently verified production guarantee.
The Research Supports a Signal, Not Mind Reading
Goodfire’s reward-hacking research, published September 17, provides context for why internal probes might detect behavior that transcript inspection misses.
Reward hacking occurs when an agent satisfies a scoring mechanism without accomplishing the intended task. An agent might exploit a grader or copy an available solution instead of solving the problem it was assigned.
Using differences between activations from simple synthetic code examples, Goodfire’s researchers identified an internal activation direction associated with concepts such as cheating, gaming a metric, and avoiding detection.
They then investigated what strongly activated the signal and what happened when they amplified it. According to the research, strongly activating text examples concerned cheating across different domains. Amplifying the signal affected model behavior in experiments involving exam stories and planted shortcuts.

Those experiments support the interpretation that the signal tracks something relevant to reward hacking, without making the probe a complete explanation of an agent’s intentions.
Goodfire itself qualifies the philosophical claim that models “know” when they are reward hacking. Its practical claim is narrower: an internal representation associated with that behavior can be measured and used for detection.
A classifier can be useful without proving conscious awareness. Even a meaningful internal signal can produce errors or fail under conditions outside its evaluation.
Baseten Availability Is a Concrete Step, With Boundaries
The release follows the Base Labs safety partnership involving Baseten, Hugging Face, and Goodfire, reported on September 17. That announcement described plans for safety evaluation and monitoring infrastructure for open-weight models. The October launch gives Baseten customers access to a concrete monitoring approach.
The validation remains limited. Launch results center on Kimi K3, and offering monitors for several risk categories does not establish that each category achieves the reported hacking-detection rate.
Model hosts considering the product should separate three questions.
Does the detector work on their workload? Authorized security testing, unusual tool use, and genuine misuse need to be distinguished in the environment where the monitor will operate. Goodfire’s benchmark is a starting point, not a substitute for that evaluation.
Does escalation preserve the economic advantage? Operators need to measure the full system, including benign alerts, reviewer calls, human review, and delays. The inexpensive first stage is valuable only if the overall pipeline meets the service’s requirements.
Does an alert arrive before the relevant action becomes irreversible? Prevention depends on how continuous monitoring connects to tool execution and other permissions. Logging an event after an external action cannot undo that action. The reported startup-latency result does not answer this intervention question.
Goodfire’s release suggests that continuous monitoring need not require an equally expensive model reading everything an agent does. Baseten customers now have access to an alternative built around internal signals and selective review.
Whether those savings survive deployment while maintaining acceptable detection and intervention performance remains the next test. Goodfire’s Kimi K3 results justify evaluating the architecture; they do not justify treating it as a universal safety layer for hosted open models.
Sources
- TechCrunch’s October 8 reporttechcrunch.com
- reward-hacking researchgoodfire.com
- Base Labs safety partnershiptechcrunch.com





