
Anthropic Cuts Eval Internet Access After Claude Exploits
The shutdown covers all internal evaluations, following four classes of unintended behavior that exposed…
Everything we have published, newest first.

The reported preparations raise a harder question: how much do public safety commitments reveal about efforts to prevent a severe incident?

Two new betas add live data and animation tools, while core artifacts reach Free users and standalone Design faces a December shutdown.

The MIT-licensed utility adds visual guidance for Claude Code and Codex while leaving clicks, credentials and approvals to people.

Illumina released code, weights and billions of variant scores, but its reported gains come with licensing restrictions and a tissue-dependent prediction weakness.

Pavel Rabtsevich reports an agent-assisted analysis, while TESS independently lists follow-up observations for the target, not confirmation of a planet.

The company links Russian and Iranian campaigns to planted articles and fabricated evidence, but its attribution and impact assessments require careful qualification.

Researchers question whether the checked Lean artifacts validate the published argument, without claiming that OpenAI’s natural-language proof is wrong.

The initiative supports critical-infrastructure defenders and offers opt-in vulnerability scans, but maintainers receive model-generated findings without human review.

Effective November 12, the revised Claude policy adds hardware safeguards, clarifies weapons and surveillance bans, and narrows restrictions on election targeting.

The October 8 workflow update separates corrections to a running task from instructions saved for later, with desktop settings and CLI shortcuts.

Baseten customers can use internal activation probes to flag risky behavior, but Goodfire’s Kimi K3 results do not establish performance across every…

StepFun’s 1M-context model gains another API access point, with explicit token prices and capabilities aimed at long-running coding and research agents.

Claude Science helped combine telescope surveys, but the project’s caveats separate a useful visualization from new observations or independently validated discoveries.

The new Lakebase guide connects isolated agent worktrees to pull request previews, while keeping migrations, data protection and cleanup central to deployment.

The downloadable FreeInference dataset exposes cache reuse, tool delays and context changes across real agent sessions, without releasing prompts or responses.
Selected by the editors. Worth your time.

The October 8 guide explains four setup steps, shared and user-specific identities, and the enterprise requirements behind agent-built Databricks Apps.

The proposed voice-to-agent ring combines meeting capture and health tracking, but Natura leaves pricing, delivery and key hardware claims unsettled.

Investor efforts to compare OpenAI With Anthropic reportedly produced the higher estimate, while different treatment of cloud-partner sales complicates the comparison.

Three dismissed researchers say unclear rules could deter safety collaboration, while OpenAI says an investigation found misconduct and denies retaliation.

Microsoft separates generally available CRM skills from public-preview sales and Teams features, while restricting its more autonomous Service Agent to selected customers.

The five-year commitment supports the Genesis Mission, but NVIDIA has not disclosed project allocations, a spending schedule or a public application process.

The three-year Genesis Mission commitment pairs Claude, coding tools, and API credits with training and technical support across more than 15 agencies.

The preliminary benchmark separates three agent failure signals, giving builders a more specific comparison tool than capability rankings alone.

The implementation guidance gives builders a migration framework for model selection, cache economics, prompt design and long-running agent work.

One persistent agent spans business apps, code, and automation, while Google separates model choice from the agent and puts industry specializations in…
The latest stories matching your interests.
Why you should use open models for everyday AI work, and when a paid API still makes sense.
Here are some projects to give you plenty of ideas to try with Claude Opus 5.5.
If you think Astra is too expensive, you now have cheaper options.
People are using Jev to review code, control computers, play games, and organize research. Here are my favorites.