Microsoft Killed Its Claude Code Licenses Because Token Billing Broke the Budget
From Microsoft's internal pilot cancellation to Uber's blown AI budget, token-based pricing is forcing a reckoning across the industry.
Explore AI news, practical guides, tutorials, reviews, and insights from Zeniteq.
From Microsoft's internal pilot cancellation to Uber's blown AI budget, token-based pricing is forcing a reckoning across the industry.
Claude can now spin up parallel agents for massive coding tasks, but this much power comes with a serious token bill.
Anthropic has doubled Claude Design's weekly token allowances for Pro, Max, Team, and Enterprise subscribers, giving designers and developers significantly more runway…
Alibaba’s 2.4-trillion-parameter MoE adds stronger agent coordination, multimodal handling, and cache rates as low as $0.17 per million tokens.
The free coding LLM has dominated OpenRouter’s rankings, while a cow-movie meme and GLM fingerprints fuel the hunt for its maker.
Gemini’s new processing mode selectively inspects video timelines, cutting costs while improving answers on long and detail-heavy footage.
The temporary API promotion reduces output token costs by one-third while leaving ChatGPT subscriptions and included usage limits unchanged.
Google releases Multi-Token Prediction drafters for Gemma 4, tripling output speed without any degradation in reasoning quality or accuracy.
The company’s guidance focuses on prompt caching, instruction cleanup, and effort tuning to lower API bills without sacrificing task quality.

Early inference tests show major efficiency and latency gains, but production scale and realistic agent workloads remain important tests.
The 320B-A18B open-weight model pairs native multimodality with a one-million-token context window and unusually low API prices.
SpaceXAI's new LLM makes its biggest gains in agentic coding and knowledge work, but xhigh reasoning consumes many more tokens.
The new model tops coding and research benchmarks, ships a Pro tier, and doubles the API price to match rising costs.
Meta says its latest agent model sustains longer workflows while using fewer tools and tokens, with more cautious user collaboration.
Fable 5’s weak token share, Opus 5’s rapid rise, and GPT-5.6 Sol’s price cut show that businesses are optimizing for value, not…
Fable 5.1 may be Anthropic's best model yet, but ridiculous quotas, token use, and safeguards make it frustrating to use.
The anonymous 1-million-token AI model shows strong agentic potential, though benchmark caveats and unresolved ownership questions matter more than the hype.
MiniMax released M3 on June 1, 2026, stacking frontier coding performance, a million-token context window, and native multimodality into a single open-weights…
Hackers are now selling Vercel’s internal database for $2M. Here’s what happened and how to protect yourself.
OpenClaw's OpenAI Codex OAuth integration lets ChatGPT Plus and Pro subscribers run GPT-5.4 agents locally without paying a single extra API token…
Qwen3.8-Flash-Next and GLM-5.3-Flash bring recent Claude-class performance to downloadable models, but “local” still means server-grade hardware.
One month after a limited preview, OpenAI's frontier models and Codex coding agent are now fully available on AWS, giving enterprises a…