Memory capacity is becoming a central AI inference constraint
Recent infrastructure coverage and systems research point to HBM capacity, memory bandwidth, model weights, and KV cache as increasingly important limits on inference cost and latency.
The latest shifts in the AI market, sourced and analyzed by our automated monitoring network and research staff.
Recent infrastructure coverage and systems research point to HBM capacity, memory bandwidth, model weights, and KV cache as increasingly important limits on inference cost and latency.
In a new MIT Technology Review essay, Bill Gates argues that AI capabilities in areas such as biology are moving faster than the guardrails meant to contain their risks.
Thomson Reuters has introduced Thomson, a legal-focused language model trained around its professional data assets and positioned as a more specialized alternative to general-purpose models.
Mistral AI and Saudi company HUMAIN are reported to be collaborating on computing infrastructure, advanced models, and Arabic-language AI applications in the region.
Apple says the new Mac Studio delivers a major jump in AI performance and is positioning its high-memory desktop as a serious platform for local AI workloads.
OpenAI has offered its clearest look yet at Jalapeño, a custom inference accelerator developed with Broadcom as the company expands beyond buying general-purpose AI hardware.
Portable Computer runs the agent harness, planner, tool router, and supported models locally on NVIDIA DGX Spark or compatible RTX/Linux systems, with cloud escalation when needed.
The Accel-backed startup has emerged from stealth with an AI-focused web index and retrieval infrastructure designed for agents that need fresh, machine-readable evidence.
Google Cloud is packaging legal skills, permission-aware MCP connectors, specialized agents, and centralized governance for law firms and corporate legal departments.
IBM has released Granite 4.2 models in 3B, 8B, and 30B sizes, with switchable thinking modes aimed at reasoning and agent-style workloads.
WSJ: Anthropic more than doubled revenue to $11.6B in Q2 with a small adjusted operating profit, nearly twice OpenAI's $6.7B (up 18% QoQ, operating loss widened to $12.3B).
After an Astra test-agent hacked Hugging Face, OpenAI paused RL for two weeks, left its largest frontier run on hold, rewrote its Preparedness Framework, and priced AI monitoring at roughly 20% of inference compute for Sol-class and Astra workloads.
Payments giant Stripe agreed to buy AI model gateway OpenRouter for over $7 billion, months after the startup raised at a $1.3B valuation. OpenRouter routes 8M+ developers to 400+ models and takes ~5% of inference spend.
DeepSeek V4 enters general release with 1M context window and peak-hour double pricing
GPT-5.6 family goes global; Terra matches 5.5 performance at half the cost
Elon Musk's xAI merges into SpaceX, simultaneously releases Grok 4.5
GPT-Live enables interruption handling and simultaneous interpretation, approaching human-like conversation
Mistral unveils robotics foundation model and announces industrial AI strategy at AI Now Summit
NVIDIA launches nvDock large language model, API-only access
DeepSeek has assembled a chip team and contacted foundries, aiming to reduce reliance on third-party hardware
Anthropic's agent work platform Cowork officially opens to Max subscribers
Muse Image is Alexander Wang's first major release since taking over Meta AI
Anthropic's enterprise strategy pays off as run-rate exceeds OpenAI by approximately $6B
Sonnet 5's $2/$10 introductory pricing is 60% below Opus 4.8; global API price war continues
Sam Altman proposes granting Washington a 5% stake in OpenAI to align regulatory and commercial interests
Claude Code tops benchmarks with Fable 5; SpaceX acquires Cursor for $60B
Commerce Department lifts export controls; Fable 5 returns July 1
Gemini 3.5 Flash launches with 4x speed improvement; 3.5 Pro pushed to late July
Google AI subscriptions restructured into three tiers; Ultra introduces Gemini Spark autonomous agent
Meta Compute will sell surplus AI compute to external customers; stock breaks $600
Microsoft 365 commercial pricing increases take effect July 1, with Frontline seeing the highest at 33%
Claude Sonnet 5 becomes default for free and Pro users, performance approaches Opus 4.8
xAI raised input from $1.25 to $2.00 (+60%) and output from $2.50 to $6.00 (+140%) per 1M tokens effective June 22.
Input dropped from $2.10 to $1.74 and output from $4.40 to $3.48 per 1M tokens, with eight new models added including MiniMax M3 and Kimi K2.7 Code.
Planned June 15 move of programmatic Claude usage to a separate API-rate credit pool was cancelled last-minute; subscriptions continue covering Agent SDK, `claude -p`, and third-party apps unchanged.
Price cut from $7.99 with storage doubled from 200GB to 400GB; Pro and Ultra tiers unchanged.
Premium Request Units replaced by token-based AI Credits on June 1; Pro stays $10/month but now includes only $10 in credits, with premium model usage metered at API rates.
5-minute minimum replaces the flat 20-minute session rate, lowering effective cost for short-lived tasks while per-minute rate stays the same.
Speculative decoding framework released June 27 boosts per-user generation on V4-Flash by 60-85% and V4-Pro by 57-78% at matched throughput; open-source MIT license on HuggingFace.
AI agent for Office apps moved from included-in-$30-Copilot-license to pay-as-you-go after June 18 GA, marking Microsoft's first major pricing restructure in nearly two decades.
1,000 free requests per day ended June 18; free-tier, AI Pro, and AI Ultra subscribers lose terminal access; only Enterprise Code Assist customers retain it.
Three-tier family at $5/$30 (Sol), $2.50/$15 (Terra), $1/$6 (Luna) per 1M tokens; limited to ~20 government-vetted organizations via API and Codex.
Moonshot AI's Kimi App introduces new pricing tiers: Standard ¥68/mo, Enhanced ¥200/mo, and Professional ¥500/mo.
GPT-5.5 listed at $5/$30 per million tokens; GPT-5.4 at $2.5/$15.
Doubao 1.6 API prices drop significantly. Input from ¥0.8/M tokens, output ¥8/M tokens — a 63% cost reduction, with the first 700K tokens free.
New Claude 4 models available now with updated API pricing.
Claude Code and other programmatic tools now billed separately from chat usage.
V4 family includes Flash, standard, and Pro variants with new pricing.
deepseek-chat and deepseek-reasoner aliases end July 24.
Student verification unlocks a full year of Cursor Pro at no cost.
Students and educators get half-price Pro access with verification.
Alibaba Cloud announces Qwen3.5-Plus pricing on Bailian: input ¥0.8/M tokens, output ¥4.8/M tokens, with a 7,000-token free tier for new users.
ByteDance reports Doubao token usage has surpassed 120 trillion, making it one of the world's largest enterprise AI platforms by token consumption.
Free tier loses access to Pro-class models; Flash remains free.
Baidu's ERNIE 4.0 API drops to ¥0.008/1K input tokens and ¥0.024/1K output tokens, with free access for basic use cases.
Tencent Hunyuan reduces API pricing: reasoning at ¥0.0055/1K input and ¥0.0165/1K output, plus a free tier of 4M tokens.
OpenAI raised the GPT-4o message caps for Plus and Team subscribers, allowing more extensive daily use.
OpenAI launched GPT-4.5 as a research preview for Pro users and above, emphasizing improved emotional intelligence and reduced hallucinations.
Showing 58 of 58 updates