🤖 AI HOT Daily · Sep 23, 2026
Anthropic ships Claude Opus 5.5 — the first model in the Claude 5.5 family, matching Fable 5.1 on most tasks at ~40% lower typical-load cost than Opus 5 ($4/$20 per 1M in/out tokens, cache reads down 60% to $0.20, output ~30% faster, Fast mode up to 2.5x speed at double token pricing), with a 1M context — and Claude Code v2.1.280 now defaults to it; Artificial Analysis crowns it #1 on its Intelligence Index at 58 (an all-time high), ties GPT-6 Astra on Terminal-Bench 4.0 (59.6%), and it enters Arena's Agent Arena; Boris Cherny has Opus 5.5 and Fable 5.1 each port HAProxy from C to Rust — both pass nearly all tests, but Opus 5.5 finishes in 9.5 hours for less; OpenAI launches GPT-6 Sol and GPT-6 Luna, applying Astra's training method to faster, cheaper models with API prices down 50% from GPT-5.6 promo rates (Sol $4→$2 in, $20→$10 out; Luna $0.20→$0.10 in, $1.20→$0.50 out), rolling out to Plus/Pro/Business/Enterprise/Edu across ChatGPT Work and Codex — plus an improved GPT-6 prompt-cache system granting up to 90% discounts on qualifying shared prefixes reused within a 30-minute window; Alibaba open-sources Qwen-Image-2.1 and tops Arena's Image Edit chart among open models at 1367 (#16 overall, 3 points from GPT-Image-1.5-high-fidelity) while also leading open entries on Text-to-Image; Hugging Face makes transformers load GGUF quantized checkpoints directly via from_pretrained(gguf_file=...), reusing ggml's Metal kernels; OpenRouter launches a Batch API that completes async work within 24 hours at typically 50% or less of normal per-token pricing across 70+ models; Kimi ships a browser extension (ex-Kimi WebBridge) enabling sidebar chats that navigate pages and fill forms, plus 'record a skill once, let Kimi replay it'; LlamaIndex's LiteParse 2.14.6 trims text-extraction time 20-25% via a self-maintained PDFium fork with built-in mimalloc (averaging 2.76ms/page); the new Mac mini (M6/M5 Pro) and Mac Studio (M5 Max/Ultra) go on sale Sep 22 with up to 4x AI performance; Bloomberg, citing an unpublished Pentagon review, reports overreliance on Palantir's Maven Smart System contributed to a Feb day-one strike on an Iranian elementary school that killed 150+ (at least 123 children); Meta's ultra-privileged Muse assistant shows a serious 0-day letting any local app or terminal command steal a user's auth token for full agent control — Patrick Wardle built write-file and camera PoCs, Meta shipped a hotfix ~12 hours after disclosure, and Amazon had already begun banning it; Epoch AI estimates the cost of achieving a given AI performance falls ~47% per quarter (~13x a year) — 66% per quarter right after SOTA status; a hands-on by Kazke concludes Grok 4.7 underwhelms while Xiaomi's MiMo V2.6 becomes the current answer to the impossible triangle of performance-price-speed; Step 5 Preview scores 44 on the AA Intelligence Index (level with Kimi K3 max, just shy of GLM-5.3 and Qwen3.8 Max's 45) at ~1/2.8 peer cost; on Banking77's 3,080 support messages, OpenRouter finds Jev 1.13 at 81.0% accuracy (3.3 points behind Claude Opus 5) but with a 175ms median latency versus 2,266ms — roughly 1/13 the time and 1/25 the cost; and NVIDIA's Nemotron 3.5 Lightning (30B total/~3B active, open-weight MoE) targets high-volume, crisp-boundary agent execution steps like tool calling and coding, paired with the 550B Nemotron 3 Ultra for hard reasoning.
1. Claude Opus 5.5: Fable-level Performance, 40% Cheaper
First of the Claude 5.5 line; ~40% lower typical-load cost than Opus 5, $4/$20 per 1M tokens, $0.20 cache reads (-60%), 30%+ faster output, 1M context.
2. Opus 5.5 Tops AA Index at 58, Ties Astra on Terminal-Bench
Highest score ever measured (58); 59.6% on Terminal-Bench 4.0 ties GPT-6 Astra — and it's live in Arena's Agent Arena.
3. Cherny's Field Test: Opus 5.5 Beats Fable 5.1 on Porting
Both C-to-Rust ports pass nearly every test; Opus 5.5 does it in 9.5 hours and costs less.
4. GPT-6 Sol and Luna: 50% Cheaper, Rolling Out Now
Astra's training recipe applied to faster, cheaper models — Sol $2/$10, Luna $0.10/$0.50 per 1M — landing in ChatGPT Work and Codex now.
5. GPT-6 Prompt Caching: Up to 90% Off
Qualifying shared prefixes reused within a 30-minute window default to higher hit rates — up to 90% off cached input tokens.
6. Qwen-Image-2.1 Open-Source, Tops Arena Edit Charts
Ranks #1 open model on Image Edit Arena at 1367 (#16 overall, 3 points off GPT-Image-1.5-high-fidelity) and leads open entries on Text-to-Image too.
7. transformers Can Now Load GGUF Quants Directly
Pass gguf_file to from_pretrained and load Hub GGUF checkpoints straight away, reusing ggml's Metal kernels for near-llama.cpp local inference.
8. OpenRouter's Batch API: Half-Price Bulk Inference
Async batches complete within 24 hours at typically 50% or less of normal per-token pricing, spanning 70+ models.
9. Kimi's Browser Extension: Record Tasks as Skills
Ex-Kimi WebBridge; chat in the sidebar, navigate pages, fill forms, and record a repeating task once for Kimi to replay.
10. LiteParse 2.14.6 Speeds Up PDF Extraction 20-25%
A self-maintained PDFium fork with built-in mimalloc averages 2.76ms/page, with markdown rendering at 3.94ms/page.
11. New Mac mini and Mac Studio Hit Shelves Today
Mac mini moves to M6 and M5 Pro with up to 4x AI performance; Mac Studio offers M5 Max and M5 Ultra.
12. Pentagon Review: Palantir AI Overreliance Led to School Strike
An unpublished review ties the Feb day-one strike on Shajarah Tayyebeh Elementary to overreliance on Palantir's Maven Smart System — 150+ dead, at least 123 of them children.
13. Meta Muse's Serious 0-Day: Any Local App Can Steal the Keys
Any local app or terminal command can lift a Muse auth token for full agent control; Wardle demoed file-write and camera PoCs, and Meta hotfixed within ~12 hours.
14. Epoch AI: The Plunging Price of Thought
Cost per fixed performance falls ~47% a quarter (~13x a year); right at SOTA it touched 66% per quarter before slowing.
15. Hands-On: Grok 4.7 Disappoints, MiMo V2.6 Answers the Triangle
Against expectations, Grok 4.7 underwhelms while Xiaomi's MiMo V2.6 becomes the current resolution to the performance-price-speed triangle.
16. Step 5 Preview Scores 44 on AA at ~1/2.8 the Cost
Intelligence Index 44 (level with Kimi K3 max, just under GLM-5.3 and Qwen3.8 Max at 45) at roughly 1/2.8 of peer cost.
17. Jev vs Opus 5 on Banking77: 13x Faster, 25x Cheaper
81.0% accuracy (3.3 points behind) yet a 175ms median versus 2,266ms — about 1/25 of the cost.
18. Nemotron 3.5 Lightning: Purpose-Built for Agent Exec
30B-total/~3B-active open MoE for crisp, high-frequency steps like tool calling and coding, teaming with 550B Nemotron 3 Ultra on hard reasoning.
Sources and verification
This legacy briefing did not retain its original source URLs. Verify safety, financial, policy, and breaking-news claims with authoritative primary sources before acting on them.