🤖 AI HOT Daily · Aug 13, 2026
xAI releases Grok 4.6 with stronger long-running agent capabilities; Alibaba opens Qwen3.8-2.4T-A95B weights — the first open Qwen-Max-class model; Microsoft debuts its first self-built reasoning model MAI-Thinking-1; LTX-2.5 generates 10s of 720P video in 6.8s; DeepSeek V4 Pro and Grok 4.6 launch the same day, both approaching Claude Fable 5; OpenRouter launches a live web-search benchmark; Claude in Chrome's sidebar becomes Claude Cowork sessions.
1. xAI Releases Grok 4.6
Building on Grok 4.5, it strengthens long-running agents and complex interactive/visual work, reaching frontier levels on agentic coding and knowledge benchmarks, tying GPT-5.6 Sol overall.
2. Alibaba Opens Qwen3.8-2.4T-A95B Weights
The first open Qwen-Max-class model: 2.4T total parameters, 95B active per token, native 262K-token context extendable to 1M tokens.
3. Microsoft Debuts MAI-Thinking-1 Reasoning Model
Microsoft's first from-scratch reasoning model, MAI-Thinking-1, is now live on Microsoft Foundry.
4. LTX-2.5: 10s of 720P Video in 6.8 Seconds
Natively integrated with ComfyUI, it renders 10s of 720P video in 6.8s on two GB200s; the Fast tier costs ~$0.90 per 10s audio-720p clip.
5. DeepSeek V4 Pro Official Release
Agent capabilities jump while pricing holds; docs hint at a coming DeepSeek Harness minimal framework that may be the key to agent power.
6. OpenRouter Launches Live Web-Search Benchmark
It benchmarks model, search engine, method, and budget combos: raising the search budget to 25 rounds nearly doubles BrowseComp scores, with model choice mattering more than engine.
7. Claude in Chrome Sidebar Becomes Cowork Sessions
Sidebar conversations now save to history, skills and connectors work in-browser, and tasks switch seamlessly across desktop, web, and mobile.
8. WhatsApp Builds On-Device Scam Alert
An optional feature running on-device ML under E2EE to flag scam messages — content never leaves the device, weights are public for verification.
9. '100% Human-Written' Research Gold Is Entirely AI
Medical-research site Research Gold claims human-only writing with PhD reviewers — but investigators found the reviewers are AI-generated, and phone 'Sarah' is also an AI.
10. Google: Recall Is the Factuality Bottleneck
A knowledge-portrait framework finds frontier LLMs' factual encoding nears saturation while recall lags — most errors are 'lost keys', not 'empty shelves' — backed by the WikiProfile benchmark.
11. How AutoGPT Manages AI-Generated PRs
Using enforced PR templates, test plans, CI coverage gates, and CLA signing, AutoGPT turns agent-submitted PRs from 'unusable' into 'usable but off-roadmap'.
Sources and verification
This legacy briefing did not retain its original source URLs. Verify safety, financial, policy, and breaking-news claims with authoritative primary sources before acting on them.