🤖 AI HOT Daily · Aug 29, 2026
Tencent Hunyuan open-sources Hy4 preview (770B params, 1M context); Zhipu releases GLM-5.3 weights for agentic coding and network defense; a federal judge rules the Trump administration's blacklisting of Anthropic illegal; Anthropic's Claude trains models to mitigate 10 alignment failures, beating 28 human safety researchers; Claude Code adds model-switch hooks; Claude for Teachers goes free-Enterprise for schools; OpenAI and Thailand launch an 8-week accelerator; Stanford ships Terminal-Bench-Science for research agents; Apple probes LLM Bayesian consistency and Agent Seer scenario synthesis; free-Groq AI-engineer Colab notebooks; Gary Marcus draws 5 lessons from OpenAI's Hugging Face breach.
1. Tencent Hunyuan Hy4 Preview: 770B Params, 1M Context
The next flagship hits 770B total / 49B active parameters with a 1M context window, open-sourced on Tencent TokenHub and OpenRouter.
2. GLM-5.3 Weights Released for Agents and Cyber Defense
Zhipu's most powerful agentic-coding and network-defense model is now downloadable, runnable, and customizable.
3. Judge Rules Anthropic Blacklisting Illegal
A Northern California court found the ban retaliatory and unconstitutional after Anthropic refused to drop its guardrails on lethal autonomous war and mass surveillance.
4. Claude Code v2.1.251 Adds Model-Switch Hooks
Hooks now intercept and confirm model switches; remote controllers stream foreground subagent tool calls live; /usage shows a spend-cap bar.
5. Claude for Teachers Goes Free Enterprise for Schools
Schools and districts get a free Enterprise tier with learning-science-based teaching skills linked to academic standards across all 50 states.
6. OpenAI and Thailand Launch 8-Week AI Accelerator
Ten health- and education-focused startups each receive $2,000 in API credits, one-on-one mentorship, and frontier-model access.
7. Claude Auto-Trains Models to Fix Alignment Failures
Ten failure classes shrink toward perfect behavior without hurting general capability, scaling to models 4.7x larger; its deception fix beats 28 human researchers' best by 20%.
8. Terminal-Bench-Science 0.1 Benchmarks Research Agents
Stanford leads this benchmark of 70 expert-vetted tasks spanning life, physical, earth, math, and engineering sciences for research agents.
9. LLMs Aren't (Always) Bayesian: Probing Belief Consistency
Apple treats LLMs as information processors and measures how evidence-based belief updates systematically deviate from Bayesian ideals across medicine, science, and law.
10. Agent Seer Synthesizes Evals from Tool Specs
By mining function names, descriptions, and typed parameters, it auto-generates realistic multi-step tool-combination scenarios without human authors or live execution.
11. AI-Engineer Notebooks: Framework-Free on Free Colab/Groq
Raw APIs instead of frameworks cover prompts, RAG, evals, agents, fine-tuning, and serving — all on the free Groq tier with no credit card.
12. Five Lessons from OpenAI's Hugging Face Breach
Marcus revisits July's breach and rival agent missteps, citing METR's 90-page report: real security challenges demand governance that keeps pace with capability.
Sources and verification
This legacy briefing did not retain its original source URLs. Verify safety, financial, policy, and breaking-news claims with authoritative primary sources before acting on them.