AI Daily

🤖 AI HOT Daily · Aug 29, 2026

Tencent Hunyuan open-sources Hy4 preview (770B params, 1M context); Zhipu releases GLM-5.3 weights for agentic coding and network defense; a federal judge rules the Trump administration's blacklisting of Anthropic illegal; Anthropic's Claude trains models to mitigate 10 alignment failures, beating 28 human safety researchers; Claude Code adds model-switch hooks; Claude for Teachers goes free-Enterprise for schools; OpenAI and Thailand launch an 8-week accelerator; Stanford ships Terminal-Bench-Science for research agents; Apple probes LLM Bayesian consistency and Agent Seer scenario synthesis; free-Groq AI-engineer Colab notebooks; Gary Marcus draws 5 lessons from OpenAI's Hugging Face breach.

  1. 1. Tencent Hunyuan Hy4 Preview: 770B Params, 1M Context

    The next flagship hits 770B total / 49B active parameters with a 1M context window, open-sourced on Tencent TokenHub and OpenRouter.

  2. 2. GLM-5.3 Weights Released for Agents and Cyber Defense

    Zhipu's most powerful agentic-coding and network-defense model is now downloadable, runnable, and customizable.

  3. 3. Judge Rules Anthropic Blacklisting Illegal

    A Northern California court found the ban retaliatory and unconstitutional after Anthropic refused to drop its guardrails on lethal autonomous war and mass surveillance.

  4. 4. Claude Code v2.1.251 Adds Model-Switch Hooks

    Hooks now intercept and confirm model switches; remote controllers stream foreground subagent tool calls live; /usage shows a spend-cap bar.

  5. 5. Claude for Teachers Goes Free Enterprise for Schools

    Schools and districts get a free Enterprise tier with learning-science-based teaching skills linked to academic standards across all 50 states.

  6. 6. OpenAI and Thailand Launch 8-Week AI Accelerator

    Ten health- and education-focused startups each receive $2,000 in API credits, one-on-one mentorship, and frontier-model access.

  7. 7. Claude Auto-Trains Models to Fix Alignment Failures

    Ten failure classes shrink toward perfect behavior without hurting general capability, scaling to models 4.7x larger; its deception fix beats 28 human researchers' best by 20%.

  8. 8. Terminal-Bench-Science 0.1 Benchmarks Research Agents

    Stanford leads this benchmark of 70 expert-vetted tasks spanning life, physical, earth, math, and engineering sciences for research agents.

  9. 9. LLMs Aren't (Always) Bayesian: Probing Belief Consistency

    Apple treats LLMs as information processors and measures how evidence-based belief updates systematically deviate from Bayesian ideals across medicine, science, and law.

  10. 10. Agent Seer Synthesizes Evals from Tool Specs

    By mining function names, descriptions, and typed parameters, it auto-generates realistic multi-step tool-combination scenarios without human authors or live execution.

  11. 11. AI-Engineer Notebooks: Framework-Free on Free Colab/Groq

    Raw APIs instead of frameworks cover prompts, RAG, evals, agents, fine-tuning, and serving — all on the free Groq tier with no credit card.

  12. 12. Five Lessons from OpenAI's Hugging Face Breach

    Marcus revisits July's breach and rival agent missteps, citing METR's 90-page report: real security challenges demand governance that keeps pace with capability.

Sources and verification

This legacy briefing did not retain its original source URLs. Verify safety, financial, policy, and breaking-news claims with authoritative primary sources before acting on them.