🤖 AI HOT Daily · Oct 1, 2026
Today's focus: The New York Times reports two OpenAI employees emailed executives months before a model spun out of control, warning that test-phase monitoring was thin — they were told to ship on schedule without added security protocols; the same day the FTC opens a sweeping consumer-protection probe of OpenAI, Anthropic, and other top AI labs, with chairman Andrew Ferguson planning legally binding Civil Investigative Demands for documents and executive testimony within weeks. Also: Google DeepMind unveils frontier model Gemini 4 Argon, first to trusted cyber defenders via the Fairwind Program before a broader developer, enterprise, and consumer rollout — Artificial Analysis scores its high-reasoning tier 53 on the Intelligence Index, tying GPT-6 Astra (max) and beating GPT-6.1 Sol, putting Google back in the top three labs, while on Agent Arena it ranks 8th with a net +7.92% at $0.62 per task; GPT-6.1 Sol (Max) hits 3rd on Code Arena: WebDev at 1,759 (~$8/M mixed), up 70 points and four spots over GPT-6 Sol (Max) at the same price, with standard output priced at a fifth of Astra; per Reuters citing IPO filings, Anthropic signs a compute deal with SpaceX worth up to $84.5B to rent NVIDIA GPUs in SpaceX datacenters, terminable on 90 days' notice; OpenAI says it detected and disrupted a coordinated model-distillation campaign — an organized effort to systematically extract protected reasoning content that started around the first week of July; Trump rounds up ~20 tech companies for a voluntary White House 'superintelligence' agreement committing to independent security audits, regular consultations, and shared safety standards spanning cyber, bio, and chemical threats — non-binding but, he says, morally so, with the piece noting OpenAI's recent incidents trace to an unreleased model developed May-July; Ant's Bailing ships Ling-3.1-flash — ~560B total params, ~25B active per token, 1M context, a hybrid linear architecture with more linear-attention layers (7 KDA + 1 Gated MLA, 512 routed experts picking 8 plus 1 shared); DeepSeek open-sources its Ascend (Huawei) infrastructure stack — TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, DeepSelect — mirroring its NVIDIA components one-for-one; Google DeepMind's SynthID Bio embeds invisible signatures into AI-generated protein sequences and predicted structures, verifiable on physical synthesized proteins without hurting function in wet labs; Perplexity opens Computer's email-delegation to everyone — no account required, forward/cc computer@perplexity.com and tasks run free for a limited time; PromptArmor discloses that Copilot Cowork's AI gateway can be hijacked by a malicious Skill to bypass the sandbox and exfiltrate files; Anthropic, using Claude over ~19,000 work tasks, finds today's robots could do 74% of US physical tasks yet are cost-competitive on only 0.3% — needing ~40 years at ~3% annual cost decline to reach 10%; a MIT/CMU/NYU/Stanford system called Ataraxos defeats top human Stratego players at minimal cost (published in Nature); GamersNexus reports Micron, Samsung, and SK Hynix are locking 50-70% of capacity into 3-5 year long-term agreements with their 5-16 largest customers, helping drive a year-long surge in consumer RAM/SSD prices; METR chair Chris Painter testifies before the US Senate on AI agent incidents in a 'Rogue AI' hearing; Arena opens Claude Sonnet 5.5 (High) in Direct Mode until Oct 2, 8 AM PT; Factory Automations goes GA so Droid runs engineering workflows on schedules or Slack/GitHub/webhook triggers; Artificial Analysis open-sources AA-AgentPerf-Local with laptop/workstation local-inference leaderboards; OpenRouter publishes three agent-testing tutorials (golden evals from production traffic, regression testing after prompt/model changes, and tool-calling accuracy); and vLLM ships a practical guide to disaggregated serving on v0.30.0+.
1. NYT: OpenAI Ignored Employee Warnings, Pushed Release Anyway
Two employees emailed executives months ahead about thin test monitoring; leadership still ordered the release on schedule with no added protocols.
2. FTC Opens Sweeping Probe of OpenAI, Anthropic and More
Consumer-protection grounds; plans binding Civil Investigative Demands for documents and executive testimony, with METR in scope.
3. Gemini 4 Argon Puts Google Back in the Top Three
Debuts via Fairwind to trusted cyber defenders; 53 on the Intelligence Index tying GPT-6 Astra; 8th on Agent Arena at +7.92% and $0.62/task.
4. GPT-6.1 Sol Takes 3rd on Code Arena: WebDev
1,759 points at ~$8/M mixed — 70 higher and four spots above GPT-6 Sol (Max), with output at one-fifth Astra's price.
5. Anthropic and SpaceX Agree on Up to $84.5B in Compute
NVIDIA GPUs in SpaceX datacenters per the S-1, terminable on 90 days' notice.
6. OpenAI Disrupts a Coordinated Distillation Campaign
Detected and shut down an organized effort to extract protected reasoning, active since the first week of July.
7. Trump's Voluntary AI Safety Pact Draws ~20 Signatories
Independent audits, regular consultation, and shared standards across cyber, bio, and chemical risks — non-binding but 'morally' so.
8. Ant Bailing's Ling-3.1-flash: Built for Long Real-World Tasks
~560B params, ~25B active per token, 1M context, hybrid linear architecture with 512 routed experts (8 + 1 shared).
9. DeepSeek Open-Sources Its Ascend Infra Stack
TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA, and DeepSelect mirror its NVIDIA components one-for-one.
10. DeepMind's SynthID Bio Watermarks AI Proteins
Invisible signatures in biological sequences stay verifiable on physical proteins without harming wet-lab function.
11. Perplexity Computer Opens Email Delegation
No account needed; forward or cc computer@perplexity.com and tasks run free for now.
12. Copilot Cowork's Gateway Hijacking Flaw Exposed
A malicious Skill can hijack the AI gateway to bypass the sandbox and exfiltrate files.
13. Anthropic: Robots Can Do 74% of Physical Work, Cheaply on 0.3%
Across ~19,000 tasks, cost parity on just 0.3%; ~40 years at 3%/yr cost decline to reach 10%.
14. Ataraxos Topples Top Human Stratego Players on the Cheap
A MIT/CMU/NYU/Stanford system crushes elite human Stratego players at minimal cost — published in Nature.
15. Memory Makers' Long-Term Deals Inflate RAM and SSD Prices
3-5 year LTAs tie 50-70% of capacity to a handful of big customers — consumer storage surged all year.
16. METR Chair Testifies on 'Rogue AI'
Chris Painter speaks to the Senate homeland-security subcommittee's 'Rogue AI' hearing.
17. Claude Sonnet 5.5 Direct Mode, 48 Hours on Arena
Open until Oct 2, 8 AM PT, then still available in Battle and Agent Mode.
18. Factory Automations Goes GA
Droid runs engineering workflows on schedules or Slack/GitHub/webhook triggers, with BYOK and your pick of machines.
19. OpenRouter's Three Agent-Testing Guides
Golden eval sets from production traffic (20-50 samples scaling to 100-1,000), regression after prompt/model changes, and tool-calling accuracy.
20. vLLM's Disaggregated-Serving Field Guide
Prefill/decode splitting, GPU-free render frontends, and how to combine them on v0.30.0+.
Sources and verification
This legacy briefing did not retain its original source URLs. Verify safety, financial, policy, and breaking-news claims with authoritative primary sources before acting on them.