AI Daily

🤖 AI HOT Daily · Sep 4, 2026

OpenAI drops GPT-6 Astra — the first model to hit the Critical cybersecurity threshold under its preparedness framework: a 1.05M context window, 72.6% on OSWorld V2-Offline, near-saturated ARC-AGI-3 (~99.9%), and broad wins over Claude Fable 5.1 at a lower price; ARC-AGI-3 saturated in just six months, twice as fast as Chollet expected; IFM open-sources six K2 Horizon models (0.9B–375B, Apache 2.0); Hugging Face ships funes, a local memory layer for coding agents; xAI launches Grok Bot and an enterprise tier (free two weeks); OpenAI's $1B Daybreak program backs frontline defenders; NVIDIA confirms its $12.93B acquisition of Hugging Face; Artificial Analysis rates Astra's coding agent on par with Fable 5 at 2.5x the price; Rohan Paul flags Astra's chain-of-thought control jumping from 16.1% to 60.9% with declining monitorability; Marcus and Chollet weigh in; Tunguz decodes Muse Spark's two-tier pricing; Google shows how to host a 24/7 agent on Cloud Run for $5.70/mo; Muse Spark 1.3 scores 68 in coding-agent index, just behind Claude.

  1. 1. GPT-6 Astra: First Model at Critical Cyberthreshold

    A 1.05M-token window, 128K max output, and 72.6% on OSWorld V2-Offline (vs 65.7% for GPT-5.6 Sol) — but gated behind the Critical cyber threshold.

  2. 2. GPT-6 Astra: SOTA Across Benchmarks, Beats Fable 5.1

    With ~99.9% on ARC-AGI-3, 100% on ExploitBench, SOTA on FrontierMath Tier 4 and TerminalBench-4.0, it tops the two-day-old Fable 5.1 across the board — cheaper too.

  3. 3. ARC-AGI-3 Saturated in 6 Months, 2x Faster Than Predicted

    Chollet expected near-saturation in about a year; it took six months — new models will upend views rooted in far older ones.

  4. 4. IFM Releases Six Open K2 Horizon Models

    Six sizes from 0.9B to 375B-A23B, all Apache 2.0; the 0.9B/3.7B/7B claim scale-class SOTA and 36B-A4B debuts the new MoVA sparse-attention architecture.

  5. 5. HF Open-Sources funes, a Memory Layer for Coders

    A local memory layer for Claude Code, Codex, pi, and Hermes that indexes past sessions into Lance datasets, so one 'funes add' lets an agent recall its provenance.

  6. 6. xAI Launches Grok Bot and an Enterprise Tier

    Built around persistent Bots with identity, memory, machines, and tools; the enterprise tier is free for Grok and Cursor Enterprise customers for two weeks.

  7. 7. OpenAI's Daybreak Puts $1B Behind Frontline Defenders

    A global program offering subsidized access, training, support, and partnerships — prioritized to water, grids, local governments, community banks, nonprofits, and OSS maintainers.

  8. 8. NVIDIA Confirms $12.9303B Acquisition of Hugging Face

    Jensen Huang confirms the deal; HF counts 18M+ developers, 3M+ models, 500k datasets, 1M apps, and 200k+ businesses.

  9. 9. Astra's Coding Agent Matches Fable 5 at Higher Price

    A 67 Coding Agent Index ties Opus 5 and Fable 5 at under half Fable 5's per-task cost, though list price hits 2.5x; token efficiency is ~70% above GPT-5.6 Sol (max).

  10. 10. Astra Card: Better Chain-of-Thought Control, Less Monitorable

    From the 117-page card: Astra's command over its own chain of thought jumps from 16.1% to 60.9%, with monitorability declining as a trade-off.

  11. 11. Astra Rolls Out: Cybersecurity First, Then All Plus

    Astra gates: Daybreak cyber customers first, then Pro, Plus, Enterprise, Business, and API within a week — with OpenAI moving 'as carefully and quickly as possible' toward all Plus users.

  12. 12. Google and Meta: Cloud-Run Agents and Muse Spark Pricing

    Google shows a 24/7 persistent agent on Cloud Run for $5.70/mo; Tunguz decodes the data-for-compute logic behind Muse Spark's two-tier API pricing.

Sources and verification

This legacy briefing did not retain its original source URLs. Verify safety, financial, policy, and breaking-news claims with authoritative primary sources before acting on them.