Friday, September 4, 2026
Coverage: 81 videos reviewed (1 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
Testers find Astra comparable to Fable 5.1 on writing, design and coding, with mixed results
Every's team gave mixed early-access verdicts on GPT-6 Astra versus Claude Fable 5.1: Dan Shipper uses Astra daily but reaches for Fable on the largest tasks. Mike Taylor reported that Astra narrowly beat Fable 5.1 across 50 blind writing comparisons on one article, and that it fully completed his Typeform-clone check. Kieran said Astra's one-shot rewrite of Every's Proof editor had errors that Fable 5 and 5.1 did not. Nate Herk found Astra's one-shot sites better designed than Fable 5.1's, and Matthew Berman saw a recurring forest-green, flat design style. 1littlecoder, who did not test the model, judged demos incremental versus Fable 5.1; The AI Advantage relayed Shipper's view that Astra is the best writing model he has tried. Each test was a single small sample.
- Evidence: 0 first-party, 3 hands-on, 2 relaying
- Watch: Every: VIBE CHECK: GPT-6 ASTRA; Nate Herk: GPT-6 Astra FINALLY Kills AI Website Slop (high hype)
Continuing stories
- OpenAI releases GPT-6 Astra on Sept. 3 with staged access - OpenAI released GPT-6 Astra on Sept. [1 first-party, 0 hands-on, 7 relaying] Watch: OpenAI: Introducing GPT-6 Astra for developers
- ARC Prize reports Astra 62.7% on ARC-AGI-3 in standard harness, 99.9% with adapter - ARC Prize reported that GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set under its standard harness at max reasoning, and 99.9% under a provider adapter that preserves private reasoning state and uses OpenAI compaction, according to AI Code King and Prompt Engineering, who relayed ARC Prize data. [0 first-party, 0 hands-on, 5 relaying] Watch: Prompt Engineering: GPT-6 Astra: The harness matters more than you think
- OpenAI prices GPT-6 Astra at $10 input and $50 output per million tokens - Astra's API price is $10 per million input tokens and $50 per million output tokens, matching Anthropic's Claude Fable 5.1 and 2.5 times GPT-5.6 Sol's $4 and $20, according to AI Code King and Theo, who relayed OpenAI's launch figures. [0 first-party, 0 hands-on, 2 relaying] Watch: Theo - t3.gg: It's Here.
- OpenAI reports Astra at 71.6% to 73% on OSWorld 2.0, versus 65.7% for GPT-5.6 Sol - OpenAI's launch materials, as relayed by four channels, put GPT-6 Astra's OSWorld 2.0 computer-use score at 72.6% versus 65.7% for GPT-5.6 Sol, with tasks finishing in about 40 minutes versus about 75. [0 first-party, 0 hands-on, 4 relaying] Watch: AI Code King: GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model an
- OpenAI classifies Astra as first model at critical cyber level in its Preparedness Framework - OpenAI said GPT-6 Astra is the first of its models to reach the critical cyber threshold in its Preparedness Framework, according to AI Code King, Fireship and Julian Goldie, who relayed the system card. [0 first-party, 0 hands-on, 4 relaying] Watch: AI Code King: GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model an
- System card: Astra shows lower chain-of-thought monitorability in OpenAI tests - OpenAI's system card for GPT-6 Astra reports lower chain-of-thought monitorability than earlier models, according to AI Explained, and OpenAI researchers cited by the speaker worry it may sandbag on safety tasks. [0 first-party, 0 hands-on, 2 relaying] Watch: AI Explained: GPT 6 Astra, so good even OpenAI are worried
Also notable
- Astra scores 97.6% to 98% on Frontier Math Tier 4; Epoch says it solved 2 of 68 unsolved problems - OpenAI's launch materials put Astra at 97.6% on Frontier Math Tier 4, against 83% for GPT-5.6 Sol and 87.8% for Fable 5.1, according to AI Code King; AI Explained gave about 98% at peak and 83% without reasoning. [0 first-party, 0 hands-on, 3 relaying] Watch: AI Code King: GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model an
- Astra scores 57.9% on Terminal Bench 4.0 versus 55.8% for Fable 5.1, with near-ties elsewhere - On Terminal Bench 4.0, GPT-6 Astra scored 57.9% versus 55.8% for Claude Fable 5.1, 52.3 for Claude Opus 5 and 37.3 for GPT-5.6 Sol, AI Code King reported from OpenAI's table. [0 first-party, 0 hands-on, 2 relaying] Watch: Theo - t3.gg: It's Here.
- Artificial Analysis rates Astra 61 on its Intelligence Index, level with GPT-5.6 Sol and below Fable 5.1 - Artificial Analysis gave GPT-6 Astra 61 on its Intelligence Index, equal to GPT-5.6 Sol and five points behind Claude Fable 5.1 at 66, according to AI Code King and Fireship. [0 first-party, 0 hands-on, 3 relaying] Watch: AI Code King: GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model an
- Testers report Astra completing long computer-use tasks, with refusals and skipped steps noted - Testers who used GPT-6 Astra's computer use reported long unattended runs. [0 first-party, 4 hands-on, 0 relaying] Watch: Nate Herk: GPT-6 Astra Made This Entire Video (high hype)
- Users report Astra Codex runs lasting days, and a repo latency cut from 800 ms to under 30 ms - Matthew Berman said he ran /goal with Astra for five straight days to build a SimCity clone that was still unfinished, with lag fixed after he asked for frame-rate optimization. [0 first-party, 2 hands-on, 1 relaying] Watch: Theo - t3.gg: It's Here.
- OpenAI presents ChatGPT Work mode with plugins for files, email and scheduled automations - OpenAI described ChatGPT Work as a switchable mode in which plugins connect to Drive, Office, email, Figma, Canva and Notion so ChatGPT can create documents, decks and apps and run scheduled computer or browser automations. [1 first-party, 1 hands-on, 0 relaying] Watch: Riley Brown: 8 ChatGPT Agents That Do My Work for Me (Steal These)
- Testers report Fable 5.1 finishing long tasks, including a 7-sheet DCF model and a Blender film - Nate B Jones tested Claude Fable 5.1 on several tasks. [0 first-party, 1 hands-on, 1 relaying] Watch: Nate B Jones: Everyone's Testing Claude Fable 5.1 On Code. It Made Me A 37-Second Fi
- Institute of Foundation Models releases six K2 Horizon open models under Apache 2 - The Institute of Foundation Models (MBZUAI) released six K2 Horizon models from sub-1B to a 375B mixture-of-experts flagship, with Apache 2 weights, training data, code and checkpoints, according to Fahd Mirza and Julian Goldie, who relayed the lab's claims. [0 first-party, 1 hands-on, 1 relaying] Watch: Fahd Mirza: K2 Horizon: 0.9B, 7B, and 32B Tested Locally, Real Results
- MiniMax M3 open model reported at roughly 400B parameters with 1M context - A MiniMax guest on the AI Engineer stream said the company released M3, an open model with roughly 400B total and 20B active parameters, vision input and 1M-token context, and is already building M3.1. [0 first-party, 0 hands-on, 1 relaying] Watch: AI Engineer: Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, M
- Alibaba's Qwen 3.8 Max described as 2.4T-parameter model with 1M-token context - Julian Goldie and GitHub's presenter said Alibaba's Qwen 3.8 Max has 2.4 trillion parameters and a one-million-token context. [0 first-party, 0 hands-on, 2 relaying]
Models & learning
- Google releases Gemini 3.8 Flash on Sept. 2 with 1M-token context and three thinking levels - Google released Gemini 3.8 Flash on Sept. [0 first-party, 0 hands-on, 2 relaying] Watch: Matt Wolfe: AI News: The Most Insane Week So Far This Year!
- Sentdex reports GLM 5.3 Flash running about 170-180 tokens per second locally - Sentdex said GLM 5.3 Flash, about 320B parameters with vision, runs at about 170 to 180 tokens per second at native precision on his RTX Pro 6000 setup, versus about 350 for DeepSeek V4 Flash 0731. [0 first-party, 1 hands-on, 1 relaying] Watch: Sentdex: All Roads Lead back To GLM!
- Hugging Face releases 200+ WebGPU kernels and a Kernels JavaScript library - Hugging Face released more than 200 open-source WebGPU kernels and a JavaScript Kernels library, shipped as Jinja templates that render WGSL for the device and load from the Hub. [1 first-party, 0 hands-on, 0 relaying] Watch: Hugging Face: We shipped 207 WebGPU Kernels for Browser AI
- MiniMax describes sparse attention and native multimodal pretraining for M3 - A MiniMax guest said M3 uses MiniMax Sparse Attention, an index branch selecting context blocks and a sparse branch computing on them, designed by an intern for 1M-token agent contexts. [0 first-party, 0 hands-on, 1 relaying] Watch: AI Engineer: Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, M
- Bijan Bowen tests Fish Audio S2.1 Pro voice cloning from 15 seconds of audio - In a sponsored test, Bijan Bowen judged Fish Audio S2.1 Pro's instant clone from a 15-second clip very convincing, though the judgement was subjective and he tested Chinese output with two takes of differing accent quality. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: Fish Audio S2.1 Pro Full Test – Building A Video Translation Pipeline!
- GPT-5.6 Sol builds a transcribe-translate-dub pipeline in Bijan Bowen's test - Bijan Bowen used GPT-5.6 Sol in the ChatGPT Mac app at extra high to build a local pipeline using a Whisper-style transcriber, Qwen 3.8 Next (Q4_K_M) for Chinese translation, Fish API dubbing and a web UI, on a 6000 Pro machine. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: Fish Audio S2.1 Pro Full Test – Building A Video Translation Pipeline!
- Reviewers discuss Meta paper on LLM research preference models for choosing experiments - Hugging Face paper reviewers discussed a Meta paper in which research preference models pick which candidate experiments to evaluate, with an inference-only variant and an agentic variant that can run pilot experiments. [0 first-party, 0 hands-on, 1 relaying]
- Prompt Engineering finds DeepSeek V4 Flash cost varied widely across nine harnesses - In his test of 20 long-horizon coding tasks across nine harnesses, Prompt Engineering's host found naive API cost of about $152 versus about $18.82 actual because roughly 97% of input tokens were cache reads. [0 first-party, 1 hands-on, 0 relaying] Watch: Prompt Engineering: GPT-6 Astra: The harness matters more than you think