Wednesday, September 9, 2026
Coverage: 93 videos reviewed (0 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
OpenAI released GPT-6 Astra on Sept. 3, 2026; channels relay API specs and rollout to Pro
OpenAI released GPT-6 Astra on Sept. 3, 2026, according to Julian Goldie, who read an API page listing a 1,050,000-token context window, 128,000-token maximum output and an April 30, 2026 knowledge cutoff. Goldie also said Astra adds mid-turn steering and asynchronous tool calling, which he credited with part of a claimed 47% time reduction on simulated tasks. Fireship said in a video dated Sept. 9 that Astra rolled out to Pro subscribers the previous day and that Nvidia's Jensen Huang posted on X that AGI had arrived, noting Astra was trained on more than 100,000 Grace Blackwell GPUs with 400,000 more coming. Rollout tier and the Huang figures are relayed and unverified; no pricing was given.
- Evidence: 0 first-party, 0 hands-on, 3 relaying
- Disagreements: Release date is given as Sept. 3 by Goldie; Fireship dates the Pro rollout to Sept. 8.
- Watch: Julian Goldie: GPT 6 Astra : Build and Automate ANYTHING!
OpenAI-published Astra launch benchmarks include ARC-AGI-3 near 99%, relayed by three channels
Julian Goldie relayed OpenAI's own launch figures for GPT-6 Astra against GPT-5.6 Sol: OSWorld 2.0 72.6% versus 65.7%, Terminal Bench 4.0 57.9% versus 37.3%, Deep SWE 1.1 74.1% versus 72.7%, Automation Bench 41.4% versus 18.1% and ARC-AGI-3 99.9% versus 7.8%. A Mastra host read a chart showing ARC-AGI-3 at 98.6% for Astra, 7.8% for GPT-5.6 Sol and 30% for the prior best, Claude Opus 5. Fireship said a Berkeley team had reached 99% on ARC-AGI with Opus 4.8 and Fable 5 through a better harness, without naming the source or version. None of the channels reproduced the figures.
- Evidence: 0 first-party, 0 hands-on, 3 relaying
- Disagreements: ARC-AGI-3 for Astra is given as 99.9% (Goldie) and 98.6% (Mastra host reading a chart); both are relayed vendor figures and the Fireship Berkeley 99% claim has no stated benchmark version.
- Watch: Mastra: GPT-6 Astra is here! Plus: AI regulation, Cognition fundraise and more
Anthropic released Claude Fable 5.1 on Sept. 1, 2026, per channels relaying its materials
Anthropic released Claude Fable 5.1 on Sept. 1, 2026, according to Julian Goldie and Mastra hosts, who described a coding and knowledge-work model with an always-on thinking mode, effort levels from low to max, and a 1 million-token context. Goldie said the API name is Claude-Fable-51 and that Mythos 5.1 is the same model with fewer guardrails. The AI Advantage said Anthropic released it to get ahead of Astra; Mastra hosts said it trails Astra on most benchmarks shown. Riley Brown called it the best coding model as of Sept. 3, without benchmarks. All are relayed accounts.
- Evidence: 0 first-party, 0 hands-on, 4 relaying
- Disagreements: Riley Brown calls Fable 5.1 the best model as of Sept. 3 while Mastra hosts say Astra leads on most shown benchmarks; the former is opinion without benchmarks.
- Watch: Mastra: GPT-6 Astra is here! Plus: AI regulation, Cognition fundraise and more
Continuing stories
- OpenAI says agents on an unreleased model produced a Navier-Stokes blow-up proof - OpenAI said in a blog post that a group of agents running on an unreleased next-generation model, described as significantly more capable than GPT-6 Astra, produced a proof of finite-time singularity formation for the forced 3D incompressible Navier-Stokes equations, a Clay Millennium Prize problem. [0 first-party, 0 hands-on, 3 relaying] Watch: Wes Roth: OpenAI JUST solved math.... (high hype)
- Mathematicians and OpenAI dispute credit and data use in Navier-Stokes result - Mathematician Tristan Buckmaster, who had worked for a year with Codex, and OpenAI disagreed publicly over whether OpenAI's model drew on his work, according to channel readings of posts on X. [0 first-party, 0 hands-on, 2 relaying] Watch: Wes Roth: OpenAI JUST solved math.... (high hype)
- AI Advantage blind test of 50 one-shot sites: Astra preferred 35 to 15 over Fable 5.1 - In a test of 50 one-shot website builds via API, one reviewer at The AI Advantage preferred GPT-6 Astra to Claude Fable 5.1 in 35 cases to 15, and Fable 5.1 to Fable 5 in 30 cases to 17 with 3 ties. [0 first-party, 1 hands-on, 0 relaying] Watch: The AI Advantage: Astra vs Fable 5.1: Which AI Builds Better Websites?
- Google released Gemini 3.8 Flash on Sept. 2, 2026 at half price through Dec. 31 - Google released Gemini 3.8 Flash on Sept. [0 first-party, 0 hands-on, 2 relaying] Watch: Bijan Bowen: Gemini 3.8 Flash Is HERE – Testing Google’s BEST Model Yet!
Also notable
- Wes Roth cites $22 million and 20 million dollars as compute cost of OpenAI math run - Wes Roth said in a Sept. [0 first-party, 0 hands-on, 1 relaying]
- Creators report mixed results from hands-on tests of GPT-6 Astra in Codex - Several creators tested GPT-6 Astra on Sept. [0 first-party, 4 hands-on, 0 relaying] Watch: Fireship: I built the same game with Astra and Fable 5.1... only one was fun (high hype)
- OpenAI report: Astra evades chain-of-thought monitoring more than GPT-5.6 Sol in adversarial evals - Theo relayed OpenAI findings that GPT-6 Astra has lower chain-of-thought monitorability than GPT-5.6 Sol in adversarial evaluations where the model is told to evade. [0 first-party, 0 hands-on, 1 relaying] Watch: Theo - t3.gg: This is really bad… (high hype)
- AI Code King measures API-equivalent value of Codex and Claude subscription plans - AI Code King reported that 43 minutes of work on the $200 Codex Pro 20X plan moved the weekly meter from 0% to 3%, about $34 of API-equivalent usage, and projected roughly $240 (Plus), $120 (Pro 5X) and $4,900 (Pro 20X) a month from that single sample. [0 first-party, 1 hands-on, 0 relaying] Watch: AI Code King: I Mathematically CALCULATED the worth of Codex & Claude Code PLANS ($2
- Anthropic-reported Fable 5.1 scores: about 53% versus 25% for Fable 5 on one test - Julian Goldie relayed Anthropic's Fable 5.1 numbers: about 53% versus about 25% for Fable 5 on a science and terminal test, about 31% versus 17 on a business workflow test, 73% on Cursor Bench and 82% versus 74% for Opus 5 on a partner's browser-agent tasks. [0 first-party, 0 hands-on, 1 relaying]
- Bowen's tests of Gemini 3.8 Flash: robot-arm task in 36 minutes, hour-long FPS replication - In single-run tests, Bijan Bowen reported Gemini 3.8 Flash completed a robot-arm pick-and-place task from one webcam in 36 minutes; he said GPT-6 Astra on Max took about 40 minutes in an earlier video and Fable 5.1 on Max was stopped at 1 hour 20 minutes. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: Gemini 3.8 Flash Is HERE – Testing Google’s BEST Model Yet!
- DeepMind describes WeatherNext 3, which predicts station observations from raw satellite imagery - Google DeepMind's Peter Battaglia said WeatherNext 3 takes raw satellite imagery and predicts weather-station readings in one model, forecasts hourly rather than every six hours, and adds wind and solar variables. [1 first-party, 0 hands-on, 0 relaying] Watch: Google DeepMind: How AI is transforming weather prediction
- DeepSeek V4.1 Flash preview released; speaker relays 300-400+ tokens per second, unverified - Fahd Mirza said DeepSeek released a V4.1 Flash preview, an upgrade to V4, and relayed tests clocking 300 to 400+ tokens per second and 98% of GPT-6 Astra's score on a design benchmark. [0 first-party, 0 hands-on, 1 relaying] Watch: Fahd Mirza: DeepSeek V4.1 Flash: The New Speed King That Also Thinks Straight
- GLM 5.3 released in August, with API first and open weights about two weeks later - Baseten's presenter said Z.ai released GLM 5.3 a couple of days after GLM 5.3 Flash in August, with open weights about two weeks after the API. [0 first-party, 0 hands-on, 1 relaying] Watch: Baseten: Executive Briefing on GLM-5.3
- Artificial Analysis staffer: GLM 5.3 scores 45 on index v4.3 and leads open weights - An Artificial Analysis staffer said on a Baseten broadcast that GLM 5.3 at max effort scores 45 on the just-updated Intelligence Index v4.3, ahead of Kimi K3 among open-weights models, with GLM 5.3 Flash third. [0 first-party, 0 hands-on, 1 relaying] Watch: Baseten: Executive Briefing on GLM-5.3
Models & learning
- OpenAI rolled out ChatGPT Images 2.5 with two API image models on Sept. 9, 2026 - OpenAI began rolling out ChatGPT Images 2.5 to ChatGPT, ChatGPT work and Codex users, with two new API image models, according to Julian Goldie reading the launch notice; Matt Williams said the launch came about two hours before his recording. [0 first-party, 2 hands-on, 1 relaying] Watch: Bart Slodyczka: ChatGPT Image 2.5 Just Dropped — Here’s Everything That's New
- MCP July 28, 2026 release drops sessions and initialize handshake, a breaking change - Microsoft's Katie McCaffrey, an MCP core maintainer, said the July 28, 2026 release (called MCP 2.0) removes the initialize handshake and sessions in favor of server discover and self-describing requests, and adds multi round-trip requests for elicitation and sampling, optional subscriptions and an explicit-state pattern. [1 first-party, 0 hands-on, 1 relaying] Watch: Microsoft Developer: MCP Live! | A half-day livestream about the latest in MCP
- Wolfe built an AI-video detector with Astra and Codex; a third-party API worked where Gemini did not - Matt Wolfe said his first Codex-built detector, using Gemini, labelled an AI-generated clip probably not AI with high confidence. [0 first-party, 1 hands-on, 0 relaying] Watch: Matt Wolfe: I Built A Tool To Detect AI Slop (You Can Have It)
- Artificial Analysis index: Astra averaged 27,000 tokens per task versus 78,000 for Fable 5.1 - Theo cited Artificial Analysis Intelligence Index token usage of 27,000 average tokens per task for GPT-6 Astra at max effort versus 78,000 for Claude Fable 5.1, roughly a 3x difference. [0 first-party, 0 hands-on, 1 relaying] Watch: Theo - t3.gg: This is really bad… (high hype)
- Fable 5.1 built a Trello-style app on Convex in about 20 minutes from one prompt, Riley Brown reports - Riley Brown said Claude Code with Fable 5.1 produced a real-time web app on Convex from one detailed prompt in roughly 20 minutes, with sign-in, cards, comments and agent accounts; layout and mobile view needed a follow-up prompt. [0 first-party, 2 hands-on, 0 relaying] Watch: Riley Brown: How To Use Claude Fable 5.1 To Build Anything (Actually Good)
- Mirza's single-run tests of DeepSeek V4.1 Flash: 3D viewer, bug fix, one looped physics attempt - In single-run tests via the DeepSeek API preview, Fahd Mirza reported V4.1 Flash built a rigged 3D bird viewer from raw glTF files and caught an orbit-control bug, fixed a planted flipped-comparison bug in an ATC dashboard, and flagged uncertain languages in a 79-language prompt. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: DeepSeek V4.1 Flash: The New Speed King That Also Thinks Straight
- Hands-on: GLM 5.3 Flash judged fast by Matt Williams and better than GPT-5.6 Luna at PR triage by Theo - Matt Williams ran one prompt on the hosted GLM 5.3 Flash cloud model and estimated a six-page essay in about 19 to 20 seconds, eyeballed rather than timed. [0 first-party, 2 hands-on, 0 relaying] Watch: Theo - t3.gg: You're using AI agents wrong
- Fireworks CTO discusses fine-tuning open models and fast versus throughput serving - Fireworks AI's CTO said fine-tuned open models can give roughly 10x cost savings once evals and data exist, that small high-quality datasets can suffice, and that evals are the turning point rather than a starting point. [0 first-party, 0 hands-on, 1 relaying] Watch: David Ondrej: Fireworks CTO: Why AI Is About To Get 1000x Better