Welcome to Jumble, your go-to source for AI news updates. This week, OpenAI used DevDay to launch dots, always-on agents with their own computers. Meanwhile, Google finally unveiled Gemini 4 Argon, and the early reviews are mixed. Let’s dive in ⬇️

In today’s newsletter:
🤖 OpenAI’s agents never clock out
🛡️ Google’s Argon gets graded
🏛️ The FTC eyes rogue AI
🧵 Amazon’s tiny decision model
🎭 Weekly Challenge: Spot fake video calls

🫧 OpenAI's DevDay Was All About Agents

OpenAI ran through 20+ launches at DevDay last week, and nearly all of them push ChatGPT toward being a place where people and AI agents work side by side. The headliner was dots, always-on agents with their own cloud computer and access to 4,000+ apps.

Would you let an AI agent work for you around the clock?

Login or Subscribe to participate

🥊 Dots vs. Meta’s Muse

Meta’s Muse beat dots to market by three weeks with the same pitch: an agent on its own cloud computer that keeps working after you close the app.

Muse is free for most people and aimed at errands like bills and travel, while dots start on paid Pro and Business plans and live where you work.

🗂️ ChatGPT Gets Its Own Office Suite

OpenAI also launched Space, a shared workspace where coworkers, ChatGPT, and their dots build on the same files and docs. Its new Pages format can even update itself, like checking a team channel and adding what it finds, with collaborative slides coming next.

💸 Sol Gets Near-Astra Smarts for Less

On the model side, OpenAI released GPT-6.1 Sol, which nearly matches Astra on coding at one-fifth the price: $2 per million input tokens and $10 per million output.

♊ Google’s Gemini 4 Argon Is Back in the Race

Google unveiled Gemini 4 Argon on September 30, replacing the long-delayed Gemini 3.5 Pro as its new flagship. For now, it’s only going to 650+ vetted cyber defenders, with paying Gemini subscribers next in line.

🛠️ What It Can Actually Do

Argon is built for long, messy jobs like codebase rewrites and legal research, and can write up to 1 million tokens in a single response. Inside Google, Argon agents are porting an 800K-line kernel to Rust and freed up 300+ TiB of data center memory.

📊 Strong on Paper, Shaky at Coding

Google’s charts have Argon winning or tying 13 of 18 benchmarks, but independent testing ties it with GPT-6 Astra, behind Claude Opus 5.5. Coding is the weak spot, and some Google employees reportedly say it looks better on benchmarks than in real work.

Weekly Scoop 🍦

🎯 Weekly Challenge: Could a Fake Video Call Fool You?

❝

Challenge: A new AI video model called Griffin fooled nearly half of study participants into thinking they’d video called a real person. This week, test your radar and fake-proof your family.

Here's what to do:

🎬 Step 1: Watch the tape. Play the demo clips on Tavus’s Griffin page and note the second anything feels off. Griffin itself is still limited to select testers.

👂 Step 2: Give it 20 seconds. Start a voice chat in ChatGPT, Gemini, or Claude and try to trip it up: interrupt, go silent, switch topics. That’s about when Griffin’s doubters caught on.

🔑 Step 3: Write a family playbook. Ask your AI: “Create a family plan for video calls asking for money: one safe word, two check questions, and a hang-up-and-call-back rule.”

📲 Step 4: Run a fire drill. Share the plan with one family member and practice it on a real call, including hanging up and calling back on a saved number.

Would you hand your workday to a dot and a shared Space, or is ChatGPT trying to do too much? And will Gemini 4 Argon win you over once it opens up, or are the coding gripes a dealbreaker? See you next time!

Stay informed, stay curious, and stay ahead with Jumble!

Zoe from Jumble