endueendue

Industry

New models, tools and research, and what they mean for teams working with agents.

20 posts

Industry · Agents

Gemini 4 Argon, reviewed for teams that run agents: what you can decide today

Three days after launch, a second look at Gemini 4 Argon from an operator's seat. Where it leads and trails by type of work, what a task really costs, what a 1M-token output and a guardrail-free tier mean for how you run agents, and what you can prepare before the API opens.

12 min read

Industry

Gemini 4 Argon briefing: Google's numbers and the first reactions

Google's first Gemini 4 model leads 12 of the 18 benchmarks it published, writes up to 1M tokens per response and starts at $2/$10 per million tokens. Only vetted cyber defenders can use it today. Here are the key numbers, what independent testers measured and where people are skeptical.

4 min read

Industry · Agents

Gemini 4 Argon, read closely: 18 benchmarks, 1M output tokens and a gated launch

Google's first Gemini 4 model leads 12 of the 18 benchmarks in its own table, and the public leaderboards mostly agree. This post goes through who ran which test, how many tokens Argon spends per task, what a 1M-token response costs, and why only cyber defenders can use it today.

16 min read

Industry · Agents

How OpenAI Dots work: an engineer's guide to the always-on agent

Three days after launch, a systems view of Dots. It covers the cloud computer each one runs on, how a goal turns into actions, the checks that sit in between, what OpenAI's own tests say about where it slips, what changed since DevDay, and what you would need to build the same thing yourself.

13 min read

Industry · Agents

How ChatGPT Space works: pages, permissions and the agents inside

An engineer's look at ChatGPT Space three days after launch. How spaces, pages and files fit together, how access is inherited, what each person's agent can see, what the Compliance API exports, and what is still missing, from export to version history.

11 min read

Industry · Engineering

GPT-6.1 Sol, one day later: what independent tests and early users found

Artificial Analysis puts GPT-6.1 Sol one point below GPT-6 Astra at under a quarter of Astra's cost per task, and at about a tenth of Claude Sonnet 5.5's despite the same list price. Early user tests, live OpenRouter data, switching gotchas and the safety numbers.

15 min read

Industry · Agents

Dots after launch day: safety questions, the privacy FAQ and a checklist

OpenAI launched Dots a day after holding back GPT-6.1 Astra and apologizing to Australia. What its own safety numbers and privacy FAQ say, how the first day went (stalled demos, a five-hour outage), and a checklist built from OpenAI's help articles.

10 min read

Industry

OpenAI DevDay 2026: all 25 announcements at a glance

OpenAI's DevDay announcements from September 29, sorted into five groups. From Dots, the always-on agents, to GPT-6.1 Sol, ChatGPT Space, Codex and the new plans, see what shipped, who gets it and when.

5 min read

Industry · Agents

OpenAI Dots explained: the always-on agent inside ChatGPT

Dots, OpenAI's headline DevDay launch, are agents with their own cloud computer that stay on in ChatGPT, Slack and Teams and keep working between conversations. What they can do, who gets them, and how permissions and approvals work.

7 min read

Industry

ChatGPT Pro 500 is here. What changes for the $200 Pro plan?

OpenAI launched a $500-a-month Pro 500 plan at DevDay and cut the Codex and Work allowance of its $200 Pro plan. The new lineup, the dates for current subscribers, and how to use your ChatGPT plan in other apps.

6 min read

Industry · Agents

ChatGPT Space and Pages: where teams and agents share one document

OpenAI launched Space, a shared workspace inside ChatGPT, and Pages, a document type built for people and agents. What they do, what ships now and later, team tasks and Slack and Teams, and the permission and privacy details worth knowing.

7 min read

Industry · Agents

Codex keeps working after you close the laptop: the DevDay 2026 updates

At DevDay 2026, OpenAI reworked Codex with shared cloud environments, a CLI with voice and an agents view, code review in the ChatGPT desktop app, and Codex Security Cloud for GitHub repositories. What changed, who can use it, and what to check first.

8 min read

Industry · Agents

OpenAI's Decisions API, side by side with Jev

At DevDay, OpenAI previewed a Decisions API built on Luna that picks one answer from options you define and accepts images. What OpenAI has shared, the first head-to-head tests, how it differs from Jev, and what is still unannounced.

5 min read

Industry · Agents

Jev in practice: cheaper LLM apps and faster 3D characters

TypeSafe's Jev picks one answer from options you define, in under half a second. How to use it as a router that cuts LLM costs, as the reflexes of a 3D character and as a guardrail, with diagrams and the limits to know first.

10 min read

Industry · Agents

Sonnet 5.5 edged past Opus. How do you choose a model now?

Claude Sonnet 5.5 scored above Opus 5.5 on one benchmark at half the token price. At some settings it still costs more per task. The published numbers, the independent ones, and an order for choosing models for your own work.

5 min read

Industry · Agents

GPT-6 Astra in Unity: what changes for 3D teams

OpenAI's new model edits Unity scenes, plays the build and fixes what breaks. What the first examples show, where the gains come from, and what to set up before you try it.

5 min read