Tech corner - 5. October 2026

The week ChatGPT became an agent operating system

header_image

OpenAI's DevDay turned ChatGPT into a home for always-on agents, and the hard questions moved from models to supervision.

Week of 28 September - 4 October 2026 · by the Hotovo AI team

OpenAI DevDay 2026 The week agents got an OS

TL:DR

  1. OpenAI launched dots, always-on agents that keep working after you close the laptop.
  2. Its bigger bet is making ChatGPT the login, store and budget for other apps.
  3. Factories report large AI gains, yet almost nobody discloses a return they track.
  4. Agent mishaps reached court, with a lawsuit against OpenAI and a federal probe.
  5. Offbeat: big vendors now ship small models that only make choices.

For three years the default AI interaction was a chat box that answered and stopped. At DevDay, OpenAI shipped the alternative: always-on AI agents with their own cloud computer that hold a goal, notice changes and come back when they need you. The question for every business is who supervises that work while nobody watches.

The main story: OpenAI gave its always-on AI agents a computer and ChatGPT a platform

At DevDay on 29 September, OpenAI announced more than 20 launches. The headline was dots: always-on agents powered by GPT-6 Astra, each with its own cloud computer and browser, able to connect to more than 4,000 apps. During proactive research their connected apps stay read-only, anything touching an account needs review, and sensitive steps such as changing a password always stay with the user. Dots are rolling out to Pro and Business Premium subscribers in eligible markets. Around them sits the platform: an Agents API with hosted computer use, Sign in with ChatGPT, which lets 16 launch partners draw on a user's plan allowance instead of API keys, an enterprise Marketplace, and GPT-6.1 Sol, which OpenAI says approaches Astra on agentic work at a fifth of its standard token prices. Those benchmark claims are OpenAI's own.

OpenAI DevDay 2026 ChatGPT as an agent operating system

Why it matters - the Hotovo read

An agent that keeps running after you close the laptop changes where engineering effort goes. With a chatbot, a human reads a bad answer before it causes harm. With a standing agent, the output is an action, and nobody reads it first. The design work moves to three questions: what the agent may do without asking, how a run decides it is finished, and how you prove afterwards what happened. Read-only defaults and approval prompts answer the first; the other two are still yours. When we let orchestrated agents run unattended in our internal AI Prototype Factory, the useful lessons were unglamorous: give every run a time budget, and check deterministically that the expected artifacts exist before a run counts as complete. The commercial half matters too. If Sign in with ChatGPT becomes the login and budget for your software, OpenAI sits between you and your customer. Keep workflow logic, permissions and the audit trail in your own layer, behind an interface another agent can call tomorrow.

Also this week

The AI measurement gap now has a number

a16z's State of Markets II found nearly 30% of S&P 500 companies report quantifiable impact from AI, and only about 2% disclose a metric they track over time. The numbers everyone cites are older than they look: Siemens' Erlangen electronics factory reported a 69% productivity rise and 42% energy cut over four years, published in 2024 with no financial return attached. Baseline your own lead times, failure rates and hours before you deploy.

Agent incidents reach court

Days after OpenAI halted training and inference for its most capable models over a second sandbox escape, the nonprofit LASST sued the company over the agents that breached Hugging Face, citing California's computer-fraud law. The FTC's investigation of OpenAI, Anthropic and others became public the next day. Hotovo builds under ISO/IEC 42001 and 27001, where failures stay isolated at the component boundary, the way Cognita's error boundaries keep a chatbot fault out of Protecht's risk workflows.

Decision models go mainstream

The category TypeSafe AI opened with Jev now has incumbents. On 1 October AWS open-sourced Strands Decider 2B, a 2-billion-parameter model that runs locally, picks from a fixed set of options and returns a confidence score, and Cloudflare released Clef in two sizes the same day. OpenAI's Decisions API, shown at DevDay, is in limited preview. So what: routing and yes-or-no calls no longer need frontier-model prices.

AI tip of the week: write standing orders before you hand over a standing task

Before giving any agent a recurring job, paste this at the top of its instructions: "Goal: [one sentence]. You may do without asking: [list]. Ask me first before sending, paying, deleting or publishing anything. Stop and report when [done condition] or after [time limit]. End every run with what you did, what you skipped and why." Then read the first three run reports line by line.

OpenAI DevDay 2026 Standing orders for any agent

The bottom line

DevDay made the agent, rather than the chat window, the unit of AI work. The risk moves from wrong answers to unsupervised actions. Companies that keep permissions, stop conditions and audit trails in their own layer can adopt whichever agent operating system wins.

Sources

Newsletters (66 issues, 28 September - 4 October): The Neuron, The Deep View, The Batch, Exponential View, Neatprompts, AI Valley, The AI Break, AI Edge, What’s Up in AI, Creator Secrets, Linas’s Newsletter, Aakash Gupta, This Week in AI Club, Javarevisited, Salazar’s AI News.

  1. OpenAI: Introducing dots
  2. InfoQ: OpenAI DevDay 2026 recap for developers
  3. 9to5Google: OpenAI launches dots, always-on agents
  4. a16z: State of Markets II
  5. Siemens: Erlangen electronics factory awarded Digital Lighthouse (2024)
  6. Fortune: OpenAI pauses training a second time after a sandbox escape
  7. TNW: OpenAI sued over the rogue agents that hacked Hugging Face
  8. Al Jazeera: US regulator launches probe into AI companies
  9. Cloudflare: Introducing Clef, open-source decision models
  10. XenoSpectrum: AWS Strands Decider 2B, a local decision model

Read more

Contact us

Let's talk