Today for AI

InfoQ 中文 · 10/9/2026, 17:55:11

OpenAI DevDay Deep Dive: Dot Agent Enables Autonomous Cloud Linux Operations; Decisions API Prototype Built in One Week

By 林绮蓓Original title: OpenAI 高管亲述:我们是怎么在一周内做出 Jev 竞品的
78AI Score
Executive Summary

OpenAI launched the Dot agent and computer use capabilities at DevDay, enabling autonomous execution of browser and desktop tasks within cloud Linux environments via multimodal perception (screenshots, accessibility APIs, code). Key breakthroughs include significantly improved error handling and retry mechanisms, alongside the new Decisions API inspired by Jev. The latter leverages Luna weights with optimized inference for millisecond-level structured decision returns. However, human oversight boundaries for critical actions like payments remain undefined.

SOURCE COVERAGEOriginal coverage

Contents6 sections

OpenAI launched Dot and computer use capabilities at its September Developer Day, enabling agents to run browsers and desktop applications in cloud-based Linux environments. The core breakthrough lies in the agents' multimodal perception (screenshots, accessibility APIs, code) and autonomous error-handling and retry capabilities. Currently, clear boundaries for human intervention in critical operations (such as payments) still need to be defined.

Agents now support multimodal interface understanding and interaction; error handling and retry mechanisms have significantly improved task completion rates; however, the boundary of human-machine responsibility (e.g., for financial operations) has not yet been standardized.

This article is suitable for AI engineers, agent architects, and API platform developers.

Translation | Lin Qibei Editing | Cai Fangfang

Computer use results are seemingly easy to verify, so why has progress in this technology been so slow? When will AI truly learn to operate computers?

In June of this year, prominent podcast host Dwarkesh Patel raised this question, sparking considerable debate within the industry.

Three months later, OpenAI showcased Dot, computer use, and a series of new developer interfaces at its Developer Day on September 29. Each Dot instance can utilize an independent cloud-based Linux computer, capable of opening browsers and running desktop applications; users can delegate tasks that originally required step-by-step manual execution in web pages or software to it. After the product demonstration, the question became even more worthy of deeper inquiry: How far can computer use currently go, and what changes have occurred in the capabilities supporting agents to complete tasks?

Following Developer Day, Latent Space’s The AI Engineer Podcast conducted in-depth discussions with OpenAI’s CUA team and API platform leads. The podcast first interviewed Ari Weinstein, co-founder of Sky and current head of Computer Use Product and Engineering at OpenAI. Ari believes that the most critical advancement over the past year is that agents have become better at troubleshooting and retrying after encountering obstacles. They are no longer limited to clicking through screenshots step by step: page structure, accessibility information, and self-written code can all serve as ways for them to operate software. However, when agents can access websites and process payments, defining where humans should intervene remains a question that products must answer.

Nikunj Handa, OpenAI’s API Product Lead, shared various practices implemented at the API layer for agent development during the interview: from allowing models to continue reasoning while tools are executing, to rapid decision-making, performance optimization, and context management. When discussing the Decisions API, Nikunj Handa acknowledged that its initiation was inspired by Jev and was not originally part of the R&D plan. Within about a week of starting the project, the team had completed a working prototype. Instead of retraining the model, they reused Luna’s weights, constrained outputs into structured results, and specifically optimized the inference system to minimize the time-to-first-decision; multiple questions could also be processed in parallel batches. It functions more like an ultra-fast classification and decision layer, applicable to customer service ticket classification, evaluation scoring, and computer use, and can also integrate with GPT Live to make agents more responsive when quick judgments are needed.

Starting from the skepticism of "why AI isn't good at operating computers yet," the two interviews gradually zoomed in: how do agents discover their mistakes, decide on the next steps, and sustain work through long and tedious real-world tasks. The Developer Day demos showed the results, while these interviews offered insights into their practices regarding computer use from different perspectives of product and platform.

TL;DR:

Host Vibhu: Dot provides each agent with an independent cloud computer. How should users explore its capabilities?

Ari Weinstein: Dot can use browsers and desktop applications on cloud-based Linux computers. A practical starting point is to identify daily computer tasks that consume significant time and try delegating suitable ones to the agent.

Host Swyx: There is a view that computer use has made little progress in the past two years. What is the actual situation?

Ari Weinstein: In the past, agents could typically initiate tasks but often stalled when encountering issues. Now, they are better at troubleshooting errors, adjusting methods, and retrying. This is one of the most notable changes in the past year.

Host Vibhu: Is the capability improvement primarily driven by the model or the agent runtime framework?

Ari Weinstein: Both play a role. Agents can now combine screenshots, accessibility information, page structure, and Playwright to operate software, and can also write code to execute multiple steps at once; the model's own capabilities and speed are also improving.

Host Vibhu: What are the main bottlenecks for computer use going forward?

Ari Weinstein: Bottlenecks are distributed across the model, inference, runtime framework, and information presentation. As agent execution speeds increase, the latency inherent in operations such as waiting for website loads becomes increasingly apparent.

Host Swyx: What should developers pay attention to when using computer use capabilities via the Agents API?

Ari Weinstein: Limit the websites and applications accessible to the agent based on the task, and seek user consent before operations with potentially significant consequences, such as payments. Reliability and appropriate safety checks are the foundation for building user trust.

Host Vibhu: How will computer use change the software development workflow?

Ari Weinstein: Agents can actually open and test the software they develop. This allows coding and result verification to form a closed loop, reducing the scenario where testing is entirely handed off to humans after development is complete.

Host Vibhu: What noteworthy changes are there in the APIs released this time?

Nikunj Handa: Asynchronous function calls allow the model to continue execution while tools are running and retrieve results later; mid-run steering allows developers to inject messages during model execution. These capabilities help handle tasks with frequent tool calls and long execution times.

Host Swyx: How did OpenAI start developing the Decisions API?

Nikunj Handa: The launch of Jev attracted attention from users and internal teams; previously, the Decisions API was not in the development plan. The team first validated the feasibility of the prototype before optimizing the speed of decision returns.

Host Vibhu: What problems is the Decisions API suitable for solving?

Nikunj Handa: The primary use case right now is rapid classification, such as processing customer support tickets. It may also be used for tasks that require quickly selecting the next action, but fast decision-making requires different capabilities than completing complex, long-horizon tasks.

Host Swyx: If we’re using Luna in both cases, what’s the difference between the Decisions API and standard structured output?

Nikunj Handa: The first version didn’t involve retraining the model. Instead, it built upon existing Luna weights by imposing constraints on the output, processing multiple questions in parallel, and specifically optimizing for time-to-first-decision.

Host Swyx: How do long-running agents handle context limits?

Nikunj Handa: Developers can set thresholds to allow the Responses API to automatically compress the context, or they can call /compact to manually control when compression occurs.

When AI Gets Its Own Computer: What Can We Delegate to It?

Host Vibhu: Today is OpenAI DevDay, and we’re recording this special interview live from the venue.

Host Swyx: We’re the first podcast you’ve done after your livestream ended.

Host Vibhu: We have Ari with us today; he leads the product and engineering team for Computer Use agents. Before diving deep into the technology, could you recap what was announced today?

Ari Weinstein: We just came out of the keynote, and there are several notable announcements related to computer use.

First is Dot, a new personal assistant product featuring some very promising computer-use capabilities.

Then there’s GPT-6.1 Sol, an excellent new model that I believe is particularly well-suited for computer use due to its advantages in cost and speed. I think we mentioned that its cost is one-fifth that of Astra; if you look specifically at computer use, the cost drops to one-seventh, which is truly remarkable.

The Agents API has also added computer-use capabilities, allowing developers to build products using the same implementation found in Codex and ChatGPT. There were also demos of existing features like app shots, which let you quickly bring content you’re working on in your computer into Codex and ChatGPT.

There’s also native computer use on Mac: Roman had it automatically take screenshots of his apps, and while it operated the applications, he could still do other things on his computer. So, the keynote was indeed impressive.

Host Swyx: Not to mention the Decisions API. Let me ask directly: Are these all powered by the same underlying model, or are they different models distilled from the same dataset? Specifically, does computer use rely on the Decisions API, or are the two relatively independent?

Ari Weinstein: The Decisions API is interesting because it introduces several new capabilities: it supports parallel inference without chain-of-thought reasoning, and it uses a smaller model than the one we use for computer use. This makes it very fast, though slightly less capable when handling complex tasks with longer execution cycles. How best to combine these approaches remains an open research question. I’m eager to see what people build with the Decisions API.

Host Vibhu: An interesting point is that now each Dot comes with its own personal computer.

Ari Weinstein: Yes.

Host Vibhu: So they seem able to maintain their work environment more persistently. You’ve been using it for a while—how should people explore its capability boundaries? What directions should they try? I personally often use it to handle customer support issues. For example, “Something broke here, I don’t want to log in or go through authentication,” so you find the relevant information and resolve the issue. Beyond that, what else should people try?

Ari Weinstein: Dot is a fascinating product because each Dot can use its own Linux virtual machine in the cloud, which differs from our other products. Previously, we typically offered cloud browsers or allowed agents to access your own computer; now, you have an entire dedicated Linux computer in the cloud. It can run full desktop applications as well as use a browser. I think computer use is so powerful and exciting precisely because it enables agents to do anything a human can do. All software in the world was originally designed for humans; now that agents can use this software too, you can delegate work to them. So, anything you do on a computer can be done by Dot. Which tasks are most useful really depends on who the end user is and what adds value to their life.

I’d suggest starting by thinking about where you spend your time, then considering whether those tasks can be delegated to an agent. Anything you do on a computer can be done by Dot. Which tasks are most useful depends on who the end user is and what adds value to their life. I recommend first reflecting on how you spend your time, then seeing if those activities can be handed off to an agent.

Host Swyx: Right, for example, booking flights, shopping, honestly, even playing games—those kinds of things are possible, right?

Ari Weinstein: Absolutely. Recently, to eat healthier, I subscribed to a meal delivery service. I love this service because it allows me to customize every meal in great detail, specifying things like “how many grams of chicken, how many grams of rice.” But the process was incredibly complex—it took me two hours to place my next order. Then I realized I could let computer use handle the ordering for me, and it finished in fifteen minutes. Using GPT-6.1 Sol, it completed the task eight times faster than I did, saving me two hours. I think this type of task really highlights its value.

Host Swyx: As a creator, I can immediately name my top use case: automating operations on YouTube. Many YouTube features don’t expose APIs, so you have to put it in a VM and let the agent run it itself. For example, A/B testing or posting community updates—these lack APIs because they hate developers.

Ari Weinstein: I’ve heard developer experience teams say the same thing. They frequently use it for YouTube-related tasks, and it works really well.

Computer Use Enters a New Phase: Better Interface Understanding and Faster Troubleshooting

Host Swyx: I want to ask a slightly sharper question. We have friends who run a highly influential AI podcast, and they have a famous take: computer use has made no progress in the past two years. That’s an interesting claim, and you are arguably one of the most qualified people in the world to discuss this topic. What exactly has progressed?

Ari Weinstein: I remember them saying that a few months ago. Hopefully, they’ve changed their minds by now, because computer use has undergone a 180-degree transformation compared to before.

Host Swyx: You’ve spent nearly your entire career working on some form of computer automation, from building Shortcuts at Apple, to Sky, and then joining OpenAI. Can you walk us through the common thread across this journey? What drives you? What was impossible before, and what milestones have been reached?

Host Vibhu: Let me add another question: From last week’s Codex computer use capabilities to today, what is the primary change? Is it the model, Dots, or the agent harness? Beyond the historical context, we’d like you to clarify exactly what today’s release changes.

Ari Weinstein: I’ve always been passionate about automation—about helping people automate tasks—because it saves time in life, allowing people to focus on what matters more to them rather than meticulously operating a computer. This is why we built those products. My experience working at Apple, founding Sky, and eventually joining OpenAI has been exciting.

Host Swyx: It feels like you were constantly trying to work around Apple’s restrictions until Apple said, “Okay, let’s just hire you so you can do these things from the inside.” Is that right?

Ari Weinstein: Working there was indeed a great experience. Looking back at Sky, it’s interesting that we were also doing computer use then, but the models were much weaker. In just the past year, models’ computer use capabilities have become incredibly strong. The biggest shift I’ve seen is that previously, they could reliably start a task but would encounter issues midway; now, they are very good at debugging errors, retrying, and assessing which approaches work and which don’t.

The field of computer use itself is evolving, and we’re adopting more techniques. Computer use often involves writing code now. If you manually expand tool calls in Codex, you’ll see that it doesn’t just execute one action at a time; instead, it writes JavaScript code for the computer to execute, sometimes completing many actions in a single go, improving both speed and capability. We’re increasingly using multimodal interaction methods like accessibility interfaces. Models might use screenshots, accessibility APIs, or Playwright. They can choose from many different mechanisms depending on the task at hand. The acceleration of the models themselves has also been staggering.

As for what’s different today, I think we’ve been continuously improving computer use, so looking at a single day’s change might not be as meaningful as looking at the past month or two. However, I believe the computer use capabilities within Dot, along with the new models released today, are very promising.

Host Vibhu: In the keynote, Tejal mentioned that computer use speed improved sevenfold, with significant gains on several benchmarks. How do you measure this? As you noted, computer use is constantly progressing. Do these improvements come from the agent harness, the model, or post-training? What changes does the new model bring?

Ari Weinstein: We actually have many ways to measure this, some of which test different combinations and configurations of the agent harness. The situation is somewhat complex because production products have more safety checks and adopt different configurations based on the current task’s needs. So, while there are many measurement methods, regardless of the approach, we see fairly consistent improvements. These gains sometimes come from the harness, and sometimes from the model. One result impressed me deeply: Compared to Astra, GPT-6.1’s cost-efficiency improvement in computer use even exceeds the reduction in its base cost. Seeing that was truly great.

Host Swyx: There was a chart in the livestream that I particularly liked, showing how you’re continuously improving the Pareto frontier of this curve. You also talked quite a bit about how improvements happen in tandem with the harness. Can you give a few examples where you had an "aha" moment? Whether the model drove harness improvements or vice versa.

Ari Weinstein: Introducing more modalities has definitely been effective. Specifically, in the past, many computer use products spent a lot of time scrolling pages: taking a screenshot, trying an action, realizing "I need to scroll down to see the next page's results," taking another screenshot, trying again, and scrolling further. Now, leveraging accessibility APIs and direct access to the Document Object Model (DOM), language models can see the entire page or application, and write code that completes multiple steps at once. These are probably the most obvious breakthroughs. Additionally, there are many smaller discoveries that aren't as flashy, but we’ve found that many speed improvements come from tackling tiny issues one by one through deep inspection.

Host Swyx: Behind this lies a massive amount of hard engineering work. For example, app snapshots—many people still don’t fully understand the difference. When you capture an app snapshot, Codex displays a nice-looking image, but they may not realize that you can actually interact with every button within it, and all text is presented to the model in a highly suitable format.

Ari Weinstein: Exactly. It’s quite interesting. If you want to dig deeper, open Codex and press the Command key twice to capture an app snapshot. This brings the content of the app you’re using into the Codex or ChatGPT chat. Then click on the attachment, and click the small button in the top right corner to see the raw text and the raw accessibility structure representation. We’ve invested a lot of work into this.

Host Swyx: So it’s exporting everything.

Ari Weinstein: Not just exporting, but doing so efficiently and saving tokens—there are many tricks involved. Technologies originally invented for people with accessibility needs who use screen readers help them navigate, and now they also help models understand interfaces.

Ari Weinstein: Exactly. For example, if you take a screenshot of a webpage containing links, the screenshot won't include the link destinations. Or if you screenshot a calendar, event titles might be truncated. But app snapshots provide the full context to the language model, enabling it to do much more.

Host Vibhu: Regarding computer-use agents, I want to ask a broader question. What you described earlier—taking screenshots, scrolling pages, and taking more screenshots—is the old state of affairs. Today, they can already automate many tasks. Where is the bottleneck? Is it the model or the runtime framework? How do you think this will evolve in two years? Can it run continuously for hours? How do we get there? What are your predictions for the future of computer use?

Ari Weinstein: I think one incredible thing the team has achieved over the past few months is that computer use now often completes tasks faster than an average human. The next breakthrough is making its performance truly surpass humans: using software at speeds equal to or exceeding those of skilled power users like us. That will have a significant impact because we can build products with much more real-time experiences. At the same time, the barrier to entry—or the friction required to start using it—will drop. For certain tasks we’re accustomed to doing manually, people will begin defaulting to agents, saving substantial amounts of time.

To achieve this, there are still many small issues and bottlenecks to solve. Some lie on the model side, some on the inference side, some within the agent runtime framework, and others involve how information is represented. As computer use gets faster, we increasingly become limited by the speed of the operations themselves. For instance, in benchmarks for computer-use tasks, a non-negligible portion of the time spent is waiting for websites to load. Suppose you’re automating a task on doordash.com; a large chunk of that time is actually just waiting for doordash.com to load itself.

Host Swyx: So you write a wait command, then execute that wait.

Ari Weinstein: Once the page finally loads, the latency between that moment and triggering the LLM to execute the next action needs to be as short as possible. This is crucial and inherently involves a set of statistical methods.

Host Swyx: Maybe you could do it in an event-driven way?

Ari Weinstein: If conditions allow, we certainly prefer an event-driven approach.

Host Swyx: JavaScript has some loading events.

Ari Weinstein: JavaScript, or rather the browser, has loading events for web navigation, but there are other types of events that genuinely cannot be handled via event-driven mechanisms. So it’s quite complex.

Host Vibhu: One example that comes to mind is chatting with customer support. They might reply in thirty seconds, or it might take three minutes.

Host Swyx: I’ve already used Codex to interact with many bots, and it works well. But sometimes I wonder if the other side knows they’re talking to a bot? Because my sentences are too complete, with perfect capitalization, and I provide full reference numbers and such details—it performs too well. But I don’t care; I just want to resolve my customer service issue.

Host Vibhu: In the prompt, I tell it: “Don’t act like a robot; act like someone who is already impatient.” I also tell it: “While waiting for a reply, invoke a sub-agent to research if there’s a better way to find the information we need.” It’s about adding a bit of human intervention.

Ari Weinstein: Great. And I think, most likely, the other side is also a bot, so now we have bots talking to each other.

Host Swyx: Yes. We’ve reached a milestone in computer use: three or four years ago, we were afraid to connect LLMs to the internet or our own devices. Now, I’ve let it configure my Domain Name System (DNS) and pay my bills for me. It’s literally tens of thousands of dollars, and I just hand it over to computer use, letting it go ahead, thinking, "What’s the worst that could happen?" You guys should also use your own product.

Now that this capability has been released via API, are there any pitfalls or advice you’d like to share with developers? Because they’re about to experience these things firsthand.

Ari Weinstein: First, I’m really glad we brought computer use into the Agents API. Many developer applications want to interact with third-party websites and services, and computer use offers this generality—it can operate anything. Now, developers can immediately use the same computer-use implementation we use ourselves to build products. If you want to build your own computer-use agent runtime framework, there’s certainly plenty of work to be done there, but it’s difficult. Moreover, our models are trained on our own computer-use framework, so using this framework—which lies within the model’s training distribution—may offer advantages in speed, cost, and accuracy.

Responding to what you said earlier, I believe everyone is still gradually adapting to this technology and building trust in it, though some of us may be moving faster than much of the world. Therefore, we have a responsibility to build this trust over time: ensuring reliability, setting appropriate safety checks, seeking user consent before actions with significant consequences like payments; or, depending on the specific application, ensuring it only accesses the websites and apps strictly necessary to complete the task. These are all worth serious consideration. I strongly encourage everyone to try the new Agents API, build interesting things with it, and share your feedback after using it.

Host Vibhu: Have you seen changes it brings to development workflows? For example, Dots can now be used in Slack, and some people use voice to build things. Roman’s demo showed having it modify an app while sending screenshots throughout the process. How are people actually integrating computer use into their programming workflows? Are there best practices worth adopting?

Ari Weinstein: One of my favorite uses, which we frequently see in practice, is having agents use computer use to actually test the software they develop. The impact is greater than it sounds. Previously, when you asked Codex to build something, you had to test it yourself afterward, effectively becoming the QA person for the agent. With computer use, the entire software development lifecycle can be closed-loop: agents can both develop and test software. I love building things this way, letting the agent test itself, so that by the time it’s handed to me, the software already works correctly.

For me, there’s an extra layer of fun because sometimes I’m developing computer use itself. So you end up with a computer-use agent operating another computer-use agent I developed, which in turn operates other things. I think this is a very valuable class of use cases.

Host Swyx: I built a visual interaction testing skill that uncovers many design issues you can’t spot just by reading the code. It’s also perfect for cloning apps: if you’re stuck using a terrible SaaS product and want to replace it, you can replicate it screen by screen. Computer use agents can navigate through the entire application, take screenshots, log actions, and then have Codex recreate everything. Anyway, thank you all for these advancements. Our time is up, but this definitely won’t be our last conversation.

Ari Weinstein: This was a great chat. Thanks for having me.

No More Waiting on Tool Calls: How the OpenAI API Enables Concurrent Execution and Reasoning

Host Vibhu: Nikunj, glad to have you here. You’ve shipped a lot of updates on the API side. As we discussed with Ari earlier, developers can now build using computer-use agents. Which API changes would you like to highlight? Also, please briefly introduce yourself and your area of responsibility.

Nikunj Handa: I’m Nikunj, and I lead product work for the API team. I’ve been here about three years, consistently involved in model releases. It feels like during my time at OpenAI, this process has never stopped. Every time we launch a new model, we work closely with the post-training and research teams to understand its new capabilities and then expose them via the API.

The key new capability in the GPT-6 release is asynchronous function calling. Many tool calls in products like Codex and Dots are time-consuming. Now, when a tool is running, the model doesn’t need to pause execution. You can initiate a tool call, let the model continue running and reasoning, and check back later for the result. We also introduced mid-turn steering, which allows injecting messages into the model’s reasoning process. This way, as soon as a tool call finishes, relevant instructions can be inserted immediately.

Host Swyx: That’s partly a model alignment capability, right? The ability itself has to be trained into the model.

Host Vibhu: I feel like this existed in applications before: you could steer the model mid-reasoning. It wasn’t optimal previously, but it’s much better now. I’m eager to see how this version performs.

Nikunj Handa: Yes. Our main goal on the API side is to wait until these capabilities are properly trained within agentic runtime frameworks before exposing them via the API, so we ensure they are robust enough. Many capabilities are powered by WebSockets, which we launched a few months ago. WebSockets enable bidirectional communication between applications and the model. I’m not referring to GPT Live here, but GPT-6. Asynchronous tool calling, asynchronous reasoning, and message injection—all of these are possible. Developing this API has been fascinating; I really enjoyed the process.

Host Swyx: This is why we’re an engineering podcast, because we talk about WebSockets. It pairs well with UltraFast, right? I recall this being the first time UltraFast was available via the API. Essentially, frontier-level models can run at theoretically maximum speed.

Nikunj Handa: Before diving into the API, I think the most interesting part of UltraFast is seeing the inference team continuously deliver results around Astra. They’ve been running Codex agents, trying to squeeze out every bit of performance. For several months, a huge amount of effort focused on improving efficiency and reducing costs, allowing us to lower Luna’s price by approximately 80%. Much of this was driven by the inference optimizations they implemented. Later, they shifted focus to making it run as fast as possible. Seeing UltraFast accelerate a model like Astra to this extent is truly exciting.

We initially launched WebSockets for GPT-5.3 Codex Spark. Obviously, tool calling is essential, and WebSockets significantly reduce the overhead of communicating back and forth with tools, making them particularly helpful.

Host Swyx: It’s always fun looking at the usage quota interface: on one side is my regular quota reset, and on the other is that Spark quota I’ve never used. But whenever I want to use it, it’s there.

Nikunj Handa: I think it finally got removed.

Host Swyx: Indeed, it’s gone. You’re slowly phasing out those legacy features.

Inspired by Jev? Beyond Structured Output: How OpenAI’s Decisions API Pursues Faster Decision-Making

Host Swyx: GPT-5.3 Spark explicitly stated it used Cerebras. You neither confirm nor deny whether UltraFast is related to Cerebras, but people are definitely discussing it, caring about it, and curious. Plus, you have your own chips. Another unavoidable topic is decision models, specifically the Decisions API. We were the first podcast to have Diogo deep-dive into Jev, and I also invited him onto AI Engineer. After seeing Jev, how quickly did you decide to follow suit?

Nikunj Handa: First, credit to Diogo and the Jev team—they inspired an entire market segment. When Jev came out, everyone was incredibly excited. Users kept reaching out to us, and internal teams said, “We need a much faster classification system.” There are some upcoming Dot features I don’t want to reveal too much about yet, but you’ll see some features built on the Decisions API that react extremely quickly. Many people at OpenAI were instantly captivated by this technical challenge and started thinking, “How do we build this? We aren’t going to retrain a model from scratch, but…”

Host Swyx: Four weeks ago, this wasn’t even on the development roadmap, right?

Nikunj Handa: Not at all. This was entirely inspired by Jev.

Host Swyx: I think you might be the first lab to replicate and adopt this approach.

Nikunj Handa: I think OpenAI has a strong hacker culture. People get excited when they see something interesting. A colleague from the inference team and a very skilled engineer from the infrastructure team said, “Let’s try building it.” They created a prototype and found it actually worked. Now we’re continuously optimizing latency, trying to make it as fast as possible, hoping to ship it in the next few days. Once we hit our latency targets, we’ll aim to release it.

Host Vibhu: It's interesting that on one hand, there's this hacker culture, and on the other, as Sam mentioned, 99%... you are also one of the most reliable APIs, likely with the highest usage volume, which is directly managed by your team. How should people view the Decisions API? Many have seen Jev or heard discussions about it but haven't actually built with it yet; you're bringing it to a broader audience. What should they consider it to be, and how should they use it?

Nikunj Handa: The primary use case we see is fast classification. Computer-use demos are impressive, but different models have their own sweet spots: having Astra write JavaScript scripts to control a computer requires different capabilities than having Luna select an action at each step. However, for certain computer-use tasks, this is sufficient, and I'm eager to see results in this area. Internally, I've seen an interesting prototype combining the Decisions API with GPT Live. GPT Live is our bidirectional real-time voice API, using an architecture where a frontend model collaborates with a backend model. GPT Live handles rapid conversation and task dispatching, while the backend can utilize models like Astra. Previously, tool calling in GPT Live felt somewhat slow. Now, some demos show GPT Live controlling computers via tool calls, making the entire experience much more responsive and natural. I look forward to seeing how people combine Live with Luna on the Decisions API; it should yield some very interesting results.

Host Swyx: I want to clarify what's happening here, especially from a product perspective, because recently many people have been creating clones of Jev—probably around a hundred in the last two weeks.

Host Vibhu: Quite a few appeared right after those first few days.

Host Swyx: People can replicate a Jev API, honestly, it's just structured output, and OpenAI was the first to do this. So I want to make clear what a decision model is and what truly matters. It's not just latency, nor just structured output, right? If the decision model has the same price as Luna, can I just use Luna, turn off reasoning, add structured output, and effectively have Jev? No, that's the real distinction.

Host Vibhu: There's also the issue of confidence scores.

Nikunj Handa: We didn't train a new model for this; we built it entirely on the existing Luna weights, so it really is Luna. On top of that, we constrain the outputs—structured output is a significant part—and specifically optimize the inference system to minimize Time-to-First-Decision (TTFD). Since you might have multiple questions, we essentially run them in parallel.

Host Swyx: As a batch.

Nikunj Handa: Yes, batch processing. People are experimenting with various inference techniques to make it as fast as possible. But at least our initial implementation, the first version, uses a zero-shot approach directly on Luna to see how it performs. Of course, we wanted to ship it first. This is OpenAI's typical iterative deployment style: release it, see what people think, and then improve the model as needed. That is the Decisions API.

Host Swyx: You also have a clear advantage: vision capabilities. They don't have vision, right?

Nikunj Handa: By using Luna, we inherently have that capability.

Host Swyx: It sounds like if it's still the same Luna weights, there are innovations coming later, such as confidence scores. We've discussed calibration on the podcast, including benchmarks for calibration. The key point is that Reinforcement Learning from Human Feedback (RLHF) tends to bias model outputs toward what you want to hear. But that doesn't necessarily reflect its actual level of certainty.

Nikunj Handa: Completely agree. I'm also keen to see how the actual results turn out. Perhaps confidence and calibration will become key areas for improvement in future model releases.

Host Swyx: There's also debate regarding the architecture. Jev hasn't disclosed specific methods, so no one knows the answer. Currently, there are two main speculations: one suggests it might use diffusion models rather than autoregressive ones. However, you achieved parallel generation with your own method. The other speculation relates to mechanistic interpretability: perhaps it analyzes model activations and then directly outputs weights. You've done research in this area too, so these guesses exist.

Host Vibhu: Demos exist for both approaches. I recall Gemini sharing Gemini Diffusion, or Gemma Diffusion, used to generate Jev-style outputs. Interpretability researchers have also tried extracting information from intermediate layers of models. However, these are all just speculations.

Host Swyx: The key is, what goal are you trying to achieve? Because creating a similar API format isn't hard; the difficulty lies in subsequent capabilities like speed, accuracy, confidence calibration, and other aspects of calibration.

Nikunj Handa: The entire field is being energized now. People will build many interesting things and learn from each other, which excites me.

From API to Agentic Products: Performance, Caching, and Context Management

Host Vibhu: As a platform team, a large part of your job is helping developers build things. What do you think people should build with the Decisions API and computer-use agents? Are there things you're working on internally that only became feasible after this shift?

Nikunj Handa: The internal use cases for the Decisions API are quite clear. For instance, user operations teams immediately started using it, saying, "We need to classify all customer support tickets." Additionally, some GPT Live demos were excellent. I imagine the Codex application team might also experiment with it. However, the project started about a week ago, so it's still in a very early stage.

Host Swyx: Put it all in a unified, Google Docs-like environment?

Nikunj Handa: These are all built on the Agents API. We just launched it, and I’m eager to see what people build with it.

Host Vibhu: You did a great job showcasing the editing space, pages, collaboration, and integrating your own Dot. There’s a lot of content there, and everyone can draw plenty of inspiration from it.

Nikunj Handa: Yes, Astra makes all of this possible. Things are moving incredibly fast right now; the speed at which people go from idea to implementation is astonishing.

Host Swyx: Are there specific areas where you’d like focused feedback? Perhaps you’ve pushed things out with several different directions, hoping developers help you decide which path to take.

Nikunj Handa: The Agents API and Decisions API are our latest products. We welcome all kinds of feedback on them to help determine our next steps. The Responses API is our flagship product, and we’re currently very focused on performance, primarily in two areas. First is latency. We’ve been rewriting the entire Responses API tech stack, targeting metrics like Time To First Token (TTFT) and Time Between Tokens (TBT), aiming to minimize latency as much as possible. This remains a key focus of our work.

Second is deep improvements in caching, especially for applications like personal agents, which essentially consist of one continuous session that never ends. We’ve been working hard to improve caching, and we can now guarantee cache hits within thirty minutes. In fact, we recently rolled out a significantly longer cache window for a user, offering a twelve-hour cache hit guarantee.

Host Swyx: Is this a public API?

Nikunj Handa: Not yet; it’s currently in preview. We aim to open it up to everyone as soon as possible. By paying slightly more for cache writes, we can guarantee cache reads over a longer period. For example, if you’re using an agent session, perform some actions, leave, and return three or four hours later to continue, you’ll still benefit from the performance advantages of caching.

Host Vibhu: With the new models, you also reduced costs significantly, right? It’s 25% cheaper.

Nikunj Handa: Cache read prices have decreased.

Host Vibhu: Developers should definitely use it, since the cost is substantially lower.

Nikunj Handa: Exactly. When building applications, consider caching thoroughly. Use our prompt diagnostics or cache diagnostics tools to identify where caching isn’t being effectively utilized. This part is crucial. I also want to mention pre-warming, which we now offer in the API. If you know a certain prompt will be used, you can pre-warm the cache: pay for the cache write upfront so it’s ready for immediate access during the next thirty minutes.

Host Swyx: Then you can create many instances based on that session, continuously tweaking the prompts.

Nikunj Handa: You can keep going, creating many instances. I particularly hope to receive feedback on underlying performance, helping us make the Responses API the most performant and reliable way to develop with LLMs. For the newer products, any kind of feedback is welcome.

Host Swyx: Try it out first, then tell us what you need.

Nikunj Handa: Please help us define the roadmap.

Host Swyx: For me, caching is obviously necessary, but eventually, you’ll hit the one million token context limit, which likely won’t change in the foreseeable future. So good compression methods are still needed. What are the best practices here?

Nikunj Handa: Absolutely correct. First, OpenAI has its proprietary compression, known as context compaction.

Host Swyx: It’s already integrated into the agents.

Host Vibhu: It’s also in the API, specifically the Agents API.

Host Swyx: You decide when to compress for us, right?

Nikunj Handa: Correct. In the Agents API, context compaction is built into the Agent Harness. If you use the Responses API, there are two approaches: one is server-side context compaction, where you tell the Responses API to automatically compress once a certain token threshold is reached, reducing the occupied context. The other is /compact, suitable for those who want full control. You can call /compact anytime, using your own logic to decide when to compress.

Host Swyx: It’s not AGI, but it is a manual override method.

Nikunj Handa: Many large coding agents prefer manual control. If you look at the implementation in the open-source Codex Harness, you’ll see they use /compact to handle this. We’re also researching new context compaction techniques, some of which are already implemented in the Codex Harness, so you can see them in action. We’re experimenting with file-based systems as well, so there’s quite a bit of ongoing work in context compaction.

Host Swyx: Time’s almost up. You covered a lot regarding performance and introduced the new APIs being released. Regarding the platform’s future, are there any other interesting directions you can share?

Nikunj Handa: What we provide now is relatively low-level. Previously at Stripe, a significant part of my work involved building higher-level infrastructure components and products on top of core payment capabilities. I’ve been thinking about how best to do this in AI. We’ve tried a few times, such as launching the Assistants API early on, but it wasn’t quite the right answer. Now we’re moving in the direction of the Agents API, which provides the Codex Harness. But exactly how much flexibility should we give users inside it? That remains an open question.

For example, how should we design Memory Walls and various higher-level API objects to further abstract and encapsulate underlying storage concepts, so that users don’t have to worry about these details? I’m eager to figure out how to design this entire domain. In AI, the current approach often involves providing a low-level API foundation, offering an example runtime framework, and then letting coding agents implement the rest. But how much of this should be built directly into the API is something I’ve been pondering. If anyone has thoughts on this, I’d love to hear them.

Host Swyx: I often think of an analogy, which we’ll use to wrap up: You are building an AI cloud platform—Sam mentioned this a year ago. You’re somewhat reliving the invention process of AWS, needing to determine one by one: this is EC2, that is S3, and so on. But you’re creating AI-native versions of each of these components.

Host Vibhu: There are indeed many parallels, such as pre-warming caches for content known to be used later. It’s great that these capabilities are open to developers, giving everyone more ways to build new things.

Host Swyx: That’s all for today.

Interview video link: https://www.latent.space/p/devday-2026