10 min read

BEP Research animation reveals a Grok Bot panel, a plus sign, the supplied cream-colored character, an equals sign, and the colorful Dots wordmark and plush figures. — by Ben Pouladian, BEP Research

Agents, personality and ongoing work. Animated editorial collage using supplied artwork.

OpenAI’s DevDay announcements put a sharper question in front of the AI infrastructure market: when useful intelligence becomes cheaper and work keeps running between prompts, how much more of the system do customers consume?

Our read is that OpenAI is flexing its muscles on cost and scale. A cheaper model, a faster premium tier, an always-on agent and reusable cloud environments are different ways to make more work practical. Together, they bring the economics of sustained AI use into clearer view.

For BEP Research, the key measure is the cost of correctly completed work. Lower token prices can expand the market. Better models and serving software can also reduce the resources each task needs. The hardware investment case depends on the balance—and on who captures the value.

We see three questions running through the announcements: how OpenAI delivers these price and speed choices, whether ongoing agents earn their place in a customer’s workflow, and what that combination means for NVIDIA, memory and networking.

How did they make it this cheap?

GPT‑6.1 Sol now has half Opus 5.5’s standard short-context token prices: $2 versus $4 per million input tokens, and $10 versus $20 for output. The price gap is clear. Its economic significance depends on how much useful work each model completes.

Standard token prices per million: GPT-6.1 Sol input $2 and output $10; Opus 5.5 input $4 and output $20. — by Ben Pouladian, BEP Research

BEP Research chart of standard short-context token prices. Token prices alone don’t establish equal task quality.

OpenAI also reports a task comparison that is worth looking at. On AutomationBench, Sol scored 31.7% at $0.19 per task, versus 29.5% at $0.65 for Opus 5.5 with fallbacks, both at medium reasoning effort. That is a higher score at roughly a third of the cost in that setting.

OpenAI AutomationBench score and cost plot. Medium-effort Sol scores 31.7% at $0.19; Opus 5.5 with fallbacks scores 29.5% at $0.65 per task. Other effort levels differ. — by Ben Pouladian, BEP Research

OpenAI’s reported benchmark, not reproduced by BEP. Original plot with BEP title, legend and source notes; fallback limitations appear on the chart.

We haven’t reproduced those results. Other reasoning settings and workloads can produce different rankings. But moving from a price per token to a cost per task gets us closer to the economic question. A cheap answer that needs several retries or substantial correction can be an expensive way to finish the work.

We made this distinction in July’s “Jensen Huang: Most Companies Will Be Built on Harnesses.” Our point was that the machinery around a model—its tools, context and retry logic—helps determine the cost of finished work. Today’s model prices are another input into that larger calculation.

We suspect the way models, serving software and NVIDIA infrastructure work together matters here. The Sol launch doesn’t disclose the hardware or serving changes behind this particular price gap. Co-design is our hypothesis about what to investigate, not an established explanation of OpenAI’s costs.

It also connects to “Free Isn’t Cheap Enough,” our September comparison of delivered inference performance. Serving software, precision and the rack configuration changed the result on the same published dashboard. That was a specific model and workload, but it showed why the system matters alongside the chip.

Better model design can reduce the computation or memory an answer needs. Better serving can use the available hardware more effectively. Pricing also reflects commercial choices. We need evidence that separates those effects before turning a customer price into an explanation of OpenAI’s costs.

Work that keeps moving

Dots makes the ongoing-work question concrete. It’s an always-on agent powered by GPT‑6 Astra, with its own cloud computer, access to connected apps and context that carries across conversations. It can work on several projects while a customer brings it new tasks.

Consider the research workflow: separating consequential AI news from the noise in a feed, preparing email replies for review, and identifying subscriptions worth dropping. These are recurring tasks with context that accumulates over time. Carrying that context forward could change the value of an agent more than a single impressive answer.

Those are uses to evaluate, not results we’re claiming today. The practical test is whether the work an agent brings back saves more time than reviewing and correcting it takes. A running agent becomes valuable when it can carry useful context forward and reliably finish something.

Dots is rolling out to eligible accounts. Today’s starting point is a primary dot; teams of dots are a future direction. Its first agent is included in qualifying paid plans, while deeper work has an allowance and tasks in Work or Codex still use their normal limits.

That distinction matters. An agent being available around the clock doesn’t mean unlimited output. We should measure how much useful work it delivers within a budget, along with the amount of supervision it requires.

Voice, cloud and premium speed

Cloud and voice make sustained work easier to reach. I enjoy talking through work in Codex, although I haven’t used Codex Cloud yet. Explaining a goal and steering the result by voice makes remote work more appealing to me. Our research question is whether that convenience leads to more tasks being started and completed.

Reusable cloud environments, access from a phone, and the refreshed CLI’s voice and agents view fit that direction. Cloud tasks can keep running while a computer is asleep. That makes the place where work runs less dependent on the device a customer is holding.

Astra Ultrafast adds a different choice. OpenAI advertises up to eight times faster token generation in Codex, reaching 300 tokens a second, and up to six times in the API. That isn’t an eightfold improvement in the time to complete every task: tools, waiting and review still matter.

Speed also consumes more usage. The new Pro 500 plan makes OpenAI’s segmentation explicit: customers can pay for a larger allowance and premium speed.

  • Plus — $20/month. Baseline paid plan.

  • Pro 100 — $100/month. Entry Pro tier; no Ultrafast.

  • Pro 200 — $200/month. More use than Pro 100; no Ultrafast.

  • Pro 500 — $500/month. 25× Plus allowance; Astra Ultrafast.

Monthly prices in USD, checked September 29. Sources: Plus pricing, Pro tiers and DevDay recap. Usage limits still apply.

These monthly subscription prices are separate from the API token prices above. Pro 200 has reopened with a lower allowance for new subscribers; eligible existing users keep their previous allowance through October 29. Astra Ultrafast also reaches eligible Enterprise plans; Sol Ultrafast is coming soon.

Our read: cheaper model calls and higher subscription tiers can coexist. The test is whether customers pay more because agents complete enough additional work to justify the budget.

The current official Astra speed pages don’t identify its hardware vendor. An earlier Ultrafast release named Cerebras for a different model. We can’t carry that attribution over to today’s Astra tier, or infer NVIDIA from speculation that it uses something else.

There’s more infrastructure around these models too: computer use in the Agents API, a Decisions API in limited preview for bounded choices, shared Spaces and Pages, and plugins that can extend where workflows live. We see a set of tools for starting, sustaining and inspecting work.

The fight for the customer

The business backdrop makes these product changes more consequential. Reuters reports OpenAI’s annualized revenue run rate approaching $70 billion, up more than 70% since the start of the third quarter, with enterprise sales more than doubling since July. That is a reported pace, not audited annual revenue or profit.

Reuters, citing Bloomberg, also reports early talks to raise at least $30 billion at about a $1.4 trillion pre-money valuation. Those are fundraising targets, not a completed deal. Revenue growth helps explain investor interest, but it doesn’t tell us the cost of supporting that growth.

Meta expanded Muse for small businesses today, adding business tools and account analytics. Meta says Muse is free for most uses, with subscriptions for people who want more. Against Dots’ inclusion in qualifying paid plans, the contrast is free entry versus an existing paid-plan relationship.

A late-afternoon quote snapshot had Meta up versus its prior close, while NVIDIA and Apple were lower. Our interpretation is that the fight over distribution and the customer relationship is becoming more important. The moves themselves don’t prove investors were responding to Muse or DevDay.

For Apple, we’re watching whether people increasingly begin a task with an agent rather than the interface they use today. For Meta, the question is whether free access brings people into a relationship they keep using. We need adoption, retention and paid-conversion evidence to judge those possibilities.

Our September 22 Muse analysis approached this from the workload up: registered agents aren’t the same as active work. Tasks per day, time spent running and the cost of model calls change the hardware bill. Those were explicit model assumptions, not Meta usage telemetry; adoption alone doesn’t settle them.

So What? Our NVIDIA Question

The hardware implications depend on what people actually do with these products. Fast interactive answers and large volumes of background work put different demands on a system. Model design shapes computation, memory and data movement; ongoing agents can add persistent context, caching, browser activity and CPU work alongside inference.

Our view is that sustained agents can burn far more tokens than a single prompt suggests. A coding, research or inbox task can involve planning, tool calls, checks and retries, each adding model calls and carrying accumulated context. Parallel jobs and background work can increase task volume as well. That gives cheaper tokens two routes to a larger bill: more tokens per workflow and more workflows running. We need total model-call cost per successfully completed task, plus the rest of the workflow’s compute. This is a mechanism to test, not BEP usage telemetry or a forecast of accelerator demand.

We’re watching cost per correctly completed task after retries and review, the amount of work running between prompts, and the share that needs premium speed. Those measures connect product adoption to demand for accelerators, memory and networking.

Cheaper useful AI could expand usage dramatically. It could also reduce the hardware needed for each task. An NVIDIA bull case requires evidence about which effect wins, where the work runs and how much of the value NVIDIA captures.

Our view after DevDay: cheaper useful AI opens more work to automation, while speed and persistent agents change how that work runs. The next evidence needs to connect actual usage to the cost and capacity of the system serving it. Subscribe free to follow our research as we track that connection.


Related BEP Research

Disclosure: I am long NVDA, META, AAPL, LITE, CRDO, TSEM, ALAB, WOLF, NOW, SMCI, BE, NBIS, ORCL (2027 LEAPS), and the AI-Science Basket as published. I have been publicly long NVDA since 2016. I advise private companies, none of which are named here. This is investment research, not investment advice. Do your own work.

Posted in

Leave a Reply

Discover more from Ben Pouladian

Subscribe now to keep reading and get access to the full archive.

Continue reading