DevNews

OpenAI's First Device Reported as a $300 Doughnut Speaker

On this page
  1. What the report describes
  2. The case for a category that already failed
  3. The camera is the actual news
  4. What the schedule says
  5. What we would take from this
  6. Sources and further reading

OpenAI has been designing hardware since it bought Jony Ive's io startup, and until this week the shape of the product was guesswork. Bloomberg reported on August 6, 2026 that the first device is a doughnut shaped smart speaker roughly the size of a hockey puck, with no display, a camera and sensors, a battery so it can move between rooms, moving parts intended to give it personality, and a price somewhere between $300 and $400. A full launch is described as 2027. None of this is confirmed by OpenAI, and the distinction between a well sourced report and an announcement is worth keeping.

The short answer

Bloomberg reported on August 6, 2026 that OpenAI's first hardware product is a doughnut shaped smart speaker about the size of a hockey puck. No screen, microphones and speaker grilles, a camera and environmental sensors, a battery so it can be carried between rooms, a premium metal body, and moving parts plus lights to signal that it is listening or responding. Jony Ive is designing it, following OpenAI's acquisition of his io startup. Full launch reported for 2027. OpenAI has confirmed none of it.

$300+reported price, described as falling between three hundred and four hundred dollars
2027reported full launch, with an introduction expected earlier
no displaythe defining choice: voice first, with a camera and sensors
Answer card: Bloomberg reported on August 6, 2026 that OpenAI's first hardware device is a doughnut shaped smart speaker about the size of a hockey puck, with no display, microphones and speaker grilles, a camera and environmental sensors, a battery for portability, moving parts and lights that signal listening and responding, a premium metal body designed by Jony Ive, and a price between three hundred and four hundred dollars, with a full launch in 2027.
What Bloomberg reports about OpenAI's first device, August 6, 2026. None of it confirmed by OpenAI. PNG

The most interesting thing about a rumoured product is usually not the product. It is what the company decided not to build.

OpenAI has been working on hardware since it acquired io, the device startup founded by former Apple design chief Jony Ive, and the speculation since then has run through every category in turn. Bloomberg reported on August 6 that the answer is none of them. The first device is a smart speaker: doughnut shaped, roughly hockey puck sized, with no display at all.

What the report describes

A rounded body in a premium metal, small enough to carry in one hand and battery powered so it moves between rooms. Speaker grilles and microphones for voice conversation with ChatGPT. A camera system and environmental sensors feeding it information about its surroundings.

Two details stand out from the usual specification list. The device is reported to have moving parts intended to give it personality, indicating when it is responding, with lights signalling when it is listening. And the price is put at somewhere between $300 and $400, which places it well above the commodity smart speaker market and roughly alongside a mid range tablet.

Timing is described as an introduction later in 2026 with a full launch in 2027.

The case for a category that already failed

Smart speakers are not a new idea, and the honest summary of the last decade is that they plateaued. Tens of millions were sold, most settled into being kitchen timers and music players, and the assistants inside them stopped improving in ways anyone noticed.

The argument for trying again rests on why that happened. Those assistants worked by intent matching: a recognised phrase mapped to a defined action. That is a good fit for setting a timer and a poor fit for anything else, and users adapted by learning the handful of commands that worked and abandoning the rest. The ceiling was in the design, not the hardware.

A device built around a current language model does not have that ceiling, because open ended conversation is the default rather than the exception.

Comparison card contrasting the previous generation of smart speakers built on intent matching, where a recognised phrase mapped to a defined action and users learned a small working vocabulary, against a device built around a current language model where open ended conversation is the default, and noting that the open question is not model capability but whether a dedicated object beats the phone already in the user's pocket.
The technical objection to the old category has an answer. The product objection does not, yet. PNG

Whether that is sufficient is a separate question, and it is not a technical one. The capability described here is available on a phone that the buyer already owns and already carries. A dedicated object has to earn its place on the table by being better in a way that matters, and always being present, always listening and having a camera is the reported answer. That is a real difference. It is also the part people will argue about.

The camera is the actual news

A speaker with a microphone is a voice interface. A speaker with a camera and environmental sensors is something that maintains a model of the room it occupies, and that is a materially different product.

It is also where the scrutiny will land, and it should. A screenless device with an always available microphone and a camera in a home raises a specific set of questions: what is processed on the device against what is sent elsewhere, what the lights and moving parts actually correspond to in terms of data capture, and how someone who did not buy it knows what it is doing. Those questions have been answered before by other product categories with mixed results. They cannot be answered here yet, because there is no documentation to answer them with.

For anyone who builds systems, there is a related and less discussed angle. A conversational device has a latency budget measured in the pause a person will tolerate before a reply feels wrong, which is short. Meeting it at scale is an infrastructure problem, and the split between what runs locally and what runs in a datacentre is the design decision that determines both the cost per device and how the thing behaves when the network is poor.

What the schedule says

An introduction in 2026 and a launch in 2027 is a long gap, and it is a recognisable one. Companies that have only ever shipped software discover the hardware problems in a predictable sequence: prototypes are not manufacturing, components have to be committed to far ahead of demand, certification consumes calendar time that no amount of effort compresses, and support, returns and warranty are an organisation rather than a feature.

Buying io brought in Ive and a group of former Apple designers, which addresses the design half of the problem convincingly. The manufacturing and operations half is not something a company can hire in a single acquisition, and the gap between introduction and launch is what that looks like on a calendar.

What we would take from this

Treat the specifics as provisional. Prices move, industrial design changes between prototype and tooling, and features that fail thermal or battery testing quietly disappear. What a report at this stage establishes is the category and the direction of the bet, and both are clear enough: screenless, voice first, stationary but portable, aware of its surroundings, priced as a considered purchase rather than an impulse one.

The part worth following is not whether the speaker sells. It is that a company built entirely on software and APIs is taking on inventory, retail logistics and hardware support, which changes what it costs to run and what it can afford to get wrong. That shift is real whether or not the doughnut turns out to be the right shape.

Sources and further reading

Frequently asked questions

Has OpenAI confirmed any of this?

No. This is Bloomberg reporting based on sources, subsequently picked up by outlets including MacRumors and 9to5Mac, and OpenAI has not announced the product or published specifications. The distinction matters more than it usually does, because hardware in development is exactly the category where details move. Prices are set late, industrial design changes between prototype and tooling, and features get cut when they fail thermal or battery testing. What a report like this establishes with reasonable confidence is the category and the general shape of the bet: a screenless, voice first, portable device rather than a phone, a wearable or a pair of glasses. The specific numbers are best read as the current state of an unfinished product rather than as a specification sheet.

Why build a speaker when smart speakers already exist and largely stalled?

Because the argument is that the previous generation failed for a reason that no longer applies. Voice assistants of the last decade were built on intent matching: they recognised a set of phrases and mapped them to actions, which worked for timers and music and fell apart on anything conversational. Users learned the small vocabulary that worked and stopped trying the rest, and the devices settled into being clocks that play podcasts. A device built around a current language model is not constrained that way, since the interaction is open ended by default. Whether that is enough to make people want a dedicated object, rather than the same capability on a phone they already carry, is the actual question, and it is not answered by the model being better. It is answered by whether there is a reason to reach for the thing on the table instead of the thing in your pocket.

What is the significance of the camera and sensors?

It is the part of the description that changes what the product is. A microphone and a speaker make a voice interface. A camera and environmental sensors make a device that has some model of the room it sits in, which is a different proposition and the reason the reported design is more interesting than the doughnut shape. It is also the part that will attract the most scrutiny, and reasonably so. A screenless device with an always available microphone and a camera in a living space raises questions about what is processed locally, what leaves the device, what the indicators actually indicate, and how a person who did not buy it knows it is there. Those are answerable questions and other product categories have answered them with varying success. They are not answered yet here, because there is no product documentation to answer them.

What does the 2027 timing tell us?

That the hardest parts are still ahead. Reports describe an introduction later in 2026 with a full launch in 2027, which is a long gap and a familiar one for a company shipping its first physical product. Software companies moving into hardware discover the same set of problems in roughly the same order: manufacturing at volume is a different discipline from building prototypes, supply chains have to be secured well in advance, certification takes calendar time that cannot be compressed, and support and returns are an operational capability rather than a feature. Hiring Jony Ive and a team of former Apple designers addresses the design half of that. The manufacturing and operations half is learned rather than hired, and the schedule reflects it.

Does any of this matter if I am not going to buy one?

The device probably does not. The direction might. If a serious attempt is made to put a language model into a dedicated always present object, the interesting consequences are downstream of the hardware: what gets processed on the device against in a datacentre, what the latency budget for a conversational response turns out to be in practice, and what that implies for the infrastructure behind it. Those questions land on people who build and run systems whether or not the product succeeds. There is also a simpler reason to watch: a company that has been purely a software and API business is committing to physical inventory, retail logistics and hardware support, which changes its cost structure and its risk profile regardless of how the speaker sells.