Bright ideas, clever finds, and the smarter way to do everyday things.

Tomorrow's News

The Tiny AI Computers That Hide Behind Your Monitor

Desktops the size of a sandwich now run serious AI models at home, with none of your files leaving the room.

A hand slides a palm sized silver computer onto a bracket behind a monitor on a clean home office desk.
A hand slides a palm sized silver computer onto a bracket behind a monitor on a clean home office desk.

There is a computer on the desk, but you cannot see it. It is clipped to the back of the monitor, roughly the size of a hardcover book, drawing about as much power as a bright lamp. It is also running an AI model that answers questions about ten years of your own files without sending a single one to the internet.

That combination is new, and it is why these little boxes suddenly have a waiting list.

What actually happened

Nvidia started it by shrinking a data center into a desk toy. Its DGX Spark pairs a GB10 chip with 128 GB of unified memory, enough to load models that used to require a tower full of graphics cards. Demand and a memory crunch pushed the price from $3,999 to $4,699 in February.

AMD came in underneath. Mini PCs built on the Ryzen AI Max chip, known as Strix Halo, put 16 processor cores, a built in graphics chip and up to 128 GB of shared memory in a box you can hold in one hand. AMD’s own developer machine lists at $3,999 with 128 GB, while boxes from other vendors run from roughly $1,600 with 64 GB up to about $3,300 fully loaded.

Apple did not even mean to be in this fight. Mac mini demand ran ahead of the company’s own forecast, driven partly by people buying them for AI work, and on May 1 Apple quietly dropped the $599 base model, moving the starting price to $799 with 512 GB of storage. The MacBook line moved to the M5, M5 Pro and M5 Max chips this spring, and a new Mac Studio slipped toward late this year.

What using one is actually like

The first surprise is how ordinary it is.

You plug the box in, install a free tool like Ollama or LM Studio, and download a model the way you would download a large game. The tool finds the hardware on its own. Then you type a question into a window that looks like every other chat window you have used, and the answer starts appearing.

The second surprise is the silence. These machines sip power compared to a gaming tower, so there is no jet engine spinning up when you ask something hard.

A late night desk where an answer streams onto the screen with the wifi switched off

The third surprise is where the speed goes. Independent benchmarking of the DGX Spark has found it chews through the reading part of a request extremely fast, processing well over a thousand tokens per second when digesting a long document, then slows to roughly 35 to 45 tokens per second while writing the answer back to you on a large model.

That is faster than most people read, which is the honest bar. It is also noticeably slower than a top tier cloud service, and you will feel the difference on long answers.

The last surprise is what you stop doing. No account. No usage meter. No wondering whether the document you just pasted is now training data somewhere.

How it works, in one breath

Big AI models need to hold billions of numbers in memory at once. These machines give the processor and the graphics chip one shared block of memory instead of a small, separate slice for each, so a model that would not fit on a normal PC fits here.

Think of it as trading a fast sports car for a cargo van. The van is not quicker, but it can carry the whole load in one trip.

Here is the mechanism that decides everything, and almost nobody puts it on the box. Once a model fits, the speed you experience is set mostly by memory bandwidth, which is how fast the chip can read those billions of numbers back, over and over, for every word it writes. A Strix Halo machine moves roughly 256 GB per second. Apple’s higher end chips move two to three times that. That gap, not the processor name, is why two machines with the same 128 GB feel different.

Fit first, bandwidth second, brand last.

Why it matters to you

Your data stays yours. Run a model locally and your tax documents, medical letters, client work and family photos never leave the desk, which is the entire reason accountants, lawyers and therapists have started buying these. It is the same instinct behind keeping your home camera footage on hardware you own.

There is no meter running. Once the box is paid for, asking it a thousand questions costs electricity, not tokens.

And it works when the internet does not. Airplane, cabin, hotel wifi that barely loads a webpage, none of it matters to a model already sitting on your drive.

What the boxes cost

MachinePriceWhat it is best at
Mac miniFrom $799The cheapest sane entry, quiet, good for dabbling
Framework Desktop with Ryzen AI MaxAround $1,600 with 64 GBRepairable, well documented, strong value
Strix Halo mini PCs from other vendorsRoughly $2,300 to $3,30096 GB to 128 GB for large open models
AMD Ryzen AI developer platform$3,999 with 128 GBA reference machine with full AMD support
Nvidia DGX Spark$4,699Nvidia’s software stack, fastest at digesting long documents

Who should pick which is simpler than the table looks. If you are curious, buy the cheapest machine with enough memory and spend nothing else. If you work with confidential files daily, the middle of this table is the sweet spot. The top of the table is for people building things, not people asking questions.

Reality check

Prices are moving the wrong way. Memory chip contract prices jumped roughly 90 percent in the first quarter of this year as AI data centers bought up supply, and that is exactly why Nvidia raised the Spark by $700 and Apple cut its cheapest configurations.

A local model is also still a step behind the best cloud models on the hardest tasks. What runs at home is very good, not state of the art.

And setup is nerdier than opening an app. Tools like Ollama and LM Studio have made it close to double click simple, but you will spend an evening reading before it feels normal.

When you can get one

All of it is buyable today. Strix Halo mini PCs from several vendors ship now, the DGX Spark sells through Nvidia partners at $4,699, and a Mac mini starts at $799.

The honest buying advice is to pick by memory, not by brand. Look at 32 GB if you want to dabble, 64 GB for daily work, and 128 GB if you want the big open models.

How this fits into your week by 2028

The likeliest future is that the box disappears into the machine you already own.

Every laptop and phone chip shipping now has a dedicated AI section, memory sizes keep climbing, and the open models that run locally keep getting smaller and sharper at the same time. The separate AI box on your desk in 2026 looks a lot like the separate graphics card of thirty years ago: essential, then absorbed.

What lasts is the habit. Once you have asked your own machine about your own files, sending them somewhere else starts to feel like an odd thing to do by default.

The quiet luxury of 2026 is a computer that answers you without telling anyone what you asked.