Nextorage aiDAPTIV Station: run a 120B AI model on a 32GB laptop

By: Anton Kratiuk | today, 13:22

Running a large language model (LLM) locally today usually means owning a workstation with a high-end GPU. Nextorage, a Japanese storage company and Phison subsidiary, wants to change that with an external device that plugs into a laptop's USB-C port. The aiDAPTIV Station, announced July 28, 2026, promises to run 120-billion-parameter models — think OpenAI GPT-OSS-120B scale — on a standard notebook with just 32GB of RAM.

How it works

The device is built around Phison aiDAPTIV+ memory offload technology. Instead of cramming an entire model into expensive DRAM or GPU VRAM, a middleware layer dynamically shifts inactive data between system RAM, GPU memory, and a high-endurance NVMe SSD cache. That SSD uses SLC NAND rated at 100 DWPD — a write endurance spec far above consumer drives — to handle the constant data churn that AI inference demands.

There is a real bottleneck to acknowledge: USB-C bandwidth is narrower than a PCIe slot on a desktop motherboard, so token generation will be slower than a dedicated GPU rig. For developers, researchers, or small businesses that want private, on-premises inference without cloud subscriptions, that trade-off may be acceptable.

Nextorage is bundling the hardware with an "AI Standard Package" — a ready-to-run software stack including an LLM execution engine, domain-specific sample data, and knowledge tools. That matters because manually quantizing and segmenting a 120B model is not a weekend project. The station supports both Windows and Ubuntu Linux.

The price and availability question

A launch in Japan is targeted for end of 2026, per the Nextorage official announcement. No UK or US release date, retail price, or distribution partner has been named. Nextorage says the price will be lower than professional GPUs with 48GB or more of VRAM — the RTX 5090 currently sells for $3,600–5,000 on the street — but that still leaves a wide range of possibilities.

The data-residency angle is worth noting. Running inference locally avoids sending sensitive queries to cloud servers, which matters for sectors under UK ICO or US data-protection scrutiny. Tom's Hardware CES 2026 coverage flagged 10x faster inference claims from Phison's demos, though independent benchmarks have yet to appear.

Worth watching

The aiDAPTIV Station is a prototype with a single confirmed market and no public pricing. Real-world throughput tests — especially over USB-C versus PCIe — will determine whether it can deliver usable inference speeds for 120B models or remains more of a proof of concept. If it hits its targets, it could offer a credible alternative to GPU scarcity and cloud lock-in for anyone running AI workloads on a budget.