September 2026

By Davide Ciffa

Running DeepSeek V4.1 Flash on the Lucebox memory hierarchy

DeepSeek V4.1 Flash ships as a 383 GB file. One Lucebox has a 32 GB GPU, 128 GB of unified memory and an SSD. We split the model's experts across all three by how often each one is used, and serve it at up to 26 tok/s when writing code, with a 128K context.

The black Lucebox tower on a wooden deck under a starry sky, with a plush DeepSeek whale next to it
The R9700 holds the hottest experts, the Strix Halo most of the rest, the SSD the long tail.
up to 26 tok/swriting code, 128K context, with speculative decoding
almost 30xfaster than keeping only the R9700 and streaming every other expert from the SSD (under 1 tok/s)
<1 minfrom start to a ready server, with about 120 GB of weights read from the SSD

TL;DR

Three tiers, three speeds

A Lucebox has three places to keep weights, and they differ by two orders of magnitude in how fast they can be read. The R9700's GDDR6 moves 640 GB/s but holds 32 GB. The Strix Halo's GPU reads all 128 GB of its unified memory at about 240 GB/s, the 64 GiB reserved for it and the rest alike. The Crucial P310 SSD holds anything but reads at about 4.5 GB/s. DeepSeek V4.1 Flash picks 6 of 384 experts per token in each of its 40 layers, so where each expert lives decides how fast a token is.

Where DeepSeek V4.1 Flash's routed experts live on one Lucebox R9700 VRAM holds about 1,000 experts (10 GiB) and reads at 640 GB/s; the Strix Halo holds about 9,400 experts (97 GiB) in its reserved and pinned memory and reads at about 240 GB/s; the SSD holds about 5,000 experts (51 GiB) and reads at about 4.5 GB/s. reads at R9700 VRAM ~10 GiB, ~1,000 experts 640 GB/s Strix Halo ~97 GiB, ~9,400 experts ~240 GB/s SSD (P310) ~51 GiB, ~5,000 experts ~4.5 GB/s
Where the 15,360 routed experts live with --profile ds41-lucebox at a 128K context, and how fast each tier is read. The Strix Halo tier continues past the 64 GiB reserved for its GPU into system memory pinned so it is never swapped out, which the GPU reads at the same speed.

The split is not by size but by use. Most tokens route to a small set of experts, so the R9700 and the Strix Halo get the ones that are used most, and the SSD gets the long tail. The result with 4K tokens of prompt:

Share of experts each tier holds against share of routed calls it serves With 4K tokens of prompt: the R9700 holds 6% of the experts and serves 24% of the calls, the Strix Halo holds 61% and serves 68%, the SSD holds 32% and serves 9%. Routed calls served Experts held R9700 24% of calls 6% of experts Strix Halo 68% of calls 61% of experts SSD 9% of calls 32% of experts
Share of routed expert calls each tier served with 4K tokens of prompt, against the share of experts it holds. The R9700 holds 6% of the experts and serves almost a quarter of the calls; the SSD holds a third and serves 9%.

How we got from under 1 to 26 tok/s

The first version kept the hottest experts on the R9700 and streamed everything else from the SSD. It answered correctly, at under 1 tok/s. Each step below kept the same answers and removed one bottleneck.

Decode speed of DeepSeek V4.1 Flash on one Lucebox, step by step Decode speed writing code, first request after start: R9700 + SSD under 1 tok/s, Strix Halo tier about 3, ranked placement about 4, expert cache about 9, fused + DSpark about 11, pinned host memory about 23, Lucebox placement about 26 tok/s. R9700 + SSD <1 tok/s + Strix Halo tier ~3 tok/s + ranked placement ~4 tok/s + expert cache ~9 tok/s + fused + DSpark ~11 tok/s + pinned memory ~23 tok/s + Lucebox placement ~26 tok/s
Decode speed writing code, first request after start, each step measured on the same Lucebox when it landed. The last bar is today's engine at the 128K default context.

Long prompts

The profile runs a 128K context. Prompts are read in 4K-token chunks, and consecutive chunks run layer by layer so each streamed expert is read once per pass rather than once per chunk. Prompt reading holds at about 120 tok/s from 4K to 20K tokens; decode slows as the context grows, because every token attends over more history.

Decode speed by prompt length at the 128K default context Short prompts about 26 tok/s, 4K tokens of prompt about 17 tok/s, 20K tokens about 14 tok/s; prompt reading about 120 tok/s at 4K and 20K. prefill short ~26 tok/s - 4K tokens ~17 tok/s ~120 tok/s ~20K tokens ~14 tok/s ~120 tok/s
Decode speed by prompt length at the 128K default context, with prompt reading speed on the right. DSpark drafter, first request after start.

Agent sessions reuse what they already read: with the prefix cache, later turns in coding sessions start answering in seconds instead of close to a minute.

Loading 383 GB

About 120 GB of that file has to be read before the first token: the dense weights and the experts that live in memory. The engine reads them on eight threads, straight from the SSD and around the operating system's file cache, so loading does not push everything else out of memory. The server is ready in under a minute.

Time from start to ready, DeepSeek V4.1 Flash on one Lucebox Page cache dropped before each start. Before: about 5 minutes. With threaded direct reads: under a minute. faster before ~5 min now <1 min ~6x
Time from start to ready with --profile ds41-lucebox, page cache dropped before each start.

Run it

On a Lucebox, open your dashboard, pick DeepSeek V4.1 Flash on the Engine page and press Install. It downloads the model and its drafter (about 390 GB, so check the free disk space), builds the engine and verifies every file; Start server then has it answering in about a minute.

Anywhere else with the same two chips, the engine runs it with one profile:

luce_server DeepSeek-V4.1-Flash-ROCMFP2S.gguf \
  --profile ds41-lucebox \
  --draft DeepSeek-V4.1-Flash-DSpark-draft-MXFP4-Q8.gguf

The model and its DSpark drafter (DeepSeek's own draft block from the checkpoint, which proposes tokens the model then checks) are on Hugging Face; the placement files ship with the engine. Run it as the Lucebox model service, so the Strix Halo tier can pin its memory.

Which model to pick

A Lucebox runs all three, one at a time, or Qwen next to DeepSeek V4 in the Lucebox Mix. DeepSeek V4.1 Flash is the largest and newest, with a 128K context, at up to 26 tok/s on code. DeepSeek V4 Flash decodes about twice as fast, over 50 tok/s on code at the same 128K context, from under a third of the disk space. Qwen 3.8 27B is the fastest, around 200 tok/s on code, for quick answers and coding agents that call the model often. Pick V4.1 for the largest model the box can run, V4 or Qwen when speed matters more.

What is next

Decode still waits on the SSD for the experts nobody predicted, and prefill runs the 2-bit expert kernels at a small fraction of what the R9700 can compute. Both are where the next speed comes from.

Hardware and setup

LuceboxAMD Radeon AI PRO R9700 32 GB + Ryzen AI MAX+ 395 (Strix Halo) with 128 GB unified memory (64 GiB reserved for the GPU), Crucial P310 2 TB NVMe, Ubuntu, ROCm 7.2
ModelDeepSeek V4.1 Flash, DeepSeek-V4.1-Flash-ROCMFP2S.gguf (2-bit experts), DSpark drafter MXFP4/Q8_0
Launch--profile ds41-lucebox: about 10 GB of experts on the R9700, the Strix Halo tier sized automatically, 128K context, 4K-token prompt chunks
WorkloadCode writing from short prompts up to about 20K tokens of prompt, 256-token answers, greedy decoding

Bottom line

No single chip in the Lucebox can hold DeepSeek V4.1 Flash, and none has to. The R9700, the Strix Halo and the SSD each take the experts that match their speed, the router leans toward the fast ones, and the slow tier is read ahead of time. The model is fitted to this machine. The result is up to 26 tok/s from a 383 GB model on one desktop box.


Measured by us on September 28 and 29, 2026, on Lucebox hardware, with the page cache dropped before each server start, and rounded. Decode speeds are the first request after start; the step-by-step chart collects measurements taken on the same box as each step landed. Expert counts vary slightly between starts with the memory the host has free.

Related

Run DeepSeek V4.1 on your Lucebox

Install it from the Engine page of your dashboard. The engine is open source.

GitHub DeepSeek V4 post Discord