September 2026
Running DeepSeek V4.1 Flash on the Lucebox memory hierarchy
DeepSeek V4.1 Flash ships as a 383 GB file. One Lucebox has a 32 GB GPU, 128 GB of unified memory and an SSD. We split the model's experts across all three by how often each one is used, and serve it at up to 26 tok/s when writing code, with a 128K context.
TL;DR
- The biggest DeepSeek on one desktop. DeepSeek V4.1 Flash is a 383 GB file. One Lucebox serves it at up to 26 tok/s when writing code, with a 128K context.
- Each expert lives where its use pays for it. The GPU holds the experts used most, the Strix Halo most of the rest, the SSD the rarely used tail. The model's router is tuned to lean toward the fast memory, and a cache and a look-ahead hide most of the SSD reads.
- One click to install, a minute to start. On a Lucebox, pick DeepSeek V4.1 Flash on the Engine page and press Install; the server is ready in under a minute.
Three tiers, three speeds
A Lucebox has three places to keep weights, and they differ by two orders of magnitude in how fast they can be read. The R9700's GDDR6 moves 640 GB/s but holds 32 GB. The Strix Halo's GPU reads all 128 GB of its unified memory at about 240 GB/s, the 64 GiB reserved for it and the rest alike. The Crucial P310 SSD holds anything but reads at about 4.5 GB/s. DeepSeek V4.1 Flash picks 6 of 384 experts per token in each of its 40 layers, so where each expert lives decides how fast a token is.
--profile ds41-lucebox at a 128K context, and how fast each tier is read. The Strix Halo tier continues past the 64 GiB reserved for its GPU into system memory pinned so it is never swapped out, which the GPU reads at the same speed.The split is not by size but by use. Most tokens route to a small set of experts, so the R9700 and the Strix Halo get the ones that are used most, and the SSD gets the long tail. The result with 4K tokens of prompt:
How we got from under 1 to 26 tok/s
The first version kept the hottest experts on the R9700 and streamed everything else from the SSD. It answered correctly, at under 1 tok/s. Each step below kept the same answers and removed one bottleneck.
- Strix Halo tier. The Strix Halo holds a second expert stack and computes it in parallel with the R9700.
- Ranked placement. Experts ranked by their measured use, the most used on the fastest tier.
- Expert cache. Streamed experts go through a slot cache on the Strix Halo, one batched graph per layer, with the next layer's experts predicted and prefetched.
- Fused graph and DSpark. A decode step runs as one graph across the three owners, and the DSpark drafter proposes tokens that the model verifies in one pass. Every verified token is bit-identical to plain decode.
- Pinned host memory. The Strix Halo tier grows past its reserved 64 GiB into pinned system memory, which moves about 4,000 experts off the SSD.
- Lucebox placement and routing. Placement fitted to this exact box, and a router that prefers experts in fast memory when two choices are close. Both ship with the engine.
Long prompts
The profile runs a 128K context. Prompts are read in 4K-token chunks, and consecutive chunks run layer by layer so each streamed expert is read once per pass rather than once per chunk. Prompt reading holds at about 120 tok/s from 4K to 20K tokens; decode slows as the context grows, because every token attends over more history.
Agent sessions reuse what they already read: with the prefix cache, later turns in coding sessions start answering in seconds instead of close to a minute.
Loading 383 GB
About 120 GB of that file has to be read before the first token: the dense weights and the experts that live in memory. The engine reads them on eight threads, straight from the SSD and around the operating system's file cache, so loading does not push everything else out of memory. The server is ready in under a minute.
--profile ds41-lucebox, page cache dropped before each start.Run it
On a Lucebox, open your dashboard, pick DeepSeek V4.1 Flash on the Engine page and press Install. It downloads the model and its drafter (about 390 GB, so check the free disk space), builds the engine and verifies every file; Start server then has it answering in about a minute.
Anywhere else with the same two chips, the engine runs it with one profile:
luce_server DeepSeek-V4.1-Flash-ROCMFP2S.gguf \
--profile ds41-lucebox \
--draft DeepSeek-V4.1-Flash-DSpark-draft-MXFP4-Q8.gguf The model and its DSpark drafter (DeepSeek's own draft block from the checkpoint, which proposes tokens the model then checks) are on Hugging Face; the placement files ship with the engine. Run it as the Lucebox model service, so the Strix Halo tier can pin its memory.
Which model to pick
A Lucebox runs all three, one at a time, or Qwen next to DeepSeek V4 in the Lucebox Mix. DeepSeek V4.1 Flash is the largest and newest, with a 128K context, at up to 26 tok/s on code. DeepSeek V4 Flash decodes about twice as fast, over 50 tok/s on code at the same 128K context, from under a third of the disk space. Qwen 3.8 27B is the fastest, around 200 tok/s on code, for quick answers and coding agents that call the model often. Pick V4.1 for the largest model the box can run, V4 or Qwen when speed matters more.
What is next
Decode still waits on the SSD for the experts nobody predicted, and prefill runs the 2-bit expert kernels at a small fraction of what the R9700 can compute. Both are where the next speed comes from.
Hardware and setup
| Lucebox | AMD Radeon AI PRO R9700 32 GB + Ryzen AI MAX+ 395 (Strix Halo) with 128 GB unified memory (64 GiB reserved for the GPU), Crucial P310 2 TB NVMe, Ubuntu, ROCm 7.2 |
|---|---|
| Model | DeepSeek V4.1 Flash, DeepSeek-V4.1-Flash-ROCMFP2S.gguf (2-bit experts), DSpark drafter MXFP4/Q8_0 |
| Launch | --profile ds41-lucebox: about 10 GB of experts on the R9700, the Strix Halo tier sized automatically, 128K context, 4K-token prompt chunks |
| Workload | Code writing from short prompts up to about 20K tokens of prompt, 256-token answers, greedy decoding |
Bottom line
No single chip in the Lucebox can hold DeepSeek V4.1 Flash, and none has to. The R9700, the Strix Halo and the SSD each take the experts that match their speed, the router leans toward the fast ones, and the slow tier is read ahead of time. The model is fitted to this machine. The result is up to 26 tok/s from a 383 GB model on one desktop box.
Measured by us on September 28 and 29, 2026, on Lucebox hardware, with the page cache dropped before each server start, and rounded. Decode speeds are the first request after start; the step-by-step chart collects measurements taken on the same box as each step landed. Expert counts vary slightly between starts with the memory the host has free.