<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>lucebox engineering blog</title><description>Engineering notes on local inference, heterogeneous computing, GPU kernels, speculative decoding, and hands-on benchmarks.</description><link>https://www.lucebox.com/</link><language>en-us</language><item><title>Ling 3.0 Flash on DGX Spark: Up to 141.9 tok/s with Adaptive DSpark and FlashKDA</title><link>https://www.lucebox.com/blog/ling3-flash-dgx-spark/</link><guid isPermaLink="true">https://www.lucebox.com/blog/ling3-flash-dgx-spark/</guid><description>Our Ling 3.0 Flash result on one DGX Spark: up to 36.4% faster prompt reading, near-tied ordinary generation, and 141.9 tok/s in the best matched DSpark case.</description><pubDate>Fri, 28 Aug 2026 12:00:00 GMT</pubDate></item><item><title>Qwen3.8-27B on the AMD R9700: up to 227 tok/s</title><link>https://www.lucebox.com/blog/qwen38-r9700/</link><guid isPermaLink="true">https://www.lucebox.com/blog/qwen38-r9700/</guid><description>The lucebox engine serves Qwen3.8-27B on a single AMD Radeon AI PRO R9700 with the DFlash2 block-diffusion drafter: up to 227 tok/s on code, 208 tok/s HumanEval average, and 3.8x llama.cpp decode running the same drafter, at 8-bit-class quality.</description><pubDate>Fri, 21 Aug 2026 12:00:00 GMT</pubDate></item><item><title>Tool prefix caching: 48× faster warm prefill for agent loops</title><link>https://www.lucebox.com/blog/tool-prefix-cache/</link><guid isPermaLink="true">https://www.lucebox.com/blog/tool-prefix-cache/</guid><description>Agent turns often resend thousands of tokens of tool definitions. Lucebox can now restore that stable prefix and process only the new conversation. On Qwen3.6-27B, median warm prefill was 1.04 seconds after a 50.35 second cold turn.</description><pubDate>Mon, 27 Jul 2026 12:00:00 GMT</pubDate></item><item><title>Lucebox (AMD Radeon AI PRO R9700 + Strix Halo) Beats NVIDIA DGX Spark by 3.63x on DeepSeek V4 Flash Decode Speed</title><link>https://www.lucebox.com/blog/deepseek-v4-asymmetric-parallelism/</link><guid isPermaLink="true">https://www.lucebox.com/blog/deepseek-v4-asymmetric-parallelism/</guid><description>51.1 tok/s on Lucebox versus our 14.09 tok/s average on one NVIDIA DGX Spark for the full 284B model. The complete $6,499 Lucebox costs 31% less than two DGX Sparks.</description><pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate></item><item><title>DeepSeek V4 Flash: 284B model, up to 32 tok/s on AMD Ryzen AI MAX+ 395</title><link>https://www.lucebox.com/blog/deepseek-v4-strix-halo/</link><guid isPermaLink="true">https://www.lucebox.com/blog/deepseek-v4-strix-halo/</guid><description>The full DeepSeek V4 Flash target runs locally on AMD Ryzen AI MAX+ 395 with 128 GB unified memory, reaching up to 32 tok/s decode and roughly 250 tok/s indexed sparse prefill.</description><pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate></item><item><title>Laguna XS 2.1 on a RTX 3090: 296 tok/s peak, 152 tok/s at 256K context</title><link>https://www.lucebox.com/blog/laguna-xs21/</link><guid isPermaLink="true">https://www.lucebox.com/blog/laguna-xs21/</guid><description>poolside&apos;s coding MoE with its official DFlash drafter: 296 tok/s peak at short context, a flat 152 tok/s at 256K tokens, and prefill at 3,500 tok/s. Lossless speculative decoding, KVFlash paging, and two model-agnostic engine optimizations run on one 24 GB card.</description><pubDate>Sat, 11 Jul 2026 12:00:00 GMT</pubDate></item><item><title>Luce KVFlash: 256K context with 72 MiB of KV on the GPU</title><link>https://www.lucebox.com/blog/kvflash/</link><guid isPermaLink="true">https://www.lucebox.com/blog/kvflash/</guid><description>On Qwen3.6-27B the KV cache costs 4.6 GiB at 256K and drags decode to 13 tok/s. KVFlash pages cold 64-token chunks to host RAM bit-exact, holding decode at 38.6 tok/s from 64K to 256K with unchanged accuracy.</description><pubDate>Fri, 12 Jun 2026 12:00:00 GMT</pubDate></item><item><title>Lucebox in a container: one image for every supported GPU</title><link>https://www.lucebox.com/blog/docker/</link><guid isPermaLink="true">https://www.lucebox.com/blog/docker/</guid><description>A prebuilt image spans the RTX 2080 Ti through RTX 5090. The fat-binary compile happens once in CI instead of on your box, with two host dependencies, self-tuning, and build provenance included.</description><pubDate>Mon, 08 Jun 2026 12:00:00 GMT</pubDate></item><item><title>Luce Spark: fit Qwen3.6 35B and Laguna XS.2 on a 16 GB GPU</title><link>https://www.lucebox.com/blog/spark/</link><guid isPermaLink="true">https://www.lucebox.com/blog/spark/</guid><description>A 33–35B MoE fires only a fraction of its experts per token but normally pays for all of them in VRAM. Spark keeps the active experts resident and swaps the rest, fitting Qwen3.6 35B-A3B and Laguna XS.2 on 16 GB GPUs.</description><pubDate>Fri, 05 Jun 2026 12:00:00 GMT</pubDate></item><item><title>Gemma 4 26B edges out DeepSeek V4 Flash (284B) on ds4-eval-92, at 5x the speed</title><link>https://www.lucebox.com/blog/gemma-vs-deepseek/</link><guid isPermaLink="true">https://www.lucebox.com/blog/gemma-vs-deepseek/</guid><description>On ds4-eval-92, Gemma 4 26B on a 24 GB RTX 5090 Laptop ties DeepSeek V4 Flash on a 192 GB Mac at 78.3% and decodes about five times faster.</description><pubDate>Wed, 20 May 2026 12:00:00 GMT</pubDate></item><item><title>Launch and tune Lucebox with real agent harnesses</title><link>https://www.lucebox.com/blog/client-harnesses/</link><guid isPermaLink="true">https://www.lucebox.com/blog/client-harnesses/</guid><description>Real-client profiles, launch scripts, and TQ3/DDTree results for OpenCode, Hermes, OpenClaw, Open WebUI, Codex, Claude Code, and Pi.</description><pubDate>Fri, 15 May 2026 12:00:00 GMT</pubDate></item><item><title>DFlash + PFlash on AMD Strix Halo: 2.5× end-to-end vs llama.cpp HIP</title><link>https://www.lucebox.com/blog/amd/</link><guid isPermaLink="true">https://www.lucebox.com/blog/amd/</guid><description>DFlash and PFlash on the Ryzen AI MAX+ 395 iGPU reach 26.85 tok/s speculative decode and 20.2 seconds of prefill at 16K, a 2.51× end-to-end speedup over vanilla llama.cpp HIP on the same silicon.</description><pubDate>Tue, 12 May 2026 12:00:00 GMT</pubDate></item><item><title>Laguna XS.2 on a 3090: 111 tok/s, 5.4x prefill, first MoE target for PFlash</title><link>https://www.lucebox.com/blog/laguna/</link><guid isPermaLink="true">https://www.lucebox.com/blog/laguna/</guid><description>Poolside Laguna XS.2 was ported into DFlash and PFlash as the first MoE target supported by PFlash, reaching about 107 tok/s decode and 5.4× faster 128K prefill than llama.cpp on one RTX 3090.</description><pubDate>Fri, 08 May 2026 12:00:00 GMT</pubDate></item><item><title>PFlash: 10× prefill speedup over llama.cpp at 128K on a RTX 3090</title><link>https://www.lucebox.com/blog/pflash/</link><guid isPermaLink="true">https://www.lucebox.com/blog/pflash/</guid><description>PFlash compresses a 128K prompt to 2.6K tokens with a small drafter before DFlash sees it, reducing cold time to first token from about 257 seconds to 24.8 seconds while preserving measured retrieval accuracy.</description><pubDate>Tue, 28 Apr 2026 12:00:00 GMT</pubDate></item><item><title>DFlash on ggml: up to 207 tok/s Qwen3.5-27B on a RTX 3090</title><link>https://www.lucebox.com/blog/dflash27b/</link><guid isPermaLink="true">https://www.lucebox.com/blog/dflash27b/</guid><description>A standalone C++ and ggml speculative decoder for Qwen3.5-27B Q4_K_M with a DFlash block-diffusion draft model and DDtree verifier, reaching up to 207 tok/s and supporting 128K context on 24 GB.</description><pubDate>Thu, 16 Apr 2026 12:00:00 GMT</pubDate></item><item><title>NVIDIA eGPU on macOS: RTX 3090 and 5090 benchmarks</title><link>https://www.lucebox.com/blog/egpu-myth/</link><guid isPermaLink="true">https://www.lucebox.com/blog/egpu-myth/</guid><description>tinygrad wrote an NVIDIA driver from scratch. We tested real models on an RTX 3090 over USB4 to measure whether an inexpensive eGPU dock can turn a Mac into an AI workstation.</description><pubDate>Tue, 14 Apr 2026 12:00:00 GMT</pubDate></item><item><title>Megakernel: Matching Apple Silicon Efficiency at 2x the Throughput on a RTX 3090</title><link>https://www.lucebox.com/blog/megakernel/</link><guid isPermaLink="true">https://www.lucebox.com/blog/megakernel/</guid><description>The first megakernel for hybrid DeltaNet and Attention LLMs fuses all 24 layers into one CUDA dispatch, reaching 1.87 tok/J and matching M5 Max efficiency at about twice the throughput on an RTX 3090.</description><pubDate>Mon, 13 Apr 2026 12:00:00 GMT</pubDate></item></channel></rss>