Backed by Y Combinator

Your Inference Workstation.

Order your Lucebox →

Partners

AMD
Google Gemma
Alibaba Qwen
poolside

Hardware

Private, always on, ready out of the box.

Unified memory
128 GB
AMD Ryzen AI MAX+ 395,
LPDDR5X-8000 at 256 GB/s
Dedicated GPU
32 GB
AMD Radeon AI PRO R9700,
GDDR6 at 640 GB/s
Storage
2 TB
NVMe, up to 7,400 MB/s read
Connectivity
5 GbE
Wi-Fi 7, Bluetooth and 2× USB4
Power supply
1,000 W
80 PLUS Platinum, SFX
Warranty
1 year
Parts and labor,
after 72 hours of burn-in
01

Your data stays private

Client files, code and unreleased ideas never reach a cloud model. Runs fully offline, stores no prompts, needs no account with us.

02

Built for coding agents

Point Claude Code, Codex, OpenCode or Pi at its Anthropic and OpenAI compatible API. Repos, tickets and docs stay on your network.

03

Always on, no meter

Research bots, automations and agent teams, running 24/7. Pay once, share it with your team, and forget per-token bills and rate limits.

04

Optimized for speed from the start

The engine, custom kernels, speculative decoding and tuned models come installed for this exact hardware. Plug it in, pair it, and it runs at full speed.

Check full specs

Performance

Two processors, one engine, built for speed.

Over 2× the reported DGX Spark decode speed.

Dense

Qwen3.8-27B

Prefill
903 tok/s
Decode
61.3 tok/s

From one user's DFlash2 run on a long prompt. DGX Spark and Mac M5 Max are public records with their own setups.

Details
Prompt processing tok/s, higher is better
  1. R9700 903
  2. DGX Spark Not reported
  3. Mac M5 Max 682.3
Token generation tok/s, higher is better
  1. R9700 61.3
  2. DGX Spark 36.59
  3. Mac M5 Max 56.4

MoE

DeepSeek V4 Flash

Prefill
788 tok/s
Decode
86 tok/s

Measured on Lucebox at 2K context. DGX Spark and Mac M5 Max are public records with their own setups.

Details
Prompt processing tok/s, higher is better
  1. Lucebox 788
  2. DGX Spark 1,076.8
  3. Mac M5 Max 790.18
Token generation tok/s, higher is better
  1. Lucebox 86
  2. DGX Spark 35.3
  3. Mac M5 Max 39.35
Built in partnership with AMD
See all results

Setup

Easy setup. Use it in the apps on your laptop.

Open setup
Loading the real setup UI…
01 / 11
Setup
Manage
UI, tabs and steps are clickable
  1. 01

    Open the setup page

    Visit lucebox.com/setup from a nearby laptop or Android device. There is no app or CLI to install.

  2. 02

    Pair and connect

    The onboarding page sends Wi-Fi, account, and optional Tailscale settings to the box over encrypted Bluetooth.

  3. 03

    Continue in Manage

    The local dashboard checks the machine, installs the qualified model profile, and starts the private API.

  4. 04

    Connect your apps

    Enable Lucebox Connect once, then open any supported app from Manage with a single click.

  5. 05

    Monitor your Lucebox

    Tokens per app and model, recent requests, and live machine meters. Prompts and replies are never stored.

Works with the tools you already use
Claude Code
Codex
OpenCode
Hermes
OpenClaw
Pi
Open WebUI
See the platform

Order

Order your Lucebox.

Developer

For individuals

$5,999 USD per machine / shipping included

  • Lucebox Engine, open source, with tuned models preinstalled
  • 1-year warranty on parts and labor
  • Worldwide shipping included
Order now
Estimated shipping: November 2026, due to high demand
WorldwideShipping included in the price
1 yearWarranty covering parts and labor

Community

Proudly open source, and built in public.

Ahmad@TheAhmadOsman

I like your stuff so far, keep going

Sudo su@sudoingX

this guy just cracked 134 tok/s on qwen 3.5-27b dense and 73 on new qwen 3.6-27b on a single 3090. open source moves at godspeed in 2026.

Lotto@LottoLabs

Interesting run w/ Dflash from the lucebox-hub guys

Rijndael@rot13maxi

speculative PREFILL?????

Riku Pasonen@Raitziger

I have tested some LLM server software for home PCs for Linux and Windows. Fastest and best for running home is Linux running 145 t/s, Lucebox. @pupposandro @luceboxai

fahd Mirza@fahdmirza

PFlash just killed the 4-minute blank screen problem. 128K token prefill in 25 seconds, same GPU, same model, no compromises

Geek Lite@QingQ77

Consumer-grade GPUs actually have sufficient hardware potential, general-purpose frameworks just waste most of it on overhead. Lucebox releases that potential through hand-written kernels, letting even a 2020 RTX 3090 rival Apple's latest chips on efficiency.

Takyon∞@Takyon

Crazy I was litteraly wondering how can I increase my token speed 10 min ago

Kyz2ren@ky2renzz

Crazy what @pupposandro just dropped on Qwen3.5-27B. 207 tok/s on a single 3090 with Q4_K_M and full DFlash speculative? Chinese labs + ggml hacks are just cooking on consumer hardware right now. This is the kind of local win I like to see.

Ivan Fioravanti@ivanfioravanti

RTX 3090 ready for a new life! Bringing it to @luceboxai team to make some experiments together

vitalik.eth@VitalikButerin

github.com/Luce-Org/lucebox-hub looks promising as a way to run "dense" models (eg. Qwen 27B) more efficiently. It's janky, but on my 5090 laptop it seems to be ~2x more tok/s than llama.cpp

CAPET@Capetlevrai

This is very very good work BRAVO. I love it

nash_su@nash_su

Nearly 10x faster! After finishing Decoding, it starts cranking through Prefill. The previous DFlash was already stunning enough, and now they've added PFlash. Speculative prefill, up to 10x speedup. Go try it right now.

AJ@ItsmeAjayKV

First time reading about speculative prefill, and it's crazy. 257s down to 24s for a 128K prompt on a single RTX 3090. Great article, definitely go ahead and give this a read.

Twon.@Web3Twon

impressive... so this is what it looks like when you focus on a set ram limit and optimizing for a single model above everything

Our open source inference engine is used by engineers at

NVIDIA
AMD
Google
Microsoft
Apple
Meta
Amazon
TSMC
Intel
SpaceX
IBM
Alibaba
Tencent
ByteDance
Huawei
Oracle
Adobe
Netflix
Uber
Cloudflare
Hugging Face
Luma AI
poolside
Modal
Supabase
55+Contributors
685+Pull requests
80K+Model downloads
View on GitHub

Blog

We share what we learn. Explore every detail in our blog.

The Lucebox tower on a wooden deck under a starry sky, the Qwen bear holding a magnifying glass over a printed bar chart and the DeepSeek whale with a printed diagram

Vision LLM inference: Lucebox has 3.2x the throughput of NVIDIA DGX Spark

At the highest load each machine served, one Lucebox answered 58 image questions a minute (Qwen3.8-27B on its AMD Radeon AI PRO R9700 and the 284B DeepSeek V4 Flash Vision on its Strix Halo, at the same time) and an NVIDIA DGX Spark running llama.cpp 18. With the same Qwen model file and eight questions at once, the R9700 alone finishes 2.4x sooner.

An AMD Radeon AI PRO R9700 and a 120 mm case fan on a wooden deck under a starry sky

Inference-aware cooling: 29% less fan speed on the AMD R9700, same temperatures

AI inference and cooling designed together on the R9700: the model forecasts the work ahead, a learned thermal model picks the slowest safe fan speed. Chat silent, agent loops 29% and batch work 25% quieter, same peak temperatures and throughput.

A Strix Halo mainboard on a wooden deck under a starry sky, with the ROCm and Vulkan logos above it

ROCm beats Vulkan on Strix Halo

DeepSeek V4 Flash on one Strix Halo, both engines on the same box on the same day: Lucebox ROCm is 17 to 47% faster on prefill and 28 to 47% faster on speculative decode than llama.cpp Vulkan at 8K, 32K and 123K, with plain decode measured on the same box too.

Cover of the DeepSeek V4 Flash report: Lucebox against the NVIDIA DGX Spark

3.63× the DGX Spark on DeepSeek V4 Flash

Asymmetric parallelism splits the full 284B model across the R9700 and the Strix Halo: 3.63× the decode speed of one DGX Spark, for 36% less than two.

Cover of the KVFlash report: paged KV cache on a 24 GB GPU

KVFlash: 256K context with 72 MiB of KV

Cold 64-token KV chunks page to host RAM bit-exact, so decode holds at 38.6 tok/s from 64K to 256K with unchanged accuracy.

Cover of the DeepSeek V4 Flash on Strix Halo report

DeepSeek V4 Flash on the Ryzen AI MAX+ 395

The full 284B model runs from 128 GB of unified memory: up to 32 tok/s decode and roughly 250 tok/s indexed sparse prefill.

Read all articles

FAQ

Frequently Asked Questions

Mission

We're here to make powerful AI something you own, not something you rent. Our mission is to commoditize access to fast, powerful local models, through hardware and software co-design.