Your Inference Computer.

Our open source inference engine is used by engineers at

AMD
NVIDIA
Google
Meta
Apple
Microsoft
Intel
AWS
Netflix
Uber
Cloudflare
Supabase
Oracle
Hugging Face
Prime Intellect
Alibaba

Local inference hardware

Built to combine unified memory capacitywith dedicated GPU performance.

01

Built for inference

128 GB of unified memory holds long context and larger models. 32 GB of GDDR6 keeps the hot path on the GPU.

02

Ready in about a minute

Engine, tuned models, and compatible APIs are installed before shipping. Plug in, pair, run.

03

Private by default

No mandatory cloud path. Your code, prompts, and model weights stay on hardware you own.

USD
$6,499
$7,900 after August 31
Limited production capacity Tell us what you are building and request a Lucebox.
Apply now

Full specs and information

An aluminium chassis engineered for compact size and sustained performance.

Finish / Black
Black Lucebox enclosure in a grey studio, side and front panel view
Black Lucebox enclosure in a grey studio, front and side view
Production specification
Unified memory 128 GB LPDDR5X-8000 · Strix Halo 256 GB/s · 40 GPU cores · up to 60 TFLOPS
Dedicated GPU 32 GB AMD Radeon AI PRO R9700 · GDDR6 640 GB/s · 64 CUs · up to 191 TFLOPS at FP16
Power supply 1,000 W Corsair SF1000 80 PLUS Platinum · ~500 W sustained · ~40 W idle
Storage 2 TB KingSpec XG7000 · NVMe 1.4 Up to 7,400 / 6,600 MB/s
Chassis & thermals
EnclosureMachined aluminium, hand-assembled in-house
Dimensions340 × 110 × 320 mm (13.4 × 4.3 × 12.6 in)
Volume11.97 L (approximately 730 in³)
Weight6 kg (13.2 lb)
CoolingAir-cooled, tuned for low noise under sustained load
Connectivity
Ports2× USB4, 2× USB-A, HDMI 2.1, 2× DisplayPort 2.1, 5 GbE Ethernet, plus 4 display outputs on the GPU
WirelessWi-Fi 7 and Bluetooth, built in
System
SoftwareUbuntu, engine and models pre-installed, OpenAI and Anthropic compatible API
Warranty1 year, parts and labor, 72-hour burn-in before shipping
Built in partnership with AMD

Our work makes Lucebox AMD 3.62× faster than one DGX Spark on DeepSeek V4 Flash.

Decode throughput / DeepSeek V4 Flash

50.8 tok/s warm decode

Our asymmetric expert-parallelism work runs the Radeon AI PRO R9700 and Strix Halo together, keeping hot experts on the faster device while the long tail uses unified memory.

3.62×faster
Published comparison The full 284B model, compared with our measured 14.04 tok/s average on one DGX Spark.
Read the article

Box-to-agent sequence

A plug-and-play solution for running models locally.

Luce CLI / setup sequenceState / ready
$ brew install lucebox && luce ░██ ░██ ░██ ░██ ░██ ░███████ ░███████ ░████████ ░███████ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░██ ░█████████ ░██ ░██ ░██ ░██ ░█████ ░██ ░██ ░███ ░██ ░██ ░██ ░███ ░██ ░██ ░██ ░██ ░██ ░██████████ ░█████░██ ░███████ ░███████ ░██░█████ ░███████ ░██ ░██ Looking for your Lucebox via Bluetooth... Found: Lucebox-A7F3 Scanning WiFi networks... Available networks: home-5g, home-2.4, guest, xfinitywifi Select (1-4): 1 Password: ●●●●●●●● Connecting to home-5g... ✓ Connected. Dashboard at http://lucebox.local ✓ Lucebox is ready. ╭──────────────────────────────────────────────────────────────╮ Lucebox Model: Qwen3.6-27B Q4_K_M LLM: ● online Endpoint: http://lucebox.local:8080 Commands /status System overview /models List models loaded /engine Runs our lucebox-hub engine /harness Switch agent harness /shell Open SSH shell /help All commands Type a message, or / for commands ╰──────────────────────────────────────────────────────────────╯ > Hey, I'm your Lucebox. How can I help you?
01

Pair

Bluetooth finds the machine and moves it onto your network without a temporary monitor or keyboard.

02

Load

The engine and selected models are installed and tuned before shipping. No driver setup or image build.

03

Route

Use the local OpenAI or Anthropic compatible endpoint from your existing agent tools and clients.

Natively supports any harness
Claude Code
Claude Code
Codex
Codex
OpenCode
OpenCode
Hermes
Hermes
OpenClaw
OpenClaw
Open WebUI
Open WebUI
Ollama
Ollama

Limited production capacity

Apply now to get yours.

Limited launch price / USD $6,499
$7,90018% launch reductionEnds August 31

Time remaining / --d --h --m --s

01
Share your use case

Tell us what you are building, the workloads involved, how many boxes you need, and your preferred timing.

02
Talk with our team

We schedule a short call to understand the work, confirm technical fit, and discuss delivery needs.

03
Confirm allocation

Capacity is limited and demand is high. If we can allocate a unit, we confirm the delivery window and send a secure purchase link after the call.

1 yearWarranty covering parts and labor
72 hoursFull-system testing before shipping
GlobalShipping included in the price

Tell us what you are building

Required fields marked *

We do not collect payment details here. If we can allocate a unit, we schedule a call and send a secure purchase link afterward.

Open development

Proudly open source, and built in public.

Public response / X

What builders said after testing our inference engine.

Ahmad@TheAhmadOsman

I like your stuff so far, keep going

Sudo su@sudoingX

this guy just cracked 134 tok/s on qwen 3.5-27b dense and 73 on new qwen 3.6-27b on a single 3090. open source moves at godspeed in 2026.

Lotto@LottoLabs

Interesting run w/ Dflash from the lucebox-hub guys

Rijndael@rot13maxi

speculative PREFILL?????

Riku Pasonen@Raitziger

I have tested some LLM server software for home PCs for Linux and Windows. Fastest and best for running home is Linux running 145 t/s, Lucebox. @pupposandro @luceboxai

fahd Mirza@fahdmirza

PFlash just killed the 4-minute blank screen problem. 128K token prefill in 25 seconds, same GPU, same model, no compromises

Geek Lite@QingQ77

Consumer-grade GPUs actually have sufficient hardware potential, general-purpose frameworks just waste most of it on overhead. Lucebox releases that potential through hand-written kernels, letting even a 2020 RTX 3090 rival Apple's latest chips on efficiency.

Takyon∞@Takyon

Crazy I was litteraly wondering how can I increase my token speed 10 min ago

Kyz2ren@ky2renzz

Crazy what @pupposandro just dropped on Qwen3.5-27B. 207 tok/s on a single 3090 with Q4_K_M and full DFlash speculative? Chinese labs + ggml hacks are just cooking on consumer hardware right now. This is the kind of local win I like to see.

Ivan Fioravanti@ivanfioravanti

RTX 3090 ready for a new life! Bringing it to @luceboxai team to make some experiments together

vitalik.eth@VitalikButerin

github.com/Luce-Org/lucebox-hub looks promising as a way to run "dense" models (eg. Qwen 27B) more efficiently. It's janky, but on my 5090 laptop it seems to be ~2x more tok/s than llama.cpp

CAPET@Capetlevrai

This is very very good work BRAVO. I love it

nash_su@nash_su

Nearly 10x faster! After finishing Decoding, it starts cranking through Prefill. The previous DFlash was already stunning enough, and now they've added PFlash. Speculative prefill, up to 10x speedup. Go try it right now.

AJ@ItsmeAjayKV

First time reading about speculative prefill, and it's crazy. 257s down to 24s for a 128K prompt on a single RTX 3090. Great article, definitely go ahead and give this a read.

Twon.@Web3Twon

impressive... so this is what it looks like when you focus on a set ram limit and optimizing for a single model above everything

50+Contributors
500+Pull requests
10+Versions released
Inspect the source

Experiments and benchmarks

We share what we learn. Explore every detail in our blog.

Questions

Frequently Asked Questions

A plug-and-play computer for local AI inference. It arrives with our open source inference engine, tuned models, custom kernels, and speculative decoding ready to use through a single CLI.
An AMD Ryzen AI MAX+ 395 with 128GB of LPDDR5X-8000 unified memory, an AMD Radeon AI PRO R9700 with 32GB of GDDR6, a 2TB KingSpec XG7000 NVMe SSD, and a Corsair SF1000 1000W 80+ Platinum power supply. The custom aluminium chassis has a volume of 11.97 liters.
$6,499 USD per machine through August 31, 2026, then $7,900. Global shipping is included. After a short call confirms fit and availability, we send a secure purchase link.
Production units ship in limited batches. Demand currently exceeds capacity, so delivery windows can vary. We confirm expected timing on a call before sending a purchase link.
Yes. Include the quantity in your application. We review larger requests based on the use case, destination, and current production capacity, so allocation and delivery timing are confirmed on the call.
Yes. Our inference work is public at github.com/Luce-Org/lucebox-hub, with 50+ contributors, 500+ pull requests, and more than 10 released versions. It includes hand-tuned kernels, DFlash speculative decoding, PFlash prefill, and hybrid-model optimizations.
Far more than fits in 32GB of VRAM. The Strix Halo's 128GB of unified memory lets you run much larger models too, like MiniMax 2.7 and DeepSeek V4-Flash. We pre-tune Qwen3.6, GLM-4.6, DeepSeek V4-Flash, and the Llama 4 family, and you can bring your own GGUF, the CLI takes care of loading it.
Yes. Lucebox exposes OpenAI-compatible and Anthropic-compatible endpoints. Point a supported client or agent harness to the Lucebox base URL and keep using its normal workflow.
Lucebox does not require either one. It exposes an OpenAI-compatible endpoint, so tools that accept a custom base URL can connect directly while Ollama and LM Studio remain available for any separate workflows you already use.
No. Lucebox runs fully offline once your model is loaded. Internet is only needed for updates and the optional cloud fallback.
The system is quiet at idle and the Radeon AI PRO R9700 cooling system becomes audible during sustained inference. The chassis directs hot air out the back, away from the desk area.
Around 500 watts under full inference load, with up to 300 W for the Radeon AI PRO R9700 and the remaining system power allocated to the Ryzen platform and components. Idle consumption is approximately 40 W. Plug it into any standard wall outlet.
One year on the full machine, parts and labor. We replace or repair any unit that fails under normal use.
Yes. Full root access. It is your machine. Open the shell, install whatever you want, run whatever you want.
If a unit arrives damaged, is dead on arrival, or develops a fault covered by the warranty, contact us and we will arrange the appropriate repair, replacement, or refund. Full terms are provided before purchase.
Yes. Global shipping is included in the listed machine price. VAT, duties, and import taxes remain separate where applicable. We confirm the destination and expected delivery window on the call.
Yes. The 128 GB of unified memory plus 32 GB of VRAM lets you keep a hot model resident in VRAM while a second sits in unified memory ready to swap in. The CLI handles loading and routing per request.

Private compute / production unit