Qwen3.8-27B FP8 on 2x 3090
27B on two consumer 3090s at ~75-85 tok/s with NVLink. Tool calling, reasoning, vision, MTP. Full deploy open-sourced.
Related builds
Qwen3.8-27B NVFP4 on DGX Spark
Day-one Qwen3.8-27B deploy: NVFP4, 262K context, one DGX Spark, one command. Full reproducible stack.
sparktop
Dashboard for multi DGX Spark setups: inference throughput, GPU processes, and more.
Qwen3.8-27B on RTX 3090
Inference stack for Qwen3.8-27B on a single RTX 3090: 195k context, 82 tps single request, 672 tps peak.
handMediaPipeHands
85-line finger-tracking toy for Physical AI video cleanup. Claude Code plus MediaPipe.