Qwen3.8-27B on RTX 3090
Inference stack for Qwen3.8-27B on a single RTX 3090: 195k context, 82 tps single request, 672 tps peak.
Related builds
cubicle
Harness that gives every AI bot its own directory, browser profile, cookie jar, and environment. Grants expire; the door closes.
sparktop
Dashboard for multi DGX Spark setups: inference throughput, GPU processes, and more.
RulesAsPrograms
Turns each agent rule into a function that checks the agent's work, instead of hoping it follows the prompt.
Slopdar
Open-sourced detector for AI-slop websites. Add a pattern, improve a detector, or build on top of it.