Qwen3.8-27B on RTX 3090
Inference stack for Qwen3.8-27B on a single RTX 3090: 195k context, 82 tps single request, 672 tps peak.
Qwen 3.8 · shipped Aug 14, 2026
Alibaba's local-size workhorse. Breakout clones, CLI tools and benches that run on a laptop, posted with a URL. 8 builds on this page.
Inference stack for Qwen3.8-27B on a single RTX 3090: 195k context, 82 tps single request, 672 tps peak.
Day-one Qwen3.8-27B deploy: NVFP4, 262K context, one DGX Spark, one command. Full reproducible stack.
27B on two consumer 3090s at ~75-85 tok/s with NVLink. Tool calling, reasoning, vision, MTP. Full deploy open-sourced.
Duplicate-file finder one-shotted by local Qwen 3.8 27B on a 128GB Mac via OpenCode.
Hermes builder tools updated after running Qwen 3.8 27B so smaller local models can break work into parts.
Game built and debugged with a DeepSeek V4 Pro / Qwen distill in Claude Code. Live on Hugging Face Spaces.
Same prompt, six models: DeepSeek V4 Flash, Kimi K3, Qwen 3.8, GLM and Grok. Playable results plus repo.
A single-file HTML canvas Breakout game written end to end by a local Qwen 3.8 27B through Hermes.
I made the fastest inference stack for Qwen3.8-27b on a RTX 3090. Up to 195k context, 82 tps single request, 672 tps peak 64 concurrent. github.com/syv-ai/qwen38-…
Day-one Qwen3.8-27B. NVFP4, 262K context, one DGX Spark, one command. Open-sourced the full reproducible deploy. Still Tuning IT ! github.com/tonyd2wild/Qwe…
A 27B model on two consumer 3090s at ~75-85 tok/s w/ NVLINK Qwen3.8-27B FP8 with tool calling, reasoning, vision, and MTP speculative decoding. NVLink or PCIe. Full deploy open-sourced. github.com/tonyd2wild/Qwe…
Last night I installed Qwen3.8-27B-4Bit inside oMLX on my MBP with 128GB RAM. It is in a word phenomenal and incredibly fast. I connected it to @opencode and it one-shotted this duplicate file detection utility: github.com/wjgilmore/dupe
Oh and if your using local models on Hermes the builder tools I made for it have been updated heavily from running the Qwen 3.8 27b , helps smaller local models break builds into manegable parts github.com/embwl0x/hermes…
Hello again, y'all! As we eagerly await the new Qwen 3.8 models, we have a new release for the sub 16GB VRAM crowd! DeepSeek V4 Pro Qwen 9B and 4B are now live! Our prior fine-tuning of DeepSeek V4 flash preview was my favorite 9B general use model; this one takes it Show more
The Big Tetris Battle 🏆 I've asked GLM-5.3, Grok 4.6, DeepSeek v4 Flash, Kimi K3, Qwen3.8-Max, and Qwen3.8-27B to create the most impressive Tetris game they can in a single shot. To keep things fair, I used ZCode for GLM-5.3, Grok Build for Grok 4.6, DeepSeek Harness for Show more
Guess what AI Model built this game? Local Qwen 3.8 27B (NVFP4) via Hermes wrote a single-file HTML canvas Breakout. This model is straight fire 🔥