🧑‍💻 Qwen2.5-Coder-14B on ZeroGPU

Qwen2.5-Coder-14B-Instruct in bf16, served on ZeroGPU — the GPU is attached per request, so the Space costs nothing to keep online and every visitor spends their own daily quota.

Generation runs at roughly 20 tokens/s. The first request after an idle period adds ~12 s of cold start while ~28 GB of weights stream into VRAM.