case studySelf-hosted GPU inference platform — 36B MoE, voice AI and a 70-GPU farm
A fully self-hosted AI stack: a 36B MoE model served concurrently with embeddings, ASR and TTS on a single 24 GB GPU — later scaled to a 70-GPU, 10-node farm that a bad VBIOS clock state nearly took down.






