- Hsinchu, Taiwan
Highlights
- Pro
Pinned Loading
-
-
glm-5.2-sglang
glm-5.2-sglang PublicServing GLM-5.2 (704 GB, FP8) with SGLang + Apptainer on a Slurm HPC cluster — 8x H200 deployment record
Shell
-
5090-inference-playbook
5090-inference-playbook PublicTuning playbook: Gemma 4 31B on a single RTX 5090 with vLLM — quantization, KV-cache dtype, context length, MTP speculative decoding
JavaScript
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

