Deep-Dive: Zhipu AI's GLM-5 and Local Open-Weight MoE Serving
Analyzing Zhipu AI's GLM-5 architecture: Mixture-of-Experts routing, controllable reasoning tokens, and local deployment using vLLM and SGLang.
Read Post →AI LEADER • VISUAL DESIGN ENTHUSIAST
A software engineer with a balanced left and right brain, specializing in pipeline development, DevOps, graphics tools, and site reliability.
Analyzing Zhipu AI's GLM-5 architecture: Mixture-of-Experts routing, controllable reasoning tokens, and local deployment using vLLM and SGLang.
Read Post →Step-by-step developer tutorial on migrating your self-hosted OpenClaw desktop backend from Claude Code to Google Antigravity and Gemini CLI on macOS.
Read Post →Under the hood of Google Gemini 3.8 Live: native speech-to-speech architectures, sub-200ms conversational latency, and asynchronous Extended Thinking with background tool execution.
Read Post →