Disaggregated Inference: Separating Prefill and Decode Nodes at Scale
Scale LLM serving with disaggregated inference. Decouple compute-heavy prefill from memory-bound decode nodes to eliminate TTFT/TPOT interference.
Read Post →AI LEADER • VISUAL DESIGN ENTHUSIAST
A software engineer with a balanced left and right brain, specializing in pipeline development, DevOps, graphics tools, and site reliability.
Scale LLM serving with disaggregated inference. Decouple compute-heavy prefill from memory-bound decode nodes to eliminate TTFT/TPOT interference.
Read Post →An exhaustive architectural, mathematical, and benchmark comparison across NeRFs, Instant-NGP, 3D Gaussian Splatting, and NVIDIA NuRec.
Read Post →Explore NVIDIA NuRec: turning drive logs into interactive 3D Gaussian digital twins with dynamic actor decomposition and cross-carline virtual sensor rig adaptation.
Read Post →