Engineering Publications
Technical Insights & Architecture Papers
Deep-dives on high-throughput backend architecture, Python & Django performance, real-time Voice AI, and resilient database modeling.
Multi-Head Speculative Decoding (Medusa): Accelerating vLLM Inference Without a Draft Model
Double LLM serving throughput without the operational burden of maintaining small draft models. Explore Medusa multi-head architecture, tree-based attention verification, and vLLM integration.