Technical Insights & Architecture Papers
Deep-dives on high-throughput backend architecture, Python & Django performance, real-time Voice AI, and resilient database modeling.
PostgreSQL Index-Only Scans & Covering Indexes: Eliminating Heap Fetches on High-Throughput Read APIs
Eliminate disk I/O bottlenecks in PostgreSQL read APIs by architecting Covering Indexes with the INCLUDE clause and tuning Visibility Maps to guarantee true Index-Only Scans.
Turn-Taking Prediction in Conversational Voice AI: Combining Acoustic VAD with Semantic End-of-Thought (EoT) Classifiers
Eliminate awkward conversational latency and premature interruptions in real-time voice agents by orchestrating acoustic Voice Activity Detection with streaming semantic End-of-Thought classifiers.
TCP BBRv3 Congestion Control in Production: Slashing Tail Latency & Bufferbloat for Real-Time LLM Token & Audio Streams
Discover how switching Linux kernel congestion control from Cubic to BBRv3 eliminates bufferbloat and slashes p99 tail latency across WebSockets, WebRTC media, and streaming LLM token delivery.
Horizontal Database Sharding at Scale: Citus Distributed Tables, Distributed Transactions, and Partition-Wise Joins
When a single PostgreSQL primary reaches write saturation and storage limits, Citus transforms PostgreSQL into a distributed cluster. Master shard keys, 2PC distributed transactions, and co-located joins.
Zero-Copy In-Memory Serialization: FlatBuffers and Cap'n Proto vs. Protocol Buffers in High-Throughput Microservices
Protocol Buffers require costly object decoding and memory allocations during serialization. Discover zero-copy serialization engines that access structured binary payloads directly in memory buffers.
Sandboxing Untrusted Code in Python with WebAssembly (Wasmtime): Zero-Container Secure Plugin Execution
Running user-submitted scripts via eval(), exec(), or Docker containers is either dangerous or resource-heavy. Implement sub-millisecond, memory-isolated Wasmtime WebAssembly sandboxes in Python.