Semantic Caching for LLMs with Redis & `pgvector`: Slashing API Costs & Sub-20ms Latency

Identical and semantically equivalent LLM queries waste massive API budgets and introduce 1.5s+ latency. Build a high-throughput semantic caching layer using embeddings, cosine distance thresholds, and Redis vector indexing for sub-20ms instant responses.

The Latency and Cost Crisis in Generative AI Pipelines

Integrating Large Language Models (LLMs) into customer-facing products introduces an entirely new performance paradigm. Unlike traditional database lookups that resolve in 5ms to 25ms, an LLM completion API call to Anthropic Claude or OpenAI GPT-4 takes anywhere from 800ms to 4,500ms to generate a complete answer. Furthermore, pricing models based on token counts mean that popular apps burn thousands of dollars monthly generating repetitive answers to virtually identical questions.

Traditional key-value HTTP caching fails completely in AI applications. If User A asks "How do I configure PgBouncer connection pooling?" and User B asks "What is the best way to set up PgBouncer for PostgreSQL?", an exact string hash match (MD5/SHA256) yields a cache miss despite both users seeking the exact same underlying knowledge. Enter Semantic Caching: indexing queries in high-dimensional vector space and returning cached answers whenever the cosine similarity crosses a predefined confidence threshold.

1. Architecture of a Semantic Caching Pipeline

A production-ready semantic cache evaluates incoming queries through a fast two-tier vector matching pipeline before making any network requests to external LLM providers:

Pipeline Stage Technology Latency Budget Operational Role
1. Text Normalization Python / Regex < 1ms Strip trailing punctuation, whitespace, and case sensitivity
2. Dense Vector Embedding FastEmbed / OpenAI `text-embedding-3-small` 12ms - 35ms Convert natural language into 384-d or 1536-d float array
3. Approximate Nearest Neighbor (ANN) Redis Vector Similarity / PostgreSQL `pgvector` HNSW 4ms - 15ms Retrieve closest cached query vector via cosine distance
4. Threshold Validation Cosine Similarity Threshold (e.g. > 0.92) < 0.5ms If similarity ≥ 0.92, return cached response immediately; else call LLM

2. Implementing Semantic Caching with Redis & Python

Redis Stack provides native vector similarity search (VSS) using HNSW (Hierarchical Navigable Small World) graphs, enabling sub-10ms nearest neighbor queries across hundreds of thousands of cached prompts:

import numpy as np
import redis
from redis.commands.search.field import VectorField, TextField
from redis.commands.search.indexDefinition import IndexDefinition, IndexType
from redis.commands.search.query import Query

class RedisSemanticCache:
    INDEX_NAME = "idx_llm_cache"
    VECTOR_DIM = 384  # Matches all-MiniLM-L6-v2 or fast local embeddings
    SIMILARITY_THRESHOLD = 0.92  # Cosine similarity cutoff

    def __init__(self, redis_url="redis://localhost:6379"):
        self.client = redis.from_url(redis_url)
        self._ensure_index()

    def _ensure_index(self):
        try:
            self.client.ft(self.INDEX_NAME).info()
        except redis.exceptions.ResponseError:
            # Create HNSW index over cosine distance metric
            schema = (
                TextField("prompt"),
                TextField("response"),
                VectorField(
                    "embedding",
                    "HNSW",
                    {
                        "TYPE": "FLOAT32",
                        "DIM": self.VECTOR_DIM,
                        "DISTANCE_METRIC": "COSINE",
                        "M": 16,
                        "EF_CONSTRUCTION": 200
                    }
                )
            )
            self.client.ft(self.INDEX_NAME).create_index(
                schema,
                definition=IndexDefinition(prefix=["cache:"], index_type=IndexType.HASH)
            )

    def lookup(self, query_embedding: list[float]) -> str | None:
        # Query Redis vector index for nearest semantic neighbor.
        query_vector = np.array(query_embedding, dtype=np.float32).tobytes()
        
        # Search for top 1 nearest vector
        q = (
            Query("*=>[KNN 1 @embedding $vec AS score]")
            .sort_by("score")
            .return_fields("prompt", "response", "score")
            .dialect(2)
        )
        results = self.client.ft(self.INDEX_NAME).search(q, query_params={"vec": query_vector})

        if results.docs:
            top_match = results.docs[0]
            # In Redis COSINE metric: score is cosine distance (1 - similarity)
            cosine_similarity = 1.0 - float(top_match.score)
            if cosine_similarity >= self.SIMILARITY_THRESHOLD:
                return top_match.response
        return None

    def store(self, prompt: str, response: str, embedding: list[float], ttl_seconds=86400):
        # Store new completion with TTL.
        import hashlib
        doc_id = f"cache:{hashlib.sha256(prompt.encode()).hexdigest()[:16]}"
        self.client.hset(
            doc_id,
            mapping={
                "prompt": prompt,
                "response": response,
                "embedding": np.array(embedding, dtype=np.float32).tobytes()
            }
        )
        self.client.expire(doc_id, ttl_seconds)

3. Fallback Alternative: PostgreSQL `pgvector` Semantic Cache

If your infrastructure already relies on PostgreSQL and you prefer to avoid hosting a dedicated Redis Stack cluster, `pgvector` provides exceptional semantic caching capabilities utilizing HNSW indexes directly in relational tables:

-- Enable pgvector extension and create semantic cache table
CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE llm_semantic_cache (
    id BIGSERIAL PRIMARY KEY,
    prompt_text TEXT NOT NULL,
    response_text TEXT NOT NULL,
    embedding vector(384) NOT NULL,
    hit_count INTEGER DEFAULT 1,
    created_at TIMESTAMPTZ DEFAULT NOW(),
    last_accessed_at TIMESTAMPTZ DEFAULT NOW()
);

-- Build high-speed HNSW index over cosine distance operator (<=>)
CREATE INDEX idx_llm_cache_hnsw 
ON llm_semantic_cache 
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 128);

-- Sub-15ms semantic match query:
SELECT 
    prompt_text,
    response_text,
    1 - (embedding <=> '[0.024, -0.015, ...]'::vector) AS cosine_similarity
FROM llm_semantic_cache
WHERE (embedding <=> '[0.024, -0.015, ...]'::vector) < 0.08  -- 0.08 distance = 0.92 similarity
ORDER BY embedding <=> '[0.024, -0.015, ...]'::vector
LIMIT 1;

For an in-depth comparison of HNSW vs IVFFlat parameter sizing on large datasets, explore our technical benchmark on PostgreSQL pgvector in Production.

4. Cache Invalidation and Dynamic Context Guardrails

Semantic caching must never be applied blindly to user-specific or stateful operations. Implement defensive guardrails before cache storage:

  • Exclude PII and Private Data: Always bypass semantic caching for authenticated user profiles, financial ledgers, or HIPAA/GDPR sensitive contexts.
  • Namespace by Model & System Prompt: Hash your system prompt version and model identifier into the cache key prefix. If you update your system prompt, old cached answers are instantly invalidated.
  • Set TTL by Volatility: Factual technical answers can comfortably sit for 7 to 30 days, while market pricing or stock updates should have a TTL of minutes or bypass caching entirely.
Architectural Continuity & Deep Dives

For related production architectures and system implementations, explore these companion guides:

Production Engineering Takeaways

  • Sub-20ms latency win: Serving 40% of queries from a semantic cache drops median API response latency from 1,800ms to 18ms.
  • Massive token cost reduction: Repetitive support questions, FAQ lookups, and technical documentation queries yield 50%+ reduction in monthly OpenAI/Anthropic API bills.
  • Threshold tuning is critical: A threshold of 0.92 - 0.95 is the industry sweet spot. Anything below 0.88 risks returning inaccurate answers to nuanced questions.
All Insights
Chat on WhatsApp