Zero-Cost Production Observability: Vector, ClickHouse & Grafana on a Single Linux VPS

SaaS telemetry pricing scales aggressively with log volume, forcing teams into dangerous log sampling. Discover how to deploy a high-throughput, zero-cost observability stack using Vector daemon, single-node ClickHouse, and Grafana on a budget Ubuntu VPS.

The SaaS Telemetry Tax

Modern distributed systems generate voluminous telemetry: Nginx access logs, structured JSON application events, database slow query traces, and OS syslog streams. For engineering teams operating on cloud providers, shipping these events to SaaS observability vendors (such as Datadog or New Relic) quickly becomes one of the largest infrastructure line items. When monthly bills spike, management inevitably mandates aggressive log sampling or short retention windows—destroying the diagnostic trail needed during production incident post-mortems.

The irony is that standard cloud virtual private servers (VPS) now provide remarkable computing density: an 8-vCPU, 16GB RAM VPS costing under $40/month possesses more than enough I/O bandwidth and CPU capacity to ingest, compress, and query tens of millions of events daily. By combining Vector (an ultra-fast Rust log processor), ClickHouse (the world's fastest columnar analytical database), and Grafana, you can construct an enterprise-tier observability platform with zero recurring vendor fees.

1. Architecture: High-Throughput Columnar Pipeline

Traditional self-hosted log stacks (like Elasticsearch or Loki) suffer from high memory footprints and CPU contention under indexing pressure. Elasticsearch requires massive JVM heaps and creates huge inverted index files. Loki, while lightweight, slows down drastically when executing complex filtering across non-indexed labels.

ClickHouse fundamentally changes this equation. Because it stores data in columnar blocks compressed with Zstandard (ZSTD) or LZ4, log data achieves 85% to 92% compression ratios. An analytical query scanning 50,000,000 rows across a 7-day period executes in less than 120 milliseconds using vectorized SIMD CPU instructions.

[ Systemd / Nginx / Apps ]
           │
           ▼
[ Vector (Rust Ingestion Daemon) ] ── (Batched JSON over HTTP) ──▶ [ ClickHouse Columnar DB ]
                                                                             │
                                                                             ▼
                                                                    [ Grafana Dashboards ]

2. ClickHouse Schema for High-Density Logging

Deploying ClickHouse on Ubuntu is straightforward via the official deb repository. Create an append-optimized log table utilizing the MergeTree engine partitioned by date:

CREATE DATABASE IF NOT EXISTS telemetry;

CREATE TABLE telemetry.application_logs (
    timestamp DateTime64(6, 'UTC') CODEC(DoubleDelta, ZSTD(1)),
    service LowCardinality(String) CODEC(ZSTD(1)),
    level LowCardinality(String) CODEC(ZSTD(1)),
    environment LowCardinality(String) CODEC(ZSTD(1)),
    ip String CODEC(ZSTD(3)),
    request_method LowCardinality(String) CODEC(ZSTD(1)),
    request_path String CODEC(ZSTD(3)),
    status_code UInt16 CODEC(ZSTD(1)),
    duration_ms Float32 CODEC(ZSTD(1)),
    message String CODEC(ZSTD(3)),
    attributes_json String CODEC(ZSTD(3))
)
ENGINE = MergeTree()
PARTITION BY toYYYYMM(timestamp)
ORDER BY (service, level, toDate(timestamp), timestamp)
TTL toDateTime(timestamp) + INTERVAL 90 DAY DELETE
SETTINGS index_granularity = 8192;

Notice the deliberate use of LowCardinality(String) for repetitive attributes like level and service. This enables dictionary encoding, reducing string storage to single-byte integer lookups in RAM.

3. Ingestion: Configuring Vector Daemon

Vector (written in Rust by Datadog/Timber) consumes negligible memory (typically under 40MB RSS) while parsing hundreds of thousands of events per second. Install Vector on your host and configure /etc/vector/vector.yaml to harvest systemd journal logs and stream them directly into ClickHouse:

sources:
  system_journal:
    type: journald
    exclude_units: ["vector.service"]

transforms:
  parse_json_logs:
    type: remap
    inputs: ["system_journal"]
    source: |
      .timestamp = parse_timestamp!(.timestamp, format: "%s")
      .service = .SYSTEMD_UNIT || "system"
      .level = .PRIORITY || "info"
      .environment = "production"
      
      # Extract structured JSON payload if present
      parsed, err = parse_json(.message)
      if err == null {
        .message = parsed.msg || .message
        .attributes_json = encode_json(parsed)
      } else {
        .attributes_json = "{}"
      }

sinks:
  clickhouse_out:
    type: clickhouse
    inputs: ["parse_json_logs"]
    endpoint: "http://127.0.0.1:8123"
    database: "telemetry"
    table: "application_logs"
    compression: gzip
    batch:
      max_bytes: 10485760  # 10MB batches
      timeout_secs: 2

Perimeter Defense: Go beyond traditional SSH port changes and password disabling by implementing zero-public SSH using WireGuard mesh networks and UFW firewalls.

Architectural Continuity & Deep Dives

For related production architectures and system implementations, explore these companion guides:

Production Takeaway

By pairing Vector's zero-overhead Rust collection with ClickHouse's columnar compression, you gain instantaneous search across billions of production events without paying a cent in SaaS telemetry bills. When an incident occurs, full un-sampled logs are queryable in milliseconds directly inside Grafana.

All Insights
Chat on WhatsApp