The SaaS Telemetry Tax
Modern distributed systems generate voluminous telemetry: Nginx access logs, structured JSON application events, database slow query traces, and OS syslog streams. For engineering teams operating on cloud providers, shipping these events to SaaS observability vendors (such as Datadog or New Relic) quickly becomes one of the largest infrastructure line items. When monthly bills spike, management inevitably mandates aggressive log sampling or short retention windows—destroying the diagnostic trail needed during production incident post-mortems.
The irony is that standard cloud virtual private servers (VPS) now provide remarkable computing density: an 8-vCPU, 16GB RAM VPS costing under $40/month possesses more than enough I/O bandwidth and CPU capacity to ingest, compress, and query tens of millions of events daily. By combining Vector (an ultra-fast Rust log processor), ClickHouse (the world's fastest columnar analytical database), and Grafana, you can construct an enterprise-tier observability platform with zero recurring vendor fees.
1. Architecture: High-Throughput Columnar Pipeline
Traditional self-hosted log stacks (like Elasticsearch or Loki) suffer from high memory footprints and CPU contention under indexing pressure. Elasticsearch requires massive JVM heaps and creates huge inverted index files. Loki, while lightweight, slows down drastically when executing complex filtering across non-indexed labels.
ClickHouse fundamentally changes this equation. Because it stores data in columnar blocks compressed with Zstandard (ZSTD) or LZ4, log data achieves 85% to 92% compression ratios. An analytical query scanning 50,000,000 rows across a 7-day period executes in less than 120 milliseconds using vectorized SIMD CPU instructions.
[ Systemd / Nginx / Apps ]
│
▼
[ Vector (Rust Ingestion Daemon) ] ── (Batched JSON over HTTP) ──▶ [ ClickHouse Columnar DB ]
│
▼
[ Grafana Dashboards ]
2. ClickHouse Schema for High-Density Logging
Deploying ClickHouse on Ubuntu is straightforward via the official deb repository. Create an append-optimized log table utilizing the MergeTree engine partitioned by date:
CREATE DATABASE IF NOT EXISTS telemetry;
CREATE TABLE telemetry.application_logs (
timestamp DateTime64(6, 'UTC') CODEC(DoubleDelta, ZSTD(1)),
service LowCardinality(String) CODEC(ZSTD(1)),
level LowCardinality(String) CODEC(ZSTD(1)),
environment LowCardinality(String) CODEC(ZSTD(1)),
ip String CODEC(ZSTD(3)),
request_method LowCardinality(String) CODEC(ZSTD(1)),
request_path String CODEC(ZSTD(3)),
status_code UInt16 CODEC(ZSTD(1)),
duration_ms Float32 CODEC(ZSTD(1)),
message String CODEC(ZSTD(3)),
attributes_json String CODEC(ZSTD(3))
)
ENGINE = MergeTree()
PARTITION BY toYYYYMM(timestamp)
ORDER BY (service, level, toDate(timestamp), timestamp)
TTL toDateTime(timestamp) + INTERVAL 90 DAY DELETE
SETTINGS index_granularity = 8192;
Notice the deliberate use of LowCardinality(String) for repetitive attributes like level and service. This enables dictionary encoding, reducing string storage to single-byte integer lookups in RAM.
3. Ingestion: Configuring Vector Daemon
Vector (written in Rust by Datadog/Timber) consumes negligible memory (typically under 40MB RSS) while parsing hundreds of thousands of events per second. Install Vector on your host and configure /etc/vector/vector.yaml to harvest systemd journal logs and stream them directly into ClickHouse:
sources:
system_journal:
type: journald
exclude_units: ["vector.service"]
transforms:
parse_json_logs:
type: remap
inputs: ["system_journal"]
source: |
.timestamp = parse_timestamp!(.timestamp, format: "%s")
.service = .SYSTEMD_UNIT || "system"
.level = .PRIORITY || "info"
.environment = "production"
# Extract structured JSON payload if present
parsed, err = parse_json(.message)
if err == null {
.message = parsed.msg || .message
.attributes_json = encode_json(parsed)
} else {
.attributes_json = "{}"
}
sinks:
clickhouse_out:
type: clickhouse
inputs: ["parse_json_logs"]
endpoint: "http://127.0.0.1:8123"
database: "telemetry"
table: "application_logs"
compression: gzip
batch:
max_bytes: 10485760 # 10MB batches
timeout_secs: 2
Perimeter Defense: Go beyond traditional SSH port changes and password disabling by implementing zero-public SSH using WireGuard mesh networks and UFW firewalls.
For related production architectures and system implementations, explore these companion guides:
- High-Density Time-Series Analytics with ClickHouse — Leverage ClickHouse AggregatingMergeTree engines to store billions of telemetry rows.
- Pragmatic Production Observability: Structured JSON Logging — Stream structured systemd logs and OpenTelemetry spans into local Vector pipelines.
- Zero-Public SSH with WireGuard Mesh Networks — Secure internal Grafana and ClickHouse administrative ports across private WireGuard tunnels.
Production Takeaway
By pairing Vector's zero-overhead Rust collection with ClickHouse's columnar compression, you gain instantaneous search across billions of production events without paying a cent in SaaS telemetry bills. When an incident occurs, full un-sampled logs are queryable in milliseconds directly inside Grafana.