gRPC-JSON Transcoding with Envoy: Serving High-Throughput gRPC and Public REST APIs from a Single Protobuf Schema

Maintaining duplicate REST and gRPC controller layers causes API drift and doubles engineering maintenance. Configure Envoy Proxy to dynamically transcode public HTTP/1.1 JSON into internal high-speed binary HTTP/2 gRPC.

The Dual-API Maintenance Crisis

Modern microservice architectures face a fundamental communication dilemma. For internal service-to-service communication, gRPC over HTTP/2 with binary Protocol Buffers is the undisputed gold standard, offering sub-millisecond serialization, strict schema enforcement, and multiplexed persistent TCP streams. However, for external public-facing clients—such as web frontends, mobile applications, and third-party developer integrations—HTTP/1.1 REST with JSON payloads remains mandatory.

Historically, engineering teams solved this by writing two separate API layers: an internal gRPC service, and a secondary "Backend for Frontend" (BFF) or REST gateway that parses incoming JSON, maps it into Protobuf objects, and proxies the call to gRPC. This approach is an architectural maintenance nightmare:

  • Duplicated Boilerplate: Engineers must write and maintain twin routing tables, input validation models, and serialization code for every single endpoint.
  • API Contract Drift: Field names, optionality rules, and error handling behaviors inevitably diverge between the REST and gRPC interfaces over time.
  • Serialization Overhead: The REST gateway constantly burns CPU cycles parsing JSON strings into memory objects only to immediately re-encode them into binary Protobuf bytes.

1. Architecture of gRPC-JSON Transcoding at the Edge

The definitive solution is gRPC-JSON Transcoding running inside the Envoy Proxy edge reverse proxy. Envoy natively implements the envoy.filters.http.grpc_json_transcoder filter, which allows Envoy to dynamically translate incoming HTTP/1.1 JSON requests into binary HTTP/2 gRPC calls and translate gRPC responses back into JSON—with zero backend application code.

[Web Browser / Mobile Client]
     │ (HTTP/1.1 JSON: GET /v1/articles/low-and-slow-proxy)
     ▼
┌────────────────────────────────────────────────────────┐
│                      Envoy Proxy                       │
│  [grpc_json_transcoder filter via Compiled Descriptor] │
└──────────────────────────┬─────────────────────────────┘
                           │ (HTTP/2 Binary gRPC: rpc GetArticle)
                           ▼
             [Internal Backend Microservice]
                (Zero REST Code Needed!)

2. Annotating Protobuf Schemas with google.api.http

To configure transcoding, developers decorate their standard .proto service definitions with HTTP mapping annotations defined by Google's API specification:

syntax = "proto3";

package devmanue.blog.v1;

import "google/api/annotations.proto";

service ArticleService {
  // Binds HTTP GET /v1/articles/{slug} to the gRPC GetArticle method
  rpc GetArticle (GetArticleRequest) returns (ArticleResponse) {
    option (google.api.http) = {
      get: "/v1/articles/{slug}"
    };
  }

  // Binds HTTP POST /v1/articles to CreateArticle with JSON body mapping
  rpc CreateArticle (CreateArticleRequest) returns (ArticleResponse) {
    option (google.api.http) = {
      post: "/v1/articles"
      body: "*"
    };
  }
}

message GetArticleRequest {
  string slug = 1;
}

message CreateArticleRequest {
  string title = 1;
  string content = 2;
  repeated string tags = 3;
}

message ArticleResponse {
  string id = 1;
  string title = 2;
  string slug = 3;
  int64 views_count = 4;
}

3. Compiling the Descriptor Set

Envoy does not read raw .proto text files directly. Instead, the Protocol Buffer compiler (protoc) compiles the schema and its dependencies into a binary Protocol Buffer Descriptor Set (.pb file):

protoc -I. -I/usr/local/include   --include_imports   --include_source_info   --descriptor_set_out=proto_descriptors.pb   article_service.proto

4. Envoy Proxy Configuration

Provide the compiled descriptor set to Envoy and activate the transcoding filter in envoy.yaml:

static_resources:
  listeners:
  - name: public_api_listener
    address:
      socket_address: { address: 0.0.0.0, port_value: 8080 }
    filter_chains:
    - filters:
      - name: envoy.filters.network.http_connection_manager
        typed_config:
          "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
          stat_prefix: grpc_json
          route_config:
            name: local_route
            virtual_hosts:
            - name: api_service
              domains: ["*"]
              routes:
              - match: { prefix: "/" }
                route: { cluster: backend_grpc_cluster, timeout: 5s }
          http_filters:
          - name: envoy.filters.http.grpc_json_transcoder
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.filters.http.grpc_json_transcoder.v3.GrpcJsonTranscoder
              proto_descriptor: "/etc/envoy/proto_descriptors.pb"
              services: ["devmanue.blog.v1.ArticleService"]
              print_options:
                add_whitespace: false
                always_print_primitive_fields: true
                preserve_proto_field_names: true
          - name: envoy.filters.http.router
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

With this configuration active, public clients can submit standard JSON to POST /v1/articles. Envoy converts the JSON into binary Protobuf, forwards it over an existing multiplexed HTTP/2 connection to the backend gRPC daemon, receives the binary Protobuf response, and streams it back to the client as valid JSON in under 0.8 milliseconds of proxy overhead.

// High-Throughput Engineering • Systems Architecture Consulting

Scaling Python & Django APIs or Resolving Concurrency Bottlenecks?

We partner with engineering founders and tech leads to architect resilient distributed systems, optimize async worker pools, design scalable databases, and eliminate production latency spikes.

All Insights
Chat on WhatsApp