Skip to main content

Deployment models

The Sidecar service can be deployed in two primary ways:
  1. Sidecar pattern: Each application instance has a Sidecar container in the same network namespace (for example, the same Kubernetes Pod), and sends REST API or gRPC requests to it locally.
  2. Standalone service: The Sidecar runs as a shared service that multiple applications can access over an exposed port within your private network.
The Sidecar is not intended to be exposed to the internet or to be accessed from a browser.
Sidecar Memory

Sidecar running with in-memory cache

Caching

When entitlement data is fetched from the Stigg API, it is stored in the configured cache:
  • In-memory cache (default): managed directly by the Sidecar
  • Redis cache: managed together with a Persistent Cache Service
Caching behavior:
  • The Sidecar uses an in-memory LRU cache capped by size.
  • You can control the cache size with the CACHE_MAX_SIZE_BYTES environment variable.
  • If not set, the Sidecar allocates up to 50% of the total available memory for the cache.
  • On a cache miss, the Sidecar fetches data from the Stigg API and updates the cache before serving future requests.
  • Only entitlements and current usage are cached. Subscriptions are not part of it.
If you enable Redis, an instance of the Persistent Cache Service must be deployed to keep Redis up to date with entitlements and usage data, as illustrated below. See persistent cache for configuration.
Sidecar Redis

Sidecar running with an external Redis cache

Redis-backed caching is particularly useful when:
  • You run in serverless environments (for example, AWS Lambda) where processes are frequently terminated
  • You have a large fleet of containers and want fresh entitlements and usage available immediately when new instances start
  • You want cached entitlements and usage to survive restarts and be shared across instances
If you do not configure Redis, the Sidecar works “as is” with its default in-memory cache. No extra services are required.

Scaling

The Sidecar is stateless and scales horizontally based on load, whether you run it as a sidecar container per pod or as a shared service. When using only in-memory cache, adding more replicas may increase the cache-miss ratio, since each instance maintains its own independent cache. For improved hit ratios and resilience, use Redis with the persistent cache service. The persistent cache service consumes a Stigg-managed SQS queue to keep Redis up to date, and is also stateless and horizontally scalable. For auto-scaling triggers, see monitoring.

Benchmarks

Based on internal benchmarks, a single instance of the Sidecar service, connected to a Redis cluster as the cache layer, can handle:
  • Throughput: 1200 requests per second (RPS) steady.
  • Latency: p95 under 10ms.
  • CPU and memory: 1 vCPU, 512MB RAM per instance is sufficient.
  • CPU usage: 70% average, commonly between 40%-75% under constant load.
  • Memory footprint: ~250MB per instance.