Deployment models
The Sidecar service can be deployed in two primary ways:- Sidecar pattern: Each application instance has a Sidecar container in the same network namespace (for example, the same Kubernetes Pod), and sends REST API or gRPC requests to it locally.
- Standalone service: The Sidecar runs as a shared service that multiple applications can access over an exposed port within your private network.

Sidecar running with in-memory cache
Caching
When entitlement data is fetched from the Stigg API, it is stored in the configured cache:- In-memory cache (default): managed directly by the Sidecar
- Redis cache: managed together with a Persistent Cache Service
- The Sidecar uses an in-memory LRU cache capped by size.
- You can control the cache size with the
CACHE_MAX_SIZE_BYTESenvironment variable. - If not set, the Sidecar allocates up to
50%of the total available memory for the cache. - On a cache miss, the Sidecar fetches data from the Stigg API and updates the cache before serving future requests.
- Only entitlements and current usage are cached. Subscriptions are not part of it.

Sidecar running with an external Redis cache
- You run in serverless environments (for example, AWS Lambda) where processes are frequently terminated
- You have a large fleet of containers and want fresh entitlements and usage available immediately when new instances start
- You want cached entitlements and usage to survive restarts and be shared across instances
Scaling
The Sidecar is stateless and scales horizontally based on load, whether you run it as a sidecar container per pod or as a shared service. When using only in-memory cache, adding more replicas may increase the cache-miss ratio, since each instance maintains its own independent cache. For improved hit ratios and resilience, use Redis with the persistent cache service. The persistent cache service consumes a Stigg-managed SQS queue to keep Redis up to date, and is also stateless and horizontally scalable. For auto-scaling triggers, see monitoring.Benchmarks
Based on internal benchmarks, a single instance of the Sidecar service, connected to a Redis cluster as the cache layer, can handle:- Throughput: 1200 requests per second (RPS) steady.
- Latency: p95 under 10ms.
- CPU and memory: 1 vCPU, 512MB RAM per instance is sufficient.
- CPU usage: 70% average, commonly between 40%-75% under constant load.
- Memory footprint: ~250MB per instance.