Health
The service exposes two endpoints accessible via HTTP on theHTTP_PORT (default 8080):
GET /livez
Returns 200 if the service is alive.
GET /readyz
Returns 200 if the service is ready.
Healthy response for both endpoints:
JSON
A Sidecar liveness probe should not trigger automatic Kubernetes pod restarts for network connectivity issues, Stigg API unreachability, or Edge failures. This is an intentional design choice to prevent the Sidecar from entering a restart loop, which would disrupt the main application whenever the upstream Stigg API is temporarily unreachable. The recommended approach is to keep the Sidecar pod alive and rely on cached reads and fail-safe fallback modes while surfacing connectivity or write errors into your own observability and alerting stack, rather than coupling them to pod restarts.
Metrics
The Sidecar exposes aGET /metrics endpoint that returns service metrics in Prometheus format. This endpoint includes both system-level and Sidecar-specific metrics.
Sidecar
sidecar_initialization_errors_total- errors during Sidecar initialization, such as an invalid API key or misconfiguration. It should always remain at 0.sidecar_invalid_api_key_errors_total- authentication failures.sidecar_network_request_errors_total- connectivity problems, increments on API errors. A rising value indicates the Stigg API may be unreachable.sidecar_redis_client_errors_total- Redis connection issues from the Sidecar.sidecar_cache_hits_totalandsidecar_cache_misses_total- cache performance.sidecar_write_through_errors_total- REST API writes whose cache update failed, labeled by route.
Persistent cache service
When Redis is enabled, the persistent cache service provides its own monitoring endpoints:GET /livez, GET /readyz, and GET /metrics. Important metrics include:
persistent_cache_write_duration_seconds- write performance. High duration can indicate a delay in propagation of changes to Redis.persistent_cache_write_errors_total- write failures, a signal of connectivity or stability issues with Redis.persistent_cache_hit_ratio- overall cluster hit rate, a good indicator of how well Redis is being utilized.persistent_cache_memory_usage_bytes- memory consumption.persistent_cache_hits_totalandpersistent_cache_misses_total- cache effectiveness.persistent_cache_messages_processed_total- throughput tracking.
Logging
The Sidecar supports configurable log levels via theLOG_LEVEL environment variable (error|warn|info|debug). Forward these logs to your logging stack and configure alerts for error-level messages, as they typically indicate issues requiring immediate attention.
Write paths, such as POST /api/v1/usage and POST /api/v1/events in the REST API or ReportUsage and ReportEvents over gRPC, return errors explicitly rather than fallback values. Wrap all SDK and API calls with try/catch blocks and log the errors.
Recommended alerts
These thresholds can be adjusted based on your production utilization patterns and SLOs. Warnings are typically routed to the application team via Slack for investigation, while critical alerts should trigger PagerDuty notifications to the on-call engineer.
Also consider alerting on:
- Readiness:
/readyzreturns a non-UP status for more than a few minutes, especially when combined with/livezfailures. - Sidecar errors: growth in any
sidecar_*_errors_totalmetric, includingsidecar_invalid_api_key_errors_total. These patterns often indicate underlying issues with connectivity, configuration, or cache effectiveness. - Sidecar cache hit ratio: the ratio of
sidecar_cache_hits_totalto total lookups falls below 70% for 15 minutes. This suggests cache configuration issues or unusual access patterns. - API error rate: the error rate exceeds 1% sustained for 15 minutes. This indicates potential problems with API connectivity or request validity that warrant investigation.
- Persistent cache consumers: the consumer count reported by the persistent cache service’s
/readyzdrops. This can indicate Redis connectivity problems or insufficient consumer capacity. - Upstream status: subscribe to the Stigg Status Page to be notified about platform-wide incidents.
- Entitlement latency (p95): measure latency at your application layer or APM around Get Entitlement and Get Entitlements calls, and alert according to your SLO. The architecture is designed for low-latency reads through pre-computed, distributed caching and edge delivery.
Auto-scaling
For auto-scaling, monitor the service’s CPU and memory metrics:Troubleshooting
Sidecar startup problems
Startup issues typically show up as a risingsidecar_initialization_errors_total metric and logs that mention invalid API keys or network connectivity problems.
Check the health and readiness endpoints (GET /livez and GET /readyz), validate that your SERVER_API_KEY environment variable is correct and active, verify network egress to the Stigg API and Edge endpoints, and confirm you’re running a supported Sidecar image version.
While the problem persists, the Sidecar keeps serving from the persistent cache if configured, or from global fallback values.
Elevated API failures
Signs of API problems include increased non-2xx responses from Stigg APIs and a climbingsidecar_network_request_errors_total metric. Entitlement check responses containing isFallback: true indicate fallback values are being used.
Correlate the timing with the Stigg Status Page to see if there’s a platform-wide issue, and verify network connectivity from your services to Stigg. Make sure global and per-check fallback values are configured for critical entitlement paths so your application keeps functioning during outages. Notify the Stigg team if the issue persists.
Sidecar entitlement checks are served from local or persistent cache. On cache miss, the Sidecar queries the Edge API. The Entitlements endpoint itself is unlimited, so entitlement lookups are effectively not rate-limited in practice. If you require higher limits for other operations, contact Stigg Support.
Redis and persistent cache issues
Redis problems show up as a risingsidecar_redis_client_errors_total metric, along with persistent cache metrics showing write errors or a declining hit ratio.
Check the persistent cache service health with GET /readyz to verify consumer and Redis state, and review GET /metrics for persistent_cache_* metrics to understand what’s failing. Double-check your Redis connection parameters, including host, port, authentication, TLS settings, and database selection. Make sure the environment prefix is aligned across your SDK, Sidecar, and persistent cache service, as mismatches are a common source of issues.