Skip to main content

Health

The service exposes two endpoints accessible via HTTP on the HTTP_PORT (default 8080): GET /livez Returns 200 if the service is alive. GET /readyz Returns 200 if the service is ready. Healthy response for both endpoints:
JSON
A Sidecar liveness probe should not trigger automatic Kubernetes pod restarts for network connectivity issues, Stigg API unreachability, or Edge failures. This is an intentional design choice to prevent the Sidecar from entering a restart loop, which would disrupt the main application whenever the upstream Stigg API is temporarily unreachable. The recommended approach is to keep the Sidecar pod alive and rely on cached reads and fail-safe fallback modes while surfacing connectivity or write errors into your own observability and alerting stack, rather than coupling them to pod restarts.

Metrics

The Sidecar exposes a GET /metrics endpoint that returns service metrics in Prometheus format. This endpoint includes both system-level and Sidecar-specific metrics.

Sidecar

  • sidecar_initialization_errors_total - errors during Sidecar initialization, such as an invalid API key or misconfiguration. It should always remain at 0.
  • sidecar_invalid_api_key_errors_total - authentication failures.
  • sidecar_network_request_errors_total - connectivity problems, increments on API errors. A rising value indicates the Stigg API may be unreachable.
  • sidecar_redis_client_errors_total - Redis connection issues from the Sidecar.
  • sidecar_cache_hits_total and sidecar_cache_misses_total - cache performance.
  • sidecar_write_through_errors_total - REST API writes whose cache update failed, labeled by route.

Persistent cache service

When Redis is enabled, the persistent cache service provides its own monitoring endpoints: GET /livez, GET /readyz, and GET /metrics. Important metrics include:
  • persistent_cache_write_duration_seconds - write performance. High duration can indicate a delay in propagation of changes to Redis.
  • persistent_cache_write_errors_total - write failures, a signal of connectivity or stability issues with Redis.
  • persistent_cache_hit_ratio - overall cluster hit rate, a good indicator of how well Redis is being utilized.
  • persistent_cache_memory_usage_bytes - memory consumption.
  • persistent_cache_hits_total and persistent_cache_misses_total - cache effectiveness.
  • persistent_cache_messages_processed_total - throughput tracking.

Logging

The Sidecar supports configurable log levels via the LOG_LEVEL environment variable (error|warn|info|debug). Forward these logs to your logging stack and configure alerts for error-level messages, as they typically indicate issues requiring immediate attention. Write paths, such as POST /api/v1/usage and POST /api/v1/events in the REST API or ReportUsage and ReportEvents over gRPC, return errors explicitly rather than fallback values. Wrap all SDK and API calls with try/catch blocks and log the errors. These thresholds can be adjusted based on your production utilization patterns and SLOs. Warnings are typically routed to the application team via Slack for investigation, while critical alerts should trigger PagerDuty notifications to the on-call engineer. Also consider alerting on:
  • Readiness: /readyz returns a non-UP status for more than a few minutes, especially when combined with /livez failures.
  • Sidecar errors: growth in any sidecar_*_errors_total metric, including sidecar_invalid_api_key_errors_total. These patterns often indicate underlying issues with connectivity, configuration, or cache effectiveness.
  • Sidecar cache hit ratio: the ratio of sidecar_cache_hits_total to total lookups falls below 70% for 15 minutes. This suggests cache configuration issues or unusual access patterns.
  • API error rate: the error rate exceeds 1% sustained for 15 minutes. This indicates potential problems with API connectivity or request validity that warrant investigation.
  • Persistent cache consumers: the consumer count reported by the persistent cache service’s /readyz drops. This can indicate Redis connectivity problems or insufficient consumer capacity.
  • Upstream status: subscribe to the Stigg Status Page to be notified about platform-wide incidents.
  • Entitlement latency (p95): measure latency at your application layer or APM around Get Entitlement and Get Entitlements calls, and alert according to your SLO. The architecture is designed for low-latency reads through pre-computed, distributed caching and edge delivery.

Auto-scaling

For auto-scaling, monitor the service’s CPU and memory metrics:

Troubleshooting

Sidecar startup problems

Startup issues typically show up as a rising sidecar_initialization_errors_total metric and logs that mention invalid API keys or network connectivity problems. Check the health and readiness endpoints (GET /livez and GET /readyz), validate that your SERVER_API_KEY environment variable is correct and active, verify network egress to the Stigg API and Edge endpoints, and confirm you’re running a supported Sidecar image version. While the problem persists, the Sidecar keeps serving from the persistent cache if configured, or from global fallback values.

Elevated API failures

Signs of API problems include increased non-2xx responses from Stigg APIs and a climbing sidecar_network_request_errors_total metric. Entitlement check responses containing isFallback: true indicate fallback values are being used. Correlate the timing with the Stigg Status Page to see if there’s a platform-wide issue, and verify network connectivity from your services to Stigg. Make sure global and per-check fallback values are configured for critical entitlement paths so your application keeps functioning during outages. Notify the Stigg team if the issue persists.
Sidecar entitlement checks are served from local or persistent cache. On cache miss, the Sidecar queries the Edge API. The Entitlements endpoint itself is unlimited, so entitlement lookups are effectively not rate-limited in practice. If you require higher limits for other operations, contact Stigg Support.

Redis and persistent cache issues

Redis problems show up as a rising sidecar_redis_client_errors_total metric, along with persistent cache metrics showing write errors or a declining hit ratio. Check the persistent cache service health with GET /readyz to verify consumer and Redis state, and review GET /metrics for persistent_cache_* metrics to understand what’s failing. Double-check your Redis connection parameters, including host, port, authentication, TLS settings, and database selection. Make sure the environment prefix is aligned across your SDK, Sidecar, and persistent cache service, as mismatches are a common source of issues.