Skip to content

[production] gateway readiness flaps while scheduler full rebuild times out in multi-Pod deployment #279

Description

@ranxi2001

Summary

The v2.9.7 gateway Pod intermittently loses readiness and repeatedly times out during scheduler snapshot rebuilds when it shares PostgreSQL/Redis with the primary instance.

Deployment

  • Repository: ranxi2001/sub2api
  • Release: v2.9.7, commit bc83ff9c367883e5b7d0140e6bb42e2e7cc5239c
  • Primary: systemd on freqtrade
  • Gateway: k3s tosky-canary/tosky-nerd, runtime.role=gateway, one Pod
  • Shared PostgreSQL and Redis are reached through the existing restricted SSH link.
  • Gateway pool settings follow the repository multi Pod guidance: SQL max 16 / idle 4, Redis pool 32 / idle 4.

Observed behavior

On 2026-10-03 Asia/Shanghai:

  • Gateway logs repeatedly reported:

    • [Scheduler] full rebuild failed: context deadline exceeded
    • [Scheduler] rebuild failed: bucket=0:antigravity:single ... context deadline exceeded
    • outbox lag warnings of 6–11 seconds
    • content_moderation.runtime_snapshot_refresh_failed ... context deadline exceeded
  • The gateway /readyz endpoint alternated between {"status":"ready"} and {"status":"not_ready"} during repeated read-only probes. The Pod did not restart.

  • The primary /readyz endpoint remained stable.

  • The current /readyz implementation uses a one second shared budget for PostgreSQL and Redis PingContext checks, so it does not expose which dependency failed.

  • A post-observation cache check still found all 39 sampled configured account cost multipliers present and matching in the database, sched:acc, and sched:meta. This confirms the issue is intermittent; it does not prove full rebuild completion.

Routing evidence

The host Nginx configuration routes requests to the primary by default and only selects the gateway when X-Sub2API-Node: nerd is explicitly supplied. Health probes confirmed both branches. No model requests reached the gateway during the observation window, so the readiness problem did not yet affect default production traffic.

Related upstream reports

  • Wei-Shaw/sub2api#4772 reports a multi Pod scheduler state problem where one Pod returns no available accounts while another Pod works with the same shared PostgreSQL/Redis.
  • Wei-Shaw/sub2api#1608 documents scheduler outbox watermark timeouts causing repeated event processing, delayed updates, and repeated full rebuilds.

Questions

  1. Should full rebuild/outbox processing be isolated from the gateway readiness dependency check, or should the shared dependency timeout be configurable for remote PostgreSQL/Redis links?
  2. Should readiness failures log a dependency-specific reason (without returning it to clients) so an intermittent 503 can be distinguished between PostgreSQL, Redis, and draining?
  3. Is there a current recommended multi Pod configuration or patch for the outbox/full-rebuild timeout path that should be applied to the owner fork?

No production configuration or data was changed while collecting this evidence.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions