Summary
The v2.9.7 gateway Pod intermittently loses readiness and repeatedly times out during scheduler snapshot rebuilds when it shares PostgreSQL/Redis with the primary instance.
Deployment
- Repository:
ranxi2001/sub2api
- Release:
v2.9.7, commit bc83ff9c367883e5b7d0140e6bb42e2e7cc5239c
- Primary: systemd on
freqtrade
- Gateway: k3s
tosky-canary/tosky-nerd, runtime.role=gateway, one Pod
- Shared PostgreSQL and Redis are reached through the existing restricted SSH link.
- Gateway pool settings follow the repository multi Pod guidance: SQL max 16 / idle 4, Redis pool 32 / idle 4.
Observed behavior
On 2026-10-03 Asia/Shanghai:
-
Gateway logs repeatedly reported:
[Scheduler] full rebuild failed: context deadline exceeded
[Scheduler] rebuild failed: bucket=0:antigravity:single ... context deadline exceeded
- outbox lag warnings of 6–11 seconds
content_moderation.runtime_snapshot_refresh_failed ... context deadline exceeded
-
The gateway /readyz endpoint alternated between {"status":"ready"} and {"status":"not_ready"} during repeated read-only probes. The Pod did not restart.
-
The primary /readyz endpoint remained stable.
-
The current /readyz implementation uses a one second shared budget for PostgreSQL and Redis PingContext checks, so it does not expose which dependency failed.
-
A post-observation cache check still found all 39 sampled configured account cost multipliers present and matching in the database, sched:acc, and sched:meta. This confirms the issue is intermittent; it does not prove full rebuild completion.
Routing evidence
The host Nginx configuration routes requests to the primary by default and only selects the gateway when X-Sub2API-Node: nerd is explicitly supplied. Health probes confirmed both branches. No model requests reached the gateway during the observation window, so the readiness problem did not yet affect default production traffic.
Related upstream reports
- Wei-Shaw/sub2api#4772 reports a multi Pod scheduler state problem where one Pod returns
no available accounts while another Pod works with the same shared PostgreSQL/Redis.
- Wei-Shaw/sub2api#1608 documents scheduler outbox watermark timeouts causing repeated event processing, delayed updates, and repeated full rebuilds.
Questions
- Should full rebuild/outbox processing be isolated from the gateway readiness dependency check, or should the shared dependency timeout be configurable for remote PostgreSQL/Redis links?
- Should readiness failures log a dependency-specific reason (without returning it to clients) so an intermittent
503 can be distinguished between PostgreSQL, Redis, and draining?
- Is there a current recommended multi Pod configuration or patch for the outbox/full-rebuild timeout path that should be applied to the owner fork?
No production configuration or data was changed while collecting this evidence.
Summary
The
v2.9.7gateway Pod intermittently loses readiness and repeatedly times out during scheduler snapshot rebuilds when it shares PostgreSQL/Redis with the primary instance.Deployment
ranxi2001/sub2apiv2.9.7, commitbc83ff9c367883e5b7d0140e6bb42e2e7cc5239cfreqtradetosky-canary/tosky-nerd,runtime.role=gateway, one PodObserved behavior
On 2026-10-03 Asia/Shanghai:
Gateway logs repeatedly reported:
[Scheduler] full rebuild failed: context deadline exceeded[Scheduler] rebuild failed: bucket=0:antigravity:single ... context deadline exceededcontent_moderation.runtime_snapshot_refresh_failed ... context deadline exceededThe gateway
/readyzendpoint alternated between{"status":"ready"}and{"status":"not_ready"}during repeated read-only probes. The Pod did not restart.The primary
/readyzendpoint remained stable.The current
/readyzimplementation uses a one second shared budget for PostgreSQL and RedisPingContextchecks, so it does not expose which dependency failed.A post-observation cache check still found all 39 sampled configured account cost multipliers present and matching in the database,
sched:acc, andsched:meta. This confirms the issue is intermittent; it does not prove full rebuild completion.Routing evidence
The host Nginx configuration routes requests to the primary by default and only selects the gateway when
X-Sub2API-Node: nerdis explicitly supplied. Health probes confirmed both branches. No model requests reached the gateway during the observation window, so the readiness problem did not yet affect default production traffic.Related upstream reports
no available accountswhile another Pod works with the same shared PostgreSQL/Redis.Questions
503can be distinguished between PostgreSQL, Redis, and draining?No production configuration or data was changed while collecting this evidence.