Skip to content

Serve Cua-S1 multimodal requests with bounded segmented CUDA Graphs - #22

Open
Levius-Fubuki wants to merge 27 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-graph-runtime
Open

Levius-Fubuki wants to merge 27 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-graph-runtime

Conversation

@Levius-Fubuki

Copy link
Copy Markdown

Purpose

The whole-language CUDA Graph experiment in #20 reduced latency but changed candidate probabilities by up to 0.065181 on eight distinct questions and 0.051902 after a long-prompt text change. Per-layer tracing on the pinned RTX 4090 stack found the first difference in the first full-attention layer. This PR adds an opt-in Graph path to the multimodal worker that captures each contiguous Gated DeltaNet run while keeping full attention eager. It updates static inputs on every replay, checks a newly captured layout against all eager vocabulary logits, and manages an LRU cache with layout, memory, and token limits plus eager fallback.

Start the worker with --graph to enable it. The default cache holds at most eight layouts and 1 GiB of measured live Graph allocations; prompts over 2,048 tokens take the eager path unless --graph-max-tokens is raised. The existing request lock prevents overlapping replays. This PR builds on #20 and therefore includes its #18/#17/#12 dependency stack until those PRs merge.

Results

Measured from clean source commit de000560ce927a0db58f4928e91234d95df2ae23 on one RTX 4090 with the pinned BF16 base and unmerged multimodal LoRA. Five cases, two runs of 20 calls per variant and case (400 synchronized engine.predict timings total). Original-image, changed-image, and same-layout changed-text candidate probabilities and choices matched eager exactly in all five cases, including both cases rejected in #20.

Case Eager p50 ms, runs 1 / 2 Graph p50 ms, runs 1 / 2 Reduction
320×240, 2 repeated questions 189.57 / 192.80 102.74 / 102.99 45.8% / 46.6%
320×240, 8 repeated questions 702.32 / 701.88 349.60 / 349.04 50.2% / 50.3%
640×480, 8 repeated questions 859.35 / 852.85 576.29 / 575.22 32.9% / 32.6%
640×480, 8 long-instruction questions 1024.08 / 1026.02 971.69 / 971.54 5.1% / 5.3%
640×480, 8 distinct questions 863.12 / 824.02 597.37 / 591.87 30.8% / 28.2%

First capture cost 329–618 ms for a repeated-question layout; the distinct-question case captured seven layouts and cost 3,493 ms in total. Those costs and the eager correctness gate are outside the timed samples. Retained Graph allocations were 92–287 MiB for these cases. This is serial engine latency, not HTTP throughput or concurrency.

Validation

  • 116 Cua-S1 tests passed on the GPU host; Ruff lint and format checks passed.
  • One-layout LRU eviction and 1 MiB budget fallback preserved candidate probabilities and enforced cache limits.
  • The --graph HTTP worker returned a ready health response and byte-identical JSON for two repeated two-question requests.
  • The independent verifier recomputed all 400 timings and checked correctness gates and cache behavior. See report, raw samples, cache checks, and reproduction steps.

Copilot AI lite review requested due to automatic review settings September 28, 2026 08:24

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants