Serve Cua-S1 multimodal requests with bounded segmented CUDA Graphs - #22
Open
Levius-Fubuki wants to merge 27 commits into
Open
Levius-Fubuki wants to merge 27 commits into
Levius-Fubuki wants to merge 27 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
The whole-language CUDA Graph experiment in #20 reduced latency but changed candidate probabilities by up to 0.065181 on eight distinct questions and 0.051902 after a long-prompt text change. Per-layer tracing on the pinned RTX 4090 stack found the first difference in the first full-attention layer. This PR adds an opt-in Graph path to the multimodal worker that captures each contiguous Gated DeltaNet run while keeping full attention eager. It updates static inputs on every replay, checks a newly captured layout against all eager vocabulary logits, and manages an LRU cache with layout, memory, and token limits plus eager fallback.
Start the worker with
--graphto enable it. The default cache holds at most eight layouts and 1 GiB of measured live Graph allocations; prompts over 2,048 tokens take the eager path unless--graph-max-tokensis raised. The existing request lock prevents overlapping replays. This PR builds on #20 and therefore includes its #18/#17/#12 dependency stack until those PRs merge.Results
Measured from clean source commit
de000560ce927a0db58f4928e91234d95df2ae23on one RTX 4090 with the pinned BF16 base and unmerged multimodal LoRA. Five cases, two runs of 20 calls per variant and case (400 synchronizedengine.predicttimings total). Original-image, changed-image, and same-layout changed-text candidate probabilities and choices matched eager exactly in all five cases, including both cases rejected in #20.First capture cost 329–618 ms for a repeated-question layout; the distinct-question case captured seven layouts and cost 3,493 ms in total. Those costs and the eager correctness gate are outside the timed samples. Retained Graph allocations were 92–287 MiB for these cases. This is serial engine latency, not HTTP throughput or concurrency.
Validation
--graphHTTP worker returned a ready health response and byte-identical JSON for two repeated two-question requests.