Skip to content

Measure CUDA Graph replay on Cua-S1 multimodal language forward - #20

Open
Levius-Fubuki wants to merge 24 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-multimodal-graph
Open

Levius-Fubuki wants to merge 24 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-multimodal-graph

Conversation

@Levius-Fubuki

Copy link
Copy Markdown

Purpose

After request-local image reuse and final-token projection, each Cua-S1 multimodal question still submits the full language forward through eager PyTorch. This PR adds a bounded CUDA Graph experiment and publishes raw evidence, including cases where replay fails the numerical gate. It does not enable Graph replay in the serving worker.

The experiment captures the multimodal language forward after vision features, embeddings and three-dimensional positions have been prepared. It replays through fixed input buffers, checks tensor layout, and compares against the original reused worker. It also changes the image and same-length question text between replays to detect stale inputs. Source is stacked on #18 (and its #12, #15 and #17 dependencies); the new work starts after c522e00.

Results

One RTX 4090, BF16 base with unmerged multimodal LoRA, serial synthetic requests. Three short-instruction cases passed the selected-choice and 0.002 absolute candidate-probability gate. Across two runs of 20 calls per variant and case (240 timed predictions), synchronized engine.predict p50 changed as follows:

Case Eager p50 ms, runs 1 / 2 Graph p50 ms, runs 1 / 2 Reduction
320×240, 2 repeated questions 181.42 / 183.28 100.31 / 100.38 44.7% / 45.2%
320×240, 8 repeated questions 669.81 / 684.26 337.57 / 338.21 49.6% / 50.6%
640×480, 8 repeated questions 825.78 / 786.77 562.90 / 561.72 31.8% / 28.6%

The eager variant returned the original worker response exactly. Graph's maximum candidate-probability difference in these cases was 0.000421 on the original inputs. First capture plus three warmups cost 347–416 ms per shape.

Two other cases were rejected before timing: eight distinct questions had a 0.065181 probability difference; a long-instruction case differed by 0.051902 after a same-length text change. Choices happened to remain the same, but this is insufficient for production use. The cause remains to be isolated. These results do not cover the native worker in #19, arbitrary shapes, concurrency or HTTP throughput.

Validation

  • 108 Cua-S1 GPU-host tests passed; Ruff lint and formatting checks passed.
  • An independent verifier recomputed all 240 saved timings and checked both rejected cases.
  • Report, raw samples, failure reports and reproduction steps.
  • The 186,201-byte complete evidence archive, including fixtures, error logs, package versions and source bundle, was copied off the GPU host. Server and local SHA-256 matched: 11831981ae8ea1c414075d563284b40b378a3471a8164c26b9a281986a12bee0.

Measured source commit: 809cacc70979b1b4e3a55b8b89be4691d09b0ce4 (clean checkout). The subsequent commit publishes evidence only.

Copilot AI lite review requested due to automatic review settings September 28, 2026 07:31

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants