Skip to content

Limit Cua-S1 reused-request output projection to final token - #18

Open
Levius-Fubuki wants to merge 21 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-post-reuse-profile
Open

Levius-Fubuki wants to merge 21 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-post-reuse-profile

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Sep 28, 2026 •

Copy link
Copy Markdown

Purpose

After request-local image reuse, each Cua-S1 question still projects every language position to the 248,320-word vocabulary, although candidate scoring reads only the final position. Pass logits_to_keep=1 on the reused multi-question path. The original single-question and full-projection reference paths remain available for exact comparison.

This branch depends on #12, #15 and #17 and contains their commits until they merge. The new implementation begins after #17's 95713a5. This is a Transformers/PEFT optimization; it does not add a native engine or custom CUDA kernel.

Test Plan

  • Verify a failing-then-passing tensor integration test for the final-token model argument; rerun the full Cua-S1 suite and repository CI lint/format commands.
  • Compare the full and final-token projections using the same reuse path. Require identical complete responses, then alternate variant order across the original 16-case matrix plus eight distinct questions. Record individual synchronized latencies, p50/p95, output-head shapes, source/environment and peak allocated memory.
  • Rerun the 13-fixture reference-versus-reuse correctness sweep and an actual eight-question HTTP worker request. Independently audit saved samples and hashes.

System1-Omni Version / Commit: clean measured source 224d88b; published evidence 89e35e9; completion note c522e00. The model and benchmark code did not change between them.

Test Result

  • GPU host: 105 Cua-S1 tests passed. Local macOS: 104 passed, 1 skipped (torch-dependent tensor test runs on the GPU host). CI-equivalent Ruff lint and formatting passed.
  • 17 cases / 2,040 timed predictions on RTX 4090. Every multi-question case improved in both runs: p50 reduction 0.5–5.3%, median 3.4% across runs. Eight distinct questions improved 852.40 → 822.68 ms and 881.26 → 851.65 ms. Four single-question controls stayed within −0.4% to +0.2%.
  • Full responses were exactly equal in all 17 timed cases. The 13-fixture sweep also found exact prepared tensors, language embeddings, 3D positions, candidate probabilities and responses across 49 question forwards per path. Actual HTTP health and eight-question inference returned 200, with HTTP response exactly equal to direct engine output.
  • For 640×480 long instructions and eight questions, peak allocated memory fell 9.034 → 8.840 GiB. The output-head input shrank from [1, 834, 2560] to [1, 1, 2560] for each question.

Full method, all raw samples, profiler groups, audit script and reproduction commands. The complete 25 MB evidence archive, including three large Chrome traces and the source bundle, was copied off the GPU host and SHA-256 verified before shutdown.

The experiment server was shut down after publication; a subsequent SSH connection was refused.

These are synthetic, concurrency-1 engine.predict measurements. They do not measure HTTP latency, throughput, p99, native kernels or Metal.

Copilot AI lite review requested due to automatic review settings September 28, 2026 06:42

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants