Limit Cua-S1 reused-request output projection to final token - #18
Open
Levius-Fubuki wants to merge 21 commits into
Open
Levius-Fubuki wants to merge 21 commits into
Levius-Fubuki wants to merge 21 commits into
Conversation
This was referenced Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
After request-local image reuse, each Cua-S1 question still projects every language position to the 248,320-word vocabulary, although candidate scoring reads only the final position. Pass
logits_to_keep=1on the reused multi-question path. The original single-question and full-projection reference paths remain available for exact comparison.This branch depends on #12, #15 and #17 and contains their commits until they merge. The new implementation begins after #17's
95713a5. This is a Transformers/PEFT optimization; it does not add a native engine or custom CUDA kernel.Test Plan
System1-Omni Version / Commit: clean measured source
224d88b; published evidence89e35e9; completion notec522e00. The model and benchmark code did not change between them.Test Result
[1, 834, 2560]to[1, 1, 2560]for each question.Full method, all raw samples, profiler groups, audit script and reproduction commands. The complete 25 MB evidence archive, including three large Chrome traces and the source bundle, was copied off the GPU host and SHA-256 verified before shutdown.
The experiment server was shut down after publication; a subsequent SSH connection was refused.
These are synthetic, concurrency-1
engine.predictmeasurements. They do not measure HTTP latency, throughput, p99, native kernels or Metal.