Skip to content

Reuse Cua-S1 image preprocessing and vision features per request - #17

Open
Levius-Fubuki wants to merge 18 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-image-reuse
Open

Levius-Fubuki wants to merge 18 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-image-reuse

Conversation

@Levius-Fubuki

Copy link
Copy Markdown

Purpose

A request with eight questions over one screenshot currently preprocesses and encodes that image eight times. This change performs image preprocessing and adapted vision encoding once per multi-question request, while preserving independent prompts, candidate order, text embeddings, 3D positions and language forwards. Single-question requests retain the original path. Every prompt is validated before inference; features remain request-local, including failure paths.

On RTX 4090, 17 paired workloads / 3,400 timed requests show 6.8–19.9% lower p50 latency for multi-question cases. Eight distinct questions improve from 1100.03 / 1094.87 ms to 887.54 / 888.45 ms across two runs. Single-question p50 remains within measured noise (−0.5% to +1.1% reduction).

Depends on #12 and #15. This branch includes their commits until they merge; this PR's new work begins after 4c605e3. It is a Transformers/PEFT request-local reuse optimization, not a native CUDA backend or custom kernel.

Full results, raw samples, tensor interface and reproduction commands.

Test Plan

  • Compare exact prepared tensors, language embeddings, attention masks, 3D positions and complete responses against the original execution path. Include distinct questions, candidate permutations, 1–26 candidates, PNG/JPEG, Unicode and consecutive image changes.
  • Alternate baseline/reuse sample order; synchronize each prediction; keep instrumentation outside timed calls; record per-variant memory peaks, source and environment.
  • Rerun the pinned independent upstream oracle, test the actual HTTP worker, and independently recompute sample counts and percentiles from saved JSON.
  • Run the full Cua-S1 test suite, Ruff lint/format checks and failure-recovery tests. Stabilize the existing header-rejection test's body-write race without changing server code.

System1-Omni Version / Commit: timed clean source c18a21b; final GPU-host CPU/oracle/HTTP checks at clean 9f4a4fe. Later commits add result evidence and documentation; measured model and benchmark code are unchanged.

Test Result

  • 104 tests passed on the GPU host; 103 passed, 1 skipped locally (torch absent locally; that tensor test passes remotely). Lint/format checks pass.
  • 13 correctness fixtures / 49 questions per path, plus validation of all 17 timed cases: exact input, embedding, position and full-response parity. Multi-question preprocessing/vision counts are N → 1; language counts remain N.
  • All 9 fresh independent upstream forwards match exactly. Actual HTTP health and eight-question inference return 200 with exact engine-response parity.
  • 3,400 timed requests / 13,600 language forwards; all counts, alternating orders, p50/p95 values and reference matches independently verified. Largest allocated GPU peak is 9.083 GiB baseline / 9.035 GiB reuse.
  • Complete evidence archive copied off the GPU host and SHA-256 verified. Results are synthetic, concurrency-1 engine latency; HTTP throughput, p99, native kernels and Metal are outside this claim.

Copilot AI lite review requested due to automatic review settings September 28, 2026 04:25

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants