Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
27 commits
Select commit Hold shift + click to select a range
74c95e5
Add Cua-S1 multimodal CUDA worker and parity recipe
Levius-Fubuki Sep 27, 2026
ee970c2
Merge main and align multimodal worker error contract
Levius-Fubuki Sep 28, 2026
767e211
Add reproducible multimodal profiling matrix
Levius-Fubuki Sep 28, 2026
ebd7adc
Distinguish profiler device annotations and support trace-only runs
Levius-Fubuki Sep 28, 2026
4c605e3
Record RTX 4090 profiling results and trace recovery
Levius-Fubuki Sep 28, 2026
159789a
docs: plan request-local image reuse experiment
Levius-Fubuki Sep 28, 2026
8c3aa20
test: add paired image reuse experiment and correctness checks
Levius-Fubuki Sep 28, 2026
7e7ed34
test: preserve unoptimized profiling baseline after reuse
Levius-Fubuki Sep 28, 2026
7874c6c
style: simplify profiling test environment stub
Levius-Fubuki Sep 28, 2026
cf899ff
Reuse image preprocessing and adapted vision features within requests
Levius-Fubuki Sep 28, 2026
c18a21b
test: cover multi-question JPEG reuse and document tensor contract
Levius-Fubuki Sep 28, 2026
4bb4035
Test image reuse cleanup after inference failures
Levius-Fubuki Sep 28, 2026
0abf31d
Avoid body-write race in transport rejection tests
Levius-Fubuki Sep 28, 2026
d0c4781
docs: link request reuse behavior and record completed reviews
Levius-Fubuki Sep 28, 2026
209a660
test: add reproducible result audit and CUDA HTTP postflight
Levius-Fubuki Sep 28, 2026
9f4a4fe
test: audit exact coverage of the paired workload matrix
Levius-Fubuki Sep 28, 2026
ccaa4fb
docs: record paired RTX 4090 image reuse results and exact parity
Levius-Fubuki Sep 28, 2026
95713a5
docs: record PR publication and verified experiment server shutdown
Levius-Fubuki Sep 28, 2026
224d88b
perf(cua-s1): project only final logits in reused path
Levius-Fubuki Sep 28, 2026
89e35e9
docs(cua-s1): publish RTX 4090 final-logits experiment
Levius-Fubuki Sep 28, 2026
c522e00
docs(cua-s1): record experiment publication and server shutdown
Levius-Fubuki Sep 28, 2026
a2abe7e
bench(cua-s1): measure multimodal CUDA Graph language forward
Levius-Fubuki Sep 28, 2026
809cacc
bench(cua-s1): validate graph replay with changed question text
Levius-Fubuki Sep 28, 2026
74b8c9e
docs(cua-s1): publish multimodal CUDA Graph evidence
Levius-Fubuki Sep 28, 2026
de00056
feat(cua-s1): add bounded segmented CUDA Graph runtime
Levius-Fubuki Sep 28, 2026
03c3acc
docs(cua-s1): publish segmented Graph runtime results
Levius-Fubuki Sep 28, 2026
9a818a7
docs(cua-s1): mark Graph runtime rollout complete
Levius-Fubuki Sep 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 30 additions & 0 deletions .github/workflows/cua-s1-multimodal.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
name: Cua-S1 multimodal CPU checks
on:
pull_request:
paths:
- 'src/models/cua_s1/multimodal/**'
- 'tests/cua_s1/**'
- 'recipe/cua_s1/**'
- '.github/workflows/cua-s1-multimodal.yml'
push:
branches: [main]
paths:
- 'src/models/cua_s1/multimodal/**'
- 'tests/cua_s1/**'
- 'recipe/cua_s1/**'
- '.github/workflows/cua-s1-multimodal.yml'
permissions:
contents: read
jobs:
cpu:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python -m pip install Pillow==11.3.0 pytest==9.1.1 ruff==0.16.8
- run: PYTHONPATH=src python -m pytest tests/cua_s1 -q
- run: ruff check --isolated --select E4,E7,E9,F,I src/models/cua_s1/multimodal tests/cua_s1 recipe/cua_s1
- run: ruff format --isolated --check src/models/cua_s1/multimodal tests/cua_s1 recipe/cua_s1
9 changes: 7 additions & 2 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,12 @@ target
# option (not recommended) you can uncomment the following to ignore the entire idea folder.
#.idea/

# Local Python worker environment
.venv/
# Python workers and local model artifacts
__pycache__/
.pytest_cache/
.ruff_cache/
.venv/
weights/

# macOS metadata
.DS_Store
11 changes: 7 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

A community-maintained inference engine for prefill-only System1-Omni models, designed around a Rust frontend, model-owned execution, and high-performance CUDA and Metal backends.

The Rust frontend forwards requests to a separately running model worker. In-repository model engines and GPU backends are not implemented yet.
The Rust frontend forwards requests to a separately running model worker. The Cua-S1 multimodal worker has been validated on CUDA through Transformers and PEFT; native GPU backends remain planned.

## Run the frontend

Expand All @@ -17,7 +17,8 @@ OMNI_JEV_BACKEND_URL=http://127.0.0.1:8000 \

Start the worker separately. See the [frontend documentation](src/frontend/README.md)
for the HTTP interface and configuration, or the [Laya recipe](recipe/laya/README.md)
for a CPU text worker and response checks.
for a CPU text worker and response checks. The [Cua-S1 recipe](recipe/cua_s1/README.md)
provides a CUDA worker for screenshot-conditioned choices.

## Architecture

Expand All @@ -42,20 +43,22 @@ Implementation code lives under `src/`; recipes and documentation stay at the re
| --- | --- |
| [`src/frontend/`](src/frontend/) | Rust serving code and the small engine interface. |
| [`src/models/laya/`](src/models/laya/) | LAYA preprocessing, batching, state, execution, and output processing. |
| [`src/models/cua_s1/multimodal/`](src/models/cua_s1/multimodal/) | Cua-S1 screenshot preprocessing, multimodal LoRA execution, and choice probabilities. |
| [`src/backends/cuda/`](src/backends/cuda/) | NVIDIA GPU operations and kernel integration. |
| [`src/backends/metal/`](src/backends/metal/) | Apple GPU operations and kernel integration. |
| [`recipe/`](recipe/) | Model setup instructions, launch commands, configuration examples, and example requests. |
| [`docs/`](docs/) | Project documentation and architecture assets. |

The frontend is a Cargo workspace member. Model and backend directories currently document planned work; they do not prescribe process boundaries.
The frontend is a Cargo workspace member. Model and backend directories document ownership, with implementations added incrementally; they do not prescribe process boundaries.

## Supported models

LAYA can run as an external Python worker for text requests. Its in-repository model engine is still planned:
Validated coverage is listed by modality and execution path:

| Model | Status |
| --- | --- |
| LAYA | [External worker](recipe/laya/README.md); model engine planned |
| [Cua-S1 4B 0.2](recipe/cua_s1/README.md) | Multimodal screenshot choices via Transformers/PEFT on RTX 4090 CUDA; [parity and measurements](recipe/cua_s1/experiments/README.md). Text serving, native CUDA kernels and Metal deferred. |

CUDA and Metal coverage will be documented per model as implementations are added and validated.

Expand Down
45 changes: 45 additions & 0 deletions docs/superpowers/plans/2026-09-28-cua-graph-runtime.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
# Cua-S1 Multimodal Graph Runtime Implementation Plan

> Inline execution in this task; the user authorized implementation, GPU experiments, PR submission, and server shutdown.

**Goal:** Serve multimodal language forwards through reusable CUDA Graphs while preserving the existing worker's candidate probabilities.

**Architecture:** The pinned Qwen3.5 model has runs of three Gated DeltaNet layers separated by full attention layers. Capture only each three-layer run; execute full attention, rotary positions, normalization, and final projection eagerly. Cache all eight graph segments as one exact-layout entry, with limits and eager fallback.

**Tech Stack:** Python 3.12, PyTorch 2.14 CUDA Graphs, Transformers 5.17, PEFT 0.21, pytest, RTX 4090.

---

### Task 1: Diagnose numerical divergence

- [x] Reproduce the two rejected cases from #20.
- [x] Compare per-layer eager and full-graph outputs; identify the first divergent operation.
- [x] Prototype segmented capture and compare complete candidate logits to the original forward.

### Task 2: Build the bounded runtime

**Files:** `src/models/cua_s1/multimodal/graph_runtime.py`, `tests/cua_s1/test_graph_runtime.py`.

- [x] Add failing tests for configuration limits, exact tensor signatures, and LRU eviction.
- [x] Implement static input buffers, side-stream warmup, capture, and replay for contiguous linear-attention layers.
- [x] Build the exact model loop with eager full attention and the original final norm and head.
- [x] Validate a newly captured shape against the original forward before caching it.
- [x] Bound the cache by layout count and retained allocation; fall back to eager on unsupported inputs and resource limits.

### Task 3: Integrate the worker

**Files:** `src/models/cua_s1/multimodal/model.py`, `src/models/cua_s1/multimodal/server.py`, `tests/cua_s1/test_image_reuse.py`.

- [x] Add a failing test that an enabled runtime receives the prepared multimodal tensors.
- [x] Route only the reused multi-question path through the runtime; retain existing default behavior.
- [x] Add server flags for opt-in Graph mode and limits; keep the server's inference lock.

### Task 4: GPU acceptance and PR

**Files:** `recipe/cua_s1/experiments/rtx4090-graph-runtime/README.md`, `recipe/cua_s1/README.md`.

- [x] On RTX 4090, verify eager versus Graph for repeated, distinct, changed-image, changed-text, and long prompts, including the two previously rejected cases.
- [x] Measure synchronized whole-request p50/p95, capture cost, retained GPU memory, and fallback counts.
- [x] Run the full Cua-S1 test suite, Ruff, and format checks.
- [x] Inspect the diff, commit, push, and open PR #22.
- [x] Stop GPU processes and shut down the user-provided server after evidence is saved.
21 changes: 21 additions & 0 deletions docs/superpowers/plans/2026-09-28-cua-image-reuse.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Cua-S1 Image Reuse Implementation Plan

> **For agentic workers:** Use subagent-driven-development for implementation and independent review; controller owns GPU experiments and publication.

**Goal:** Complete request-local image reuse with reproducible correctness and speed evidence, submit a PR, then shut down the GPU server.

**Architecture:** Keep existing prepare/score as the reference path. Add isolated image preparation reuse and shared visual features for multi-question predict. Build per-question embeddings and positions using pinned model helpers.

**Tech Stack:** Python, PyTorch 2.14, Transformers 5.17, PEFT 0.21, RTX 4090.

- [x] Baseline: run all tests/cua_s1; preserve existing PR #15 data.
- [x] Implementation: tests/cua_s1/test_image_reuse.py first; prove failures; implement src/models/cua_s1/multimodal/model.py and focused helper module if needed. Preserve preflight validation, distinct prompts, no shared mutable state, PEFT vision layers and unchanged single-question behavior.
- [x] Harness: recipe/cua_s1/benchmark_image_reuse.py plus CPU tests. Check exact tensors, positions and probabilities before timed paired measurements; alternate order; record each sample and independent per-variant peak memory. Retain failed reports. Reuse profile_multimodal fixture and provenance helpers.
- [x] GPU: use clean committed checkout, run CPU tests, correctness smoke then 16-case paired experiment (5 warmups, 2 x 50 iterations per variant), plus distinct questions. Separate count/trace verification from all timed calls.
- [x] Review: independent spec review followed by quality review; resolve findings and re-run affected checks.
- [x] Publication: document measured results and tensor contract, retain raw JSON, run lint/full tests, commit/push and create linked PR identifying #12/#15 dependencies.
- [x] Shutdown: copy experiment artifacts locally, verify hashes, confirm no required GPU work remains, shut down server and verify provider/connection state.

## Completion evidence

Published non-draft [PR #17](https://github.com/ThinkFlowLab/system1-omni/pull/17). The complete experiment archive was copied locally and its SHA-256 matched the GPU host. After publication, the platform shutdown command exited successfully; a subsequent SSH attempt was refused. No experiment process remained on the GPU before shutdown.
24 changes: 24 additions & 0 deletions docs/superpowers/plans/2026-09-28-cua-post-reuse-optimization.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
# Cua-S1 Post-Reuse Optimization Implementation Plan

> **For agentic workers:** Execute these steps in order in the isolated worktree; preserve independent profiling and timed results.

**Goal:** Select and verify one optimization after image reuse, publish a reproducible PR, and shut down the GPU server.

**Architecture:** Use the existing reference and reuse paths. Measure current kernel groups, then limit the output projection to one token only if the paired experiment shows a benefit. Store raw experiment evidence and a readable analysis under `recipe/cua_s1/experiments/`.

**Tech Stack:** Python 3.12, PyTorch 2.14, Transformers 5.17, PEFT 0.21, RTX 4090.

---

- [x] Profile the unchanged reuse path with 1 and 8 questions; retain raw kernel groups, module invocation counts, traces, and source revision.
- [x] Run a temporary paired exploration of full projection versus `logits_to_keep=1`, checking complete responses before timing. Select the candidate only if its effect is credible.
- [x] Add a failing tensor integration test in `tests/cua_s1/test_image_reuse.py` proving the reused path passes `logits_to_keep=1`; run it on the GPU host or a torch-equipped environment.
- [x] Modify `score_reused` in `src/models/cua_s1/multimodal/model.py` and clean up a nearby Ruff guard; rerun the focused test and the full Cua-S1 suite.
- [x] Add a paired benchmark recipe with explicit baseline selection, alternating order, raw samples, response parity, source and environment metadata, and peak allocated memory.
- [x] Commit a clean source revision, transfer it to the GPU host, and run correctness and timed measurements. Review raw samples, variability, kernel changes, and single-question behavior.
- [x] Publish result data and reproduction commands, run verification, push the branch, create and attach the PR.
- [x] Archive and checksum complete GPU evidence locally, ensure no experiment is running, shut down the server, and verify it no longer accepts SSH.

## Completion evidence

Published [PR #18](https://github.com/ThinkFlowLab/system1-omni/pull/18) with the 17-case, 2,040-prediction dataset. The independent audit recomputed every saved p50/p95 and passed the 13-fixture, 49-question exact correctness and eight-question HTTP checks. Clean measured source `224d88b` passed 105 Cua-S1 tests on the GPU host; local tests passed 104 with one torch-dependent skip. The full trace archive SHA-256 matched on the server and local host: `774d952e89a9586709419c6c8583e4aa7ee2d27b4bf836253c8c01881c4dccd9`. The platform shutdown command exited successfully, and a subsequent SSH attempt was refused.
9 changes: 9 additions & 0 deletions docs/superpowers/specs/2026-09-28-cua-image-reuse-design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# Request-local Cua-S1 image reuse

Approved scope: implement the proposed single-request preprocessing and vision reuse, verify exact correctness and paired RTX 4090 performance, publish a PR, back up results and shut down the experiment server.

Reuse the single image within a parsed Request. Preserve each question's full processor tokenization, candidate readout, token accounting and three-dimensional positions. Validate every processed length before any vision or language execution. Keep the unmodified single-question path and an explicit baseline path for experiments. The shared visual encoder includes all active PEFT adapters. Never cache text embeddings, positions or outputs across questions or images. Do not introduce a cross-request cache, batching, adapter merging or approximate arithmetic.

Preprocessing should reuse the pinned processor's image result through a request-local processor copy, without changing global modules. Vision reuse should compute the image features once, build question-specific embeddings and explicit positions through the pinned model helpers, then invoke the existing PEFT model with inputs_embeds. Record the concrete tensor interface and ownership in the recipe documentation.

Acceptance: exact prepared tensor and candidate probability parity across distinct questions, varying candidate counts/order, image sizes/formats and consecutive different-image requests; one preprocessing and one vision execution per multi-question request; failure cleanup; no single-question regression outside measurement noise. Measure the existing 16-case matrix using paired alternating baseline/candidate calls, plus a distinct-question workload. Save raw samples, environment, source revisions, memory peaks and commands. Counts and optional trace verification run outside latency measurements. Publish without claiming native backend or kernel optimization.
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
# Cua-S1 optimization after request-local image reuse

The user approved continuing the experiment after PR #17, publishing a complete PR, backing up evidence, and shutting down the GPU server.

Profile the existing request-local reuse path on the same pinned RTX 4090 environment. Preserve an unmodified response and latency reference, and examine one- and eight-question cases with both repeated and distinct questions. CPU annotations, CUDA kernels, and synchronized end-to-end latency are different measurements and must not be added together.

Three candidates follow from the existing implementation: compute only the final output-token projection, replace the PyTorch Gated DeltaNet/causal-convolution fallbacks, or capture a stable request shape with CUDA Graphs. The first is the smallest change and can be tested against the existing model contract. Select it only if a clean paired experiment shows a useful end-to-end gain and parity. Otherwise report the evidence and choose the next measured hotspot rather than claiming a speedup.

If selected, pass `logits_to_keep=1` only in the reused multi-question path. The model still performs one language forward per question; the parameter limits the output head to the final hidden state, from which candidate probabilities are already read. Retain the old reference path for paired comparison. Confirm exact input embeddings and 3D positions, candidate probabilities within the pinned oracle tolerance, and identical decisions and complete responses. Benchmark 1/2/4/8 questions with alternating order, save individual samples and peak allocated memory, and record source/environment provenance. Document limitations and restore the GPU host to a stopped state after verified backup.
Loading