Skip to content

[Bug]: Linux, sibling engines: --auto-tier silently drops the VRAM tier that coli plan/coli doctor advertise (qwen36 CUDA build, 11.8 → 21 tok/s with --gpu auto) #1581

Description

@troturier

Commit or version

v1.11.0 — main at a8f2ca6. The code path is unchanged on dev (b5968eb, c/coli line 579).

Environment

Ubuntu 26.04.1, i7-11700K (8C/16T), 62 GB RAM, RTX 4060 Ti 8 GB (driver 595.91.07), CUDA Toolkit 12.4 (Ubuntu package, NVCC=/usr/bin/nvcc NVCC_CCBIN=g++-13), gcc 15.2, Samsung 980 PRO NVMe. Model: Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64.

Reproduction steps

make -C c qwen36 CUDA=1 CUDA_ARCH=native NVCC=/usr/bin/nvcc NVCC_CCBIN=g++-13
make -C c colibri CUDA=1 CUDA_ARCH=native NVCC=/usr/bin/nvcc NVCC_CCBIN=g++-13   # so cuda_binary() passes on main

# 1. the planner and the doctor both advertise a VRAM tier, no flag needed
python3 c/coli plan   --model <qwen36_i4_gs64>
python3 c/coli doctor --model <qwen36_i4_gs64>

# 2. apply the plan
COLI_MODEL=<qwen36_i4_gs64> python3 c/coli serve --auto-tier --model-id qwen36
# ... one chat completion over the API, nvidia-smi during decode

# 3. same, with --gpu auto
COLI_MODEL=<qwen36_i4_gs64> python3 c/coli serve --auto-tier --gpu auto --model-id qwen36

Expected behavior

--auto-tier is documented as "automatically apply the RAM/VRAM plan". The plan it applies is the one coli plan prints, which on this box includes a VRAM tier:

VRAM   5.2 GB hot tier · ~2963 experts · 0:NVIDIA GeForce RTX 4060 Ti
limit  CPU expert tail and GPU compute

and coli doctor reports [ ok] accelerator.gpu GPU engine and devices are available (doctor checks the qwen36 binary itself). I expected step 2 to start the tier, or at least to say that it was not going to.

Actual behavior and logs

Step 2 starts CPU-only, silently. The engine banner has no [CUDA] / [qtier] line, VRAM stays at the desktop's 897 MiB during decode, and the serve log has no [PLAN] line at all (the sibling-engine path prints none), so nothing tells the user the VRAM part of the plan was dropped:

== qwen36 Phase-2 engine | cache=256/layer bits=4 ctx=8192 pilot=0 wide=1 hot=0 smooth=0.30 conf=0.92 ==
resident weights loaded in 6.7s | RSS after load: 9.23 GB
OpenAI-compatible API listening on http://127.0.0.1:8000/v1

Step 3 works as intended:

[CUDA] device 0: NVIDIA GeForce RTX 4060 Ti, 8.2 GB VRAM, sm_89
[place] auto: dev 0 holds 1207.4 MB of trunk (lmhead yes, 30 dnproj layers), 4.05 GB left for experts
[qtier] CUDA VRAM expert tier active: 1 device(s), 1.80 MB/expert

Same 115-token completion, temperature=0, identical output text:

launch request (warm) tools/datapoint.py rotating median TTFT
--auto-tier 11.5 s 11.8 tok/s 2.4–3.4 s
--auto-tier --gpu auto 5.9 s 21.0 tok/s 1.0 s

So a Linux user of a sibling engine who does what coli plan suggests gets half the throughput and no hint why. I only found it because I was watching nvidia-smi.

Where it comes from

In env_for() the two paths disagree:

  • GLM (c/coli ~line 767 on main): has_cuda=cuda_binary(); e=environment_for_plan(plan,e,has_cuda)--auto-tier alone enables the tier when the binary is a CUDA build.
  • sibling engines (~line 572 on main, 579 on dev): environment_for_plan(plan,env,cuda_enabled=env.get("COLI_CUDA") == "1"), with the comment "merely asking for auto-tier must not turn a CPU-only sibling binary into an attempted CUDA launch".

That concern was legitimate when cuda_binary() could only inspect the GLM binary. Since #1533 / #1537 dev has cuda_binary(engine) and engine_for_gpu_check(a), so the sibling path can now make the same decision as the GLM path, from the right binary. (On main v1.11.0, --gpu auto additionally refuses with "the engine binary is CPU-only" unless c/colibri is also a CUDA build — that is #1533, already fixed on dev, not re-reporting it.)

The bare-coli chat auto-enable block is Windows-only on purpose (#363, "Linux already has working detection + the explicit-flag UX"); this report is not about bare chat, it is about --auto-tier dropping part of the plan it claims to apply.

Suggested fix

Either of:

  1. In the sibling-engine path, cuda_enabled = env.get("COLI_CUDA") != "0" and cuda_binary(engine_for_gpu_check(a)) — apply the plan's VRAM tier when the family engine is a CUDA build, keep --gpu none / COLI_CUDA=0 as the off-switch, same contract as the GLM path.
  2. If the explicit-flag contract is to stay for sibling engines: print the [PLAN] RAM … · CPU line on that path too, and have coli plan / coli doctor say that the VRAM tier needs --gpu auto on Linux, so the printed plan and the applied plan cannot silently differ.

Happy to test a patch on this box (single 8 GB card, qwen36 and deepseek_v4 CUDA builds available).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions