Skip to content

Doom demo: SauerkrautLM-Doom on Core ML, and GLiClass on ViZDoom - #20

Merged
Alex-Wengg merged 8 commits into
mainfrom
feat/vizdoom-gliclass
Sep 27, 2026
Merged

Alex-Wengg merged 8 commits into
mainfrom
feat/vizdoom-gliclass

Conversation

@Alex-Wengg

@Alex-Wengg Alex-Wengg commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

ViZDoom defend_the_center demo: a Core ML port of SauerkrautLM-Doom-MultiVec-1.3M (VAGO solutions, Apache 2.0), plus a headless check of stock GLiClass on the same game.

  • Tools/doom/sauerkraut/: split-screen demo (game, the 40×25 depth grid the model reads, action probabilities, ms per decision), waits for space, resizable. demo.sh also opens a Ghostty window with sudo asitop and the decision log. Model downloads from FluidInference/sauerkrautlm-doom-coreml. Conversion and PyTorch/Core ML evaluation scripts included.
  • Tools/doom/defend_the_center.py + GLiClassServe (GLiClass over JSON lines): stock GLiClass vs a hand-coded aimer.

Seeds 10000–10099, M5 Pro:

Player Kills Survived ms / decision
Sauerkraut, Core ML fp16 GPU 20.54 50.6 s 1.2
Sauerkraut, PyTorch 20.42 50.5 s 8.4 (MPS fp16), 10.5 (MPS fp32)
Hand-coded aimer 13.05 24.8 s —
GLiClass, consequence labels 11.98 23.1 s 4.5

Core ML fp32 matches PyTorch kill for kill on 100/100 seeds. Notes for reviewers:

  • Upstream's ASCII channel overflows uint8 and is always @, so the model plays from depth bins; the demo says "depth grid".
  • Rendering every tic changes ViZDoom's RNG, so the window doesn't replay benchmark seeds exactly (play.py --check does); play is equally strong (19.70 vs 20.37 kills over 30 seeds).
  • Python tooling only, apart from GLiClassServe.

🤖 Generated with Claude Code

GLiClassServe exposes GLiClassManager over JSON lines on stdin/stdout so
non-Swift harnesses can call it. Tools/doom drives ViZDoom with text state
from the labels buffer (no pixels) and compares GLiClass against a
hand-coded aimer and random on seeds 1-20.

Bare labels: 0 kills (turns correctly, never fires). Consequence labels:
14.00 kills vs aimer 14.80 at ~4 ms/call, 6/5/9 W/T/L per seed.
…ter)

Core ML port of VAGOsolutions/SauerkrautLM-Doom-MultiVec-1.3M (Apache 2.0)
plus a split-screen pygame viewer: game, the 40x25 depth grid the model
reads, action probabilities, ms per decision. demo.sh also opens Terminal
windows for sudo asitop and tail -f of the ANSI decision log.

Conversion re-implements the ModernBERT forward with static masks since
HF mask construction does not trace through coremltools. fp32 Core ML is
identical to PyTorch on 100/100 seeds (20.42 kills, 50.5 s); fp16 L1026
on GPU 20.54 kills at 1.5-3.4 ms vs 57.7 ms PyTorch CPU.

Runtime needs no PyTorch: tokenizer and upstream convert_with_depth are
re-implemented, including its uint8 truncation of downscaled depth
(needed for seed-exact parity). Upstream's ASCII channel overflows uint8
and is always '@', so the panel draws depth bins.

Rendering every tic consumes ViZDoom game randomness, so the viewer does
not replay headless seeds exactly; play is equally strong (19.70 vs
20.37 kills on seeds 10000-10029).
tmux session doom-demo: sudo asitop on top, tail -f of the log below,
focus on the asitop pane for the password. Falls back to two windows
without tmux. Log lines shortened to fit the pane.
open -na Ghostty.app -e runs a generated launcher that starts the tmux
session (asitop top, log bottom). Terminal.app windows remain the
fallback when Ghostty or tmux is missing.
Game on top, depth grid and probabilities below (680x815 layout, scaled
into a resizable window, fits above the Dock). Each episode waits on its
first frame until space; --autostart (and headless recording) skip it.
Ghostty gets the terminal via --command=, which avoids its execute
confirmation prompt.
Default model downloads from FluidInference/sauerkrautlm-doom-coreml into
~/Library/Caches/FluidUse (real files: Core ML fails to compile a
symlinked .mlpackage from the hub cache). --model still takes a local
path. Headless check reproduces the benchmark seeds from the download.
Idle-machine medians on the same frame: PyTorch MPS 10.5 ms fp32 / 8.4 ms
fp16 vs Core ML GPU 2.5 ms fp32 / 1.2 ms fp16. The overlay now shows the
MPS fp16 figure next to the fp16 Core ML model.
Pressing n mid-episode showed "survived" on the end card. The *.mp4
ignore also matched Media/, where the repo tracks demo videos.
@Alex-Wengg
Alex-Wengg merged commit f45d3d4 into main Sep 27, 2026
1 check passed
@Alex-Wengg
Alex-Wengg deleted the feat/vizdoom-gliclass branch September 27, 2026 04:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant