Skip to content

chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 0c09831 ) - #2289

Merged
1Solon merged 1 commit into
mainfrom
renovate/ghcr.io-ggml-org-llama.cpp-server-cuda
Sep 8, 2026
Merged

chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 0c09831 )#2289
1Solon merged 1 commit into
mainfrom
renovate/ghcr.io-ggml-org-llama.cpp-server-cuda

Conversation

@renovate

@renovate renovate Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

This PR contains the following updates:

Package Update Change
ghcr.io/ggml-org/llama.cpp digest e755b930c09831

Configuration

📅 Schedule: (UTC)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

Rebasing: Whenever PR is behind base branch, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about these updates again.


  • If you want to rebase/retry this PR, check this box

This PR was generated by Mend Renovate. View the repository job log.

@renovate
renovate Bot requested a review from 1Solon September 4, 2026 11:41
@renovate
renovate Bot force-pushed the renovate/ghcr.io-ggml-org-llama.cpp-server-cuda branch from e6c6b3f to 12a4abc Compare September 5, 2026 16:37
@renovate renovate Bot changed the title chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → fc76b62 ) chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 82a6723 ) Sep 5, 2026
@renovate renovate Bot changed the title chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 82a6723 ) chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 84a9f77 ) Sep 6, 2026
@renovate
renovate Bot force-pushed the renovate/ghcr.io-ggml-org-llama.cpp-server-cuda branch 2 times, most recently from 8bbc41a to bff14f1 Compare September 6, 2026 16:26
@renovate renovate Bot changed the title chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 84a9f77 ) chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 4ae7aeb ) Sep 7, 2026
@renovate
renovate Bot force-pushed the renovate/ghcr.io-ggml-org-llama.cpp-server-cuda branch from bff14f1 to f2185d1 Compare September 7, 2026 08:06
@renovate renovate Bot changed the title chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 4ae7aeb ) chore(container): update image ghcr.io/ggml-org/llama.cpp ( e755b93 → 0c09831 ) Sep 8, 2026
@renovate
renovate Bot force-pushed the renovate/ghcr.io-ggml-org-llama.cpp-server-cuda branch 9 times, most recently from 6d046c5 to 409950b Compare September 8, 2026 15:53
@renovate
renovate Bot force-pushed the renovate/ghcr.io-ggml-org-llama.cpp-server-cuda branch from 409950b to 6956c40 Compare September 8, 2026 15:58

@1Solon 1Solon left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact digest-only diff at 6956c40 and upstream 67a17c17caa95742186f8b1ecadd1b5abd6d5ebb...9dcf84e5ae2718947188b539aab8b9c2b15d3ba1 (78 commits). Qwen QKV/MTP changes preserve separate projections; GDN normalization is corrected; recurrent-layer metadata is converter-side; no required current-model/cache/CLI migration found. Router LRU changes do not apply to this fixed-model invocation. Inherits verified 34Gi request/limit from #2293. Server dry run passed. Proceeding with sequential rollout and loopback inference/OOM checks.

@1Solon
1Solon merged commit 0f4291b into main Sep 8, 2026
@1Solon
1Solon deleted the renovate/ghcr.io-ggml-org-llama.cpp-server-cuda branch September 8, 2026 16:02
@1Solon

1Solon commented Sep 8, 2026

Copy link
Copy Markdown
Owner

Post-merge verification passed at 0f4291b: live imageID sha256:0c09831866497d97802a57a36c4010253d3d43e692e521b947eec5fd4ef0af4d, fingerprint b10853-9dcf84e5a. Pod, Model, InferenceService and Flux Kustomizations ready. Request/limit remain 34Gi; memory.max=36507222016, observed peak=19827515392 bytes. Two loopback chat tests returned OK, both 2/2 MTP drafts accepted; repeat reused 13 prompt tokens. Health OK, zero restarts, all memory.events counters zero. kube5 Ready without memory pressure; NVIDIA GPU reporting works. Startup included unexplained nonfatal ERROR: init 250 result=11 and duplicate n-gpu-layers deprecation warning; neither prevented loading/inference. These are smoke checks, not a full-cache or long-context stress test. No storage/database changes or rollback performed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant