Skip to content

Mode B with Ollama thinking models (qwen3.x, DeepSeek): LLM fact extraction silently falls back to Mode A; plus daemon port-conflict false-ready, unrecoverable migration drift, and lease deadlock after daemon kill #128

Description

@tianjueak

Environment

  • SLM version: 4.1.11 (installed via npm, run through the bundled .slm-venv)
  • OS: Windows 11 24H2
  • Runtime: Node 22.22.2 (launcher) + Python 3.12 (daemon process)
  • LLM: Ollama 0.x (fresh reinstall), model qwen3.5:9b (Q4_K_M) — a thinking model
  • Embedding: Ollama nomic-embed-text (768-dim)
  • Mode: B (client_driven_agentic: true, agentic_max_rounds: 3)
  • Access: MCP stdio (SLM_DAEMON_PORT=8767) + CLI (slm-npm)

Bug 1 (critical): Thinking models return their whole answer in message.thinking, leaving message.content empty — LLM fact extraction silently degrades to Mode A

Symptom: In Mode B, remember produces facts with entities: [], importance: 0.5 (the Mode-A default), and fact_type often misclassified. daemon.log shows:

LLM extraction returned no facts, falling back to local.

Root cause: qwen3.x-family thinking models (and likely DeepSeek-R1-style models) put the entire generation in the message.thinking field of the native /api/chat response; message.content is an empty string when the task is complex enough to trigger reasoning. LLMBackbone._extract_text (llm/backbone.py) reads only message.content:

return data.get("message", {}).get("content", "").strip()

So raw is "" → _parse_llm_response hits the if not raw or not raw.strip(): return [] branch (silent, no log) → chunk extraction falls back to _extract_local. The LLM is invoked (you can see the /api/chat call in the Ollama log, taking ~25s) but its output is discarded.

Verified with a direct /api/chat call using the exact same _SYSTEM_PROMPT + prompt from fact_extractor.py:

=== message keys: ['role', 'content', 'thinking']
=== content len: 0
=== thinking len: 3182

Workaround used:

  1. In _build_ollama, add the top-level "think": False to the payload — qwen3.x supports disabling thinking entirely, and then content carries the JSON answer. (Other models ignore this key.)
  2. In _extract_text, fall back to message.thinking when content is empty, as a belt-and-braces measure.

Suggested fix: at minimum, make _extract_text fall back to thinking when content is empty; optionally allow a config flag to disable thinking (think: False) for known thinking models. Also log a warning instead of silently returning [] when raw is empty — the current silent path makes Mode B look "working" while being pure Mode A.


Bug 2 (packaging): npm source tree and .slm-venv/Lib/site-packages contain two divergent copies of the package

Symptom: Patching node_modules/superlocalmemory/src/superlocalmemory/llm/backbone.py had no effect on the daemon. The daemon is launched from the bundled venv and actually imports the pip-installed copy at .slm-venv/Lib/site-packages/superlocalmemory/, which is a byte-for-byte separate file.

Impact: Any user patching the npm source tree (as the docs suggest for hotfixes) will silently patch the wrong copy. Took several hours to find that two copies existed at all. It also means the npm tarball carries a full second copy of the Python package inside the venv — worth confirming this is intentional and documenting which one is authoritative.

Suggested fix: document clearly which copy is loaded (or make the venv a proper editable install of the src tree so there is a single source of truth).


Bug 3 (reliability): daemon_port conflict produces a "false-ready" daemon and a falsely-successful CLI status

Symptom: With daemon_port: 8765 in the config and port 8765 already occupied by another service (in my case, my own admin executor), slm serve start failed, but slm status returned success: true (mode b, fact_count 0) — and every write (remember) failed with DAEMON_UNAVAILABLE (request_rejected). /health on the daemon port answered "ok" even though the daemon never bound a usable port.

Root cause: the config default daemon_port=8765 collided with an unrelated service. The daemon's legacy-port redirect logic (Port 8767 in use (old daemon?), skipping legacy redirect) plus a stale daemon.pid made the CLI believe an owned daemon existed and was healthy, while requests were rejected. status reads what looks like a live HTTP endpoint (possibly the unrelated service answering JSON) and reports success.

Impact: hours of confusion; the fix was deleting stale daemon.pid/daemon.port and starting with an explicit SLM_DAEMON_PORT=8767 env var. status should validate the daemon identity (capability/descriptor match) before reporting success, and a port-conflict should fail loudly at startup rather than half-start.


Bug 4 (recovery): completed migrations that fail verify() cannot self-heal — daemon rejects ALL writes forever

Symptom: after DROP TABLE atomic_facts + re-create (my own fault while cleaning test data), /health reported:

{"ready": false, "readiness": {"migrations": false,
 "migration_failures": ["M011_archive_and_merge","M013_bi_temporal_columns","M014_v345_scale_ready","M016_add_scope_support","M025_perf_indexes"], ...}}

Every write returned {"error":"Service unavailable: required schema migrations have not completed."}. The migration log still marks these as complete, but verify() fails because the recreated table is missing the migration-added columns/indexes.

Root cause: _migration_internals.py logs schema incomplete for completed migration X; automatic replay is disabled when the module has no repair(conn) hook — and M011/M013/M014 have no repair function at all (they only ship verify + a DDL string). So a drifted schema is permanently unrecoverable through the daemon; I had to stop the daemon and hand-apply the missing ALTER TABLE ... ADD COLUMN / CREATE INDEX statements.

Suggested fix: ship repair(conn) (or re-run the idempotent DDL steps) for every migration that only adds columns/indexes — these are inherently safe to re-apply. At minimum, expose a CLI repair command so users don't have to hand-edit the database.


Bug 5 (boundary): killing the daemon mid-enrichment leaves ingestion_operations stuck in enriching forever (lease deadlock)

Symptom: after force-killing the daemon, new operations stayed enriching with a stale non-empty lease_owner and never progressed, even after the lease expiry time passed. Manual UPDATE ingestion_operations SET lease_owner='', lease_expires_at=0 WHERE state='enriching' was required to unblock.

Impact: any crash/kill of the daemon during enrichment wedges the write pipeline until manual DB surgery. A stale-lease reaper that runs even when no worker process exists would make this self-healing.


Bug 6 (minor, misleading): DELETE FROM on sqlite-vec virtual tables (fact_embeddings*) throws database disk image is malformed while PRAGMA integrity_check returns ok

Symptom: clearing data with DELETE FROM fact_embeddings (a vec0 virtual table) errored with database disk image is malformed; PRAGMA integrity_check/quick_check both returned ok. Not a real corruption — likely a vec0 limitation. Worth either documenting or guarding so operators don't panic and nuke their DB.


Bug 7 (minor, protocol): MCP stdio server rejects notifications/initialized with Method not found

Symptom: standard MCP clients that send the notifications/initialized notification after initialize receive {"code":-32601,"message":"Method not found"}. Worked around by skipping the notification and calling tools directly. Newer MCP client libraries send this by default, so this may break some clients on Windows/stdio.


Notes

  • All of the above were hit on Windows with the npm distribution; some (Bug 1, Bug 2) are platform-independent.
  • Bug 1 is the one that most affects real users of Mode B with local Ollama: as shipped, Mode B with any thinking-capable local model silently behaves like Mode A on the write path while still paying the LLM inference cost.
  • Happy to provide the full daemon.log excerpts or the exact reproduction scripts (MCP test client, direct /api/chat payload) for any of these.

Environment recap: SLM 4.1.11 · Windows 11 24H2 · Node 22.22.2 · Python 3.12 · Ollama (qwen3.5:9b / nomic-embed-text) · MCP stdio + CLI, daemon on port 8767.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions