Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 1 addition & 3 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -33,9 +33,7 @@ jobs:
--output-file /tmp/slm-runtime-requirements.txt
uvx --from pip-audit==2.9.0 pip-audit \
--requirement /tmp/slm-runtime-requirements.txt \
--no-deps --disable-pip \
--ignore-vuln GHSA-rrmf-rvhw-rf47 \
--ignore-vuln PYSEC-2026-3447
--no-deps --disable-pip
- name: Reject high-confidence high-severity Python findings
run: uvx --from bandit==1.8.6 bandit -r src/superlocalmemory -q -lll -iii

Expand Down
26 changes: 25 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,12 +5,36 @@
</picture>
</p>

<h1 align="center">SuperLocalMemory V4.1.17</h1>
<h1 align="center">SuperLocalMemory — local-first memory for AI agents</h1>

<h2 align="center">Rent the LLM. Own the memory.</h2>

<p align="center"><em>Rent an LLM — but own the memory, for your company and for your industry.</em></p>

Store context once and recall it across agent sessions through the CLI or MCP. Start with local memory, then choose the operating mode and integrations your workspace needs. SuperLocalMemory is part of Qualixar's AI Reliability Engineering work.

**[Research and evidence](https://www.superlocalmemory.com/research)** · **[Qualixar product overview](https://qualixar.com/products/superlocalmemory)** · **[Author and research context](https://varunpratap.com/products/superlocalmemory)**

**[Install](https://www.superlocalmemory.com/install)** · **[Product walkthrough](https://www.superlocalmemory.com/demo)** · **[CLI proof](docs/QUICK_PROOF.md)** · **[Release notes](CHANGELOG.md)**

```bash
npm install -g superlocalmemory # Primary CLI install path
slm setup # Select mode A for local-only operation; review integration choices.
slm doctor
slm remember "Synthetic demo: the release checklist requires a human approval after tests pass." --tags demo --json --sync
slm recall "release checklist human approval" --json
```

The synthetic CLI proof returned the stored sentence in mode A on version 4.1.17. It tests store-and-recall behavior; it is not a retrieval accuracy benchmark.
[See the commands, scope and isolation settings](docs/QUICK_PROOF.md).

If the proof is useful for your agent workflow, [star the repository](https://github.com/qualixar/superlocalmemory) to find the project again. Stars are optional; installation and documentation are open without one.

**Source security repair:** dependency changes on this branch do not update already published packages. [See dependency security and the temporary optional-backend restriction](docs/dependency-security.md).

## Governance, architecture and research


<p align="center"><strong>The governed memory layer for AI agents: local-first, auditable, and built for the compliance obligations teams now actually carry.</strong><br/>
Models are interchangeable and rented by the token. What your agents <em>remember</em> is
yours — it is your customers' data, your retention obligations, and your audit trail. SLM
Expand Down
20 changes: 20 additions & 0 deletions docs/QUICK_PROOF.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Store and recall one synthetic fact

Install using the [Quick Start](../README.md#quick-start), choose operating mode
A in `slm setup`, and run `slm doctor` before using the CLI.

```bash
slm remember "Synthetic demo: the release checklist requires a human approval after tests pass." --tags demo --json --sync
slm recall "release checklist human approval" --json
```

On 1 October 2026, source version 4.1.17 returned a durable `queryable` receipt
and recalled that exact synthetic sentence. Recall returned one result with
`calibration_status: uncalibrated` and `answer_confidence: null`. A ranking score
is not an accuracy measurement. This one-fact check is not a retrieval benchmark.

The verification used a separate `SLM_DATA_DIR`, `SLM_DAEMON_PORT=8893` and
`SLM_DISABLE_LEGACY_PORT=1` to avoid an existing memory service. It did not write
real customer data or change host hooks. Use an unused port and a dedicated
absolute data directory if you reproduce that isolation. Setup and integrations
remain explicit choices; review them before changing an existing workspace.
34 changes: 34 additions & 0 deletions docs/dependency-security.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# Dependency security repair in source

The October 2026 repair updates affected runtime dependency floors and the lock
without adding a blanket audit exception. It covers sentence-transformers,
transformers, NLTK removal, HTTP clients, AnyIO, PyJWT, accelerate and urllib3.
The build-system setuptools floor is 83.0.0 or newer. This is a source repair;
a published package does not acquire these changes until a new release ships.

## Optional aggressive compression

NLTK through 3.10.3 has an unpatched High advisory,
[GHSA-8mgp-746c-j5xp](https://github.com/advisories/GHSA-8mgp-746c-j5xp).
LLMLingua 0.2.2 depends on NLTK. These packages are temporarily omitted from
shipped requirements rather than hidden by an audit ignore or moved into an
unsafe extra. The LLMLingua implementation and selected model remain in source.

An existing installation cannot activate this backend with a known-affected,
missing, malformed or prerelease NLTK version: the guard runs before importing
the backend. Setup uses the same guarded constructor. The router retains its
lossless fallback and records why the optional backend is unavailable.
Local memory, retrieval, cache, safe normalization and reversible storage are
separate capabilities. No new aggressive compression ratio is claimed.

## Native dependency and integration updates

Torch 2.13 removes the prior Low Torch advisory and permits a patched setuptools
runtime. Optional ingestion dependencies receive floors for httplib2, icalendar,
oauthlib and pyasn1, so an all-extras install follows the repaired graph too.
There is no blanket vulnerability ignore. Exact lock, runtime and all-extras
audit results must be recorded before release.

Before release, verify the exact wheel/lock, runtime and optional-backend behavior,
embedding/reranker compatibility, tests, and both filtered and unfiltered audits.
Do not present source-only changes as already fixed in the published PyPI/npm package.
41 changes: 24 additions & 17 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -68,12 +68,12 @@ dependencies = [
# sentence-transformers[onnx] — its extras pull optimum which
# overrides the sentence-transformers pin via transitive deps.
# Pin all three explicitly instead.
"sentence-transformers==5.3.0",
"sentence-transformers==5.6.0",
"optimum==2.1.0",
"onnxruntime==1.24.4",
"transformers==5.5.4",
"transformers==5.10.1",
"huggingface_hub==1.5.0",
"torch==2.11.0",
"torch==2.13.0",
"scikit-learn==1.8.0",
# Vector KNN extension for the semantic channel.
"sqlite-vec==0.1.9",
Expand All @@ -86,14 +86,18 @@ dependencies = [
# bindings. Existing SLM graph stores were created by 0.3.0; upgrading the
# wheel without an export/import migration makes daemon startup panic.
"pycozo[embedded]==0.3.0",
# v3.6.10: LLMLingua-2 prose compression (aggressive mode, opt-in at runtime).
# Hard dependency so `pip install superlocalmemory` ships compression-ready;
# the ~560MB model downloads on first setup/warmup (fail-open). Verified the
# pin does NOT disturb the transformers/torch/numpy pins above.
"llmlingua==0.2.2",
# LLMLingua pulls NLTK transitively. Pin above the disclosed traversal
# advisories so clean installs cannot resolve to the vulnerable 3.9.x line.
"nltk==3.10.0",
# LLMLingua/NLTK are temporarily not shipped: the latest NLTK has an
# unpatched High advisory. Default/lossless compression stays available.
# See docs/dependency-security.md before restoring this optional backend.
"packaging>=24.0",
# Security floors apply to pip installations as well as the uv lock.
"accelerate>=1.15.0,<2",
"anyio==4.14.2",
"httpcore2>=2.10.0,<3",
"httpx2>=2.12.0,<3",
"pyjwt>=2.15.1,<3",
"urllib3>=2.8.0,<3",
"setuptools>=83.0.0",
]

[project.optional-dependencies]
Expand All @@ -102,10 +106,10 @@ dependencies = [
# works but installs nothing extra.
search = [
# Same hard pin as core deps — no [onnx] extra.
"sentence-transformers==5.3.0",
"sentence-transformers==5.6.0",
"optimum==2.1.0",
"einops==0.8.2",
"torch==2.11.0",
"torch==2.13.0",
"scikit-learn==1.8.0",
"onnxruntime==1.24.4",
]
Expand Down Expand Up @@ -139,7 +143,10 @@ ingestion = [
"keyring>=25.0.0",
"google-auth-oauthlib>=1.2.0",
"google-api-python-client>=2.100.0",
"icalendar>=6.0.0",
"icalendar>=7.1.3",
"httplib2>=0.32.0",
"oauthlib>=4.0.0,<5",
"pyasn1>=0.6.4",
]
full = [
"superlocalmemory[search,ui,learning,performance,ingestion,injection,scale]",
Expand All @@ -151,7 +158,7 @@ dev = [
"sqlite-vec>=0.1.6",
# Starlette TestClient's maintained synchronous backend. Test-only: the
# product runtime continues to use the separately pinned httpx client.
"httpx2==2.5.0",
"httpx2==2.13.1",
]

[project.urls]
Expand All @@ -164,7 +171,7 @@ Issues = "https://github.com/qualixar/superlocalmemory/issues"
slm = "superlocalmemory.cli.main:main"

[build-system]
requires = ["setuptools>=77.0.3", "wheel"]
requires = ["setuptools>=83.0.0", "wheel"]
build-backend = "setuptools.build_meta"

[tool.setuptools.packages.find]
Expand Down Expand Up @@ -242,7 +249,7 @@ select = ["E", "F", "I", "W"]
[dependency-groups]
dev = [
"build>=1.4.0",
"httpx2==2.5.0",
"httpx2==2.13.1",
"pytest>=9.0.2",
"pytest-asyncio>=0.21",
"twine>=6.2.0",
Expand Down
2 changes: 1 addition & 1 deletion src/superlocalmemory/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@
__version__ = "4.1.17"

_REQUIRED_VERSIONS = {
"sentence_transformers": "5.3.0",
"sentence_transformers": "5.6.0",
"onnxruntime": "1.24.4",
}

Expand Down
15 changes: 7 additions & 8 deletions src/superlocalmemory/cli/setup_wizard.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ def _resolve_slm_home() -> Path:
return canonical_data_root()
_EMBED_MODEL = "nomic-ai/nomic-embed-text-v1.5"
_RERANKER_MODEL = "cross-encoder/ms-marco-MiniLM-L-12-v2"
# v3.6.10: compulsory LLMLingua-2 prose compression model (~560MB, aggressive mode).
# Selected optional LLMLingua model; backend restricted pending dependency review.
_COMPRESSOR_MODEL = "microsoft/llmlingua-2-xlm-roberta-large-meetingbank"


Expand Down Expand Up @@ -211,17 +211,16 @@ def _download_compressor(model_name: str) -> bool:
"""Download the LLMLingua-2 prose compression model (v3.6.10).

Mirrors _download_reranker: a subprocess forces the HF download with visible
progress. Fail-open — a network hiccup must NOT break setup; the model also
lazy-downloads on first use in prose_llmlingua.py.
progress. The shared backend guard runs before model loading; a restricted
backend stays unavailable and does not trigger an automatic dependency install.
"""
print(f"\n Downloading compression model: {model_name}")
print(f" (LLMLingua-2 prose compressor, ~560MB — aggressive mode only)\n")

# H-03: model name via argv, never interpolated into executed source.
script = (
"import sys; from llmlingua import PromptCompressor; "
"PromptCompressor(model_name=sys.argv[1], use_llmlingua2=True, "
"device_map='cpu'); "
"import sys; from superlocalmemory.optimize.compress.prose_llmlingua import LLMLinguaCompressor; "
"LLMLinguaCompressor(model_name=sys.argv[1], device_map='cpu'); "
"print('OK')"
)

Expand All @@ -241,10 +240,10 @@ def _download_compressor(model_name: str) -> bool:
if result.returncode == 0:
print(f" ✓ Compression model ready")
return True
print(f" ✗ Compression model download failed (will lazy-download on first use)")
print(f" ✗ Compression backend unavailable; default/lossless compression remains available")
return False
except ImportError:
print(f" ⚠ llmlingua not installed — compression model will download on first use")
print(f" ⚠ Compression backend not shipped pending upstream security fix; lossless compression remains available")
return False
except Exception as exc:
print(f" ✗ Compression model error: {exc}")
Expand Down
9 changes: 9 additions & 0 deletions src/superlocalmemory/core/component_registry.py
Original file line number Diff line number Diff line change
Expand Up @@ -368,6 +368,15 @@ def probe_cozo() -> Component:


def probe_llmlingua() -> Component:
from superlocalmemory.optimize.compress.prose_llmlingua import backend_restriction

restriction = backend_restriction()
if restriction is not None:
return Component(
key="llmlingua", label="LLMLingua compressor lib",
category=CATEGORY_OPTIONAL, status=STATUS_DEGRADED,
detail=restriction, fix_cmd="", auto_fixable=False, last_checked=time.time(),
)
return _probe_optional_pkg(
"llmlingua", "LLMLingua compressor lib", "llmlingua", "llmlingua",
"not installed (optional prose compression)",
Expand Down
4 changes: 2 additions & 2 deletions src/superlocalmemory/core/engine_wiring.py
Original file line number Diff line number Diff line change
Expand Up @@ -331,9 +331,9 @@ def _init_vector_store(config: SLMConfig) -> Any | None:
if vs.available:
logger.info("VectorStore initialized (sqlite-vec KNN enabled)")
return vs
logger.debug("VectorStore unavailable; using ANNIndex fallback")
logger.warning("VectorStore unavailable; using ANNIndex fallback")
except Exception as exc:
logger.debug("VectorStore init failed: %s", exc)
logger.warning("VectorStore init failed: %s", exc)
return None


Expand Down
31 changes: 30 additions & 1 deletion src/superlocalmemory/optimize/compress/prose_llmlingua.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,15 +18,40 @@
from __future__ import annotations

import logging
from importlib.metadata import PackageNotFoundError, version
from typing import Any

from packaging.version import InvalidVersion, Version

logger = logging.getLogger("slm.optimize.compress.llmlingua")

_DEFAULT_RATE: float = 0.5
_MODEL_BERT: str = "microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank"
_MODEL_XLM: str = "microsoft/llmlingua-2-xlm-roberta-large-meetingbank"


# Empty until an upstream patched release has been independently reviewed.
_REVIEWED_NLTK_VERSIONS: frozenset[str] = frozenset()


def backend_restriction() -> str | None:
"""Check metadata only; never import model code or download a model."""
try:
installed = Version(version("nltk"))
except (PackageNotFoundError, InvalidVersion):
return "Optional LLMLingua backend unavailable: NLTK safety version cannot be verified."
if (
installed.is_prerelease
or installed.is_devrelease
or str(installed) not in _REVIEWED_NLTK_VERSIONS
):
return (
"Optional LLMLingua backend restricted pending a reviewed NLTK fix "
"for GHSA-8mgp-746c-j5xp. Default/lossless compression remains available."
)
return None


class LLMLinguaCompressor:
"""Opt-in LLMLingua-2 prose compressor."""

Expand All @@ -36,11 +61,15 @@ def __init__(
device_map: str = "cpu",
rate: float = _DEFAULT_RATE,
) -> None:
restriction = backend_restriction()
if restriction is not None:
raise ImportError(restriction)
try:
from llmlingua import PromptCompressor # type: ignore[import]
except ImportError as e:
raise ImportError(
"llmlingua package not installed. Install: pip install llmlingua."
"LLMLingua backend is not shipped while its NLTK dependency has an "
"unpatched High advisory. Default/lossless compression remains available."
) from e

logger.info("Loading LLMLingua-2 model=%s device=%s", model_name, device_map)
Expand Down
4 changes: 2 additions & 2 deletions src/superlocalmemory/optimize/compress/router.py
Original file line number Diff line number Diff line change
Expand Up @@ -369,8 +369,8 @@ def _get_llmlingua_compressor(self) -> "LLMLinguaCompressor | None": # pragma:
try:
from superlocalmemory.optimize.compress.prose_llmlingua import LLMLinguaCompressor
self._llmlingua_compressor = LLMLinguaCompressor()
except ImportError:
logger.warning("LLMLinguaCompressor not available — prose compression disabled")
except ImportError as exc:
logger.warning("LLMLinguaCompressor unavailable — using lossless compression: %s", exc)
return None
return self._llmlingua_compressor

Expand Down
2 changes: 1 addition & 1 deletion src/superlocalmemory/retrieval/vector_store.py
Original file line number Diff line number Diff line change
Expand Up @@ -243,7 +243,7 @@ def _ensure_vec0_table(self) -> None:
conn.execute(row_map_idx)
conn.commit()
except Exception as exc:
logger.debug("vec0 table creation failed: %s", exc)
logger.warning("vec0 table creation failed: %s", exc)
self._available = False

# -- Serialization ------------------------------------------------------
Expand Down
4 changes: 3 additions & 1 deletion src/superlocalmemory/storage/backup.py
Original file line number Diff line number Diff line change
Expand Up @@ -206,7 +206,9 @@ def _backup_via_sqlite_api(src: Path, dest: Path) -> None:
# Verify and durably flush BEFORE the rename, so the final name never
# appears over incomplete or corrupt content.
try:
fd = os.open(str(staging), os.O_RDONLY)
# Windows durable flush requires write access. This is our temporary
# snapshot, not the source database; do not suppress a failed flush.
fd = os.open(str(staging), os.O_RDWR | getattr(os, "O_BINARY", 0))
try:
os.fsync(fd)
finally:
Expand Down
Loading
Loading