Motivation
The FAISS backend today uses IndexIVFFlat, which stores every vector at full float32 precision. For the default embedding model (Qwen/Qwen3-VL-Embedding-2B, dim = 2048), a 30 M-vector corpus requires roughly 245 GB of RAM just to hold the index — far outside the reach of air-gapped or Docker-free deployments that cannot run a Qdrant service.
FAISS ships IndexIVFPQ (Product Quantization) out of the box. Switching to PQ with M=64, nbits=8 compresses each vector from 8 192 bytes to 64 bytes (128× reduction, ~1.9 GB for 30 M vectors) at a typical recall-at-10 drop of ≤ 5–8% on inner-product tasks — a trade-off Qdrant users can already make today via --qdrant-quantization-config. Adding an equivalent flag to the FAISS path closes that gap for users who cannot run a Docker sidecar.
Requested change
Add opt-in --pq-m / --pq-nbits flags to the faiss build path so users can select IndexIVFPQ instead of IndexIVFFlat.
CLI — embed/src/pixelrag_embed/index.py
# inside the 'build' subparser, faiss-only flags
+p_build.add_argument(
+ "--pq-m",
+ type=int,
+ default=0,
+ help="PQ sub-quantizers (0 = IVFFlat, no compression). "
+ "Must divide the embedding dim evenly. "
+ "Typical: dim/32 — e.g. 64 for dim 2048.",
+)
+p_build.add_argument(
+ "--pq-nbits",
+ type=int,
+ default=8,
+ choices=[4, 8],
+ help="Bits per PQ sub-quantizer (default: 8).",
+)
In build_ivf(), when pq_m > 0, replace:
index = faiss.IndexIVFFlat(quantizer, dim, nlist, metric_type)
with:
index = faiss.IndexIVFPQ(quantizer, dim, nlist, pq_m, pq_nbits)
Config — pixelrag.yaml
index:
backend: faiss
pq_m: 64 # sub-quantizers; must divide dim evenly
pq_nbits: 8 # 4 or 8 (default 8)
Pipeline forwarding — index/src/pixelrag_index/pipelines.py
The faiss branch of the cmd build in build() already passes --nlist. It would additionally pass --pq-m / --pq-nbits when they appear in index_cfg.
Why not just use Qdrant?
Qdrant scalar/binary quantization via --qdrant-quantization-config is an excellent solution for connected deployments. IVF+PQ covers the complementary case: no Docker available, no network service permitted, fully air-gapped lab or edge environments. Both users deserve a compression path; this closes the gap on the FAISS side.
Memory impact at a glance
| Index type |
M (sub-quantizers) |
Bytes / vector |
30 M vectors |
Recall@10 (approx.) |
| IVFFlat (current) |
— |
8 192 |
~245 GB |
baseline |
| IVFPQ nbits=8 |
128 |
128 |
~3.8 GB |
~96% |
| IVFPQ nbits=8 |
64 |
64 |
~1.9 GB |
~92% |
| IVFPQ nbits=4 |
64 |
32 |
~0.95 GB |
~87% |
Figures are illustrative for inner-product search on visual-embedding benchmarks; actual recall depends on the corpus. Even the most aggressive setting (M=64, nbits=4) makes a 30 M-vector corpus fit on a workstation with 4 GB of free RAM.
Scope
embed/src/pixelrag_embed/index.py: build_ivf() function + arg parser — approximately 25 lines changed/added.
index/src/pixelrag_index/pipelines.py: forward pq_m / pq_nbits from index_cfg — approximately 6 lines.
- No new dependencies:
faiss-cpu and faiss-gpu already include IndexIVFPQ.
Happy to open a PR if this direction looks good to the maintainers.
Motivation
The FAISS backend today uses
IndexIVFFlat, which stores every vector at full float32 precision. For the default embedding model (Qwen/Qwen3-VL-Embedding-2B, dim = 2048), a 30 M-vector corpus requires roughly 245 GB of RAM just to hold the index — far outside the reach of air-gapped or Docker-free deployments that cannot run a Qdrant service.FAISS ships
IndexIVFPQ(Product Quantization) out of the box. Switching to PQ withM=64, nbits=8compresses each vector from 8 192 bytes to 64 bytes (128× reduction, ~1.9 GB for 30 M vectors) at a typical recall-at-10 drop of ≤ 5–8% on inner-product tasks — a trade-off Qdrant users can already make today via--qdrant-quantization-config. Adding an equivalent flag to the FAISS path closes that gap for users who cannot run a Docker sidecar.Requested change
Add opt-in
--pq-m/--pq-nbitsflags to thefaissbuild path so users can selectIndexIVFPQinstead ofIndexIVFFlat.CLI —
embed/src/pixelrag_embed/index.pyIn
build_ivf(), whenpq_m > 0, replace:with:
Config —
pixelrag.yamlPipeline forwarding —
index/src/pixelrag_index/pipelines.pyThe
faissbranch of thecmdbuild inbuild()already passes--nlist. It would additionally pass--pq-m/--pq-nbitswhen they appear inindex_cfg.Why not just use Qdrant?
Qdrant scalar/binary quantization via
--qdrant-quantization-configis an excellent solution for connected deployments. IVF+PQ covers the complementary case: no Docker available, no network service permitted, fully air-gapped lab or edge environments. Both users deserve a compression path; this closes the gap on the FAISS side.Memory impact at a glance
Figures are illustrative for inner-product search on visual-embedding benchmarks; actual recall depends on the corpus. Even the most aggressive setting (M=64, nbits=4) makes a 30 M-vector corpus fit on a workstation with 4 GB of free RAM.
Scope
embed/src/pixelrag_embed/index.py:build_ivf()function + arg parser — approximately 25 lines changed/added.index/src/pixelrag_index/pipelines.py: forwardpq_m/pq_nbitsfromindex_cfg— approximately 6 lines.faiss-cpuandfaiss-gpualready includeIndexIVFPQ.Happy to open a PR if this direction looks good to the maintainers.