mirror of
https://gitverse.ru/kpa39l/rag-pipeline-docker.git
synced 2026-09-29 09:15:11 +00:00
Initial commit: Hermes skill rag-pipeline-docker
This commit is contained in:
@@ -0,0 +1,81 @@
|
||||
# Qdrant: version matching + multiple collections (verified 2026-09-04)
|
||||
|
||||
## qdrant-client must match the server minor version
|
||||
|
||||
Symptom chain with client 1.19.0 vs server 1.17.1 (`qdrant/qdrant:v1.17.1`):
|
||||
- `Qdrant client version 1.19.0 is incompatible with server version 1.17.1` warning,
|
||||
- `create_collection` with a bare `models.VectorParams(size=1024, ...)` silently
|
||||
creates an ANONYMOUS vector (name `""`), NOT `dense`,
|
||||
- the subsequent `upsert` fails: `400 ... Not existing vector name error: dense`.
|
||||
|
||||
Fix — pin the client to the server minor:
|
||||
```bash
|
||||
pip install "qdrant-client==1.17.1" # match qdrant/qdrant:v1.17.1
|
||||
```
|
||||
|
||||
Always pass NAMED vectors so the config works regardless of client version:
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
c = QdrantClient("http://localhost:6333")
|
||||
c.create_collection(
|
||||
collection_name=NAME,
|
||||
vectors_config={
|
||||
"dense": models.VectorParams(size=1024, distance=models.Distance.COSINE),
|
||||
},
|
||||
sparse_vectors_config={
|
||||
"sparse": models.SparseVectorParams(index=models.SparseIndexParams(on_disk=True)),
|
||||
},
|
||||
)
|
||||
```
|
||||
|
||||
## Query API in qdrant-client 1.17
|
||||
|
||||
- `client.search(...)` does NOT exist in 1.17.
|
||||
- `query_points(..., query_vector=...)` → `AssertionError: Unknown arguments: ['query_vector']`.
|
||||
- Working call:
|
||||
```python
|
||||
res = c.query_points(
|
||||
collection_name=NAME,
|
||||
query=<dense_embedding_list>, # list[float] from Ollama /api/embeddings
|
||||
using="dense", # named-vector selector
|
||||
limit=5,
|
||||
with_payload=True,
|
||||
)
|
||||
for pt in res.points:
|
||||
print(pt.score, pt.payload.get("text"))
|
||||
```
|
||||
|
||||
## One collection per project/domain (multi-collection design)
|
||||
|
||||
For a distinct document set (batch of PDF protocol/files), create a SEPARATE
|
||||
collection with the SAME schema (`dense` 1024d COSINE + sparse `sparse` BM25) and the
|
||||
SAME embedder (bge-m3). Keep search query embeddings compatible by using the same
|
||||
embedder for all collections.
|
||||
Benefits: independent re-index, per-domain context search, no pollution of the
|
||||
general KB. Example: `skc_vinny_gorod` alongside `knowledge_base`.
|
||||
|
||||
## Do NOT retarget context_enhancer to a second collection
|
||||
|
||||
`context_enhancer.py` binds `COLLECTION = os.environ.get("QDRANT_COLLECTION",
|
||||
"knowledge_base")` at IMPORT time (module level). Swapping the env var at runtime
|
||||
does NOT retarget it — the module-level constant is already fixed.
|
||||
|
||||
To search a second collection, write a STANDALONE REST search:
|
||||
1. `POST http://ollama:11434/api/embeddings` `{"model": "bge-m3", "prompt": text}` → `embedding` (1024d).
|
||||
2. `POST http://qdrant:6333/collections/<NAME>/points/query` with `{"vector": emb, "limit": N, "with_payload": true, "using": "dense"}`.
|
||||
3. Read `result.points[*].payload` + `.score`.
|
||||
|
||||
This same pattern is the foundation for a per-project context-injector hook
|
||||
(e.g. NetBox project context) — hit the secondary collection directly, don't go
|
||||
through the shared KB search.
|
||||
|
||||
## Scanned PDFs: pymupdf returns empty text (no text layer)
|
||||
|
||||
`page.get_text("text")` returns `""` for image-only pages — verified on a real
|
||||
51-page government PDF (0 text on every page). It is a genuine scan, not a glitch.
|
||||
Mark scanned pages `[SCANNED_PAGE]` and route to OCR (marker-pdf / vision);
|
||||
do NOT report "no content" or fabricate text. pymupdf's built-in Type1 fonts
|
||||
(times-roman, helv, cour, tiro) do NOT contain the Cyrillic glyph map — inserting
|
||||
Cyrillic with them renders as dots and re-extracts as dots. Real PDFs (generated
|
||||
from Word/CAD) embed proper fonts and extract Cyrillic fine; the byte test above is
|
||||
only a pymupdf-font artifact, not a real-PDF problem.
|
||||
Reference in New Issue
Block a user