Initial commit: Hermes skill rag-pipeline-docker

This commit is contained in:
estorozhenko
2026-09-06 13:51:06 +00:00
commit bdc03a5d59
6 changed files with 656 additions and 0 deletions
@@ -0,0 +1,81 @@
# Qdrant: version matching + multiple collections (verified 2026-09-04)
## qdrant-client must match the server minor version
Symptom chain with client 1.19.0 vs server 1.17.1 (`qdrant/qdrant:v1.17.1`):
- `Qdrant client version 1.19.0 is incompatible with server version 1.17.1` warning,
- `create_collection` with a bare `models.VectorParams(size=1024, ...)` silently
creates an ANONYMOUS vector (name `""`), NOT `dense`,
- the subsequent `upsert` fails: `400 ... Not existing vector name error: dense`.
Fix — pin the client to the server minor:
```bash
pip install "qdrant-client==1.17.1" # match qdrant/qdrant:v1.17.1
```
Always pass NAMED vectors so the config works regardless of client version:
```python
from qdrant_client import QdrantClient, models
c = QdrantClient("http://localhost:6333")
c.create_collection(
collection_name=NAME,
vectors_config={
"dense": models.VectorParams(size=1024, distance=models.Distance.COSINE),
},
sparse_vectors_config={
"sparse": models.SparseVectorParams(index=models.SparseIndexParams(on_disk=True)),
},
)
```
## Query API in qdrant-client 1.17
- `client.search(...)` does NOT exist in 1.17.
- `query_points(..., query_vector=...)` → `AssertionError: Unknown arguments: ['query_vector']`.
- Working call:
```python
res = c.query_points(
collection_name=NAME,
query=<dense_embedding_list>, # list[float] from Ollama /api/embeddings
using="dense", # named-vector selector
limit=5,
with_payload=True,
)
for pt in res.points:
print(pt.score, pt.payload.get("text"))
```
## One collection per project/domain (multi-collection design)
For a distinct document set (batch of PDF protocol/files), create a SEPARATE
collection with the SAME schema (`dense` 1024d COSINE + sparse `sparse` BM25) and the
SAME embedder (bge-m3). Keep search query embeddings compatible by using the same
embedder for all collections.
Benefits: independent re-index, per-domain context search, no pollution of the
general KB. Example: `skc_vinny_gorod` alongside `knowledge_base`.
## Do NOT retarget context_enhancer to a second collection
`context_enhancer.py` binds `COLLECTION = os.environ.get("QDRANT_COLLECTION",
"knowledge_base")` at IMPORT time (module level). Swapping the env var at runtime
does NOT retarget it — the module-level constant is already fixed.
To search a second collection, write a STANDALONE REST search:
1. `POST http://ollama:11434/api/embeddings` `{"model": "bge-m3", "prompt": text}` → `embedding` (1024d).
2. `POST http://qdrant:6333/collections/<NAME>/points/query` with `{"vector": emb, "limit": N, "with_payload": true, "using": "dense"}`.
3. Read `result.points[*].payload` + `.score`.
This same pattern is the foundation for a per-project context-injector hook
(e.g. NetBox project context) — hit the secondary collection directly, don't go
through the shared KB search.
## Scanned PDFs: pymupdf returns empty text (no text layer)
`page.get_text("text")` returns `""` for image-only pages — verified on a real
51-page government PDF (0 text on every page). It is a genuine scan, not a glitch.
Mark scanned pages `[SCANNED_PAGE]` and route to OCR (marker-pdf / vision);
do NOT report "no content" or fabricate text. pymupdf's built-in Type1 fonts
(times-roman, helv, cour, tiro) do NOT contain the Cyrillic glyph map — inserting
Cyrillic with them renders as dots and re-extracts as dots. Real PDFs (generated
from Word/CAD) embed proper fonts and extract Cyrillic fine; the byte test above is
only a pymupdf-font artifact, not a real-PDF problem.