# Ingest Debugging Sessions Chronological reproduction of the pipeline from zero to working search. Each session covers issues encountered and fixes applied. --- ## Session 1: 2026-06-20 to 2026-06-21 (Initial Pipeline Bring-Up) Full reproduction of the pipeline from zero to working search. ## Initial State - Qdrant collection `knowledge_base` exists with dense (768d) + sparse vectors - `wiki_ingest_state.json` has 467 files, all with `ingested_at` set - Qdrant points_count = 0 — nothing actually stored - ARQ queue empty — stale jobs consumed - Worker container healthy but was crashing on DNS ## Issue 1: Worker DNS Failure **Symptom:** Worker logs show `Temporary failure in name resolution` when accessing `ollama:11434` **Root cause:** Worker container not connected to `ollama_default` Docker network. `docker compose up -d` recreating the container detaches secondary networks. **Fix:** ```bash docker network connect ollama_default docker-worker-1 ``` Verify: ```bash docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys' ``` Should show both `docker_default` and `ollama_default`. ## Issue 2: Sparse Vector Name Mismatch **Symptom:** Worker completes but Qdrant points_count stays at 0. Logs show HTTP 400 errors. **Root cause:** The Qdrant collection was created with sparse vectors named `"bm25"`, but the ingest worker sends sparse vectors keyed as `"sparse"`: ```python # What the collection had: sparse_vectors_config={'bm25': SparseVectorParams(...)} # What the worker sends: {'points': [..., 'vector': {'sparse': ..., 'dense': ...}]} ``` **Fix:** Recreate collection with consistent naming: ```python c.delete_collection('knowledge_base') c.create_collection( collection_name='knowledge_base', vectors_config=VectorParams(size=768, distance=Distance.COSINE), sparse_vectors_config={'sparse': SparseVectorParams( index=SparseIndexParams(on_disk=False, full_scan_threshold=10000) )} ) ``` ## Issue 3: Dense Embedding Model Mismatch (Search) **Symptom:** After ingest works and Qdrant has 683 points, `context_enhancer.py` search returns only low-score or no results, falling back to lexical search. **Root cause:** `context_enhancer.py` was configured to use OpenRouter `qwen/qwen3-embedding-8b` for query embedding, but the ingested collection uses Ollama `nomic-embed-text:latest` (768d). Different models produce incompatible vector spaces — cosine similarity is near zero. **Fix in context_enhancer.py:** ```python # Before (OpenRouter): EMBEDDING_MODEL = "qwen/qwen3-embedding-8b" resp = requests.post("https://openrouter.ai/api/v1/embeddings", ...) # After (local Ollama): OLLAMA_EMBEDDING_URL = "http://localhost:11434" OLLAMA_EMBEDDING_MODEL = "nomic-embed-text:latest" resp = requests.post(f"{OLLAMA_EMBEDDING_URL}/api/embeddings", json={"model": OLLAMA_EMBEDDING_MODEL, "prompt": text}) ``` Also clean up icarus/hooks.py which had OPENROUTER_API_KEY env manipulation that became dead code. ## Issue 4: FastEmbed BM25 Subprocess Failure **Symptom:** Sparse embedding returns None silently. `context_enhancer.py` falls back to dense-only or lexical. **Root cause:** The subprocess relies on `FASTEMBED_SITEPKGS` env var pointing to the ai-lab venv site-packages, but subprocess.run() does NOT inherit the parent's env vars automatically when the env was set in Python (not the shell). ```python # BROKEN — subprocess doesn't see FASTEMBED_SITEPKGS: result = subprocess.run( [_FASTEMBED_PYTHON, "-c", "...import fastembed..."], input=text, capture_output=True, text=True, timeout=15 ) # FIXED — explicitly pass env: _env = os.environ.copy() _env.setdefault("FASTEMBED_SITEPKGS", _FASTEMBED_SITEPKGS) result = subprocess.run( [_FASTEMBED_PYTHON, "-c", "..."], input=text, capture_output=True, text=True, timeout=15, env=_env ) ``` ## Issue 5: Score Threshold Too High **Symptom:** nomic-embed-text results have scores in the 0.30-0.56 range, filter out most results. **Fix:** Lowered default threshold from 0.55 to 0.35. ## Healthy Config (Final) ``` Qdrant: localhost:6333, collection="knowledge_base", 683 points Redis: 127.0.0.1:6379 (authenticated) Ollama: localhost:11434, model="nomic-embed-text:latest" (768d) Worker: docker-worker-1, connected to ollama_default network FastEmbed: BM25 via ai-lab venv subprocess Cron: wiki-ingest-sync, every 10 minutes State: ~/.hermes/wiki_ingest_state.json (tracking 467 files) ``` --- ## Session 2: 2026-07-15 (Three New Blockers — python-dotenv, Sparse Embed Path, Hermes Hook Integration) Reproduced the stack from scratch. Three blocking issues found on top of the original five. ### Blocker 1: `python-dotenv` Missing in Cron Environment **Symptom:** cron `wiki-ingest-sync` job fails silently. Logs: `ModuleNotFoundError: No module named 'dotenv'`. **Root cause:** The cron job runs under the system environment, which does NOT have `python-dotenv` installed. The `wiki_continuous_ingest.py` script imports `from dotenv import load_dotenv` at the top. **Fix:** Install `python-dotenv` system-wide: ```bash pip install python-dotenv ``` Or patch the script to load `.env` via `os.environ` + `open()` instead of `python-dotenv`. **Verification:** ```bash python3 -c "from dotenv import load_dotenv; print('OK')" ``` ### Blocker 2: Sparse Embedding Returns `Expecting value: line 1 column 1` **Symptom:** `context_enhancer.py` sparse embedding crashes mid-search with `json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)`. **Root cause:** The FastEmbed subprocess (`_FASTEMBED_PYTHON`) runs a Python script that tries to `import fastembed` and `import dotenv`. The subprocess's python path points to `/usr/bin/python3` which does NOT have `fastembed` or `dotenv` in its site-packages. The subprocess fails silently (prints nothing to stdout), and the parent tries `json.loads(result.stdout)` on empty output. The existing fix from Session 1 (Issue #4 — `FASTEMBED_SITEPKGS` env var) may have been applied, but the fundamental problem is the subprocess PYTHON PATH itself, not the env var. If `/usr/bin/python3` cannot `import fastembed` at all, setting the env var won't help. **Fix:** Use the ai-lab venv python directly: ```python _FASTEMBED_PYTHON = "/opt/ai-lab/.venv/bin/python3" ``` instead of: ```python _FASTEMBED_PYTHON = "/usr/bin/python3" ``` This venv has both `fastembed` and `python-dotenv` installed. **Diagnostic:** ```bash /opt/ai-lab/.venv/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')" /usr/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')" ``` ### Blocker 3: Hermes Hook (`icarus/hooks.py`) Not Connected to Hermes **Symptom:** The file `/opt/hermes/.hermes/plugins/icarus/hooks.py` exists with all the right logic (search Qdrant, inject context into Hermes responses), but it's NEVER called. No Hermes config, plugin, or hook registration activates it. **Root cause:** `icarus/hooks.py` relies on being loaded by Hermes as a user plugin, but it's not enabled. The plugin dir is at `/opt/hermes/.hermes/plugins/icarus/` with its own `plugin.yaml` (v0.3.0, 16 tools, 4 hooks). Two things were needed: 1. `hermes plugins enable icarus` — the proper CLI command to activate a user plugin 2. PYTHONPATH fix inside `hooks.py` — the `_search_qdrant()` function does `from scripts.context_enhancer import ...` but `/opt/hermes/memory-os` is not on `sys.path`. Must add `import sys` at top and `sys.path.insert(0, '/opt/hermes/memory-os')` inside `_search_qdrant()` before the import. Without this, the import raises `ModuleNotFoundError` and `_search_qdrant()` returns empty list silently (fail-open). **Fix — two steps:** Step 1 — PYTHONPATH in hooks.py: ```python # Add at top of hooks.py: import sys # Add inside _search_qdrant(), before the import: _MEMORY_OS = "/opt/hermes/memory-os" if _MEMORY_OS not in sys.path: sys.path.insert(0, _MEMORY_OS) ``` Step 2 — Enable plugin: ```bash hermes plugins enable icarus # Takes effect on next session. No config.yaml editing needed. ``` **Verification of PYTHONPATH fix:** ```bash python3 -c " import sys sys.path.insert(0, '/opt/hermes/memory-os') from scripts.context_enhancer import embed_query, embed_query_sparse print('context_enhancer import: OK') dense = embed_query('test') print(f'Dense: {len(dense)} dims') sparse = embed_query_sparse('test') print(f'Sparse: {len(sparse)} values') " ``` **Verification of plugin status:** ```bash hermes plugins list | grep icarus # Should show: icarus │ enabled │ 0.3.0 ``` ## Qdrant Collection Verification ```bash python3 -c " from qdrant_client import QdrantClient, models c = QdrantClient('http://localhost:6333') info = c.get_collection('knowledge_base') print(f'Points: {info.points_count}') print(f'Status: {info.status}') print(f'Dense dims: {info.config.params.vectors.size}') print(f'Sparse keys: {list(info.config.params.sparse_vectors.keys())}') " ```