Files
memory-os/references/ingest-debug-session.md
T
2026-09-06 13:51:22 +00:00

8.8 KiB

Ingest Debugging Sessions

Chronological reproduction of the pipeline from zero to working search. Each session covers issues encountered and fixes applied.


Session 1: 2026-06-20 to 2026-06-21 (Initial Pipeline Bring-Up)

Full reproduction of the pipeline from zero to working search.

Initial State

  • Qdrant collection knowledge_base exists with dense (768d) + sparse vectors
  • wiki_ingest_state.json has 467 files, all with ingested_at set
  • Qdrant points_count = 0 — nothing actually stored
  • ARQ queue empty — stale jobs consumed
  • Worker container healthy but was crashing on DNS

Issue 1: Worker DNS Failure

Symptom: Worker logs show Temporary failure in name resolution when accessing ollama:11434

Root cause: Worker container not connected to ollama_default Docker network. docker compose up -d recreating the container detaches secondary networks.

Fix:

docker network connect ollama_default docker-worker-1

Verify:

docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys'

Should show both docker_default and ollama_default.

Issue 2: Sparse Vector Name Mismatch

Symptom: Worker completes but Qdrant points_count stays at 0. Logs show HTTP 400 errors.

Root cause: The Qdrant collection was created with sparse vectors named "bm25", but the ingest worker sends sparse vectors keyed as "sparse":

# What the collection had:
sparse_vectors_config={'bm25': SparseVectorParams(...)}

# What the worker sends:
{'points': [..., 'vector': {'sparse': ..., 'dense': ...}]}

Fix: Recreate collection with consistent naming:

c.delete_collection('knowledge_base')
c.create_collection(
    collection_name='knowledge_base',
    vectors_config=VectorParams(size=768, distance=Distance.COSINE),
    sparse_vectors_config={'sparse': SparseVectorParams(
        index=SparseIndexParams(on_disk=False, full_scan_threshold=10000)
    )}
)

Symptom: After ingest works and Qdrant has 683 points, context_enhancer.py search returns only low-score or no results, falling back to lexical search.

Root cause: context_enhancer.py was configured to use OpenRouter qwen/qwen3-embedding-8b for query embedding, but the ingested collection uses Ollama nomic-embed-text:latest (768d). Different models produce incompatible vector spaces — cosine similarity is near zero.

Fix in context_enhancer.py:

# Before (OpenRouter):
EMBEDDING_MODEL = "qwen/qwen3-embedding-8b"
resp = requests.post("https://openrouter.ai/api/v1/embeddings", ...)

# After (local Ollama):
OLLAMA_EMBEDDING_URL = "http://localhost:11434"
OLLAMA_EMBEDDING_MODEL = "nomic-embed-text:latest"
resp = requests.post(f"{OLLAMA_EMBEDDING_URL}/api/embeddings",
    json={"model": OLLAMA_EMBEDDING_MODEL, "prompt": text})

Also clean up icarus/hooks.py which had OPENROUTER_API_KEY env manipulation that became dead code.

Issue 4: FastEmbed BM25 Subprocess Failure

Symptom: Sparse embedding returns None silently. context_enhancer.py falls back to dense-only or lexical.

Root cause: The subprocess relies on FASTEMBED_SITEPKGS env var pointing to the ai-lab venv site-packages, but subprocess.run() does NOT inherit the parent's env vars automatically when the env was set in Python (not the shell).

# BROKEN — subprocess doesn't see FASTEMBED_SITEPKGS:
result = subprocess.run(
    [_FASTEMBED_PYTHON, "-c", "...import fastembed..."],
    input=text, capture_output=True, text=True, timeout=15
)

# FIXED — explicitly pass env:
_env = os.environ.copy()
_env.setdefault("FASTEMBED_SITEPKGS", _FASTEMBED_SITEPKGS)
result = subprocess.run(
    [_FASTEMBED_PYTHON, "-c", "..."],
    input=text, capture_output=True, text=True, timeout=15, env=_env
)

Issue 5: Score Threshold Too High

Symptom: nomic-embed-text results have scores in the 0.30-0.56 range, filter out most results.

Fix: Lowered default threshold from 0.55 to 0.35.

Healthy Config (Final)

Qdrant:     localhost:6333, collection="knowledge_base", 683 points
Redis:      127.0.0.1:6379 (authenticated)
Ollama:     localhost:11434, model="nomic-embed-text:latest" (768d)
Worker:     docker-worker-1, connected to ollama_default network
FastEmbed:  BM25 via ai-lab venv subprocess
Cron:       wiki-ingest-sync, every 10 minutes
State:      ~/.hermes/wiki_ingest_state.json (tracking 467 files)

Session 2: 2026-07-15 (Three New Blockers — python-dotenv, Sparse Embed Path, Hermes Hook Integration)

Reproduced the stack from scratch. Three blocking issues found on top of the original five.

Blocker 1: python-dotenv Missing in Cron Environment

Symptom: cron wiki-ingest-sync job fails silently. Logs: ModuleNotFoundError: No module named 'dotenv'.

Root cause: The cron job runs under the system environment, which does NOT have python-dotenv installed. The wiki_continuous_ingest.py script imports from dotenv import load_dotenv at the top.

Fix: Install python-dotenv system-wide:

pip install python-dotenv

Or patch the script to load .env via os.environ + open() instead of python-dotenv.

Verification:

python3 -c "from dotenv import load_dotenv; print('OK')"

Blocker 2: Sparse Embedding Returns Expecting value: line 1 column 1

Symptom: context_enhancer.py sparse embedding crashes mid-search with json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0).

Root cause: The FastEmbed subprocess (_FASTEMBED_PYTHON) runs a Python script that tries to import fastembed and import dotenv. The subprocess's python path points to /usr/bin/python3 which does NOT have fastembed or dotenv in its site-packages. The subprocess fails silently (prints nothing to stdout), and the parent tries json.loads(result.stdout) on empty output.

The existing fix from Session 1 (Issue #4 — FASTEMBED_SITEPKGS env var) may have been applied, but the fundamental problem is the subprocess PYTHON PATH itself, not the env var. If /usr/bin/python3 cannot import fastembed at all, setting the env var won't help.

Fix: Use the ai-lab venv python directly:

_FASTEMBED_PYTHON = "/opt/ai-lab/.venv/bin/python3"

instead of:

_FASTEMBED_PYTHON = "/usr/bin/python3"

This venv has both fastembed and python-dotenv installed.

Diagnostic:

/opt/ai-lab/.venv/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
/usr/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"

Blocker 3: Hermes Hook (icarus/hooks.py) Not Connected to Hermes

Symptom: The file /opt/hermes/.hermes/plugins/icarus/hooks.py exists with all the right logic (search Qdrant, inject context into Hermes responses), but it's NEVER called. No Hermes config, plugin, or hook registration activates it.

Root cause: icarus/hooks.py relies on being loaded by Hermes as a user plugin, but it's not enabled. The plugin dir is at /opt/hermes/.hermes/plugins/icarus/ with its own plugin.yaml (v0.3.0, 16 tools, 4 hooks). Two things were needed:

  1. hermes plugins enable icarus — the proper CLI command to activate a user plugin
  2. PYTHONPATH fix inside hooks.py — the _search_qdrant() function does from scripts.context_enhancer import ... but /opt/hermes/memory-os is not on sys.path. Must add import sys at top and sys.path.insert(0, '/opt/hermes/memory-os') inside _search_qdrant() before the import. Without this, the import raises ModuleNotFoundError and _search_qdrant() returns empty list silently (fail-open).

Fix — two steps:

Step 1 — PYTHONPATH in hooks.py:

# Add at top of hooks.py:
import sys

# Add inside _search_qdrant(), before the import:
_MEMORY_OS = "/opt/hermes/memory-os"
if _MEMORY_OS not in sys.path:
    sys.path.insert(0, _MEMORY_OS)

Step 2 — Enable plugin:

hermes plugins enable icarus
# Takes effect on next session. No config.yaml editing needed.

Verification of PYTHONPATH fix:

python3 -c "
import sys
sys.path.insert(0, '/opt/hermes/memory-os')
from scripts.context_enhancer import embed_query, embed_query_sparse
print('context_enhancer import: OK')
dense = embed_query('test')
print(f'Dense: {len(dense)} dims')
sparse = embed_query_sparse('test')
print(f'Sparse: {len(sparse)} values')
"

Verification of plugin status:

hermes plugins list | grep icarus
# Should show: icarus │ enabled │ 0.3.0

Qdrant Collection Verification

python3 -c "
from qdrant_client import QdrantClient, models
c = QdrantClient('http://localhost:6333')
info = c.get_collection('knowledge_base')
print(f'Points: {info.points_count}')
print(f'Status: {info.status}')
print(f'Dense dims: {info.config.params.vectors.size}')
print(f'Sparse keys: {list(info.config.params.sparse_vectors.keys())}')
"