mirror of
https://gitverse.ru/kpa39l/memory-os.git
synced 2026-09-29 09:35:05 +00:00
Initial commit: Hermes skill memory-os
This commit is contained in:
@@ -0,0 +1,241 @@
|
||||
# Ingest Debugging Sessions
|
||||
|
||||
Chronological reproduction of the pipeline from zero to working search. Each session
|
||||
covers issues encountered and fixes applied.
|
||||
|
||||
---
|
||||
|
||||
## Session 1: 2026-06-20 to 2026-06-21 (Initial Pipeline Bring-Up)
|
||||
|
||||
Full reproduction of the pipeline from zero to working search.
|
||||
|
||||
## Initial State
|
||||
|
||||
- Qdrant collection `knowledge_base` exists with dense (768d) + sparse vectors
|
||||
- `wiki_ingest_state.json` has 467 files, all with `ingested_at` set
|
||||
- Qdrant points_count = 0 — nothing actually stored
|
||||
- ARQ queue empty — stale jobs consumed
|
||||
- Worker container healthy but was crashing on DNS
|
||||
|
||||
## Issue 1: Worker DNS Failure
|
||||
|
||||
**Symptom:** Worker logs show `Temporary failure in name resolution` when accessing `ollama:11434`
|
||||
|
||||
**Root cause:** Worker container not connected to `ollama_default` Docker network.
|
||||
`docker compose up -d` recreating the container detaches secondary networks.
|
||||
|
||||
**Fix:**
|
||||
```bash
|
||||
docker network connect ollama_default docker-worker-1
|
||||
```
|
||||
Verify:
|
||||
```bash
|
||||
docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys'
|
||||
```
|
||||
Should show both `docker_default` and `ollama_default`.
|
||||
|
||||
## Issue 2: Sparse Vector Name Mismatch
|
||||
|
||||
**Symptom:** Worker completes but Qdrant points_count stays at 0. Logs show HTTP 400 errors.
|
||||
|
||||
**Root cause:** The Qdrant collection was created with sparse vectors named `"bm25"`,
|
||||
but the ingest worker sends sparse vectors keyed as `"sparse"`:
|
||||
|
||||
```python
|
||||
# What the collection had:
|
||||
sparse_vectors_config={'bm25': SparseVectorParams(...)}
|
||||
|
||||
# What the worker sends:
|
||||
{'points': [..., 'vector': {'sparse': ..., 'dense': ...}]}
|
||||
```
|
||||
|
||||
**Fix:** Recreate collection with consistent naming:
|
||||
```python
|
||||
c.delete_collection('knowledge_base')
|
||||
c.create_collection(
|
||||
collection_name='knowledge_base',
|
||||
vectors_config=VectorParams(size=768, distance=Distance.COSINE),
|
||||
sparse_vectors_config={'sparse': SparseVectorParams(
|
||||
index=SparseIndexParams(on_disk=False, full_scan_threshold=10000)
|
||||
)}
|
||||
)
|
||||
```
|
||||
|
||||
## Issue 3: Dense Embedding Model Mismatch (Search)
|
||||
|
||||
**Symptom:** After ingest works and Qdrant has 683 points, `context_enhancer.py` search
|
||||
returns only low-score or no results, falling back to lexical search.
|
||||
|
||||
**Root cause:** `context_enhancer.py` was configured to use OpenRouter `qwen/qwen3-embedding-8b`
|
||||
for query embedding, but the ingested collection uses Ollama `nomic-embed-text:latest` (768d).
|
||||
Different models produce incompatible vector spaces — cosine similarity is near zero.
|
||||
|
||||
**Fix in context_enhancer.py:**
|
||||
```python
|
||||
# Before (OpenRouter):
|
||||
EMBEDDING_MODEL = "qwen/qwen3-embedding-8b"
|
||||
resp = requests.post("https://openrouter.ai/api/v1/embeddings", ...)
|
||||
|
||||
# After (local Ollama):
|
||||
OLLAMA_EMBEDDING_URL = "http://localhost:11434"
|
||||
OLLAMA_EMBEDDING_MODEL = "nomic-embed-text:latest"
|
||||
resp = requests.post(f"{OLLAMA_EMBEDDING_URL}/api/embeddings",
|
||||
json={"model": OLLAMA_EMBEDDING_MODEL, "prompt": text})
|
||||
```
|
||||
|
||||
Also clean up icarus/hooks.py which had OPENROUTER_API_KEY env manipulation that
|
||||
became dead code.
|
||||
|
||||
## Issue 4: FastEmbed BM25 Subprocess Failure
|
||||
|
||||
**Symptom:** Sparse embedding returns None silently. `context_enhancer.py` falls back to
|
||||
dense-only or lexical.
|
||||
|
||||
**Root cause:** The subprocess relies on `FASTEMBED_SITEPKGS` env var pointing to the
|
||||
ai-lab venv site-packages, but subprocess.run() does NOT inherit the parent's env vars
|
||||
automatically when the env was set in Python (not the shell).
|
||||
|
||||
```python
|
||||
# BROKEN — subprocess doesn't see FASTEMBED_SITEPKGS:
|
||||
result = subprocess.run(
|
||||
[_FASTEMBED_PYTHON, "-c", "...import fastembed..."],
|
||||
input=text, capture_output=True, text=True, timeout=15
|
||||
)
|
||||
|
||||
# FIXED — explicitly pass env:
|
||||
_env = os.environ.copy()
|
||||
_env.setdefault("FASTEMBED_SITEPKGS", _FASTEMBED_SITEPKGS)
|
||||
result = subprocess.run(
|
||||
[_FASTEMBED_PYTHON, "-c", "..."],
|
||||
input=text, capture_output=True, text=True, timeout=15, env=_env
|
||||
)
|
||||
```
|
||||
|
||||
## Issue 5: Score Threshold Too High
|
||||
|
||||
**Symptom:** nomic-embed-text results have scores in the 0.30-0.56 range, filter out
|
||||
most results.
|
||||
|
||||
**Fix:** Lowered default threshold from 0.55 to 0.35.
|
||||
|
||||
## Healthy Config (Final)
|
||||
|
||||
```
|
||||
Qdrant: localhost:6333, collection="knowledge_base", 683 points
|
||||
Redis: 127.0.0.1:6379 (authenticated)
|
||||
Ollama: localhost:11434, model="nomic-embed-text:latest" (768d)
|
||||
Worker: docker-worker-1, connected to ollama_default network
|
||||
FastEmbed: BM25 via ai-lab venv subprocess
|
||||
Cron: wiki-ingest-sync, every 10 minutes
|
||||
State: ~/.hermes/wiki_ingest_state.json (tracking 467 files)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Session 2: 2026-07-15 (Three New Blockers — python-dotenv, Sparse Embed Path, Hermes Hook Integration)
|
||||
|
||||
Reproduced the stack from scratch. Three blocking issues found on top of the original five.
|
||||
|
||||
### Blocker 1: `python-dotenv` Missing in Cron Environment
|
||||
|
||||
**Symptom:** cron `wiki-ingest-sync` job fails silently. Logs: `ModuleNotFoundError: No module named 'dotenv'`.
|
||||
|
||||
**Root cause:** The cron job runs under the system environment, which does NOT have `python-dotenv` installed. The `wiki_continuous_ingest.py` script imports `from dotenv import load_dotenv` at the top.
|
||||
|
||||
**Fix:** Install `python-dotenv` system-wide:
|
||||
```bash
|
||||
pip install python-dotenv
|
||||
```
|
||||
Or patch the script to load `.env` via `os.environ` + `open()` instead of `python-dotenv`.
|
||||
|
||||
**Verification:**
|
||||
```bash
|
||||
python3 -c "from dotenv import load_dotenv; print('OK')"
|
||||
```
|
||||
|
||||
### Blocker 2: Sparse Embedding Returns `Expecting value: line 1 column 1`
|
||||
|
||||
**Symptom:** `context_enhancer.py` sparse embedding crashes mid-search with `json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)`.
|
||||
|
||||
**Root cause:** The FastEmbed subprocess (`_FASTEMBED_PYTHON`) runs a Python script that tries to `import fastembed` and `import dotenv`. The subprocess's python path points to `/usr/bin/python3` which does NOT have `fastembed` or `dotenv` in its site-packages. The subprocess fails silently (prints nothing to stdout), and the parent tries `json.loads(result.stdout)` on empty output.
|
||||
|
||||
The existing fix from Session 1 (Issue #4 — `FASTEMBED_SITEPKGS` env var) may have been applied, but the fundamental problem is the subprocess PYTHON PATH itself, not the env var. If `/usr/bin/python3` cannot `import fastembed` at all, setting the env var won't help.
|
||||
|
||||
**Fix:** Use the ai-lab venv python directly:
|
||||
```python
|
||||
_FASTEMBED_PYTHON = "/opt/ai-lab/.venv/bin/python3"
|
||||
```
|
||||
instead of:
|
||||
```python
|
||||
_FASTEMBED_PYTHON = "/usr/bin/python3"
|
||||
```
|
||||
|
||||
This venv has both `fastembed` and `python-dotenv` installed.
|
||||
|
||||
**Diagnostic:**
|
||||
```bash
|
||||
/opt/ai-lab/.venv/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
|
||||
/usr/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
|
||||
```
|
||||
|
||||
### Blocker 3: Hermes Hook (`icarus/hooks.py`) Not Connected to Hermes
|
||||
|
||||
**Symptom:** The file `/opt/hermes/.hermes/plugins/icarus/hooks.py` exists with all the right logic (search Qdrant, inject context into Hermes responses), but it's NEVER called. No Hermes config, plugin, or hook registration activates it.
|
||||
|
||||
**Root cause:** `icarus/hooks.py` relies on being loaded by Hermes as a user plugin, but it's not enabled. The plugin dir is at `/opt/hermes/.hermes/plugins/icarus/` with its own `plugin.yaml` (v0.3.0, 16 tools, 4 hooks). Two things were needed:
|
||||
|
||||
1. `hermes plugins enable icarus` — the proper CLI command to activate a user plugin
|
||||
2. PYTHONPATH fix inside `hooks.py` — the `_search_qdrant()` function does `from scripts.context_enhancer import ...` but `/opt/hermes/memory-os` is not on `sys.path`. Must add `import sys` at top and `sys.path.insert(0, '/opt/hermes/memory-os')` inside `_search_qdrant()` before the import. Without this, the import raises `ModuleNotFoundError` and `_search_qdrant()` returns empty list silently (fail-open).
|
||||
|
||||
**Fix — two steps:**
|
||||
|
||||
Step 1 — PYTHONPATH in hooks.py:
|
||||
```python
|
||||
# Add at top of hooks.py:
|
||||
import sys
|
||||
|
||||
# Add inside _search_qdrant(), before the import:
|
||||
_MEMORY_OS = "/opt/hermes/memory-os"
|
||||
if _MEMORY_OS not in sys.path:
|
||||
sys.path.insert(0, _MEMORY_OS)
|
||||
```
|
||||
|
||||
Step 2 — Enable plugin:
|
||||
```bash
|
||||
hermes plugins enable icarus
|
||||
# Takes effect on next session. No config.yaml editing needed.
|
||||
```
|
||||
|
||||
**Verification of PYTHONPATH fix:**
|
||||
```bash
|
||||
python3 -c "
|
||||
import sys
|
||||
sys.path.insert(0, '/opt/hermes/memory-os')
|
||||
from scripts.context_enhancer import embed_query, embed_query_sparse
|
||||
print('context_enhancer import: OK')
|
||||
dense = embed_query('test')
|
||||
print(f'Dense: {len(dense)} dims')
|
||||
sparse = embed_query_sparse('test')
|
||||
print(f'Sparse: {len(sparse)} values')
|
||||
"
|
||||
```
|
||||
|
||||
**Verification of plugin status:**
|
||||
```bash
|
||||
hermes plugins list | grep icarus
|
||||
# Should show: icarus │ enabled │ 0.3.0
|
||||
```
|
||||
|
||||
## Qdrant Collection Verification
|
||||
|
||||
```bash
|
||||
python3 -c "
|
||||
from qdrant_client import QdrantClient, models
|
||||
c = QdrantClient('http://localhost:6333')
|
||||
info = c.get_collection('knowledge_base')
|
||||
print(f'Points: {info.points_count}')
|
||||
print(f'Status: {info.status}')
|
||||
print(f'Dense dims: {info.config.params.vectors.size}')
|
||||
print(f'Sparse keys: {list(info.config.params.sparse_vectors.keys())}')
|
||||
"
|
||||
```
|
||||
Reference in New Issue
Block a user