24 KiB
name, description, version, author, metadata
| name | description | version | author | metadata | |||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| memory-os | Qdrant-based RAG pipeline for wiki/note vault ingestion — sync, parse, chunk, embed (dense: bge-m3 + sparse: BM25), store in Qdrant (1024d COSINE), and auto-inject retrieved context into Hermes responses. | 1.8.0 | Hermes Agent |
|
Memory OS — Wiki RAG Pipeline
A production RAG pipeline that continuously syncs an Obsidian vault (or any markdown wiki) into Qdrant for semantic + BM25 retrieval, and auto-injects relevant context into every Hermes response.
Architecture:
Obsidian vault (WebDAV: /mnt/yandex-disk/obsidian/mozg/ — Yandex Disk davfs2)
↓ (sync_obsidian_to_wiki.py — every 30 min via Hermes cron, job_id: 2346a68b601d)
wiki-raw/ directory (80 .md files, 73 ingested + 7 empty skipped)
↓ (wiki_continuous_ingest.py — ARQ worker)
Parse → Chunk → Embed (dense: bge-m3 1024d, sparse: BM25)
↓
Qdrant collection "knowledge_base" (dense 1024d COSINE + sparse "sparse")
↓
context_enhancer.py (CLI search tool — 4-level fallback cascade)
↓
icarus/hooks.py (auto-inject into Hermes responses)
When This Skill Activates
When the user:
- Asks about Memory OS setup, debugging, or configuration
- Reports that sync_obsidian_to_wiki.py or the ingress pipeline isn't running
- Asks about alternative Obsidian vault mounts (WebDAV vs s3fs) for the pipeline
- Reports Qdrant returning no results or empty searches
- Asks to reingest, reset, or rebuild the search index
- Reports embedding mismatches or search quality issues
- Deploys or updates the ingest pipeline
- Asks about the wiki ingestion cron job
Obsidian Vault Source Path
The pipeline picks up notes from an Obsidian vault. The canonical source is:
| Mount | Path | Status |
|---|---|---|
| WebDAV (Yandex Disk) | /mnt/yandex-disk/obsidian/mozg/ |
✅ Active — davfs2, file_mode=600 |
| s3fs (Garage) | /opt/hermes/obsidian-vault/ |
❌ Defunct — bucket empty at /etc/fstab entry |
READ-ONLY guarantee: sync_obsidian_to_wiki.py copies FROM vault TO wiki target. It NEVER modifies or deletes source files. The vault is mounted via davfs2 which is inherently read-only by filesystem mode, and the code only calls copy2() and unlink() on the target path.
s3fs/obsidian-vault defunct — do NOT use /opt/hermes/obsidian-vault/. The old s3fs mount is empty. The canonical vault is WebDAV at /mnt/yandex-disk/obsidian/mozg/. Any cron job or script referencing /opt/hermes/obsidian-vault/ is stale — redirect to the WebDAV path. The stale mount has 8 root-owned leftover files; they are NOT the real vault.
WebDAV switch-over (2026-07-16): The s3fs mount at s3.nixg.ru was empty. Changed OBSIDIAN_VAULT from /opt/hermes/obsidian-vault to /mnt/yandex-disk/obsidian/mozg/. First sync: 74 files in 44s, 0 errors.
Hermes Cron Setup
The sync is a managed Hermes cron job:
| Property | Value |
|---|---|
| Job ID | 2346a68b601d |
| Schedule | */30 * * * * |
| Runner | Wrapper script at ~/.hermes/scripts/obsidian-sync.sh |
| no_agent | true — pure shell execution, no LLM |
| Script | cd /opt/hermes/memory-os && .venv/bin/python3 scripts/sync_obsidian_to_wiki.py |
The wrapper exists because Hermes cron's script field requires a path under ~/.hermes/scripts/. It calls .venv/bin/python3 scripts/sync_obsidian_to_wiki.py.
sync → ingest chain: sync_obsidian_to_wiki.py internally calls wiki_continuous_ingest.py when files change. Must use .venv/bin/python3 (system python3 lacks arq). Fixed 2026-07-16.
no_agent optimization: The cron job uses no_agent=true — pure script execution, no LLM reasoning needed. Avoids wasting tokens on every 30-min tick. Verified 2026-07-19.
WebDAV vs s3fs Comparison
| Aspect | WebDAV (davfs2) | s3fs (Garage) |
|---|---|---|
| Works | ✅ Active | ❌ Empty mount |
| POSIX | ⚠️ file_mode=600 |
✅ Normal perms |
| Safety | ✅ Read-only by fuse | ✅ Script guard |
| Config | webdav.yandex.ru at /mnt/yandex-disk |
s3.nixg.ru at /opt/hermes/obsidian-vault |
| Component | Path | Role |
|---|---|---|
| State tracker | /opt/hermes/email/state/wiki_ingest_state.json |
Tracks which files were queued |
| Failures log | /opt/hermes/email/state/wiki_ingest_failures.json |
Records ingest errors |
| Sync state | /opt/hermes/email/state/obsidian_sync_state.json |
Tracks Obsidian vault sync state (mtime/size per file) |
| Sync script | scripts/sync_obsidian_to_wiki.py |
Syncs Obsidian vault → wiki-raw/ |
| Ingest worker | scripts/wiki_continuous_ingest.py |
ARQ worker: parse → chunk → embed → Qdrant |
| Bulk ingest | scripts/bulk_wiki_ingest_ollama.py |
One-shot re-index: reads all .md, chunks (headings→paragraphs→words), embeds via Ollama bge-m3, stores in Qdrant. Replaces manual pipeline reset. |
| Search CLI | scripts/context_enhancer.py |
4-level fallback search (hybrid → dense → lexical → SQLite) |
| Hermes hook | icarus/hooks.py |
Auto-injects Qdrant context into responses |
| Docker env | docker/.env |
Settings for Ollama, Qdrant, Redis |
Key Environment Variables
# Qdrant
QDRANT_URL=http://localhost:6333
QDRANT_COLLECTION=knowledge_base
# Embedding (dense — local Ollama bge-m3 1024d)
OLLAMA_EMBEDDING_URL=http://localhost:11434
OLLAMA_EMBEDDING_MODEL=bge-m3:latest
# Embedding dimension (must match collection)
EMBEDDING_DIMS=1024
# Embedding (sparse — BM25 via FastEmbed)
# Uses subprocess to ai-lab venv at _FASTEMBED_PYTHON path in context_enhancer.py
# Redis (for ARQ queue)
REDIS_HOST=127.0.0.1
REDIS_PORT=6379
REDIS_PASSWORD=<password>
Key Commands
# Reset everything (re-ingest from scratch)
cd /opt/hermes/memory-os
# 1. Clear state so all files are considered new
echo '{}' > /opt/hermes/email/state/wiki_ingest_state.json
# 2. Drop and recreate Qdrant collection (1024d COSINE + sparse)
python3 -c "
from qdrant_client import QdrantClient, models
c = QdrantClient('http://localhost:6333')
c.delete_collection('knowledge_base')
c.create_collection(
collection_name='knowledge_base',
vectors_config=models.VectorParams(size=1024, distance=models.Distance.COSINE),
sparse_vectors_config={'sparse': models.SparseVectorParams(
index=models.SparseIndexParams(on_disk=False, full_scan_threshold=10000)
)}
)
print('Created 1024d COSINE + sparse')
"
# 3. Bulk re-index all files (with chunking)
/opt/hermes/memory-os/.venv/bin/python3 scripts/bulk_wiki_ingest_ollama.py
# 4. Search
python3 scripts/context_enhancer.py "your query" --top-k 5 --format markdown
python3 scripts/context_enhancer.py "your query" --top-k 10 --threshold 0.35 --format compact
Structure of Qdrant Collection
The collection knowledge_base must have:
- dense vectors: 1024 dimensions, COSINE distance (bge-m3 via Ollama)
- sparse vectors: BM25 (named
sparse, NOTbm25)
models.VectorParams(size=1024, distance=models.Distance.COSINE)
models.SparseVectorParams(
index=models.SparseIndexParams(on_disk=False, full_scan_threshold=10000)
)
The sparse vector name MUST be "sparse" — this matches what the ingest worker sends. Creating it as "bm25" causes a 400 Bad Request error on every write.
Historical note: This was previously 768d (nomic-embed-text). Migrated to 1024d (bge-m3) on 2026-07-16 for multilingual support. See references/chunking-implementation.md for the current chunking approach.
Debugging Checklist (when search returns nothing or poor results)
1. Check Qdrant has points
python3 -c "
from qdrant_client import QdrantClient
c = QdrantClient('http://localhost:6333')
info = c.get_collection('knowledge_base')
print(f'Points: {info.points_count}')
print(f'Status: {info.status}')
"
If points_count = 0, the collection was never ingested or was dropped.
2. Check sparse vector name matches
python3 -c "
from qdrant_client import QdrantClient
c = QdrantClient('http://localhost:6333')
info = c.get_collection('knowledge_base')
print('Sparse vectors:', info.config.params.sparse_vectors)
"
If sparse vectors key is "bm25" but the worker sends "sparse", you get 400 errors. Recreate the collection with the correct name.
3. Check dense embedding model matches
The context_enhancer.py's embed_query() must use the SAME model that was used during ingest. If OpenRouter was used for search but local Ollama for ingest (or vice versa), dense vectors won't match and search degrades to lexical fallback.
Check what the search client uses vs what was used during ingest:
# Search client config
grep -n "OLLAMA_EMBEDDING_MODEL\|EMBEDDING_MODEL" scripts/context_enhancer.py
# Ingest config (docker/.env or wiki_continuous_ingest.py)
grep -n "embedding\|model" docker/.env
grep -rn "embed\|nomic" scripts/wiki_continuous_ingest.py
4. Check FastEmbed BM25 works (sparse engine)
python3 -c "
from fastembed.sparse import SparseTextEmbedding
model = SparseTextEmbedding(model_name='Qdrant/bm25')
sparse = list(model.embed(['test query']))[0]
print(f'indices: {len(sparse.indices)}, values: {len(sparse.values)}')
"
If this fails, the FastEmbed subprocess in context_enhancer.py needs the FASTEMBED_SITEPKGS env var.
5. Check worker logs
docker logs docker-worker-1 --tail 50
Look for 400 errors (bad sparse name), DNS failures (ollama unreachable), or timeouts.
6. Check Docker networking
After every docker compose up -d, the worker container loses secondary networks:
docker network connect ollama_default docker-worker-1
docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys'
7. Check ARQ queue isn't stale
redis-cli -a <password> keys 'arq:*' | head -20
redis-cli -a <password> llen 'arq:queue:health-check'
If there are stale jobs in the queue, clear them with:
redis-cli -a <password> del 'arq:queue:health-check'
8. Test dense embedding directly
curl -s http://localhost:11434/api/embeddings \
-d '{"model":"bge-m3:latest","prompt":"test query"}' | \
python3 -c "import sys, json; d=json.load(sys.stdin); print(f'Dims: {len(d[\"embedding\"])}')"
# Should print "Dims: 1024"
9. Check collection dimensions match .env embedding settings (400 Bad Request)
When you get a 400 Client Error: Bad Request for url: http://localhost:6333/collections/knowledge_base/points/query, the most likely cause is the collection was created with one embedding dimension (e.g. 768d for nomic-embed-text) but .env now points to a different model with different dimensions (e.g. 1024d for bge-m3).
Cross-check:
# What the collection expects:
python3 -c "
from qdrant_client import QdrantClient
c = QdrantClient('http://localhost:6333')
info = c.get_collection('knowledge_base')
print(f'Collection dims: {info.config.params.vectors.size}')
print(f'Distance: {info.config.params.vectors.distance}')
"
# What .env says:
grep -E 'EMBEDDING_DIMS|EMBEDDING_MODEL' /opt/hermes/memory-os/.env
If they don't match, you must either:
- (a) Recreate the collection with the correct dimension (destructive — lose all points). Use the current model's dimension.
- (b) Change
.envto match the collection (switch back to the old model). - (c) Create a second collection for the new dimension, keep the old one.
Current config (2026-07-16): bge-m3 1024d COSINE. See
references/chunking-implementation.mdfor the chunking approach.
See also
references/dimension-mismatch-s3fs-blockers.mdfor full reproduction transcript.
Pitfalls
-
All state files live under
/opt/hermes/email/state/, NOT~/.hermes/. Three scripts reference state files — sync_obsidian_to_wiki.py, wiki_continuous_ingest.py, dlq_manager.py. All use/opt/hermes/email/state/. If any script still usesPath.home() / ".hermes"oros.path.expanduser("~/.hermes")for state paths, patch it. The old paths at~/.hermes/obsidian_sync_state.json,~/.hermes/wiki_ingest_state.json, and~/.hermes/wiki_ingest_failures.jsonare stale. -
Missing wrapper script blocks cron execution silently. The Hermes cron job
obsidian-sync(job_id2346a68b601d) runs~/.hermes/scripts/obsidian-sync.sh. If this file is missing, the cron tick produces no output and no error — the job just does nothing. On 2026-07-19 the script was documented in the skill but never created on disk. Fix:mkdir -p ~/.hermes/scripts; write the wrapper;chmod +x. Then switch the job tono_agent=trueso pure script execution doesn't burn LLM tokens. -
**
~expansion differs between Python and the shell when running from/opt/hermes/memory-os/. -
~expansion differs between Python and the shell when running from/opt/hermes/memory-os/.context_enhancer.pyusesos.path.expanduser("~/.hermes/state.db")forLINEAGE_DB. When the script runs via.venv/bin/pythonfrom/opt/hermes/memory-os/,~resolves to/opt/hermes/— NOT/home/estorozhenko/. This meansLINEAGE_DBbecomes/opt/hermes/.hermes/state.db— a different Hermes state DB that doesn't have thelineagetable. Theregister_lineage()call then silently fails with[LINEAGE-WARNING] Failed to register lineage: no such table: lineage. Fix: create thelineagetable in/opt/hermes/.hermes/state.dbtoo, or setSTATE_DB_PATH=/home/estorozhenko/.hermes/state.dbin.envor process environment. Symptom: search results appear but[LINEAGE-WARNING]is printed to stderr. -
Fabric directory (
~/fabric/) does not exist by default. The Icarus plugin writes fabric entries toFABRIC_DIR = Path.home() / "fabric"(line 19 ofstate.py). This directory is never auto-created beforewrite_entry()is called (it callsmkdir(parents=True, exist_ok=True)inwrite_entry()itself, so writes succeed, butread_recent()andread_cross_agent()return empty ifFABRIC_DIR.exists()is False). Fix:mkdir -p ~/fabric/to ensure the directory exists before any Icarus session starts. Without this, the first session after enabling Icarus will have no[fabric]context injected. -
HERMES_AGENT_NAMEis not set. IfHERMES_AGENT_NAMEis missing from.envor environment,state.AGENT_NAMEis empty string, and the logger warns:"icarus: HERMES_AGENT_NAME not set — fabric entries will use agent=\"agent\"". This means all fabric entries are tagged withagent: agentinstead of a meaningful name. Fix: addHERMES_AGENT_NAME=hermes(or another name) to/opt/hermes/.hermes/config.yamlor/opt/hermes/.hermes/.env. The name is used in fabric filenames (agent-entry_type-slug-id.md) and in multi-agent deployments. -
Fabric entries are NOT injected until mid-session — they only appear starting from turn 2+.
pre_llm_callinhooks.pycallsstate.recall()which reads fabric entries. On the first turn of a session,is_first_turn=Trueadditionally triggers_search_facts()for durable facts. Fabric entries from a previous session are injected as[fabric]blocks. If you started a new session and see no[fabric]block, the directory may not exist, or no entries were written by the previous session (becauseon_session_endwrites them, and if the session was closed abnormally, the hook never fired). -
memory_store.db(durable facts) is empty by default. Thefactstable in~/.hermes/memory_store.dbstarts with 0 rows. It's populated only by explicitmemorytool calls. A_search_facts()call always returns empty until the first fact is saved. This is normal — no fix needed. -
Session history FTS5 (
_search_sessions) needs a session_id exclusion to avoid self-referencing. The_search_sessions()function inhooks.py(line 331) filters out the current session bysession_idto avoid injecting the current conversation's own messages. If this filter breaks, the agent will see its own replies from the same session and recursively inject them. The filter usesWHERE session_id != ?in the SQL. -
Session history injection produces
[sessions]blocks, NOT[fabric]or[qdrant]. Thepre_llm_callhook injects four separate blocks:[fabric](from fabric entries),[qdrant](from Qdrant vault search),[sessions](from FTS5 session history), and[facts](from durable facts). Each has its own dedup set (_injected_fabric,_injected_qdrant,_injected_sessions,_injected_facts). If you see a[sessions]block, it came from FTS5, not Qdrant. -
CRITICAL: icarus/hooks.py threshold must be 0.30, not 0.55.
_search_qdrant()passesthreshold=0.55tosearch_with_fallback(). RRF fusion (hybrid dense+sparse) returns scores 0.33-0.50 even for good matches — the old gate silently filters EVERYTHING. This is fundamentally different from dense-only search which returns 0.90+ for the same queries. A threshold of 0.35 still blocks some RRF results (0.33 scores). The safe floor is 0.30. Verified 2026-07-16: hybrid search with threshold=0.55 returned 0 results for every query tested; 0.30 returned 4 results. Without this fix,search_with_fallbackhitslevel="none", cascades into SQLite fallback ([CE-FALLBACK] SQLite search failed: no such table: lineage), pollutes logs, and Icarus injects nothing on every turn — the injection code runs, finds nothing, and returns silently. -
FASTEMBED_VENVmust be set in.envor the sparse BM25 subprocess uses system python without fastembed. Ifcontext_enhancer.pydoesn't findFASTEMBED_VENVin env, it falls back tosys.executable(system python). The.envmust containFASTEMBED_VENV=/opt/hermes/memory-os/.venv/bin/python3andFASTEMBED_SITEPKGS=/opt/hermes/memory-os/.venv/lib/python3.12/site-packages. Alsocontext_enhancer.pyneedsload_dotenv()to read.env(added 2026-07-16). -
python-dotenvmust be installed system-wide if any script usesload_dotenv(). Install withpip3 install --break-system-packages python-dotenv. -
Sparse vector name mismatch is the most common ingest failure. The collection must name the sparse config
"sparse", not"bm25"or anything else. The worker always sends to"sparse". -
Dense model mismatch kills semantic search. The search client must use the exact same embedding model as the ingest pipeline. Mixing local Ollama with OpenRouter embeddings (or different models) produces near-zero semantic similarity scores.
-
Docker compose up -d disconnects secondary networks. After any
docker compose up -d, the worker loses connection toollama_default. Always re-rundocker network connect. -
FASTEMBED_SITEPKGS must be in subprocess env. The sparse embedding runs in a subprocess; the env var must be explicitly passed via
env=_envwith_env.setdefault(). -
Subprocess python path must be the venv, not system python. If
_FASTEMBED_PYTHONpoints to/usr/bin/python3butfastembedis installed in/opt/ai-lab/.venv/bin/python3, the subprocess silently fails (empty stdout → JSON parse error). Always verify which python can import fastembed, and set the path accordingly. -
python-dotenvmay be missing in cron/system environment. Thewiki_continuous_ingest.pyscript importsfrom dotenv import load_dotenv. If the script runs via cron or systemd (not the ai-lab venv), it crashes withModuleNotFoundError. Install system-wide:pip install python-dotenv, or patch the script to avoid the dependency. -
Cron entry for ingest must use .venv python, not system python. The cron job
0 * * * * cd /opt/hermes/memory-os && python3 scripts/wiki_continuous_ingest.pyfails withModuleNotFoundErrorbecausedotenvandarqare only in the project's.venv/. Fix: use/opt/hermes/memory-os/.venv/bin/python3instead ofpython3. Verified 2026-07-16: after the fix, the script runs cleanly (⏭️ Nada novo. 4 arquivos rastreados, 4 inalterados.). -
icarus/hooks.py is dead code until registered as a Hermes user plugin. The plugin lives at
/opt/hermes/.hermes/plugins/icarus/(NOT in the memory-os project dir). Register withhermes plugins enable icarus— NOT by editing config.yaml manually. The plugin has its ownplugin.yaml(v0.3.0) with 16 tools and 4 hooks (on_session_start, pre_llm_call, post_llm_call, on_session_end). Withouthermes plugins enable, the file is never executed. -
icarus hooks.py needs a PYTHONPATH fix before Qdrant search works. The
_search_qdrant()function at line 285 doesfrom scripts.context_enhancer import ...but/opt/hermes/memory-osis not onsys.path. Fix: addimport sysat the top of hooks.py, then insertsys.path.insert(0, '/opt/hermes/memory-os')inside_search_qdrant()before the import. Without this,_search_qdrant()always raisesModuleNotFoundErrorand returns an empty list silently (fail-open). -
bge-m3 scores are higher than nomic-embed-text. With nomic-embed-text (768d), hybrid scores were 0.33-0.50. With bge-m3 (1024d), scores are 0.48-0.65 for Russian queries. The Icarus threshold was raised from 0.30 to 0.40. CLI default stays at 0.35 (fine for dense-only).
-
bge-m3 context limit is ~6000 characters. The Ollama /api/embeddings endpoint returns 500 for texts longer than ~6000 chars. The
bulk_wiki_ingest_ollama.pyscript chunks files viachunk_text()— recursive split by ##/###/#### headings → paragraphs → words (last resort).MAX_CHUNK_SIZE=5000,CHUNK_OVERLAP=300. Each chunk becomes a separate Qdrant point withchunk_indexandchunk_totalin payload. -
Switching embedding models requires a full re-index. When changing from nomic-embed-text (768d) to bge-m3 (1024d): delete old Qdrant collection, recreate with correct dims, re-index all files. Both .env and context_enhancer.py must be updated.
-
Score threshold has two separate regimes. Dense-only search (via search_knowledge_base) returns COSINE scores 0.90+ for good matches — threshold 0.35 works fine. Hybrid search (via search_with_fallback with sparse vector, used by Icarus) uses RRF fusion which normalises to 0.33-0.50. The Icarus threshold in hooks.py must be set to 0.30, not 0.35 or 0.55. If you see level="none" and SQLite fallback errors, the threshold is too high for RRF.
-
Dimension mismatch between Qdrant collection and .env embedding config is invisible until search time. If
.envwas changed to a different embedding model with different dimensions (e.g. nomic-embed-text 768d -> bge-m3 1024d) but the Qdrant collection still has the old dimensions,context_enhancer.pyreturns400 Bad Requestat query time. Cross-check:collection.config.params.vectors.sizevsEMBEDDING_DIMSin.env. -
sync_obsidian_to_wiki.py has a safety guard against empty vaults. If the WebDAV vault mount is empty but the state file has entries, the script skips deletion (
VAULT ПУСТ... пропускаю удаление). This prevents data loss when the mount is temporarily disconnected. The canonical vault is now WebDAV at/mnt/yandex-disk/obsidian/mozg/— s3fs is defunct. -
sync → ingest chain needs .venv python. Inside
sync_obsidian_to_wiki.py, the post-sync call towiki_continuous_ingest.pyusesos.system(f"python3 ..."). System python lacksarq. Must use.venv/bin/python3. Fixed 2026-07-16. -
State file tracks queue time, not ingest completion. After resetting the state file and enqueuing, you must wait for the ARQ worker to process before Qdrant has points. Check
wiki_ingest_failures.jsonfor processing errors. -
Failures.json is append-only — it accumulates errors across runs. To get a clean view, read it with jq or truncate it when resetting.
-
search-api
fileandsourcefields depend on which ingest path stored the data. The Docker worker storesfile_path(/wiki/homelab/foo.md) andsource(wiki-homelab). The bulk ingest script storespath(full filesystem path) andfilename(basename), but nosource. The search-api was looking forpayload.get("file")which neither path stores — always returned null. Fixed by cascading throughfile_path→path→filename. Forsource, if missing or "unknown", the fix extracts thewiki-directory component fromfile_path. Seereferences/search-api-payload-fields.md. After changing search-api code, rebuild and restart:docker compose build search-api && docker compose up -d search-api.