Files
memory-os/SKILL.md
T
2026-09-06 13:51:22 +00:00

24 KiB

name, description, version, author, metadata
name description version author metadata
memory-os Qdrant-based RAG pipeline for wiki/note vault ingestion — sync, parse, chunk, embed (dense: bge-m3 + sparse: BM25), store in Qdrant (1024d COSINE), and auto-inject retrieved context into Hermes responses. 1.8.0 Hermes Agent
hermes
tags category related_skills sources
rag
qdrant
vector-search
wiki-ingest
memory
embedding
bm25
ollama
bge-m3
multilingual
research
llm-wiki
obsidian
llama-cpp
references/chunking-implementation.md
references/dimension-mismatch-s3fs-blockers.md
references/ingest-debug-session.md
references/verify-threshold-bug.md
references/webdav-migration.md
references/search-api-payload-fields.md

Memory OS — Wiki RAG Pipeline

A production RAG pipeline that continuously syncs an Obsidian vault (or any markdown wiki) into Qdrant for semantic + BM25 retrieval, and auto-injects relevant context into every Hermes response.

Architecture:

Obsidian vault (WebDAV: /mnt/yandex-disk/obsidian/mozg/ — Yandex Disk davfs2)
    ↓ (sync_obsidian_to_wiki.py — every 30 min via Hermes cron, job_id: 2346a68b601d)
wiki-raw/ directory (80 .md files, 73 ingested + 7 empty skipped)
    ↓ (wiki_continuous_ingest.py — ARQ worker)
Parse → Chunk → Embed (dense: bge-m3 1024d, sparse: BM25)
    ↓
Qdrant collection "knowledge_base" (dense 1024d COSINE + sparse "sparse")
    ↓
context_enhancer.py (CLI search tool — 4-level fallback cascade)
    ↓
icarus/hooks.py (auto-inject into Hermes responses)

When This Skill Activates

When the user:

  • Asks about Memory OS setup, debugging, or configuration
  • Reports that sync_obsidian_to_wiki.py or the ingress pipeline isn't running
  • Asks about alternative Obsidian vault mounts (WebDAV vs s3fs) for the pipeline
  • Reports Qdrant returning no results or empty searches
  • Asks to reingest, reset, or rebuild the search index
  • Reports embedding mismatches or search quality issues
  • Deploys or updates the ingest pipeline
  • Asks about the wiki ingestion cron job

Obsidian Vault Source Path

The pipeline picks up notes from an Obsidian vault. The canonical source is:

Mount Path Status
WebDAV (Yandex Disk) /mnt/yandex-disk/obsidian/mozg/ ✅ Active — davfs2, file_mode=600
s3fs (Garage) /opt/hermes/obsidian-vault/ ❌ Defunct — bucket empty at /etc/fstab entry

READ-ONLY guarantee: sync_obsidian_to_wiki.py copies FROM vault TO wiki target. It NEVER modifies or deletes source files. The vault is mounted via davfs2 which is inherently read-only by filesystem mode, and the code only calls copy2() and unlink() on the target path.

s3fs/obsidian-vault defunct — do NOT use /opt/hermes/obsidian-vault/. The old s3fs mount is empty. The canonical vault is WebDAV at /mnt/yandex-disk/obsidian/mozg/. Any cron job or script referencing /opt/hermes/obsidian-vault/ is stale — redirect to the WebDAV path. The stale mount has 8 root-owned leftover files; they are NOT the real vault.

WebDAV switch-over (2026-07-16): The s3fs mount at s3.nixg.ru was empty. Changed OBSIDIAN_VAULT from /opt/hermes/obsidian-vault to /mnt/yandex-disk/obsidian/mozg/. First sync: 74 files in 44s, 0 errors.

Hermes Cron Setup

The sync is a managed Hermes cron job:

Property Value
Job ID 2346a68b601d
Schedule */30 * * * *
Runner Wrapper script at ~/.hermes/scripts/obsidian-sync.sh
no_agent true — pure shell execution, no LLM
Script cd /opt/hermes/memory-os && .venv/bin/python3 scripts/sync_obsidian_to_wiki.py

The wrapper exists because Hermes cron's script field requires a path under ~/.hermes/scripts/. It calls .venv/bin/python3 scripts/sync_obsidian_to_wiki.py.

sync → ingest chain: sync_obsidian_to_wiki.py internally calls wiki_continuous_ingest.py when files change. Must use .venv/bin/python3 (system python3 lacks arq). Fixed 2026-07-16.

no_agent optimization: The cron job uses no_agent=true — pure script execution, no LLM reasoning needed. Avoids wasting tokens on every 30-min tick. Verified 2026-07-19.

WebDAV vs s3fs Comparison

Aspect WebDAV (davfs2) s3fs (Garage)
Works ✅ Active ❌ Empty mount
POSIX ⚠️ file_mode=600 ✅ Normal perms
Safety ✅ Read-only by fuse ✅ Script guard
Config webdav.yandex.ru at /mnt/yandex-disk s3.nixg.ru at /opt/hermes/obsidian-vault
Component Path Role
State tracker /opt/hermes/email/state/wiki_ingest_state.json Tracks which files were queued
Failures log /opt/hermes/email/state/wiki_ingest_failures.json Records ingest errors
Sync state /opt/hermes/email/state/obsidian_sync_state.json Tracks Obsidian vault sync state (mtime/size per file)
Sync script scripts/sync_obsidian_to_wiki.py Syncs Obsidian vault → wiki-raw/
Ingest worker scripts/wiki_continuous_ingest.py ARQ worker: parse → chunk → embed → Qdrant
Bulk ingest scripts/bulk_wiki_ingest_ollama.py One-shot re-index: reads all .md, chunks (headings→paragraphs→words), embeds via Ollama bge-m3, stores in Qdrant. Replaces manual pipeline reset.
Search CLI scripts/context_enhancer.py 4-level fallback search (hybrid → dense → lexical → SQLite)
Hermes hook icarus/hooks.py Auto-injects Qdrant context into responses
Docker env docker/.env Settings for Ollama, Qdrant, Redis

Key Environment Variables

# Qdrant
QDRANT_URL=http://localhost:6333
QDRANT_COLLECTION=knowledge_base

# Embedding (dense — local Ollama bge-m3 1024d)
OLLAMA_EMBEDDING_URL=http://localhost:11434
OLLAMA_EMBEDDING_MODEL=bge-m3:latest

# Embedding dimension (must match collection)
EMBEDDING_DIMS=1024

# Embedding (sparse — BM25 via FastEmbed)
# Uses subprocess to ai-lab venv at _FASTEMBED_PYTHON path in context_enhancer.py

# Redis (for ARQ queue)
REDIS_HOST=127.0.0.1
REDIS_PORT=6379
REDIS_PASSWORD=<password>

Key Commands

# Reset everything (re-ingest from scratch)
cd /opt/hermes/memory-os

# 1. Clear state so all files are considered new
echo '{}' > /opt/hermes/email/state/wiki_ingest_state.json

# 2. Drop and recreate Qdrant collection (1024d COSINE + sparse)
python3 -c "
from qdrant_client import QdrantClient, models
c = QdrantClient('http://localhost:6333')
c.delete_collection('knowledge_base')
c.create_collection(
    collection_name='knowledge_base',
    vectors_config=models.VectorParams(size=1024, distance=models.Distance.COSINE),
    sparse_vectors_config={'sparse': models.SparseVectorParams(
        index=models.SparseIndexParams(on_disk=False, full_scan_threshold=10000)
    )}
)
print('Created 1024d COSINE + sparse')
"

# 3. Bulk re-index all files (with chunking)
/opt/hermes/memory-os/.venv/bin/python3 scripts/bulk_wiki_ingest_ollama.py

# 4. Search
python3 scripts/context_enhancer.py "your query" --top-k 5 --format markdown
python3 scripts/context_enhancer.py "your query" --top-k 10 --threshold 0.35 --format compact

Structure of Qdrant Collection

The collection knowledge_base must have:

  • dense vectors: 1024 dimensions, COSINE distance (bge-m3 via Ollama)
  • sparse vectors: BM25 (named sparse, NOT bm25)
models.VectorParams(size=1024, distance=models.Distance.COSINE)
models.SparseVectorParams(
    index=models.SparseIndexParams(on_disk=False, full_scan_threshold=10000)
)

The sparse vector name MUST be "sparse" — this matches what the ingest worker sends. Creating it as "bm25" causes a 400 Bad Request error on every write.

Historical note: This was previously 768d (nomic-embed-text). Migrated to 1024d (bge-m3) on 2026-07-16 for multilingual support. See references/chunking-implementation.md for the current chunking approach.

Debugging Checklist (when search returns nothing or poor results)

1. Check Qdrant has points

python3 -c "
from qdrant_client import QdrantClient
c = QdrantClient('http://localhost:6333')
info = c.get_collection('knowledge_base')
print(f'Points: {info.points_count}')
print(f'Status: {info.status}')
"

If points_count = 0, the collection was never ingested or was dropped.

2. Check sparse vector name matches

python3 -c "
from qdrant_client import QdrantClient
c = QdrantClient('http://localhost:6333')
info = c.get_collection('knowledge_base')
print('Sparse vectors:', info.config.params.sparse_vectors)
"

If sparse vectors key is "bm25" but the worker sends "sparse", you get 400 errors. Recreate the collection with the correct name.

3. Check dense embedding model matches

The context_enhancer.py's embed_query() must use the SAME model that was used during ingest. If OpenRouter was used for search but local Ollama for ingest (or vice versa), dense vectors won't match and search degrades to lexical fallback.

Check what the search client uses vs what was used during ingest:

# Search client config
grep -n "OLLAMA_EMBEDDING_MODEL\|EMBEDDING_MODEL" scripts/context_enhancer.py

# Ingest config (docker/.env or wiki_continuous_ingest.py)
grep -n "embedding\|model" docker/.env
grep -rn "embed\|nomic" scripts/wiki_continuous_ingest.py

4. Check FastEmbed BM25 works (sparse engine)

python3 -c "
from fastembed.sparse import SparseTextEmbedding
model = SparseTextEmbedding(model_name='Qdrant/bm25')
sparse = list(model.embed(['test query']))[0]
print(f'indices: {len(sparse.indices)}, values: {len(sparse.values)}')
"

If this fails, the FastEmbed subprocess in context_enhancer.py needs the FASTEMBED_SITEPKGS env var.

5. Check worker logs

docker logs docker-worker-1 --tail 50

Look for 400 errors (bad sparse name), DNS failures (ollama unreachable), or timeouts.

6. Check Docker networking

After every docker compose up -d, the worker container loses secondary networks:

docker network connect ollama_default docker-worker-1
docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys'

7. Check ARQ queue isn't stale

redis-cli -a <password> keys 'arq:*' | head -20
redis-cli -a <password> llen 'arq:queue:health-check'

If there are stale jobs in the queue, clear them with:

redis-cli -a <password> del 'arq:queue:health-check'

8. Test dense embedding directly

curl -s http://localhost:11434/api/embeddings \
  -d '{"model":"bge-m3:latest","prompt":"test query"}' | \
  python3 -c "import sys, json; d=json.load(sys.stdin); print(f'Dims: {len(d[\"embedding\"])}')"
# Should print "Dims: 1024"

9. Check collection dimensions match .env embedding settings (400 Bad Request)

When you get a 400 Client Error: Bad Request for url: http://localhost:6333/collections/knowledge_base/points/query, the most likely cause is the collection was created with one embedding dimension (e.g. 768d for nomic-embed-text) but .env now points to a different model with different dimensions (e.g. 1024d for bge-m3).

Cross-check:

# What the collection expects:
python3 -c "
from qdrant_client import QdrantClient
c = QdrantClient('http://localhost:6333')
info = c.get_collection('knowledge_base')
print(f'Collection dims: {info.config.params.vectors.size}')
print(f'Distance: {info.config.params.vectors.distance}')
"

# What .env says:
grep -E 'EMBEDDING_DIMS|EMBEDDING_MODEL' /opt/hermes/memory-os/.env

If they don't match, you must either:

  • (a) Recreate the collection with the correct dimension (destructive — lose all points). Use the current model's dimension.
  • (b) Change .env to match the collection (switch back to the old model).
  • (c) Create a second collection for the new dimension, keep the old one.

Current config (2026-07-16): bge-m3 1024d COSINE. See references/chunking-implementation.md for the chunking approach.

See also references/dimension-mismatch-s3fs-blockers.md for full reproduction transcript.

Pitfalls

  • All state files live under /opt/hermes/email/state/, NOT ~/.hermes/. Three scripts reference state files — sync_obsidian_to_wiki.py, wiki_continuous_ingest.py, dlq_manager.py. All use /opt/hermes/email/state/. If any script still uses Path.home() / ".hermes" or os.path.expanduser("~/.hermes") for state paths, patch it. The old paths at ~/.hermes/obsidian_sync_state.json, ~/.hermes/wiki_ingest_state.json, and ~/.hermes/wiki_ingest_failures.json are stale.

  • Missing wrapper script blocks cron execution silently. The Hermes cron job obsidian-sync (job_id 2346a68b601d) runs ~/.hermes/scripts/obsidian-sync.sh. If this file is missing, the cron tick produces no output and no error — the job just does nothing. On 2026-07-19 the script was documented in the skill but never created on disk. Fix: mkdir -p ~/.hermes/scripts; write the wrapper; chmod +x. Then switch the job to no_agent=true so pure script execution doesn't burn LLM tokens.

  • **~ expansion differs between Python and the shell when running from /opt/hermes/memory-os/.

  • ~ expansion differs between Python and the shell when running from /opt/hermes/memory-os/. context_enhancer.py uses os.path.expanduser("~/.hermes/state.db") for LINEAGE_DB. When the script runs via .venv/bin/python from /opt/hermes/memory-os/, ~ resolves to /opt/hermes/ — NOT /home/estorozhenko/. This means LINEAGE_DB becomes /opt/hermes/.hermes/state.db — a different Hermes state DB that doesn't have the lineage table. The register_lineage() call then silently fails with [LINEAGE-WARNING] Failed to register lineage: no such table: lineage. Fix: create the lineage table in /opt/hermes/.hermes/state.db too, or set STATE_DB_PATH=/home/estorozhenko/.hermes/state.db in .env or process environment. Symptom: search results appear but [LINEAGE-WARNING] is printed to stderr.

  • Fabric directory (~/fabric/) does not exist by default. The Icarus plugin writes fabric entries to FABRIC_DIR = Path.home() / "fabric" (line 19 of state.py). This directory is never auto-created before write_entry() is called (it calls mkdir(parents=True, exist_ok=True) in write_entry() itself, so writes succeed, but read_recent() and read_cross_agent() return empty if FABRIC_DIR.exists() is False). Fix: mkdir -p ~/fabric/ to ensure the directory exists before any Icarus session starts. Without this, the first session after enabling Icarus will have no [fabric] context injected.

  • HERMES_AGENT_NAME is not set. If HERMES_AGENT_NAME is missing from .env or environment, state.AGENT_NAME is empty string, and the logger warns: "icarus: HERMES_AGENT_NAME not set — fabric entries will use agent=\"agent\"". This means all fabric entries are tagged with agent: agent instead of a meaningful name. Fix: add HERMES_AGENT_NAME=hermes (or another name) to /opt/hermes/.hermes/config.yaml or /opt/hermes/.hermes/.env. The name is used in fabric filenames (agent-entry_type-slug-id.md) and in multi-agent deployments.

  • Fabric entries are NOT injected until mid-session — they only appear starting from turn 2+. pre_llm_call in hooks.py calls state.recall() which reads fabric entries. On the first turn of a session, is_first_turn=True additionally triggers _search_facts() for durable facts. Fabric entries from a previous session are injected as [fabric] blocks. If you started a new session and see no [fabric] block, the directory may not exist, or no entries were written by the previous session (because on_session_end writes them, and if the session was closed abnormally, the hook never fired).

  • memory_store.db (durable facts) is empty by default. The facts table in ~/.hermes/memory_store.db starts with 0 rows. It's populated only by explicit memory tool calls. A _search_facts() call always returns empty until the first fact is saved. This is normal — no fix needed.

  • Session history FTS5 (_search_sessions) needs a session_id exclusion to avoid self-referencing. The _search_sessions() function in hooks.py (line 331) filters out the current session by session_id to avoid injecting the current conversation's own messages. If this filter breaks, the agent will see its own replies from the same session and recursively inject them. The filter uses WHERE session_id != ? in the SQL.

  • Session history injection produces [sessions] blocks, NOT [fabric] or [qdrant]. The pre_llm_call hook injects four separate blocks: [fabric] (from fabric entries), [qdrant] (from Qdrant vault search), [sessions] (from FTS5 session history), and [facts] (from durable facts). Each has its own dedup set (_injected_fabric, _injected_qdrant, _injected_sessions, _injected_facts). If you see a [sessions] block, it came from FTS5, not Qdrant.

  • CRITICAL: icarus/hooks.py threshold must be 0.30, not 0.55. _search_qdrant() passes threshold=0.55 to search_with_fallback(). RRF fusion (hybrid dense+sparse) returns scores 0.33-0.50 even for good matches — the old gate silently filters EVERYTHING. This is fundamentally different from dense-only search which returns 0.90+ for the same queries. A threshold of 0.35 still blocks some RRF results (0.33 scores). The safe floor is 0.30. Verified 2026-07-16: hybrid search with threshold=0.55 returned 0 results for every query tested; 0.30 returned 4 results. Without this fix, search_with_fallback hits level="none", cascades into SQLite fallback ([CE-FALLBACK] SQLite search failed: no such table: lineage), pollutes logs, and Icarus injects nothing on every turn — the injection code runs, finds nothing, and returns silently.

  • FASTEMBED_VENV must be set in .env or the sparse BM25 subprocess uses system python without fastembed. If context_enhancer.py doesn't find FASTEMBED_VENV in env, it falls back to sys.executable (system python). The .env must contain FASTEMBED_VENV=/opt/hermes/memory-os/.venv/bin/python3 and FASTEMBED_SITEPKGS=/opt/hermes/memory-os/.venv/lib/python3.12/site-packages. Also context_enhancer.py needs load_dotenv() to read .env (added 2026-07-16).

  • python-dotenv must be installed system-wide if any script uses load_dotenv(). Install with pip3 install --break-system-packages python-dotenv.

  • Sparse vector name mismatch is the most common ingest failure. The collection must name the sparse config "sparse", not "bm25" or anything else. The worker always sends to "sparse".

  • Dense model mismatch kills semantic search. The search client must use the exact same embedding model as the ingest pipeline. Mixing local Ollama with OpenRouter embeddings (or different models) produces near-zero semantic similarity scores.

  • Docker compose up -d disconnects secondary networks. After any docker compose up -d, the worker loses connection to ollama_default. Always re-run docker network connect.

  • FASTEMBED_SITEPKGS must be in subprocess env. The sparse embedding runs in a subprocess; the env var must be explicitly passed via env=_env with _env.setdefault().

  • Subprocess python path must be the venv, not system python. If _FASTEMBED_PYTHON points to /usr/bin/python3 but fastembed is installed in /opt/ai-lab/.venv/bin/python3, the subprocess silently fails (empty stdout → JSON parse error). Always verify which python can import fastembed, and set the path accordingly.

  • python-dotenv may be missing in cron/system environment. The wiki_continuous_ingest.py script imports from dotenv import load_dotenv. If the script runs via cron or systemd (not the ai-lab venv), it crashes with ModuleNotFoundError. Install system-wide: pip install python-dotenv, or patch the script to avoid the dependency.

  • Cron entry for ingest must use .venv python, not system python. The cron job 0 * * * * cd /opt/hermes/memory-os && python3 scripts/wiki_continuous_ingest.py fails with ModuleNotFoundError because dotenv and arq are only in the project's .venv/. Fix: use /opt/hermes/memory-os/.venv/bin/python3 instead of python3. Verified 2026-07-16: after the fix, the script runs cleanly (⏭️ Nada novo. 4 arquivos rastreados, 4 inalterados.).

  • icarus/hooks.py is dead code until registered as a Hermes user plugin. The plugin lives at /opt/hermes/.hermes/plugins/icarus/ (NOT in the memory-os project dir). Register with hermes plugins enable icarus — NOT by editing config.yaml manually. The plugin has its own plugin.yaml (v0.3.0) with 16 tools and 4 hooks (on_session_start, pre_llm_call, post_llm_call, on_session_end). Without hermes plugins enable, the file is never executed.

  • icarus hooks.py needs a PYTHONPATH fix before Qdrant search works. The _search_qdrant() function at line 285 does from scripts.context_enhancer import ... but /opt/hermes/memory-os is not on sys.path. Fix: add import sys at the top of hooks.py, then insert sys.path.insert(0, '/opt/hermes/memory-os') inside _search_qdrant() before the import. Without this, _search_qdrant() always raises ModuleNotFoundError and returns an empty list silently (fail-open).

  • bge-m3 scores are higher than nomic-embed-text. With nomic-embed-text (768d), hybrid scores were 0.33-0.50. With bge-m3 (1024d), scores are 0.48-0.65 for Russian queries. The Icarus threshold was raised from 0.30 to 0.40. CLI default stays at 0.35 (fine for dense-only).

  • bge-m3 context limit is ~6000 characters. The Ollama /api/embeddings endpoint returns 500 for texts longer than ~6000 chars. The bulk_wiki_ingest_ollama.py script chunks files via chunk_text() — recursive split by ##/###/#### headings → paragraphs → words (last resort). MAX_CHUNK_SIZE=5000, CHUNK_OVERLAP=300. Each chunk becomes a separate Qdrant point with chunk_index and chunk_total in payload.

  • Switching embedding models requires a full re-index. When changing from nomic-embed-text (768d) to bge-m3 (1024d): delete old Qdrant collection, recreate with correct dims, re-index all files. Both .env and context_enhancer.py must be updated.

  • Score threshold has two separate regimes. Dense-only search (via search_knowledge_base) returns COSINE scores 0.90+ for good matches — threshold 0.35 works fine. Hybrid search (via search_with_fallback with sparse vector, used by Icarus) uses RRF fusion which normalises to 0.33-0.50. The Icarus threshold in hooks.py must be set to 0.30, not 0.35 or 0.55. If you see level="none" and SQLite fallback errors, the threshold is too high for RRF.

  • Dimension mismatch between Qdrant collection and .env embedding config is invisible until search time. If .env was changed to a different embedding model with different dimensions (e.g. nomic-embed-text 768d -> bge-m3 1024d) but the Qdrant collection still has the old dimensions, context_enhancer.py returns 400 Bad Request at query time. Cross-check: collection.config.params.vectors.size vs EMBEDDING_DIMS in .env.

  • sync_obsidian_to_wiki.py has a safety guard against empty vaults. If the WebDAV vault mount is empty but the state file has entries, the script skips deletion (VAULT ПУСТ... пропускаю удаление). This prevents data loss when the mount is temporarily disconnected. The canonical vault is now WebDAV at /mnt/yandex-disk/obsidian/mozg/ — s3fs is defunct.

  • sync → ingest chain needs .venv python. Inside sync_obsidian_to_wiki.py, the post-sync call to wiki_continuous_ingest.py uses os.system(f"python3 ..."). System python lacks arq. Must use .venv/bin/python3. Fixed 2026-07-16.

  • State file tracks queue time, not ingest completion. After resetting the state file and enqueuing, you must wait for the ARQ worker to process before Qdrant has points. Check wiki_ingest_failures.json for processing errors.

  • Failures.json is append-only — it accumulates errors across runs. To get a clean view, read it with jq or truncate it when resetting.

  • search-api file and source fields depend on which ingest path stored the data. The Docker worker stores file_path (/wiki/homelab/foo.md) and source (wiki-homelab). The bulk ingest script stores path (full filesystem path) and filename (basename), but no source. The search-api was looking for payload.get("file") which neither path stores — always returned null. Fixed by cascading through file_path → path → filename. For source, if missing or "unknown", the fix extracts the wiki- directory component from file_path. See references/search-api-payload-fields.md. After changing search-api code, rebuild and restart: docker compose build search-api && docker compose up -d search-api.