mirror of
https://gitverse.ru/kpa39l/memory-os.git
synced 2026-09-29 09:35:05 +00:00
Initial commit: Hermes skill memory-os
This commit is contained in:
@@ -0,0 +1,319 @@
|
|||||||
|
---
|
||||||
|
name: memory-os
|
||||||
|
description: "Qdrant-based RAG pipeline for wiki/note vault ingestion — sync, parse, chunk, embed (dense: bge-m3 + sparse: BM25), store in Qdrant (1024d COSINE), and auto-inject retrieved context into Hermes responses."
|
||||||
|
version: 1.8.0
|
||||||
|
author: Hermes Agent
|
||||||
|
metadata:
|
||||||
|
hermes:
|
||||||
|
tags: [rag, qdrant, vector-search, wiki-ingest, memory, embedding, bm25, ollama, bge-m3, multilingual]
|
||||||
|
category: research
|
||||||
|
related_skills: [llm-wiki, obsidian, llama-cpp]
|
||||||
|
sources:
|
||||||
|
- references/chunking-implementation.md
|
||||||
|
- references/dimension-mismatch-s3fs-blockers.md
|
||||||
|
- references/ingest-debug-session.md
|
||||||
|
- references/verify-threshold-bug.md
|
||||||
|
- references/webdav-migration.md
|
||||||
|
- references/search-api-payload-fields.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# Memory OS — Wiki RAG Pipeline
|
||||||
|
|
||||||
|
A production RAG pipeline that continuously syncs an Obsidian vault (or any markdown wiki) into Qdrant for semantic + BM25 retrieval, and auto-injects relevant context into every Hermes response.
|
||||||
|
|
||||||
|
**Architecture:**
|
||||||
|
|
||||||
|
```
|
||||||
|
Obsidian vault (WebDAV: /mnt/yandex-disk/obsidian/mozg/ — Yandex Disk davfs2)
|
||||||
|
↓ (sync_obsidian_to_wiki.py — every 30 min via Hermes cron, job_id: 2346a68b601d)
|
||||||
|
wiki-raw/ directory (80 .md files, 73 ingested + 7 empty skipped)
|
||||||
|
↓ (wiki_continuous_ingest.py — ARQ worker)
|
||||||
|
Parse → Chunk → Embed (dense: bge-m3 1024d, sparse: BM25)
|
||||||
|
↓
|
||||||
|
Qdrant collection "knowledge_base" (dense 1024d COSINE + sparse "sparse")
|
||||||
|
↓
|
||||||
|
context_enhancer.py (CLI search tool — 4-level fallback cascade)
|
||||||
|
↓
|
||||||
|
icarus/hooks.py (auto-inject into Hermes responses)
|
||||||
|
```
|
||||||
|
|
||||||
|
## When This Skill Activates
|
||||||
|
|
||||||
|
When the user:
|
||||||
|
- Asks about Memory OS setup, debugging, or configuration
|
||||||
|
- Reports that sync_obsidian_to_wiki.py or the ingress pipeline isn't running
|
||||||
|
- Asks about alternative Obsidian vault mounts (WebDAV vs s3fs) for the pipeline
|
||||||
|
- Reports Qdrant returning no results or empty searches
|
||||||
|
- Asks to reingest, reset, or rebuild the search index
|
||||||
|
- Reports embedding mismatches or search quality issues
|
||||||
|
- Deploys or updates the ingest pipeline
|
||||||
|
- Asks about the wiki ingestion cron job
|
||||||
|
|
||||||
|
## Obsidian Vault Source Path
|
||||||
|
|
||||||
|
The pipeline picks up notes from an Obsidian vault. The canonical source is:
|
||||||
|
|
||||||
|
| Mount | Path | Status |
|
||||||
|
|---|---|---|
|
||||||
|
| **WebDAV (Yandex Disk)** | `/mnt/yandex-disk/obsidian/mozg/` | ✅ Active — davfs2, file_mode=600 |
|
||||||
|
| s3fs (Garage) | `/opt/hermes/obsidian-vault/` | ❌ Defunct — bucket empty at /etc/fstab entry |
|
||||||
|
|
||||||
|
**READ-ONLY guarantee:** sync_obsidian_to_wiki.py copies FROM vault TO wiki target. It NEVER modifies or deletes source files. The vault is mounted via davfs2 which is inherently read-only by filesystem mode, and the code only calls `copy2()` and `unlink()` on the target path.
|
||||||
|
|
||||||
|
**s3fs/obsidian-vault defunct — do NOT use `/opt/hermes/obsidian-vault/`.** The old s3fs mount is empty. The canonical vault is WebDAV at `/mnt/yandex-disk/obsidian/mozg/`. Any cron job or script referencing `/opt/hermes/obsidian-vault/` is stale — redirect to the WebDAV path. The stale mount has 8 root-owned leftover files; they are NOT the real vault.
|
||||||
|
|
||||||
|
**WebDAV switch-over (2026-07-16):** The s3fs mount at `s3.nixg.ru` was empty. Changed `OBSIDIAN_VAULT` from `/opt/hermes/obsidian-vault` to `/mnt/yandex-disk/obsidian/mozg/`. First sync: 74 files in 44s, 0 errors.
|
||||||
|
|
||||||
|
## Hermes Cron Setup
|
||||||
|
|
||||||
|
The sync is a managed Hermes cron job:
|
||||||
|
|
||||||
|
| Property | Value |
|
||||||
|
|---|---|
|
||||||
|
| Job ID | `2346a68b601d` |
|
||||||
|
| Schedule | `*/30 * * * *` |
|
||||||
|
| Runner | Wrapper script at `~/.hermes/scripts/obsidian-sync.sh` |
|
||||||
|
| no_agent | `true` — pure shell execution, no LLM |
|
||||||
|
| Script | `cd /opt/hermes/memory-os && .venv/bin/python3 scripts/sync_obsidian_to_wiki.py` |
|
||||||
|
|
||||||
|
The wrapper exists because Hermes cron's script field requires a path under `~/.hermes/scripts/`. It calls `.venv/bin/python3 scripts/sync_obsidian_to_wiki.py`.
|
||||||
|
|
||||||
|
**sync → ingest chain:** `sync_obsidian_to_wiki.py` internally calls `wiki_continuous_ingest.py` when files change. Must use `.venv/bin/python3` (system python3 lacks `arq`). Fixed 2026-07-16.
|
||||||
|
|
||||||
|
**no_agent optimization:** The cron job uses `no_agent=true` — pure script execution, no LLM reasoning needed. Avoids wasting tokens on every 30-min tick. Verified 2026-07-19.
|
||||||
|
|
||||||
|
## WebDAV vs s3fs Comparison
|
||||||
|
|
||||||
|
| Aspect | WebDAV (davfs2) | s3fs (Garage) |
|
||||||
|
|---|---|---|
|
||||||
|
| Works | ✅ Active | ❌ Empty mount |
|
||||||
|
| POSIX | ⚠️ `file_mode=600` | ✅ Normal perms |
|
||||||
|
| Safety | ✅ Read-only by fuse | ✅ Script guard |
|
||||||
|
| Config | `webdav.yandex.ru` at `/mnt/yandex-disk` | `s3.nixg.ru` at `/opt/hermes/obsidian-vault` |
|
||||||
|
|
||||||
|
| Component | Path | Role |
|
||||||
|
|---|---|---|
|
||||||
|
| State tracker | `/opt/hermes/email/state/wiki_ingest_state.json` | Tracks which files were queued |
|
||||||
|
| Failures log | `/opt/hermes/email/state/wiki_ingest_failures.json` | Records ingest errors |
|
||||||
|
| Sync state | `/opt/hermes/email/state/obsidian_sync_state.json` | Tracks Obsidian vault sync state (mtime/size per file) |
|
||||||
|
| Sync script | `scripts/sync_obsidian_to_wiki.py` | Syncs Obsidian vault → wiki-raw/ |
|
||||||
|
| Ingest worker | `scripts/wiki_continuous_ingest.py` | ARQ worker: parse → chunk → embed → Qdrant |
|
||||||
|
| **Bulk ingest** | `scripts/bulk_wiki_ingest_ollama.py` | One-shot re-index: reads all .md, chunks (headings→paragraphs→words), embeds via Ollama bge-m3, stores in Qdrant. Replaces manual pipeline reset. |
|
||||||
|
| Search CLI | `scripts/context_enhancer.py` | 4-level fallback search (hybrid → dense → lexical → SQLite) |
|
||||||
|
| Hermes hook | `icarus/hooks.py` | Auto-injects Qdrant context into responses |
|
||||||
|
| Docker env | `docker/.env` | Settings for Ollama, Qdrant, Redis |
|
||||||
|
|
||||||
|
## Key Environment Variables
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Qdrant
|
||||||
|
QDRANT_URL=http://localhost:6333
|
||||||
|
QDRANT_COLLECTION=knowledge_base
|
||||||
|
|
||||||
|
# Embedding (dense — local Ollama bge-m3 1024d)
|
||||||
|
OLLAMA_EMBEDDING_URL=http://localhost:11434
|
||||||
|
OLLAMA_EMBEDDING_MODEL=bge-m3:latest
|
||||||
|
|
||||||
|
# Embedding dimension (must match collection)
|
||||||
|
EMBEDDING_DIMS=1024
|
||||||
|
|
||||||
|
# Embedding (sparse — BM25 via FastEmbed)
|
||||||
|
# Uses subprocess to ai-lab venv at _FASTEMBED_PYTHON path in context_enhancer.py
|
||||||
|
|
||||||
|
# Redis (for ARQ queue)
|
||||||
|
REDIS_HOST=127.0.0.1
|
||||||
|
REDIS_PORT=6379
|
||||||
|
REDIS_PASSWORD=<password>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Key Commands
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Reset everything (re-ingest from scratch)
|
||||||
|
cd /opt/hermes/memory-os
|
||||||
|
|
||||||
|
# 1. Clear state so all files are considered new
|
||||||
|
echo '{}' > /opt/hermes/email/state/wiki_ingest_state.json
|
||||||
|
|
||||||
|
# 2. Drop and recreate Qdrant collection (1024d COSINE + sparse)
|
||||||
|
python3 -c "
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
c = QdrantClient('http://localhost:6333')
|
||||||
|
c.delete_collection('knowledge_base')
|
||||||
|
c.create_collection(
|
||||||
|
collection_name='knowledge_base',
|
||||||
|
vectors_config=models.VectorParams(size=1024, distance=models.Distance.COSINE),
|
||||||
|
sparse_vectors_config={'sparse': models.SparseVectorParams(
|
||||||
|
index=models.SparseIndexParams(on_disk=False, full_scan_threshold=10000)
|
||||||
|
)}
|
||||||
|
)
|
||||||
|
print('Created 1024d COSINE + sparse')
|
||||||
|
"
|
||||||
|
|
||||||
|
# 3. Bulk re-index all files (with chunking)
|
||||||
|
/opt/hermes/memory-os/.venv/bin/python3 scripts/bulk_wiki_ingest_ollama.py
|
||||||
|
|
||||||
|
# 4. Search
|
||||||
|
python3 scripts/context_enhancer.py "your query" --top-k 5 --format markdown
|
||||||
|
python3 scripts/context_enhancer.py "your query" --top-k 10 --threshold 0.35 --format compact
|
||||||
|
```
|
||||||
|
|
||||||
|
## Structure of Qdrant Collection
|
||||||
|
|
||||||
|
The collection `knowledge_base` must have:
|
||||||
|
- **dense** vectors: **1024 dimensions**, COSINE distance (bge-m3 via Ollama)
|
||||||
|
- **sparse** vectors: BM25 (named `sparse`, NOT `bm25`)
|
||||||
|
|
||||||
|
```python
|
||||||
|
models.VectorParams(size=1024, distance=models.Distance.COSINE)
|
||||||
|
models.SparseVectorParams(
|
||||||
|
index=models.SparseIndexParams(on_disk=False, full_scan_threshold=10000)
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
The sparse vector name MUST be `"sparse"` — this matches what the ingest worker sends. Creating it as `"bm25"` causes a 400 Bad Request error on every write.
|
||||||
|
|
||||||
|
**Historical note:** This was previously 768d (nomic-embed-text). Migrated to 1024d (bge-m3) on 2026-07-16 for multilingual support. See `references/chunking-implementation.md` for the current chunking approach.
|
||||||
|
|
||||||
|
## Debugging Checklist (when search returns nothing or poor results)
|
||||||
|
|
||||||
|
### 1. Check Qdrant has points
|
||||||
|
```bash
|
||||||
|
python3 -c "
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
c = QdrantClient('http://localhost:6333')
|
||||||
|
info = c.get_collection('knowledge_base')
|
||||||
|
print(f'Points: {info.points_count}')
|
||||||
|
print(f'Status: {info.status}')
|
||||||
|
"
|
||||||
|
```
|
||||||
|
If points_count = 0, the collection was never ingested or was dropped.
|
||||||
|
|
||||||
|
### 2. Check sparse vector name matches
|
||||||
|
```bash
|
||||||
|
python3 -c "
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
c = QdrantClient('http://localhost:6333')
|
||||||
|
info = c.get_collection('knowledge_base')
|
||||||
|
print('Sparse vectors:', info.config.params.sparse_vectors)
|
||||||
|
"
|
||||||
|
```
|
||||||
|
If sparse vectors key is `"bm25"` but the worker sends `"sparse"`, you get 400 errors. Recreate the collection with the correct name.
|
||||||
|
|
||||||
|
### 3. Check dense embedding model matches
|
||||||
|
The `context_enhancer.py`'s `embed_query()` must use the SAME model that was used during ingest. If OpenRouter was used for search but local Ollama for ingest (or vice versa), dense vectors won't match and search degrades to lexical fallback.
|
||||||
|
|
||||||
|
Check what the search client uses vs what was used during ingest:
|
||||||
|
```bash
|
||||||
|
# Search client config
|
||||||
|
grep -n "OLLAMA_EMBEDDING_MODEL\|EMBEDDING_MODEL" scripts/context_enhancer.py
|
||||||
|
|
||||||
|
# Ingest config (docker/.env or wiki_continuous_ingest.py)
|
||||||
|
grep -n "embedding\|model" docker/.env
|
||||||
|
grep -rn "embed\|nomic" scripts/wiki_continuous_ingest.py
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. Check FastEmbed BM25 works (sparse engine)
|
||||||
|
```bash
|
||||||
|
python3 -c "
|
||||||
|
from fastembed.sparse import SparseTextEmbedding
|
||||||
|
model = SparseTextEmbedding(model_name='Qdrant/bm25')
|
||||||
|
sparse = list(model.embed(['test query']))[0]
|
||||||
|
print(f'indices: {len(sparse.indices)}, values: {len(sparse.values)}')
|
||||||
|
"
|
||||||
|
```
|
||||||
|
If this fails, the FastEmbed subprocess in context_enhancer.py needs the `FASTEMBED_SITEPKGS` env var.
|
||||||
|
|
||||||
|
### 5. Check worker logs
|
||||||
|
```bash
|
||||||
|
docker logs docker-worker-1 --tail 50
|
||||||
|
```
|
||||||
|
Look for 400 errors (bad sparse name), DNS failures (ollama unreachable), or timeouts.
|
||||||
|
|
||||||
|
### 6. Check Docker networking
|
||||||
|
After every `docker compose up -d`, the worker container loses secondary networks:
|
||||||
|
```bash
|
||||||
|
docker network connect ollama_default docker-worker-1
|
||||||
|
docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys'
|
||||||
|
```
|
||||||
|
|
||||||
|
### 7. Check ARQ queue isn't stale
|
||||||
|
```bash
|
||||||
|
redis-cli -a <password> keys 'arq:*' | head -20
|
||||||
|
redis-cli -a <password> llen 'arq:queue:health-check'
|
||||||
|
```
|
||||||
|
If there are stale jobs in the queue, clear them with:
|
||||||
|
```bash
|
||||||
|
redis-cli -a <password> del 'arq:queue:health-check'
|
||||||
|
```
|
||||||
|
|
||||||
|
### 8. Test dense embedding directly
|
||||||
|
```bash
|
||||||
|
curl -s http://localhost:11434/api/embeddings \
|
||||||
|
-d '{"model":"bge-m3:latest","prompt":"test query"}' | \
|
||||||
|
python3 -c "import sys, json; d=json.load(sys.stdin); print(f'Dims: {len(d[\"embedding\"])}')"
|
||||||
|
# Should print "Dims: 1024"
|
||||||
|
```
|
||||||
|
|
||||||
|
### 9. Check collection dimensions match `.env` embedding settings (400 Bad Request)
|
||||||
|
When you get a `400 Client Error: Bad Request for url: http://localhost:6333/collections/knowledge_base/points/query`, the most likely cause is the collection was created with one embedding dimension (e.g. 768d for nomic-embed-text) but `.env` now points to a different model with different dimensions (e.g. 1024d for bge-m3).
|
||||||
|
|
||||||
|
Cross-check:
|
||||||
|
```bash
|
||||||
|
# What the collection expects:
|
||||||
|
python3 -c "
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
c = QdrantClient('http://localhost:6333')
|
||||||
|
info = c.get_collection('knowledge_base')
|
||||||
|
print(f'Collection dims: {info.config.params.vectors.size}')
|
||||||
|
print(f'Distance: {info.config.params.vectors.distance}')
|
||||||
|
"
|
||||||
|
|
||||||
|
# What .env says:
|
||||||
|
grep -E 'EMBEDDING_DIMS|EMBEDDING_MODEL' /opt/hermes/memory-os/.env
|
||||||
|
```
|
||||||
|
|
||||||
|
If they don't match, you must either:
|
||||||
|
- (a) **Recreate the collection** with the correct dimension (destructive — lose all points). Use the current model's dimension.
|
||||||
|
- (b) **Change `.env`** to match the collection (switch back to the old model).
|
||||||
|
- (c) **Create a second collection** for the new dimension, keep the old one.
|
||||||
|
|
||||||
|
> **Current config (2026-07-16):** bge-m3 1024d COSINE. See `references/chunking-implementation.md` for the chunking approach.
|
||||||
|
|
||||||
|
> See also `references/dimension-mismatch-s3fs-blockers.md` for full reproduction transcript.
|
||||||
|
|
||||||
|
## Pitfalls
|
||||||
|
|
||||||
|
- **All state files live under `/opt/hermes/email/state/`, NOT `~/.hermes/`.** Three scripts reference state files — sync_obsidian_to_wiki.py, wiki_continuous_ingest.py, dlq_manager.py. All use `/opt/hermes/email/state/`. If any script still uses `Path.home() / ".hermes"` or `os.path.expanduser("~/.hermes")` for state paths, patch it. The old paths at `~/.hermes/obsidian_sync_state.json`, `~/.hermes/wiki_ingest_state.json`, and `~/.hermes/wiki_ingest_failures.json` are stale.
|
||||||
|
- **Missing wrapper script blocks cron execution silently.** The Hermes cron job `obsidian-sync` (job_id `2346a68b601d`) runs `~/.hermes/scripts/obsidian-sync.sh`. If this file is missing, the cron tick produces no output and no error — the job just does nothing. On 2026-07-19 the script was documented in the skill but never created on disk. **Fix:** `mkdir -p ~/.hermes/scripts`; write the wrapper; `chmod +x`. Then switch the job to `no_agent=true` so pure script execution doesn't burn LLM tokens.
|
||||||
|
- **`~` expansion differs between Python and the shell when running from `/opt/hermes/memory-os/`.
|
||||||
|
- **`~` expansion differs between Python and the shell when running from `/opt/hermes/memory-os/`.** `context_enhancer.py` uses `os.path.expanduser("~/.hermes/state.db")` for `LINEAGE_DB`. When the script runs via `.venv/bin/python` from `/opt/hermes/memory-os/`, `~` resolves to `/opt/hermes/` — NOT `/home/estorozhenko/`. This means `LINEAGE_DB` becomes `/opt/hermes/.hermes/state.db` — a different Hermes state DB that doesn't have the `lineage` table. The `register_lineage()` call then silently fails with `[LINEAGE-WARNING] Failed to register lineage: no such table: lineage`. **Fix:** create the `lineage` table in `/opt/hermes/.hermes/state.db` too, or set `STATE_DB_PATH=/home/estorozhenko/.hermes/state.db` in `.env` or process environment. Symptom: search results appear but `[LINEAGE-WARNING]` is printed to stderr.
|
||||||
|
- **Fabric directory (`~/fabric/`) does not exist by default.** The Icarus plugin writes fabric entries to `FABRIC_DIR = Path.home() / "fabric"` (line 19 of `state.py`). This directory is never auto-created before `write_entry()` is called (it calls `mkdir(parents=True, exist_ok=True)` in `write_entry()` itself, so writes succeed, but `read_recent()` and `read_cross_agent()` return empty if `FABRIC_DIR.exists()` is False). **Fix:** `mkdir -p ~/fabric/` to ensure the directory exists before any Icarus session starts. Without this, the first session after enabling Icarus will have no `[fabric]` context injected.
|
||||||
|
- **`HERMES_AGENT_NAME` is not set.** If `HERMES_AGENT_NAME` is missing from `.env` or environment, `state.AGENT_NAME` is empty string, and the logger warns: `"icarus: HERMES_AGENT_NAME not set — fabric entries will use agent=\"agent\""`. This means all fabric entries are tagged with `agent: agent` instead of a meaningful name. **Fix:** add `HERMES_AGENT_NAME=hermes` (or another name) to `/opt/hermes/.hermes/config.yaml` or `/opt/hermes/.hermes/.env`. The name is used in fabric filenames (`agent-entry_type-slug-id.md`) and in multi-agent deployments.
|
||||||
|
- **Fabric entries are NOT injected until mid-session — they only appear starting from turn 2+.** `pre_llm_call` in `hooks.py` calls `state.recall()` which reads fabric entries. On the first turn of a session, `is_first_turn=True` additionally triggers `_search_facts()` for durable facts. Fabric entries from a previous session are injected as `[fabric]` blocks. If you started a new session and see no `[fabric]` block, the directory may not exist, or no entries were written by the previous session (because `on_session_end` writes them, and if the session was closed abnormally, the hook never fired).
|
||||||
|
- **`memory_store.db` (durable facts) is empty by default.** The `facts` table in `~/.hermes/memory_store.db` starts with 0 rows. It's populated only by explicit `memory` tool calls. A `_search_facts()` call always returns empty until the first fact is saved. This is normal — no fix needed.
|
||||||
|
- **Session history FTS5 (`_search_sessions`) needs a session_id exclusion to avoid self-referencing.** The `_search_sessions()` function in `hooks.py` (line 331) filters out the current session by `session_id` to avoid injecting the current conversation's own messages. If this filter breaks, the agent will see its own replies from the same session and recursively inject them. The filter uses `WHERE session_id != ?` in the SQL.
|
||||||
|
- **Session history injection produces `[sessions]` blocks, NOT `[fabric]` or `[qdrant]`.** The `pre_llm_call` hook injects four separate blocks: `[fabric]` (from fabric entries), `[qdrant]` (from Qdrant vault search), `[sessions]` (from FTS5 session history), and `[facts]` (from durable facts). Each has its own dedup set (`_injected_fabric`, `_injected_qdrant`, `_injected_sessions`, `_injected_facts`). If you see a `[sessions]` block, it came from FTS5, not Qdrant.
|
||||||
|
|
||||||
|
- **CRITICAL: icarus/hooks.py threshold must be 0.30, not 0.55.** `_search_qdrant()` passes `threshold=0.55` to `search_with_fallback()`. RRF fusion (hybrid dense+sparse) returns scores 0.33-0.50 even for good matches — the old gate silently filters EVERYTHING. This is fundamentally different from dense-only search which returns 0.90+ for the same queries. A threshold of 0.35 still blocks some RRF results (0.33 scores). The safe floor is 0.30. Verified 2026-07-16: hybrid search with threshold=0.55 returned 0 results for every query tested; 0.30 returned 4 results. Without this fix, `search_with_fallback` hits `level="none"`, cascades into SQLite fallback (`[CE-FALLBACK] SQLite search failed: no such table: lineage`), pollutes logs, and Icarus injects nothing on every turn — the injection code runs, finds nothing, and returns silently.
|
||||||
|
- **`FASTEMBED_VENV` must be set in `.env` or the sparse BM25 subprocess uses system python without fastembed.** If `context_enhancer.py` doesn't find `FASTEMBED_VENV` in env, it falls back to `sys.executable` (system python). The `.env` must contain `FASTEMBED_VENV=/opt/hermes/memory-os/.venv/bin/python3` and `FASTEMBED_SITEPKGS=/opt/hermes/memory-os/.venv/lib/python3.12/site-packages`. Also `context_enhancer.py` needs `load_dotenv()` to read `.env` (added 2026-07-16).
|
||||||
|
- **`python-dotenv` must be installed system-wide** if any script uses `load_dotenv()`. Install with `pip3 install --break-system-packages python-dotenv`.
|
||||||
|
- **Sparse vector name mismatch is the most common ingest failure.** The collection must name the sparse config `"sparse"`, not `"bm25"` or anything else. The worker always sends to `"sparse"`.
|
||||||
|
- **Dense model mismatch kills semantic search.** The search client must use the exact same embedding model as the ingest pipeline. Mixing local Ollama with OpenRouter embeddings (or different models) produces near-zero semantic similarity scores.
|
||||||
|
- **Docker compose up -d disconnects secondary networks.** After any `docker compose up -d`, the worker loses connection to `ollama_default`. Always re-run `docker network connect`.
|
||||||
|
- **FASTEMBED_SITEPKGS must be in subprocess env.** The sparse embedding runs in a subprocess; the env var must be explicitly passed via `env=_env` with `_env.setdefault()`.
|
||||||
|
- **Subprocess python path must be the venv, not system python.** If `_FASTEMBED_PYTHON` points to `/usr/bin/python3` but `fastembed` is installed in `/opt/ai-lab/.venv/bin/python3`, the subprocess silently fails (empty stdout → JSON parse error). Always verify which python can import fastembed, and set the path accordingly.
|
||||||
|
- **`python-dotenv` may be missing in cron/system environment.** The `wiki_continuous_ingest.py` script imports `from dotenv import load_dotenv`. If the script runs via cron or systemd (not the ai-lab venv), it crashes with `ModuleNotFoundError`. Install system-wide: `pip install python-dotenv`, or patch the script to avoid the dependency.
|
||||||
|
- **Cron entry for ingest must use .venv python, not system python.** The cron job `0 * * * * cd /opt/hermes/memory-os && python3 scripts/wiki_continuous_ingest.py` fails with `ModuleNotFoundError` because `dotenv` and `arq` are only in the project's `.venv/`. Fix: use `/opt/hermes/memory-os/.venv/bin/python3` instead of `python3`. Verified 2026-07-16: after the fix, the script runs cleanly (`⏭️ Nada novo. 4 arquivos rastreados, 4 inalterados.`).
|
||||||
|
- **icarus/hooks.py is dead code until registered as a Hermes user plugin.** The plugin lives at `/opt/hermes/.hermes/plugins/icarus/` (NOT in the memory-os project dir). Register with `hermes plugins enable icarus` — NOT by editing config.yaml manually. The plugin has its own `plugin.yaml` (v0.3.0) with 16 tools and 4 hooks (on_session_start, pre_llm_call, post_llm_call, on_session_end). Without `hermes plugins enable`, the file is never executed.
|
||||||
|
- **icarus hooks.py needs a PYTHONPATH fix before Qdrant search works.** The `_search_qdrant()` function at line 285 does `from scripts.context_enhancer import ...` but `/opt/hermes/memory-os` is not on `sys.path`. Fix: add `import sys` at the top of hooks.py, then insert `sys.path.insert(0, '/opt/hermes/memory-os')` inside `_search_qdrant()` before the import. Without this, `_search_qdrant()` always raises `ModuleNotFoundError` and returns an empty list silently (fail-open).
|
||||||
|
- **bge-m3 scores are higher than nomic-embed-text.** With nomic-embed-text (768d), hybrid scores were 0.33-0.50. With bge-m3 (1024d), scores are 0.48-0.65 for Russian queries. The Icarus threshold was raised from 0.30 to 0.40. CLI default stays at 0.35 (fine for dense-only).
|
||||||
|
- **bge-m3 context limit is ~6000 characters.** The Ollama /api/embeddings endpoint returns 500 for texts longer than ~6000 chars. The `bulk_wiki_ingest_ollama.py` script chunks files via `chunk_text()` — recursive split by ##/###/#### headings → paragraphs → words (last resort). `MAX_CHUNK_SIZE=5000`, `CHUNK_OVERLAP=300`. Each chunk becomes a separate Qdrant point with `chunk_index` and `chunk_total` in payload.
|
||||||
|
- **Switching embedding models requires a full re-index.** When changing from nomic-embed-text (768d) to bge-m3 (1024d): delete old Qdrant collection, recreate with correct dims, re-index all files. Both .env and context_enhancer.py must be updated.
|
||||||
|
- **Score threshold has two separate regimes.** Dense-only search (via search_knowledge_base) returns COSINE scores 0.90+ for good matches — threshold 0.35 works fine. Hybrid search (via search_with_fallback with sparse vector, used by Icarus) uses RRF fusion which normalises to 0.33-0.50. The Icarus threshold in hooks.py must be set to 0.30, not 0.35 or 0.55. If you see level="none" and SQLite fallback errors, the threshold is too high for RRF.
|
||||||
|
- **Dimension mismatch between Qdrant collection and .env embedding config is invisible until search time.** If `.env` was changed to a different embedding model with different dimensions (e.g. nomic-embed-text 768d -> bge-m3 1024d) but the Qdrant collection still has the old dimensions, `context_enhancer.py` returns `400 Bad Request` at query time. Cross-check: `collection.config.params.vectors.size` vs `EMBEDDING_DIMS` in `.env`.
|
||||||
|
- **sync_obsidian_to_wiki.py has a safety guard against empty vaults.** If the WebDAV vault mount is empty but the state file has entries, the script skips deletion (`VAULT ПУСТ... пропускаю удаление`). This prevents data loss when the mount is temporarily disconnected. The canonical vault is now WebDAV at `/mnt/yandex-disk/obsidian/mozg/` — s3fs is defunct.
|
||||||
|
- **sync → ingest chain needs .venv python.** Inside `sync_obsidian_to_wiki.py`, the post-sync call to `wiki_continuous_ingest.py` uses `os.system(f"python3 ...")`. System python lacks `arq`. Must use `.venv/bin/python3`. Fixed 2026-07-16.
|
||||||
|
- **State file tracks queue time, not ingest completion.** After resetting the state file and enqueuing, you must wait for the ARQ worker to process before Qdrant has points. Check `wiki_ingest_failures.json` for processing errors.
|
||||||
|
- **Failures.json is append-only** — it accumulates errors across runs. To get a clean view, read it with jq or truncate it when resetting.
|
||||||
|
- **search-api `file` and `source` fields depend on which ingest path stored the data.** The Docker worker stores `file_path` (`/wiki/homelab/foo.md`) and `source` (`wiki-homelab`). The bulk ingest script stores `path` (full filesystem path) and `filename` (basename), but no `source`. The search-api was looking for `payload.get("file")` which neither path stores — always returned null. Fixed by cascading through `file_path` → `path` → `filename`. For `source`, if missing or "unknown", the fix extracts the `wiki-` directory component from `file_path`. See `references/search-api-payload-fields.md`. After changing search-api code, rebuild and restart: `docker compose build search-api && docker compose up -d search-api`.
|
||||||
@@ -0,0 +1,97 @@
|
|||||||
|
# Chunking Implementation — 2026-07-16
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
bge-m3 via Ollama `/api/embeddings` rejects inputs >~6000 chars with HTTP 500.
|
||||||
|
The old `bulk_wiki_ingest_ollama.py` truncated files to 6000 chars, losing
|
||||||
|
mid/end content from large documents.
|
||||||
|
|
||||||
|
Affected files (before chunking):
|
||||||
|
- Configuration backup.md (240KB) — only first 6K of 240KB indexed
|
||||||
|
- Конфигурация компакт Микротика.md (19KB) — Ollama 500 on chunk 2 (16083 chars, no `\n\n`)
|
||||||
|
- Any file with base64 images, ANSI escapes, or control characters caused 500s
|
||||||
|
|
||||||
|
## Solution: `chunk_text()` in `bulk_wiki_ingest_ollama.py`
|
||||||
|
|
||||||
|
Three-level recursive split:
|
||||||
|
|
||||||
|
```
|
||||||
|
1. Headings: split by ##, then ###, then ####
|
||||||
|
2. Paragraphs: split by \n\n (within blocks > MAX_CHUNK_SIZE)
|
||||||
|
3. Words: split by space (last resort, when a single paragraph exceeds MAX_CHUNK_SIZE)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Config
|
||||||
|
|
||||||
|
```python
|
||||||
|
MAX_CHUNK_SIZE = 5000 # safe buffer below bge-m3 ~6000 limit
|
||||||
|
CHUNK_OVERLAP = 300 # overlap between consecutive chunks
|
||||||
|
```
|
||||||
|
|
||||||
|
### `sanitize_text()` — Pre-embedding Cleanup
|
||||||
|
|
||||||
|
Removes content that causes Ollama 500:
|
||||||
|
|
||||||
|
- Base64 images (`data:image/...`) → `[IMAGE]` placeholder
|
||||||
|
- Control characters (except `\n`, `\r`, `\t`) → stripped
|
||||||
|
- ANSI escape sequences (`\x1b[...`) → stripped
|
||||||
|
- Excessive whitespace → collapsed
|
||||||
|
|
||||||
|
### Point ID Scheme
|
||||||
|
|
||||||
|
```python
|
||||||
|
unique_id = hashlib.md5(f"{filepath}:chunk:{chunk_index}".encode()).hexdigest()
|
||||||
|
point_id = int(unique_id[:8], 16) & 0x7fffffff
|
||||||
|
```
|
||||||
|
|
||||||
|
### Payload Fields
|
||||||
|
|
||||||
|
```python
|
||||||
|
{
|
||||||
|
"filename": filepath.name,
|
||||||
|
"path": str(filepath),
|
||||||
|
"content": clean_text[:1000], # preview
|
||||||
|
"length": len(clean_text),
|
||||||
|
"title": title,
|
||||||
|
"chunk_index": chunk_index,
|
||||||
|
"chunk_total": len(chunks),
|
||||||
|
"indexed_at": datetime.now().isoformat(),
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Results
|
||||||
|
|
||||||
|
| Metric | Before | After |
|
||||||
|
|---|---|---|
|
||||||
|
| Total points | 149 | 159 |
|
||||||
|
| Files indexed | 75/85 | 78/85 |
|
||||||
|
| Ollama 500 errors | 2 | 0 |
|
||||||
|
| Configuration backup.md | truncated (1 chunk) | 53 chunks |
|
||||||
|
| Конфигурация компакт Микротика.md | 500 error | 6 chunks |
|
||||||
|
| Цифровой паспорт.md | 2 chunks | 3 chunks |
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Three searches confirmed deep content retrieval:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. "заселение по биометрии" → Цифровой паспорт (score 0.61)
|
||||||
|
# 2. "TemperatureLimit DCMIConfiguration thermal" → Configuration backup (score 0.73)
|
||||||
|
# 3. "MikroTik компактный экспорт" → Конфигурация Микротика (score 0.62)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Pitfalls Encountered
|
||||||
|
|
||||||
|
1. **Ollama 500 on large paragraphs:** Some files (like конфигурация Микротика) have
|
||||||
|
code blocks with no `\n\n` for hundreds of lines. The paragraph split returned
|
||||||
|
one element > MAX_CHUNK_SIZE. Fix: added word-level split as last resort.
|
||||||
|
|
||||||
|
2. **Base64 images in markdown:** Files like Цифровой паспорт.md contain forwarded
|
||||||
|
emails with embedded base64 images. Ollama 500 on the raw data. Fix: `sanitize_text()`
|
||||||
|
regex replaces `data:image/...;base64,...` with `[base64-data]`.
|
||||||
|
|
||||||
|
3. **ANSI escape sequences:** Config files with terminal control codes. Fix: strip
|
||||||
|
`\x1b\[[0-9;]*[a-zA-Z]` patterns.
|
||||||
|
|
||||||
|
4. **Empty files produce Ollama error:** 7 files in vault are 0 bytes. Skip with
|
||||||
|
`if not content.strip(): return 0`.
|
||||||
@@ -0,0 +1,61 @@
|
|||||||
|
# Dimension Mismatch & s3fs Blockers — 2026-07-16
|
||||||
|
|
||||||
|
## Issue 1: Qdrant 400 Bad Request (Dimension Mismatch)
|
||||||
|
|
||||||
|
**Symptom:**
|
||||||
|
```
|
||||||
|
[CE-FALLBACK] Qdrant general error (400 Client Error: Bad Request
|
||||||
|
for url: http://localhost:6333/collections/knowledge_base/points/query),
|
||||||
|
falling back to lexical.
|
||||||
|
[CE-FALLBACK] All fallback levels exhausted.
|
||||||
|
```
|
||||||
|
|
||||||
|
**Root cause:**
|
||||||
|
The Qdrant collection `knowledge_base` was created with 768d vectors (nomic-embed-text via Ollama), but `.env` was later changed to point at Polza.ai with `qwen3-embedding-8b` (4096d). The `context_enhancer.py` reads `.env` and sends 4096d query vectors to a 768d collection — Qdrant rejects with 400.
|
||||||
|
|
||||||
|
**Cross-check:**
|
||||||
|
```
|
||||||
|
# Collection says: 768d
|
||||||
|
# .env says: 4096d
|
||||||
|
```
|
||||||
|
|
||||||
|
**Resolution options:**
|
||||||
|
- (a) Change `.env` back to Ollama nomic-embed-text 768d (local, free, matches collection)
|
||||||
|
- (b) Recreate the collection at 4096d (destructive — lose all 683 points)
|
||||||
|
- (c) Create a second collection for 4096d, keep the old one
|
||||||
|
|
||||||
|
**Note:** The `wiki_continuous_ingest.py` worker reads from the same `.env` — points may also be upserted with wrong dimensions. Also the `docker/.env` file may differ from the project root `.env`.
|
||||||
|
|
||||||
|
## Issue 2: s3fs Obsidian Vault Mount Empty
|
||||||
|
|
||||||
|
**Symptom:**
|
||||||
|
```
|
||||||
|
ls -la /opt/hermes/obsidian-vault/
|
||||||
|
total 8
|
||||||
|
drwxr-xr-x 2 estorozhenko estorozhenko 4096 ...
|
||||||
|
drwxr-xr-x 6 estorozhenko estorozhenko 4096 ...
|
||||||
|
```
|
||||||
|
|
||||||
|
Only `.` and `..` — the directory exists but is empty.
|
||||||
|
|
||||||
|
**fstab entry:**
|
||||||
|
```
|
||||||
|
s3fs#obsidian /opt/hermes/obsidian-vault fuse _netdev,allow_other,
|
||||||
|
passwd_file=/etc/passwd-s3fs,use_cache=/tmp,
|
||||||
|
url=https://s3.nixg.ru,endpoint=garage,
|
||||||
|
use_path_request_style,sigv4 0 0
|
||||||
|
```
|
||||||
|
|
||||||
|
This is a Garage S3 bucket mounted via s3fs. The mount point exists but no data is visible.
|
||||||
|
|
||||||
|
**Impact:**
|
||||||
|
- `sync_obsidian_to_wiki.py` cannot sync (no source files)
|
||||||
|
- The script has a safety guard: if vault is empty but state has files, it skips deletion (`VAULT ПУСТ — пропускаю удаление`). This prevents data loss if the mount is temporarily disconnected.
|
||||||
|
- C1 (schedule sync) is blocked until mount is restored
|
||||||
|
|
||||||
|
**Diagnosis steps:**
|
||||||
|
1. `mount | grep s3fs` — check if mount is active
|
||||||
|
2. `systemctl status` for s3fs — check if automount service is running
|
||||||
|
3. `sudo journalctl -u` for s3fs — check for errors
|
||||||
|
4. Direct HTTP check against `s3.nixg.ru` — test S3 endpoint reachability
|
||||||
|
5. Verify `passwd-s3fs` file exists and has correct credentials
|
||||||
@@ -0,0 +1,241 @@
|
|||||||
|
# Ingest Debugging Sessions
|
||||||
|
|
||||||
|
Chronological reproduction of the pipeline from zero to working search. Each session
|
||||||
|
covers issues encountered and fixes applied.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Session 1: 2026-06-20 to 2026-06-21 (Initial Pipeline Bring-Up)
|
||||||
|
|
||||||
|
Full reproduction of the pipeline from zero to working search.
|
||||||
|
|
||||||
|
## Initial State
|
||||||
|
|
||||||
|
- Qdrant collection `knowledge_base` exists with dense (768d) + sparse vectors
|
||||||
|
- `wiki_ingest_state.json` has 467 files, all with `ingested_at` set
|
||||||
|
- Qdrant points_count = 0 — nothing actually stored
|
||||||
|
- ARQ queue empty — stale jobs consumed
|
||||||
|
- Worker container healthy but was crashing on DNS
|
||||||
|
|
||||||
|
## Issue 1: Worker DNS Failure
|
||||||
|
|
||||||
|
**Symptom:** Worker logs show `Temporary failure in name resolution` when accessing `ollama:11434`
|
||||||
|
|
||||||
|
**Root cause:** Worker container not connected to `ollama_default` Docker network.
|
||||||
|
`docker compose up -d` recreating the container detaches secondary networks.
|
||||||
|
|
||||||
|
**Fix:**
|
||||||
|
```bash
|
||||||
|
docker network connect ollama_default docker-worker-1
|
||||||
|
```
|
||||||
|
Verify:
|
||||||
|
```bash
|
||||||
|
docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys'
|
||||||
|
```
|
||||||
|
Should show both `docker_default` and `ollama_default`.
|
||||||
|
|
||||||
|
## Issue 2: Sparse Vector Name Mismatch
|
||||||
|
|
||||||
|
**Symptom:** Worker completes but Qdrant points_count stays at 0. Logs show HTTP 400 errors.
|
||||||
|
|
||||||
|
**Root cause:** The Qdrant collection was created with sparse vectors named `"bm25"`,
|
||||||
|
but the ingest worker sends sparse vectors keyed as `"sparse"`:
|
||||||
|
|
||||||
|
```python
|
||||||
|
# What the collection had:
|
||||||
|
sparse_vectors_config={'bm25': SparseVectorParams(...)}
|
||||||
|
|
||||||
|
# What the worker sends:
|
||||||
|
{'points': [..., 'vector': {'sparse': ..., 'dense': ...}]}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Fix:** Recreate collection with consistent naming:
|
||||||
|
```python
|
||||||
|
c.delete_collection('knowledge_base')
|
||||||
|
c.create_collection(
|
||||||
|
collection_name='knowledge_base',
|
||||||
|
vectors_config=VectorParams(size=768, distance=Distance.COSINE),
|
||||||
|
sparse_vectors_config={'sparse': SparseVectorParams(
|
||||||
|
index=SparseIndexParams(on_disk=False, full_scan_threshold=10000)
|
||||||
|
)}
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Issue 3: Dense Embedding Model Mismatch (Search)
|
||||||
|
|
||||||
|
**Symptom:** After ingest works and Qdrant has 683 points, `context_enhancer.py` search
|
||||||
|
returns only low-score or no results, falling back to lexical search.
|
||||||
|
|
||||||
|
**Root cause:** `context_enhancer.py` was configured to use OpenRouter `qwen/qwen3-embedding-8b`
|
||||||
|
for query embedding, but the ingested collection uses Ollama `nomic-embed-text:latest` (768d).
|
||||||
|
Different models produce incompatible vector spaces — cosine similarity is near zero.
|
||||||
|
|
||||||
|
**Fix in context_enhancer.py:**
|
||||||
|
```python
|
||||||
|
# Before (OpenRouter):
|
||||||
|
EMBEDDING_MODEL = "qwen/qwen3-embedding-8b"
|
||||||
|
resp = requests.post("https://openrouter.ai/api/v1/embeddings", ...)
|
||||||
|
|
||||||
|
# After (local Ollama):
|
||||||
|
OLLAMA_EMBEDDING_URL = "http://localhost:11434"
|
||||||
|
OLLAMA_EMBEDDING_MODEL = "nomic-embed-text:latest"
|
||||||
|
resp = requests.post(f"{OLLAMA_EMBEDDING_URL}/api/embeddings",
|
||||||
|
json={"model": OLLAMA_EMBEDDING_MODEL, "prompt": text})
|
||||||
|
```
|
||||||
|
|
||||||
|
Also clean up icarus/hooks.py which had OPENROUTER_API_KEY env manipulation that
|
||||||
|
became dead code.
|
||||||
|
|
||||||
|
## Issue 4: FastEmbed BM25 Subprocess Failure
|
||||||
|
|
||||||
|
**Symptom:** Sparse embedding returns None silently. `context_enhancer.py` falls back to
|
||||||
|
dense-only or lexical.
|
||||||
|
|
||||||
|
**Root cause:** The subprocess relies on `FASTEMBED_SITEPKGS` env var pointing to the
|
||||||
|
ai-lab venv site-packages, but subprocess.run() does NOT inherit the parent's env vars
|
||||||
|
automatically when the env was set in Python (not the shell).
|
||||||
|
|
||||||
|
```python
|
||||||
|
# BROKEN — subprocess doesn't see FASTEMBED_SITEPKGS:
|
||||||
|
result = subprocess.run(
|
||||||
|
[_FASTEMBED_PYTHON, "-c", "...import fastembed..."],
|
||||||
|
input=text, capture_output=True, text=True, timeout=15
|
||||||
|
)
|
||||||
|
|
||||||
|
# FIXED — explicitly pass env:
|
||||||
|
_env = os.environ.copy()
|
||||||
|
_env.setdefault("FASTEMBED_SITEPKGS", _FASTEMBED_SITEPKGS)
|
||||||
|
result = subprocess.run(
|
||||||
|
[_FASTEMBED_PYTHON, "-c", "..."],
|
||||||
|
input=text, capture_output=True, text=True, timeout=15, env=_env
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Issue 5: Score Threshold Too High
|
||||||
|
|
||||||
|
**Symptom:** nomic-embed-text results have scores in the 0.30-0.56 range, filter out
|
||||||
|
most results.
|
||||||
|
|
||||||
|
**Fix:** Lowered default threshold from 0.55 to 0.35.
|
||||||
|
|
||||||
|
## Healthy Config (Final)
|
||||||
|
|
||||||
|
```
|
||||||
|
Qdrant: localhost:6333, collection="knowledge_base", 683 points
|
||||||
|
Redis: 127.0.0.1:6379 (authenticated)
|
||||||
|
Ollama: localhost:11434, model="nomic-embed-text:latest" (768d)
|
||||||
|
Worker: docker-worker-1, connected to ollama_default network
|
||||||
|
FastEmbed: BM25 via ai-lab venv subprocess
|
||||||
|
Cron: wiki-ingest-sync, every 10 minutes
|
||||||
|
State: ~/.hermes/wiki_ingest_state.json (tracking 467 files)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Session 2: 2026-07-15 (Three New Blockers — python-dotenv, Sparse Embed Path, Hermes Hook Integration)
|
||||||
|
|
||||||
|
Reproduced the stack from scratch. Three blocking issues found on top of the original five.
|
||||||
|
|
||||||
|
### Blocker 1: `python-dotenv` Missing in Cron Environment
|
||||||
|
|
||||||
|
**Symptom:** cron `wiki-ingest-sync` job fails silently. Logs: `ModuleNotFoundError: No module named 'dotenv'`.
|
||||||
|
|
||||||
|
**Root cause:** The cron job runs under the system environment, which does NOT have `python-dotenv` installed. The `wiki_continuous_ingest.py` script imports `from dotenv import load_dotenv` at the top.
|
||||||
|
|
||||||
|
**Fix:** Install `python-dotenv` system-wide:
|
||||||
|
```bash
|
||||||
|
pip install python-dotenv
|
||||||
|
```
|
||||||
|
Or patch the script to load `.env` via `os.environ` + `open()` instead of `python-dotenv`.
|
||||||
|
|
||||||
|
**Verification:**
|
||||||
|
```bash
|
||||||
|
python3 -c "from dotenv import load_dotenv; print('OK')"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Blocker 2: Sparse Embedding Returns `Expecting value: line 1 column 1`
|
||||||
|
|
||||||
|
**Symptom:** `context_enhancer.py` sparse embedding crashes mid-search with `json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)`.
|
||||||
|
|
||||||
|
**Root cause:** The FastEmbed subprocess (`_FASTEMBED_PYTHON`) runs a Python script that tries to `import fastembed` and `import dotenv`. The subprocess's python path points to `/usr/bin/python3` which does NOT have `fastembed` or `dotenv` in its site-packages. The subprocess fails silently (prints nothing to stdout), and the parent tries `json.loads(result.stdout)` on empty output.
|
||||||
|
|
||||||
|
The existing fix from Session 1 (Issue #4 — `FASTEMBED_SITEPKGS` env var) may have been applied, but the fundamental problem is the subprocess PYTHON PATH itself, not the env var. If `/usr/bin/python3` cannot `import fastembed` at all, setting the env var won't help.
|
||||||
|
|
||||||
|
**Fix:** Use the ai-lab venv python directly:
|
||||||
|
```python
|
||||||
|
_FASTEMBED_PYTHON = "/opt/ai-lab/.venv/bin/python3"
|
||||||
|
```
|
||||||
|
instead of:
|
||||||
|
```python
|
||||||
|
_FASTEMBED_PYTHON = "/usr/bin/python3"
|
||||||
|
```
|
||||||
|
|
||||||
|
This venv has both `fastembed` and `python-dotenv` installed.
|
||||||
|
|
||||||
|
**Diagnostic:**
|
||||||
|
```bash
|
||||||
|
/opt/ai-lab/.venv/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
|
||||||
|
/usr/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
|
||||||
|
```
|
||||||
|
|
||||||
|
### Blocker 3: Hermes Hook (`icarus/hooks.py`) Not Connected to Hermes
|
||||||
|
|
||||||
|
**Symptom:** The file `/opt/hermes/.hermes/plugins/icarus/hooks.py` exists with all the right logic (search Qdrant, inject context into Hermes responses), but it's NEVER called. No Hermes config, plugin, or hook registration activates it.
|
||||||
|
|
||||||
|
**Root cause:** `icarus/hooks.py` relies on being loaded by Hermes as a user plugin, but it's not enabled. The plugin dir is at `/opt/hermes/.hermes/plugins/icarus/` with its own `plugin.yaml` (v0.3.0, 16 tools, 4 hooks). Two things were needed:
|
||||||
|
|
||||||
|
1. `hermes plugins enable icarus` — the proper CLI command to activate a user plugin
|
||||||
|
2. PYTHONPATH fix inside `hooks.py` — the `_search_qdrant()` function does `from scripts.context_enhancer import ...` but `/opt/hermes/memory-os` is not on `sys.path`. Must add `import sys` at top and `sys.path.insert(0, '/opt/hermes/memory-os')` inside `_search_qdrant()` before the import. Without this, the import raises `ModuleNotFoundError` and `_search_qdrant()` returns empty list silently (fail-open).
|
||||||
|
|
||||||
|
**Fix — two steps:**
|
||||||
|
|
||||||
|
Step 1 — PYTHONPATH in hooks.py:
|
||||||
|
```python
|
||||||
|
# Add at top of hooks.py:
|
||||||
|
import sys
|
||||||
|
|
||||||
|
# Add inside _search_qdrant(), before the import:
|
||||||
|
_MEMORY_OS = "/opt/hermes/memory-os"
|
||||||
|
if _MEMORY_OS not in sys.path:
|
||||||
|
sys.path.insert(0, _MEMORY_OS)
|
||||||
|
```
|
||||||
|
|
||||||
|
Step 2 — Enable plugin:
|
||||||
|
```bash
|
||||||
|
hermes plugins enable icarus
|
||||||
|
# Takes effect on next session. No config.yaml editing needed.
|
||||||
|
```
|
||||||
|
|
||||||
|
**Verification of PYTHONPATH fix:**
|
||||||
|
```bash
|
||||||
|
python3 -c "
|
||||||
|
import sys
|
||||||
|
sys.path.insert(0, '/opt/hermes/memory-os')
|
||||||
|
from scripts.context_enhancer import embed_query, embed_query_sparse
|
||||||
|
print('context_enhancer import: OK')
|
||||||
|
dense = embed_query('test')
|
||||||
|
print(f'Dense: {len(dense)} dims')
|
||||||
|
sparse = embed_query_sparse('test')
|
||||||
|
print(f'Sparse: {len(sparse)} values')
|
||||||
|
"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Verification of plugin status:**
|
||||||
|
```bash
|
||||||
|
hermes plugins list | grep icarus
|
||||||
|
# Should show: icarus │ enabled │ 0.3.0
|
||||||
|
```
|
||||||
|
|
||||||
|
## Qdrant Collection Verification
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python3 -c "
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
c = QdrantClient('http://localhost:6333')
|
||||||
|
info = c.get_collection('knowledge_base')
|
||||||
|
print(f'Points: {info.points_count}')
|
||||||
|
print(f'Status: {info.status}')
|
||||||
|
print(f'Dense dims: {info.config.params.vectors.size}')
|
||||||
|
print(f'Sparse keys: {list(info.config.params.sparse_vectors.keys())}')
|
||||||
|
"
|
||||||
|
```
|
||||||
@@ -0,0 +1,90 @@
|
|||||||
|
# Search API — Payload Field Mapping (2026-07-16)
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
`POST /search` returned `"file": null` and `"source": "unknown"` for all results.
|
||||||
|
The search-api looked for `payload.file` and `payload.source`, but neither field
|
||||||
|
exists in the Qdrant payload as stored by the ingest pipeline.
|
||||||
|
|
||||||
|
## Payload Fields by Source
|
||||||
|
|
||||||
|
### Docker Worker (`docker/worker/tasks/file_ingestion.py`)
|
||||||
|
|
||||||
|
Stores these payload fields (line 286-311):
|
||||||
|
|
||||||
|
| Field | Example | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `text` | `"Скрипт настройки..."` | Chunk content |
|
||||||
|
| `source` | `"wiki-homelab"` | Derived from `get_source_tag()` — path relative to WIKI_PATH |
|
||||||
|
| `file_path` | `"/wiki/homelab/MikroTik RouterBOARD.md"` | Full path inside the container |
|
||||||
|
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
|
||||||
|
| `tags` | `["networking", "router"]` | Optional |
|
||||||
|
| `chunk_index` | `0` | Zero-based chunk number |
|
||||||
|
| `chunk_total` | `5` | Total chunks for this file |
|
||||||
|
|
||||||
|
**NO `file` field.** NO `filename` field. NO `path` field.
|
||||||
|
|
||||||
|
### Bulk Ingest (`scripts/bulk_wiki_ingest_ollama.py`)
|
||||||
|
|
||||||
|
Stores different payload fields (line 299-304):
|
||||||
|
|
||||||
|
| Field | Example | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| `filename` | `"MikroTik RouterBOARD RBD520-5HacD2HnD.md"` | `filepath.name` |
|
||||||
|
| `path` | `"/opt/hermes/memory-os/wiki-raw/homelab/MikroTik.md"` | `str(filepath)` |
|
||||||
|
| `content` | `"Скрипт настройки..."` | Truncated to 1000 chars preview |
|
||||||
|
| `length` | `15832` | Full chunk length |
|
||||||
|
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
|
||||||
|
| `chunk_index` | `0` | Zero-based chunk number |
|
||||||
|
| `chunk_total` | `5` | Total chunks for this file |
|
||||||
|
|
||||||
|
**NO `file` field.** NO `source` field. NO `file_path` field.
|
||||||
|
|
||||||
|
## The Fix
|
||||||
|
|
||||||
|
The search-api `main.py` results loop was changed from:
|
||||||
|
|
||||||
|
```python
|
||||||
|
file=payload.get("file"), # always None
|
||||||
|
source=payload.get("source", "unknown"), # always "unknown" for bulk ingest
|
||||||
|
```
|
||||||
|
|
||||||
|
To:
|
||||||
|
|
||||||
|
```python
|
||||||
|
file_path = payload.get("file_path") or payload.get("path") or ""
|
||||||
|
filename = os.path.basename(file_path) if file_path else payload.get("filename")
|
||||||
|
|
||||||
|
source = payload.get("source")
|
||||||
|
if not source or source == "unknown":
|
||||||
|
if file_path:
|
||||||
|
parts = file_path.split("/")
|
||||||
|
wiki_idx = next((i for i, p in enumerate(parts) if p.startswith("wiki-")), -1)
|
||||||
|
if wiki_idx >= 0:
|
||||||
|
source = parts[wiki_idx]
|
||||||
|
else:
|
||||||
|
source = "wiki"
|
||||||
|
else:
|
||||||
|
source = "unknown"
|
||||||
|
```
|
||||||
|
|
||||||
|
## Future-Proofing
|
||||||
|
|
||||||
|
If a new ingest path is added (e.g. a CLI tool or API), it should store EITHER:
|
||||||
|
- `file_path` (full path) — the search-api extracts basename
|
||||||
|
- `filename` + `path` (separate fields) — the search-api prefers `filename` if no path available
|
||||||
|
|
||||||
|
The search-api is now tolerant of both schemes. If a new field is introduced, add it to the
|
||||||
|
`file_path or payload.get("path")` chain in `_extract_filename()`. But the cleaner approach
|
||||||
|
is to standardise all ingest paths on `file_path`.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -s -X POST http://localhost:8000/search \
|
||||||
|
-H "Content-Type: application/json" \
|
||||||
|
-d '{"query":"настройка VPN", "top_k": 3}' | python3 -m json.tool
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected: `"file": "MikroTik RouterBOARD RBD520-5HacD2HnD.md"` (not null),
|
||||||
|
`"source": "wiki"` (not "unknown").
|
||||||
@@ -0,0 +1,90 @@
|
|||||||
|
# Icarus Threshold Bug — 2026-07-16 (v2)
|
||||||
|
|
||||||
|
## The Problem
|
||||||
|
|
||||||
|
`icarus/hooks.py` calls `_search_qdrant(query, top_k=2, threshold=0.55)` from `pre_llm_call()`.
|
||||||
|
This passes `score_threshold=0.55` to `search_with_fallback()` in `context_enhancer.py`.
|
||||||
|
|
||||||
|
RRF fusion (hybrid dense+sparse) returns scores 0.33–0.50 even for good matches.
|
||||||
|
At 0.55, every query is filtered out. The search cascade falls to `level="none"`,
|
||||||
|
then hits SQLite fallback (`[CE-FALLBACK] SQLite search failed: no such table: lineage`),
|
||||||
|
and finally returns an empty list. Icarus injects nothing.
|
||||||
|
|
||||||
|
This is fundamentally different from dense-only search, which returns COSINE scores
|
||||||
|
0.90+ for the same queries. The threshold bug was invisible because:
|
||||||
|
- `_search_qdrant()` is fail-open (returns `[]` on any error)
|
||||||
|
- The plugin loads and runs without crashing
|
||||||
|
- "Runs without crashing" ≠ "returns useful results"
|
||||||
|
|
||||||
|
## Root Cause: RRF vs Dense Score Regimes
|
||||||
|
|
||||||
|
RRF (Reciprocal Rank Fusion) normalises scores from two independent retrievers
|
||||||
|
(dense and sparse) into a shared 0–1 range via `1/(k + rank)`. This inherently
|
||||||
|
produces clustered scores around 0.33–0.50 regardless of the underlying semantic
|
||||||
|
similarity. This is **not a bug in RRF** — it's how RRF works.
|
||||||
|
|
||||||
|
Dense-only search returns raw COSINE similarity (0.90+ for good matches).
|
||||||
|
|
||||||
|
The same query at different thresholds:
|
||||||
|
|
||||||
|
```
|
||||||
|
Query: "XRay VPS VPN"
|
||||||
|
|
||||||
|
=== Hybrid (RRF), threshold 0.55 ===
|
||||||
|
Level: none, Results: 0
|
||||||
|
Falls through to SQLite → [CE-FALLBACK] All fallback levels exhausted.
|
||||||
|
|
||||||
|
=== Dense-only, threshold 0.35 ===
|
||||||
|
Level: dense-only, Results: 3
|
||||||
|
[0.9938] Пароль Юлии Зозули
|
||||||
|
[0.9078] Nagios
|
||||||
|
[0.9017] Клавиатуры
|
||||||
|
|
||||||
|
=== Hybrid (RRF), threshold 0.35 ===
|
||||||
|
Level: hybrid, Results: 2
|
||||||
|
[0.5000] Инструкция как безопасно расширить кластер Garage до 3+ нод (v2.1)
|
||||||
|
[0.5000] Пароль Юлии Зозули
|
||||||
|
|
||||||
|
=== Hybrid (RRF), threshold 0.30 ===
|
||||||
|
Level: hybrid, Results: 4
|
||||||
|
[0.5000] Garage кластер
|
||||||
|
[0.5000] Пароль Юлии Зозули
|
||||||
|
[0.3333] Двухфакторная VPN
|
||||||
|
[0.3333] Nagios
|
||||||
|
```
|
||||||
|
|
||||||
|
At 0.35, RRF still filters the 0.33 results (2 out of 4 are lost).
|
||||||
|
At 0.30, all 4 are returned.
|
||||||
|
|
||||||
|
## The Fix
|
||||||
|
|
||||||
|
Lower `threshold` in `_search_qdrant()` in `/opt/hermes/.hermes/plugins/icarus/hooks.py`
|
||||||
|
from 0.55 to 0.30.
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Line 722 — NOW:
|
||||||
|
qdrant_results = _search_qdrant(user_message, top_k=2, threshold=0.30)
|
||||||
|
```
|
||||||
|
|
||||||
|
The comment above it was also updated to explain the RRF vs dense score gap.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
After fix, all queries return hybrid results within 6ms Qdrant time:
|
||||||
|
|
||||||
|
```
|
||||||
|
XRay VPS VPN → hybrid, 4 results
|
||||||
|
hermes plugin icarus → hybrid, 3 results
|
||||||
|
telegram → hybrid, 4 results
|
||||||
|
wireguard → hybrid, 4 results
|
||||||
|
```
|
||||||
|
|
||||||
|
No SQLite fallback noise. No `level="none"`.
|
||||||
|
|
||||||
|
## Key Insight
|
||||||
|
|
||||||
|
When debugging an empty `[qdrant]` block in Icarus context injection:
|
||||||
|
1. First check if `_search_qdrant()` even runs (no import error → PYTHONPATH fix)
|
||||||
|
2. Then check if it returns results (threshold too high for RRF)
|
||||||
|
3. These are SEPARATE bugs — the first was fixed 2026-07-14 (PYTHONPATH),
|
||||||
|
the second on 2026-07-16 (threshold 0.55→0.30)
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# WebDAV Migration & Sync Pipeline Fix — 2026-07-16
|
||||||
|
|
||||||
|
## Change: s3fs → WebDAV as Obsidian Vault Source
|
||||||
|
|
||||||
|
**Problem:** The s3fs mount at `/opt/hermes/obsidian-vault/` (via Garage S3 `s3.nixg.ru`) was empty — only `.` and `..` in the directory, no actual files. sync_obsidian_to_wiki.py couldn't copy anything. C1 was blocked.
|
||||||
|
|
||||||
|
**Diagnosis:**
|
||||||
|
- `/etc/fstab` had s3fs entry pointing at `s3.nixg.ru` (Garage)
|
||||||
|
- `mount | grep s3fs` showed the mount was active but directory contained 0 files
|
||||||
|
- `/mnt/yandex-disk/` (Yandex Disk via davfs2 WebDAV) had a working `obsidian/mozg/` vault with 65 .md files
|
||||||
|
|
||||||
|
**Fix:** Changed `OBSIDIAN_VAULT` from `/opt/hermes/obsidian-vault` to `/mnt/yandex-disk/obsidian/mozg/`.
|
||||||
|
|
||||||
|
**First run:** 74 files copied in 44 seconds, 0 errors.
|
||||||
|
|
||||||
|
## Blocker 2: sync → ingest chain broken
|
||||||
|
|
||||||
|
**Symptom:** The sync script finishes copying, then tries to call `wiki_continuous_ingest.py` with bare `python3` — crashes with `ModuleNotFoundError: No module named 'arq'`.
|
||||||
|
|
||||||
|
**Root cause:** `sync_obsidian_to_wiki.py` line 162 does `os.system(f"python3 {ingest_script}")`. System python3 does NOT have `arq` (only installed in project `.venv/`).
|
||||||
|
|
||||||
|
**Fix:** Changed to `/opt/hermes/memory-os/.venv/bin/python3 {ingest_script}`.
|
||||||
|
|
||||||
|
## Safety Guarantee
|
||||||
|
|
||||||
|
The user explicitly asked for protection against Hermes deleting files in their Obsidian vault. Confirmed:
|
||||||
|
|
||||||
|
- **Code-level:** sync_obsidian_to_wiki.py only reads source (via `os.walk()` + `copy2()`). All deletions (`unlink()`) target `WIKI_TARGET`, never `OBSIDIAN_VAULT`. This was true before the change and remains true after.
|
||||||
|
- **Filesystem-level:** davfs2 mounts with `file_mode=600` — the user process itself can only write with explicit sudo. The sync script runs as user `estorozhenko`.
|
||||||
|
- **Memory-level:** Added read-only guarantee to user profile.
|
||||||
|
|
||||||
|
**Lesson:** When the user expresses concern about data safety, (1) show code evidence that it's already safe, (2) add explicit protection in memory/skill, (3) explain trade-offs (read-only mount vs full POSIX) so they can make informed decisions.
|
||||||
Reference in New Issue
Block a user