mirror of
https://gitverse.ru/kpa39l/memory-os.git
synced 2026-09-29 09:35:05 +00:00
Initial commit: Hermes skill memory-os
This commit is contained in:
@@ -0,0 +1,97 @@
|
||||
# Chunking Implementation — 2026-07-16
|
||||
|
||||
## Problem
|
||||
|
||||
bge-m3 via Ollama `/api/embeddings` rejects inputs >~6000 chars with HTTP 500.
|
||||
The old `bulk_wiki_ingest_ollama.py` truncated files to 6000 chars, losing
|
||||
mid/end content from large documents.
|
||||
|
||||
Affected files (before chunking):
|
||||
- Configuration backup.md (240KB) — only first 6K of 240KB indexed
|
||||
- Конфигурация компакт Микротика.md (19KB) — Ollama 500 on chunk 2 (16083 chars, no `\n\n`)
|
||||
- Any file with base64 images, ANSI escapes, or control characters caused 500s
|
||||
|
||||
## Solution: `chunk_text()` in `bulk_wiki_ingest_ollama.py`
|
||||
|
||||
Three-level recursive split:
|
||||
|
||||
```
|
||||
1. Headings: split by ##, then ###, then ####
|
||||
2. Paragraphs: split by \n\n (within blocks > MAX_CHUNK_SIZE)
|
||||
3. Words: split by space (last resort, when a single paragraph exceeds MAX_CHUNK_SIZE)
|
||||
```
|
||||
|
||||
### Config
|
||||
|
||||
```python
|
||||
MAX_CHUNK_SIZE = 5000 # safe buffer below bge-m3 ~6000 limit
|
||||
CHUNK_OVERLAP = 300 # overlap between consecutive chunks
|
||||
```
|
||||
|
||||
### `sanitize_text()` — Pre-embedding Cleanup
|
||||
|
||||
Removes content that causes Ollama 500:
|
||||
|
||||
- Base64 images (`data:image/...`) → `[IMAGE]` placeholder
|
||||
- Control characters (except `\n`, `\r`, `\t`) → stripped
|
||||
- ANSI escape sequences (`\x1b[...`) → stripped
|
||||
- Excessive whitespace → collapsed
|
||||
|
||||
### Point ID Scheme
|
||||
|
||||
```python
|
||||
unique_id = hashlib.md5(f"{filepath}:chunk:{chunk_index}".encode()).hexdigest()
|
||||
point_id = int(unique_id[:8], 16) & 0x7fffffff
|
||||
```
|
||||
|
||||
### Payload Fields
|
||||
|
||||
```python
|
||||
{
|
||||
"filename": filepath.name,
|
||||
"path": str(filepath),
|
||||
"content": clean_text[:1000], # preview
|
||||
"length": len(clean_text),
|
||||
"title": title,
|
||||
"chunk_index": chunk_index,
|
||||
"chunk_total": len(chunks),
|
||||
"indexed_at": datetime.now().isoformat(),
|
||||
}
|
||||
```
|
||||
|
||||
## Results
|
||||
|
||||
| Metric | Before | After |
|
||||
|---|---|---|
|
||||
| Total points | 149 | 159 |
|
||||
| Files indexed | 75/85 | 78/85 |
|
||||
| Ollama 500 errors | 2 | 0 |
|
||||
| Configuration backup.md | truncated (1 chunk) | 53 chunks |
|
||||
| Конфигурация компакт Микротика.md | 500 error | 6 chunks |
|
||||
| Цифровой паспорт.md | 2 chunks | 3 chunks |
|
||||
|
||||
## Verification
|
||||
|
||||
Three searches confirmed deep content retrieval:
|
||||
|
||||
```bash
|
||||
# 1. "заселение по биометрии" → Цифровой паспорт (score 0.61)
|
||||
# 2. "TemperatureLimit DCMIConfiguration thermal" → Configuration backup (score 0.73)
|
||||
# 3. "MikroTik компактный экспорт" → Конфигурация Микротика (score 0.62)
|
||||
```
|
||||
|
||||
## Pitfalls Encountered
|
||||
|
||||
1. **Ollama 500 on large paragraphs:** Some files (like конфигурация Микротика) have
|
||||
code blocks with no `\n\n` for hundreds of lines. The paragraph split returned
|
||||
one element > MAX_CHUNK_SIZE. Fix: added word-level split as last resort.
|
||||
|
||||
2. **Base64 images in markdown:** Files like Цифровой паспорт.md contain forwarded
|
||||
emails with embedded base64 images. Ollama 500 on the raw data. Fix: `sanitize_text()`
|
||||
regex replaces `data:image/...;base64,...` with `[base64-data]`.
|
||||
|
||||
3. **ANSI escape sequences:** Config files with terminal control codes. Fix: strip
|
||||
`\x1b\[[0-9;]*[a-zA-Z]` patterns.
|
||||
|
||||
4. **Empty files produce Ollama error:** 7 files in vault are 0 bytes. Skip with
|
||||
`if not content.strip(): return 0`.
|
||||
@@ -0,0 +1,61 @@
|
||||
# Dimension Mismatch & s3fs Blockers — 2026-07-16
|
||||
|
||||
## Issue 1: Qdrant 400 Bad Request (Dimension Mismatch)
|
||||
|
||||
**Symptom:**
|
||||
```
|
||||
[CE-FALLBACK] Qdrant general error (400 Client Error: Bad Request
|
||||
for url: http://localhost:6333/collections/knowledge_base/points/query),
|
||||
falling back to lexical.
|
||||
[CE-FALLBACK] All fallback levels exhausted.
|
||||
```
|
||||
|
||||
**Root cause:**
|
||||
The Qdrant collection `knowledge_base` was created with 768d vectors (nomic-embed-text via Ollama), but `.env` was later changed to point at Polza.ai with `qwen3-embedding-8b` (4096d). The `context_enhancer.py` reads `.env` and sends 4096d query vectors to a 768d collection — Qdrant rejects with 400.
|
||||
|
||||
**Cross-check:**
|
||||
```
|
||||
# Collection says: 768d
|
||||
# .env says: 4096d
|
||||
```
|
||||
|
||||
**Resolution options:**
|
||||
- (a) Change `.env` back to Ollama nomic-embed-text 768d (local, free, matches collection)
|
||||
- (b) Recreate the collection at 4096d (destructive — lose all 683 points)
|
||||
- (c) Create a second collection for 4096d, keep the old one
|
||||
|
||||
**Note:** The `wiki_continuous_ingest.py` worker reads from the same `.env` — points may also be upserted with wrong dimensions. Also the `docker/.env` file may differ from the project root `.env`.
|
||||
|
||||
## Issue 2: s3fs Obsidian Vault Mount Empty
|
||||
|
||||
**Symptom:**
|
||||
```
|
||||
ls -la /opt/hermes/obsidian-vault/
|
||||
total 8
|
||||
drwxr-xr-x 2 estorozhenko estorozhenko 4096 ...
|
||||
drwxr-xr-x 6 estorozhenko estorozhenko 4096 ...
|
||||
```
|
||||
|
||||
Only `.` and `..` — the directory exists but is empty.
|
||||
|
||||
**fstab entry:**
|
||||
```
|
||||
s3fs#obsidian /opt/hermes/obsidian-vault fuse _netdev,allow_other,
|
||||
passwd_file=/etc/passwd-s3fs,use_cache=/tmp,
|
||||
url=https://s3.nixg.ru,endpoint=garage,
|
||||
use_path_request_style,sigv4 0 0
|
||||
```
|
||||
|
||||
This is a Garage S3 bucket mounted via s3fs. The mount point exists but no data is visible.
|
||||
|
||||
**Impact:**
|
||||
- `sync_obsidian_to_wiki.py` cannot sync (no source files)
|
||||
- The script has a safety guard: if vault is empty but state has files, it skips deletion (`VAULT ПУСТ — пропускаю удаление`). This prevents data loss if the mount is temporarily disconnected.
|
||||
- C1 (schedule sync) is blocked until mount is restored
|
||||
|
||||
**Diagnosis steps:**
|
||||
1. `mount | grep s3fs` — check if mount is active
|
||||
2. `systemctl status` for s3fs — check if automount service is running
|
||||
3. `sudo journalctl -u` for s3fs — check for errors
|
||||
4. Direct HTTP check against `s3.nixg.ru` — test S3 endpoint reachability
|
||||
5. Verify `passwd-s3fs` file exists and has correct credentials
|
||||
@@ -0,0 +1,241 @@
|
||||
# Ingest Debugging Sessions
|
||||
|
||||
Chronological reproduction of the pipeline from zero to working search. Each session
|
||||
covers issues encountered and fixes applied.
|
||||
|
||||
---
|
||||
|
||||
## Session 1: 2026-06-20 to 2026-06-21 (Initial Pipeline Bring-Up)
|
||||
|
||||
Full reproduction of the pipeline from zero to working search.
|
||||
|
||||
## Initial State
|
||||
|
||||
- Qdrant collection `knowledge_base` exists with dense (768d) + sparse vectors
|
||||
- `wiki_ingest_state.json` has 467 files, all with `ingested_at` set
|
||||
- Qdrant points_count = 0 — nothing actually stored
|
||||
- ARQ queue empty — stale jobs consumed
|
||||
- Worker container healthy but was crashing on DNS
|
||||
|
||||
## Issue 1: Worker DNS Failure
|
||||
|
||||
**Symptom:** Worker logs show `Temporary failure in name resolution` when accessing `ollama:11434`
|
||||
|
||||
**Root cause:** Worker container not connected to `ollama_default` Docker network.
|
||||
`docker compose up -d` recreating the container detaches secondary networks.
|
||||
|
||||
**Fix:**
|
||||
```bash
|
||||
docker network connect ollama_default docker-worker-1
|
||||
```
|
||||
Verify:
|
||||
```bash
|
||||
docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys'
|
||||
```
|
||||
Should show both `docker_default` and `ollama_default`.
|
||||
|
||||
## Issue 2: Sparse Vector Name Mismatch
|
||||
|
||||
**Symptom:** Worker completes but Qdrant points_count stays at 0. Logs show HTTP 400 errors.
|
||||
|
||||
**Root cause:** The Qdrant collection was created with sparse vectors named `"bm25"`,
|
||||
but the ingest worker sends sparse vectors keyed as `"sparse"`:
|
||||
|
||||
```python
|
||||
# What the collection had:
|
||||
sparse_vectors_config={'bm25': SparseVectorParams(...)}
|
||||
|
||||
# What the worker sends:
|
||||
{'points': [..., 'vector': {'sparse': ..., 'dense': ...}]}
|
||||
```
|
||||
|
||||
**Fix:** Recreate collection with consistent naming:
|
||||
```python
|
||||
c.delete_collection('knowledge_base')
|
||||
c.create_collection(
|
||||
collection_name='knowledge_base',
|
||||
vectors_config=VectorParams(size=768, distance=Distance.COSINE),
|
||||
sparse_vectors_config={'sparse': SparseVectorParams(
|
||||
index=SparseIndexParams(on_disk=False, full_scan_threshold=10000)
|
||||
)}
|
||||
)
|
||||
```
|
||||
|
||||
## Issue 3: Dense Embedding Model Mismatch (Search)
|
||||
|
||||
**Symptom:** After ingest works and Qdrant has 683 points, `context_enhancer.py` search
|
||||
returns only low-score or no results, falling back to lexical search.
|
||||
|
||||
**Root cause:** `context_enhancer.py` was configured to use OpenRouter `qwen/qwen3-embedding-8b`
|
||||
for query embedding, but the ingested collection uses Ollama `nomic-embed-text:latest` (768d).
|
||||
Different models produce incompatible vector spaces — cosine similarity is near zero.
|
||||
|
||||
**Fix in context_enhancer.py:**
|
||||
```python
|
||||
# Before (OpenRouter):
|
||||
EMBEDDING_MODEL = "qwen/qwen3-embedding-8b"
|
||||
resp = requests.post("https://openrouter.ai/api/v1/embeddings", ...)
|
||||
|
||||
# After (local Ollama):
|
||||
OLLAMA_EMBEDDING_URL = "http://localhost:11434"
|
||||
OLLAMA_EMBEDDING_MODEL = "nomic-embed-text:latest"
|
||||
resp = requests.post(f"{OLLAMA_EMBEDDING_URL}/api/embeddings",
|
||||
json={"model": OLLAMA_EMBEDDING_MODEL, "prompt": text})
|
||||
```
|
||||
|
||||
Also clean up icarus/hooks.py which had OPENROUTER_API_KEY env manipulation that
|
||||
became dead code.
|
||||
|
||||
## Issue 4: FastEmbed BM25 Subprocess Failure
|
||||
|
||||
**Symptom:** Sparse embedding returns None silently. `context_enhancer.py` falls back to
|
||||
dense-only or lexical.
|
||||
|
||||
**Root cause:** The subprocess relies on `FASTEMBED_SITEPKGS` env var pointing to the
|
||||
ai-lab venv site-packages, but subprocess.run() does NOT inherit the parent's env vars
|
||||
automatically when the env was set in Python (not the shell).
|
||||
|
||||
```python
|
||||
# BROKEN — subprocess doesn't see FASTEMBED_SITEPKGS:
|
||||
result = subprocess.run(
|
||||
[_FASTEMBED_PYTHON, "-c", "...import fastembed..."],
|
||||
input=text, capture_output=True, text=True, timeout=15
|
||||
)
|
||||
|
||||
# FIXED — explicitly pass env:
|
||||
_env = os.environ.copy()
|
||||
_env.setdefault("FASTEMBED_SITEPKGS", _FASTEMBED_SITEPKGS)
|
||||
result = subprocess.run(
|
||||
[_FASTEMBED_PYTHON, "-c", "..."],
|
||||
input=text, capture_output=True, text=True, timeout=15, env=_env
|
||||
)
|
||||
```
|
||||
|
||||
## Issue 5: Score Threshold Too High
|
||||
|
||||
**Symptom:** nomic-embed-text results have scores in the 0.30-0.56 range, filter out
|
||||
most results.
|
||||
|
||||
**Fix:** Lowered default threshold from 0.55 to 0.35.
|
||||
|
||||
## Healthy Config (Final)
|
||||
|
||||
```
|
||||
Qdrant: localhost:6333, collection="knowledge_base", 683 points
|
||||
Redis: 127.0.0.1:6379 (authenticated)
|
||||
Ollama: localhost:11434, model="nomic-embed-text:latest" (768d)
|
||||
Worker: docker-worker-1, connected to ollama_default network
|
||||
FastEmbed: BM25 via ai-lab venv subprocess
|
||||
Cron: wiki-ingest-sync, every 10 minutes
|
||||
State: ~/.hermes/wiki_ingest_state.json (tracking 467 files)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Session 2: 2026-07-15 (Three New Blockers — python-dotenv, Sparse Embed Path, Hermes Hook Integration)
|
||||
|
||||
Reproduced the stack from scratch. Three blocking issues found on top of the original five.
|
||||
|
||||
### Blocker 1: `python-dotenv` Missing in Cron Environment
|
||||
|
||||
**Symptom:** cron `wiki-ingest-sync` job fails silently. Logs: `ModuleNotFoundError: No module named 'dotenv'`.
|
||||
|
||||
**Root cause:** The cron job runs under the system environment, which does NOT have `python-dotenv` installed. The `wiki_continuous_ingest.py` script imports `from dotenv import load_dotenv` at the top.
|
||||
|
||||
**Fix:** Install `python-dotenv` system-wide:
|
||||
```bash
|
||||
pip install python-dotenv
|
||||
```
|
||||
Or patch the script to load `.env` via `os.environ` + `open()` instead of `python-dotenv`.
|
||||
|
||||
**Verification:**
|
||||
```bash
|
||||
python3 -c "from dotenv import load_dotenv; print('OK')"
|
||||
```
|
||||
|
||||
### Blocker 2: Sparse Embedding Returns `Expecting value: line 1 column 1`
|
||||
|
||||
**Symptom:** `context_enhancer.py` sparse embedding crashes mid-search with `json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)`.
|
||||
|
||||
**Root cause:** The FastEmbed subprocess (`_FASTEMBED_PYTHON`) runs a Python script that tries to `import fastembed` and `import dotenv`. The subprocess's python path points to `/usr/bin/python3` which does NOT have `fastembed` or `dotenv` in its site-packages. The subprocess fails silently (prints nothing to stdout), and the parent tries `json.loads(result.stdout)` on empty output.
|
||||
|
||||
The existing fix from Session 1 (Issue #4 — `FASTEMBED_SITEPKGS` env var) may have been applied, but the fundamental problem is the subprocess PYTHON PATH itself, not the env var. If `/usr/bin/python3` cannot `import fastembed` at all, setting the env var won't help.
|
||||
|
||||
**Fix:** Use the ai-lab venv python directly:
|
||||
```python
|
||||
_FASTEMBED_PYTHON = "/opt/ai-lab/.venv/bin/python3"
|
||||
```
|
||||
instead of:
|
||||
```python
|
||||
_FASTEMBED_PYTHON = "/usr/bin/python3"
|
||||
```
|
||||
|
||||
This venv has both `fastembed` and `python-dotenv` installed.
|
||||
|
||||
**Diagnostic:**
|
||||
```bash
|
||||
/opt/ai-lab/.venv/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
|
||||
/usr/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
|
||||
```
|
||||
|
||||
### Blocker 3: Hermes Hook (`icarus/hooks.py`) Not Connected to Hermes
|
||||
|
||||
**Symptom:** The file `/opt/hermes/.hermes/plugins/icarus/hooks.py` exists with all the right logic (search Qdrant, inject context into Hermes responses), but it's NEVER called. No Hermes config, plugin, or hook registration activates it.
|
||||
|
||||
**Root cause:** `icarus/hooks.py` relies on being loaded by Hermes as a user plugin, but it's not enabled. The plugin dir is at `/opt/hermes/.hermes/plugins/icarus/` with its own `plugin.yaml` (v0.3.0, 16 tools, 4 hooks). Two things were needed:
|
||||
|
||||
1. `hermes plugins enable icarus` — the proper CLI command to activate a user plugin
|
||||
2. PYTHONPATH fix inside `hooks.py` — the `_search_qdrant()` function does `from scripts.context_enhancer import ...` but `/opt/hermes/memory-os` is not on `sys.path`. Must add `import sys` at top and `sys.path.insert(0, '/opt/hermes/memory-os')` inside `_search_qdrant()` before the import. Without this, the import raises `ModuleNotFoundError` and `_search_qdrant()` returns empty list silently (fail-open).
|
||||
|
||||
**Fix — two steps:**
|
||||
|
||||
Step 1 — PYTHONPATH in hooks.py:
|
||||
```python
|
||||
# Add at top of hooks.py:
|
||||
import sys
|
||||
|
||||
# Add inside _search_qdrant(), before the import:
|
||||
_MEMORY_OS = "/opt/hermes/memory-os"
|
||||
if _MEMORY_OS not in sys.path:
|
||||
sys.path.insert(0, _MEMORY_OS)
|
||||
```
|
||||
|
||||
Step 2 — Enable plugin:
|
||||
```bash
|
||||
hermes plugins enable icarus
|
||||
# Takes effect on next session. No config.yaml editing needed.
|
||||
```
|
||||
|
||||
**Verification of PYTHONPATH fix:**
|
||||
```bash
|
||||
python3 -c "
|
||||
import sys
|
||||
sys.path.insert(0, '/opt/hermes/memory-os')
|
||||
from scripts.context_enhancer import embed_query, embed_query_sparse
|
||||
print('context_enhancer import: OK')
|
||||
dense = embed_query('test')
|
||||
print(f'Dense: {len(dense)} dims')
|
||||
sparse = embed_query_sparse('test')
|
||||
print(f'Sparse: {len(sparse)} values')
|
||||
"
|
||||
```
|
||||
|
||||
**Verification of plugin status:**
|
||||
```bash
|
||||
hermes plugins list | grep icarus
|
||||
# Should show: icarus │ enabled │ 0.3.0
|
||||
```
|
||||
|
||||
## Qdrant Collection Verification
|
||||
|
||||
```bash
|
||||
python3 -c "
|
||||
from qdrant_client import QdrantClient, models
|
||||
c = QdrantClient('http://localhost:6333')
|
||||
info = c.get_collection('knowledge_base')
|
||||
print(f'Points: {info.points_count}')
|
||||
print(f'Status: {info.status}')
|
||||
print(f'Dense dims: {info.config.params.vectors.size}')
|
||||
print(f'Sparse keys: {list(info.config.params.sparse_vectors.keys())}')
|
||||
"
|
||||
```
|
||||
@@ -0,0 +1,90 @@
|
||||
# Search API — Payload Field Mapping (2026-07-16)
|
||||
|
||||
## Problem
|
||||
|
||||
`POST /search` returned `"file": null` and `"source": "unknown"` for all results.
|
||||
The search-api looked for `payload.file` and `payload.source`, but neither field
|
||||
exists in the Qdrant payload as stored by the ingest pipeline.
|
||||
|
||||
## Payload Fields by Source
|
||||
|
||||
### Docker Worker (`docker/worker/tasks/file_ingestion.py`)
|
||||
|
||||
Stores these payload fields (line 286-311):
|
||||
|
||||
| Field | Example | Notes |
|
||||
|---|---|---|
|
||||
| `text` | `"Скрипт настройки..."` | Chunk content |
|
||||
| `source` | `"wiki-homelab"` | Derived from `get_source_tag()` — path relative to WIKI_PATH |
|
||||
| `file_path` | `"/wiki/homelab/MikroTik RouterBOARD.md"` | Full path inside the container |
|
||||
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
|
||||
| `tags` | `["networking", "router"]` | Optional |
|
||||
| `chunk_index` | `0` | Zero-based chunk number |
|
||||
| `chunk_total` | `5` | Total chunks for this file |
|
||||
|
||||
**NO `file` field.** NO `filename` field. NO `path` field.
|
||||
|
||||
### Bulk Ingest (`scripts/bulk_wiki_ingest_ollama.py`)
|
||||
|
||||
Stores different payload fields (line 299-304):
|
||||
|
||||
| Field | Example | Notes |
|
||||
|---|---|---|
|
||||
| `filename` | `"MikroTik RouterBOARD RBD520-5HacD2HnD.md"` | `filepath.name` |
|
||||
| `path` | `"/opt/hermes/memory-os/wiki-raw/homelab/MikroTik.md"` | `str(filepath)` |
|
||||
| `content` | `"Скрипт настройки..."` | Truncated to 1000 chars preview |
|
||||
| `length` | `15832` | Full chunk length |
|
||||
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
|
||||
| `chunk_index` | `0` | Zero-based chunk number |
|
||||
| `chunk_total` | `5` | Total chunks for this file |
|
||||
|
||||
**NO `file` field.** NO `source` field. NO `file_path` field.
|
||||
|
||||
## The Fix
|
||||
|
||||
The search-api `main.py` results loop was changed from:
|
||||
|
||||
```python
|
||||
file=payload.get("file"), # always None
|
||||
source=payload.get("source", "unknown"), # always "unknown" for bulk ingest
|
||||
```
|
||||
|
||||
To:
|
||||
|
||||
```python
|
||||
file_path = payload.get("file_path") or payload.get("path") or ""
|
||||
filename = os.path.basename(file_path) if file_path else payload.get("filename")
|
||||
|
||||
source = payload.get("source")
|
||||
if not source or source == "unknown":
|
||||
if file_path:
|
||||
parts = file_path.split("/")
|
||||
wiki_idx = next((i for i, p in enumerate(parts) if p.startswith("wiki-")), -1)
|
||||
if wiki_idx >= 0:
|
||||
source = parts[wiki_idx]
|
||||
else:
|
||||
source = "wiki"
|
||||
else:
|
||||
source = "unknown"
|
||||
```
|
||||
|
||||
## Future-Proofing
|
||||
|
||||
If a new ingest path is added (e.g. a CLI tool or API), it should store EITHER:
|
||||
- `file_path` (full path) — the search-api extracts basename
|
||||
- `filename` + `path` (separate fields) — the search-api prefers `filename` if no path available
|
||||
|
||||
The search-api is now tolerant of both schemes. If a new field is introduced, add it to the
|
||||
`file_path or payload.get("path")` chain in `_extract_filename()`. But the cleaner approach
|
||||
is to standardise all ingest paths on `file_path`.
|
||||
|
||||
## Verification
|
||||
|
||||
```bash
|
||||
curl -s -X POST http://localhost:8000/search \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"query":"настройка VPN", "top_k": 3}' | python3 -m json.tool
|
||||
```
|
||||
|
||||
Expected: `"file": "MikroTik RouterBOARD RBD520-5HacD2HnD.md"` (not null),
|
||||
`"source": "wiki"` (not "unknown").
|
||||
@@ -0,0 +1,90 @@
|
||||
# Icarus Threshold Bug — 2026-07-16 (v2)
|
||||
|
||||
## The Problem
|
||||
|
||||
`icarus/hooks.py` calls `_search_qdrant(query, top_k=2, threshold=0.55)` from `pre_llm_call()`.
|
||||
This passes `score_threshold=0.55` to `search_with_fallback()` in `context_enhancer.py`.
|
||||
|
||||
RRF fusion (hybrid dense+sparse) returns scores 0.33–0.50 even for good matches.
|
||||
At 0.55, every query is filtered out. The search cascade falls to `level="none"`,
|
||||
then hits SQLite fallback (`[CE-FALLBACK] SQLite search failed: no such table: lineage`),
|
||||
and finally returns an empty list. Icarus injects nothing.
|
||||
|
||||
This is fundamentally different from dense-only search, which returns COSINE scores
|
||||
0.90+ for the same queries. The threshold bug was invisible because:
|
||||
- `_search_qdrant()` is fail-open (returns `[]` on any error)
|
||||
- The plugin loads and runs without crashing
|
||||
- "Runs without crashing" ≠ "returns useful results"
|
||||
|
||||
## Root Cause: RRF vs Dense Score Regimes
|
||||
|
||||
RRF (Reciprocal Rank Fusion) normalises scores from two independent retrievers
|
||||
(dense and sparse) into a shared 0–1 range via `1/(k + rank)`. This inherently
|
||||
produces clustered scores around 0.33–0.50 regardless of the underlying semantic
|
||||
similarity. This is **not a bug in RRF** — it's how RRF works.
|
||||
|
||||
Dense-only search returns raw COSINE similarity (0.90+ for good matches).
|
||||
|
||||
The same query at different thresholds:
|
||||
|
||||
```
|
||||
Query: "XRay VPS VPN"
|
||||
|
||||
=== Hybrid (RRF), threshold 0.55 ===
|
||||
Level: none, Results: 0
|
||||
Falls through to SQLite → [CE-FALLBACK] All fallback levels exhausted.
|
||||
|
||||
=== Dense-only, threshold 0.35 ===
|
||||
Level: dense-only, Results: 3
|
||||
[0.9938] Пароль Юлии Зозули
|
||||
[0.9078] Nagios
|
||||
[0.9017] Клавиатуры
|
||||
|
||||
=== Hybrid (RRF), threshold 0.35 ===
|
||||
Level: hybrid, Results: 2
|
||||
[0.5000] Инструкция как безопасно расширить кластер Garage до 3+ нод (v2.1)
|
||||
[0.5000] Пароль Юлии Зозули
|
||||
|
||||
=== Hybrid (RRF), threshold 0.30 ===
|
||||
Level: hybrid, Results: 4
|
||||
[0.5000] Garage кластер
|
||||
[0.5000] Пароль Юлии Зозули
|
||||
[0.3333] Двухфакторная VPN
|
||||
[0.3333] Nagios
|
||||
```
|
||||
|
||||
At 0.35, RRF still filters the 0.33 results (2 out of 4 are lost).
|
||||
At 0.30, all 4 are returned.
|
||||
|
||||
## The Fix
|
||||
|
||||
Lower `threshold` in `_search_qdrant()` in `/opt/hermes/.hermes/plugins/icarus/hooks.py`
|
||||
from 0.55 to 0.30.
|
||||
|
||||
```python
|
||||
# Line 722 — NOW:
|
||||
qdrant_results = _search_qdrant(user_message, top_k=2, threshold=0.30)
|
||||
```
|
||||
|
||||
The comment above it was also updated to explain the RRF vs dense score gap.
|
||||
|
||||
## Verification
|
||||
|
||||
After fix, all queries return hybrid results within 6ms Qdrant time:
|
||||
|
||||
```
|
||||
XRay VPS VPN → hybrid, 4 results
|
||||
hermes plugin icarus → hybrid, 3 results
|
||||
telegram → hybrid, 4 results
|
||||
wireguard → hybrid, 4 results
|
||||
```
|
||||
|
||||
No SQLite fallback noise. No `level="none"`.
|
||||
|
||||
## Key Insight
|
||||
|
||||
When debugging an empty `[qdrant]` block in Icarus context injection:
|
||||
1. First check if `_search_qdrant()` even runs (no import error → PYTHONPATH fix)
|
||||
2. Then check if it returns results (threshold too high for RRF)
|
||||
3. These are SEPARATE bugs — the first was fixed 2026-07-14 (PYTHONPATH),
|
||||
the second on 2026-07-16 (threshold 0.55→0.30)
|
||||
@@ -0,0 +1,32 @@
|
||||
# WebDAV Migration & Sync Pipeline Fix — 2026-07-16
|
||||
|
||||
## Change: s3fs → WebDAV as Obsidian Vault Source
|
||||
|
||||
**Problem:** The s3fs mount at `/opt/hermes/obsidian-vault/` (via Garage S3 `s3.nixg.ru`) was empty — only `.` and `..` in the directory, no actual files. sync_obsidian_to_wiki.py couldn't copy anything. C1 was blocked.
|
||||
|
||||
**Diagnosis:**
|
||||
- `/etc/fstab` had s3fs entry pointing at `s3.nixg.ru` (Garage)
|
||||
- `mount | grep s3fs` showed the mount was active but directory contained 0 files
|
||||
- `/mnt/yandex-disk/` (Yandex Disk via davfs2 WebDAV) had a working `obsidian/mozg/` vault with 65 .md files
|
||||
|
||||
**Fix:** Changed `OBSIDIAN_VAULT` from `/opt/hermes/obsidian-vault` to `/mnt/yandex-disk/obsidian/mozg/`.
|
||||
|
||||
**First run:** 74 files copied in 44 seconds, 0 errors.
|
||||
|
||||
## Blocker 2: sync → ingest chain broken
|
||||
|
||||
**Symptom:** The sync script finishes copying, then tries to call `wiki_continuous_ingest.py` with bare `python3` — crashes with `ModuleNotFoundError: No module named 'arq'`.
|
||||
|
||||
**Root cause:** `sync_obsidian_to_wiki.py` line 162 does `os.system(f"python3 {ingest_script}")`. System python3 does NOT have `arq` (only installed in project `.venv/`).
|
||||
|
||||
**Fix:** Changed to `/opt/hermes/memory-os/.venv/bin/python3 {ingest_script}`.
|
||||
|
||||
## Safety Guarantee
|
||||
|
||||
The user explicitly asked for protection against Hermes deleting files in their Obsidian vault. Confirmed:
|
||||
|
||||
- **Code-level:** sync_obsidian_to_wiki.py only reads source (via `os.walk()` + `copy2()`). All deletions (`unlink()`) target `WIKI_TARGET`, never `OBSIDIAN_VAULT`. This was true before the change and remains true after.
|
||||
- **Filesystem-level:** davfs2 mounts with `file_mode=600` — the user process itself can only write with explicit sudo. The sync script runs as user `estorozhenko`.
|
||||
- **Memory-level:** Added read-only guarantee to user profile.
|
||||
|
||||
**Lesson:** When the user expresses concern about data safety, (1) show code evidence that it's already safe, (2) add explicit protection in memory/skill, (3) explain trade-offs (read-only mount vs full POSIX) so they can make informed decisions.
|
||||
Reference in New Issue
Block a user