Initial commit: Hermes skill memory-os

This commit is contained in:
estorozhenko
2026-09-06 13:51:22 +00:00
commit a775acbe59
7 changed files with 930 additions and 0 deletions
+97
View File
@@ -0,0 +1,97 @@
# Chunking Implementation — 2026-07-16
## Problem
bge-m3 via Ollama `/api/embeddings` rejects inputs >~6000 chars with HTTP 500.
The old `bulk_wiki_ingest_ollama.py` truncated files to 6000 chars, losing
mid/end content from large documents.
Affected files (before chunking):
- Configuration backup.md (240KB) — only first 6K of 240KB indexed
- Конфигурация компакт Микротика.md (19KB) — Ollama 500 on chunk 2 (16083 chars, no `\n\n`)
- Any file with base64 images, ANSI escapes, or control characters caused 500s
## Solution: `chunk_text()` in `bulk_wiki_ingest_ollama.py`
Three-level recursive split:
```
1. Headings: split by ##, then ###, then ####
2. Paragraphs: split by \n\n (within blocks > MAX_CHUNK_SIZE)
3. Words: split by space (last resort, when a single paragraph exceeds MAX_CHUNK_SIZE)
```
### Config
```python
MAX_CHUNK_SIZE = 5000 # safe buffer below bge-m3 ~6000 limit
CHUNK_OVERLAP = 300 # overlap between consecutive chunks
```
### `sanitize_text()` — Pre-embedding Cleanup
Removes content that causes Ollama 500:
- Base64 images (`data:image/...`) → `[IMAGE]` placeholder
- Control characters (except `\n`, `\r`, `\t`) → stripped
- ANSI escape sequences (`\x1b[...`) → stripped
- Excessive whitespace → collapsed
### Point ID Scheme
```python
unique_id = hashlib.md5(f"{filepath}:chunk:{chunk_index}".encode()).hexdigest()
point_id = int(unique_id[:8], 16) & 0x7fffffff
```
### Payload Fields
```python
{
"filename": filepath.name,
"path": str(filepath),
"content": clean_text[:1000], # preview
"length": len(clean_text),
"title": title,
"chunk_index": chunk_index,
"chunk_total": len(chunks),
"indexed_at": datetime.now().isoformat(),
}
```
## Results
| Metric | Before | After |
|---|---|---|
| Total points | 149 | 159 |
| Files indexed | 75/85 | 78/85 |
| Ollama 500 errors | 2 | 0 |
| Configuration backup.md | truncated (1 chunk) | 53 chunks |
| Конфигурация компакт Микротика.md | 500 error | 6 chunks |
| Цифровой паспорт.md | 2 chunks | 3 chunks |
## Verification
Three searches confirmed deep content retrieval:
```bash
# 1. "заселение по биометрии" → Цифровой паспорт (score 0.61)
# 2. "TemperatureLimit DCMIConfiguration thermal" → Configuration backup (score 0.73)
# 3. "MikroTik компактный экспорт" → Конфигурация Микротика (score 0.62)
```
## Pitfalls Encountered
1. **Ollama 500 on large paragraphs:** Some files (like конфигурация Микротика) have
code blocks with no `\n\n` for hundreds of lines. The paragraph split returned
one element > MAX_CHUNK_SIZE. Fix: added word-level split as last resort.
2. **Base64 images in markdown:** Files like Цифровой паспорт.md contain forwarded
emails with embedded base64 images. Ollama 500 on the raw data. Fix: `sanitize_text()`
regex replaces `data:image/...;base64,...` with `[base64-data]`.
3. **ANSI escape sequences:** Config files with terminal control codes. Fix: strip
`\x1b\[[0-9;]*[a-zA-Z]` patterns.
4. **Empty files produce Ollama error:** 7 files in vault are 0 bytes. Skip with
`if not content.strip(): return 0`.
@@ -0,0 +1,61 @@
# Dimension Mismatch & s3fs Blockers — 2026-07-16
## Issue 1: Qdrant 400 Bad Request (Dimension Mismatch)
**Symptom:**
```
[CE-FALLBACK] Qdrant general error (400 Client Error: Bad Request
for url: http://localhost:6333/collections/knowledge_base/points/query),
falling back to lexical.
[CE-FALLBACK] All fallback levels exhausted.
```
**Root cause:**
The Qdrant collection `knowledge_base` was created with 768d vectors (nomic-embed-text via Ollama), but `.env` was later changed to point at Polza.ai with `qwen3-embedding-8b` (4096d). The `context_enhancer.py` reads `.env` and sends 4096d query vectors to a 768d collection — Qdrant rejects with 400.
**Cross-check:**
```
# Collection says: 768d
# .env says: 4096d
```
**Resolution options:**
- (a) Change `.env` back to Ollama nomic-embed-text 768d (local, free, matches collection)
- (b) Recreate the collection at 4096d (destructive — lose all 683 points)
- (c) Create a second collection for 4096d, keep the old one
**Note:** The `wiki_continuous_ingest.py` worker reads from the same `.env` — points may also be upserted with wrong dimensions. Also the `docker/.env` file may differ from the project root `.env`.
## Issue 2: s3fs Obsidian Vault Mount Empty
**Symptom:**
```
ls -la /opt/hermes/obsidian-vault/
total 8
drwxr-xr-x 2 estorozhenko estorozhenko 4096 ...
drwxr-xr-x 6 estorozhenko estorozhenko 4096 ...
```
Only `.` and `..` — the directory exists but is empty.
**fstab entry:**
```
s3fs#obsidian /opt/hermes/obsidian-vault fuse _netdev,allow_other,
passwd_file=/etc/passwd-s3fs,use_cache=/tmp,
url=https://s3.nixg.ru,endpoint=garage,
use_path_request_style,sigv4 0 0
```
This is a Garage S3 bucket mounted via s3fs. The mount point exists but no data is visible.
**Impact:**
- `sync_obsidian_to_wiki.py` cannot sync (no source files)
- The script has a safety guard: if vault is empty but state has files, it skips deletion (`VAULT ПУСТ — пропускаю удаление`). This prevents data loss if the mount is temporarily disconnected.
- C1 (schedule sync) is blocked until mount is restored
**Diagnosis steps:**
1. `mount | grep s3fs` — check if mount is active
2. `systemctl status` for s3fs — check if automount service is running
3. `sudo journalctl -u` for s3fs — check for errors
4. Direct HTTP check against `s3.nixg.ru` — test S3 endpoint reachability
5. Verify `passwd-s3fs` file exists and has correct credentials
+241
View File
@@ -0,0 +1,241 @@
# Ingest Debugging Sessions
Chronological reproduction of the pipeline from zero to working search. Each session
covers issues encountered and fixes applied.
---
## Session 1: 2026-06-20 to 2026-06-21 (Initial Pipeline Bring-Up)
Full reproduction of the pipeline from zero to working search.
## Initial State
- Qdrant collection `knowledge_base` exists with dense (768d) + sparse vectors
- `wiki_ingest_state.json` has 467 files, all with `ingested_at` set
- Qdrant points_count = 0 — nothing actually stored
- ARQ queue empty — stale jobs consumed
- Worker container healthy but was crashing on DNS
## Issue 1: Worker DNS Failure
**Symptom:** Worker logs show `Temporary failure in name resolution` when accessing `ollama:11434`
**Root cause:** Worker container not connected to `ollama_default` Docker network.
`docker compose up -d` recreating the container detaches secondary networks.
**Fix:**
```bash
docker network connect ollama_default docker-worker-1
```
Verify:
```bash
docker inspect docker-worker-1 | jq '.[].NetworkSettings.Networks | keys'
```
Should show both `docker_default` and `ollama_default`.
## Issue 2: Sparse Vector Name Mismatch
**Symptom:** Worker completes but Qdrant points_count stays at 0. Logs show HTTP 400 errors.
**Root cause:** The Qdrant collection was created with sparse vectors named `"bm25"`,
but the ingest worker sends sparse vectors keyed as `"sparse"`:
```python
# What the collection had:
sparse_vectors_config={'bm25': SparseVectorParams(...)}
# What the worker sends:
{'points': [..., 'vector': {'sparse': ..., 'dense': ...}]}
```
**Fix:** Recreate collection with consistent naming:
```python
c.delete_collection('knowledge_base')
c.create_collection(
collection_name='knowledge_base',
vectors_config=VectorParams(size=768, distance=Distance.COSINE),
sparse_vectors_config={'sparse': SparseVectorParams(
index=SparseIndexParams(on_disk=False, full_scan_threshold=10000)
)}
)
```
## Issue 3: Dense Embedding Model Mismatch (Search)
**Symptom:** After ingest works and Qdrant has 683 points, `context_enhancer.py` search
returns only low-score or no results, falling back to lexical search.
**Root cause:** `context_enhancer.py` was configured to use OpenRouter `qwen/qwen3-embedding-8b`
for query embedding, but the ingested collection uses Ollama `nomic-embed-text:latest` (768d).
Different models produce incompatible vector spaces — cosine similarity is near zero.
**Fix in context_enhancer.py:**
```python
# Before (OpenRouter):
EMBEDDING_MODEL = "qwen/qwen3-embedding-8b"
resp = requests.post("https://openrouter.ai/api/v1/embeddings", ...)
# After (local Ollama):
OLLAMA_EMBEDDING_URL = "http://localhost:11434"
OLLAMA_EMBEDDING_MODEL = "nomic-embed-text:latest"
resp = requests.post(f"{OLLAMA_EMBEDDING_URL}/api/embeddings",
json={"model": OLLAMA_EMBEDDING_MODEL, "prompt": text})
```
Also clean up icarus/hooks.py which had OPENROUTER_API_KEY env manipulation that
became dead code.
## Issue 4: FastEmbed BM25 Subprocess Failure
**Symptom:** Sparse embedding returns None silently. `context_enhancer.py` falls back to
dense-only or lexical.
**Root cause:** The subprocess relies on `FASTEMBED_SITEPKGS` env var pointing to the
ai-lab venv site-packages, but subprocess.run() does NOT inherit the parent's env vars
automatically when the env was set in Python (not the shell).
```python
# BROKEN — subprocess doesn't see FASTEMBED_SITEPKGS:
result = subprocess.run(
[_FASTEMBED_PYTHON, "-c", "...import fastembed..."],
input=text, capture_output=True, text=True, timeout=15
)
# FIXED — explicitly pass env:
_env = os.environ.copy()
_env.setdefault("FASTEMBED_SITEPKGS", _FASTEMBED_SITEPKGS)
result = subprocess.run(
[_FASTEMBED_PYTHON, "-c", "..."],
input=text, capture_output=True, text=True, timeout=15, env=_env
)
```
## Issue 5: Score Threshold Too High
**Symptom:** nomic-embed-text results have scores in the 0.30-0.56 range, filter out
most results.
**Fix:** Lowered default threshold from 0.55 to 0.35.
## Healthy Config (Final)
```
Qdrant: localhost:6333, collection="knowledge_base", 683 points
Redis: 127.0.0.1:6379 (authenticated)
Ollama: localhost:11434, model="nomic-embed-text:latest" (768d)
Worker: docker-worker-1, connected to ollama_default network
FastEmbed: BM25 via ai-lab venv subprocess
Cron: wiki-ingest-sync, every 10 minutes
State: ~/.hermes/wiki_ingest_state.json (tracking 467 files)
```
---
## Session 2: 2026-07-15 (Three New Blockers — python-dotenv, Sparse Embed Path, Hermes Hook Integration)
Reproduced the stack from scratch. Three blocking issues found on top of the original five.
### Blocker 1: `python-dotenv` Missing in Cron Environment
**Symptom:** cron `wiki-ingest-sync` job fails silently. Logs: `ModuleNotFoundError: No module named 'dotenv'`.
**Root cause:** The cron job runs under the system environment, which does NOT have `python-dotenv` installed. The `wiki_continuous_ingest.py` script imports `from dotenv import load_dotenv` at the top.
**Fix:** Install `python-dotenv` system-wide:
```bash
pip install python-dotenv
```
Or patch the script to load `.env` via `os.environ` + `open()` instead of `python-dotenv`.
**Verification:**
```bash
python3 -c "from dotenv import load_dotenv; print('OK')"
```
### Blocker 2: Sparse Embedding Returns `Expecting value: line 1 column 1`
**Symptom:** `context_enhancer.py` sparse embedding crashes mid-search with `json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)`.
**Root cause:** The FastEmbed subprocess (`_FASTEMBED_PYTHON`) runs a Python script that tries to `import fastembed` and `import dotenv`. The subprocess's python path points to `/usr/bin/python3` which does NOT have `fastembed` or `dotenv` in its site-packages. The subprocess fails silently (prints nothing to stdout), and the parent tries `json.loads(result.stdout)` on empty output.
The existing fix from Session 1 (Issue #4 — `FASTEMBED_SITEPKGS` env var) may have been applied, but the fundamental problem is the subprocess PYTHON PATH itself, not the env var. If `/usr/bin/python3` cannot `import fastembed` at all, setting the env var won't help.
**Fix:** Use the ai-lab venv python directly:
```python
_FASTEMBED_PYTHON = "/opt/ai-lab/.venv/bin/python3"
```
instead of:
```python
_FASTEMBED_PYTHON = "/usr/bin/python3"
```
This venv has both `fastembed` and `python-dotenv` installed.
**Diagnostic:**
```bash
/opt/ai-lab/.venv/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
/usr/bin/python3 -c "from fastembed.sparse import SparseTextEmbedding; print('OK')"
```
### Blocker 3: Hermes Hook (`icarus/hooks.py`) Not Connected to Hermes
**Symptom:** The file `/opt/hermes/.hermes/plugins/icarus/hooks.py` exists with all the right logic (search Qdrant, inject context into Hermes responses), but it's NEVER called. No Hermes config, plugin, or hook registration activates it.
**Root cause:** `icarus/hooks.py` relies on being loaded by Hermes as a user plugin, but it's not enabled. The plugin dir is at `/opt/hermes/.hermes/plugins/icarus/` with its own `plugin.yaml` (v0.3.0, 16 tools, 4 hooks). Two things were needed:
1. `hermes plugins enable icarus` — the proper CLI command to activate a user plugin
2. PYTHONPATH fix inside `hooks.py` — the `_search_qdrant()` function does `from scripts.context_enhancer import ...` but `/opt/hermes/memory-os` is not on `sys.path`. Must add `import sys` at top and `sys.path.insert(0, '/opt/hermes/memory-os')` inside `_search_qdrant()` before the import. Without this, the import raises `ModuleNotFoundError` and `_search_qdrant()` returns empty list silently (fail-open).
**Fix — two steps:**
Step 1 — PYTHONPATH in hooks.py:
```python
# Add at top of hooks.py:
import sys
# Add inside _search_qdrant(), before the import:
_MEMORY_OS = "/opt/hermes/memory-os"
if _MEMORY_OS not in sys.path:
sys.path.insert(0, _MEMORY_OS)
```
Step 2 — Enable plugin:
```bash
hermes plugins enable icarus
# Takes effect on next session. No config.yaml editing needed.
```
**Verification of PYTHONPATH fix:**
```bash
python3 -c "
import sys
sys.path.insert(0, '/opt/hermes/memory-os')
from scripts.context_enhancer import embed_query, embed_query_sparse
print('context_enhancer import: OK')
dense = embed_query('test')
print(f'Dense: {len(dense)} dims')
sparse = embed_query_sparse('test')
print(f'Sparse: {len(sparse)} values')
"
```
**Verification of plugin status:**
```bash
hermes plugins list | grep icarus
# Should show: icarus │ enabled │ 0.3.0
```
## Qdrant Collection Verification
```bash
python3 -c "
from qdrant_client import QdrantClient, models
c = QdrantClient('http://localhost:6333')
info = c.get_collection('knowledge_base')
print(f'Points: {info.points_count}')
print(f'Status: {info.status}')
print(f'Dense dims: {info.config.params.vectors.size}')
print(f'Sparse keys: {list(info.config.params.sparse_vectors.keys())}')
"
```
+90
View File
@@ -0,0 +1,90 @@
# Search API — Payload Field Mapping (2026-07-16)
## Problem
`POST /search` returned `"file": null` and `"source": "unknown"` for all results.
The search-api looked for `payload.file` and `payload.source`, but neither field
exists in the Qdrant payload as stored by the ingest pipeline.
## Payload Fields by Source
### Docker Worker (`docker/worker/tasks/file_ingestion.py`)
Stores these payload fields (line 286-311):
| Field | Example | Notes |
|---|---|---|
| `text` | `"Скрипт настройки..."` | Chunk content |
| `source` | `"wiki-homelab"` | Derived from `get_source_tag()` — path relative to WIKI_PATH |
| `file_path` | `"/wiki/homelab/MikroTik RouterBOARD.md"` | Full path inside the container |
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
| `tags` | `["networking", "router"]` | Optional |
| `chunk_index` | `0` | Zero-based chunk number |
| `chunk_total` | `5` | Total chunks for this file |
**NO `file` field.** NO `filename` field. NO `path` field.
### Bulk Ingest (`scripts/bulk_wiki_ingest_ollama.py`)
Stores different payload fields (line 299-304):
| Field | Example | Notes |
|---|---|---|
| `filename` | `"MikroTik RouterBOARD RBD520-5HacD2HnD.md"` | `filepath.name` |
| `path` | `"/opt/hermes/memory-os/wiki-raw/homelab/MikroTik.md"` | `str(filepath)` |
| `content` | `"Скрипт настройки..."` | Truncated to 1000 chars preview |
| `length` | `15832` | Full chunk length |
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
| `chunk_index` | `0` | Zero-based chunk number |
| `chunk_total` | `5` | Total chunks for this file |
**NO `file` field.** NO `source` field. NO `file_path` field.
## The Fix
The search-api `main.py` results loop was changed from:
```python
file=payload.get("file"), # always None
source=payload.get("source", "unknown"), # always "unknown" for bulk ingest
```
To:
```python
file_path = payload.get("file_path") or payload.get("path") or ""
filename = os.path.basename(file_path) if file_path else payload.get("filename")
source = payload.get("source")
if not source or source == "unknown":
if file_path:
parts = file_path.split("/")
wiki_idx = next((i for i, p in enumerate(parts) if p.startswith("wiki-")), -1)
if wiki_idx >= 0:
source = parts[wiki_idx]
else:
source = "wiki"
else:
source = "unknown"
```
## Future-Proofing
If a new ingest path is added (e.g. a CLI tool or API), it should store EITHER:
- `file_path` (full path) — the search-api extracts basename
- `filename` + `path` (separate fields) — the search-api prefers `filename` if no path available
The search-api is now tolerant of both schemes. If a new field is introduced, add it to the
`file_path or payload.get("path")` chain in `_extract_filename()`. But the cleaner approach
is to standardise all ingest paths on `file_path`.
## Verification
```bash
curl -s -X POST http://localhost:8000/search \
-H "Content-Type: application/json" \
-d '{"query":"настройка VPN", "top_k": 3}' | python3 -m json.tool
```
Expected: `"file": "MikroTik RouterBOARD RBD520-5HacD2HnD.md"` (not null),
`"source": "wiki"` (not "unknown").
+90
View File
@@ -0,0 +1,90 @@
# Icarus Threshold Bug — 2026-07-16 (v2)
## The Problem
`icarus/hooks.py` calls `_search_qdrant(query, top_k=2, threshold=0.55)` from `pre_llm_call()`.
This passes `score_threshold=0.55` to `search_with_fallback()` in `context_enhancer.py`.
RRF fusion (hybrid dense+sparse) returns scores 0.33–0.50 even for good matches.
At 0.55, every query is filtered out. The search cascade falls to `level="none"`,
then hits SQLite fallback (`[CE-FALLBACK] SQLite search failed: no such table: lineage`),
and finally returns an empty list. Icarus injects nothing.
This is fundamentally different from dense-only search, which returns COSINE scores
0.90+ for the same queries. The threshold bug was invisible because:
- `_search_qdrant()` is fail-open (returns `[]` on any error)
- The plugin loads and runs without crashing
- "Runs without crashing" ≠ "returns useful results"
## Root Cause: RRF vs Dense Score Regimes
RRF (Reciprocal Rank Fusion) normalises scores from two independent retrievers
(dense and sparse) into a shared 0–1 range via `1/(k + rank)`. This inherently
produces clustered scores around 0.33–0.50 regardless of the underlying semantic
similarity. This is **not a bug in RRF** — it's how RRF works.
Dense-only search returns raw COSINE similarity (0.90+ for good matches).
The same query at different thresholds:
```
Query: "XRay VPS VPN"
=== Hybrid (RRF), threshold 0.55 ===
Level: none, Results: 0
Falls through to SQLite → [CE-FALLBACK] All fallback levels exhausted.
=== Dense-only, threshold 0.35 ===
Level: dense-only, Results: 3
[0.9938] Пароль Юлии Зозули
[0.9078] Nagios
[0.9017] Клавиатуры
=== Hybrid (RRF), threshold 0.35 ===
Level: hybrid, Results: 2
[0.5000] Инструкция как безопасно расширить кластер Garage до 3+ нод (v2.1)
[0.5000] Пароль Юлии Зозули
=== Hybrid (RRF), threshold 0.30 ===
Level: hybrid, Results: 4
[0.5000] Garage кластер
[0.5000] Пароль Юлии Зозули
[0.3333] Двухфакторная VPN
[0.3333] Nagios
```
At 0.35, RRF still filters the 0.33 results (2 out of 4 are lost).
At 0.30, all 4 are returned.
## The Fix
Lower `threshold` in `_search_qdrant()` in `/opt/hermes/.hermes/plugins/icarus/hooks.py`
from 0.55 to 0.30.
```python
# Line 722 — NOW:
qdrant_results = _search_qdrant(user_message, top_k=2, threshold=0.30)
```
The comment above it was also updated to explain the RRF vs dense score gap.
## Verification
After fix, all queries return hybrid results within 6ms Qdrant time:
```
XRay VPS VPN → hybrid, 4 results
hermes plugin icarus → hybrid, 3 results
telegram → hybrid, 4 results
wireguard → hybrid, 4 results
```
No SQLite fallback noise. No `level="none"`.
## Key Insight
When debugging an empty `[qdrant]` block in Icarus context injection:
1. First check if `_search_qdrant()` even runs (no import error → PYTHONPATH fix)
2. Then check if it returns results (threshold too high for RRF)
3. These are SEPARATE bugs — the first was fixed 2026-07-14 (PYTHONPATH),
the second on 2026-07-16 (threshold 0.55→0.30)
+32
View File
@@ -0,0 +1,32 @@
# WebDAV Migration & Sync Pipeline Fix — 2026-07-16
## Change: s3fs → WebDAV as Obsidian Vault Source
**Problem:** The s3fs mount at `/opt/hermes/obsidian-vault/` (via Garage S3 `s3.nixg.ru`) was empty — only `.` and `..` in the directory, no actual files. sync_obsidian_to_wiki.py couldn't copy anything. C1 was blocked.
**Diagnosis:**
- `/etc/fstab` had s3fs entry pointing at `s3.nixg.ru` (Garage)
- `mount | grep s3fs` showed the mount was active but directory contained 0 files
- `/mnt/yandex-disk/` (Yandex Disk via davfs2 WebDAV) had a working `obsidian/mozg/` vault with 65 .md files
**Fix:** Changed `OBSIDIAN_VAULT` from `/opt/hermes/obsidian-vault` to `/mnt/yandex-disk/obsidian/mozg/`.
**First run:** 74 files copied in 44 seconds, 0 errors.
## Blocker 2: sync → ingest chain broken
**Symptom:** The sync script finishes copying, then tries to call `wiki_continuous_ingest.py` with bare `python3` — crashes with `ModuleNotFoundError: No module named 'arq'`.
**Root cause:** `sync_obsidian_to_wiki.py` line 162 does `os.system(f"python3 {ingest_script}")`. System python3 does NOT have `arq` (only installed in project `.venv/`).
**Fix:** Changed to `/opt/hermes/memory-os/.venv/bin/python3 {ingest_script}`.
## Safety Guarantee
The user explicitly asked for protection against Hermes deleting files in their Obsidian vault. Confirmed:
- **Code-level:** sync_obsidian_to_wiki.py only reads source (via `os.walk()` + `copy2()`). All deletions (`unlink()`) target `WIKI_TARGET`, never `OBSIDIAN_VAULT`. This was true before the change and remains true after.
- **Filesystem-level:** davfs2 mounts with `file_mode=600` — the user process itself can only write with explicit sudo. The sync script runs as user `estorozhenko`.
- **Memory-level:** Added read-only guarantee to user profile.
**Lesson:** When the user expresses concern about data safety, (1) show code evidence that it's already safe, (2) add explicit protection in memory/skill, (3) explain trade-offs (read-only mount vs full POSIX) so they can make informed decisions.