mirror of
https://gitverse.ru/kpa39l/memory-os.git
synced 2026-09-29 09:35:05 +00:00
90 lines
3.2 KiB
Markdown
90 lines
3.2 KiB
Markdown
# Search API — Payload Field Mapping (2026-07-16)
|
|
|
|
## Problem
|
|
|
|
`POST /search` returned `"file": null` and `"source": "unknown"` for all results.
|
|
The search-api looked for `payload.file` and `payload.source`, but neither field
|
|
exists in the Qdrant payload as stored by the ingest pipeline.
|
|
|
|
## Payload Fields by Source
|
|
|
|
### Docker Worker (`docker/worker/tasks/file_ingestion.py`)
|
|
|
|
Stores these payload fields (line 286-311):
|
|
|
|
| Field | Example | Notes |
|
|
|---|---|---|
|
|
| `text` | `"Скрипт настройки..."` | Chunk content |
|
|
| `source` | `"wiki-homelab"` | Derived from `get_source_tag()` — path relative to WIKI_PATH |
|
|
| `file_path` | `"/wiki/homelab/MikroTik RouterBOARD.md"` | Full path inside the container |
|
|
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
|
|
| `tags` | `["networking", "router"]` | Optional |
|
|
| `chunk_index` | `0` | Zero-based chunk number |
|
|
| `chunk_total` | `5` | Total chunks for this file |
|
|
|
|
**NO `file` field.** NO `filename` field. NO `path` field.
|
|
|
|
### Bulk Ingest (`scripts/bulk_wiki_ingest_ollama.py`)
|
|
|
|
Stores different payload fields (line 299-304):
|
|
|
|
| Field | Example | Notes |
|
|
|---|---|---|
|
|
| `filename` | `"MikroTik RouterBOARD RBD520-5HacD2HnD.md"` | `filepath.name` |
|
|
| `path` | `"/opt/hermes/memory-os/wiki-raw/homelab/MikroTik.md"` | `str(filepath)` |
|
|
| `content` | `"Скрипт настройки..."` | Truncated to 1000 chars preview |
|
|
| `length` | `15832` | Full chunk length |
|
|
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
|
|
| `chunk_index` | `0` | Zero-based chunk number |
|
|
| `chunk_total` | `5` | Total chunks for this file |
|
|
|
|
**NO `file` field.** NO `source` field. NO `file_path` field.
|
|
|
|
## The Fix
|
|
|
|
The search-api `main.py` results loop was changed from:
|
|
|
|
```python
|
|
file=payload.get("file"), # always None
|
|
source=payload.get("source", "unknown"), # always "unknown" for bulk ingest
|
|
```
|
|
|
|
To:
|
|
|
|
```python
|
|
file_path = payload.get("file_path") or payload.get("path") or ""
|
|
filename = os.path.basename(file_path) if file_path else payload.get("filename")
|
|
|
|
source = payload.get("source")
|
|
if not source or source == "unknown":
|
|
if file_path:
|
|
parts = file_path.split("/")
|
|
wiki_idx = next((i for i, p in enumerate(parts) if p.startswith("wiki-")), -1)
|
|
if wiki_idx >= 0:
|
|
source = parts[wiki_idx]
|
|
else:
|
|
source = "wiki"
|
|
else:
|
|
source = "unknown"
|
|
```
|
|
|
|
## Future-Proofing
|
|
|
|
If a new ingest path is added (e.g. a CLI tool or API), it should store EITHER:
|
|
- `file_path` (full path) — the search-api extracts basename
|
|
- `filename` + `path` (separate fields) — the search-api prefers `filename` if no path available
|
|
|
|
The search-api is now tolerant of both schemes. If a new field is introduced, add it to the
|
|
`file_path or payload.get("path")` chain in `_extract_filename()`. But the cleaner approach
|
|
is to standardise all ingest paths on `file_path`.
|
|
|
|
## Verification
|
|
|
|
```bash
|
|
curl -s -X POST http://localhost:8000/search \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query":"настройка VPN", "top_k": 3}' | python3 -m json.tool
|
|
```
|
|
|
|
Expected: `"file": "MikroTik RouterBOARD RBD520-5HacD2HnD.md"` (not null),
|
|
`"source": "wiki"` (not "unknown"). |