Files
memory-os/references/search-api-payload-fields.md
T
2026-09-06 13:51:22 +00:00

90 lines
3.2 KiB
Markdown

# Search API — Payload Field Mapping (2026-07-16)
## Problem
`POST /search` returned `"file": null` and `"source": "unknown"` for all results.
The search-api looked for `payload.file` and `payload.source`, but neither field
exists in the Qdrant payload as stored by the ingest pipeline.
## Payload Fields by Source
### Docker Worker (`docker/worker/tasks/file_ingestion.py`)
Stores these payload fields (line 286-311):
| Field | Example | Notes |
|---|---|---|
| `text` | `"Скрипт настройки..."` | Chunk content |
| `source` | `"wiki-homelab"` | Derived from `get_source_tag()` — path relative to WIKI_PATH |
| `file_path` | `"/wiki/homelab/MikroTik RouterBOARD.md"` | Full path inside the container |
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
| `tags` | `["networking", "router"]` | Optional |
| `chunk_index` | `0` | Zero-based chunk number |
| `chunk_total` | `5` | Total chunks for this file |
**NO `file` field.** NO `filename` field. NO `path` field.
### Bulk Ingest (`scripts/bulk_wiki_ingest_ollama.py`)
Stores different payload fields (line 299-304):
| Field | Example | Notes |
|---|---|---|
| `filename` | `"MikroTik RouterBOARD RBD520-5HacD2HnD.md"` | `filepath.name` |
| `path` | `"/opt/hermes/memory-os/wiki-raw/homelab/MikroTik.md"` | `str(filepath)` |
| `content` | `"Скрипт настройки..."` | Truncated to 1000 chars preview |
| `length` | `15832` | Full chunk length |
| `title` | `"MikroTik RouterBOARD RBD520"` | From frontmatter or filename stem |
| `chunk_index` | `0` | Zero-based chunk number |
| `chunk_total` | `5` | Total chunks for this file |
**NO `file` field.** NO `source` field. NO `file_path` field.
## The Fix
The search-api `main.py` results loop was changed from:
```python
file=payload.get("file"), # always None
source=payload.get("source", "unknown"), # always "unknown" for bulk ingest
```
To:
```python
file_path = payload.get("file_path") or payload.get("path") or ""
filename = os.path.basename(file_path) if file_path else payload.get("filename")
source = payload.get("source")
if not source or source == "unknown":
if file_path:
parts = file_path.split("/")
wiki_idx = next((i for i, p in enumerate(parts) if p.startswith("wiki-")), -1)
if wiki_idx >= 0:
source = parts[wiki_idx]
else:
source = "wiki"
else:
source = "unknown"
```
## Future-Proofing
If a new ingest path is added (e.g. a CLI tool or API), it should store EITHER:
- `file_path` (full path) — the search-api extracts basename
- `filename` + `path` (separate fields) — the search-api prefers `filename` if no path available
The search-api is now tolerant of both schemes. If a new field is introduced, add it to the
`file_path or payload.get("path")` chain in `_extract_filename()`. But the cleaner approach
is to standardise all ingest paths on `file_path`.
## Verification
```bash
curl -s -X POST http://localhost:8000/search \
-H "Content-Type: application/json" \
-d '{"query":"настройка VPN", "top_k": 3}' | python3 -m json.tool
```
Expected: `"file": "MikroTik RouterBOARD RBD520-5HacD2HnD.md"` (not null),
`"source": "wiki"` (not "unknown").