Files
memory-os/references/chunking-implementation.md
2026-09-06 13:51:22 +00:00

97 lines
3.2 KiB
Markdown

# Chunking Implementation — 2026-07-16
## Problem
bge-m3 via Ollama `/api/embeddings` rejects inputs >~6000 chars with HTTP 500.
The old `bulk_wiki_ingest_ollama.py` truncated files to 6000 chars, losing
mid/end content from large documents.
Affected files (before chunking):
- Configuration backup.md (240KB) — only first 6K of 240KB indexed
- Конфигурация компакт Микротика.md (19KB) — Ollama 500 on chunk 2 (16083 chars, no `\n\n`)
- Any file with base64 images, ANSI escapes, or control characters caused 500s
## Solution: `chunk_text()` in `bulk_wiki_ingest_ollama.py`
Three-level recursive split:
```
1. Headings: split by ##, then ###, then ####
2. Paragraphs: split by \n\n (within blocks > MAX_CHUNK_SIZE)
3. Words: split by space (last resort, when a single paragraph exceeds MAX_CHUNK_SIZE)
```
### Config
```python
MAX_CHUNK_SIZE = 5000 # safe buffer below bge-m3 ~6000 limit
CHUNK_OVERLAP = 300 # overlap between consecutive chunks
```
### `sanitize_text()` — Pre-embedding Cleanup
Removes content that causes Ollama 500:
- Base64 images (`data:image/...`) → `[IMAGE]` placeholder
- Control characters (except `\n`, `\r`, `\t`) → stripped
- ANSI escape sequences (`\x1b[...`) → stripped
- Excessive whitespace → collapsed
### Point ID Scheme
```python
unique_id = hashlib.md5(f"{filepath}:chunk:{chunk_index}".encode()).hexdigest()
point_id = int(unique_id[:8], 16) & 0x7fffffff
```
### Payload Fields
```python
{
"filename": filepath.name,
"path": str(filepath),
"content": clean_text[:1000], # preview
"length": len(clean_text),
"title": title,
"chunk_index": chunk_index,
"chunk_total": len(chunks),
"indexed_at": datetime.now().isoformat(),
}
```
## Results
| Metric | Before | After |
|---|---|---|
| Total points | 149 | 159 |
| Files indexed | 75/85 | 78/85 |
| Ollama 500 errors | 2 | 0 |
| Configuration backup.md | truncated (1 chunk) | 53 chunks |
| Конфигурация компакт Микротика.md | 500 error | 6 chunks |
| Цифровой паспорт.md | 2 chunks | 3 chunks |
## Verification
Three searches confirmed deep content retrieval:
```bash
# 1. "заселение по биометрии" → Цифровой паспорт (score 0.61)
# 2. "TemperatureLimit DCMIConfiguration thermal" → Configuration backup (score 0.73)
# 3. "MikroTik компактный экспорт" → Конфигурация Микротика (score 0.62)
```
## Pitfalls Encountered
1. **Ollama 500 on large paragraphs:** Some files (like конфигурация Микротика) have
code blocks with no `\n\n` for hundreds of lines. The paragraph split returned
one element > MAX_CHUNK_SIZE. Fix: added word-level split as last resort.
2. **Base64 images in markdown:** Files like Цифровой паспорт.md contain forwarded
emails with embedded base64 images. Ollama 500 on the raw data. Fix: `sanitize_text()`
regex replaces `data:image/...;base64,...` with `[base64-data]`.
3. **ANSI escape sequences:** Config files with terminal control codes. Fix: strip
`\x1b\[[0-9;]*[a-zA-Z]` patterns.
4. **Empty files produce Ollama error:** 7 files in vault are 0 bytes. Skip with
`if not content.strip(): return 0`.