# Chunking Implementation — 2026-07-16 ## Problem bge-m3 via Ollama `/api/embeddings` rejects inputs >~6000 chars with HTTP 500. The old `bulk_wiki_ingest_ollama.py` truncated files to 6000 chars, losing mid/end content from large documents. Affected files (before chunking): - Configuration backup.md (240KB) — only first 6K of 240KB indexed - Конфигурация компакт Микротика.md (19KB) — Ollama 500 on chunk 2 (16083 chars, no `\n\n`) - Any file with base64 images, ANSI escapes, or control characters caused 500s ## Solution: `chunk_text()` in `bulk_wiki_ingest_ollama.py` Three-level recursive split: ``` 1. Headings: split by ##, then ###, then #### 2. Paragraphs: split by \n\n (within blocks > MAX_CHUNK_SIZE) 3. Words: split by space (last resort, when a single paragraph exceeds MAX_CHUNK_SIZE) ``` ### Config ```python MAX_CHUNK_SIZE = 5000 # safe buffer below bge-m3 ~6000 limit CHUNK_OVERLAP = 300 # overlap between consecutive chunks ``` ### `sanitize_text()` — Pre-embedding Cleanup Removes content that causes Ollama 500: - Base64 images (`data:image/...`) → `[IMAGE]` placeholder - Control characters (except `\n`, `\r`, `\t`) → stripped - ANSI escape sequences (`\x1b[...`) → stripped - Excessive whitespace → collapsed ### Point ID Scheme ```python unique_id = hashlib.md5(f"{filepath}:chunk:{chunk_index}".encode()).hexdigest() point_id = int(unique_id[:8], 16) & 0x7fffffff ``` ### Payload Fields ```python { "filename": filepath.name, "path": str(filepath), "content": clean_text[:1000], # preview "length": len(clean_text), "title": title, "chunk_index": chunk_index, "chunk_total": len(chunks), "indexed_at": datetime.now().isoformat(), } ``` ## Results | Metric | Before | After | |---|---|---| | Total points | 149 | 159 | | Files indexed | 75/85 | 78/85 | | Ollama 500 errors | 2 | 0 | | Configuration backup.md | truncated (1 chunk) | 53 chunks | | Конфигурация компакт Микротика.md | 500 error | 6 chunks | | Цифровой паспорт.md | 2 chunks | 3 chunks | ## Verification Three searches confirmed deep content retrieval: ```bash # 1. "заселение по биометрии" → Цифровой паспорт (score 0.61) # 2. "TemperatureLimit DCMIConfiguration thermal" → Configuration backup (score 0.73) # 3. "MikroTik компактный экспорт" → Конфигурация Микротика (score 0.62) ``` ## Pitfalls Encountered 1. **Ollama 500 on large paragraphs:** Some files (like конфигурация Микротика) have code blocks with no `\n\n` for hundreds of lines. The paragraph split returned one element > MAX_CHUNK_SIZE. Fix: added word-level split as last resort. 2. **Base64 images in markdown:** Files like Цифровой паспорт.md contain forwarded emails with embedded base64 images. Ollama 500 on the raw data. Fix: `sanitize_text()` regex replaces `data:image/...;base64,...` with `[base64-data]`. 3. **ANSI escape sequences:** Config files with terminal control codes. Fix: strip `\x1b\[[0-9;]*[a-zA-Z]` patterns. 4. **Empty files produce Ollama error:** 7 files in vault are 0 bytes. Skip with `if not content.strip(): return 0`.