3.2 KiB
Chunking Implementation — 2026-07-16
Problem
bge-m3 via Ollama /api/embeddings rejects inputs >~6000 chars with HTTP 500.
The old bulk_wiki_ingest_ollama.py truncated files to 6000 chars, losing
mid/end content from large documents.
Affected files (before chunking):
- Configuration backup.md (240KB) — only first 6K of 240KB indexed
- Конфигурация компакт Микротика.md (19KB) — Ollama 500 on chunk 2 (16083 chars, no
\n\n) - Any file with base64 images, ANSI escapes, or control characters caused 500s
Solution: chunk_text() in bulk_wiki_ingest_ollama.py
Three-level recursive split:
1. Headings: split by ##, then ###, then ####
2. Paragraphs: split by \n\n (within blocks > MAX_CHUNK_SIZE)
3. Words: split by space (last resort, when a single paragraph exceeds MAX_CHUNK_SIZE)
Config
MAX_CHUNK_SIZE = 5000 # safe buffer below bge-m3 ~6000 limit
CHUNK_OVERLAP = 300 # overlap between consecutive chunks
sanitize_text() — Pre-embedding Cleanup
Removes content that causes Ollama 500:
- Base64 images (
data:image/...) →[IMAGE]placeholder - Control characters (except
\n,\r,\t) → stripped - ANSI escape sequences (
\x1b[...) → stripped - Excessive whitespace → collapsed
Point ID Scheme
unique_id = hashlib.md5(f"{filepath}:chunk:{chunk_index}".encode()).hexdigest()
point_id = int(unique_id[:8], 16) & 0x7fffffff
Payload Fields
{
"filename": filepath.name,
"path": str(filepath),
"content": clean_text[:1000], # preview
"length": len(clean_text),
"title": title,
"chunk_index": chunk_index,
"chunk_total": len(chunks),
"indexed_at": datetime.now().isoformat(),
}
Results
| Metric | Before | After |
|---|---|---|
| Total points | 149 | 159 |
| Files indexed | 75/85 | 78/85 |
| Ollama 500 errors | 2 | 0 |
| Configuration backup.md | truncated (1 chunk) | 53 chunks |
| Конфигурация компакт Микротика.md | 500 error | 6 chunks |
| Цифровой паспорт.md | 2 chunks | 3 chunks |
Verification
Three searches confirmed deep content retrieval:
# 1. "заселение по биометрии" → Цифровой паспорт (score 0.61)
# 2. "TemperatureLimit DCMIConfiguration thermal" → Configuration backup (score 0.73)
# 3. "MikroTik компактный экспорт" → Конфигурация Микротика (score 0.62)
Pitfalls Encountered
-
Ollama 500 on large paragraphs: Some files (like конфигурация Микротика) have code blocks with no
\n\nfor hundreds of lines. The paragraph split returned one element > MAX_CHUNK_SIZE. Fix: added word-level split as last resort. -
Base64 images in markdown: Files like Цифровой паспорт.md contain forwarded emails with embedded base64 images. Ollama 500 on the raw data. Fix:
sanitize_text()regex replacesdata:image/...;base64,...with[base64-data]. -
ANSI escape sequences: Config files with terminal control codes. Fix: strip
\x1b\[[0-9;]*[a-zA-Z]patterns. -
Empty files produce Ollama error: 7 files in vault are 0 bytes. Skip with
if not content.strip(): return 0.