Files
memory-os/references/chunking-implementation.md
2026-09-06 13:51:22 +00:00

3.2 KiB

Chunking Implementation — 2026-07-16

Problem

bge-m3 via Ollama /api/embeddings rejects inputs >~6000 chars with HTTP 500. The old bulk_wiki_ingest_ollama.py truncated files to 6000 chars, losing mid/end content from large documents.

Affected files (before chunking):

  • Configuration backup.md (240KB) — only first 6K of 240KB indexed
  • Конфигурация компакт Микротика.md (19KB) — Ollama 500 on chunk 2 (16083 chars, no \n\n)
  • Any file with base64 images, ANSI escapes, or control characters caused 500s

Solution: chunk_text() in bulk_wiki_ingest_ollama.py

Three-level recursive split:

1. Headings: split by ##, then ###, then ####
2. Paragraphs: split by \n\n (within blocks > MAX_CHUNK_SIZE)
3. Words: split by space (last resort, when a single paragraph exceeds MAX_CHUNK_SIZE)

Config

MAX_CHUNK_SIZE = 5000   # safe buffer below bge-m3 ~6000 limit
CHUNK_OVERLAP = 300     # overlap between consecutive chunks

sanitize_text() — Pre-embedding Cleanup

Removes content that causes Ollama 500:

  • Base64 images (data:image/...) → [IMAGE] placeholder
  • Control characters (except \n, \r, \t) → stripped
  • ANSI escape sequences (\x1b[...) → stripped
  • Excessive whitespace → collapsed

Point ID Scheme

unique_id = hashlib.md5(f"{filepath}:chunk:{chunk_index}".encode()).hexdigest()
point_id = int(unique_id[:8], 16) & 0x7fffffff

Payload Fields

{
    "filename": filepath.name,
    "path": str(filepath),
    "content": clean_text[:1000],  # preview
    "length": len(clean_text),
    "title": title,
    "chunk_index": chunk_index,
    "chunk_total": len(chunks),
    "indexed_at": datetime.now().isoformat(),
}

Results

Metric Before After
Total points 149 159
Files indexed 75/85 78/85
Ollama 500 errors 2 0
Configuration backup.md truncated (1 chunk) 53 chunks
Конфигурация компакт Микротика.md 500 error 6 chunks
Цифровой паспорт.md 2 chunks 3 chunks

Verification

Three searches confirmed deep content retrieval:

# 1. "заселение по биометрии" → Цифровой паспорт (score 0.61)
# 2. "TemperatureLimit DCMIConfiguration thermal" → Configuration backup (score 0.73)
# 3. "MikroTik компактный экспорт" → Конфигурация Микротика (score 0.62)

Pitfalls Encountered

  1. Ollama 500 on large paragraphs: Some files (like конфигурация Микротика) have code blocks with no \n\n for hundreds of lines. The paragraph split returned one element > MAX_CHUNK_SIZE. Fix: added word-level split as last resort.

  2. Base64 images in markdown: Files like Цифровой паспорт.md contain forwarded emails with embedded base64 images. Ollama 500 on the raw data. Fix: sanitize_text() regex replaces data:image/...;base64,... with [base64-data].

  3. ANSI escape sequences: Config files with terminal control codes. Fix: strip \x1b\[[0-9;]*[a-zA-Z] patterns.

  4. Empty files produce Ollama error: 7 files in vault are 0 bytes. Skip with if not content.strip(): return 0.