mirror of
https://gitverse.ru/kpa39l/memory-os.git
synced 2026-09-29 09:35:05 +00:00
97 lines
3.2 KiB
Markdown
97 lines
3.2 KiB
Markdown
# Chunking Implementation — 2026-07-16
|
|
|
|
## Problem
|
|
|
|
bge-m3 via Ollama `/api/embeddings` rejects inputs >~6000 chars with HTTP 500.
|
|
The old `bulk_wiki_ingest_ollama.py` truncated files to 6000 chars, losing
|
|
mid/end content from large documents.
|
|
|
|
Affected files (before chunking):
|
|
- Configuration backup.md (240KB) — only first 6K of 240KB indexed
|
|
- Конфигурация компакт Микротика.md (19KB) — Ollama 500 on chunk 2 (16083 chars, no `\n\n`)
|
|
- Any file with base64 images, ANSI escapes, or control characters caused 500s
|
|
|
|
## Solution: `chunk_text()` in `bulk_wiki_ingest_ollama.py`
|
|
|
|
Three-level recursive split:
|
|
|
|
```
|
|
1. Headings: split by ##, then ###, then ####
|
|
2. Paragraphs: split by \n\n (within blocks > MAX_CHUNK_SIZE)
|
|
3. Words: split by space (last resort, when a single paragraph exceeds MAX_CHUNK_SIZE)
|
|
```
|
|
|
|
### Config
|
|
|
|
```python
|
|
MAX_CHUNK_SIZE = 5000 # safe buffer below bge-m3 ~6000 limit
|
|
CHUNK_OVERLAP = 300 # overlap between consecutive chunks
|
|
```
|
|
|
|
### `sanitize_text()` — Pre-embedding Cleanup
|
|
|
|
Removes content that causes Ollama 500:
|
|
|
|
- Base64 images (`data:image/...`) → `[IMAGE]` placeholder
|
|
- Control characters (except `\n`, `\r`, `\t`) → stripped
|
|
- ANSI escape sequences (`\x1b[...`) → stripped
|
|
- Excessive whitespace → collapsed
|
|
|
|
### Point ID Scheme
|
|
|
|
```python
|
|
unique_id = hashlib.md5(f"{filepath}:chunk:{chunk_index}".encode()).hexdigest()
|
|
point_id = int(unique_id[:8], 16) & 0x7fffffff
|
|
```
|
|
|
|
### Payload Fields
|
|
|
|
```python
|
|
{
|
|
"filename": filepath.name,
|
|
"path": str(filepath),
|
|
"content": clean_text[:1000], # preview
|
|
"length": len(clean_text),
|
|
"title": title,
|
|
"chunk_index": chunk_index,
|
|
"chunk_total": len(chunks),
|
|
"indexed_at": datetime.now().isoformat(),
|
|
}
|
|
```
|
|
|
|
## Results
|
|
|
|
| Metric | Before | After |
|
|
|---|---|---|
|
|
| Total points | 149 | 159 |
|
|
| Files indexed | 75/85 | 78/85 |
|
|
| Ollama 500 errors | 2 | 0 |
|
|
| Configuration backup.md | truncated (1 chunk) | 53 chunks |
|
|
| Конфигурация компакт Микротика.md | 500 error | 6 chunks |
|
|
| Цифровой паспорт.md | 2 chunks | 3 chunks |
|
|
|
|
## Verification
|
|
|
|
Three searches confirmed deep content retrieval:
|
|
|
|
```bash
|
|
# 1. "заселение по биометрии" → Цифровой паспорт (score 0.61)
|
|
# 2. "TemperatureLimit DCMIConfiguration thermal" → Configuration backup (score 0.73)
|
|
# 3. "MikroTik компактный экспорт" → Конфигурация Микротика (score 0.62)
|
|
```
|
|
|
|
## Pitfalls Encountered
|
|
|
|
1. **Ollama 500 on large paragraphs:** Some files (like конфигурация Микротика) have
|
|
code blocks with no `\n\n` for hundreds of lines. The paragraph split returned
|
|
one element > MAX_CHUNK_SIZE. Fix: added word-level split as last resort.
|
|
|
|
2. **Base64 images in markdown:** Files like Цифровой паспорт.md contain forwarded
|
|
emails with embedded base64 images. Ollama 500 on the raw data. Fix: `sanitize_text()`
|
|
regex replaces `data:image/...;base64,...` with `[base64-data]`.
|
|
|
|
3. **ANSI escape sequences:** Config files with terminal control codes. Fix: strip
|
|
`\x1b\[[0-9;]*[a-zA-Z]` patterns.
|
|
|
|
4. **Empty files produce Ollama error:** 7 files in vault are 0 bytes. Skip with
|
|
`if not content.strip(): return 0`. |