Files
2026-09-06 13:51:06 +00:00

308 lines
9.9 KiB
Markdown

---
name: rag-pipeline-docker
title: RAG Pipeline in Docker Compose
category: devops
description: Deploying a RAG pipeline with Docker Compose — Qdrant (vector DB), Redis (job queue), ARQ worker (chunking + embedding), connected to an external Ollama instance for embeddings and LLM inference. Covers env var management, cross-network connectivity, cron-based sync+ingest, and troubleshooting.
triggers:
- qdrant docker
- arq worker
- rag pipeline deploy
- memory-os setup
- vector db docker compose
- ollama + qdrant integration
- wiki ingest pipeline
- obsidian sync qdrant
---
# RAG Pipeline in Docker Compose
## Architecture
```
Obsidian vault (host)
↓ sync_obsidian_to_wiki.py (cron: every 10m)
Wiki path (host)
↓ ARQ queue → worker container
↓ chunk_text() + get_embedding() + get_sparse_embedding()
Qdrant (localhost:6333 / collection: knowledge_base)
↓ dense: nomic-embed-text (768d, Cosine)
↓ sparse: BM25 (on_disk)
```
## Quick Start
### 1. Docker Compose Stack
```yaml
services:
redis:
image: redis:7-alpine
restart: unless-stopped
# password via REDIS_PASSWORD env
healthcheck: [CMD-SHELL, "redis-cli ${REDIS_PASSWORD:+-a $REDIS_PASSWORD} ping"]
qdrant:
image: qdrant/qdrant:v1.17.1
restart: unless-stopped
ports: ["127.0.0.1:6333:6333"]
volumes: [qdrant_data:/qdrant/storage]
healthcheck: [CMD, sh, -c, "grep -q ':18BD' /proc/net/tcp"]
worker:
build: ./worker
restart: unless-stopped
depends_on: [qdrant, redis]
# see env section below
volumes:
- wiki_path:/wiki:ro
- hermes_home:/hermes:rw
```
### 2. Environment Variables
**LLM for reflection / reasoning (inside worker):**
```env
OLLAMA_BASE_URL=http://ollama:11434
OLLAMA_MODEL=qwen3-8b-64k
```
**Embedding (inside worker):**
```env
EMBEDDING_API_BASE=http://ollama:11434/v1
EMBEDDING_MODEL=nomic-embed-text:latest
EMBEDDING_DIMS=768
EMBEDDING_API_KEY=
```
**Redis:**
```env
REDIS_PASSWORD=<your-password>
REDIS_HOST=redis
REDIS_PORT=6379
```
**Qdrant:**
```env
QDRANT_HOST=qdrant
QDRANT_PORT=6333
COLLECTION_NAME=knowledge_base
```
### 3. Cross-Stack Network
If Qdrant/Redis/worker are in one compose stack and Ollama is in another, the worker needs access to both networks:
```yaml
services:
worker:
networks:
- default # memory-os_default — for Redis + Qdrant
- ollama_default # external — for Ollama DNS
networks:
default:
name: memory-os_default
ollama_default:
external: true
```
> **Critical:** `host.docker.internal` does NOT work on Linux (Docker Desktop only). Use `ollama:11434` (via shared network) or `172.17.0.1:11434` (host gateway) instead.
### 4. Verify Connectivity
```bash
# DNS resolution
docker exec <worker> getent hosts ollama
# Ollama API
docker exec <worker> python3 -c "
import urllib.request, json
req = urllib.request.Request('http://ollama:11434/api/tags')
resp = urllib.request.urlopen(req, timeout=10)
data = json.loads(resp.read())
print(f'Models: {len(data[\"models\"])}')
"
# Qdrant collection
curl -s http://127.0.0.1:6333/collections/knowledge_base | python3 -c "
import sys,json; d=json.load(sys.stdin)
print(f'points: {d[\"result\"][\"points_count\"]}')
"
```
## Search API (FastAPI)
Add a search API layer that accepts text queries and returns results from Qdrant:
### Docker Compose Service
```yaml
search-api:
build:
context: ../search_api # relative to docker/ directory
dockerfile: Dockerfile
restart: unless-stopped
depends_on:
qdrant:
condition: service_healthy
networks:
- default
- ollama_default
environment:
OLLAMA_URL: http://ollama:11434
OLLAMA_EMBEDDING_MODEL: nomic-embed-text:latest
QDRANT_URL: http://qdrant:6333
COLLECTION_NAME: ${COLLECTION_NAME:-knowledge_base}
ports:
- "127.0.0.1:8000:8000"
healthcheck:
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/health', timeout=5)"]
interval: 15s
timeout: 5s
retries: 5
start_period: 10s
```
### FastAPI App Structure
```
search_api/
├── Dockerfile
├── requirements.txt # fastapi, uvicorn, httpx, pydantic
└── main.py
```
### Endpoints
- `GET /health` — returns `{"status": "ok", "qdrant": true, "ollama": true}`
- `POST /search` — accepts `{"query": "...", "top_k": 5}`, returns `{"query": "...", "results": [...], "total": N}`
### Flow
1. Receive text query → POST to Ollama `/api/embeddings` (nomic-embed-text)
2. Use returned dense vector → POST to Qdrant `/collections/{name}/points/search`
3. Return results with score, text, source
### Verify
```bash
# Health
curl http://127.0.0.1:8000/health
# Search
curl -X POST http://127.0.0.1:8000/search \
-H 'Content-Type: application/json' \
-d '{"query":"your search text","top_k":3}'
```
## Periodic Tasks
### Sync + Ingest (every 10 min)
Set up a cron job that runs every 10 minutes:
1. **Sync script** — copies new/changed `.md` files from Obsidian vault to wiki path, tracking state via JSON file
2. **Ingest script** — detects new/modified files, enqueues them to ARQ worker for chunking + embedding
```bash
# Manual run
python3 /path/to/scripts/sync_obsidian_to_wiki.py
python3 /path/to/scripts/wiki_continuous_ingest.py
```
Via Hermes cronjob (LLM-driven — uses `no_agent: false`):
```
hermes cron create \
--name "memory-os sync+ingest" \
--schedule "every 10m" \
--prompt "Run: python3 /path/to/sync_obsidian_to_wiki.py"
```
### Micro-Reflection Trigger (every 5 min, silent)
An ARQ worker can have a `process_micro_reflection` function that runs idle-time reflection. To trigger it on a schedule **without LLM overhead**, use a `no_agent: true` watchdog cronjob that runs a script inside the worker container.
**Pre-requisite:** Mount the scripts directory into the worker container:
```yaml
services:
worker:
volumes:
- ../scripts:/app/scripts:ro # relative to docker/ directory
```
**Script** (`reflection_trigger.py`): checks if the ARQ worker is idle (no pending/executing jobs), respects a per-hour budget, and enqueues `process_micro_reflection` via Redis.
**Cronjob (no_agent, silent, local):**
```
hermes cron create \
--name "memory-os micro-reflection" \
--schedule "*/5 * * * *" \
--script "docker exec <worker> python3 /app/scripts/reflection_trigger.py" \
--no-agent
hermes cron update \
--job-id <id> \
--deliver local
```
Key points:
- `no_agent: true` — no LLM tokens consumed, just runs the script and delivers stdout verbatim
- `deliver: local` — suppresses Telegram/Discord notifications; the job runs silently
- Empty stdout = silent (no message sent), error output = alert delivered
- The script must be on the host filesystem AND mounted into the container via `volumes:`
## Checking Worker Health
```bash
# Container status
docker ps --filter name=worker
# Worker logs
docker logs <worker> --tail 50
# Check for errors
docker logs <worker> 2>&1 | grep -i "error\|traceback\|exception" | head -10
# ARQ stats (from worker logs)
docker logs <worker> 2>&1 | grep "j_complete\|j_failed"
```
## Pitfalls
### `host.docker.internal` on Linux
`host.docker.internal` is a Docker Desktop feature (macOS/Windows). On Linux, it does not resolve. Use one of:
- Container name on shared network: `http://ollama:11434`
- Host gateway: `http://172.17.0.1:11434`
### Env vars not propagated to container
Variables defined in `.env` are NOT automatically available inside containers — they must be explicitly listed in `docker-compose.yml` under `services.worker.environment`. `docker compose config` can verify the effective config.
### Redis password mismatch
If the worker uses `redis.asyncio` or `arq.connections.RedisSettings`, ensure the password matches what's in `redis.conf`. Test with `redis-cli -a $PASSWORD ping`.
### Network detachment on recreate
When a container is recreated via `docker compose up -d --force-recreate`, it may lose connections to external networks. The fix is to declare the network in `docker-compose.yml` with `external: true` and add it to the service's `networks:` list.
### Qdrant healthcheck on custom port
The default Qdrant healthcheck greps `/proc/net/tcp` for `:18BD` (port 6333 in hex). If using a non-standard port, update the healthcheck.
### ARQ worker timeout
The `ollama_chat` function in reflection tasks may timeout if the model is large or generating long responses. Set `ARQ_JOB_TIMEOUT` high enough (e.g., 300s) and ensure `httpx.AsyncClient(timeout=120)` matches.
### `no_agent` cron script must be on host filesystem
A `no_agent: true` cronjob's `--script` runs on the host, not inside the container. If the script only exists inside the container (e.g., at `/app/scripts/`), the cronjob will fail. Mount the scripts directory into the container AND keep the script accessible on the host, or use `docker exec` to run it inside the container:
```
--script "docker exec <container> python3 /app/scripts/script.py"
```
### `reflection_trigger.py` paths hardcoded to old project
The `reflection_trigger.py` script was originally written for a different project (`~/.ai-stack/`). The `.env` path and log paths must be updated to match the new project layout before the script works after a copy. Search for `Path.home() / "ai-stack"` or similar hardcoded paths and update them to the new project root.
### Volume paths in docker-compose are relative to compose file
When adding a `volumes:` mount like `- ../scripts:/app/scripts:ro`, the path is relative to the `docker-compose.yml` file's directory, not the project root. If the compose file is in `docker/`, then `../scripts` resolves to `project/scripts/`.
## Support Files
- **`references/memory-os-session.md`** — session-specific details from the Memory OS deployment (env files, state files, error transcripts, search API code)
- **`scripts/test_qdrant_search.py`** — standalone test script: gets embedding from Ollama, searches Qdrant, prints top-5 results