--- name: rag-pipeline-docker title: RAG Pipeline in Docker Compose category: devops description: Deploying a RAG pipeline with Docker Compose — Qdrant (vector DB), Redis (job queue), ARQ worker (chunking + embedding), connected to an external Ollama instance for embeddings and LLM inference. Covers env var management, cross-network connectivity, cron-based sync+ingest, and troubleshooting. triggers: - qdrant docker - arq worker - rag pipeline deploy - memory-os setup - vector db docker compose - ollama + qdrant integration - wiki ingest pipeline - obsidian sync qdrant --- # RAG Pipeline in Docker Compose ## Architecture ``` Obsidian vault (host) ↓ sync_obsidian_to_wiki.py (cron: every 10m) Wiki path (host) ↓ ARQ queue → worker container ↓ chunk_text() + get_embedding() + get_sparse_embedding() Qdrant (localhost:6333 / collection: knowledge_base) ↓ dense: nomic-embed-text (768d, Cosine) ↓ sparse: BM25 (on_disk) ``` ## Quick Start ### 1. Docker Compose Stack ```yaml services: redis: image: redis:7-alpine restart: unless-stopped # password via REDIS_PASSWORD env healthcheck: [CMD-SHELL, "redis-cli ${REDIS_PASSWORD:+-a $REDIS_PASSWORD} ping"] qdrant: image: qdrant/qdrant:v1.17.1 restart: unless-stopped ports: ["127.0.0.1:6333:6333"] volumes: [qdrant_data:/qdrant/storage] healthcheck: [CMD, sh, -c, "grep -q ':18BD' /proc/net/tcp"] worker: build: ./worker restart: unless-stopped depends_on: [qdrant, redis] # see env section below volumes: - wiki_path:/wiki:ro - hermes_home:/hermes:rw ``` ### 2. Environment Variables **LLM for reflection / reasoning (inside worker):** ```env OLLAMA_BASE_URL=http://ollama:11434 OLLAMA_MODEL=qwen3-8b-64k ``` **Embedding (inside worker):** ```env EMBEDDING_API_BASE=http://ollama:11434/v1 EMBEDDING_MODEL=nomic-embed-text:latest EMBEDDING_DIMS=768 EMBEDDING_API_KEY= ``` **Redis:** ```env REDIS_PASSWORD= REDIS_HOST=redis REDIS_PORT=6379 ``` **Qdrant:** ```env QDRANT_HOST=qdrant QDRANT_PORT=6333 COLLECTION_NAME=knowledge_base ``` ### 3. Cross-Stack Network If Qdrant/Redis/worker are in one compose stack and Ollama is in another, the worker needs access to both networks: ```yaml services: worker: networks: - default # memory-os_default — for Redis + Qdrant - ollama_default # external — for Ollama DNS networks: default: name: memory-os_default ollama_default: external: true ``` > **Critical:** `host.docker.internal` does NOT work on Linux (Docker Desktop only). Use `ollama:11434` (via shared network) or `172.17.0.1:11434` (host gateway) instead. ### 4. Verify Connectivity ```bash # DNS resolution docker exec getent hosts ollama # Ollama API docker exec python3 -c " import urllib.request, json req = urllib.request.Request('http://ollama:11434/api/tags') resp = urllib.request.urlopen(req, timeout=10) data = json.loads(resp.read()) print(f'Models: {len(data[\"models\"])}') " # Qdrant collection curl -s http://127.0.0.1:6333/collections/knowledge_base | python3 -c " import sys,json; d=json.load(sys.stdin) print(f'points: {d[\"result\"][\"points_count\"]}') " ``` ## Search API (FastAPI) Add a search API layer that accepts text queries and returns results from Qdrant: ### Docker Compose Service ```yaml search-api: build: context: ../search_api # relative to docker/ directory dockerfile: Dockerfile restart: unless-stopped depends_on: qdrant: condition: service_healthy networks: - default - ollama_default environment: OLLAMA_URL: http://ollama:11434 OLLAMA_EMBEDDING_MODEL: nomic-embed-text:latest QDRANT_URL: http://qdrant:6333 COLLECTION_NAME: ${COLLECTION_NAME:-knowledge_base} ports: - "127.0.0.1:8000:8000" healthcheck: test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/health', timeout=5)"] interval: 15s timeout: 5s retries: 5 start_period: 10s ``` ### FastAPI App Structure ``` search_api/ ├── Dockerfile ├── requirements.txt # fastapi, uvicorn, httpx, pydantic └── main.py ``` ### Endpoints - `GET /health` — returns `{"status": "ok", "qdrant": true, "ollama": true}` - `POST /search` — accepts `{"query": "...", "top_k": 5}`, returns `{"query": "...", "results": [...], "total": N}` ### Flow 1. Receive text query → POST to Ollama `/api/embeddings` (nomic-embed-text) 2. Use returned dense vector → POST to Qdrant `/collections/{name}/points/search` 3. Return results with score, text, source ### Verify ```bash # Health curl http://127.0.0.1:8000/health # Search curl -X POST http://127.0.0.1:8000/search \ -H 'Content-Type: application/json' \ -d '{"query":"your search text","top_k":3}' ``` ## Periodic Tasks ### Sync + Ingest (every 10 min) Set up a cron job that runs every 10 minutes: 1. **Sync script** — copies new/changed `.md` files from Obsidian vault to wiki path, tracking state via JSON file 2. **Ingest script** — detects new/modified files, enqueues them to ARQ worker for chunking + embedding ```bash # Manual run python3 /path/to/scripts/sync_obsidian_to_wiki.py python3 /path/to/scripts/wiki_continuous_ingest.py ``` Via Hermes cronjob (LLM-driven — uses `no_agent: false`): ``` hermes cron create \ --name "memory-os sync+ingest" \ --schedule "every 10m" \ --prompt "Run: python3 /path/to/sync_obsidian_to_wiki.py" ``` ### Micro-Reflection Trigger (every 5 min, silent) An ARQ worker can have a `process_micro_reflection` function that runs idle-time reflection. To trigger it on a schedule **without LLM overhead**, use a `no_agent: true` watchdog cronjob that runs a script inside the worker container. **Pre-requisite:** Mount the scripts directory into the worker container: ```yaml services: worker: volumes: - ../scripts:/app/scripts:ro # relative to docker/ directory ``` **Script** (`reflection_trigger.py`): checks if the ARQ worker is idle (no pending/executing jobs), respects a per-hour budget, and enqueues `process_micro_reflection` via Redis. **Cronjob (no_agent, silent, local):** ``` hermes cron create \ --name "memory-os micro-reflection" \ --schedule "*/5 * * * *" \ --script "docker exec python3 /app/scripts/reflection_trigger.py" \ --no-agent hermes cron update \ --job-id \ --deliver local ``` Key points: - `no_agent: true` — no LLM tokens consumed, just runs the script and delivers stdout verbatim - `deliver: local` — suppresses Telegram/Discord notifications; the job runs silently - Empty stdout = silent (no message sent), error output = alert delivered - The script must be on the host filesystem AND mounted into the container via `volumes:` ## Checking Worker Health ```bash # Container status docker ps --filter name=worker # Worker logs docker logs --tail 50 # Check for errors docker logs 2>&1 | grep -i "error\|traceback\|exception" | head -10 # ARQ stats (from worker logs) docker logs 2>&1 | grep "j_complete\|j_failed" ``` ## Pitfalls ### `host.docker.internal` on Linux `host.docker.internal` is a Docker Desktop feature (macOS/Windows). On Linux, it does not resolve. Use one of: - Container name on shared network: `http://ollama:11434` - Host gateway: `http://172.17.0.1:11434` ### Env vars not propagated to container Variables defined in `.env` are NOT automatically available inside containers — they must be explicitly listed in `docker-compose.yml` under `services.worker.environment`. `docker compose config` can verify the effective config. ### Redis password mismatch If the worker uses `redis.asyncio` or `arq.connections.RedisSettings`, ensure the password matches what's in `redis.conf`. Test with `redis-cli -a $PASSWORD ping`. ### Network detachment on recreate When a container is recreated via `docker compose up -d --force-recreate`, it may lose connections to external networks. The fix is to declare the network in `docker-compose.yml` with `external: true` and add it to the service's `networks:` list. ### Qdrant healthcheck on custom port The default Qdrant healthcheck greps `/proc/net/tcp` for `:18BD` (port 6333 in hex). If using a non-standard port, update the healthcheck. ### ARQ worker timeout The `ollama_chat` function in reflection tasks may timeout if the model is large or generating long responses. Set `ARQ_JOB_TIMEOUT` high enough (e.g., 300s) and ensure `httpx.AsyncClient(timeout=120)` matches. ### `no_agent` cron script must be on host filesystem A `no_agent: true` cronjob's `--script` runs on the host, not inside the container. If the script only exists inside the container (e.g., at `/app/scripts/`), the cronjob will fail. Mount the scripts directory into the container AND keep the script accessible on the host, or use `docker exec` to run it inside the container: ``` --script "docker exec python3 /app/scripts/script.py" ``` ### `reflection_trigger.py` paths hardcoded to old project The `reflection_trigger.py` script was originally written for a different project (`~/.ai-stack/`). The `.env` path and log paths must be updated to match the new project layout before the script works after a copy. Search for `Path.home() / "ai-stack"` or similar hardcoded paths and update them to the new project root. ### Volume paths in docker-compose are relative to compose file When adding a `volumes:` mount like `- ../scripts:/app/scripts:ro`, the path is relative to the `docker-compose.yml` file's directory, not the project root. If the compose file is in `docker/`, then `../scripts` resolves to `project/scripts/`. ## Support Files - **`references/memory-os-session.md`** — session-specific details from the Memory OS deployment (env files, state files, error transcripts, search API code) - **`scripts/test_qdrant_search.py`** — standalone test script: gets embedding from Ollama, searches Qdrant, prints top-5 results