Files
rag-pipeline-docker/SKILL.md
T
2026-09-06 13:51:06 +00:00

9.9 KiB

name, title, category, description, triggers
name title category description triggers
rag-pipeline-docker RAG Pipeline in Docker Compose devops Deploying a RAG pipeline with Docker Compose — Qdrant (vector DB), Redis (job queue), ARQ worker (chunking + embedding), connected to an external Ollama instance for embeddings and LLM inference. Covers env var management, cross-network connectivity, cron-based sync+ingest, and troubleshooting.
qdrant docker
arq worker
rag pipeline deploy
memory-os setup
vector db docker compose
ollama + qdrant integration
wiki ingest pipeline
obsidian sync qdrant

RAG Pipeline in Docker Compose

Architecture

Obsidian vault (host)
  ↓ sync_obsidian_to_wiki.py (cron: every 10m)
Wiki path (host)
  ↓ ARQ queue → worker container
  ↓ chunk_text() + get_embedding() + get_sparse_embedding()
Qdrant (localhost:6333 / collection: knowledge_base)
  ↓ dense: nomic-embed-text (768d, Cosine)
  ↓ sparse: BM25 (on_disk)

Quick Start

1. Docker Compose Stack

services:
  redis:
    image: redis:7-alpine
    restart: unless-stopped
    # password via REDIS_PASSWORD env
    healthcheck: [CMD-SHELL, "redis-cli ${REDIS_PASSWORD:+-a $REDIS_PASSWORD} ping"]

  qdrant:
    image: qdrant/qdrant:v1.17.1
    restart: unless-stopped
    ports: ["127.0.0.1:6333:6333"]
    volumes: [qdrant_data:/qdrant/storage]
    healthcheck: [CMD, sh, -c, "grep -q ':18BD' /proc/net/tcp"]

  worker:
    build: ./worker
    restart: unless-stopped
    depends_on: [qdrant, redis]
    # see env section below
    volumes:
      - wiki_path:/wiki:ro
      - hermes_home:/hermes:rw

2. Environment Variables

LLM for reflection / reasoning (inside worker):

OLLAMA_BASE_URL=http://ollama:11434
OLLAMA_MODEL=qwen3-8b-64k

Embedding (inside worker):

EMBEDDING_API_BASE=http://ollama:11434/v1
EMBEDDING_MODEL=nomic-embed-text:latest
EMBEDDING_DIMS=768
EMBEDDING_API_KEY=

Redis:

REDIS_PASSWORD=<your-password>
REDIS_HOST=redis
REDIS_PORT=6379

Qdrant:

QDRANT_HOST=qdrant
QDRANT_PORT=6333
COLLECTION_NAME=knowledge_base

3. Cross-Stack Network

If Qdrant/Redis/worker are in one compose stack and Ollama is in another, the worker needs access to both networks:

services:
  worker:
    networks:
      - default          # memory-os_default — for Redis + Qdrant
      - ollama_default   # external — for Ollama DNS

networks:
  default:
    name: memory-os_default
  ollama_default:
    external: true

Critical: host.docker.internal does NOT work on Linux (Docker Desktop only). Use ollama:11434 (via shared network) or 172.17.0.1:11434 (host gateway) instead.

4. Verify Connectivity

# DNS resolution
docker exec <worker> getent hosts ollama

# Ollama API
docker exec <worker> python3 -c "
import urllib.request, json
req = urllib.request.Request('http://ollama:11434/api/tags')
resp = urllib.request.urlopen(req, timeout=10)
data = json.loads(resp.read())
print(f'Models: {len(data[\"models\"])}')
"

# Qdrant collection
curl -s http://127.0.0.1:6333/collections/knowledge_base | python3 -c "
import sys,json; d=json.load(sys.stdin)
print(f'points: {d[\"result\"][\"points_count\"]}')
"

Search API (FastAPI)

Add a search API layer that accepts text queries and returns results from Qdrant:

Docker Compose Service

search-api:
    build:
      context: ../search_api       # relative to docker/ directory
      dockerfile: Dockerfile
    restart: unless-stopped
    depends_on:
      qdrant:
        condition: service_healthy
    networks:
      - default
      - ollama_default
    environment:
      OLLAMA_URL: http://ollama:11434
      OLLAMA_EMBEDDING_MODEL: nomic-embed-text:latest
      QDRANT_URL: http://qdrant:6333
      COLLECTION_NAME: ${COLLECTION_NAME:-knowledge_base}
    ports:
      - "127.0.0.1:8000:8000"
    healthcheck:
      test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/health', timeout=5)"]
      interval: 15s
      timeout: 5s
      retries: 5
      start_period: 10s

FastAPI App Structure

search_api/
├── Dockerfile
├── requirements.txt   # fastapi, uvicorn, httpx, pydantic
└── main.py

Endpoints

  • GET /health — returns {"status": "ok", "qdrant": true, "ollama": true}
  • POST /search — accepts {"query": "...", "top_k": 5}, returns {"query": "...", "results": [...], "total": N}

Flow

  1. Receive text query → POST to Ollama /api/embeddings (nomic-embed-text)
  2. Use returned dense vector → POST to Qdrant /collections/{name}/points/search
  3. Return results with score, text, source

Verify

# Health
curl http://127.0.0.1:8000/health

# Search
curl -X POST http://127.0.0.1:8000/search \
  -H 'Content-Type: application/json' \
  -d '{"query":"your search text","top_k":3}'

Periodic Tasks

Sync + Ingest (every 10 min)

Set up a cron job that runs every 10 minutes:

  1. Sync script — copies new/changed .md files from Obsidian vault to wiki path, tracking state via JSON file
  2. Ingest script — detects new/modified files, enqueues them to ARQ worker for chunking + embedding
# Manual run
python3 /path/to/scripts/sync_obsidian_to_wiki.py
python3 /path/to/scripts/wiki_continuous_ingest.py

Via Hermes cronjob (LLM-driven — uses no_agent: false):

hermes cron create \
  --name "memory-os sync+ingest" \
  --schedule "every 10m" \
  --prompt "Run: python3 /path/to/sync_obsidian_to_wiki.py"

Micro-Reflection Trigger (every 5 min, silent)

An ARQ worker can have a process_micro_reflection function that runs idle-time reflection. To trigger it on a schedule without LLM overhead, use a no_agent: true watchdog cronjob that runs a script inside the worker container.

Pre-requisite: Mount the scripts directory into the worker container:

services:
  worker:
    volumes:
      - ../scripts:/app/scripts:ro   # relative to docker/ directory

Script (reflection_trigger.py): checks if the ARQ worker is idle (no pending/executing jobs), respects a per-hour budget, and enqueues process_micro_reflection via Redis.

Cronjob (no_agent, silent, local):

hermes cron create \
  --name "memory-os micro-reflection" \
  --schedule "*/5 * * * *" \
  --script "docker exec <worker> python3 /app/scripts/reflection_trigger.py" \
  --no-agent
hermes cron update \
  --job-id <id> \
  --deliver local

Key points:

  • no_agent: true — no LLM tokens consumed, just runs the script and delivers stdout verbatim
  • deliver: local — suppresses Telegram/Discord notifications; the job runs silently
  • Empty stdout = silent (no message sent), error output = alert delivered
  • The script must be on the host filesystem AND mounted into the container via volumes:

Checking Worker Health

# Container status
docker ps --filter name=worker

# Worker logs
docker logs <worker> --tail 50

# Check for errors
docker logs <worker> 2>&1 | grep -i "error\|traceback\|exception" | head -10

# ARQ stats (from worker logs)
docker logs <worker> 2>&1 | grep "j_complete\|j_failed"

Pitfalls

host.docker.internal on Linux

host.docker.internal is a Docker Desktop feature (macOS/Windows). On Linux, it does not resolve. Use one of:

  • Container name on shared network: http://ollama:11434
  • Host gateway: http://172.17.0.1:11434

Env vars not propagated to container

Variables defined in .env are NOT automatically available inside containers — they must be explicitly listed in docker-compose.yml under services.worker.environment. docker compose config can verify the effective config.

Redis password mismatch

If the worker uses redis.asyncio or arq.connections.RedisSettings, ensure the password matches what's in redis.conf. Test with redis-cli -a $PASSWORD ping.

Network detachment on recreate

When a container is recreated via docker compose up -d --force-recreate, it may lose connections to external networks. The fix is to declare the network in docker-compose.yml with external: true and add it to the service's networks: list.

Qdrant healthcheck on custom port

The default Qdrant healthcheck greps /proc/net/tcp for :18BD (port 6333 in hex). If using a non-standard port, update the healthcheck.

ARQ worker timeout

The ollama_chat function in reflection tasks may timeout if the model is large or generating long responses. Set ARQ_JOB_TIMEOUT high enough (e.g., 300s) and ensure httpx.AsyncClient(timeout=120) matches.

no_agent cron script must be on host filesystem

A no_agent: true cronjob's --script runs on the host, not inside the container. If the script only exists inside the container (e.g., at /app/scripts/), the cronjob will fail. Mount the scripts directory into the container AND keep the script accessible on the host, or use docker exec to run it inside the container:

--script "docker exec <container> python3 /app/scripts/script.py"

reflection_trigger.py paths hardcoded to old project

The reflection_trigger.py script was originally written for a different project (~/.ai-stack/). The .env path and log paths must be updated to match the new project layout before the script works after a copy. Search for Path.home() / "ai-stack" or similar hardcoded paths and update them to the new project root.

Volume paths in docker-compose are relative to compose file

When adding a volumes: mount like - ../scripts:/app/scripts:ro, the path is relative to the docker-compose.yml file's directory, not the project root. If the compose file is in docker/, then ../scripts resolves to project/scripts/.

Support Files

  • references/memory-os-session.md — session-specific details from the Memory OS deployment (env files, state files, error transcripts, search API code)
  • scripts/test_qdrant_search.py — standalone test script: gets embedding from Ollama, searches Qdrant, prints top-5 results