mirror of
https://gitverse.ru/kpa39l/rag-pipeline-docker.git
synced 2026-09-29 09:15:11 +00:00
308 lines
9.9 KiB
Markdown
308 lines
9.9 KiB
Markdown
---
|
|
name: rag-pipeline-docker
|
|
title: RAG Pipeline in Docker Compose
|
|
category: devops
|
|
description: Deploying a RAG pipeline with Docker Compose — Qdrant (vector DB), Redis (job queue), ARQ worker (chunking + embedding), connected to an external Ollama instance for embeddings and LLM inference. Covers env var management, cross-network connectivity, cron-based sync+ingest, and troubleshooting.
|
|
triggers:
|
|
- qdrant docker
|
|
- arq worker
|
|
- rag pipeline deploy
|
|
- memory-os setup
|
|
- vector db docker compose
|
|
- ollama + qdrant integration
|
|
- wiki ingest pipeline
|
|
- obsidian sync qdrant
|
|
---
|
|
|
|
# RAG Pipeline in Docker Compose
|
|
|
|
## Architecture
|
|
|
|
```
|
|
Obsidian vault (host)
|
|
↓ sync_obsidian_to_wiki.py (cron: every 10m)
|
|
Wiki path (host)
|
|
↓ ARQ queue → worker container
|
|
↓ chunk_text() + get_embedding() + get_sparse_embedding()
|
|
Qdrant (localhost:6333 / collection: knowledge_base)
|
|
↓ dense: nomic-embed-text (768d, Cosine)
|
|
↓ sparse: BM25 (on_disk)
|
|
```
|
|
|
|
## Quick Start
|
|
|
|
### 1. Docker Compose Stack
|
|
|
|
```yaml
|
|
services:
|
|
redis:
|
|
image: redis:7-alpine
|
|
restart: unless-stopped
|
|
# password via REDIS_PASSWORD env
|
|
healthcheck: [CMD-SHELL, "redis-cli ${REDIS_PASSWORD:+-a $REDIS_PASSWORD} ping"]
|
|
|
|
qdrant:
|
|
image: qdrant/qdrant:v1.17.1
|
|
restart: unless-stopped
|
|
ports: ["127.0.0.1:6333:6333"]
|
|
volumes: [qdrant_data:/qdrant/storage]
|
|
healthcheck: [CMD, sh, -c, "grep -q ':18BD' /proc/net/tcp"]
|
|
|
|
worker:
|
|
build: ./worker
|
|
restart: unless-stopped
|
|
depends_on: [qdrant, redis]
|
|
# see env section below
|
|
volumes:
|
|
- wiki_path:/wiki:ro
|
|
- hermes_home:/hermes:rw
|
|
```
|
|
|
|
### 2. Environment Variables
|
|
|
|
**LLM for reflection / reasoning (inside worker):**
|
|
```env
|
|
OLLAMA_BASE_URL=http://ollama:11434
|
|
OLLAMA_MODEL=qwen3-8b-64k
|
|
```
|
|
|
|
**Embedding (inside worker):**
|
|
```env
|
|
EMBEDDING_API_BASE=http://ollama:11434/v1
|
|
EMBEDDING_MODEL=nomic-embed-text:latest
|
|
EMBEDDING_DIMS=768
|
|
EMBEDDING_API_KEY=
|
|
```
|
|
|
|
**Redis:**
|
|
```env
|
|
REDIS_PASSWORD=<your-password>
|
|
REDIS_HOST=redis
|
|
REDIS_PORT=6379
|
|
```
|
|
|
|
**Qdrant:**
|
|
```env
|
|
QDRANT_HOST=qdrant
|
|
QDRANT_PORT=6333
|
|
COLLECTION_NAME=knowledge_base
|
|
```
|
|
|
|
### 3. Cross-Stack Network
|
|
|
|
If Qdrant/Redis/worker are in one compose stack and Ollama is in another, the worker needs access to both networks:
|
|
|
|
```yaml
|
|
services:
|
|
worker:
|
|
networks:
|
|
- default # memory-os_default — for Redis + Qdrant
|
|
- ollama_default # external — for Ollama DNS
|
|
|
|
networks:
|
|
default:
|
|
name: memory-os_default
|
|
ollama_default:
|
|
external: true
|
|
```
|
|
|
|
> **Critical:** `host.docker.internal` does NOT work on Linux (Docker Desktop only). Use `ollama:11434` (via shared network) or `172.17.0.1:11434` (host gateway) instead.
|
|
|
|
### 4. Verify Connectivity
|
|
|
|
```bash
|
|
# DNS resolution
|
|
docker exec <worker> getent hosts ollama
|
|
|
|
# Ollama API
|
|
docker exec <worker> python3 -c "
|
|
import urllib.request, json
|
|
req = urllib.request.Request('http://ollama:11434/api/tags')
|
|
resp = urllib.request.urlopen(req, timeout=10)
|
|
data = json.loads(resp.read())
|
|
print(f'Models: {len(data[\"models\"])}')
|
|
"
|
|
|
|
# Qdrant collection
|
|
curl -s http://127.0.0.1:6333/collections/knowledge_base | python3 -c "
|
|
import sys,json; d=json.load(sys.stdin)
|
|
print(f'points: {d[\"result\"][\"points_count\"]}')
|
|
"
|
|
```
|
|
|
|
## Search API (FastAPI)
|
|
|
|
Add a search API layer that accepts text queries and returns results from Qdrant:
|
|
|
|
### Docker Compose Service
|
|
|
|
```yaml
|
|
search-api:
|
|
build:
|
|
context: ../search_api # relative to docker/ directory
|
|
dockerfile: Dockerfile
|
|
restart: unless-stopped
|
|
depends_on:
|
|
qdrant:
|
|
condition: service_healthy
|
|
networks:
|
|
- default
|
|
- ollama_default
|
|
environment:
|
|
OLLAMA_URL: http://ollama:11434
|
|
OLLAMA_EMBEDDING_MODEL: nomic-embed-text:latest
|
|
QDRANT_URL: http://qdrant:6333
|
|
COLLECTION_NAME: ${COLLECTION_NAME:-knowledge_base}
|
|
ports:
|
|
- "127.0.0.1:8000:8000"
|
|
healthcheck:
|
|
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/health', timeout=5)"]
|
|
interval: 15s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 10s
|
|
```
|
|
|
|
### FastAPI App Structure
|
|
|
|
```
|
|
search_api/
|
|
├── Dockerfile
|
|
├── requirements.txt # fastapi, uvicorn, httpx, pydantic
|
|
└── main.py
|
|
```
|
|
|
|
### Endpoints
|
|
|
|
- `GET /health` — returns `{"status": "ok", "qdrant": true, "ollama": true}`
|
|
- `POST /search` — accepts `{"query": "...", "top_k": 5}`, returns `{"query": "...", "results": [...], "total": N}`
|
|
|
|
### Flow
|
|
|
|
1. Receive text query → POST to Ollama `/api/embeddings` (nomic-embed-text)
|
|
2. Use returned dense vector → POST to Qdrant `/collections/{name}/points/search`
|
|
3. Return results with score, text, source
|
|
|
|
### Verify
|
|
|
|
```bash
|
|
# Health
|
|
curl http://127.0.0.1:8000/health
|
|
|
|
# Search
|
|
curl -X POST http://127.0.0.1:8000/search \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"query":"your search text","top_k":3}'
|
|
```
|
|
|
|
## Periodic Tasks
|
|
|
|
### Sync + Ingest (every 10 min)
|
|
|
|
Set up a cron job that runs every 10 minutes:
|
|
|
|
1. **Sync script** — copies new/changed `.md` files from Obsidian vault to wiki path, tracking state via JSON file
|
|
2. **Ingest script** — detects new/modified files, enqueues them to ARQ worker for chunking + embedding
|
|
|
|
```bash
|
|
# Manual run
|
|
python3 /path/to/scripts/sync_obsidian_to_wiki.py
|
|
python3 /path/to/scripts/wiki_continuous_ingest.py
|
|
```
|
|
|
|
Via Hermes cronjob (LLM-driven — uses `no_agent: false`):
|
|
```
|
|
hermes cron create \
|
|
--name "memory-os sync+ingest" \
|
|
--schedule "every 10m" \
|
|
--prompt "Run: python3 /path/to/sync_obsidian_to_wiki.py"
|
|
```
|
|
|
|
### Micro-Reflection Trigger (every 5 min, silent)
|
|
|
|
An ARQ worker can have a `process_micro_reflection` function that runs idle-time reflection. To trigger it on a schedule **without LLM overhead**, use a `no_agent: true` watchdog cronjob that runs a script inside the worker container.
|
|
|
|
**Pre-requisite:** Mount the scripts directory into the worker container:
|
|
|
|
```yaml
|
|
services:
|
|
worker:
|
|
volumes:
|
|
- ../scripts:/app/scripts:ro # relative to docker/ directory
|
|
```
|
|
|
|
**Script** (`reflection_trigger.py`): checks if the ARQ worker is idle (no pending/executing jobs), respects a per-hour budget, and enqueues `process_micro_reflection` via Redis.
|
|
|
|
**Cronjob (no_agent, silent, local):**
|
|
```
|
|
hermes cron create \
|
|
--name "memory-os micro-reflection" \
|
|
--schedule "*/5 * * * *" \
|
|
--script "docker exec <worker> python3 /app/scripts/reflection_trigger.py" \
|
|
--no-agent
|
|
hermes cron update \
|
|
--job-id <id> \
|
|
--deliver local
|
|
```
|
|
|
|
Key points:
|
|
- `no_agent: true` — no LLM tokens consumed, just runs the script and delivers stdout verbatim
|
|
- `deliver: local` — suppresses Telegram/Discord notifications; the job runs silently
|
|
- Empty stdout = silent (no message sent), error output = alert delivered
|
|
- The script must be on the host filesystem AND mounted into the container via `volumes:`
|
|
|
|
## Checking Worker Health
|
|
|
|
```bash
|
|
# Container status
|
|
docker ps --filter name=worker
|
|
|
|
# Worker logs
|
|
docker logs <worker> --tail 50
|
|
|
|
# Check for errors
|
|
docker logs <worker> 2>&1 | grep -i "error\|traceback\|exception" | head -10
|
|
|
|
# ARQ stats (from worker logs)
|
|
docker logs <worker> 2>&1 | grep "j_complete\|j_failed"
|
|
```
|
|
|
|
## Pitfalls
|
|
|
|
### `host.docker.internal` on Linux
|
|
`host.docker.internal` is a Docker Desktop feature (macOS/Windows). On Linux, it does not resolve. Use one of:
|
|
- Container name on shared network: `http://ollama:11434`
|
|
- Host gateway: `http://172.17.0.1:11434`
|
|
|
|
### Env vars not propagated to container
|
|
Variables defined in `.env` are NOT automatically available inside containers — they must be explicitly listed in `docker-compose.yml` under `services.worker.environment`. `docker compose config` can verify the effective config.
|
|
|
|
### Redis password mismatch
|
|
If the worker uses `redis.asyncio` or `arq.connections.RedisSettings`, ensure the password matches what's in `redis.conf`. Test with `redis-cli -a $PASSWORD ping`.
|
|
|
|
### Network detachment on recreate
|
|
When a container is recreated via `docker compose up -d --force-recreate`, it may lose connections to external networks. The fix is to declare the network in `docker-compose.yml` with `external: true` and add it to the service's `networks:` list.
|
|
|
|
### Qdrant healthcheck on custom port
|
|
The default Qdrant healthcheck greps `/proc/net/tcp` for `:18BD` (port 6333 in hex). If using a non-standard port, update the healthcheck.
|
|
|
|
### ARQ worker timeout
|
|
The `ollama_chat` function in reflection tasks may timeout if the model is large or generating long responses. Set `ARQ_JOB_TIMEOUT` high enough (e.g., 300s) and ensure `httpx.AsyncClient(timeout=120)` matches.
|
|
|
|
### `no_agent` cron script must be on host filesystem
|
|
A `no_agent: true` cronjob's `--script` runs on the host, not inside the container. If the script only exists inside the container (e.g., at `/app/scripts/`), the cronjob will fail. Mount the scripts directory into the container AND keep the script accessible on the host, or use `docker exec` to run it inside the container:
|
|
|
|
```
|
|
--script "docker exec <container> python3 /app/scripts/script.py"
|
|
```
|
|
|
|
### `reflection_trigger.py` paths hardcoded to old project
|
|
The `reflection_trigger.py` script was originally written for a different project (`~/.ai-stack/`). The `.env` path and log paths must be updated to match the new project layout before the script works after a copy. Search for `Path.home() / "ai-stack"` or similar hardcoded paths and update them to the new project root.
|
|
|
|
### Volume paths in docker-compose are relative to compose file
|
|
When adding a `volumes:` mount like `- ../scripts:/app/scripts:ro`, the path is relative to the `docker-compose.yml` file's directory, not the project root. If the compose file is in `docker/`, then `../scripts` resolves to `project/scripts/`.
|
|
|
|
## Support Files
|
|
|
|
- **`references/memory-os-session.md`** — session-specific details from the Memory OS deployment (env files, state files, error transcripts, search API code)
|
|
- **`scripts/test_qdrant_search.py`** — standalone test script: gets embedding from Ollama, searches Qdrant, prints top-5 results |