mirror of
https://gitverse.ru/kpa39l/rag-pipeline-docker.git
synced 2026-09-29 09:15:11 +00:00
Initial commit: Hermes skill rag-pipeline-docker
This commit is contained in:
@@ -0,0 +1,308 @@
|
||||
---
|
||||
name: rag-pipeline-docker
|
||||
title: RAG Pipeline in Docker Compose
|
||||
category: devops
|
||||
description: Deploying a RAG pipeline with Docker Compose — Qdrant (vector DB), Redis (job queue), ARQ worker (chunking + embedding), connected to an external Ollama instance for embeddings and LLM inference. Covers env var management, cross-network connectivity, cron-based sync+ingest, and troubleshooting.
|
||||
triggers:
|
||||
- qdrant docker
|
||||
- arq worker
|
||||
- rag pipeline deploy
|
||||
- memory-os setup
|
||||
- vector db docker compose
|
||||
- ollama + qdrant integration
|
||||
- wiki ingest pipeline
|
||||
- obsidian sync qdrant
|
||||
---
|
||||
|
||||
# RAG Pipeline in Docker Compose
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
Obsidian vault (host)
|
||||
↓ sync_obsidian_to_wiki.py (cron: every 10m)
|
||||
Wiki path (host)
|
||||
↓ ARQ queue → worker container
|
||||
↓ chunk_text() + get_embedding() + get_sparse_embedding()
|
||||
Qdrant (localhost:6333 / collection: knowledge_base)
|
||||
↓ dense: nomic-embed-text (768d, Cosine)
|
||||
↓ sparse: BM25 (on_disk)
|
||||
```
|
||||
|
||||
## Quick Start
|
||||
|
||||
### 1. Docker Compose Stack
|
||||
|
||||
```yaml
|
||||
services:
|
||||
redis:
|
||||
image: redis:7-alpine
|
||||
restart: unless-stopped
|
||||
# password via REDIS_PASSWORD env
|
||||
healthcheck: [CMD-SHELL, "redis-cli ${REDIS_PASSWORD:+-a $REDIS_PASSWORD} ping"]
|
||||
|
||||
qdrant:
|
||||
image: qdrant/qdrant:v1.17.1
|
||||
restart: unless-stopped
|
||||
ports: ["127.0.0.1:6333:6333"]
|
||||
volumes: [qdrant_data:/qdrant/storage]
|
||||
healthcheck: [CMD, sh, -c, "grep -q ':18BD' /proc/net/tcp"]
|
||||
|
||||
worker:
|
||||
build: ./worker
|
||||
restart: unless-stopped
|
||||
depends_on: [qdrant, redis]
|
||||
# see env section below
|
||||
volumes:
|
||||
- wiki_path:/wiki:ro
|
||||
- hermes_home:/hermes:rw
|
||||
```
|
||||
|
||||
### 2. Environment Variables
|
||||
|
||||
**LLM for reflection / reasoning (inside worker):**
|
||||
```env
|
||||
OLLAMA_BASE_URL=http://ollama:11434
|
||||
OLLAMA_MODEL=qwen3-8b-64k
|
||||
```
|
||||
|
||||
**Embedding (inside worker):**
|
||||
```env
|
||||
EMBEDDING_API_BASE=http://ollama:11434/v1
|
||||
EMBEDDING_MODEL=nomic-embed-text:latest
|
||||
EMBEDDING_DIMS=768
|
||||
EMBEDDING_API_KEY=
|
||||
```
|
||||
|
||||
**Redis:**
|
||||
```env
|
||||
REDIS_PASSWORD=<your-password>
|
||||
REDIS_HOST=redis
|
||||
REDIS_PORT=6379
|
||||
```
|
||||
|
||||
**Qdrant:**
|
||||
```env
|
||||
QDRANT_HOST=qdrant
|
||||
QDRANT_PORT=6333
|
||||
COLLECTION_NAME=knowledge_base
|
||||
```
|
||||
|
||||
### 3. Cross-Stack Network
|
||||
|
||||
If Qdrant/Redis/worker are in one compose stack and Ollama is in another, the worker needs access to both networks:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
worker:
|
||||
networks:
|
||||
- default # memory-os_default — for Redis + Qdrant
|
||||
- ollama_default # external — for Ollama DNS
|
||||
|
||||
networks:
|
||||
default:
|
||||
name: memory-os_default
|
||||
ollama_default:
|
||||
external: true
|
||||
```
|
||||
|
||||
> **Critical:** `host.docker.internal` does NOT work on Linux (Docker Desktop only). Use `ollama:11434` (via shared network) or `172.17.0.1:11434` (host gateway) instead.
|
||||
|
||||
### 4. Verify Connectivity
|
||||
|
||||
```bash
|
||||
# DNS resolution
|
||||
docker exec <worker> getent hosts ollama
|
||||
|
||||
# Ollama API
|
||||
docker exec <worker> python3 -c "
|
||||
import urllib.request, json
|
||||
req = urllib.request.Request('http://ollama:11434/api/tags')
|
||||
resp = urllib.request.urlopen(req, timeout=10)
|
||||
data = json.loads(resp.read())
|
||||
print(f'Models: {len(data[\"models\"])}')
|
||||
"
|
||||
|
||||
# Qdrant collection
|
||||
curl -s http://127.0.0.1:6333/collections/knowledge_base | python3 -c "
|
||||
import sys,json; d=json.load(sys.stdin)
|
||||
print(f'points: {d[\"result\"][\"points_count\"]}')
|
||||
"
|
||||
```
|
||||
|
||||
## Search API (FastAPI)
|
||||
|
||||
Add a search API layer that accepts text queries and returns results from Qdrant:
|
||||
|
||||
### Docker Compose Service
|
||||
|
||||
```yaml
|
||||
search-api:
|
||||
build:
|
||||
context: ../search_api # relative to docker/ directory
|
||||
dockerfile: Dockerfile
|
||||
restart: unless-stopped
|
||||
depends_on:
|
||||
qdrant:
|
||||
condition: service_healthy
|
||||
networks:
|
||||
- default
|
||||
- ollama_default
|
||||
environment:
|
||||
OLLAMA_URL: http://ollama:11434
|
||||
OLLAMA_EMBEDDING_MODEL: nomic-embed-text:latest
|
||||
QDRANT_URL: http://qdrant:6333
|
||||
COLLECTION_NAME: ${COLLECTION_NAME:-knowledge_base}
|
||||
ports:
|
||||
- "127.0.0.1:8000:8000"
|
||||
healthcheck:
|
||||
test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/health', timeout=5)"]
|
||||
interval: 15s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 10s
|
||||
```
|
||||
|
||||
### FastAPI App Structure
|
||||
|
||||
```
|
||||
search_api/
|
||||
├── Dockerfile
|
||||
├── requirements.txt # fastapi, uvicorn, httpx, pydantic
|
||||
└── main.py
|
||||
```
|
||||
|
||||
### Endpoints
|
||||
|
||||
- `GET /health` — returns `{"status": "ok", "qdrant": true, "ollama": true}`
|
||||
- `POST /search` — accepts `{"query": "...", "top_k": 5}`, returns `{"query": "...", "results": [...], "total": N}`
|
||||
|
||||
### Flow
|
||||
|
||||
1. Receive text query → POST to Ollama `/api/embeddings` (nomic-embed-text)
|
||||
2. Use returned dense vector → POST to Qdrant `/collections/{name}/points/search`
|
||||
3. Return results with score, text, source
|
||||
|
||||
### Verify
|
||||
|
||||
```bash
|
||||
# Health
|
||||
curl http://127.0.0.1:8000/health
|
||||
|
||||
# Search
|
||||
curl -X POST http://127.0.0.1:8000/search \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"query":"your search text","top_k":3}'
|
||||
```
|
||||
|
||||
## Periodic Tasks
|
||||
|
||||
### Sync + Ingest (every 10 min)
|
||||
|
||||
Set up a cron job that runs every 10 minutes:
|
||||
|
||||
1. **Sync script** — copies new/changed `.md` files from Obsidian vault to wiki path, tracking state via JSON file
|
||||
2. **Ingest script** — detects new/modified files, enqueues them to ARQ worker for chunking + embedding
|
||||
|
||||
```bash
|
||||
# Manual run
|
||||
python3 /path/to/scripts/sync_obsidian_to_wiki.py
|
||||
python3 /path/to/scripts/wiki_continuous_ingest.py
|
||||
```
|
||||
|
||||
Via Hermes cronjob (LLM-driven — uses `no_agent: false`):
|
||||
```
|
||||
hermes cron create \
|
||||
--name "memory-os sync+ingest" \
|
||||
--schedule "every 10m" \
|
||||
--prompt "Run: python3 /path/to/sync_obsidian_to_wiki.py"
|
||||
```
|
||||
|
||||
### Micro-Reflection Trigger (every 5 min, silent)
|
||||
|
||||
An ARQ worker can have a `process_micro_reflection` function that runs idle-time reflection. To trigger it on a schedule **without LLM overhead**, use a `no_agent: true` watchdog cronjob that runs a script inside the worker container.
|
||||
|
||||
**Pre-requisite:** Mount the scripts directory into the worker container:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
worker:
|
||||
volumes:
|
||||
- ../scripts:/app/scripts:ro # relative to docker/ directory
|
||||
```
|
||||
|
||||
**Script** (`reflection_trigger.py`): checks if the ARQ worker is idle (no pending/executing jobs), respects a per-hour budget, and enqueues `process_micro_reflection` via Redis.
|
||||
|
||||
**Cronjob (no_agent, silent, local):**
|
||||
```
|
||||
hermes cron create \
|
||||
--name "memory-os micro-reflection" \
|
||||
--schedule "*/5 * * * *" \
|
||||
--script "docker exec <worker> python3 /app/scripts/reflection_trigger.py" \
|
||||
--no-agent
|
||||
hermes cron update \
|
||||
--job-id <id> \
|
||||
--deliver local
|
||||
```
|
||||
|
||||
Key points:
|
||||
- `no_agent: true` — no LLM tokens consumed, just runs the script and delivers stdout verbatim
|
||||
- `deliver: local` — suppresses Telegram/Discord notifications; the job runs silently
|
||||
- Empty stdout = silent (no message sent), error output = alert delivered
|
||||
- The script must be on the host filesystem AND mounted into the container via `volumes:`
|
||||
|
||||
## Checking Worker Health
|
||||
|
||||
```bash
|
||||
# Container status
|
||||
docker ps --filter name=worker
|
||||
|
||||
# Worker logs
|
||||
docker logs <worker> --tail 50
|
||||
|
||||
# Check for errors
|
||||
docker logs <worker> 2>&1 | grep -i "error\|traceback\|exception" | head -10
|
||||
|
||||
# ARQ stats (from worker logs)
|
||||
docker logs <worker> 2>&1 | grep "j_complete\|j_failed"
|
||||
```
|
||||
|
||||
## Pitfalls
|
||||
|
||||
### `host.docker.internal` on Linux
|
||||
`host.docker.internal` is a Docker Desktop feature (macOS/Windows). On Linux, it does not resolve. Use one of:
|
||||
- Container name on shared network: `http://ollama:11434`
|
||||
- Host gateway: `http://172.17.0.1:11434`
|
||||
|
||||
### Env vars not propagated to container
|
||||
Variables defined in `.env` are NOT automatically available inside containers — they must be explicitly listed in `docker-compose.yml` under `services.worker.environment`. `docker compose config` can verify the effective config.
|
||||
|
||||
### Redis password mismatch
|
||||
If the worker uses `redis.asyncio` or `arq.connections.RedisSettings`, ensure the password matches what's in `redis.conf`. Test with `redis-cli -a $PASSWORD ping`.
|
||||
|
||||
### Network detachment on recreate
|
||||
When a container is recreated via `docker compose up -d --force-recreate`, it may lose connections to external networks. The fix is to declare the network in `docker-compose.yml` with `external: true` and add it to the service's `networks:` list.
|
||||
|
||||
### Qdrant healthcheck on custom port
|
||||
The default Qdrant healthcheck greps `/proc/net/tcp` for `:18BD` (port 6333 in hex). If using a non-standard port, update the healthcheck.
|
||||
|
||||
### ARQ worker timeout
|
||||
The `ollama_chat` function in reflection tasks may timeout if the model is large or generating long responses. Set `ARQ_JOB_TIMEOUT` high enough (e.g., 300s) and ensure `httpx.AsyncClient(timeout=120)` matches.
|
||||
|
||||
### `no_agent` cron script must be on host filesystem
|
||||
A `no_agent: true` cronjob's `--script` runs on the host, not inside the container. If the script only exists inside the container (e.g., at `/app/scripts/`), the cronjob will fail. Mount the scripts directory into the container AND keep the script accessible on the host, or use `docker exec` to run it inside the container:
|
||||
|
||||
```
|
||||
--script "docker exec <container> python3 /app/scripts/script.py"
|
||||
```
|
||||
|
||||
### `reflection_trigger.py` paths hardcoded to old project
|
||||
The `reflection_trigger.py` script was originally written for a different project (`~/.ai-stack/`). The `.env` path and log paths must be updated to match the new project layout before the script works after a copy. Search for `Path.home() / "ai-stack"` or similar hardcoded paths and update them to the new project root.
|
||||
|
||||
### Volume paths in docker-compose are relative to compose file
|
||||
When adding a `volumes:` mount like `- ../scripts:/app/scripts:ro`, the path is relative to the `docker-compose.yml` file's directory, not the project root. If the compose file is in `docker/`, then `../scripts` resolves to `project/scripts/`.
|
||||
|
||||
## Support Files
|
||||
|
||||
- **`references/memory-os-session.md`** — session-specific details from the Memory OS deployment (env files, state files, error transcripts, search API code)
|
||||
- **`scripts/test_qdrant_search.py`** — standalone test script: gets embedding from Ollama, searches Qdrant, prints top-5 results
|
||||
Reference in New Issue
Block a user