From 9f5aad905f1c97ba9fe957859a769c09ff4885aa Mon Sep 17 00:00:00 2001 From: estorozhenko Date: Sun, 6 Sep 2026 13:51:23 +0000 Subject: [PATCH] Initial commit: Hermes skill news-aggregation-pipeline --- SKILL.md | 100 +++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 100 insertions(+) create mode 100644 SKILL.md diff --git a/SKILL.md b/SKILL.md new file mode 100644 index 0000000..7e1f02c --- /dev/null +++ b/SKILL.md @@ -0,0 +1,100 @@ +--- +name: news-aggregation-pipeline +description: "Use when building news/RSS aggregation pipelines." +--- + +# News Aggregation Pipeline + +Build domain news aggregators / news DBs with custom crawlers (RSS + HTML), +LLM classification, and vector search. Covers the /opt/news project (IT + +hospitality IT news) and any similar future domain pipeline. + +## Current project (source of truth) + +- PRD: `/opt/news/PRD.md` — architecture, DB schema, taxonomy, phases 0–3, + Definition of Done, open questions. READ IT FIRST when working on /opt/news. +- Stack decisions: Python 3.12 venv, feedparser/httpx/selectolax/trafilatura, + SQLite canonical DB, Qdrant collection `news`, Ollama (bge-m3 embeddings, + qwen3:8b classifier), Hermes cron, Telegram for digests. +- Status: phase 0 (PRD done; next: sources.yaml ~30 sources, validate_sources.py, + keywords.yaml, git init on gitea). + +## Architecture principles + +- **Canonical DB = SQLite**, vector index (Qdrant) = derived layer. SQLite holds + articles, dedup keys, crawler state (etag/modified/last_error/status). +- **Source registry = single YAML in git** (`sources.yaml`): slug, name, url, + feed_url (null → html_crawler), crawler, domain, lang, priority (P0/P1/P2 → + cadence 15min/1h/4h), tags, scrape_selector, enabled. +- **Never physically delete sources** — disable with `enabled: false` + (user rule: no deleting data). Registry is the only source of truth; SQLite + and statuses are derived. +- **Feed health**: every fetch writes last_fetch/last_error/status; after N=5 + consecutive errors mark `dead`; watchdog alerts ONLY on alive↔dead transitions + (user hates routine noise). +- **Domain taxonomy with crossover**: `it_general`, `it_hospitality`, `other` + (dropped or relevance=low); IT news with clear hospitality application gets + both domains (e.g. cyber breach of hotel chain). + +## Crawler patterns + +### rss_crawler (~80% of volume) +1. Read sources.yaml, group by priority. +2. GET with `If-None-Match`/`If-Modified-Since` from stored etag/modified → + 304 = skip (polite + cheap). +3. Normalize URL (strip UTM params, anchors) → `canonical_url` = dedup key. +4. If feed has summary-only entries, fetch full text via html fetch. +5. Dedup: `sha256(canonical_url)` unique in SQLite; existing → skip. +6. hashlib sha256, python-dateutil for published_at. + +### html_crawler (~15%) +- httpx + selectolax (parsing) + trafilatura (text extraction). +- `scrape_selector` per source (e.g. `article h2 a`), auto-link discovery fallback. +- Rate-limit 1–2s, honest User-Agent (`news-crawler/0.1 (+https://nixg.ru)`), + simple robots.txt cache; robots denies → skip source. + +### api_crawler (phase 2) +- Telegram channels via MTProto gateway (slidge/tg), HN/Reddit APIs. + +## LLM classification (qwen3:8b via Ollama) + +- Fast keyword pre-filter (keywords.yaml, hard words per domain; ≥1 hard word → + candidate), then LLM for final domain/topics/relevance/summary. +- Prompt: title + first ~800 chars; output strict JSON: + `domain, topics, keywords, relevance (critical/high/low), summary`. +- **PITFALL: qwen3:8b requires `extra_body {"think": false}`** — otherwise it + emits ~3778 tokens of thinking with empty content (known quirk, applies to + all text aux calls on this box). +- critical/high → Qdrant index; low → SQLite only. + +## Indexing + +- Embeddings via Ollama: bge-m3 (1024d, already wired in memory-os search-api) + or nomic-embed-text (768d). Verify model present in Ollama before first run. +- Qdrant collection `news` — SEPARATE from memory-os `knowledge_base`. +- Dense + sparse BM25 payload; payload: article_id, url, title, published_at, + domain, topics, source_slug. +- Semantic dedup (phase 2): cosine > 0.95 against existing point → duplicate. + +## Infra reuse on this box + +- Qdrant docker-qdrant-1 :6333, Redis docker-redis-1 :6379, ARQ worker, + Ollama :11434 (localhost), Hermes cron, Telegram gateway, gitea.nixg.ru. +- If sharing Redis with memory-os, use separate namespace `news:*`. +- `host.docker.internal` does NOT work on Linux — use service names on shared + networks or `172.17.0.1`. + +## Scheduling (user's watchdog preferences) + +- Mechanical crawler runs: plain scripts via Hermes cron (no LLM). +- Watchdogs: `no_agent: true`, `deliver: local`, empty stdout = silent; + script must stay quiet unless a real threshold/state change fires (e.g. + check_icq_cert.sh pattern: only speaks at <=45 days). +- Daily digest: LLM-driven cron, delivery to Telegram. + +## Definition of Done (from PRD) + +- ≥100 articles/day with zero dupes (count(sha256) == count(id)). +- Hand-check 20 random articles: ≥90% correct domain. +- Qdrant returns sensible hits for 10 control queries. +- Watchdog silent for 3 consecutive days when healthy. \ No newline at end of file