mirror of
https://gitverse.ru/kpa39l/news-aggregation-pipeline.git
synced 2026-09-29 09:15:02 +00:00
Initial commit: Hermes skill news-aggregation-pipeline
This commit is contained in:
@@ -0,0 +1,100 @@
|
|||||||
|
---
|
||||||
|
name: news-aggregation-pipeline
|
||||||
|
description: "Use when building news/RSS aggregation pipelines."
|
||||||
|
---
|
||||||
|
|
||||||
|
# News Aggregation Pipeline
|
||||||
|
|
||||||
|
Build domain news aggregators / news DBs with custom crawlers (RSS + HTML),
|
||||||
|
LLM classification, and vector search. Covers the /opt/news project (IT +
|
||||||
|
hospitality IT news) and any similar future domain pipeline.
|
||||||
|
|
||||||
|
## Current project (source of truth)
|
||||||
|
|
||||||
|
- PRD: `/opt/news/PRD.md` — architecture, DB schema, taxonomy, phases 0–3,
|
||||||
|
Definition of Done, open questions. READ IT FIRST when working on /opt/news.
|
||||||
|
- Stack decisions: Python 3.12 venv, feedparser/httpx/selectolax/trafilatura,
|
||||||
|
SQLite canonical DB, Qdrant collection `news`, Ollama (bge-m3 embeddings,
|
||||||
|
qwen3:8b classifier), Hermes cron, Telegram for digests.
|
||||||
|
- Status: phase 0 (PRD done; next: sources.yaml ~30 sources, validate_sources.py,
|
||||||
|
keywords.yaml, git init on gitea).
|
||||||
|
|
||||||
|
## Architecture principles
|
||||||
|
|
||||||
|
- **Canonical DB = SQLite**, vector index (Qdrant) = derived layer. SQLite holds
|
||||||
|
articles, dedup keys, crawler state (etag/modified/last_error/status).
|
||||||
|
- **Source registry = single YAML in git** (`sources.yaml`): slug, name, url,
|
||||||
|
feed_url (null → html_crawler), crawler, domain, lang, priority (P0/P1/P2 →
|
||||||
|
cadence 15min/1h/4h), tags, scrape_selector, enabled.
|
||||||
|
- **Never physically delete sources** — disable with `enabled: false`
|
||||||
|
(user rule: no deleting data). Registry is the only source of truth; SQLite
|
||||||
|
and statuses are derived.
|
||||||
|
- **Feed health**: every fetch writes last_fetch/last_error/status; after N=5
|
||||||
|
consecutive errors mark `dead`; watchdog alerts ONLY on alive↔dead transitions
|
||||||
|
(user hates routine noise).
|
||||||
|
- **Domain taxonomy with crossover**: `it_general`, `it_hospitality`, `other`
|
||||||
|
(dropped or relevance=low); IT news with clear hospitality application gets
|
||||||
|
both domains (e.g. cyber breach of hotel chain).
|
||||||
|
|
||||||
|
## Crawler patterns
|
||||||
|
|
||||||
|
### rss_crawler (~80% of volume)
|
||||||
|
1. Read sources.yaml, group by priority.
|
||||||
|
2. GET with `If-None-Match`/`If-Modified-Since` from stored etag/modified →
|
||||||
|
304 = skip (polite + cheap).
|
||||||
|
3. Normalize URL (strip UTM params, anchors) → `canonical_url` = dedup key.
|
||||||
|
4. If feed has summary-only entries, fetch full text via html fetch.
|
||||||
|
5. Dedup: `sha256(canonical_url)` unique in SQLite; existing → skip.
|
||||||
|
6. hashlib sha256, python-dateutil for published_at.
|
||||||
|
|
||||||
|
### html_crawler (~15%)
|
||||||
|
- httpx + selectolax (parsing) + trafilatura (text extraction).
|
||||||
|
- `scrape_selector` per source (e.g. `article h2 a`), auto-link discovery fallback.
|
||||||
|
- Rate-limit 1–2s, honest User-Agent (`news-crawler/0.1 (+https://nixg.ru)`),
|
||||||
|
simple robots.txt cache; robots denies → skip source.
|
||||||
|
|
||||||
|
### api_crawler (phase 2)
|
||||||
|
- Telegram channels via MTProto gateway (slidge/tg), HN/Reddit APIs.
|
||||||
|
|
||||||
|
## LLM classification (qwen3:8b via Ollama)
|
||||||
|
|
||||||
|
- Fast keyword pre-filter (keywords.yaml, hard words per domain; ≥1 hard word →
|
||||||
|
candidate), then LLM for final domain/topics/relevance/summary.
|
||||||
|
- Prompt: title + first ~800 chars; output strict JSON:
|
||||||
|
`domain, topics, keywords, relevance (critical/high/low), summary`.
|
||||||
|
- **PITFALL: qwen3:8b requires `extra_body {"think": false}`** — otherwise it
|
||||||
|
emits ~3778 tokens of thinking with empty content (known quirk, applies to
|
||||||
|
all text aux calls on this box).
|
||||||
|
- critical/high → Qdrant index; low → SQLite only.
|
||||||
|
|
||||||
|
## Indexing
|
||||||
|
|
||||||
|
- Embeddings via Ollama: bge-m3 (1024d, already wired in memory-os search-api)
|
||||||
|
or nomic-embed-text (768d). Verify model present in Ollama before first run.
|
||||||
|
- Qdrant collection `news` — SEPARATE from memory-os `knowledge_base`.
|
||||||
|
- Dense + sparse BM25 payload; payload: article_id, url, title, published_at,
|
||||||
|
domain, topics, source_slug.
|
||||||
|
- Semantic dedup (phase 2): cosine > 0.95 against existing point → duplicate.
|
||||||
|
|
||||||
|
## Infra reuse on this box
|
||||||
|
|
||||||
|
- Qdrant docker-qdrant-1 :6333, Redis docker-redis-1 :6379, ARQ worker,
|
||||||
|
Ollama :11434 (localhost), Hermes cron, Telegram gateway, gitea.nixg.ru.
|
||||||
|
- If sharing Redis with memory-os, use separate namespace `news:*`.
|
||||||
|
- `host.docker.internal` does NOT work on Linux — use service names on shared
|
||||||
|
networks or `172.17.0.1`.
|
||||||
|
|
||||||
|
## Scheduling (user's watchdog preferences)
|
||||||
|
|
||||||
|
- Mechanical crawler runs: plain scripts via Hermes cron (no LLM).
|
||||||
|
- Watchdogs: `no_agent: true`, `deliver: local`, empty stdout = silent;
|
||||||
|
script must stay quiet unless a real threshold/state change fires (e.g.
|
||||||
|
check_icq_cert.sh pattern: only speaks at <=45 days).
|
||||||
|
- Daily digest: LLM-driven cron, delivery to Telegram.
|
||||||
|
|
||||||
|
## Definition of Done (from PRD)
|
||||||
|
|
||||||
|
- ≥100 articles/day with zero dupes (count(sha256) == count(id)).
|
||||||
|
- Hand-check 20 random articles: ≥90% correct domain.
|
||||||
|
- Qdrant returns sensible hits for 10 control queries.
|
||||||
|
- Watchdog silent for 3 consecutive days when healthy.
|
||||||
Reference in New Issue
Block a user