Files
2026-09-06 13:51:23 +00:00

4.6 KiB
Raw Permalink Blame History

name, description
name description
news-aggregation-pipeline Use when building news/RSS aggregation pipelines.

News Aggregation Pipeline

Build domain news aggregators / news DBs with custom crawlers (RSS + HTML), LLM classification, and vector search. Covers the /opt/news project (IT + hospitality IT news) and any similar future domain pipeline.

Current project (source of truth)

  • PRD: /opt/news/PRD.md — architecture, DB schema, taxonomy, phases 0–3, Definition of Done, open questions. READ IT FIRST when working on /opt/news.
  • Stack decisions: Python 3.12 venv, feedparser/httpx/selectolax/trafilatura, SQLite canonical DB, Qdrant collection news, Ollama (bge-m3 embeddings, qwen3:8b classifier), Hermes cron, Telegram for digests.
  • Status: phase 0 (PRD done; next: sources.yaml ~30 sources, validate_sources.py, keywords.yaml, git init on gitea).

Architecture principles

  • Canonical DB = SQLite, vector index (Qdrant) = derived layer. SQLite holds articles, dedup keys, crawler state (etag/modified/last_error/status).
  • Source registry = single YAML in git (sources.yaml): slug, name, url, feed_url (null → html_crawler), crawler, domain, lang, priority (P0/P1/P2 → cadence 15min/1h/4h), tags, scrape_selector, enabled.
  • Never physically delete sources — disable with enabled: false (user rule: no deleting data). Registry is the only source of truth; SQLite and statuses are derived.
  • Feed health: every fetch writes last_fetch/last_error/status; after N=5 consecutive errors mark dead; watchdog alerts ONLY on alive↔dead transitions (user hates routine noise).
  • Domain taxonomy with crossover: it_general, it_hospitality, other (dropped or relevance=low); IT news with clear hospitality application gets both domains (e.g. cyber breach of hotel chain).

Crawler patterns

rss_crawler (~80% of volume)

  1. Read sources.yaml, group by priority.
  2. GET with If-None-Match/If-Modified-Since from stored etag/modified → 304 = skip (polite + cheap).
  3. Normalize URL (strip UTM params, anchors) → canonical_url = dedup key.
  4. If feed has summary-only entries, fetch full text via html fetch.
  5. Dedup: sha256(canonical_url) unique in SQLite; existing → skip.
  6. hashlib sha256, python-dateutil for published_at.

html_crawler (~15%)

  • httpx + selectolax (parsing) + trafilatura (text extraction).
  • scrape_selector per source (e.g. article h2 a), auto-link discovery fallback.
  • Rate-limit 1–2s, honest User-Agent (news-crawler/0.1 (+https://nixg.ru)), simple robots.txt cache; robots denies → skip source.

api_crawler (phase 2)

  • Telegram channels via MTProto gateway (slidge/tg), HN/Reddit APIs.

LLM classification (qwen3:8b via Ollama)

  • Fast keyword pre-filter (keywords.yaml, hard words per domain; ≥1 hard word → candidate), then LLM for final domain/topics/relevance/summary.
  • Prompt: title + first ~800 chars; output strict JSON: domain, topics, keywords, relevance (critical/high/low), summary.
  • PITFALL: qwen3:8b requires extra_body {"think": false} — otherwise it emits ~3778 tokens of thinking with empty content (known quirk, applies to all text aux calls on this box).
  • critical/high → Qdrant index; low → SQLite only.

Indexing

  • Embeddings via Ollama: bge-m3 (1024d, already wired in memory-os search-api) or nomic-embed-text (768d). Verify model present in Ollama before first run.
  • Qdrant collection news — SEPARATE from memory-os knowledge_base.
  • Dense + sparse BM25 payload; payload: article_id, url, title, published_at, domain, topics, source_slug.
  • Semantic dedup (phase 2): cosine > 0.95 against existing point → duplicate.

Infra reuse on this box

  • Qdrant docker-qdrant-1 :6333, Redis docker-redis-1 :6379, ARQ worker, Ollama :11434 (localhost), Hermes cron, Telegram gateway, gitea.nixg.ru.
  • If sharing Redis with memory-os, use separate namespace news:*.
  • host.docker.internal does NOT work on Linux — use service names on shared networks or 172.17.0.1.

Scheduling (user's watchdog preferences)

  • Mechanical crawler runs: plain scripts via Hermes cron (no LLM).
  • Watchdogs: no_agent: true, deliver: local, empty stdout = silent; script must stay quiet unless a real threshold/state change fires (e.g. check_icq_cert.sh pattern: only speaks at <=45 days).
  • Daily digest: LLM-driven cron, delivery to Telegram.

Definition of Done (from PRD)

  • ≥100 articles/day with zero dupes (count(sha256) == count(id)).
  • Hand-check 20 random articles: ≥90% correct domain.
  • Qdrant returns sensible hits for 10 control queries.
  • Watchdog silent for 3 consecutive days when healthy.