Files
news-aggregation-pipeline/SKILL.md
T
2026-09-06 13:51:23 +00:00

100 lines
4.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: news-aggregation-pipeline
description: "Use when building news/RSS aggregation pipelines."
---
# News Aggregation Pipeline
Build domain news aggregators / news DBs with custom crawlers (RSS + HTML),
LLM classification, and vector search. Covers the /opt/news project (IT +
hospitality IT news) and any similar future domain pipeline.
## Current project (source of truth)
- PRD: `/opt/news/PRD.md` — architecture, DB schema, taxonomy, phases 0–3,
Definition of Done, open questions. READ IT FIRST when working on /opt/news.
- Stack decisions: Python 3.12 venv, feedparser/httpx/selectolax/trafilatura,
SQLite canonical DB, Qdrant collection `news`, Ollama (bge-m3 embeddings,
qwen3:8b classifier), Hermes cron, Telegram for digests.
- Status: phase 0 (PRD done; next: sources.yaml ~30 sources, validate_sources.py,
keywords.yaml, git init on gitea).
## Architecture principles
- **Canonical DB = SQLite**, vector index (Qdrant) = derived layer. SQLite holds
articles, dedup keys, crawler state (etag/modified/last_error/status).
- **Source registry = single YAML in git** (`sources.yaml`): slug, name, url,
feed_url (null → html_crawler), crawler, domain, lang, priority (P0/P1/P2 →
cadence 15min/1h/4h), tags, scrape_selector, enabled.
- **Never physically delete sources** — disable with `enabled: false`
(user rule: no deleting data). Registry is the only source of truth; SQLite
and statuses are derived.
- **Feed health**: every fetch writes last_fetch/last_error/status; after N=5
consecutive errors mark `dead`; watchdog alerts ONLY on alive↔dead transitions
(user hates routine noise).
- **Domain taxonomy with crossover**: `it_general`, `it_hospitality`, `other`
(dropped or relevance=low); IT news with clear hospitality application gets
both domains (e.g. cyber breach of hotel chain).
## Crawler patterns
### rss_crawler (~80% of volume)
1. Read sources.yaml, group by priority.
2. GET with `If-None-Match`/`If-Modified-Since` from stored etag/modified →
304 = skip (polite + cheap).
3. Normalize URL (strip UTM params, anchors) → `canonical_url` = dedup key.
4. If feed has summary-only entries, fetch full text via html fetch.
5. Dedup: `sha256(canonical_url)` unique in SQLite; existing → skip.
6. hashlib sha256, python-dateutil for published_at.
### html_crawler (~15%)
- httpx + selectolax (parsing) + trafilatura (text extraction).
- `scrape_selector` per source (e.g. `article h2 a`), auto-link discovery fallback.
- Rate-limit 1–2s, honest User-Agent (`news-crawler/0.1 (+https://nixg.ru)`),
simple robots.txt cache; robots denies → skip source.
### api_crawler (phase 2)
- Telegram channels via MTProto gateway (slidge/tg), HN/Reddit APIs.
## LLM classification (qwen3:8b via Ollama)
- Fast keyword pre-filter (keywords.yaml, hard words per domain; ≥1 hard word →
candidate), then LLM for final domain/topics/relevance/summary.
- Prompt: title + first ~800 chars; output strict JSON:
`domain, topics, keywords, relevance (critical/high/low), summary`.
- **PITFALL: qwen3:8b requires `extra_body {"think": false}`** — otherwise it
emits ~3778 tokens of thinking with empty content (known quirk, applies to
all text aux calls on this box).
- critical/high → Qdrant index; low → SQLite only.
## Indexing
- Embeddings via Ollama: bge-m3 (1024d, already wired in memory-os search-api)
or nomic-embed-text (768d). Verify model present in Ollama before first run.
- Qdrant collection `news` — SEPARATE from memory-os `knowledge_base`.
- Dense + sparse BM25 payload; payload: article_id, url, title, published_at,
domain, topics, source_slug.
- Semantic dedup (phase 2): cosine > 0.95 against existing point → duplicate.
## Infra reuse on this box
- Qdrant docker-qdrant-1 :6333, Redis docker-redis-1 :6379, ARQ worker,
Ollama :11434 (localhost), Hermes cron, Telegram gateway, gitea.nixg.ru.
- If sharing Redis with memory-os, use separate namespace `news:*`.
- `host.docker.internal` does NOT work on Linux — use service names on shared
networks or `172.17.0.1`.
## Scheduling (user's watchdog preferences)
- Mechanical crawler runs: plain scripts via Hermes cron (no LLM).
- Watchdogs: `no_agent: true`, `deliver: local`, empty stdout = silent;
script must stay quiet unless a real threshold/state change fires (e.g.
check_icq_cert.sh pattern: only speaks at <=45 days).
- Daily digest: LLM-driven cron, delivery to Telegram.
## Definition of Done (from PRD)
- ≥100 articles/day with zero dupes (count(sha256) == count(id)).
- Hand-check 20 random articles: ≥90% correct domain.
- Qdrant returns sensible hits for 10 control queries.
- Watchdog silent for 3 consecutive days when healthy.