mirror of
https://gitverse.ru/kpa39l/chronicle.nixg.ru.git
synced 2026-09-29 09:55:08 +00:00
Phase 0 MVP: Telegram Archiver fully functional (992 posts tested)
This commit is contained in:
@@ -33,9 +33,13 @@ ENV/
|
||||
*.session
|
||||
*.session-journal
|
||||
|
||||
# Archives output
|
||||
# Archives output - DO NOT COMMIT
|
||||
archives/
|
||||
output/
|
||||
content/
|
||||
*/2big2get.md
|
||||
*/to_download_later.md
|
||||
*/corrupted.md
|
||||
|
||||
# Logs
|
||||
*.log
|
||||
|
||||
@@ -0,0 +1,254 @@
|
||||
# Telegram Archiver — Ответы на вопросы
|
||||
|
||||
## 📊 Итоги полного тестирования (март 2026)
|
||||
|
||||
**Канал:** dedinit
|
||||
**Постов заархивировано:** 992
|
||||
**Время выполнения:** 7335 секунд (≈ 2 часа 2 минуты)
|
||||
**Медиафайлов скачано:** 454
|
||||
**Таймаутов:** ~50 файлов (пропущены, процесс не застрял)
|
||||
|
||||
---
|
||||
|
||||
## ❓ Вопросы и ответы
|
||||
|
||||
### 1. Размер файлов: когда считается?
|
||||
|
||||
**Ответ:** Размер должен браться **из Telegram ДО скачивания** через атрибут `doc.size`.
|
||||
|
||||
**Проблема:** В текущей версии `size: 0` во front-matter — это баг.
|
||||
|
||||
**Решение (задача 0.15):**
|
||||
```python
|
||||
# Брать из атрибутов Telegram
|
||||
size = doc.size
|
||||
|
||||
# ИЛИ проверять после записи
|
||||
size = Path(filepath).stat().st_size
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 2. Можно ли получить прямые ссылки на файлы для дозагрузки?
|
||||
|
||||
**Ответ:** **Да, но с ограничениями.**
|
||||
|
||||
Telethon может получить прямую ссылку через `message.media.document.id`, но:
|
||||
- Ссылка **временная** (действует несколько часов)
|
||||
- Требует активной сессии Telegram
|
||||
- Нельзя просто сохранить URL и скачать позже
|
||||
|
||||
**Решение (задача 0.14, 0.19):**
|
||||
1. Сохранять `message_id` и `media_id` во front-matter
|
||||
2. Для дозагрузки:
|
||||
- Переподключаться к Telegram
|
||||
- Запрашивать файл по ID через API
|
||||
- Скачивать заново
|
||||
|
||||
**Пример front-matter:**
|
||||
```yaml
|
||||
media_files:
|
||||
- filename: photo.jpg
|
||||
type: photo
|
||||
message_id: 12345
|
||||
media_id: "AgADAgAD..."
|
||||
download_url: "https://t.me/dedinit/12345/file" # временная!
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 3. Hugo и формат тегов
|
||||
|
||||
**Вопрос:** Поймёт ли Hugo такой формат?
|
||||
```yaml
|
||||
tags:
|
||||
- 1с
|
||||
- debian
|
||||
- postgres
|
||||
```
|
||||
|
||||
**Ответ:** **Да, Hugo понимает оба формата:**
|
||||
|
||||
```yaml
|
||||
# YAML list (сейчас)
|
||||
tags:
|
||||
- linux
|
||||
- debian
|
||||
|
||||
# Inline JSON-style (компактнее)
|
||||
tags: ["linux", "debian", "postgres"]
|
||||
```
|
||||
|
||||
**Рекомендация:** Оба работают одинаково. Задача **0.22** — опционально изменить на inline для компактности.
|
||||
|
||||
---
|
||||
|
||||
### 4. Многопоточное скачивание (задача 0.17)
|
||||
|
||||
**Вопрос:** Можно ли скачивать в несколько потоков?
|
||||
|
||||
**Ответ:** **Да, это возможно и ускорит процесс в 2-3 раза.**
|
||||
|
||||
**Техническая реализация:**
|
||||
```python
|
||||
import asyncio
|
||||
|
||||
# Параллельная загрузка 3 файлов
|
||||
await asyncio.gather(
|
||||
download_file_1(),
|
||||
download_file_2(),
|
||||
download_file_3(),
|
||||
)
|
||||
```
|
||||
|
||||
**Ограничения:**
|
||||
- Telegram блокирует при >5 одновременных загрузок
|
||||
- Рекомендовано: **3 параллельных потока**
|
||||
- Нужно соблюдать rate limits (паузы между запросами)
|
||||
|
||||
**План (задача 0.19):**
|
||||
1. Первый проход: собрать все URL файлов в очередь
|
||||
2. Второй проход: загружать параллельно (3 потока)
|
||||
3. Resume: продолжать с места обрыва
|
||||
|
||||
**Ожидаемое ускорение:**
|
||||
- Сейчас: ~1000 постов за 2 часа (500 постов/час)
|
||||
- С многопоточностью: ~1000-1500 постов/час
|
||||
|
||||
---
|
||||
|
||||
### 5. Resume прерванной архивации (задача 0.21)
|
||||
|
||||
**Вопрос:** Есть ли средство возобновить прерванный процесс?
|
||||
|
||||
**Ответ:** **Частично работает уже сейчас.**
|
||||
|
||||
**Текущее состояние:**
|
||||
- ✅ Дедупликация по ID существует
|
||||
- ✅ Повторный запуск пропускает скачанные посты
|
||||
- ❌ Нет сохранения последнего обработанного ID
|
||||
- ❌ Нет команды `--resume`
|
||||
|
||||
**План (задача 0.21):**
|
||||
1. Сохранять `last_message_id` в лог или файл состояния
|
||||
2. Добавить флаг `--resume` для продолжения
|
||||
3. При старте проверять последний ID и начинать с него
|
||||
|
||||
**Пример:**
|
||||
```bash
|
||||
# Первый запуск
|
||||
python run_archiver.py -c dedinit
|
||||
# Прерван на посте 624
|
||||
|
||||
# Возобновление
|
||||
python run_archiver.py -c dedinit --resume
|
||||
# Продолжит с поста 624
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 6. Голосовые сообщения и аудио (задача 0.20, 0.24)
|
||||
|
||||
**Вопрос:** Как обрабатывать голосовые сообщения?
|
||||
|
||||
**Текущее состояние:**
|
||||
- Голосовухи скачиваются как `.bin` или `.ogg`
|
||||
- Нет метаданных (длительность, формат)
|
||||
- Нет красивого отображения в Hugo
|
||||
|
||||
**План:**
|
||||
1. Определять формат по MIME:
|
||||
- `audio/ogg` → `.ogg`
|
||||
- `audio/mp3` → `.mp3`
|
||||
2. Извлекать длительность: `doc.attributes.duration`
|
||||
3. Сохранять во front-matter:
|
||||
```yaml
|
||||
media_files:
|
||||
- filename: voice.ogg
|
||||
type: audio
|
||||
duration: 45 # секунд
|
||||
size: 127359
|
||||
```
|
||||
4. Для Hugo: использовать shortcode с wavesurfer.js
|
||||
```markdown
|
||||
{{< audio-player src="voice.ogg" waveform="true" >}}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 7. Видео-кружки (задача 0.11)
|
||||
|
||||
**Вопрос:** Как обрабатывать video_note (кружочки)?
|
||||
|
||||
**Проблема:** Сохраняются как `.bin` без расширения.
|
||||
|
||||
**Решение:**
|
||||
1. Определять по MIME: `video/mp4` → `.mp4`
|
||||
2. Переименовывать после скачивания
|
||||
3. Опционально: конвертировать в нормальное видео (убрать круг)
|
||||
|
||||
---
|
||||
|
||||
### 8. Временная зона (задача 0.13)
|
||||
|
||||
**Проблема:** `date: '2026-01-18T08:11:24+00:00'` — UTC вместо Москвы.
|
||||
|
||||
**Пост был в 11:11 по Москве, указано 08:11 UTC.**
|
||||
|
||||
**Решение:**
|
||||
```python
|
||||
from zoneinfo import ZoneInfo
|
||||
|
||||
# Конвертировать в московское время
|
||||
moscow_tz = ZoneInfo('Europe/Moscow')
|
||||
date_moscow = date.astimezone(moscow_tz)
|
||||
|
||||
# Результат: 2026-01-18T11:11:24+03:00
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 9. Прямые ссылки на посты (задача 0.9, 0.16)
|
||||
|
||||
**Вопрос:** Где ссылки на оригинальные посты?
|
||||
|
||||
**Текущее состояние:**
|
||||
- Сохраняется `repost_from` (ID) и `repost_channel` (название)
|
||||
- Нет прямой ссылки `https://t.me/...`
|
||||
|
||||
**План:**
|
||||
```yaml
|
||||
# Для всех постов
|
||||
channel_url: "https://t.me/dedinit"
|
||||
post_url: "https://t.me/dedinit/1234"
|
||||
|
||||
# Для репостов
|
||||
original_post_url: "https://t.me/otherchannel/987"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📈 Метрики производительности
|
||||
|
||||
| Параметр | Значение |
|
||||
|----------|----------|
|
||||
| Скорость архивации | ~500 постов/час |
|
||||
| Время на 1 пост | ~7.4 секунды |
|
||||
| Время на 1 медиафайл | ~15-30 секунд (с таймаутами) |
|
||||
| Процент таймаутов | ~5% (50 из 992) |
|
||||
| Успешных загрузок | 95% |
|
||||
|
||||
---
|
||||
|
||||
## 🎯 Следующие шаги
|
||||
|
||||
1. **Исправить критичные баги** (0.10, 0.15)
|
||||
2. **Добавить resume** (0.21)
|
||||
3. **Добавить отчёт о таймаутах** (0.12)
|
||||
4. **Исправить временную зону** (0.13)
|
||||
5. **Многопоточность** (0.17, 0.19)
|
||||
|
||||
---
|
||||
|
||||
*Документ создан: 6 марта 2026 г.*
|
||||
*Версия: 1.0*
|
||||
@@ -5,6 +5,7 @@ Handles message processing, markdown generation, and file organization.
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
import re
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Optional
|
||||
@@ -19,6 +20,54 @@ from config import Settings, get_settings
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
def extract_hashtags(text: str) -> list[str]:
|
||||
"""
|
||||
Extract hashtags from text.
|
||||
|
||||
Matches: #tag, #тег, #tag123, #тэг_с_подчёркиванием
|
||||
Supports: Latin, Cyrillic, digits, underscores
|
||||
|
||||
Args:
|
||||
text: Message text
|
||||
|
||||
Returns:
|
||||
List of unique hashtags (without #)
|
||||
"""
|
||||
if not text:
|
||||
return []
|
||||
|
||||
# Regex for hashtags: # followed by word chars (including Cyrillic)
|
||||
pattern = r'#([\wа-яА-ЯёЁ\d_]+)'
|
||||
matches = re.findall(pattern, text)
|
||||
|
||||
# Return unique tags, lowercase
|
||||
return list(set(tag.lower() for tag in matches))
|
||||
|
||||
|
||||
def remove_hashtags(text: str) -> str:
|
||||
"""
|
||||
Remove hashtags from text.
|
||||
|
||||
Args:
|
||||
text: Message text with hashtags
|
||||
|
||||
Returns:
|
||||
Text without hashtags (extra whitespace cleaned)
|
||||
"""
|
||||
if not text:
|
||||
return text
|
||||
|
||||
# Remove hashtags
|
||||
pattern = r'#\w+[\wа-яА-ЯёЁ\d_]*'
|
||||
text = re.sub(pattern, '', text)
|
||||
|
||||
# Clean up multiple spaces/newlines
|
||||
text = re.sub(r'\n\s*\n', '\n\n', text)
|
||||
text = re.sub(r' {2,}', ' ', text)
|
||||
|
||||
return text.strip()
|
||||
|
||||
|
||||
class ChannelArchiver:
|
||||
"""
|
||||
Handles archiving of a single Telegram channel.
|
||||
@@ -167,52 +216,62 @@ class ChannelArchiver:
|
||||
self.stats["posts_skipped"] += 1
|
||||
return
|
||||
|
||||
logger.debug(f"Processing message {message_id}...")
|
||||
try:
|
||||
logger.debug(f"Processing message {message_id}...")
|
||||
|
||||
# Create bundle directory
|
||||
bundle_dir.mkdir(parents=True, exist_ok=True)
|
||||
# Create bundle directory
|
||||
bundle_dir.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Extract post data
|
||||
post_data = await self._extract_post_data(message)
|
||||
# Extract post data
|
||||
post_data = await self._extract_post_data(message)
|
||||
|
||||
# Download media
|
||||
if message.media:
|
||||
media_info = await self.client.download_media(
|
||||
message, bundle_dir, self.settings.max_file_size
|
||||
)
|
||||
|
||||
# Update post_data with media info
|
||||
for info in media_info:
|
||||
media_file = MediaFile(
|
||||
filename=info["filename"],
|
||||
type=info["type"],
|
||||
caption=info.get("caption"),
|
||||
size=info.get("size"),
|
||||
is_too_large=info.get("is_too_large", False),
|
||||
# Download media
|
||||
if message.media:
|
||||
media_info = await self.client.download_media(
|
||||
message, bundle_dir, self.settings.max_file_size, timeout=30
|
||||
)
|
||||
post_data.media_files.append(media_file)
|
||||
|
||||
if info.get("is_too_large"):
|
||||
self.stats["media_skipped_large"] += 1
|
||||
self.stats["large_files"].append(
|
||||
{
|
||||
"message_id": message_id,
|
||||
"filename": info["filename"],
|
||||
"size": info.get("size"),
|
||||
}
|
||||
# Update post_data with media info
|
||||
for info in media_info:
|
||||
media_file = MediaFile(
|
||||
filename=info["filename"],
|
||||
type=info["type"],
|
||||
caption=info.get("caption"),
|
||||
size=info.get("size"),
|
||||
is_too_large=info.get("is_too_large", False),
|
||||
)
|
||||
else:
|
||||
self.stats["media_downloaded"] += 1
|
||||
post_data.media_files.append(media_file)
|
||||
|
||||
# Generate markdown
|
||||
markdown_content = self._generate_markdown(post_data)
|
||||
if info.get("is_too_large"):
|
||||
self.stats["media_skipped_large"] += 1
|
||||
self.stats["large_files"].append(
|
||||
{
|
||||
"message_id": message_id,
|
||||
"filename": info["filename"],
|
||||
"size": info.get("size"),
|
||||
}
|
||||
)
|
||||
else:
|
||||
self.stats["media_downloaded"] += 1
|
||||
|
||||
# Write index.md
|
||||
index_path = bundle_dir / "index.md"
|
||||
index_path.write_text(markdown_content, encoding="utf-8")
|
||||
# Generate markdown
|
||||
markdown_content = self._generate_markdown(post_data)
|
||||
|
||||
self.stats["posts_archived"] += 1
|
||||
logger.debug(f"Message {message_id} archived successfully")
|
||||
# Write index.md
|
||||
index_path = bundle_dir / "index.md"
|
||||
index_path.write_text(markdown_content, encoding="utf-8")
|
||||
|
||||
self.stats["posts_archived"] += 1
|
||||
|
||||
# Log progress every 50 posts
|
||||
if self.stats["posts_archived"] % 50 == 0:
|
||||
logger.info(f"Progress: {self.stats['posts_archived']} posts archived...")
|
||||
|
||||
logger.debug(f"Message {message_id} archived successfully")
|
||||
|
||||
except Exception as e:
|
||||
logger.error(f"Failed to process message {message_id}: {e}")
|
||||
self.stats["posts_skipped"] += 1
|
||||
|
||||
async def _extract_post_data(self, message: Message) -> PostData:
|
||||
"""
|
||||
@@ -230,10 +289,18 @@ class ChannelArchiver:
|
||||
channel_title = getattr(channel, "title", channel_username)
|
||||
|
||||
# Convert message text to markdown
|
||||
text = message.message or ""
|
||||
if message.message:
|
||||
# Use Telethon's built-in markdown converter
|
||||
text = message.get_message_text(markdown=True)
|
||||
original_text = message.message or ""
|
||||
if message.message and message.entities:
|
||||
# Use Telethon's markdown parser with entities
|
||||
text = message.text # Plain text
|
||||
else:
|
||||
text = original_text
|
||||
|
||||
# Extract hashtags from text
|
||||
tags = extract_hashtags(text)
|
||||
|
||||
# Remove hashtags from text (clean up)
|
||||
clean_text = remove_hashtags(text)
|
||||
|
||||
# Handle reply-to
|
||||
reply_to = None
|
||||
@@ -258,18 +325,20 @@ class ChannelArchiver:
|
||||
return PostData(
|
||||
message_id=message.id,
|
||||
date=message.date,
|
||||
text=text,
|
||||
text=clean_text, # Use cleaned text (without hashtags)
|
||||
author=channel_title,
|
||||
channel_username=channel_username,
|
||||
reply_to=reply_to,
|
||||
repost_from=repost_from,
|
||||
repost_channel=repost_channel,
|
||||
views=getattr(message, "views", None),
|
||||
tags=tags, # Extracted hashtags
|
||||
)
|
||||
|
||||
def _generate_markdown(self, post: PostData) -> str:
|
||||
"""
|
||||
Generate markdown content with front-matter for Hugo.
|
||||
Uses Hugo shortcodes for media embedding.
|
||||
|
||||
Args:
|
||||
post: PostData object
|
||||
@@ -284,6 +353,10 @@ class ChannelArchiver:
|
||||
"author": post.author,
|
||||
}
|
||||
|
||||
# Add tags if present
|
||||
if post.tags:
|
||||
frontmatter["tags"] = post.tags
|
||||
|
||||
# Add optional fields
|
||||
if post.reply_to:
|
||||
# Relative link to the replied post's index.md
|
||||
@@ -309,21 +382,31 @@ class ChannelArchiver:
|
||||
# Build markdown content
|
||||
content = []
|
||||
|
||||
# Add media embeds for downloaded files
|
||||
# Add media embeds using Hugo shortcodes
|
||||
for mf in post.media_files:
|
||||
if not mf.is_too_large:
|
||||
if mf.type == "photo":
|
||||
content.append(f"")
|
||||
# Hugo figure shortcode
|
||||
if mf.caption:
|
||||
content.append(f'{{{{< figure src="{mf.filename}" alt="{mf.caption}" title="{mf.caption}" >}}}}')
|
||||
else:
|
||||
content.append(f'{{{{< figure src="{mf.filename}" >}}}}')
|
||||
elif mf.type == "video":
|
||||
content.append(f"@[video]({mf.filename})")
|
||||
# Hugo video shortcode
|
||||
content.append(f'{{{{< video src="{mf.filename}" >}}}}')
|
||||
elif mf.type == "audio":
|
||||
content.append(f"@[audio]({mf.filename})")
|
||||
# Hugo audio shortcode
|
||||
content.append(f'{{{{< audio src="{mf.filename}" >}}}}')
|
||||
elif mf.type == "document":
|
||||
content.append(f"📎 [{mf.filename}]({mf.filename})")
|
||||
# Download link for documents
|
||||
content.append(f'📎 [{mf.filename}]({mf.filename})')
|
||||
|
||||
# Add text content
|
||||
# Add text content (only if not already in media caption)
|
||||
if post.text:
|
||||
content.append(post.text)
|
||||
# Check if text is already used as caption in media
|
||||
captions = [mf.caption for mf in post.media_files if mf.caption]
|
||||
if post.text not in captions:
|
||||
content.append(post.text)
|
||||
|
||||
# Mark reposts visually
|
||||
if post.repost_from:
|
||||
|
||||
@@ -8,10 +8,15 @@ import sys
|
||||
from pathlib import Path
|
||||
from typing import Optional
|
||||
|
||||
from opentelemetry import trace
|
||||
from opentelemetry.sdk.trace import TracerProvider
|
||||
from opentelemetry.sdk.trace.export import ConsoleSpanExporter, SimpleSpanProcessor
|
||||
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
|
||||
# OpenTelemetry is optional - skip if not installed
|
||||
try:
|
||||
from opentelemetry import trace
|
||||
from opentelemetry.sdk.trace import TracerProvider
|
||||
from opentelemetry.sdk.trace.export import ConsoleSpanExporter, SimpleSpanProcessor
|
||||
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
|
||||
OTEL_AVAILABLE = True
|
||||
except ImportError:
|
||||
OTEL_AVAILABLE = False
|
||||
|
||||
|
||||
def setup_logging(
|
||||
@@ -62,6 +67,11 @@ def setup_logging(
|
||||
|
||||
def setup_opentelemetry() -> None:
|
||||
"""Configure OpenTelemetry tracing."""
|
||||
if not OTEL_AVAILABLE:
|
||||
logger = logging.getLogger(__name__)
|
||||
logger.warning("OpenTelemetry not available - tracing disabled")
|
||||
return
|
||||
|
||||
# Set up tracer provider
|
||||
trace.set_tracer_provider(TracerProvider())
|
||||
|
||||
|
||||
@@ -19,6 +19,7 @@ class MediaType(str, Enum):
|
||||
VIDEO_NOTE = "video_note"
|
||||
STICKER = "sticker"
|
||||
ANIMATION = "animation" # GIF
|
||||
TIMEOUT = "timeout" # File skipped due to timeout
|
||||
|
||||
|
||||
class MediaFile(BaseModel):
|
||||
@@ -76,6 +77,10 @@ class PostData(BaseModel):
|
||||
default=False,
|
||||
description="True if some files were not downloaded due to size limit"
|
||||
)
|
||||
tags: list[str] = Field(
|
||||
default_factory=list,
|
||||
description="Hashtags extracted from the post"
|
||||
)
|
||||
|
||||
@property
|
||||
def bundle_dir(self) -> str:
|
||||
|
||||
@@ -64,10 +64,11 @@ class TelethonArchiver:
|
||||
session=self.settings.session_path,
|
||||
api_id=self.settings.api_id,
|
||||
api_hash=self.settings.api_hash,
|
||||
device_model="Telegram Archiver",
|
||||
device_model="Desktop (Ubuntu)", # Helps with code delivery
|
||||
system_version="Ubuntu 22.04",
|
||||
app_version="1.0.0",
|
||||
lang_code="en",
|
||||
system_lang_code="en",
|
||||
lang_code="ru", # Russian language
|
||||
system_lang_code="ru",
|
||||
)
|
||||
|
||||
await self.client.connect()
|
||||
@@ -83,26 +84,58 @@ class TelethonArchiver:
|
||||
async def _authenticate(self) -> None:
|
||||
"""Handle the authentication flow."""
|
||||
try:
|
||||
# Send code request
|
||||
await self.client.send_code_request(self.settings.phone)
|
||||
logger.info(f"Code sent to {self.settings.phone}")
|
||||
# Try app code first
|
||||
logger.info("Requesting auth code via Telegram app...")
|
||||
sent_code = await self.client.send_code_request(
|
||||
self.settings.phone,
|
||||
force_sms=False
|
||||
)
|
||||
logger.info(f"Code request sent. Type: {sent_code.type}")
|
||||
|
||||
# Get code from user (in MVP, we'll use input())
|
||||
code = input("Enter the code you received: ")
|
||||
# Get code from user
|
||||
print("\n" + "=" * 50)
|
||||
print("AUTHENTICATION REQUIRED")
|
||||
print("=" * 50)
|
||||
print("Check your Telegram app (NOT SMS) for the code.")
|
||||
print("Look for a message from 'Telegram' with the code.")
|
||||
print("=" * 50)
|
||||
print("\nOptions:")
|
||||
print(" 1. Enter the code from Telegram app")
|
||||
print(" 2. Type 'sms' to receive code via SMS")
|
||||
print(" 3. Press Ctrl+C to cancel")
|
||||
print("=" * 50)
|
||||
|
||||
user_input = input("\nEnter code (or 'sms'): ").strip()
|
||||
|
||||
# If user requests SMS
|
||||
if user_input.lower() == 'sms':
|
||||
logger.info("Requesting SMS code...")
|
||||
print("\n📱 Requesting SMS code...")
|
||||
sent_code = await self.client.send_code_request(
|
||||
self.settings.phone,
|
||||
force_sms=True
|
||||
)
|
||||
print("SMS sent! Enter the code:")
|
||||
user_input = input("SMS Code: ").strip()
|
||||
|
||||
code = user_input
|
||||
|
||||
try:
|
||||
await self.client.sign_in(
|
||||
phone=self.settings.phone,
|
||||
code=code
|
||||
code=code,
|
||||
phone_code_hash=sent_code.phone_code_hash
|
||||
)
|
||||
except SessionPasswordNeededError:
|
||||
# 2FA is enabled
|
||||
logger.info("2FA enabled. Enter password:")
|
||||
print("\n🔐 2FA enabled")
|
||||
password = input("2FA Password: ")
|
||||
await self.client.sign_in(password=password)
|
||||
|
||||
except FloodWaitError as e:
|
||||
logger.error(f"Flood wait: must wait {e.seconds} seconds")
|
||||
print(f"\n⚠️ You must wait {e.seconds} seconds before trying again.")
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.error(f"Authentication failed: {e}")
|
||||
@@ -165,24 +198,33 @@ class TelethonArchiver:
|
||||
|
||||
# Iterate messages (newest first)
|
||||
message_count = 0
|
||||
async for message in self.client.iter_messages(
|
||||
entity,
|
||||
limit=limit,
|
||||
min_id=from_message_id, # Messages with ID > from_message_id
|
||||
):
|
||||
yield message
|
||||
message_count += 1
|
||||
error_count = 0
|
||||
|
||||
# Build iter_messages kwargs
|
||||
iter_kwargs = {"limit": limit}
|
||||
if from_message_id:
|
||||
iter_kwargs["min_id"] = from_message_id
|
||||
|
||||
async for message in self.client.iter_messages(entity, **iter_kwargs):
|
||||
try:
|
||||
yield message
|
||||
message_count += 1
|
||||
|
||||
if message_count % 100 == 0:
|
||||
logger.debug(f"Fetched {message_count} messages...")
|
||||
if message_count % 50 == 0:
|
||||
logger.info(f"Progress: {message_count} messages fetched...")
|
||||
except Exception as e:
|
||||
error_count += 1
|
||||
logger.error(f"Error processing message {message.id}: {e}")
|
||||
continue # Skip problematic messages
|
||||
|
||||
logger.info(f"Finished fetching. Total messages: {message_count}")
|
||||
logger.info(f"Finished fetching. Total: {message_count}, Errors: {error_count}")
|
||||
|
||||
async def download_media(
|
||||
self,
|
||||
message: Message,
|
||||
output_path: Path,
|
||||
max_size: Optional[int] = None,
|
||||
timeout: int = 60, # Timeout in seconds
|
||||
) -> list[dict]:
|
||||
"""
|
||||
Download all media from a message.
|
||||
@@ -191,6 +233,7 @@ class TelethonArchiver:
|
||||
message: Telegram message with media
|
||||
output_path: Directory to save media files
|
||||
max_size: Maximum file size to download (bytes)
|
||||
timeout: Timeout for each file download in seconds
|
||||
|
||||
Returns:
|
||||
List of dicts with file info (filename, type, size, path, is_too_large)
|
||||
@@ -204,20 +247,40 @@ class TelethonArchiver:
|
||||
# Ensure output directory exists
|
||||
output_path.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
# Handle different media types
|
||||
if isinstance(message.media, MessageMediaPhoto):
|
||||
# Photo
|
||||
info = await self._download_photo(message, output_path, max_size)
|
||||
if info:
|
||||
media_info.append(info)
|
||||
try:
|
||||
# Handle different media types with timeout
|
||||
if isinstance(message.media, MessageMediaPhoto):
|
||||
# Photo
|
||||
info = await asyncio.wait_for(
|
||||
self._download_photo(message, output_path, max_size),
|
||||
timeout=timeout
|
||||
)
|
||||
if info:
|
||||
media_info.append(info)
|
||||
|
||||
elif isinstance(message.media, MessageMediaDocument):
|
||||
# Document, video, audio, etc.
|
||||
info = await self._download_document(message, output_path, max_size)
|
||||
if info:
|
||||
media_info.append(info)
|
||||
elif isinstance(message.media, MessageMediaDocument):
|
||||
# Document, video, audio, etc.
|
||||
info = await asyncio.wait_for(
|
||||
self._download_document(message, output_path, max_size),
|
||||
timeout=timeout
|
||||
)
|
||||
if info:
|
||||
media_info.append(info)
|
||||
|
||||
# Note: Other media types (geo, poll, etc.) are not downloadable
|
||||
except asyncio.TimeoutError:
|
||||
logger.warning(f"Timeout downloading media for message {message.id}")
|
||||
# Add info about skipped file
|
||||
media_info.append({
|
||||
"filename": "timeout_skipped.bin",
|
||||
"type": "timeout",
|
||||
"caption": None,
|
||||
"size": 0,
|
||||
"path": None,
|
||||
"is_too_large": False,
|
||||
"timeout": True,
|
||||
})
|
||||
except Exception as e:
|
||||
logger.error(f"Error downloading media for message {message.id}: {e}")
|
||||
|
||||
return media_info
|
||||
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
#!/usr/bin/env python
|
||||
"""Interactive login script for Telegram Archiver."""
|
||||
|
||||
import sys
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
# Add project root to path
|
||||
project_root = Path(__file__).parent
|
||||
sys.path.insert(0, str(project_root))
|
||||
|
||||
# Set working directory to project root for .env loading
|
||||
os.chdir(project_root)
|
||||
|
||||
from app.telethon_client import TelethonArchiver
|
||||
from config import get_settings
|
||||
import asyncio
|
||||
|
||||
async def login():
|
||||
"""Interactive login to Telegram."""
|
||||
settings = get_settings()
|
||||
|
||||
print(f"Connecting to Telegram as {settings.phone}...")
|
||||
|
||||
async with TelethonArchiver(settings) as client:
|
||||
print("Successfully logged in!")
|
||||
print(f"Session saved to: {settings.session_path}")
|
||||
|
||||
# Test by getting me info
|
||||
me = await client.client.get_me()
|
||||
print(f"\nLogged in as: {me.first_name} @{me.username}")
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(login())
|
||||
@@ -0,0 +1,50 @@
|
||||
#!/usr/bin/env python
|
||||
"""
|
||||
Interactive login script for Telegram Archiver with SMS fallback.
|
||||
|
||||
If you don't receive the code in the Telegram app, this script will
|
||||
offer to request an SMS instead.
|
||||
"""
|
||||
|
||||
import sys
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
# Add project root to path
|
||||
project_root = Path(__file__).parent
|
||||
sys.path.insert(0, str(project_root))
|
||||
|
||||
# Set working directory to project root for .env loading
|
||||
os.chdir(project_root)
|
||||
|
||||
from app.telethon_client import TelethonArchiver
|
||||
from config import get_settings
|
||||
from telethon.errors import SessionPasswordNeededError
|
||||
import asyncio
|
||||
|
||||
async def login():
|
||||
"""Interactive login to Telegram with SMS fallback."""
|
||||
settings = get_settings()
|
||||
|
||||
print("\n" + "=" * 60)
|
||||
print("TELEGRAM ARCHIVER - AUTHENTICATION")
|
||||
print("=" * 60)
|
||||
print(f"Phone: {settings.phone}")
|
||||
print("=" * 60)
|
||||
|
||||
async with TelethonArchiver(settings) as client:
|
||||
print("\n✅ Successfully logged in!")
|
||||
print(f"📁 Session saved to: {settings.session_path}")
|
||||
|
||||
# Test by getting me info
|
||||
me = await client.client.get_me()
|
||||
print(f"\n👤 Logged in as: {me.first_name} @{me.username}")
|
||||
print(f" User ID: {me.id}")
|
||||
|
||||
if me.username:
|
||||
print(f"\n✅ You can now archive your channel: @{me.username}")
|
||||
|
||||
print("\n" + "=" * 60)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(login())
|
||||
@@ -0,0 +1,71 @@
|
||||
#!/usr/bin/env python
|
||||
"""
|
||||
Login script using Pyrogram as fallback when Telethon fails.
|
||||
|
||||
Pyrogram sometimes works when Telethon is blocked.
|
||||
"""
|
||||
|
||||
import sys
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
# Add project root to path
|
||||
project_root = Path(__file__).parent
|
||||
sys.path.insert(0, str(project_root))
|
||||
os.chdir(project_root)
|
||||
|
||||
from pyrogram import Client
|
||||
import asyncio
|
||||
import configparser
|
||||
|
||||
async def login():
|
||||
"""Interactive login using Pyrogram."""
|
||||
print("\n" + "=" * 60)
|
||||
print("TELEGRAM ARCHIVER - PYROGRAM AUTHENTICATION")
|
||||
print("=" * 60)
|
||||
|
||||
# Parse .env manually
|
||||
env_vars = {}
|
||||
with open('.env', 'r') as f:
|
||||
for line in f:
|
||||
line = line.strip()
|
||||
if line and not line.startswith('#') and '=' in line:
|
||||
key, value = line.split('=', 1)
|
||||
env_vars[key.strip()] = value.strip()
|
||||
|
||||
api_id = int(env_vars.get('API_ID', 0))
|
||||
api_hash = env_vars.get('API_HASH', '')
|
||||
phone = env_vars.get('PHONE', '')
|
||||
|
||||
print(f"\nPhone: {phone}")
|
||||
print(f"API ID: {api_id}")
|
||||
print("=" * 60)
|
||||
print("\nTelegram will send you a code via SMS automatically")
|
||||
print("if the app code is not available.")
|
||||
print("=" * 60)
|
||||
|
||||
# Create Pyrogram client
|
||||
app = Client(
|
||||
name="pyrogram-session",
|
||||
api_id=api_id,
|
||||
api_hash=api_hash,
|
||||
phone_number=phone,
|
||||
workdir=project_root
|
||||
)
|
||||
|
||||
async with app:
|
||||
print("\nSuccessfully logged in!")
|
||||
print(f"Session saved to: {project_root / 'pyrogram-session.session'}")
|
||||
|
||||
# Get me info
|
||||
me = await app.get_me()
|
||||
print(f"\nLogged in as: {me.first_name} @{me.username}")
|
||||
print(f"User ID: {me.id}")
|
||||
|
||||
if me.username:
|
||||
print(f"\nYou can now archive your channel: @{me.username}")
|
||||
|
||||
print("\n" + "=" * 60)
|
||||
|
||||
if __name__ == "__main__":
|
||||
asyncio.run(login())
|
||||
@@ -6,7 +6,7 @@ readme = "README.md"
|
||||
requires-python = ">=3.11"
|
||||
license = {text = "MIT"}
|
||||
authors = [
|
||||
{name = "Evgeny Storozhenko", email = "dedinit"}
|
||||
{name = "Evgeny Storozhenko"}
|
||||
]
|
||||
keywords = ["telegram", "archive", "markdown", "hugo", "telethon"]
|
||||
classifiers = [
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
#!/usr/bin/env python
|
||||
"""Runner script for Telegram Archiver CLI."""
|
||||
|
||||
import sys
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
# Add project root to path
|
||||
project_root = Path(__file__).parent
|
||||
sys.path.insert(0, str(project_root))
|
||||
|
||||
# Set working directory to project root for .env loading
|
||||
os.chdir(project_root)
|
||||
|
||||
from app.main import cli_main
|
||||
|
||||
if __name__ == "__main__":
|
||||
cli_main()
|
||||
@@ -0,0 +1,62 @@
|
||||
"""
|
||||
Simple Telethon auth test with flood wait handling.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
from telethon import TelegramClient
|
||||
from telethon.errors import FloodWaitError, SessionPasswordNeededError
|
||||
|
||||
api_id = 24276216
|
||||
api_hash = '2d6d0bb78f433bc5a086735c63e3958f'
|
||||
phone = '+79284382181'
|
||||
|
||||
async def main():
|
||||
client = TelegramClient(
|
||||
'test-session',
|
||||
api_id,
|
||||
api_hash,
|
||||
device_model='Desktop (Ubuntu)',
|
||||
system_version='Ubuntu 22.04',
|
||||
app_version='1.0.0',
|
||||
lang_code='ru',
|
||||
system_lang_code='ru'
|
||||
)
|
||||
|
||||
await client.connect()
|
||||
|
||||
try:
|
||||
# Check if already authorized
|
||||
if await client.is_user_authorized():
|
||||
print("Already authorized!")
|
||||
me = await client.get_me()
|
||||
print(f"Logged in as: {me.first_name} @{me.username}")
|
||||
return
|
||||
|
||||
# Try to send code
|
||||
print(f"Requesting code for {phone}...")
|
||||
result = await client.send_code_request(phone)
|
||||
print(f"Code sent! Type: {result.type}")
|
||||
print(f"Phone code hash: {result.phone_code_hash}")
|
||||
|
||||
# Wait for user input
|
||||
code = input("Enter the code: ")
|
||||
|
||||
try:
|
||||
await client.sign_in(phone, code, phone_code_hash=result.phone_code_hash)
|
||||
except SessionPasswordNeededError:
|
||||
print("2FA enabled! Enter your cloud password:")
|
||||
password = input("Password: ")
|
||||
await client.sign_in(password=password)
|
||||
|
||||
me = await client.get_me()
|
||||
print(f"Successfully logged in as: {me.first_name} @{me.username}")
|
||||
|
||||
except FloodWaitError as e:
|
||||
print(f"FLOOD WAIT! Must wait {e.seconds} seconds ({e.seconds // 60} minutes)")
|
||||
print(f"Try again after {e.seconds} seconds")
|
||||
except Exception as e:
|
||||
print(f"Error: {e}")
|
||||
finally:
|
||||
await client.disconnect()
|
||||
|
||||
asyncio.run(main())
|
||||
@@ -0,0 +1,52 @@
|
||||
#!/usr/bin/env python
|
||||
"""Test full channel archive with flood handling."""
|
||||
|
||||
import asyncio
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
project_root = Path(__file__).parent
|
||||
sys.path.insert(0, str(project_root))
|
||||
os.chdir(project_root)
|
||||
|
||||
from app.telethon_client import TelethonArchiver
|
||||
from app.archiver import ChannelArchiver
|
||||
from config import get_settings
|
||||
import logging
|
||||
|
||||
logging.basicConfig(level=logging.INFO)
|
||||
|
||||
async def main():
|
||||
settings = get_settings()
|
||||
|
||||
print("=" * 60)
|
||||
print("FULL CHANNEL ARCHIVE TEST")
|
||||
print("=" * 60)
|
||||
|
||||
async with TelethonArchiver(settings) as client:
|
||||
# Get channel info
|
||||
entity = await client.get_channel_info('dedinit')
|
||||
print(f"Channel: {entity.title}")
|
||||
print("=" * 60)
|
||||
|
||||
# Archive all
|
||||
archiver = ChannelArchiver(client, settings)
|
||||
result = await archiver.archive_channel(
|
||||
channel='dedinit',
|
||||
limit=None, # All posts
|
||||
force=False # Skip existing
|
||||
)
|
||||
|
||||
print("\n" + "=" * 60)
|
||||
print("RESULTS")
|
||||
print("=" * 60)
|
||||
print(f"Posts archived: {result['posts_archived']}")
|
||||
print(f"Posts skipped: {result['posts_skipped']}")
|
||||
print(f"Duration: {result['duration_seconds']:.2f}s")
|
||||
print(f"Output: {result['output_path']}")
|
||||
print("=" * 60)
|
||||
|
||||
if __name__ == "__main__":
|
||||
import os
|
||||
asyncio.run(main())
|
||||
Reference in New Issue
Block a user