Files
file-tree-catalog/SKILL.md
T
2026-09-06 13:51:19 +00:00

34 lines
3.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: file-tree-catalog
description: Catalog a project/file tree into Markdown programmatically.
---
# File-Tree Catalog
Create a human-navigable Markdown catalog of a large local file tree (engineering project archives, design-document sets, media libraries). This is **mechanical filesystem work — do it with Python, do NOT burn tokens on an LLM**. Deliverables: a `КАТАЛОГ_проекта.md` in the root of the tree.
## When to use
- User asks to "каталогизировать", "сделать каталог/структуру/содержимое" of a folder, "понять что лежит в проекте X".
- User hands you a path to an archive of documents (СКС/ЛВС/КТСБ drawings, PDF sets, etc.) and wants structure + contents overview.
## Workflow
1. **Confirm the target** (what kind of catalog) with one `clarify` if ambiguous — but structural cataloging is usually unambiguous: structure + contents. Don't over-ask.
2. **Enumerate** the tree with `execute_code` (Python, not shell): walk `os.walk`, collect extensions, sizes, counts. Get a feel for scale first (files, GB, top-level dirs, distribution by extension) before building the document.
3. **Build the catalog procedurally** (see bundled script). Key decisions:
- Group by immediate subdirectories; recurse 2-3 levels max so the doc stays readable.
- Summarize `note+dwg/` (or similar bulk design subfolders) as `N файлов` instead of listing 50+ CAD files.
- Tag each file with its *role* derived from extension + naming: PDF="итоговый комплект / итоговый PDF", DOCX="список изменений", XLSX="замечания/ответы", TXT="изменения".
- Prepend an **Общая статистика** section (total files, total size, per-extension counts) and a **Состав проекта** section (what the top-level systems/objects are).
4. **Write** the `.md` into the tree root with `open(out,"w")`. Report the absolute path.
## Pitfalls
- **Don't use an LLM for the catalog itself** — building a `dict`/tree + markdown join is deterministic and faster. Reserve a local model only for *semantic* additions (expanding abbreviations), and only if it actually answers.
- **Don't fabricate expansions of unknown codes.** If a шифр/abbreviation (e.g. `Н-КВ`) isn't verifiable, state the *fact* ("belongs to the Кубань-Вино object") and explicitly note it's unverified — do NOT invent "новый корпус" etc.
- **Filter service/junk files** or the doc drowns. Consistent engineering-archive junk on this host: `Thumbs.db`, `~$…` (Office lock files), `.dwl`/`.dwl2` (AutoCAD locks), `.bak`, `plot.log`, `.db`. Count them separately in stats but omit from the body.
- **Ordinal/numbering**: sort dirs with `key=str.lower` for natural locale ordering (КТСБ.1 < КТСБ.2, зоны in order).
- Be careful writing the markdown with `%`-format strings or plain concatenation — **nested f-strings with quotes inside a list comprehension raise a SyntaxError**; use `%`-formatting or build lines imperatively.
- Optional user-request: "use the local model to save tokens" — for mechanical cataloging the honest answer is "no model needed"; say so rather than calling ollama pointlessly. Local qwen3 on this host returns empty `content` (reasoning-only replies), so don't count on it for prose.
## Files
- `scripts/build_catalog.py` — reusable generator (stats + tree + role tagging + junk filter). Edit the constant `ROOT` and run; adapt `maxdepth`/`junk` sets as needed.
- `references/engineering-doc-archive.md` — conventions and domain notes for Russian design/СКС-ЛВС-КТСБ document archives.