Files
2026-09-06 13:51:20 +00:00

50 lines
3.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Spec/структура parser — engineering PDF text layer
Pattern that worked on the wine-city СКС project (Спецификация оборудования ТШ шкафов из 90-page DWG-render PDF, 384KB extracted text). Reusable for any "оборудование по шкафам/узлам" spec table whose text layer is radially extracted.
## Goal
Turn the extracted .txt (with `===== СТРАНИЦА N =====` page markers) into structured items WITHOUT reading 2000 lines/spec into the LLM.
## Heuristic-обозначение parser
```python
import re
def parse_spec(lines, start, end):
"""items = [ {oboz, name, qty}, ... ]"""
items, i = [], start
while i < end:
s = lines[i-1].strip()
if not s:
i += 1; continue
# обозначение: one token, upper/alnum/dash/dot, >3 chars, not starting with digit
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', s) and len(s) > 3 and not s[0].isdigit():
cur = {"oboz": s, "name": [], "qty": None}
j = i + 1; name_lines = []
while j < end and len(name_lines) < 8: # cap: наименование ≤8 строк
t = lines[j-1].strip()
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', t) and len(t) > 3 and not t[0].isdigit():
break # next обозначение
if re.fullmatch(r'\d+', t):
cur["qty"] = int(t); j += 1; break # qty = bare int
if t and not re.fullmatch(r'\d+', t):
name_lines.append(t)
j += 1
cur["name"] = " ".join(name_lines)[:140]
items.append(cur); i = j
else:
i += 1
return items
```
## Known noise to filter (DWG-render text layer)
- Rack U-number grids: long runs of bare ints (1..47) — compiler may capture as qty=None items; ignore `qty is None` rows unless they carry a real name (they won't: name stays empty).
- SFP1..4 / QSFP1..6 / "резерв" / "220V Port 1..." label rows leak in from switch faceplate text. Drop tokens in {SFP*, QSFP*} and empty-name items.
- Real spec items are the ones with BOTH a multi-char обозначение AND a non-empty name AND a qty — keep those as the equipment table.
## Diff across nearly-identical sections (e.g. ТШ-A.1..A.4)
Since соседние spec-блоки почти одинаковы, don't re-read each: parse all to JSON, then compare counts per (oboz, name) to spot the deltas (A.1 vs A.2 vs A.3 vs A.4) and only narrate what differs.
## Finding section boundaries without full reads
Scan the .txt lines for header regexes once (cheap): `СПЕЦИФИКАЦИЯ`, `ВЕДОМОСТЬ`, `СХЕМА`, `ОБЩИЕ ДАННЫЕ`, `СОДЕРЖАНИЕ`; record their line numbers, then `read_file(offset=…)` ONLY the interesting ranges.
## Verification
Before trusting counts, `read_file` 1-2 rows at their known line and match против исходной таблицы on the actual PDF page (`page.get_text()` for the page). Spec numbers are consumable-accurate (закупочные количества) — worth the check.