# Spec/структура parser — engineering PDF text layer Pattern that worked on the wine-city СКС project (Спецификация оборудования ТШ шкафов из 90-page DWG-render PDF, 384KB extracted text). Reusable for any "оборудование по шкафам/узлам" spec table whose text layer is radially extracted. ## Goal Turn the extracted .txt (with `===== СТРАНИЦА N =====` page markers) into structured items WITHOUT reading 2000 lines/spec into the LLM. ## Heuristic-обозначение parser ```python import re def parse_spec(lines, start, end): """items = [ {oboz, name, qty}, ... ]""" items, i = [], start while i < end: s = lines[i-1].strip() if not s: i += 1; continue # обозначение: one token, upper/alnum/dash/dot, >3 chars, not starting with digit if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', s) and len(s) > 3 and not s[0].isdigit(): cur = {"oboz": s, "name": [], "qty": None} j = i + 1; name_lines = [] while j < end and len(name_lines) < 8: # cap: наименование ≤8 строк t = lines[j-1].strip() if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', t) and len(t) > 3 and not t[0].isdigit(): break # next обозначение if re.fullmatch(r'\d+', t): cur["qty"] = int(t); j += 1; break # qty = bare int if t and not re.fullmatch(r'\d+', t): name_lines.append(t) j += 1 cur["name"] = " ".join(name_lines)[:140] items.append(cur); i = j else: i += 1 return items ``` ## Known noise to filter (DWG-render text layer) - Rack U-number grids: long runs of bare ints (1..47) — compiler may capture as qty=None items; ignore `qty is None` rows unless they carry a real name (they won't: name stays empty). - SFP1..4 / QSFP1..6 / "резерв" / "220V Port 1..." label rows leak in from switch faceplate text. Drop tokens in {SFP*, QSFP*} and empty-name items. - Real spec items are the ones with BOTH a multi-char обозначение AND a non-empty name AND a qty — keep those as the equipment table. ## Diff across nearly-identical sections (e.g. ТШ-A.1..A.4) Since соседние spec-блоки почти одинаковы, don't re-read each: parse all to JSON, then compare counts per (oboz, name) to spot the deltas (A.1 vs A.2 vs A.3 vs A.4) and only narrate what differs. ## Finding section boundaries without full reads Scan the .txt lines for header regexes once (cheap): `СПЕЦИФИКАЦИЯ`, `ВЕДОМОСТЬ`, `СХЕМА`, `ОБЩИЕ ДАННЫЕ`, `СОДЕРЖАНИЕ`; record their line numbers, then `read_file(offset=…)` ONLY the interesting ranges. ## Verification Before trusting counts, `read_file` 1-2 rows at their known line and match против исходной таблицы on the actual PDF page (`page.get_text()` for the page). Spec numbers are consumable-accurate (закупочные количества) — worth the check.