Files
2026-09-06 13:51:20 +00:00

3.1 KiB
Raw Permalink Blame History

Spec/структура parser — engineering PDF text layer

Pattern that worked on the wine-city СКС project (Спецификация оборудования ТШ шкафов из 90-page DWG-render PDF, 384KB extracted text). Reusable for any "оборудование по шкафам/узлам" spec table whose text layer is radially extracted.

Goal

Turn the extracted .txt (with ===== СТРАНИЦА N ===== page markers) into structured items WITHOUT reading 2000 lines/spec into the LLM.

Heuristic-обозначение parser

import re
def parse_spec(lines, start, end):
    """items = [ {oboz, name, qty}, ... ]"""
    items, i = [], start
    while i < end:
        s = lines[i-1].strip()
        if not s:
            i += 1; continue
        # обозначение: one token, upper/alnum/dash/dot, >3 chars, not starting with digit
        if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', s) and len(s) > 3 and not s[0].isdigit():
            cur = {"oboz": s, "name": [], "qty": None}
            j = i + 1; name_lines = []
            while j < end and len(name_lines) < 8:     # cap: наименование ≤8 строк
                t = lines[j-1].strip()
                if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', t) and len(t) > 3 and not t[0].isdigit():
                    break                              # next обозначение
                if re.fullmatch(r'\d+', t):
                    cur["qty"] = int(t); j += 1; break # qty = bare int
                if t and not re.fullmatch(r'\d+', t):
                    name_lines.append(t)
                j += 1
            cur["name"] = " ".join(name_lines)[:140]
            items.append(cur); i = j
        else:
            i += 1
    return items

Known noise to filter (DWG-render text layer)

  • Rack U-number grids: long runs of bare ints (1..47) — compiler may capture as qty=None items; ignore qty is None rows unless they carry a real name (they won't: name stays empty).
  • SFP1..4 / QSFP1..6 / "резерв" / "220V Port 1..." label rows leak in from switch faceplate text. Drop tokens in {SFP*, QSFP*} and empty-name items.
  • Real spec items are the ones with BOTH a multi-char обозначение AND a non-empty name AND a qty — keep those as the equipment table.

Diff across nearly-identical sections (e.g. ТШ-A.1..A.4)

Since соседние spec-блоки почти одинаковы, don't re-read each: parse all to JSON, then compare counts per (oboz, name) to spot the deltas (A.1 vs A.2 vs A.3 vs A.4) and only narrate what differs.

Finding section boundaries without full reads

Scan the .txt lines for header regexes once (cheap): СПЕЦИФИКАЦИЯ, ВЕДОМОСТЬ, СХЕМА, ОБЩИЕ ДАННЫЕ, СОДЕРЖАНИЕ; record their line numbers, then read_file(offset=…) ONLY the interesting ranges.

Verification

Before trusting counts, read_file 1-2 rows at their known line and match против исходной таблицы on the actual PDF page (page.get_text() for the page). Spec numbers are consumable-accurate (закупочные количества) — worth the check.