mirror of
https://gitverse.ru/kpa39l/project-doc-processing.git
synced 2026-09-29 09:15:04 +00:00
50 lines
3.1 KiB
Markdown
50 lines
3.1 KiB
Markdown
# Spec/структура parser — engineering PDF text layer
|
||
|
||
Pattern that worked on the wine-city СКС project (Спецификация оборудования ТШ шкафов из 90-page DWG-render PDF, 384KB extracted text). Reusable for any "оборудование по шкафам/узлам" spec table whose text layer is radially extracted.
|
||
|
||
## Goal
|
||
Turn the extracted .txt (with `===== СТРАНИЦА N =====` page markers) into structured items WITHOUT reading 2000 lines/spec into the LLM.
|
||
|
||
## Heuristic-обозначение parser
|
||
```python
|
||
import re
|
||
def parse_spec(lines, start, end):
|
||
"""items = [ {oboz, name, qty}, ... ]"""
|
||
items, i = [], start
|
||
while i < end:
|
||
s = lines[i-1].strip()
|
||
if not s:
|
||
i += 1; continue
|
||
# обозначение: one token, upper/alnum/dash/dot, >3 chars, not starting with digit
|
||
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', s) and len(s) > 3 and not s[0].isdigit():
|
||
cur = {"oboz": s, "name": [], "qty": None}
|
||
j = i + 1; name_lines = []
|
||
while j < end and len(name_lines) < 8: # cap: наименование ≤8 строк
|
||
t = lines[j-1].strip()
|
||
if re.fullmatch(r'[A-ZА-Я0-9][A-ZА-Я0-9\-.]+', t) and len(t) > 3 and not t[0].isdigit():
|
||
break # next обозначение
|
||
if re.fullmatch(r'\d+', t):
|
||
cur["qty"] = int(t); j += 1; break # qty = bare int
|
||
if t and not re.fullmatch(r'\d+', t):
|
||
name_lines.append(t)
|
||
j += 1
|
||
cur["name"] = " ".join(name_lines)[:140]
|
||
items.append(cur); i = j
|
||
else:
|
||
i += 1
|
||
return items
|
||
```
|
||
|
||
## Known noise to filter (DWG-render text layer)
|
||
- Rack U-number grids: long runs of bare ints (1..47) — compiler may capture as qty=None items; ignore `qty is None` rows unless they carry a real name (they won't: name stays empty).
|
||
- SFP1..4 / QSFP1..6 / "резерв" / "220V Port 1..." label rows leak in from switch faceplate text. Drop tokens in {SFP*, QSFP*} and empty-name items.
|
||
- Real spec items are the ones with BOTH a multi-char обозначение AND a non-empty name AND a qty — keep those as the equipment table.
|
||
|
||
## Diff across nearly-identical sections (e.g. ТШ-A.1..A.4)
|
||
Since соседние spec-блоки почти одинаковы, don't re-read each: parse all to JSON, then compare counts per (oboz, name) to spot the deltas (A.1 vs A.2 vs A.3 vs A.4) and only narrate what differs.
|
||
|
||
## Finding section boundaries without full reads
|
||
Scan the .txt lines for header regexes once (cheap): `СПЕЦИФИКАЦИЯ`, `ВЕДОМОСТЬ`, `СХЕМА`, `ОБЩИЕ ДАННЫЕ`, `СОДЕРЖАНИЕ`; record their line numbers, then `read_file(offset=…)` ONLY the interesting ranges.
|
||
|
||
## Verification
|
||
Before trusting counts, `read_file` 1-2 rows at their known line and match против исходной таблицы on the actual PDF page (`page.get_text()` for the page). Spec numbers are consumable-accurate (закупочные количества) — worth the check. |