13 KiB
Refactor Plan: Prescreen Worker → Hybrid Extraction (Heuristic + LLM Fallback)
Status: delivered through Phase 4; Phase 5 rollout pending merge Scope:
src/contract_check/worker_prescreen/,src/contract_check/core/llm/,src/contract_check/core/config.py,src/contract_check/core/metrics.pySupersedes: the regex-only extractor described inworker_prescreen/extractor.py(regex-v1)
1. Goal
Replace the monolithic regex extractor (extractor.py, ~260 lines of Russian/BY legal
regexes) with a hybrid two-stage extractor:
- Stage 1 — Heuristic (deterministic, zero-cost): keyword dictionaries, positional windows, sentence scanning, per-field confidence weights. No regex — plain string operations only.
- Stage 2 — LLM fallback (only when Stage 1 confidence is low): reuse the
existing
core/llmprovider infrastructure (Ollama Cloud / YandexGPT) with a JSON-schema-constrained extraction prompt.
Non-goals:
- No changes to routing semantics (
router.pythresholds/decisions stay as-is). - No DB schema migration (
prescreen_results.extractor_versionalready exists,String(32)fits"heuristic-v2"/"llm-v1"). - No changes to MQ topology, messages (
PrescreenCompletedshape unchanged), or downstreamworker_analyze. - Auto-approve stays disabled by default.
2. Current State
worker_prescreen/
├── consumer.py # consumes prescreen.q → handler
├── handler.py # Stage 4: asyncio.to_thread(extract_contract_meta, text)
├── extractor.py # PrescreenContractMeta + 10 regex patterns ← TO REPLACE
├── router.py # decides auto_approve|manual_review|deep_analysis
└── config.py
Problems:
- Regexes are brittle (line-noise, OCR artifacts, whitespace variants) and hard to extend (every new contract type = new regex).
- No recovery path: low-confidence extraction always ends in
manual_review. prescreen_results.extractor_versionis never written by the handler INSERT — it silently relies on the DB default'regex-v1'.
3. Target Architecture
┌────────────────────────────────────────────┐
│ handler.py │
│ Stage 4: meta = await extractor.extract() │
└───────────────┬────────────────────────────┘
│
HybridMetaExtractor (orchestrator)
│
┌────────────────────┴──────────────────────┐
▼ ▼
HeuristicExtractor (sync, to_thread) LLMPrescreenExtractor
keyword + positional + sentence wraps core/llm provider
scanning, confidence-weighted extract_prescreen(text)
│ │
confidence ≥ threshold? ──── no ────► run LLM, validate JSON,
│ yes merge over heuristic meta
▼
PrescreenContractMeta → router.decide() → persist (with extractor_version)
3.1 New module layout (flat, matches existing package conventions)
worker_prescreen/
├── extractor.py # PrescreenContractMeta (model, unchanged shape)
│ # + MetadataExtractor protocol + re-export shim
├── extractor_heuristic.py# HeuristicExtractor (Stage 1)
├── extractor_llm.py # LLMPrescreenExtractor (Stage 2, dict → pydantic validation)
├── extractor_hybrid.py # HybridMetaExtractor (threshold + merge + kill-switch)
├── handler.py # Stage 4 rewired; INSERT gains extractor_version
└── router.py # unchanged
3.2 Heuristic stage design (regex-free)
| Field | Method | Notes |
|---|---|---|
contract_type |
Phrase dictionary match on normalized text | {"supply": ["договор поставки", "договор купли-продажи", ...], ...} — lowercase + whitespace-collapse once, then in checks |
party_a/b |
Positional: scan only the header window (text up to first 1. / ПРЕДМЕТ heading, capped ~1500 chars); token-scan for entity-form tokens (ООО, ИП, АО, …) then capture until quote »/" close or line end |
Replaces _ENTITY_RE |
total_amount + currency |
Trigger-word scan (составляет, стоимость, цена, …) then manual digit-window parse (_scan_number helper walking chars, handling spaces/commas); currency via token lookup in the following ~50 chars |
Replaces _AMOUNT_RE |
start/end_date |
Trigger-word scan (действует с, с … по …) then _parse_ddmmyyyy manual splitter (already regex-free in v1 — keep) |
Replaces _DATE_RE/_END_DATE_RE |
| boolean clauses | Sentence segmentation via str.split on . / \n, keyword membership per sentence (неустойка, штраф, расторгнуть, арбитражн, …) |
Replaces _PENALTY_RE etc. |
Confidence scoring becomes weighted instead of flat field coverage:
FIELD_WEIGHTS = {
"contract_type": 0.25,
"party_a": 0.15,
"party_b": 0.15,
"total_amount": 0.15,
"currency": 0.05,
"start_date": 0.10,
"end_date": 0.05,
"has_penalty_clause": 0.05,
"has_termination_clause": 0.05,
"has_arbitration": 0.05,
} # sums to 1.0
Optional per-field method bonus (exact phrase match = full weight, positional window hit = full weight, fuzzy tail hit = ×0.7) — start simple, weights constant, tune later with real data.
3.3 LLM fallback design
- Protocol extension (
core/llm/port.py):async def extract_prescreen(self, text: str) -> dict[str, Any]— returns raw JSON dict; no import of worker_prescreen (keeps layering clean). Implemented by bothOllamaCloudProviderandYandexGPTProvideras a thin wrapper over their existing JSON-chat + repair-loop machinery (_chat_json/responseFormat=json_schema) with a dedicatedPRESCREEN_SYSTEMprompt: extract the 10 meta fields, cite nothing, JSON only. - Worker-side wrapper (
extractor_llm.py):dict→PrescreenContractMetavia pydantic (rejects hallucinated fields, re-Nones unknown enum values), then recomputes weighted confidence. - Input cap: only the first
prescreen_llm_max_chars(default 20 000) chars are sent — meta lives in the header for templated contracts; cost control. - Merge rule: LLM values override heuristic
Nones and low-confidence fields; boolean clauses become OR(heuristic, llm) — both sources are presence-checks, false positives are cheap, false negatives route wrong.
3.4 Failure semantics
- LLM fallback error (quota, timeout, invalid JSON after repair):
do not fail the message. Log + keep heuristic meta,
extractor_version="heuristic-v2", record the error inprescreen_results.auto_findings({"llm_fallback_error": "..."}) and theprescreen_fallback_runs_total{outcome="failed"}counter. - Heuristic stage is pure string ops — its only failure mode is
Nonefields, which is already handled by low confidence → fallback / manual_review.
3.5 Config additions (core/config.py, --- prescreen stage ---)
prescreen_llm_fallback_enabled: bool = Field(default=False, ...) # kill-switch; flip to True after burn-in
prescreen_llm_fallback_threshold: float = Field(default=0.75, ...) # ≤ router threshold
prescreen_llm_max_chars: int = Field(default=20_000, ...)
3.6 Metrics additions (core/metrics.py)
prescreen_fallback_runs = Counter(
"contract_check_prescreen_fallback_runs_total",
"...", labelnames=["outcome"], # used | failed | skipped | disabled
)
prescreen_extraction_stage = Histogram(
"contract_check_prescreen_extraction_stage_seconds",
"...", labelnames=["stage"], # heuristic | llm
)
extractor_version values: "regex-v1" (legacy shim flag), "heuristic-v2"
(LLM not run / disabled / failed) and "hybrid-llm-v1" (LLM result merged).
4. Implementation Phases
Each phase is independently shippable and reverted by config/env flag.
Phase 1 — Port + heuristic extractor (no behavior change for routing)
extractor.py: keepPrescreenContractMeta+ weighted_score_confidence; addMetadataExtractorprotocol (def extract(text: str) -> PrescreenContractMeta); keepextract_contract_metaas a delegating shim sotests/unit/test_prescreen_extractor.pyandworker_prescreen/__init__.pykeep importing it.- New
extractor_heuristic.pyimplementing all helpers from §3.2 (_scan_number,_find_after_trigger, header-window splitter, sentence splitter, dictionaries as module-level constants). - Port existing unit-test fixtures (
SIMPLE_SUPPLY,MINIMAL, boolean-flag parametrize) ontoHeuristicExtractor; they must pass with identical expected values. extractor_versionshim reports"heuristic-v2".
Phase 2 — LLM extract_prescreen on providers
- Extend
LLMProviderprotocol + both adapters (ollama_cloud.py,yandex_gpt.py): newPRESCREEN_SYSTEMprompt, JSON schema for the 10 fields, reuse repair loop, truncate input tochunk_size-independent small cap. - Unit tests with a fake transport (both adapters already have this pattern):
valid dict, invalid enum →
None, malformed JSON → repair once → fail. extractor_llm.py: dict →PrescreenContractMetavalidation + weighted confidence.
Phase 3 — Orchestrator + handler wiring
extractor_hybrid.py: threshold check, LLM call (async), merge rules, error swallowing, extractor_version selection, kill-switch.handler.pyStage 4: replaceasyncio.to_thread(extract_contract_meta, …)withawait self._extractor.extract(contract_text)whereself._extractoris injectable (constructor arg, defaults to hybrid) — mirrors the existingproviderinjection pattern inAnalyzeHandler.- Handler Stage 6 INSERT: bind
extractor_version. - Config + metrics from §3.5/§3.6.
Phase 4 — Tests + verification
- Unit: orchestrator matrix — high confidence skips LLM; low confidence
merges; LLM failure → heuristic result +
outcome="failed"; disabled →outcome="disabled". - Integration (
tests/integration/test_prescreen_worker.pypattern): run handler in-process with a stub provider; assertprescreen_results.extractor_versionpersisted, routing unchanged. - Backfill spot-check: replayed 8 RU/BY-style contract samples through regex-v1 vs heuristic-v2. Confidence deltas: all neutral or better, heuristic fixes regex under-parsing on NBSP/noise and captures fuller party names. Record below in PR description.
ruff check . && mypy src && pytest tests/unit tests/integration -k prescreen. (ruff/mypy findings outside the touched files are from pre-existing uncommitted changes; prescreen tests pass.)
Phase 5 — Rollout
- Merge with
prescreen_llm_fallback_enabled=false(pure heuristic). - Observe
prescreen_fallback_runs/prescreen_confidencefor a few days. - Document in
docs/ARCHITECTURE.mdprescreen section +.env.examplecompleted. The regex notes inSPIKE_PHASE0.mdare no longer referenced from the extractor docstring. Staging/prod enablement remains an operational step after merge.
5. Risks & Mitigations
| Risk | Mitigation |
|---|---|
| Heuristic drops accuracy vs regex on edge templates | Port 100% of v1 unit tests before deleting regex; keep extractor.py regex code one release behind a flag (PRESCREEN_KEEP_REGEX=true env, temporary) |
| LLM hallucinates fields | pydantic validation whitelists contract_type enums; numeric/date parse checks; booleans only OR-merged |
| LLM latency blows up prescreen SLA | 20k char cap, existing provider timeouts, prescreen_extraction_stage histogram; fallback failures never block the pipeline |
| Cost creep on fallback rate | prescreen_fallback_runs_total alert; threshold tunable without deploy via env |
| Layering violation (core ← worker import) | Protocol returns plain dict; pydantic model stays in worker_prescreen |
6. Rollback
- Config:
PRESCREEN_LLM_FALLBACK_ENABLED=false→ deterministic heuristic only. - Full: revert merge — no DB migration to undo;
extractor_versionstrings are informational only.