5.4 KiB
Phase 0 Spike Report — Needle 2 prescreen & RustFS storage
Date: 2026-08-16. Evidence from live testing in isolated environments (
/tmp/opencode/needle-spike, ephemeralrustfs/rustfs:latestcontainer). Gates the Phase 1 (extraction refactor) and Phase 2 (prescreen stage) plans.
Spike 1: cactus-needle (Needle 2) for RU contract prescreen — NOT VIABLE
What was verified
cactus-needle==2.0.5is real (Cactus Compute, Apache-2.0, first release 2026-08-10). Installs cleanly on Python 3.13. The spec's claims hold: 14MB engine (fetched once from HF, cached offline),needle.extract(text, PydanticModel), calibratedconfidenceon responses,.cactLoRA fine-tuning pipeline. Note: PyPI package name iscactus-needle(import needle); bareneedleon PyPI is an unrelated dead 2017 CSS testing tool.- Engine session RAM measured at 37.8MB (simple schema) to 119.7MB (larger schema) — fine for a 256M container. Python deps pull the JAX/Flax/optax stack (~200MB+ image impact) because fine-tuning support ships in the core dependency list.
Results (base weights, buffer_size default)
| Case | Result |
|---|---|
| EN invoice (README example) | Perfect extraction, confidence=0.9484 |
| EN off-topic sentence | Correctly refused: empty call [], confidence=0.7427 |
| RU contract, 2.5K chars | Total hallucination: parties → "Supplier"/"Supplier", amount → 12.0 USD (real: 1 234 567,80 RUB), dates → 2024 defaults. confidence=0.0 |
| RU contract, 43K chars | Same failure mode, different hallucinations. confidence=0.0 |
| RU garbage text | Not refused — fabricated "Party A/Party B/150 USD". confidence=0.0 |
Verdict
The base model is English-centric; it does not understand Russian. The calibrated confidence head behaves correctly (0.0 = "I can't read this"), so a naive deployment would have been safe (100% → manual_review) but useless.
Fine-tuning cannot rescue this for v1:
- Calibrated confidence is disabled on tuned weights (reports
None) — and the spec's entire routing design hinges on confidence gating. - The model failed even trivial verbatim copy tasks (could not copy «ООО «Ромашка»» from adjacent text) — this is tokenizer/training coverage, not prompt tuning.
- Building a RU legal extraction dataset + LoRA pipeline is its own project.
Plan adjustment (Phase 2)
worker-prescreenships with a deterministic regex+pydantic extractor instead of Needle: RU/BY contract boilerplate is highly templated («договор поставки», «именуемое в дальнейшем», «сумма договора составляет N рублей», «неустойка», «арбитражный суд», DD.MM.YYYY ranges). Zero marginal cost, no hallucination, p95 ≈ ms. Absence of a match =None→manual_review, which matches the routing philosophy.confidence_scoreis redefined as a deterministic field-coverage score (fraction of target fields extracted), not model calibration. Router thresholds operate on coverage.needle_client.py/model_weights/NEEDLE_*env are dropped from the plan; keepprescreen_results.extractor_versionfor future engine swaps.- Optional fast-follow (not v1): small-LLM prescreen via the existing
core/llmport for fields regex missed. - Revisit Needle only if/when a multilingual or RU-tuned base ships with calibration intact.
Spike 2: RustFS as MinIO alternative — VIABLE, but stay on MinIO for now
Tested with the project's own MinioStorage adapter (minio SDK) against
rustfs/rustfs:latest, single node, named volume.
| Check | Result |
|---|---|
Health (/minio/health/live) |
200 — MinIO-compatible probe works |
| bucket create/exists, put/get/stat/delete | OK (incl. Cyrillic UTF-8 payloads) |
set/get/delete_bucket_lifecycle (ILM config API) |
OK — rule round-trips |
ILM expiration behavior (RUSTFS_ILM_DEBUG_DAY_SECS=2, rule days=1) |
Object under users/ purged in ~10s; control object outside prefix untouched |
| Perf (4KB put+get ×20) | ~2.0ms/op |
| Docker ergonomics | Runs as UID 10001 — bind mounts must be chowned; named volumes are the easy path |
Verdict
The critical blocker (152-ФЗ TTL purge via ILM) works in practice — the "Under Testing" label upstream refers to CI gating, not absence. Remaining reasons to keep MinIO as the default for production:
- Project is weeks old (1.0.0-rc era); lifecycle, distributed mode, and KMS are all still marked "Under Testing" upstream.
- SSE migration trap: objects written by MinIO with SSE-S3/KMS/C are not
readable by RustFS — a later migration requires re-putting objects through
the app (we run SSE-S3=off in dev, on in prod per
S3_SERVER_SIDE_ENCRYPTION). - No operational history; MinIO is battle-tested for this exact compose setup.
Because everything S3-shaped is already behind core/s3/port.py, switching
later is a compose + env change, not a code change. Action item: revisit when
RustFS marks Lifecycle + KMS GA (tracked in ARCHITECTURE notes, not code).
Phase gate outcomes
- Phase 1 (extraction refactor): GO — unchanged.
- Phase 2 (prescreen stage): GO with amendments — regex extractor instead
of Needle; coverage-based confidence; no needle deps/env; conservative
routing (
PRESCREEN_AUTO_APPROVE=falsedefault) as already planned. - Storage: no change — MinIO stays; RustFS is a credible fast-follow.