OCR-SPRIN-SERVICE

adrian/OCR-SPRIN-SERVICE

Fork 0

Commit Graph

Author	SHA1	Message	Date
Devin AI	45fbfdabb7	Phase 5: hybrid LLM extraction (Ollama) for header gaps Adds a small Ollama HTTP client (httpx-based, no extra runtime deps), prompt builders, and a hybrid header extractor that runs after the deterministic regex layer. The merger never overwrites a regex-filled field — the LLM only fills gaps. If LLM_ENABLED=false (the default), or the Ollama server is unreachable, the pipeline degrades gracefully: - LLM_ENABLED=false -> no LLM call at all, no flag. - LLM_ENABLED=true, header complete -> no LLM call. - LLM_ENABLED=true, header has gaps, LLM responded ok -> merge + LLM_FALLBACK flag (review hint). - LLM_ENABLED=true, header has gaps, LLM unavailable -> keep regex result + LLM_UNAVAILABLE flag. Default model qwen2.5:1.5b on http://localhost:11434 — chosen for CPU throughput (~5-15s per call) at acceptable accuracy. The LLM only fills the header (nomor, tanggal, satuan, perihal, dasar). Personnel rows stay with PP-Structure since that's more accurate and doesn't need LLM. Tests: - test_llm_client.py: httpx MockTransport-driven tests for the wire format, error paths (HTTP 5xx, malformed JSON, missing envelope, ConnectError), and request shape. - test_llm_extractor.py: merge policy + None-on-unavailable behaviour. - test_orchestrator_llm.py: end-to-end orchestrator wiring with stubs for ingest/preprocess/OCR/table — verifies LLM is skipped when disabled, skipped when header is complete, called and flagged when gaps exist, and marked unavailable when the client returns None. 162 unit tests pass total (was 146). Co-Authored-By: adrian kuman firmansah <adriancuman@gmail.com>	2026-04-25 16:56:43 +00:00

Author

SHA1

Message

Date

Devin AI

45fbfdabb7

Phase 5: hybrid LLM extraction (Ollama) for header gaps

Adds a small Ollama HTTP client (httpx-based, no extra runtime deps),
prompt builders, and a hybrid header extractor that runs *after* the
deterministic regex layer. The merger never overwrites a regex-filled
field — the LLM only fills gaps. If LLM_ENABLED=false (the default), or
the Ollama server is unreachable, the pipeline degrades gracefully:

  - LLM_ENABLED=false  ->  no LLM call at all, no flag.
  - LLM_ENABLED=true,
    header complete    ->  no LLM call.
  - LLM_ENABLED=true,
    header has gaps,
    LLM responded ok   ->  merge + LLM_FALLBACK flag (review hint).
  - LLM_ENABLED=true,
    header has gaps,
    LLM unavailable    ->  keep regex result + LLM_UNAVAILABLE flag.

Default model qwen2.5:1.5b on http://localhost:11434 — chosen for CPU
throughput (~5-15s per call) at acceptable accuracy. The LLM only fills
the *header* (nomor, tanggal, satuan, perihal, dasar). Personnel rows
stay with PP-Structure since that's more accurate and doesn't need LLM.

Tests:
 - test_llm_client.py: httpx MockTransport-driven tests for the wire
   format, error paths (HTTP 5xx, malformed JSON, missing envelope,
   ConnectError), and request shape.
 - test_llm_extractor.py: merge policy + None-on-unavailable behaviour.
 - test_orchestrator_llm.py: end-to-end orchestrator wiring with stubs
   for ingest/preprocess/OCR/table — verifies LLM is skipped when
   disabled, skipped when header is complete, called and flagged when
   gaps exist, and marked unavailable when the client returns None.

162 unit tests pass total (was 146).

Co-Authored-By: adrian kuman firmansah <adriancuman@gmail.com>

2026-04-25 16:56:43 +00:00

1 Commits