Skip to main content

RAG Queue Recovery — MC #106978

RAG Queue Recovery — MC #106978

Status

S1–S4 are complete. The S5 offline package was rebuilt by MC #107082 for the reviewed d38f... client successor and exact 508e... server. It remains offline, has no authority receipt or UTC window, and is unauthorized for production effects. S6 requires a 19-day recurrence soak after a separately authorized successful G4; S7 closes the postmortem only after Proveo PASS.

Data-accounting result

  • 47,322 source records are accounted for.
  • 14,003 were recovered, 3 were replayable, 22,336 were semantic duplicates, and 10,980 malformed records remain byte-preserved.
  • All 82 uncertain records have evidence-backed dispositions.
  • Unexplained loss: 0. Irretrievable loss: 0.
  • The untrusted historical generation is custody evidence only and is never a service or rollback source.

Frozen S5 inputs

  • Final client: d38f399067113dc1b29e38a315442362e3f89ee2, direct successor of 33d6e30545bdc6de171732a069316c6a05ec5f9e (patch 9b94181cce3c30f8be3946f7f0eb91bee5768deb0a51ff5a08450071c960f7fe).
  • Final server endpoint: 508e618b6c071564e07bcea3530c6826c5c142b9.
  • Preactivated candidate: f26fae393e46e40d01721f9423fc8036cea250a89570d4fb5619638a78d61740.
  • Safe rollback: f50eeb7a3d8d575a016f3f00eb906417aa275f529842d32d1b424045b07185b6.
  • Candidate accounting: 112 receipt-backed accepted rows and 13,894 untouched queued rows.
  • Receipt set: 112 private schema-v3 records bound to the actual public runtime endpoint.

Independent round-13 QA, a separate reliability witness, native task-bound P2P, and clean post-commit verification all passed for d38f... (294/294 tests under both umask 0022 and 0077). These are offline readiness verdicts, not deployment authority. The actual qualified topology remains local Docker through the Cloudflare tunnel; the Azure VM is deallocated and the historical Azure body below its supersession notice remains documentation drift.

G4 sequence

  1. Record exact CEO authority and a maximum 45-minute UTC control window.
  2. Deploy the exact server image first; prove authentication, one bounded document-ID lookup, no enumeration, and deterministic rollback.
  3. Coherently deploy all staged runtime/helper bytes: the five changed paths (agents/hivemind/hivemind.js, lib/rag-db-writer-lock.js, lib/rag-outbox.js, tools/lightrag-auth-helper.js, and tools/rag-drain-worker.js) and their unchanged staged receipt/spool/schema/config/HiveMind dependencies. Publish the private schema-v3 receipts and preactivated candidate with descriptor, file, and parent-directory fsync.
  4. Only after the future authority is consumed, quarantine legacy system/config/.lightrag-token-cache.json if it is a UID-owned, mode-0600, single-link regular file. The runtime cache policy is exclusively state/runtime-auth/lightrag-token-cache.json under a mode-0700 parent; absence is valid and any present file must be UID-owned, mode 0600, single-link, descriptor-bound. Rollback restores both cache prestates byte-for-byte.
  5. Reconcile the 112 accepted rows using authenticated bounded GET only. This phase hard-fences every POST path, fallback, history/global-artifact access, global permit-state access, and the optional backlog nudge.
  6. Start the watchdog and require a hard local-audit acknowledgement plus a side-effect-free HiveMind local-gate-ack.
  7. After the future receipt/window exists, generate exactly three private selector files, one each for Mission Control, filesystem, and BookStack. Each binds task #107020, the authority digest, exact delivery, and the maximum 45-minute window. Reserve the exact DB row before one-use selector consumption; never put selector or delivery identity in argv or public evidence.
  8. Start the persistent drain at no more than 20 POSTs per rolling 60 seconds, then restore the three source adapters serially. Login, nudge, Slack, canary, adapter, replay, and drain POSTs share the same durable permit state across process restart. Local audit and HiveMind are hard; Slack is retry-soft only under the literal policy. The superseded legacy outbox-ingest job remains disabled.
  9. Replay the seven authoritative MC outcomes plus a mandatory final structured delta sweep.

Terminal semantics

  • Only exact PROCESSED may mark a row done.
  • Exact FAILED is terminal.
  • Absent, malformed, oversized, unavailable, PENDING, and PROCESSING remain accepted and are never posted again.
  • MD5 is used only to derive the deployed server lookup key; SHA-256 content and delivery identities remain semantic authority.

Rollback

Any wrong hash/image/router, unauthorized method, accepted-only POST, queued/history/global mutation, activation or auth-cache-policy failure, writer overlap, SQLite failure, missing audit/HiveMind acknowledgement, canary reservation/consumption failure or duplicate, privacy leak, unexplained replay delta, or window expiry triggers rollback. Stop all restored jobs, preserve the failed main/WAL/SHM coherently, restore exact f50e..., restore all prior staged-path bytes, both cache prestates, receipts, and the pinned prior server image, then prove descriptor, fsync, offline/job, router, and health fences. Never use the untrusted historical generation.

Monitoring and closure

After a separately authorized successful S5, Proveo owns S6 verification for the required 19-day recurrence interval. The soak repeatedly checks immutable snapshots, queue progress, one-writer topology, durable permit capacity, daemon health, corruption signatures, auth-cache privacy, and downstream canaries. Main task #106978 remains open until S6 and S7 pass without force. BookStack publication remains blocked/drifted; this local file is canonical until an authenticated publication is separately authorized and verified.