Skip to main content

Ollama Port-11434 Race — com.alai.ollama-serve-v2 vs homebrew.mxcl.ollama (MC #9813)

Ollama Port-11434 Race — com.alai.ollama-serve-v2 vs homebrew.mxcl.ollama (MC #9813)

Owner: Proveo (QA finding) → FlowForge (fix, MC #106486) | Found: 2026-07-28 | MC: #9813, follow-up #106486 Source of truth: live launchctl print + ps eww on ANVIL, /opt/homebrew/var/log/ollama.log


1. What #8519 claimed

MC #8519 ("Proveo E2E validation: Ollama contention fix") passed on 2026-05-05 with DoD:

"Proveo L2 PASS: llama3.1:8b expires 2318, inference 3.2s, MAX_LOADED_MODELS=2 in plist, no pattern fallback."

The fix: ~/Library/LaunchAgents/com.alai.ollama-serve-v2.plist sets OLLAMA_MAX_LOADED_MODELS=2 (plus OLLAMA_HOST=127.0.0.1:11434 and a restricted OLLAMA_ORIGINS), intended to cap concurrently loaded models and stop Ollama memory/contention thrash on the orchestrator Mac. This claim was accurate at the time — not a fabrication.

2. What the #9813 QA re-check (2026-07-28) found — REGRESSION

Two LaunchAgents both try to bind port 11434:

Agent State Notes
com.alai.ollama-serve-v2 spawn scheduled, last exit code 1 Has the hardened env (MAX_LOADED_MODELS=2) but never successfully binds.
homebrew.mxcl.ollama running, live PID Homebrew's default RunAtLoad service. Wins the race at every boot. Env = only OLLAMA_FLASH_ATTENTION=1, OLLAMA_KV_CACHE_TYPE=q8_0no MAX_LOADED_MODELS cap at all.

/opt/homebrew/var/log/ollama.log:

Error: listen tcp 127.0.0.1:11434: bind: address already in use

Net effect: the actual Ollama traffic on this host is served by the un-hardened homebrew.mxcl.ollama service. The #8519 contention fix is not in effect — it is correctly configured on disk but the LaunchAgent carrying it never wins the port.

3. Fix (tracked in follow-up MC #106486, not applied by this QA pass — Proveo is read-only)

One of:

  • (a) launchctl bootout gui/$(id -u)/homebrew.mxcl.ollama + remove/rename its plist so com.alai.ollama-serve-v2 owns the port, or
  • (b) Port OLLAMA_MAX_LOADED_MODELS=2 + the other hardened env vars (from B5, MC #104484) directly into homebrew.mxcl.ollama.plist, retire com.alai.ollama-serve-v2, then launchctl kickstart -k.

4. Not re-verified / open gaps

  • Whether contention actually recurs in practice under 3+ concurrently loaded models was not reproduced live in this pass.
  • com.john.ollama-guard (ollama-memory-guard.js, 60s interval, "max 1 model active, auto-unload on memory pressure") may be independently mitigating symptoms — not evaluated here.

5. Process lesson

This evidence directory (~/system/evidence/9813/) briefly contained contradictory artifacts: an accurate qa-review-8519.md (FAIL) alongside a summary.json/ verification.json claiming PASS, written by an earlier failed autowork attempt in the same session, based on a non-probative check ("ollama ps shows ≤2 models loaded" — true by coincidence of current usage, not because any cap was enforced). Corrected to FAIL across all three files after independent live re-verification. Lesson: a model-count snapshot from ollama ps is not evidence a cap is enforced — check which process is actually bound to the port and its live env, not just the config file.