Distributed Factory

ANVIL/FORGE distributed AI factory — host resolution, bootstrap, rollout

OLLAMA_HOST Resolution Strategy — ALAI Fleet (MC #8476)

OLLAMA_HOST Resolution Strategy — ALAI Fleet

Owner: Skillforge | Verified: 2026-07-28 | Re-verified: 2026-08-16 | MC: #8476 Source of truth: ~/system/tools/ollama-host.sh (implemented 2026-04-20, CodeCraft) + ~/system/architecture/distributed-ai-factory-plan.md §4


1. Why This Exists

ALAI agent scripts call Ollama for local inference (classifiers, embeddings, session summarization). Different machines reach Ollama differently:

ollama-host.sh centralizes this so scripts don't hardcode an endpoint. Instead they read $OLLAMA_HOST (or, in JS, process.env.OLLAMA_HOST), which the wrapper resolves and exports once per shell session.


2. Resolution Order (as implemented in ollama-host.sh)

The live script's actual order is cache-first, then hostname, then a network probe ladder — this is slightly more defensive than the original architecture-plan pseudocode (§4 of distributed-ai-factory-plan.md), which omits the cache and the local-probe step. Documenting the real script here:

1. Explicit override   — if $OLLAMA_HOST is already set in the environment, use it
                          as-is. No probing. (Lets a human or CI job pin an endpoint.)
2. Cache check          — if /tmp/ollama-host.cache exists and is < 5 min old:
                            - cached value present and != "NONE"  → export it, done
                            - cached value == "NONE"              → warn, leave unset, done
3. Hostname match       — if `hostname` contains "makinja-sin-mac-studio" or
                          "makinja.local" (i.e. we ARE ANVIL) → localhost:11434
4. Local probe          — curl localhost:11434/api/version, 1s timeout
                          (catches "Ollama running locally for some other reason")
5. Tailscale ANVIL      — curl http://100.103.49.98:11434/api/version, 2s timeout
6. FORGE (LAN only)     — curl http://10.0.0.2:11434/api/version, 1s timeout
                          (fast-fails on non-LAN hosts, e.g. remote/cloud machines)
7. Graceful degrade     — write "NONE" to the cache, print one WARNING to stderr,
                          leave $OLLAMA_HOST unset. Agents with a Claude API fallback
                          continue; Ollama-only agents fail explicitly downstream
                          rather than hanging on an unreachable host.

Every successful resolution (steps 3–6) writes its result to the cache file before exporting, so the next shell in the same 5-minute window skips the probe ladder entirely (step 2 short-circuits).


3. Cache: /tmp/ollama-host.cache

ollama_flush_cache()

ollama_flush_cache() {
  rm -f "$_OLLAMA_CACHE_FILE"
  echo "[ollama-host] Cache cleared. Run 'source ~/system/tools/ollama-host.sh' to re-resolve."
}

Deletes /tmp/ollama-host.cache and prints a reminder to re-source. Call this manually after:

It is exported as a shell function (export -f ollama_flush_cache) so it's callable from any subshell in the sourcing session, not just interactively.

There is also ollama_available(), a one-line guard scripts can call before an Ollama-only code path:

ollama_available() {
  [[ -n "${OLLAMA_HOST:-}" ]] && [[ "${OLLAMA_HOST}" != "NONE" ]]
}

4. Consumption in Node/JS Agent Scripts

Scripts that were refactored off hardcoded endpoints read the resolved value via process.env.OLLAMA_HOST (the shell wrapper exports it, so any JS process spawned from a shell that sourced ollama-host.sh inherits it).

Verified counts (grep -rlE over ~/system, .js/.sh/.ts/.mjs):

Date Files reading process.env.OLLAMA_HOST
2026-04-20 (task description estimate, unverified) 52
2026-07-28 (first live grep) 149
2026-08-16 (re-verified live grep) 848

⚠️ Correction to MC #8476 task description: the task text (written 2026-04-20, the day ollama-host.sh was implemented) states "52 refactored files now read process.env.OLLAMA_HOST". No historical record (evidence/reports) of a "52 files" sign-off was found via discover.js search, so that original number cannot be reconstructed or verified — treat it as a same-day estimate, not a target to reconcile against.

The count itself is not stable — it grew 149 → 848 (5.7x) between 2026-07-28 and 2026-08-16, three weeks apart, tracking the system's overall growth rate rather than a single discrete refactor. Practical implication: do not cite an absolute count from this runbook as current without re-running the grep below; cite the trend (rapid, ongoing adoption of the wrapper pattern) instead.

grep -rlE "process\.env\.OLLAMA_HOST" ~/system --include="*.js" --include="*.sh" --include="*.ts" --include="*.mjs" | wc -l

Not every hardcoded reference is necessarily a bug: some are inside ollama-host.sh itself (the probe URLs), health-probe/monitoring scripts that intentionally target a specific host, or docs/comments. A hardcode count is a signal for follow-up, not an automatic defect list — each hit needs eyeballing before "refactor" is filed as a task.


5. Shell Integration

Expected integration (per ollama-host.sh header comment):

[ -f ~/system/tools/ollama-host.sh ] && source ~/system/tools/ollama-host.sh

added to ~/.zshrc.

⚠️ Verified gap (re-checked 2026-08-16, unchanged since 2026-07-28): ~/.zshrc on this machine (ANVIL) still has no reference to ollama-host.sh or "ollama" at all (grep -in ollama ~/.zshrc returns nothing, 158-line file). The wrapper is NOT currently auto-sourced on interactive shell start. It must be invoked explicitly (source ~/system/tools/ollama-host.sh) or sourced by whatever launches an agent process today.

This is a real gap, not a documentation omission — filed as a follow-up rather than silently assumed fixed. See §6. Note: a different overlay mechanism exists (~/system/config/role-overrides/workstation/env.sh, added by the workstation role's bootstrap.sh for non-ANVIL machines like a Mac Air test host — MC #8562) which sets its own OLLAMA_HOST and is appended to ~/.zshrc/~/.bashrc by that bootstrap path, but only when the workstation role is provisioned. It targets a stale fallback IP (100.104.164.86 for FORGE, offline since ~2026-06-14 per ~/system/CLAUDE.md) and is a separate code path from ollama-host.sh — do not conflate the two when auditing shell integration.


6. Follow-ups Identified While Writing This Runbook

Ollama Port-11434 Race — com.alai.ollama-serve-v2 vs homebrew.mxcl.ollama (MC #9813)

Ollama Port-11434 Race — com.alai.ollama-serve-v2 vs homebrew.mxcl.ollama (MC #9813)

Owner: Proveo (QA finding) → FlowForge (fix, MC #106486) | Found: 2026-07-28 | MC: #9813, follow-up #106486 Source of truth: live launchctl print + ps eww on ANVIL, /opt/homebrew/var/log/ollama.log


1. What #8519 claimed

MC #8519 ("Proveo E2E validation: Ollama contention fix") passed on 2026-05-05 with DoD:

"Proveo L2 PASS: llama3.1:8b expires 2318, inference 3.2s, MAX_LOADED_MODELS=2 in plist, no pattern fallback."

The fix: ~/Library/LaunchAgents/com.alai.ollama-serve-v2.plist sets OLLAMA_MAX_LOADED_MODELS=2 (plus OLLAMA_HOST=127.0.0.1:11434 and a restricted OLLAMA_ORIGINS), intended to cap concurrently loaded models and stop Ollama memory/contention thrash on the orchestrator Mac. This claim was accurate at the time — not a fabrication.

2. What the #9813 QA re-check (2026-07-28) found — REGRESSION

Two LaunchAgents both try to bind port 11434:

Agent State Notes
com.alai.ollama-serve-v2 spawn scheduled, last exit code 1 Has the hardened env (MAX_LOADED_MODELS=2) but never successfully binds.
homebrew.mxcl.ollama running, live PID Homebrew's default RunAtLoad service. Wins the race at every boot. Env = only OLLAMA_FLASH_ATTENTION=1, OLLAMA_KV_CACHE_TYPE=q8_0no MAX_LOADED_MODELS cap at all.

/opt/homebrew/var/log/ollama.log:

Error: listen tcp 127.0.0.1:11434: bind: address already in use

Net effect: the actual Ollama traffic on this host is served by the un-hardened homebrew.mxcl.ollama service. The #8519 contention fix is not in effect — it is correctly configured on disk but the LaunchAgent carrying it never wins the port.

3. Fix (tracked in follow-up MC #106486, not applied by this QA pass — Proveo is read-only)

One of:

4. Not re-verified / open gaps

5. Process lesson

This evidence directory (~/system/evidence/9813/) briefly contained contradictory artifacts: an accurate qa-review-8519.md (FAIL) alongside a summary.json/ verification.json claiming PASS, written by an earlier failed autowork attempt in the same session, based on a non-probative check ("ollama ps shows ≤2 models loaded" — true by coincidence of current usage, not because any cap was enforced). Corrected to FAIL across all three files after independent live re-verification. Lesson: a model-count snapshot from ollama ps is not evidence a cap is enforced — check which process is actually bound to the port and its live env, not just the config file.