No description
  • Python 38.7%
  • Nix 36.2%
  • HTML 13.3%
  • Shell 6.9%
  • JavaScript 4.2%
  • Other 0.7%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-08-12 10:01:43 -07:00
lora-configs scrub 2026-07-29 20:16:50 -07:00
patches updates 2026-08-04 22:48:59 -07:00
templates model updates 2026-07-20 16:29:50 -07:00
test-data fixes 2026-08-04 18:01:05 -07:00
.gitignore added models and lora infra 2026-07-21 01:20:19 -07:00
admin.html mcp fixes; admin ui fixes 2026-08-11 00:45:33 -07:00
AGENTS.md Storage refactor: unified model store, nothing in home folders, declarative manifests 2026-07-21 12:53:28 -07:00
CLAUDE.md librechat: a definition, so the deployment can choose the machine 2026-08-10 18:48:26 -07:00
code-sandbox-server.py updates 2026-07-30 22:47:19 -07:00
code-sandbox.sh updates 2026-07-30 22:47:19 -07:00
code.html kiosk viz working 2026-08-04 21:07:25 -07:00
comfyui.sh move the content policy out, and grow the seams it needed 2026-08-10 10:02:50 -07:00
cute-dog.jpg edit_image: normalize input to PNG via ffmpeg (sd.cpp can't load WebP) 2026-07-18 14:37:27 -07:00
dashboard-server.py added rssmonster 2026-08-12 09:27:10 -07:00
docs.nix studio refactor 2026-08-03 18:03:59 -07:00
doctor.sh added remote model cache 2026-08-04 10:26:58 -07:00
embed-media.py move the content policy out, and grow the seams it needed 2026-08-10 10:02:50 -07:00
eval-runner.py final phase completed 2026-07-31 18:02:57 -07:00
evals.html final phase completed 2026-07-31 18:02:57 -07:00
faceswap-video-workflow.json added remote model cache 2026-08-04 10:26:58 -07:00
faceswap-workflow.json face swap updates 2026-07-31 00:09:46 -07:00
flake.lock moved out personal config 2026-08-06 16:07:06 -07:00
flake.nix librechat: the MCP hop needs the identity header too, and the default model is not ours to name 2026-08-10 19:50:02 -07:00
gallery.html several fixes 2026-08-06 01:40:20 -07:00
gallery.js studio fixes 2026-08-10 21:08:12 -07:00
genai-fetch-media.sh move the content policy out, and grow the seams it needed 2026-08-10 10:02:50 -07:00
gpu-claim.sh added gpu lock 2026-08-09 17:57:11 -07:00
gpu.html style updates 2026-08-09 18:13:36 -07:00
hand-tuned.nix fixes 2026-08-04 22:23:06 -07:00
HARDWARE.md model updates 2026-08-03 00:40:55 -07:00
harness-models.nix fixes 2026-08-04 22:23:06 -07:00
hwscan.sh final phase completed 2026-07-31 18:02:57 -07:00
image-server.py offer the LoRAs the config declares, not the ones on disk 2026-08-10 11:32:14 -07:00
kiosk-daemon.py kiosk viz working 2026-08-04 21:07:25 -07:00
librechat.nix librechat: tell it how big the context is, or it sends no messages at all 2026-08-10 21:01:00 -07:00
llama-swap.nix llama-swap: actually build the dashboard 2026-08-01 02:20:41 -07:00
lora-add.sh move the content policy out, and grow the seams it needed 2026-08-10 10:02:50 -07:00
lora-caption.py lora updates 2026-07-20 15:15:49 -07:00
lora-list.sh generalized 2026-07-27 23:04:50 -07:00
lora-train.sh move the content policy out, and grow the seams it needed 2026-08-10 10:02:50 -07:00
ltx23-i2v-workflow.json model fixes 2026-08-09 23:45:53 -07:00
ltx23-lipsync-workflow.json model fixes 2026-08-09 23:45:53 -07:00
ltx23-reference-workflow.json model fixes 2026-08-09 23:45:53 -07:00
ltx23-t2v-workflow.json model fixes 2026-08-09 23:45:53 -07:00
magentic-ui.sh added magentic 2026-07-23 20:08:59 -07:00
mcp-gateway-server.py mcp fixes; admin ui fixes 2026-08-11 00:45:33 -07:00
media-tools-server.py mcp fixes; admin ui fixes 2026-08-11 00:45:33 -07:00
memory-server.py updates 2026-07-30 22:47:19 -07:00
memory.html kiosk viz working 2026-08-04 21:07:25 -07:00
minimax-h3-i2v-workflow.json studio refactor 2026-08-03 18:03:59 -07:00
minimax-h3-r2v-workflow.json studio refactor 2026-08-03 18:03:59 -07:00
minimax-h3-t2v-workflow.json studio refactor 2026-08-03 18:03:59 -07:00
minimax-h3-v2v-workflow.json added remote model cache 2026-08-04 10:26:58 -07:00
models.html kiosk viz working 2026-08-04 21:07:25 -07:00
module.nix removed api auth 2026-08-12 10:01:43 -07:00
narrate.sh final phase completed 2026-07-31 18:02:57 -07:00
ollama-shim-server.py upgrades 2026-07-30 16:26:47 -07:00
open-webui-seed.py per-person tool menus in Open WebUI, without a rebuild in the loop 2026-08-10 17:38:06 -07:00
options.nix added rssmonster 2026-08-12 09:27:10 -07:00
prefetch.sh model fixes 2026-08-09 23:45:53 -07:00
profiles.nix model updates 2026-08-03 00:40:55 -07:00
prune.sh added remote model cache 2026-08-04 10:26:58 -07:00
pulid-flux-workflow.json Storage refactor: unified model store, nothing in home folders, declarative manifests 2026-07-21 12:53:28 -07:00
pulid-swap-full-workflow.json updates 2026-07-30 21:15:36 -07:00
pulid-swap-workflow.json upgrades 2026-07-29 20:08:11 -07:00
rag-server.py updates 2026-07-30 21:15:36 -07:00
rag.html kiosk viz working 2026-08-04 21:07:25 -07:00
README.md studio fixes 2026-08-10 21:08:12 -07:00
ROADMAP.md voice updates 2026-07-31 21:47:06 -07:00
sampling.nix model fixes 2026-08-09 23:45:53 -07:00
search-tool-server.py Add web search: SearXNG + OpenAPI search tool server, wire into Open WebUI 2026-07-17 14:57:19 -07:00
segment-server.py image model updates 2026-07-25 00:55:46 -07:00
store-overlay.sh mediaStore.overlays: the same split for renders 2026-08-10 13:42:26 -07:00
store-sync.sh added remote model cache 2026-08-04 10:26:58 -07:00
studio.html studio fixes 2026-08-10 21:08:12 -07:00
TODOs.md store overlays: a subtree that is part of the store only while claimed 2026-08-10 13:16:36 -07:00
transcribe-video.sh final phase completed 2026-07-31 18:02:57 -07:00
tts-hq-server.py Free GPU VRAM when media services idle (fixes LLM OOM / 502s) 2026-07-18 13:49:17 -07:00
tts-server.py SOTA audit implementation: HQ TTS + cloning, music, media services 2026-07-17 20:51:27 -07:00
tune-runner.py autoresearch updates 2026-08-06 12:48:26 -07:00
tuning.md several fixes 2026-08-06 01:40:20 -07:00
viz.html viz updates 2026-08-06 19:11:56 -07:00
vm-test-wyoming.nix voice updates 2026-07-31 21:47:06 -07:00
vm-test.nix librechat: on by default, and name the unit the runtime actually makes 2026-08-10 16:21:40 -07:00
voice-server.py added voice functionality 2026-07-31 19:50:49 -07:00
voice.html kiosk viz working 2026-08-04 21:07:25 -07:00
wan22-i2v-workflow.json added remote model cache 2026-08-04 10:26:58 -07:00
wan22-t2v-workflow.json added remote model cache 2026-08-04 10:26:58 -07:00
wyoming-openai.nix final phase completed 2026-07-31 18:02:57 -07:00

genai-server

NixOS flake for a local GenAI serving stack, sized for a 32GB NVIDIA GPU + 128GB RAM (built for logistikon: RTX 5090 / Ryzen 9700X).

The portal — start here

http://<host>:8897 (replace <host> with e.g. logistikon.lan) is the one address worth bookmarking. It is the stack's front door and its monitoring surface: what's wrong, what the hardware is doing, what you can open, and what models are on disk — on one page, with a shared nav across every page it hosts.

Portal page What
/ Overview — problems strip, live GPU/CPU/RAM gauges, the service directory (grouped, with up/down + health dots), and a fleet summary
/viz 3D system view — live hardware and data-flow visualization
/studio Studio — make and change pictures, video and sound; one form over image-server, the media tools and published ComfyUI workflows
/rag Knowledge — document collections: search, ingest by URL or paste, inspect and delete
/code Code sandbox — run Python / JavaScript / shell in a disposable VM, with input and output files
/memory Memory — the durable facts models carry between conversations: search, add, forget, and the reconciliation queue
/models Model fleet — every LLM with its capabilities, context, KV quant, offload, group and TTL; filter, download, enable, remove

It also fronts the APIs, so a client needs one host:port for the whole box:

Route What
/v1/* OpenAI API → llama-swap, filtered to ready-and-enabled models
/svc/<name>/* this flake's own tool servers: image, media, search, tts, segment, rag, memory, code, mcp — e.g. /svc/media/openapi.json
/api/health one call: {ok, problems[], services, gpus, cpu, mem, models} (503 when unhealthy, so curl -f works)
/api/portal the portal surface — hosted pages + grouped service directory
/api/activity one line: what the box is working on right now ({activity, model, kind, detail})
/metrics Prometheus: GPU/CPU/RAM, per-service up + health, per-model ready/enabled/size

Hosts extend it through portal.extraPages, portal.extraServices and portal.proxy.extraUpstreams — see Usage. Anything this flake adds in future registers here rather than becoming another bookmark (ROADMAP.md Stage 2a, "one front door").

Web UIs

The third-party interfaces the portal links to (all reachable directly, and all probed on the portal's Overview page):

URL What
http://<host>:3000 Open WebUI — main chat: LLMs, vision, web search, mic/read-aloud, image generation, agentic media tools
http://<host>:8188 ComfyUI — node-graph image / video / music generation (FLUX.2, Wan 2.2, LTX 2.3, ACE-Step)
http://<host>:8895 MagenticLite — human-in-the-loop web agent (Fara drives a QEMU-sandboxed browser; approve/steer/take over at any point)
http://<host>:8888 SearXNG — private metasearch

Minor built-in pages that ship with their tools: :8080 llama-swap model manager (its Transcription tab uploads audio to the asr model), :4000 LiteLLM admin. The remaining ports (:8892/8893/8894/8896, and the search/tool servers) are JSON APIs with no browser UI — reach them through /svc/<name>/ rather than by port.

Services (all ports)

Port Service What
8080 llama-swap OpenAI-compatible LLM API + web UI; hot-swaps models on demand
8897 portal the front door: pages (/, /viz, /studio, /models, /rag, /memory, /code, /evals), service directory + health (/api/health), model admin, and the API proxies (/v1/* → llama-swap, /svc/<name>/* → the tool servers below)
3000 Open WebUI chat frontend wired to llama-swap
8188 ComfyUI image/video/music generation (systemd service, genai user; version pinned by comfyui.rev; the bulk model sets come from comfyui.modelSets)
8888 SearXNG local metasearch; JSON API via /search?q=...&format=json
8891 search tool OpenAPI web-search tool server (SearXNG backend)
4000 LiteLLM Anthropic-protocol bridge (/v1/messages) for Claude Code etc.
8892 tts-server OpenAI-compatible /v1/audio/speech; routes Piper / Kokoro / Chatterbox by voice
8893 image-server image gen (/generations, engines from imageServer.engines: Z-Image default, plus "model": "krea-2" / "flux-dev" / "hidream-o1" (in-image text; "hidream-o1-dev" for the distilled draft build) / "pony" / "epicrealism" / "juggernaut") + edit (/edits, engines from imageServer.editEngines: Kontext default, "model": "qwen-edit" for text/identity that has to survive repeated passes), reimagine (/reimagine schnell), inpaint (/inpaint, masked region only — source composited back outside the mask; engines from imageServer.inpaintEngines: Fill default, "model": "pony" stylized SDXL-inpainting); optional "seed" on every sampling path (omit for random; the seed used is always resolved server-side, returned as data[i].seed and recorded in the genai chunk, so any render can be repeated after the fact — a batch walks it, seed, seed+1, …); LoRAs via <lora:name:0.9> prompt tags on every path (on /generations a registered LoRA's base auto-selects the engine; the source-image paths each run one fixed model, so a LoRA from another architecture is refused there instead of rerouted); the three source-image paths render at the source's resolution, so a larger one is downscaled to imageServer.editMaxPixels (2MP) first — per-request "max_pixels" overrides that up to editMaxPixelsCeiling (4MP), and /inpaint still returns the full-resolution original outside the mask; a render is killed when its requester disconnects (Open WebUI's stop button) or via POST /cancel
8894 media-tools OpenAPI tool server: generate_image (with installed subject LoRAs), generate_image_with_face (PuLID via ComfyUI), swap_face / swap_face_full (PuLID onto an existing photo — masked face region, or the whole frame), swap_face_fast (ReActor/inswapper pixel transplant, seconds), edit_image, reimagine_image, inpaint_image, smart_edit (auto-mask: SAM 3 segment + inpaint), create_mask (segmentation preview: red-tint overlay + reusable mask), transcribe_audio, text_to_speech; an optional seed on all eight sampling tools, always reported back in the result so a chat model can answer "same picture but …"; lora/loras/lora_strength on every tool that runs a diffusion model — generate_image, the four edit tools and the two PuLID face tools (generate_image_with_face, swap_face_full), each offering only the LoRAs its model can bind, and the face tools splicing ComfyUI loader nodes rather than prompt tags; a photo that gets re-sampled as a latent (the edit tools' source, the two PuLID swaps' target) is downscaled to imageServer.editMaxPixels first — the identity/reference photos are not — with max_pixels to spend more on one render and swap_face's full_resolution to composite the face back over the untouched original; artifacts under /files/
8895 MagenticLite web-agent UI (magentic-ui 0.2.x): Fara browser-use + orchestrator, browser sandboxed in a Quicksand QEMU micro-VM (KVM-accelerated, no Docker). The live browser view (noVNC, per-session password) uses ports 8860-8879, firewalled alongside the service ports
8896 tts-hq internal HQ TTS backend (Kokoro narration, Chatterbox cloning)
8903 code-sandbox run Python / JavaScript / shell in a throwaway QEMU microVM (exec-sandbox on KVM): no host filesystem, no state between runs, no network by default. openapi.json so chat and MCP both get it (codeSandbox.enable)
8902 memory-server durable cross-session memory: remember / recall / forget, reinforcement and near-duplicate consolidation. Namespaced by a trusted identity header (memory.enable)
8900 rag-server knowledge collections: hybrid BM25 + semantic search, ingest by text/URL/path, openapi.json so Open WebUI and the MCP gateway pick it up. State in /var/lib/genai-rag (rag.enable)
8899 mcp-gateway MCP over Streamable HTTP at /mcp — every tool server's OpenAPI operation as an MCP tool (web search, image gen/edit, segmentation, speech). GET /tools is the resolved inventory. Also at :8897/svc/mcp/mcp (mcp.enable)
11434 ollama-shim Ollama API dialect — /api/tags, /api/show, /api/chat, /api/generate, /api/embed, /api/ps, /api/pull, NDJSON streaming. Translates to the portal, so Ollama-only clients (Home Assistant, JetBrains, mobile apps) use the same fleet. Models stay declarative: pulling a name that isn't in llmModels errors, and create/copy/push/delete answer 501 (ollama.enable)
10300 wyoming-openai Wyoming protocol bridge for Home Assistant (STT + TTS). Off by default; adds no models — forwards to the asr model and tts-server. No authentication, so wyoming.openFirewall is separate from the global one
8901 voice-server Realtime speech-to-speech over WebSocket (/ws), plus /health. Off by default. The chat model runs on the CPU, so this service uses no VRAM
8898 segment-server SAM 3 text-prompted mask generation (/segment); the auto-mask half of smart_edit, CPU-resident. Weights are license-gated: accept at hf.co/facebook/sam3 and set hfTokenFile

STT is served OpenAI-style at :8080/v1/audio/transcriptions — llama-swap's asr model (Qwen3-ASR-1.7B, ggml-org GGUF; also answers to the whisper-1 alias), so it swaps and unloads like any other model instead of running as a second daemon. response_format must be json: llama-server does not implement the text/srt/vtt variants whisper.cpp had, so subtitle files need a separate conversion step. Open WebUI is wired to all three (mic button = asr, read-aloud = Piper, image toggle in the chat input = FLUX.1-dev at ~20 steps — the quality default, ~40-60s/image; ask the chat for a Z-Image generation when you want the ~5s fast path); the seeding service enforces the engine config. The image server unloads llama-swap models before generating (full GPU), so expect the next LLM request to re-load for a few seconds.

Ports are opened only on firewallInterfaces (default tailscale0); loopback always works.

Models (select by name via the API / UI)

Name Model Footprint / speed Reach for it when
coder-pro Qwen3-Coder-Next 80B-A3B card + experts in host RAM, ~20-35 tok/s, 256k ctx The Claude Code / OpenCode model. Purpose-trained for long-horizon agent loops and recovery from failed steps, and the only model here whose window clears Claude Code's assumed ~200k — so long sessions auto-compact instead of overflowing.
ds4-flash DeepSeek-V4-Flash-0731 284B-A13B ~19GB VRAM + ~80GB RAM, ~12-16 tok/s generation (estimated, not yet measured), 128k ctx The agentic escalation tier: a 284B model with a 128k window that costs almost no KV, tuned for tool use and long agent loops rather than for single answers. Reach for it when coder-pro has lost the thread on a long session — not for a quick question, because it generates at roughly 1/20th of qwen.
fara Fara1.5-27B (computer-use agent) ~24GB VRAM (~17.5GB weights + ~1GB mmproj), ~45 tok/s, 128k ctx Browser automation from screenshots. NOT a chat or coding model — drive it with fara-cli or Magentic-UI pointed at this endpoint. Picking it in a chat window can only fail.
fast-cpu Qwen3-4B-Instruct-2507 with tools (on the CPU) zero VRAM, ~2.5GB host RAM, 22 tok/s generate / 223 tok/s prefill measured on 8 threads, 16k ctx A background model for coding harnesses: the slot that writes conversation titles and does small classification calls, moved off the GPU so it stops queueing behind the conversation it is describing.
glm-flash GLM-4.7-Flash 30B-A3B ~24GB VRAM, entirely on card (~17.5GB weights + ~3.8GB KV), 128k ctx Snappy agentic coding entirely in VRAM — no experts streaming from host RAM, so no CPU-offload latency. The iteration tier: cycle against it, escalate when it stalls.
minimax MiniMax-M2.7 229B-A10B ~19GB VRAM + ~65GB RAM, ~17 tok/s generation, 728 tok/s prefill, 64k ctx A second opinion, not a better coder. It is the biggest model in the fleet and reads like the strongest one, but against qwen-dense it is a PEER — reach for it when qwen-dense-long is confidently stuck and you want a genuinely independent take, not when you want more power.
ornith Ornith-1.0-35B-A3B (thinking) 22548MiB measured at 128k (~1.70GB of it q8 KV), 288 tok/s generation, 9387 tok/s prefill, entirely on card, 128k ctx coder-pro's benchmarks without coder-pro's host RAM, and the fastest model on the box: a 35B MoE with 3B active that fits on the card whole, measured at 288 tok/s against qwen's 270 and coder-pro's 20-35. Coexists with the resident set. Still to be A/B'd against coder-pro on real agent work.
qwen Qwen3.6-35B-A3B (thinking) ~20GB VRAM, ~270 tok/s, entirely on card, 256k ctx The daily driver, coding included — the reflex choice unless you specifically need the best code, the missing refusals, or a long agent loop. Entirely in VRAM at ~270 tok/s, the only model here fast enough that you stop noticing latency.
qwen-dense Qwen3.6-27B MTP (thinking) — off by default 25744MiB measured, ~80 tok/s (MTP spec decode), f16 KV, 80k ctx The best coder on the box by benchmark, and the most expensive to keep warm: it cannot share a 32GB card with transcription, RAG and memory. Off by default for capacity, not for quality.
qwen-dense-long Qwen3.6-27B (thinking), no MTP — the default 27B 23518MiB measured (~17GB weights + ~4.3GB KV), ~47 tok/s, q8 KV, 128k ctx The 27B you get unless you name another one, and the best coding quality available without evicting anything. qwen-dense's weights and answers with 128k instead of MTP speed.
qwen-dense-uc Qwen3.6-27B abliterated (thinking) ~19GB VRAM, ~47 tok/s, 128k ctx The uncensored dense 27B. Worth picking over qwen-dense-long only when refusals are the actual problem — it is a lossy edit of exactly those weights.
qwen-uc Qwen3.6-35B-A3B abliterated (thinking) ~21GB VRAM, ~260 tok/s, entirely on card, 256k ctx qwen with refusal-direction removal — same architecture, same speed, same 262k window. For when a refusal is blocking legitimate work, not as a general upgrade.
research gpt-oss-120b (reasoning: high) ~26GB VRAM + ~40GB RAM, ~30 tok/s, 64k ctx Knowledge-heavy queries and a second opinion: 117B total params carry broad world knowledge the 27-35B models do not have. No longer the reasoning escalation — the Qwen3.6 generation took that.
voice Qwen3-4B-Instruct-2507 (on the CPU) zero VRAM (runs on the CPU), 0.15s to first token, 12-14 tok/s, 8k ctx The realtime voice model, and it runs on the CPU — which is the whole reason a spoken turn never has to fight coder-pro for the card.
asr Qwen3-ASR-1.7B 3708MiB measured, always resident, ~23 tok/s, 4k ctx Speech-to-text behind Open WebUI's mic button and media-tools' transcribe_audio. Resident, so a transcription never evicts a warm coding session.
embed Qwen3-Embedding-0.6B 1698MiB measured, always resident, 2k ctx RAG embeddings. Resident by design, so a chat-model swap never evicts it and retrieval keeps working while the big models come and go.
rerank Qwen3-Reranker-0.6B 1698MiB measured, always resident (the old "~0.7GB" was weights only), 2k ctx Cross-encoder reranking for hybrid RAG — it scores the candidates that BM25 and embeddings surfaced. Open WebUI's hybrid search reranks through it.
coder-pro — Qwen3-Coder-Next 80B-A3B

The Claude Code / OpenCode model. Purpose-trained for long-horizon agent loops and recovery from failed steps, and the only model here whose window clears Claude Code's assumed ~200k — so long sessions auto-compact instead of overflowing.

Best at

  • 256k context: the practical reason to pick it, not a spec-sheet number.
  • Trained for tool use and multi-step recovery, which single-shot benchmarks understate.
  • Not a thinking model, so no reasoning latency before the first tool call — the Qwen3.6 thinking variants now survive the Anthropic bridge too, but they make you wait for the reasoning.
  • ttl 7200 keeps a warm prompt cache across long idles; anything needing the whole card calls /unload explicitly.

Costs and limits

  • 70.6 SWE-bench Verified and 36.2 Terminal-Bench 2.0 — below qwen-dense-long on both.
  • Experts stream from host RAM: ~20-35 tok/s, and it wants the RAM floor to be real.

Compared with

  • glm-flash — glm-flash is entirely in VRAM and snappier to iterate against; coder-pro has 2x the window and is what you escalate to when it stalls.
  • minimax — 4x the window and the model that actually drives agents. minimax is for one hard question, not a session.
  • qwen-dense-long — That one scores higher on coding (77.2/59.3 vs 70.6/36.2) but has half the window and no agent-loop training. Drive agents with coder-pro; ask hard coding questions of qwen-dense-long.
ds4-flash — DeepSeek-V4-Flash-0731 284B-A13B

The agentic escalation tier: a 284B model with a 128k window that costs almost no KV, tuned for tool use and long agent loops rather than for single answers. Reach for it when coder-pro has lost the thread on a long session — not for a quick question, because it generates at roughly 1/20th of qwen.

Outgrowing its window reroutes to coder-pro rather than hard-erroring.

Best at

  • Purpose-built for agentic work: the 0731 release is a large tool-use and coding upgrade over the V4 preview, and beats DeepSeek-V4-Pro (Preview) on benchmarks at a fraction of the active parameters.
  • 128k window for ~6.5GB of f16 KV — hybrid CSA+HCA attention with 1 KV head, so long context costs less here than minimax's 64k does.
  • A third lineage on the box: not a Qwen and not MiniMax, so it fails differently from both.
  • Only 13B active of 284B, so it generates at roughly MiniMax speed despite being 55B bigger.
  • Reasoning-effort levels are native to its chat template rather than bolted on by a prompt.

Costs and limits

  • UD-IQ2_M — 2-bit, the lowest-precision quant in the fleet. Forced by 277B of experts against 123GB of RAM, not chosen.
  • ~12-16 tok/s ESTIMATED and not yet measured on this box; the nCpuMoe and KV figures behind that are derived arithmetic.
  • ~80GB of host RAM while loaded — the largest RAM footprint here, and it will evict page cache other services were using.
  • Needs llama-cpp >= b10254. On anything older it answers, re-prefills every agentic turn, and looks slow rather than broken.
  • q8 KV corrupts it silently, so it cannot buy VRAM back the way every other model here can.
  • 1M native context is not reachable: the window is capped at 131072 by VRAM, not by the model.

Compared with

  • coder-pro — coder-pro drives sessions at ~10x the speed with a 256k window and is where agent loops belong. Come here when it is stuck, not to replace it.
  • minimax — The same job — one hard question off an independent lineage — with 2x the window and 55B more total parameters, against minimax's higher-fidelity 3-bit quant. minimax is MEASURED at ~17 tok/s; this one is not measured at all yet. Prefer minimax until this entry's numbers are real.
  • qwen-dense — Not a comparison worth making on speed. This is a far bigger model at 2-bit; qwen-dense answers ~20x faster and wins most questions that fit its window.
  • research — Both non-Qwen second opinions. research is a knowledge cross-check that fits on the card; this is an agentic coder that does not.
fara — Fara1.5-27B (computer-use agent)

Browser automation from screenshots. NOT a chat or coding model — drive it with fara-cli or Magentic-UI pointed at this endpoint. Picking it in a chat window can only fail.

Also answers to cua.

Best at

  • 72.3 Online-Mind2Web: form filling, shopping, bookings, information gathering.
  • 128k, because screenshots are token-hungry and computer-use traces run long.
  • MIT licensed.

Costs and limits

  • Vision-only perception and a click/type/scroll action space — no files, no shell, no code.
  • Its model card scopes it to web tasks; it is not a general agent.

Compared with

  • qwen — Unrelated jobs that both involve images: qwen understands a picture you attach, fara acts on a browser it is looking at.
fast-cpu — Qwen3-4B-Instruct-2507 with tools (on the CPU)

A background model for coding harnesses: the slot that writes conversation titles and does small classification calls, moved off the GPU so it stops queueing behind the conversation it is describing.

Best at

  • Zero VRAM, so it never competes with a chat model and never triggers an eviction.
  • Tool calls work, which is what separates it from voice and what a harness's background slot needs.
  • Reuses a GGUF the store already has, so enabling it downloads nothing.

Costs and limits

  • PREFILL is the constraint: ~600 tokens in 3.8s, but 4000 tokens takes 26s. A client that sends long prompts to its background slot will feel that.
  • A 4B instruct model. It is a helper, not a second opinion — never point a SUBAGENT slot at it, because subagents do real work.

Compared with

  • voice — The same weights and the same CPU. voice is the spoken-reply path (8k, no tools); this one has tools and a bigger window for a harness's background traffic.
glm-flash — GLM-4.7-Flash 30B-A3B

Snappy agentic coding entirely in VRAM — no experts streaming from host RAM, so no CPU-offload latency. The iteration tier: cycle against it, escalate when it stalls.

Also answers to glm.

Outgrowing its window reroutes to coder-pro rather than hard-erroring.

Best at

  • Fully resident in VRAM (~17.5GB of weights), so latency is GPU-bound rather than RAM-bandwidth-bound.
  • MLA attention makes context nearly free: 47 layers of 576-elem latent ≈ 28KB/tok, so 128k is only ~3.8GB.
  • coder-pro-class agentic behaviour at in-VRAM speed.
  • MIT licensed.

Costs and limits

  • A capacity tier below coder-pro — this is the model you escalate FROM.
  • ~24GB loaded leaves little room beside the resident set.
  • Native context is 202752; 128k is what fits here, not what the model can do.

Compared with

  • coder-pro — coder-pro has 2x the window and long-horizon agent training. glm-flash is quicker to iterate against; escalate when it gets stuck rather than starting there.
  • qwen — qwen is faster still with a 2x window, but glm-flash is the more agentic of the two — better at tool loops.
minimax — MiniMax-M2.7 229B-A10B

A second opinion, not a better coder. It is the biggest model in the fleet and reads like the strongest one, but against qwen-dense it is a PEER — reach for it when qwen-dense-long is confidently stuck and you want a genuinely independent take, not when you want more power.

Outgrowing its window reroutes to coder-pro rather than hard-erroring.

Best at

  • A different lineage from every Qwen here, so it fails differently. That independence is the actual reason to run it.
  • Wins SWE-bench Multilingual (76.5), SWE-Bench Pro (56.2), NL2Repo and GDPval-AA.
  • Good on multi-language repos and long-horizon spec->repo work.
  • 728 tok/s prefill: it reads a big prompt quickly even though it generates slowly.

Costs and limits

  • Not an upgrade over qwen-dense: it loses Terminal-Bench 2.0 (57.0 vs 59.3), and its wins are full-precision numbers this deployment does not run at.
  • IQ3_XXS — roughly 3 points of Aider pass rate below 4-bit on MoE models, and the quant is forced by architecture, not chosen.
  • ~17 tok/s, about 1/4.7 of qwen-dense. Below the interactive bar on purpose.
  • 64k, the smallest window of the coding-capable models: minimax-m2 is plain full attention (62 layers x 8 KV heads), so KV costs a measured 8432MiB at 64k and the native 196608 would want ~25GB of card.
  • ~65GB of host RAM while loaded.

Compared with

  • coder-pro — coder-pro has 4x the window and is the model that drives agent loops. This answers one hard question; it does not run a session.
  • qwen-dense — A peer, not an upgrade. It wins SWE-bench Multilingual and SWE-Bench Pro, loses Terminal-Bench 2.0, and runs 3-bit at ~1/4.7 the speed. Its value is independence, not capability.
  • qwen-dense-long — Start there. Come here only when that one is stuck — the same conclusion in a quarter the time beats a different one slowly.
  • research — Both second opinions off a non-Qwen lineage. This one codes better; research knows more.
ornith — Ornith-1.0-35B-A3B (thinking)

coder-pro's benchmarks without coder-pro's host RAM, and the fastest model on the box: a 35B MoE with 3B active that fits on the card whole, measured at 288 tok/s against qwen's 270 and coder-pro's 20-35. Coexists with the resident set. Still to be A/B'd against coder-pro on real agent work.

Outgrowing its window reroutes to coder-pro rather than hard-erroring.

Best at

  • 288 tok/s generation measured — the fastest here, ahead of qwen's 270, and ~10x coder-pro for the same class of coding work.
  • 9387 tok/s prefill over a 28k-token prompt (3.0s): nothing streams from host RAM, where coder-pro manages 2298 tok/s at 99k.
  • 75.6 SWE-bench Verified — above coder-pro's 70.6, from a model that needs no CPU offload to run.
  • 22548MiB at 128k, so it coexists with asr + embed + rerank with ~2.9GB to spare — verified live, unlike qwen-dense which locks them out.
  • Self-scaffolding RL post-training: it learns the task harness alongside the solution, which is the part single-shot benchmarks understate for agentic work.
  • MIT licensed, with no regional restrictions.

Costs and limits

  • Terminal-Bench 2.1 (64.2) is a different harness from the 2.0 numbers quoted for qwen-dense-long and coder-pro. It cannot be ranked against them as published, and no real-task A/B has been run here yet.
  • Every published Ornith score is the vendor's own and independent verification was still pending as of 2026-08. Reviewers who skipped the benchmarks for held-out tasks found it genuinely strong rather than benchmaxxed — reassuring, but not a reproduced number.
  • Over-gates. Reviewers report it stalling on straightforward, fully-disclosed requests by demanding access or prerequisites it has already been given. In an unattended loop that reads as a hang, and it is the opposite of what coder-pro's failed-step recovery training buys.
  • Reported to hit a ceiling on genuinely long jobs — one tester had it botch a ~100-iteration kernel implementation that larger models completed. Strong on short and mid-length agent chains is the consistent finding; long-horizon is where it stops.
  • Half of coder-pro's window (128k vs 256k), which is below Claude Code's assumed ~200k — long sessions reroute rather than compact.
  • 256k does not fit: the measured KV slope (13.6KB/tok at q8) puts it ~1.2GB from the ceiling with the residents loaded, under the 1.5GB margin.
  • A thinking model, so there is reasoning latency before the first tool call. coder-pro has none.
  • No MTP wired, so no spec-decode speedup: the head plus the f16 KV it needs would overrun the card next to the resident set.
  • Generation falls to ~245 tok/s by 28k of context — still the fastest here, but the headline number is a short-context one.

Compared with

  • coder-pro — The comparison this entry exists for, and the measurements favour this one hard: ~10x the generation speed, ~4x the prefill, better published SWE-bench (75.6 vs 70.6), and on the card instead of streaming 80B of experts from host RAM. It stays the challenger anyway, because NOBODY HAS RUN THIS COMPARISON — the published Ornith write-ups are all against Qwen3.6-35B-A3B, which is qwen here, not against a purpose-built agent driver. coder-pro keeps 2x the window (the part that clears Claude Code's assumed ~200k) and the failed-step recovery training, and the two weaknesses reviewers do report — over-gating and a ceiling on very long jobs — land exactly on that axis. Default stays coder-pro until a real-task A/B says otherwise.
  • glm-flash — The other agentic coder that fits on the card whole. glm-flash is smaller and thinking-free; this one is bigger, reasons first, benchmarks higher and measures faster.
  • qwen — Same architecture and size class; this one is measured slightly FASTER (288 vs 270) and coding-tuned. qwen keeps 2x the window and native vision.
  • qwen-dense-long — Still the higher SWE-bench score (77.2 vs 75.6), but dense and ~47 tok/s. This is the MoE bet, and measurement settled it: ~6x the throughput for 1.6 points.
qwen — Qwen3.6-35B-A3B (thinking)

The daily driver, coding included — the reflex choice unless you specifically need the best code, the missing refusals, or a long agent loop. Entirely in VRAM at ~270 tok/s, the only model here fast enough that you stop noticing latency.

Also answers to default, general.

Best at

  • ~270 tok/s fully in VRAM: ~5.7x qwen-dense-long, ~16x minimax.
  • Full native 262k window for only ~2.7GB of KV — a GDN hybrid, so just 10 of 40 layers carry KV.
  • 73.4 SWE-bench Verified — genuinely strong agentic coding, not a consolation prize for picking the fast model.
  • Native vision — attach images directly, no separate model.
  • Strong Japanese<->English translation.

Costs and limits

  • Out-coded by qwen-dense-long (77.2 vs 73.4 SWE-bench Verified) and out-driven by coder-pro on long agent loops.
  • 35B total but only ~3B active per token, so it has less depth on hard single-shot reasoning than the dense 27B.

Compared with

  • coder-pro — Same 256k window, ~10x the speed. coder-pro is purpose-trained to DRIVE agent loops and recover from failed steps; qwen is the generalist you ask directly.
  • qwen-dense-long — qwen is ~5.7x faster with a 2x window; qwen-dense-long is the better coder. Speed versus code quality, and for most questions speed wins.
  • qwen-uc — Same architecture, same speed, refusals removed at a small fidelity cost. Stay here unless a refusal is the actual problem.
qwen-dense — Qwen3.6-27B MTP (thinking) — off by default

The best coder on the box by benchmark, and the most expensive to keep warm: it cannot share a 32GB card with transcription, RAG and memory. Off by default for capacity, not for quality.

Outgrowing its window reroutes to qwen rather than hard-erroring.

Best at

  • 77.2 SWE-bench Verified and 59.3 Terminal-Bench 2.0 — the highest coding scores here, above coder-pro's 70.6/36.2.
  • MTP speculative decoding: ~80 tok/s, ~1.7x the same weights without it, and lossless by construction.
  • Official Qwen3.6 "precise coding" sampling (temp 0.6).
  • Native vision.
  • Claude Code can drive it: since 2026-08-02 the Anthropic bridge turns reasoning_content into real thinking blocks, so being a thinking model no longer rules it out.

Costs and limits

  • 25744MiB measured — overruns the card next to the 7104MiB resident set. Enabling it trades transcription, voice, RAG and memory for the speed.
  • 80k, the smallest window of the Qwen3.6 variants: f16 KV is what keeps MTP draft acceptance at ~90%, and f16 is what costs the context.
  • Half of coder-pro's window (80k vs 256k), which is below Claude Code's assumed ~200k — long sessions overflow rather than compacting.

Compared with

  • coder-pro — Better benchmarks (77.2 vs 70.6 SWE-bench Verified) and faster, but a third of the window. coder-pro's 256k is what clears Claude Code's assumed ~200k so sessions compact instead of overflowing.
  • qwen-dense-long — Identical weights, identical answers — speculative decoding verifies every token. This one is ~1.7x faster; that one has 128k instead of 80k and coexists with the resident set. You are trading speed against everything else on the box.

⚠ 25744MiB measured — cannot be loaded alongside the resident set (asr + embed + rerank, 7104MiB) on a 32GB card. Whichever loads second exits with "upstream command exited prematurely", so enabling this trades transcription, voice, RAG and memory for MTP's ~1.7x speed. qwen-dense-long is the same weights at 128k and coexists with everything.

qwen-dense-long — Qwen3.6-27B (thinking), no MTP — the default 27B

The 27B you get unless you name another one, and the best coding quality available without evicting anything. qwen-dense's weights and answers with 128k instead of MTP speed.

Also answers to dense.

Outgrowing its window reroutes to qwen rather than hard-erroring.

Best at

  • Identical output to qwen-dense — dropping speculative decoding costs throughput and nothing else.
  • 77.2 SWE-bench Verified / 59.3 Terminal-Bench 2.0, the top coding scores here.
  • 128k at q8 KV costs ~4.3GB — LESS than the ~5.4GB qwen-dense spends on 80k at f16.
  • 23518MiB leaves ~1.9GB beside the full resident set, so transcription, RAG and memory keep working.
  • Claude Code and OpenCode both drive it — the anthropic-bridge eval suite runs against this model precisely because it thinks.

Costs and limits

  • ~47 tok/s against qwen-dense's ~80. That is the entire cost of the trade.
  • Half of coder-pro's window, and no long-horizon agent training.
  • It thinks, so first-token latency is higher than a non-thinking coder's — the reasoning arrives intact through the Anthropic bridge, but you wait for it.

Compared with

  • coder-pro — Better raw coding scores here, but coder-pro has 2x the window and is trained for long agent loops. Use coder-pro to drive an agent, this to answer a hard coding question.
  • minimax — This one is ~4.7x faster and a peer on benchmarks. Go to minimax only when this is confidently stuck and you want an independent lineage.
  • qwen — qwen is ~5.7x faster with a 2x window; this is the better coder (77.2 vs 73.4 SWE-bench Verified). Ask qwen first, escalate here when the code has to be right.
  • qwen-dense — Same weights, same answers. That one is ~1.7x faster and locks out transcription, voice, RAG and memory; this one is 128k and shares the card.
qwen-dense-uc — Qwen3.6-27B abliterated (thinking)

The uncensored dense 27B. Worth picking over qwen-dense-long only when refusals are the actual problem — it is a lossy edit of exactly those weights.

Outgrowing its window reroutes to qwen-uc rather than hard-erroring.

Best at

  • Dense 27B reasoning depth without the refusal behaviour.
  • Q5_K_M — a higher quant than the 4-bit variants here.
  • 128k for ~4.3GB of KV, native vision included.

Costs and limits

  • Abliteration is lossy: qwen-dense-long is the higher-fidelity version of the same weights at the same 128k.
  • No MTP, so ~47 tok/s with no spec-decode speedup.

Compared with

  • qwen-dense-long — Same weights, not abliterated, same 128k window. Prefer that one unless refusals are why you are here — the context was never what made this one worth picking.
  • qwen-uc — This is the dense 27B (deeper, ~47 tok/s, 128k); that is the 35B-A3B MoE (~260 tok/s, 256k). Same abliteration lineage, opposite tradeoff.
qwen-uc — Qwen3.6-35B-A3B abliterated (thinking)

qwen with refusal-direction removal — same architecture, same speed, same 262k window. For when a refusal is blocking legitimate work, not as a general upgrade.

Also answers to uncensored.

Best at

  • The cleanest abliteration of the field: KL 0.0074 against the base model, smallest capability deltas in the Abliterlitics benchmarks.
  • ~260 tok/s and 256k ctx — it costs essentially nothing against qwen.
  • Fiction, red-teaming, and prompts that trip false refusals.

Costs and limits

  • Every abliteration is a lossy edit of the base weights. qwen is the higher-fidelity model when refusals are not the problem.
  • That KL number ranks it against other abliterations, not against the base model — it is the best of a lossy field, not lossless.

Compared with

  • qwen — Same model and speed with refusals removed, at a small fidelity cost. Default to qwen and come here when it refuses something it should not.
  • qwen-dense-uc — This is the fast MoE (~260 tok/s, 256k); that is the dense 27B — deeper on hard reasoning, ~5x slower, half the window.
research — gpt-oss-120b (reasoning: high)

Knowledge-heavy queries and a second opinion: 117B total params carry broad world knowledge the 27-35B models do not have. No longer the reasoning escalation — the Qwen3.6 generation took that.

Also answers to gpt-oss-120b.

Outgrowing its window reroutes to qwen rather than hard-erroring.

Best at

  • 117B total params: breadth of factual coverage is what it is actually for.
  • f16 KV per the official gpt-oss guide, which never quantizes it — a report there measured KV quant halving throughput.
  • Sliding-window attention halves KV cost, which is how 64k fits at all.

Costs and limits

  • AA reasoning index 24 against qwen-dense's 37 — passed by a model a quarter its size.
  • 64k window; 128k would OOM next to the resident services.
  • ~61GB MXFP4 with experts in RAM, for a model that is now a specialist rather than a default.

Compared with

  • minimax — Both are second-opinion models on a different lineage from the Qwens. minimax is the stronger coder; research has the broader factual base.
  • qwen-dense-long — That one out-reasons this at a fraction of the size. Come here for breadth of world knowledge, not for reasoning power.
voice — Qwen3-4B-Instruct-2507 (on the CPU)

The realtime voice model, and it runs on the CPU — which is the whole reason a spoken turn never has to fight coder-pro for the card.

Best at

  • Zero VRAM. Always available, whatever is loaded on the GPU.
  • 0.15s to first token and 12-14 tok/s measured: speech is spoken at ~3 words/second, so it stays three times ahead of the speaker.
  • Serves a GGUF the store already has (the Z-Image text encoder) rather than downloading a second copy.

Costs and limits

  • A 4B instruct model — it handles voice turns, not hard questions.
  • Completion only: no tools, no vision, no thinking.

Compared with

  • qwen — qwen is vastly more capable and needs the GPU. This exists so that answering "turn on the lights" out loud never evicts somebody's warm session.
asr — Qwen3-ASR-1.7B

Speech-to-text behind Open WebUI's mic button and media-tools' transcribe_audio. Resident, so a transcription never evicts a warm coding session.

Also answers to whisper-1.

Best at

  • Serves OpenAI's /v1/audio/transcriptions natively, and answers to the whisper-1 alias, so stock clients work unchanged.
  • Its audio encoder ships as an mmproj in the same repo, so it is just another llama-swap model rather than a second STT daemon.

Costs and limits

  • 3708MiB resident at 4k ctx — the single largest permanent claim on the card, and the reason the context has been cut twice.
  • 4k is ~3 minutes of audio. media-tools' transcribe_audio sends whole files unchunked, so longer clips fail there — that is genai-transcribe's job.
  • response_format must be json (no srt/vtt), and it cannot emit timestamps at all, so subtitles go through genai-transcribe.

Compared with

  • voice — asr is the ear, voice is the mouth: this transcribes on the GPU, voice answers on the CPU.
embed — Qwen3-Embedding-0.6B

RAG embeddings. Resident by design, so a chat-model swap never evicts it and retrieval keeps working while the big models come and go.

Best at

  • Multilingual, and small enough (1698MiB) to stay loaded permanently.
  • -c 2048 caps the KV and compute buffers; uncapped it allocates for the model's full context and balloons to ~5GB.

Costs and limits

  • Not a chat model — hidden from model pickers because selecting it can only fail.
  • Resident means never evicted, not VRAM reserved for free: its 1698MiB is permanently unavailable to the biggest chat model.

Compared with

  • rerank — The two halves of hybrid retrieval: embed finds candidates by meaning, rerank puts them in order.
rerank — Qwen3-Reranker-0.6B

Cross-encoder reranking for hybrid RAG — it scores the candidates that BM25 and embeddings surfaced. Open WebUI's hybrid search reranks through it.

Best at

  • Jina-compatible /v1/rerank: {model, query, documents[, top_n]} -> relevance-scored list.
  • A cross-encoder reads query and document together, so it catches relevance that independent embeddings miss.

Costs and limits

  • 1698MiB measured, not the ~0.7GB the weights suggest — and it is resident, so that is card the big models never get.
  • Not a chat model; hidden from pickers.

Compared with

  • embed — embed retrieves, rerank orders. Hybrid RAG runs both, which is why they share the resident group.

asr, embed and rerank serve non-chat endpoints (they sort last in the table above for that reason), so webui.utilityModels hides them from the dashboard proxy's /v1/models — otherwise they show up as selectable models in Open WebUI's chat picker (and LiteLLM's model list), where picking one can only fail. Hiding is listing-only: they stay fully routable, which is what lets Open WebUI's reranker keep POSTing {"model": "rerank"} to that same proxy. webui.utilityModels = [ ] shows everything again; adding "fara" also drops it from the picker, at the cost of breaking MagenticLite, which reaches it through this proxy.

That section is generated, and this is the only copy of it. Everything above between the model-fleet markers is rendered by docs.nix from each catalog entry's serve.guide in options.nix; the same fields are published by /api/status and render as the expandable panels on the portal's /models page. Advice about which model to pick used to live here and in the module, which is exactly the kind of duplication that comes apart quietly — the interesting entries are the ones that changed after a measurement (qwen-dense went off by default when the resident set grew; rerank turned out to cost 1698MiB rather than the ~0.7GB of its weights), and an edit in one place was invisible from the other.

So: edit serve.guide, never this section. Then

nix run .#update-readme          # re-render the block
nix build .#checks.x86_64-linux.docs   # or just `nix flake check`

The check fails the build if the committed text and the catalog disagree, in either direction, and prints the diff plus the command that fixes it. Two smaller guards ride along: versus keys must name real catalog entries (a module assertion, so a renamed model cannot leave dangling advice), and every shipped model must carry a summary (or it would silently vanish from the table). Context windows are not written into guide.footprint — renderers append them from serve.context, which module.nix already pins to the -c in the shipped llama-swap command.

Rule of thumb: for Claude Code use coder-pro (256k window, agent-trained, non-thinking — the Anthropic bridge drops reasoning_content, so thinking models degrade there). For OpenCode, coder-pro is the default and qwen-dense-long is the challenger: it out-benches coder-pro (77.2 vs 70.6 SWE-V, 59.3 vs 36.2 Terminal-Bench 2.0) and MTP makes it ~2-4× faster, but it thinks (higher first-token latency) and coder-pro is the more battle-hardened pure agent — A/B on real tasks. glm-flash when iteration speed matters more than depth. qwen is the fast default for everything else (73.4 SWE-V at ~6× dense speed); research for knowledge-heavy queries. All chat models have web-search + media tools attached and use native function calling. (The former coder Qwen3-Coder-30B slot was retired: obsoleted by qwen-dense/glm-flash, and its two 64k slots silently truncated long agent prompts.)

fara is the odd one out: a computer-use agent, not a chat model. The intended frontend is MagenticLite on :8895 (the magentic-ui service): it runs Fara as the browser-use model and qwen as the orchestrator (tunable via the magenticUi.* options), with the agent's browser inside a Quicksand QEMU micro-VM. The two roles swap on llama-swap at agent-round boundaries — a few seconds each from page cache; set magenticUi.agentMode = "websurfer_only" to eliminate swapping entirely. Ad-hoc alternative: uvx fara-cli --base_url http://<host>:8080/v1 --api_key none --model fara "book a table for two". fara has no LiteLLM context fallback on purpose — a CUA session degrading to a chat model would emit garbage browser actions.

Beyond /v1/chat/completions, /v1/embeddings and /v1/rerank, llama-server also exposes legacy /v1/completions and /tokenize — all routed by model name through llama-swap (:8080) and the dashboard filter proxy (:8897).

Models land in /var/lib/genai-models/llm (part of the unified model store — see Storage below). On activation/boot the genai-models-prefetch service downloads any missing model blobs with aria2 (16 parallel connections — llama-server's own downloader is single-stream and much slower), so a fresh box warms itself automatically. genai-prefetch <repo[:tag]> runs the same tool manually. Small mmproj sidecars are still fetched by llama-server on first use.

Quant authors (esp. unsloth) sometimes re-upload fixed GGUFs after llama.cpp correctness/tool-parsing fixes — a stale file silently degrades coding. To force-refresh one model: sudo rm -rf /var/lib/genai-models/llm/models--<org>--<repo> then re-run genai-prefetch <repo:tag> (or wait for the boot prefetch).

Qwen thinking models run with preserve_thinking=true so agentic loops keep prior reasoning in context. MoE models bigger than VRAM use --n-cpu-moe N (lower N until CUDA OOM, then back off). Sampling flags follow vendor recommendations — see comments in module.nix.

All Qwen 3.5/3.6 chat models run with a fixed chat template (vendored in templates/, from froggeric/Qwen-Fixed-Chat-Templates): the stock template makes the model occasionally emit an empty tool call, which agent clients (Claude Code, OpenCode) read as "task complete" and silently stop mid-session. Requires llama.cpp ≥ b9925 (llama-server --version; a NixOS assertion enforces this at eval time).

The resident set is a standing VRAM tax, and it does not fit with everything

embed, rerank and asr are in the resident group, which means never evicted — so their cost is subtracted from every chat model, permanently. Measured 2026-08-03 on the 32607MiB card:

measured
embed (-c 2048) 1698MiB
rerank (-c 2048) 1698MiB
asr (4k ctx) 3708MiB
resident total 7104MiB
leaves for a chat model 25503MiB

Against that budget, qwen-dense needs 25784MiB (80k f16 KV, which f16 is there to keep MTP draft acceptance high) and qwen-dense-long needs 23526MiB (measured alone on an empty card). So:

  • qwen-dense and asr cannot both be loaded. Whichever arrives second dies with upstream command exited prematurely — llama.cpp fast-failing a CUDA allocation, not a llama-swap fault. resident means "never evicted", not "VRAM reserved", so this is order-dependent: transcribe first and the 27B is locked out; load the 27B first and transcription is. qwen-dense therefore ships disabled. Turn it on with:
services.genai-server.llmModels.qwen-dense.serve.enable = true;

accepting that transcription, voice, RAG and memory lock it out (and it locks them out) for as long as either side is resident. asr's share of that budget is transcription.maxAudioMinutes — ~152MiB of VRAM per minute of audio, defaulting to 3.

  • qwen-dense-long fits with ~1977MiB spare, inside the 1.5GB margin this repo keeps. It was 537MiB before asr dropped from 16k to 4k.
  • Everything smaller (qwen, glm-flash, coder-pro, minimax) coexists fine, because their weights or their CPU offload leave more room.

Dropping embed and rerank from -c 8192 to -c 2048 recovered 1356MiB of this and is why qwen-dense-long fits at all. There is no similar free win left: qwen-dense's weights alone are ~20GB, so no context setting makes it coexist without also giving up MTP. That is a deliberate open tradeoff, not an oversight — the speech-under-load eval exists to keep it visible.

Virtual model IDs (selectors)

A selector is a name clients ask for that resolves to a real model per request (needs llamaSwap.useNewerBuild until nixpkgs ships >= v242):

services.genai-server.llamaSwap.selectors.coder = {
  strategy = "warm";                       # warm | pin | spillover
  targets = [ "coder-pro" "qwen-dense" ];
  description = "Whichever coding model is already loaded";
};

Two uses. A/B without touching clients: reorder targets to promote a challenger and Claude Code, OpenCode and Open WebUI all follow. Not evicting a warm model: strategy = "warm" picks a target already in VRAM, which on a single shared card is usually worth more than the difference between two good models.

Selectors show up in /v1/models (tagged type: selector) and are added to the portal's allowlist automatically — they are not catalog entries, so otherwise the availability filter would hide them.

Adding an LLM

A catalog entry, not a code change. Set serve.enable = true and the llama-swap block is generated from serve.*:

services.genai-server.llmModels.my-model = {
  repo = "unsloth/Some-Model-GGUF";
  tag  = "Q4_K_M";
  serve = {
    enable   = true;
    context  = 65536;          # -c; the VRAM lever (KV scales linearly)
    preset   = "qwen-thinking";# vendor sampling preset
    kvQuant  = "q8_0";         # halves KV VRAM, near-lossless
    nCpuMoe  = 24;             # MoE experts kept in host RAM
    aliases  = [ "mine" ];
    ttl      = 900;
    capabilities = [ "completion" "tools" "vision" ];
  };
};

It joins serve.group (main swaps, resident stays loaded), gets prefetched with everything else, and shows up on /models with its capabilities.

A catalog entry is not the whole job for a chat model. Open WebUI keeps tool attachments and vision flags in its own config database, seeded from two lists that serve.capabilities does not drive:

declared also add it to or it will
"tools" webui.toolModels get no tool servers — no image, video, or web search
"vision" webui.visionModels fail on attached images instead of seeing them

These are separate lists on purpose — fara is tool-capable and excluded because the media tools collide with its browser-action space — but the cost is that nothing warns you. A model missing from toolModels does not error; it answers "I don't have image-to-image or video tools," which is true of what it was handed and false about the box. Both lists take effect on the next rebuild, when the seeder runs.

Other fields: type (chat / embedding / rerank / transcription — picks the server mode), mmprojUrl (only for repos whose HF manifest hides their projector), shardFile (split GGUFs — see Big models), minVramGB / minRamGB (hardware floors below which the entry is omitted rather than left to fail), draft (MTP or a draft repo for speculative decoding), stripPenalties (for models whose tool-calling breaks under client-injected penalty samplers), extraFlags (appended last, so it wins) and rawEntry (llama-swap keys the submodule doesn't model).

The shipped models keep hand-written llama-swap entries in module.nix because their flags encode measured VRAM math that a submodule can't carry — they still fill in serve.* as the metadata /models, /api/status and Ollama-dialect clients publish. Those two could drift, so a NixOS assertion checks every declared serve.context against the -c in the command that actually runs. If you set serve.enable on a name that is hand-written, the hand-written entry wins and you get a warning saying so.

Big models: sharded GGUFs

Past roughly 50GB a quant is published as a split GGUF, and such a repo has no HF manifest at all — the endpoint answers 400 The specified repository contains sharded GGUF. That is the same endpoint genai-prefetch and llama.cpp's own -hf resolve through, so repo:tag cannot name these models. Declare the first shard instead:

services.genai-server.llmModels.big = {
  repo = "unsloth/MiniMax-M2.7-GGUF";
  tag  = "UD-IQ3_XXS";                 # the quant DIRECTORY, not a tag
  serve.shardFile =
    "UD-IQ3_XXS/MiniMax-M2.7-UD-IQ3_XXS-00001-of-00003.gguf";
};

genai-prefetch enumerates the remaining parts from the tree API into /var/lib/genai-models/llm/sharded/<repo>/<quant>/, and the model is served with -m <first shard> — llama.cpp finds the siblings by name. An assertion checks that shardFile lives in the directory tag names, since those two are one path spelled in two places. The portal reports such a model ready only when every part is present; a partial set reads absent, because llama.cpp opens the first shard and then dies demanding the rest.

Downloads resolve the CDN redirect before handing the URL to aria2, and that detail is load-bearing: HF's Xet-backed CDN signs a redirect with a byte-range condition whenever the request that triggered it carried a Range header, so aria2's 16 parallel ranges against one resolved URL get 403s and eventually stall outright. Resolving with a plain request yields a URL valid for every range. Don't "optimise" that step away.

Sizing is not file size. Two numbers decide whether a big MoE is usable, and the second one is the one people miss:

  • Total size → whether it loads. serve.minRamGB is the guard, and it is separate from minVramGB on purpose: --n-cpu-moe makes a model fit on a small card by construction, so it clears the VRAM floor and then wants 60-90GB of host RAM nothing asked about. Below that floor the failure is swap thrash, not a clean OOM — it reads as "the model is slow".
  • KV cache architecture → how much card is left for weights. A hybrid or MLA model (coder-pro, glm-flash, the Qwen3.6s) carries almost no KV, so context is nearly free. A full-attention model does not: minimax spends 8.3GB on 64k, which is what forces it to the 3-bit quant. Read the GGUF metadata (block_count, attention.head_count_kv, key_length) before assuming a quant fits.
  • Active params → speed. Only active experts are read per token, so throughput is roughly host RAM bandwidth ÷ active bytes. On a single-CCD Zen 5 that bandwidth is ~60-65GB/s regardless of the DDR5 rating: 3B active ≈ 20-35 tok/s, 5B ≈ 30, 10B ≈ 17 measured. Treat that arithmetic as a lower bound — it ignores the layers still resident on the card, which is why minimax measured 17 tok/s against a predicted 10-14.

Derive nCpuMoe from measurement, not arithmetic. The component costs are what matter, and on this box they are: KV as computed from the GGUF header, compute buffers ~1.1GB at -ub 2048, non-expert weights ~3.2GB, and ~1.2GB per layer of experts. A first guess that ignored a warm ComfyUI's ~3.9GB residual put minimax at 51 and it OOM'd on the KV allocation; the real number is 57. Tune against a warm box, not a freshly rebooted one, or you ship a model that loads in the morning and OOMs by evening.

Knowledge collections (RAG)

http://<host>:8897/rag manages them; :8900 is the API. Retrieval is hybrid — BM25 catches exact terms and identifiers, embeddings catch paraphrase — fused with reciprocal-rank fusion and reranked by the resident cross-encoder.

curl :8900/ingest_url  -d '{"collection":"docs","url":"https://example/page"}'
curl :8900/ingest_text -d '{"collection":"docs","text":"...","uri":"note:1"}'
curl :8900/search      -d '{"query":"how do I free the GPU?","limit":5}'

Results carry score and reranked. The score is the reranker's relevance when reranking ran and the fusion score otherwise — it always explains the order — and reranked: false means the cross-encoder was unavailable and you are seeing fusion order, rather than silently pretending.

Re-ingesting the same uri replaces that document instead of duplicating it, so refreshing a source is just running the ingest again. Filesystem ingest is off unless rag.ingestRoots lists directories; paths are resolved with realpath, so .. and symlinks cannot escape them.

It publishes openapi.json, so Open WebUI can register it as a knowledge tool and the MCP gateway exposes every operation automatically — including drop_collection, which is marked destructive and needs confirm: true.

Running code

:8903 executes code in a disposable QEMU microVM — a fresh VM per run, destroyed afterwards. Chat models and MCP clients both get it as a tool, so they can compute an answer instead of guessing one.

curl :8903/run_code -d '{"code":"print(sum(range(101)))"}'
curl :8903/run_code -d '{"code":"...","language":"bash"}'

Pass data in with files ({"in.csv": "..."}) and set return_files to get back what the code wrote. ~45ms per run on a warm-pool hit.

What it cannot do, by construction: see this machine's filesystem, keep state between runs, or reach the network. Network needs two gates — the host permitting it (codeSandbox.allowNetwork, default false) and the request asking. Think before opening that: it is model-written code, from text a user or web page supplied, gaining outbound access from inside your network. codeSandbox.allowedDomains narrows it.

Requires /dev/kvm. Without it the service answers 503 rather than quietly running code unsandboxed.

Memory

:8902 stores short durable facts a model can recall in a later conversation — distinct from the document collections above.

curl :8902/remember -d '{"text":"Prefers metric units and 24-hour time"}'
curl :8902/recall   -d '{"query":"what units should I use?"}'

Better still, hand it a conversation and let it decide what is worth keeping:

curl :8902/observe -d '{"text":"User: I always use NixOS, never Ubuntu..."}'

Corrections replace what they correct. Each new fact is shown to a model alongside the memories nearest to it, which picks ADD / UPDATE / DELETE / NOOP — so "switched to imperial" deletes "prefers metric" instead of sitting next to it. A similarity threshold cannot do this: those two are only 0.73 apart, and the stale one would otherwise be returned first.

It never takes the GPU from you. That reconciliation runs only against a model llama-swap already has resident; otherwise the work queues and a cheap 60s poll picks it up when the card is next warm. Facts are stored before they are reconciled, so a busy GPU delays a correction but never loses a write. forget takes an id or exact text — never a fuzzy match — and is gated behind confirm: true over MCP.

One store, not two. Open WebUI 0.11 ships its own memory, and running both would give the box two divergent sets — chat writing to one while Claude Code reads the other. So rag-style, exactly one is on: this service is registered with Open WebUI as a tool server and ENABLE_MEMORIES=False turns its built-in memory off. Chat models still get remember/recall; Claude Code over MCP sees the same memories. Browse and edit them at :8897/memory. Reverse the choice with services.genai-server.memory.useAsWebuiMemory = false.

Namespacing is only as strong as the identity you can trust. Memories are scoped by a header (memory.identityHeaders, default including Open WebUI's X-OpenWebUI-User-Id), never by a field the model fills in — a model that can name its own namespace can read every other one. Open WebUI does not currently forward those headers over MCP connections, so anything arriving that way shares memory.defaultNamespace. That is fine for a single-user box; it is not isolation between users who distrust each other.

MCP clients (Claude Code, Claude Desktop, Zed)

The same tools chat models use, over MCP — no second implementation. Point a client at http://<host>:8899/mcp (or http://<host>:8897/svc/mcp/mcp to keep one origin for the whole box):

claude mcp add --transport http genai http://<host>:8899/mcp
curl http://<host>:8899/tools     # what's bridged right now

Tools come from the tool servers' OpenAPI documents — web_search, generate_image, generate_image_with_face, swap_face, swap_face_full, swap_face_fast, edit_image, reimagine_image, inpaint_image, smart_edit, create_mask, transcribe_audio, text_to_speech — so adding a tool to a tool server makes it appear here with no MCP-side change.

Destructive operations get MCP's destructiveHint annotation and a required confirm: true argument, so a model can't trigger one by accident. Nothing shipped is destructive; mcp.upstreams.<name>.destructive marks the ones that are when a host adds its own tool server:

services.genai-server.mcp.upstreams.rag = {
  url = "http://127.0.0.1:8900";
  prefix = "rag_";                    # avoids operationId clashes
  destructive = [ "drop_collection" ];
};

Browsers must be same-origin or listed in mcp.allowedOrigins (the spec's DNS-rebinding guard); non-browser clients send no Origin and are unaffected.

Ollama-only clients

Point anything that speaks Ollama at http://<host>:11434 — Home Assistant's conversation agent, the JetBrains AI plugin, mobile apps. It is a translation layer on top of the same llama-swap fleet, not a second model server, so ollama is never installed and the model store is not duplicated.

curl http://<host>:11434/api/tags                       # the ready+enabled catalog
curl http://<host>:11434/api/show -d '{"model":"qwen"}' # capabilities, context
curl http://<host>:11434/api/chat \
  -d '{"model":"qwen","messages":[{"role":"user","content":"hi"}]}'

/api/show reports each model's capabilities (completion, tools, vision, thinking, embedding, rerank) from its serve block, which is what Ollama clients auto-route on — a vision model that didn't advertise vision would simply never be sent an image. Aliases work here too (dense, whisper-1, uncensored).

Models stay declarative. /api/pull starts a download only for a name already in llmModels; anything else returns an error naming the option to add it to. /api/create, /api/copy, /api/push and /api/delete answer 501 — the git-tracked catalog is the gallery. Disable the whole dialect with services.genai-server.ollama.enable = false.

Home Assistant voice (Wyoming)

Home Assistant speaks Wyoming, not OpenAI, so it gets a protocol bridge:

services.genai-server.wyoming = {
  enable = true;
  openFirewall = true;          # HA is normally on another machine
};

Then in Home Assistant: Settings → Devices & Services → Add Integration → Wyoming Protocol, host <host>, port 10300. Speech-to-text and text-to-speech both appear, ready to drop into an Assist pipeline.

This adds no models. STT forwards to the asr model llama-swap already serves; TTS to the same tts-server facade behind /v1/audio/speech. nixpkgs' wyoming-faster-whisper and wyoming-piper would each mean a second copy of a model this box already has, competing for the same card.

Voices default to Piper, which runs on the CPU — the right default for a voice assistant on a shared GPU, since answering "what's the weather" should not evict a 46GB coding model. Kokoro voices sound better and do use the GPU; add them deliberately:

services.genai-server.wyoming.ttsVoices = [ "en_US-lessac-medium" "af_heart" ];

Wyoming has no authentication. Anything that can reach the port can transcribe and synthesize, so openFirewall is separate from openFirewallGlobally and this should never be port-forwarded.

STT goes through the portal's /v1 proxy rather than llama-swap directly, because that is where the ASR preamble is stripped — see below.

Transcripts and the ASR preamble

Qwen3-ASR emits a structured header and llama-server passes it into the text field, so a raw transcription reads:

language English<asr_text>The quick brown fox jumps over the lazy dog.

The portal strips it on /v1/audio/transcriptions, which is why anything doing speech-to-text should go through :8897 rather than llama-swap directly — Open WebUI's microphone, the media-tools STT tool and Wyoming all do. Tune or disable it with portal.asrTextCleanup (a regex; "" passes the model's raw output through):

services.genai-server.portal.asrTextCleanup = "^\\s*language\\s+[A-Za-z]+\\s*<asr_text>\\s*";

Realtime voice (:8901)

services.genai-server.voice.enable = true;

Speech in, speech out, over one WebSocket — then open :8897/voice and talk. Distinct from /v1/audio/*, which is request/response: the value here is in the seams — noticing you stopped talking, answering before the reply is finished, and shutting up when you interrupt.

It costs no VRAM. The chat model (voice, a 4B) is served with -ngl 0, so it lives in system RAM and is permanently warm without taking anything from the card. Measured: 0.15s to first token at 12-14 tok/s — about three times faster than speech is spoken, which is the only throughput voice actually needs. It also reuses a GGUF already in the model store (the Z-Image text encoder) rather than downloading the same weights twice.

Measured end to end on this box, warm:

leg
transcription (asr, GPU, resident) ~30ms
chat + first-sentence synthesis ~960ms
detection → first audio ~990ms
plus the VAD hangover you actually wait through +600ms default

So roughly 1.6s from falling silent to hearing a reply. The page shows both numbers, because the smaller one flatters the experience.

Two behaviours worth knowing:

  • Synthesis starts at the first sentence, not the last. Waiting for a complete answer would blow the latency budget on its own.
  • Barge-in. Talking over the reply cancels it mid-sentence. An assistant that keeps talking while being interrupted is worse than a slow one.

voice.vadHangoverMs (default 600) is how long the server waits before deciding you finished. Both directions of that trade are bad — too short cuts you off mid-thought, too long adds dead air — and it errs long because real speech pauses more than synthesized speech does.

Microphone access needs a secure context. Browsers refuse getUserMedia over plain HTTP from another machine, so /voice works at http://localhost:8897/voice on the box itself, or behind TLS. The page says so rather than failing mysteriously.

Coding agents (Claude Code / OpenCode)

Claude Code — point it at the LiteLLM bridge:

export ANTHROPIC_BASE_URL=http://logistikon:4000
export ANTHROPIC_AUTH_TOKEN=dummy
export ANTHROPIC_MODEL=coder-pro             # primary; glm-flash for snappier loops.
                                             # NEVER a thinking model (qwen/qwen-dense):
                                             # the bridge drops reasoning_content
export ANTHROPIC_DEFAULT_HAIKU_MODEL=coder-pro  # haiku-tier background calls → same warm
                                             # model (otherwise they hit llama-swap as
                                             # claude-haiku-* -> 404 / pointless swaps)
export ANTHROPIC_SMALL_FAST_MODEL=coder-pro  # older name of the same knob, kept for
                                             # back-compat with older Claude Code
export CLAUDE_CODE_SUBAGENT_MODEL=coder-pro  # subagents stay on the warm model too
export CLAUDE_CODE_ATTRIBUTION_HEADER=0      # the attribution block mutates the prompt
                                             # prefix and silently defeats llama-server's
                                             # prefix cache (full re-prefill every turn)
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1  # cut parallel background calls:
                                             # they serialize on the single slot and
                                             # evict the prompt cache (minutes/turn)
export CLAUDE_CODE_MAX_OUTPUT_TOKENS=16384   # smaller output reservation = more
                                             # prompt headroom before overflow
claude

Caveat: Claude Code estimates context with Anthropic's tokenizer, which undercounts Qwen tokens on code — a session it believes is at ~185k can really be at ~230k+, past even coder-pro's 256k. When that happens the LiteLLM fallbacks entry converts the overflow into a properly-typed context error so Claude Code compacts and continues (instead of a retry loop that pins the GPU — see the litellm comments in module.nix).

Claude Code assumes a ~200k context window and will not self-limit. coder-pro (256k) clears that, so sessions auto-compact normally. On the smaller-window models the LiteLLM context_window_fallbacks ladder is the safety net: an overflowing session silently continues on a larger-window model (qwen-dense → qwen, qwen-dense-uc → qwen-uc, glm-flash → coder-pro, research → qwen).

OpenCode — declare the true per-model windows so it compacts before overflowing instead of dying; put this in ~/.config/opencode/opencode.json (routes through LiteLLM to keep the fallback net):

Per-model options matter: opencode sends NO temperature for custom models (the server-side --temp would apply) but force-sends top_p: 1.0 for any model id containing "qwen" — the explicit options below pin the vendor sampling and neutralize that. autoupdate off (nix manages the binary), compaction.prune trims old tool outputs before compacting, and disabling the title agent removes a concurrent request that evicts the single slot's prefix cache at session start. Do NOT set small_model — titles already run on the session's model; pinning one would create llama-swap churn.

{
  "$schema": "https://opencode.ai/config.json",
  "autoupdate": false,
  "compaction": { "prune": true },
  "agent": { "title": { "disable": true } },
  "provider": {
    "logistikon": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "logistikon (local)",
      "options": { "baseURL": "http://logistikon:4000/v1" },
      "models": {
        "coder-pro":     { "name": "Qwen3-Coder-Next 80B (256k)",     "limit": { "context": 262144, "output": 32768 }, "options": { "temperature": 1.0, "top_p": 0.95 } },
        "qwen-dense":    { "name": "Qwen3.6-27B MTP (80k, top coder)", "limit": { "context": 81920,  "output": 32768 }, "options": { "temperature": 0.6, "top_p": 0.95 } },
        "glm-flash":     { "name": "GLM-4.7-Flash (128k)",            "limit": { "context": 131072, "output": 32768 }, "options": { "temperature": 0.7, "top_p": 1.0 } },
        "qwen":          { "name": "Qwen3.6-35B-A3B (256k)",          "limit": { "context": 262144, "output": 32768 }, "options": { "temperature": 0.6, "top_p": 0.95 } },
        "qwen-uc":       { "name": "Qwen3.6-35B huihui (256k)",       "limit": { "context": 262144, "output": 32768 }, "options": { "temperature": 0.6, "top_p": 0.95 } },
        "qwen-dense-uc": { "name": "Qwen3.6-27B huihui (128k)",       "limit": { "context": 131072, "output": 32768 }, "options": { "temperature": 0.6, "top_p": 0.95 } },
        "research":      { "name": "gpt-oss-120b (64k)",              "limit": { "context": 65536,  "output": 16384 }, "options": { "temperature": 1.0, "top_p": 1.0 } }
      }
    }
  }
}

Speech (TTS)

One OpenAI endpoint (:8892/v1/audio/speech), three engines picked by voice:

  • en_US-lessac-medium (default) — Piper: CPU, faster-than-realtime; for UI read-aloud and anything latency-sensitive.
  • af_heart, bm_george, jf_alpha, ... — Kokoro-82M: narration quality; use for audiobook-style output.
  • clone:<name>Chatterbox zero-shot voice cloning (MIT; beat ElevenLabs 63.75% in blind tests). Register a reference first: curl -T sample.wav http://<host>:8896/voices/<name> (10-30s of clean speech), then request voice clone:<name>.

Video subtitles (genai-transcribe)

genai-transcribe talk.mkv                 # -> talk.srt
genai-transcribe -l ja ~/Videos           # recursive; language forced
genai-transcribe -V noisy-recording.mp4   # add Silero VAD (rarely needed)

Directories are searched recursively for video files. A video is skipped if a .srt with the same basename already exists, and the .srt is only written after a fully successful run — an interrupted run leaves a .whisper.partial and resumes from the last decoded segment.

Why this uses whisper and not the asr model. Subtitles need timestamps, and asr has none: llama-server returns {"text": ...} and rejects verbose_json/srt/vtt with "Only 'json' response_format is supported". Everything that does not need timing — the mic button, transcribe_audio, Wyoming — still uses asr, which is better at it and always resident. This is a second engine for a capability the first cannot express, not a second way to do the same job.

It runs on the CPU, deliberately: a subtitle job lasts hours and the GPU is shared. Nothing here loads a model onto the card, and there is no daemon — it is a command, not a service.

Two behaviours worth knowing, because they are the difference between usable subtitles and subtly wrong ones:

  • Audio is split at silences longer than 2.5s and each region transcribed separately. Whisper will otherwise emit a single cue whose text spans a long pause, so a line appears tens of seconds before it is spoken. It cannot do that if it never sees across the pause.
  • -mc 0 disables context carry-over, which is what stops whisper falling into repetition loops over music and silence.

The model (large-v3-turbo, 1.5GB) is a mediaModels entry in its own whisper set, so a host that never transcribes a video never downloads it; the command fetches the set on first use via genai-fetch-media.

Audiobooks (genai-narrate)

genai-narrate book.epub              # -> book.m4b, chaptered
genai-narrate -v bm_george novel.md  # British male; .txt/.md converted first
genai-narrate -L                     # list voices

Kokoro-82M via audiblez: e-book in, chaptered .m4b out, next to the input. Chapter WAVs accumulate in a .narration.tmp/ work dir, kept on failure so a long run is not lost, removed on success.

On the CPU by default (--cuda is deliberately not passed). A novel runs an hour or two at ~60-100 characters/sec, and holding the shared card that long is not a fair trade. The environment self-installs on first use — torch and Kokoro, several GB — into /var/lib/genai-narrate, with the model weights going to the shared HF_HOME so Kokoro is fetched once for the box.

There is no "speak this aloud" mode: a shared server has no audio sink, and /v1/audio/speech already synthesizes Kokoro on demand — ask for a Kokoro voice (af_heart, bm_george, ...) and tts-server routes it. This command is for turning a book into a file.

Repairing older audiobooks (-r)

genai-narrate -r ~/Audiobooks        # directory, recursive
genai-narrate -r book.m4b            # or one file

Audiblez writes raw pcm_s16le into the MP4 container, which the ipod muxer refuses — so a .m4b produced before this was fixed is ~385 kbps of PCM that many audiobook players reject, still tagged with the placeholder title .. -r re-encodes to 64k AAC and restores the title from the filename, taking a minute per book instead of re-narrating for hours.

It only touches files that are actually broken (already-correct AAC with a real title is left alone, so it is safe to re-run), writes to a temporary file and only replaces the original on success, and refuses to replace anything if the chapter count came out lower than it went in — losing chapters turns an audiobook back into one long track. It needs only ffmpeg, so it never triggers the multi-GB environment install.

Non-epub input is converted with pandoc first, with a placeholder title: audiblez always narrates "{title} {author}." before chapter 1, and an empty title is spoken aloud as the literal word "None". The real title is restored in the .m4b tags afterwards.

Media generation (ComfyUI)

ComfyUI runs as a service on :8188 (self-installs to /var/lib/comfyui as the genai user on first start; model weights come from the shared store via extra-model-paths). comfyui.modelSets = [ "comfy" ] declares the recommended node-graph set (~90GB), which the unit fetches before it serves: FLUX.2 dev fp8 (image, best prompt adherence), Wan 2.2 14B t2v+i2v (video), LTX 2.3 distilled (fast video with synchronized audio). Use the built-in workflow templates. Z-Image stays on :8893 for fast API image gen. Music generation also runs in ComfyUI (native ACE-Step nodes) — no separate service.

The ComfyUI version is pinned by comfyui.rev (default v0.30.1), and the unit reconciles the checkout against it on every start — fetch, check the ref out, and reinstall requirements.txt only if the commit actually moved. Before this option the version was pinned by accident: the launcher cloned the default branch on first start and never touched it again, so what a box ran was whatever upstream HEAD was on the day it was first built, and moving it meant a git pull under /var/lib that no rebuild reproduces. Safe because everything the service owns (models/, custom_nodes/, user/, output/, .venv/) is in ComfyUI's .gitignore — only tracked source moves. Node packs pin separately, so raising this can land one on a core it does not support; the symptom is a node that stops loading, logged in journalctl -u comfyui.

Video with native audio: MiniMax H3 (comfyui.modelSets = [ "h3" ], ~63GB, and not to be confused with the minimax LLM — same vendor, unrelated model). A diffusion transformer that emits video and stereo audio — dialogue, effects and music — in one forward pass, driven by the stock video_minimax_h3_{t2v,i2v,r2v} templates. Needs comfyui.revv0.30.0. 42.5GB of weights run on a 32GB card because the stages are sequential and --disable-smart-memory frees between them: text encoder (15.7GB, the NVFP4 build — this box is Blackwell), then the DiT (21GB), then the two VAEs. Native canvas is 768px on the short edge (capped 768×1344) at 24fps for ~15s; 2K is an upscale on top of that, not the sampler's working resolution. A clip takes minutes and holds the whole card — the launcher asks llama-swap to unload first, so expect a warm chat model to be evicted.

Two levels of idle reclaim, because they reach different memory. comfyui.idleReleaseMinutes (default 10) calls /free once the queue has been empty that long, which returns what torch allocated. What it cannot return is what torch never allocated: measured here, ComfyUI idle for two days after its last render reported 528MB of torch_vram_total while the process held 3.9GB of VRAM and 13GB of RAM — the CUDA primary context and the cuBLAS/cuDNN/torch kernel images loaded into it. No CUDA call frees those; only process exit does. (--disable-cuda-malloc is already on and is why this is 3.9GB rather than the 6-7GB it once was.)

So comfyui.idleStopMinutes (default 0, off) stops the service outright after that long with no client connected, and a systemd socket starts it again on the next request — measured 7 seconds. Turning it on moves ComfyUI to loopback behind that socket; the portal card links to the public port as before but health-probes the server directly, because probing an activation socket is indistinguishable from using it and would keep the service alive forever. A stopped ComfyUI shows as idle, not as a fault. Idleness is connections, not queue depth — read the option before enabling it, since a render whose client has disconnected can be cut short.

Custom nodes are declarative. comfyui.customNodes is a table of node packs; on every start the unit clones or fast-forwards each one, installs its deps into the venv, links the model categories it reads by bare filename, and fetches its modelSet — so a node pack is a rebuild, not a setup command, and a restart with nothing changed does no pip work. The shipped set is ComfyUI-GGUF, PuLID-FLUX and ReActor (the last two are what the face tools on :8894 drive). Adding one:

services.genai-server.comfyui.customNodes."ComfyUI-Frame-Interpolation" = {
  repo = "https://github.com/Fannovel16/ComfyUI-Frame-Interpolation";
};

Install failures are logged and skipped rather than taking ComfyUI down — journalctl -u comfyui is where a node that did not build shows up.

Single-photo identity (PuLID-FLUX)

The "upload one selfie, get images of that person" technique commercial sites use: zero-shot identity conditioning, no training. InsightFace extracts a face embedding from one reference photo and PuLID injects it during FLUX.1-dev generation. Instant, but it tends to mirror the reference's angle/hairstyle and captures only the face — a trained LoRA (below) still wins on full likeness and pose variety, and the two stack (add a LoraLoaderModelOnly node between the UNET loader and ApplyPulidFlux with a flux-trained LoRA — the shared LoRA store is a ComfyUI loras path, so trained/added LoRAs appear in its loader nodes). Same courtesy applies as for LoRA training: get the subject's okay.

Nothing to set up: PuLID is a shipped comfyui.customNodes entry, so the comfyui unit clones the node into its venv and fetches the pulid model set (~19GB — PuLID needs FLUX.1; the FLUX.2 already in ComfyUI is a different architecture) before it starts serving. Drop it on a host that does not want the download:

services.genai-server.comfyui.customNodes."ComfyUI_PuLID_Flux_ll".enable = false;

Use: open http://<host>:8188, load pulid-flux-workflow.json from this repo (drag it onto the canvas), upload a clear frontal face photo in the "Reference face" node, edit the prompt (photo-caption style), Queue. The first queue auto-downloads EVA-CLIP + facexlib (~1.2GB), so it sits loading for a few minutes once; after that ~30-60s per image. Knobs on ApplyPulidFlux: weight 0.8-1.0 = identity strength; raise start_at toward 0.2 for more prompt freedom at some likeness cost.

From chat (Open WebUI): attach a photo of the person and ask in plain language — "make an image of this person skiing in the Alps". The model calls the generate_image_with_face tool, which drives the same PuLID workflow through ComfyUI's API (reference upload, generation, ~1-2 min; the LLM is unloaded during the job and reloads on your next message). Say "stronger/weaker likeness" to adjust the identity weight. Needs the comfyui service up with its node installed — the tool returns a clear error otherwise. For people you generate often, a trained LoRA (below) still gives better likeness than one-photo transfer.

A style LoRA can ride along: pass lora (or several in loras) and the bridge splices LoraLoaderModelOnly nodes between the graph's UNet loader and ApplyPulidFlux, so the LoRA patches the base weights and the identity adapter goes on top. Only flux LoRAs qualify — every PuLID base is FLUX.1-dev, and anything else is refused with the list of ones that fit. Both bases take it: measured at 20 steps / 1024², ~15s on the fp8 safetensors and ~24s on the Q8 GGUF (ComfyUI-GGUF patches the quantized UNet), against ~18s with no LoRA at all.

Face swapping (putting a face into an existing photo)

PuLID generates a new image of someone. Three tools instead put a face into a photo you already have — same two arguments (url = the photo being edited, source_url = whose face), different trade-offs:

Tool How Time What survives
swap_face SAM 3 masks the face, FLUX regenerates that region under PuLID identity, composites it back ~1-2 min everything outside the mask, byte-for-byte; lighting and skin tone adapt
swap_face_full no mask — the whole frame is re-sampled img2img under PuLID identity ~1-2 min pose and composition only; clothing, background and detail are regenerated
swap_face_fast ReActor: inswapper generates a 128px face and pastes it in, then a restorer sharpens it seconds the entire photo except the face box — no diffusion model is loaded, so the chat model stays resident (it is evicted only if the ONNX sessions cannot find ~2GB, then the swap retries)

Reach for swap_face_fast first: it is the classic face swap, it is ~100x cheaper, and it cannot drift because nothing is re-rendered. The PuLID paths earn their minutes when the two photos disagree on lighting or angle and a transplant looks pasted on; swap_face_full is the last resort for a face the mask cannot cover cleanly. All three are in chat, the studio form (:8897/studio) and the MCP gateway.

All three come up with the stack: the PuLID and ReActor node packs are shipped comfyui.customNodes entries, and their weights (~19GB and ~2.8GB) are fetched by the comfyui unit before it serves. No setup command.

swap_face_fast is also the most adjustable of the three, because ReActor exposes the parts a transplant is made of:

Argument What it decides
model the swap network — inswapper (baseline), reswapper-256 (open reimplementation, 256px), hyperswap-1c (best at an angle)
weight how much of the swap to keep. The node returns the untouched photo alongside the swapped one and they are identical outside the face, so this crossfades the face alone — 0.5 is half-swapped, 0 is the original photo back
identity_mix a different thing: averages the two people's face embeddings so inswapper aims at a third identity rather than dissolving between two renders. Measured on this box it is a weak dial — inswapper rebuilds the entire face patch from its own prior either way, so 12.5% and 100% land close together. weight is the knob for "less swapped"
restore_model / restore_strength which restorer repairs the swapped face, and how much of its output is blended back. Drop the strength when the result looks airbrushed — GFPGAN smooths skin, gpen-512 keeps more of it
restore_fidelity CodeFormer's quality/fidelity dial; ignored by the ONNX restorers
face_boost upscale the swapped crop and repair it at full resolution instead of patching 128px in place. The quality knob; ignored by hyperswap
face_index / source_face_index which face, as a comma list — "0,2" swaps two people in one pass, pairing them with the source faces in order
face_order / source_face_order what those indices count through. The default is largest face first, so left-right is how you say "the second person from the left"
face_gender / source_gender only touch faces detected as female/male — picking someone out of a group without counting at all
detector retinaface (default) or YOLOv5, for a small or angled face that is missed entirely

The swap networks live in faceSwap.models and the restorers in faceSwap.restoreModels — both merge with host entries like every other catalog, and both are fetched by the faceswap set. Which store directory a swapper sits in (insightface/, reswapper/, hyperswap/) is what tells ReActor how to run it, so a new one is a mediaModels entry plus a faceSwap.models entry, never a code change.

Two things to know about ReActor. Its inswapper and buffalo_l weights are InsightFace's, released for non-commercial research use only (ReSwapper is the way around that: same job, trained from scratch). And upstream carries its own check over the input photos that replaces the result with a 512x512 black frame when it fires, saying nothing about why — a GitHub-policy artifact rather than a technical requirement, and one that costs a ~350MB download on first use.

This stack ships that node as its authors do. swap_face_fast recognizes the black placeholder by shape and reports it rather than handing you a black image, and points at swap_face (the PuLID path), which has no such gate. Turning the gate off is not something this repo decides: it is a comfyui.customNodes."ComfyUI-ReActor".patches entry, and whichever module a host imports to make that call owns it. Being declared means it is also reverted — drop the declaration and the next comfyui start puts the file back.

Same courtesy as everywhere else here: these are real people's faces.

Training LoRAs of people (family photos)

lora-train fine-tunes an image model on your own photos so it can generate arbitrary new images of specific people (LoRA subject training, via ostris/ai-toolkit — self-installs to /var/lib/ai-toolkit on first use; jobs live in /var/lib/lora-jobs). Run it as a user in the genai group. Everything runs and stays on this box. Get the subjects' okay first — these are real people.

Three base models, one trainer:

Train time (5090) Likeness Generate with
flux (FLUX.1-dev) ~2-3h best, most proven :8893 with "model": "flux-dev", or ComfyUI
zimage (Z-Image-Turbo) ~1h very good :8893 default engine = the chat image button
pony (Pony V6 XL / SDXL) ~1-1.5h stylized/anime — not for photoreal :8893 with "model": "pony"

LoRAs are architecture-bound: a LoRA only works with the base family it was trained on (the install sidecar records this, and every path — chat tool, image button, raw API — routes to the right engine automatically; the "Generate with" column is where each base lands, not something to type). Community LoRAs from CivitAI etc. are declared, not installed — see Adding a LoRA below. Either way they are usable from chat by name or via <lora:name:0.8> prompt tags. Remove a locally trained one with lora-train uninstall <name>.

One-time for flux only: the base model is gated — accept the license at huggingface.co/black-forest-labs/FLUX.1-dev and export HF_TOKEN=hf_... before lora-train run. (pony downloads its shared-store checkpoint the image server already uses — one download; not gated.)

Walkthrough (one person)

lora-train new dad flux erx_dad        # scaffold /var/lib/lora-jobs/dad/
cp /path/to/photos/*.jpg /var/lib/lora-jobs/dad/dataset/
lora-train caption dad                 # captions via the local qwen-dense vision model
lora-train run dad                     # train; checkpoints + samples in .../dad/output/
lora-train install dad                 # deploy it for chat + API use

Dataset: 15-30 photos, varied — closeups and half/full body, different angles, lighting, expressions, clothing, backgrounds. Crop other people out. The trigger word is what binds the identity; pick something that isn't a real word (erx_dad, not dad). Review the generated .txt captions before training — they should describe scene/pose/clothing, never identity traits (hand-edit freely; re-running caption keeps existing files). During training, sample images land in the output dir every 500 steps — if likeness is good early, you can stop and use the latest checkpoint.

Training wants the whole GPU for hours: run stops llama-swap for the duration (chat and STT are both down while it trains — STT is a llama-swap model) and restarts it when training finishes, crashes, or is interrupted. Tune anything else by editing /var/lib/lora-jobs/<name>/config.yaml before run.

(ai-toolkit also ships a Next.js web UI. It isn't wired up here: its npm install pulls native deps — Prisma engines, sqlite3, sharp — that assume an FHS system and don't resolve on NixOS. The CLI covers the whole workflow.)

Using the result

lora-train install dad     # newest checkpoint -> /var/lib/genai-models/loras

That copies the newest .safetensors from the job's output and writes a dad.json sidecar recording the trigger word and base model. Both the image server and the media tool server read that shared directory. (For ComfyUI, ComfyUI reads the same store — no extra copy needed.)

From chat (easiest). Installed LoRAs are listed in the generate_image tool description, so any tool-enabled model can pick one by name — just ask in plain language:

make a picture of dad jumping on a trampoline in the backyard

The model passes lora: "dad", and the server prepends the <lora:dad:0.9> tag, injects the erx_dad trigger word, and selects the right base model automatically. Say "use a stronger/weaker likeness" to nudge lora_strength (0.1-1.5, default 0.9). Asking for several installed subjects in one scene works too — the model passes loras: ["dad", "mom"] and every tag and trigger word is injected; per-subject weights ride along as loras: ["dad:0.7", "mom"] ("make dad's likeness weaker"), with lora_strength as the default for entries without their own. The LoRAs must share the same base model, and expect some identity bleed between subjects — stacking person LoRAs is a known diffusion weak spot. After installing a new LoRA, restart Open WebUI if the model doesn't seem to know about it — the tool spec is fetched per connection and may be cached.

Direct prompting (Open WebUI's image button, or the API) takes the tag yourself, and there the trigger word is mandatory:

curl -X POST http://<host>:8893/v1/images/generations \
     -H 'Content-Type: application/json' -d '{
       "prompt": "<lora:dad:0.9> photo of erx_dad sailing a boat at sunset"}'

"model" can be omitted: a tag naming a registered LoRA selects the engine its base needs (here flux → flux-dev), overriding whatever model the request carried — a mismatched combo would only make sd-cli skip the weights. Tags naming unregistered files (no .json sidecar) leave the requested model alone.

Write prompts as photo captions (subject, pose, setting, lighting, framing) rather than instructions — that matches the caption style the LoRA was trained on. Strength 0.8-1.0; lower it if outputs get stiff or over-baked.

Grainy LoRA output? Z-Image-Turbo generates in 8 steps, but a LoRA bends its distilled sampling path and 8 steps leave that path under-denoised — the residual is grain. The server auto-raises any LoRA-tagged Z-Image request to LORA_STEPS (16), which clears it while keeping bare generations fast; tune that env in module.nix. Raise it further per request with "steps": N on the API (or "make it more detailed" in chat, which sets the steps tool param). If grain persists only at high strength but low strength looks baked, the LoRA likely overfit — try an earlier checkpoint, or use the FLUX path, which samples at 20 steps and is inherently cleaner for faces.

LoRAs are architecture-bound (a flux LoRA loads on any FLUX.1-dev engine, but not on FLUX.2 — different architecture), but for registered LoRAs the server enforces the match itself: the tag's base picks the engine everywhere — the generate_image tool, the chat image button (which always sends its one configured model — the tag overrides it), and raw API calls alike.

Editing and face tools take LoRAs too. edit_image, reimagine_image, inpaint_image, smart_edit and the two PuLID tools (generate_image_with_face, swap_face_full) all accept the same lora / loras / lora_strength arguments, and the studio's start from an image tab shows a picker for them. One difference matters: these run one fixed model each — Kontext, schnell, the inpaint engine you picked, or the PuLID graph's FLUX.1-dev base — so there is nothing to reroute to and a LoRA from another architecture is refused with the reason rather than switched around. A flux LoRA for the FLUX paths, a pony/sdxl LoRA for an SDXL inpaint engine. Trigger words are still injected for you.

On the two face tools, prefer a style LoRA. A LoRA of a person is a second identity source aimed at one face. Measured on this box the photo won — rita_flux at 0.8 under identity strength 0.9 changed framing, wardrobe and colour grade while the face stayed the reference's, on the fp8 and Q8-GGUF bases alike — but when it goes the other way you get a convincing third face rather than an error, so check the likeness before trusting one.

Managing installed LoRAs

lora-train list              # what's installed, with trigger word + base
lora-list                    # same, plus per-LoRA usage examples: the webui
                             # chat phrasing and a ready-to-send API request
lora-train uninstall dad     # remove it (needs sudo)

Uninstalling deletes only the deployed copy; /var/lib/lora-jobs/dad/ keeps the dataset, config, and every checkpoint, so you can retrain or install again later. Both servers read the directory per request, so removal takes effect immediately with no restart — though Open WebUI may keep advertising the name until it refetches the tool spec, and a stale request naming a removed LoRA just returns an error listing the valid ones. If you also copied the file into ComfyUI's own models dir at some point, remove that copy too.

Not happy with the results? Before discarding it, try dropping the strength (<lora:dad:0.6>, or "weaker likeness" in chat) — over-baked output that ignores your prompt is usually too much LoRA, not a bad LoRA. Pin the seed first, or you are comparing two different rolls and cannot tell a fixed LoRA from a lucky one: take the seed off a render (the studio caption, the gallery's metadata, or the seed a tool returns), put it in the Seed box, and change only the strength. Mangled anatomy that clears up at a lower weight was over-driving; mangled anatomy that survives every weight at the same seed is the LoRA. Failing that, an earlier checkpoint from the job's output/ is often better than the final one, since person LoRAs overfit late in training.

Multiple people in one image: stacking two person-LoRAs blends faces. Either train one LoRA on a joint dataset with a distinct trigger word per person, or generate the scene with one person and fix the other's face via /inpaint (mask the face, prompt with that person's LoRA + trigger).

Portal (:8897)

The single entry point. One nav across every page it hosts, one directory of everything else, one health answer.

Overview (/)

Problems first: anything down, wedged, or enabled-but-not-downloaded appears in a strip at the top (and in the health pill in the nav, on every page). Below that, live GPU / CPU / RAM gauges; the service directory grouped into Chat & agents, Generation, Serving and Tools & APIs, each card carrying a TCP probe and an HTTP health check where the service has one (a green dot means answering, amber means listening but failing its health path); and the model fleet.

Backends with no browser UI (tts-hq, segment-server) render as status chips rather than links — they are monitored, not clickable.

ComfyUI's card is administrators-only (comfyui.adminOnly, on by default, and portal.extraServices.*.admin for a host's own links). Not about the software: it is a full node editor with no identity of its own, where every checkpoint and LoRA on the box is one dropdown away, so it is the one surface here that cannot be filtered per person. That makes it DISCOVERY rather than access — the portal links to ComfyUI rather than proxying it, so anyone holding the URL still reaches it, and a gate that holds belongs in front of the service. Turn it off once ComfyUI runs per user.

Who counts as an administrator is the same question everywhere in this stack, answered once: the per-person setting on /admin first, then the platform's role header, then identity.admins. An override of off beats a platform role of admin, which is the case that makes it an override rather than a hint.

Model fleet (/models)

Every configured model with its state (ready / downloading / absent), size, an enable toggle, and download/remove buttons — plus what it can do (capability chips) and how it is served (context, KV quant, MoE offload, group, TTL). Downloads use the parallel prefetcher. State choices persist in /var/lib/genai-dashboard/enabled.json (runtime, survives restarts; a from-scratch rebuild starts from the flake defaults). The Overview page keeps a two-card summary (swappable / always-on) that links here.

The serving detail is not a description of the config — it is the config: it comes from each catalog entry's serve block (see Adding an LLM), and for the hand-tuned models an assertion checks the declared context against the command that actually runs. Models that came from that declarative path are badged generated.

Evals (/evals)

What was measured on this box, rather than what the model cards claim: the latest genai-eval report with its failures spelled out, any model comparison as a per-model table, and the run history — click a run to load it. Reports are plain JSON in /var/lib/genai-eval, so anything on this page is equally available to a script.

Read-only on purpose. The page never starts a run: suites load models, and the process serving this page is the one answering the health probes. Runs come from genai-eval on the box or a timer.

Model comparison gets a warning banner when no model cleared the suite, because that is as likely to mean the harness is broken as that the models are bad — which is exactly what happened the first time this ran (see ROADMAP Stage 14, and the note under Adding an LLM about thinking models and token budgets).

API surface

  • /v1/* — the filter proxy Open WebUI and LiteLLM point at instead of llama-swap directly: only ready+enabled models appear in /v1/models, and a chat request for an unavailable model returns a clean 503 instead of triggering a slow single-stream load. (Open WebUI's model dropdown may still list tool-wired models that aren't downloaded — selecting one 503s.)
  • /svc/<name>/* — reverse proxy to this flake's own stdlib services: image, media, search, tts, segment. So curl :8897/svc/media/openapi.json reaches media-tools, and an API client needs one origin for the whole box. Strict allowlist, built from the port map — an unknown name 404s. JSON APIs only; the portal is a stdlib HTTP server and does not proxy websockets, which is why the third-party UIs stay links.
  • /api/health{ok, problems[], services, gpus, cpu, mem, models}, and a 503 status when not ok, so curl -f :8897/api/health is a valid uptime check. This is the thing to poll.
  • /api/portal — the hosted pages and the grouped service directory, for anything that wants to render or discover the surface.
  • /api/evals — the newest genai-eval report in full plus a summary of the rest; /api/evals/<file> fetches one by name (matched against the directory listing, so the name cannot escape it).
  • /svc/mcp/mcp — the MCP endpoint (see below), so an MCP client points at the portal like everything else.

Extending it

services.genai-server.portal = {
  title = "logistikon";
  extraPages = [
    { path = "https://grafana.lan"; name = "Grafana"; desc = "long-term metrics"; }
  ];
  extraServices = [
    { name = "Jupyter"; port = 8899; kind = "ui"; group = "Tools & APIs";
      desc = "notebooks"; health = "/api/status"; }
  ];
  proxy.extraUpstreams.notebooks = "http://127.0.0.1:8899";
};

Health probes, /api/health and the genai_service_* metrics pick up added cards automatically.

Everything the box has ever generated, newest first — studio renders, chat tool results and published-workflow clips alike, because all three write one store (mediaStore.dir, below) and this is a listing of it rather than a copy of it. Search, engine, media type, LoRA, date and sort-order filters run server-side against the whole store; Load more pages through it.

Images, video, and what made them are two controls, not one. Images & video narrows to one or the other and is derived from the file itself, so nothing has to be recorded for it to work — which matters, because "anything that moves" must not be a list of engine names somebody keeps up to date as graphs are published. The engine picker is that list, and it now holds the published workflows as well: a clip's engine is its graph (wan22_video, ltx23_animate, …), grouped apart from the image engines in the dropdown. That is derived from the operation the record already carries rather than written into a new field, which is why the clips already in the store have it without being re-rendered. The two compose, like every other filter here: video plus ltx23_animate plus a prompt word is one server-side query.

Reaching the rest of the archive. Paging alone does not: a few thousand renders is thirty-odd presses from the oldest one and no number of presses from the middle. Two controls answer that. Newest/oldest first flips the whole listing, so the beginning of the store is one click rather than thirty. The date picker offers only months the store actually has, each with its own count — so it doubles as a map of where the renders are — and paging then works inside the month. Both narrow the same server-side query the search box does, which is what makes them compose: July plus flux plus a prompt word is one request, not a page filtered afterwards.

Each month travels with the timestamps that define it and the page sends those back unchanged, rather than turning "2026-03" into a window on its own clock. The browser is not necessarily in the box's timezone, and a month it bucketed itself would disagree at both edges — by a few hours, on exactly the renders nearest the boundary. The windows are half-open ([from, to)), so adjacent months never claim the same midnight and the counts add up to the store.

The date window also reaches Select all matching, which is the half that is easy to miss: a bulk action has to be computed from the query that drew the grid, or it ticks — and Delete is one button along — renders that were filtered away and never shown.

Searching by what is in a picture (portal.semantic.enable, off by default). The search box above matches the PROMPT, which is the request rather than the result: silent about everything the model decided, wrong wherever the model ignored it, and absent altogether from the entries that say no parameters were recorded — which on a long-lived box is most of its history. A second box looks at the picture instead. Type a description, press Enter, and the listing comes back closest-first; any entry's viewer also offers Find similar, which is the same ranking taken from that render's own vector.

It runs on a dual encoder — SigLIP 2 by default, portal.semantic.model to change it — on the CPU, and deliberately: the card is shared and an index build is exactly the long hold that must not take it. Each encoder keeps its own index, so trying another one is free and switching back rebuilds nothing.

Staying current (portal.semantic.autoIndex, on by default) has two halves, separately switchable, and the split is the point. A render the PORTAL writes — anything made in the studio — is embedded within seconds of finishing (indexOnWrite). That covers what the portal sees, and is deliberately not trusted to cover the store, because it does not: ComfyUI writes the media directory directly and so does media-tools, neither through this process. A hook on the portal's own write path alone would leave whole categories of file permanently unfindable while looking complete. So a periodic sweep (sweepMinutes) is the backstop, keyed on the directory the way the gallery listing is, which is what makes it catch a file whatever wrote it. Each pass is bounded (sweepBatch) so a first sweep over a full store spreads across passes instead of holding the CPU on a machine somebody is working on; it converges without bookkeeping, since what is left over is simply found again next time.

The sweep is the load-bearing half and indexOnWrite is the one to reach for if either misbehaves — hence its own switch. It is worth being honest about what it earns: it buys freshness for the renders least likely to need a content search (you are looking at the thing you just made), it spends its CPU at the busiest moment, and it is a second code path that can fall silent without reporting anything. What it does buy is a clean "N of M indexed" line, which otherwise reappears on every render until the next sweep. A no-op sweep costs 7ms here against the full store, so a short sweepMinutes is a genuine alternative to switching it on at all.

A full build is also available on demand — Build search index…, administrators only — and is store-wide rather than per-viewer, because the index is machinery like the thumbnail cache: one built from one person's view would leave everybody else's search silently empty. It grants no visibility; results still pass the ownership filter and whatever any optional module keeps back.

Measured on this box (2,396 renders, 485 of them video, SigLIP 2 base): 8.8 minutes to build the whole index at 4.5 files/sec — videos are the slow part at ~1.6/s, since each costs three ffmpeg keyframe seeks. The index is 7.4MB (2,396 × 768 float32). A typed query is ~3.4s, nearly all of it loading the encoder, which is why the box searches on Enter rather than per keystroke; Find similar loads no model and is effectively instant.

Retrieval was checked against an independent labeller rather than against its own model card: a set of per-file scores this store already carried, produced by a different architecture, is a free labelled set. Ranking agreement (AUC, 0.5 being chance): 0.982 for a query matching those labels, 0.259 — correctly inverted — for one that should not. The embeddings encode real picture content on this store.

Describe the picture, not the medium. The encoder sees frames, so it has no notion of motion or of a file "being a video": a video of a person moving is a category error and returns nothing useful, while a person mid-stride works. Clips are not second-class otherwise — measured here, content queries return 58-67% video in their top 100 against a store that is 20% video, so they rank if anything better than stills.

Two things worth knowing about how it behaves. The ranking is computed over the whole index and intersected with what you may see afterwards, so a top-N means N of your renders rather than whatever survived a cutoff taken over somebody else's — and a search can only ever remove entries, never reveal one. And a typed query loads the encoder per call, a second or three, which is why it searches on Enter rather than per keystroke; Find similar loads no model at all and is effectively instant, since the vector it ranks against is already in the index.

If the index does not cover the whole store yet, the gallery says so under the count. That line matters more than it looks: a search over a third of the store returns thin results that read as a bad model rather than as an unfinished build, and nothing else on the page would ever say which.

Tiles are thumbnails, generated once with ffmpeg and cached in the portal's state dir. That is not a nicety: a grid of sixty full-size renders is tens of megabytes to draw postage stamps (31.5MB measured, against 1.0MB now), and a video tile has no frame to show at all until enough of the clip has arrived — which is why clips appeared as black boxes that played perfectly when opened. Every tile is a still now, with a marker on the ones that are clips; the viewer still loads the real thing. If ffmpeg is missing the tile falls back to the original, so the picture is late rather than absent.

"Load more" appends below what you are reading and does not move the page: the picture you were looking at stays where it is and the new ones arrive under it. (An earlier version held the BUTTON under the pointer instead, by scrolling down past everything it had just added — which put you at the bottom of the page every time you asked for more of it.)

One caveat that took three wrong diagnoses to find: the portal now speaks HTTP/1.1 with keep-alive and listens with a backlog of 128. The stdlib defaults are HTTP/1.0 (a new connection per response) and a backlog of five — so a page asking for sixty thumbnails opened sixty connections and the kernel dropped whatever did not fit, before the process could see them. Blank tiles, and an access log showing nothing but 200s, because a connection that is never accepted is never logged. Tiles also retry once if their picture fails, which covers the same class of hiccup in a proxy.

Tiles are not lazy-loaded, which is deliberate: a page of thumbnails is about a megabyte, less than one of the renders it depicts, and deferring them broke the grid. "Load more" appends a page below what you are reading, so with lazy loading those tiles could sit unloaded until scrolled to — a tile showing the background is indistinguishable from a failed render, and it opened perfectly when clicked.

Each entry shows what produced it, read back out of the file itself: the genai record this stack stamps into a PNG (and writes beside anything that cannot hold one), falling back to stable-diffusion.cpp's own parameters chunk. Entries that carry neither are still listed, saying so — most of a long-lived box's history predates any of this, and an entry that says the parameters were not kept is worth more than an entry that does not exist. Where the record is complete, Regenerate re-runs it exactly (same seed, same inputs) and Open in studio hands it to the form to change first. Both go through the studio's form — Regenerate is the second one plus the Run — so a re-run is a thing you can see the settings of before it lands, and there is one path to a render rather than two implementations of it. A re-run that could not be restored faithfully (an engine or LoRA the box no longer has) is loaded but not started, and says which; neither is one fired into somebody else's render while the card is busy. Clicking the picture itself opens it at full size in a new tab — a plain link, so middle-click and "open in new tab" behave normally. A clip is not a link: there, a click belongs to the player.

An edit, a face swap or a workflow also shows what it was made from: a thumbnail per input under the parameters, each a link opening that file at full size in a new tab. Those files are the ones kept in sources/ so the render could be repeated, and the gallery deliberately does not list them — the raw material of a render is not a render — so this is the only place they can be reached from. A reference clip gets a frame like any other video; a voice sample has no frame to take and shows as a link with a mark instead. Nothing is stored for this: the thumbnails are derived from the record every edit already carries, so entries made long before the panel existed have it too. An input whose file has since been deleted by hand says so where its name is, rather than leaving a broken picture to mean it.

Duplicates from before the stores were merged. Every tool result used to be written twice — media-tools made it, the portal fetched it back and stored its own copy — and only the portal's copy carried a record. Both are in the store now, which is why an entry could sit next to an identical one saying nothing was recorded. The portal reconciles them on start: copies are paired by their image data (PNG text chunks ignored, since the stamped copy differs from the plain one by exactly that), the pair keeps whichever record either half had, and the redundant copy is marked so the listing shows the render once. Nothing is deleted and no name disappears — a chat message linking to the other copy keeps working — and the space is reclaimed, if you want it, through Prune's duplicate copies rule. New renders are written once, so this converges to a no-op.

Nothing ages out. A render is the one artifact here that no rebuild can reproduce, so deleting is always something a person did:

  • Select → tick entries → Delete, which names the count and the bytes before it does anything.
  • Delete in the viewer, for the one you are looking at — including the studio's strip, which opens the same viewer with the same buttons.
  • Prune applies one rule — older than N days, keep only the newest N, everything with no recorded parameters, masks and previews, or duplicate copies — and previews it first: the button that deletes carries the count the server just reported, and changing the rule takes that count away again. Input images kept for repeating an edit are never selected, and neither is the copy of a duplicated render that the gallery lists.

When the disk holding the store drops below mediaStore.diskWarnGB (default 20GB) a banner says so, on this page and in the studio. It never deletes anything to get back above it.

Studio (:8897/studio)

One form over three different backends, grouped by what you are trying to do rather than by which service does it — Create, Start from an image, People. Pick an operation, fill in the form, press the single Run button in the bar at the top. That bar is also where you see what is running and for how long, and it disables Run whenever the GPU is busy — including work started somewhere else entirely, like Open WebUI or a chat tool.

The separation the page hides is kept at the boundary, in studio_op_target():

operation backend
generate_image image-server (:8893) — engines, LoRAs, sampler knobs
the nine edit/face operations media-tools (:8894) — fixed graphs, hand-written forms
workflow_* published comfyui.workflows — fields generated from the graph's own schema, minutes-long, possibly video

It is a whitelist per family, not a proxy: transcription and TTS live on :8894 too and are deliberately unreachable from here.

Reloading loses nothing. Renders are jobs on the server, so a reload rejoins one already running; the form's contents are saved as you type and come back with it — including what you attached.

Attachments are uploaded when you pick them, not when you press Run, and what the form keeps is the link. That is the whole trick: the box has to receive these bytes anyway (every input is written into the media store so the render can be repeated), so doing it at pick time means the form holds a URL a few dozen bytes long. It goes in localStorage with everything else, the picture is there after a reload — or in another tab, or on your phone — and the request that starts the render carries a link instead of megabytes of base64. Nothing large is kept in the browser at all. A render opened from the gallery skips the upload: its inputs are already in the store, so the form points straight at them.

Video results land in the same place as images: one media store, listed in the gallery.

Below the form, Recent is the gallery itself — the same component the Gallery page mounts, showing the newest dozen with a link to the rest. It is server-backed, so it is what the BOX has made rather than what this browser remembers: a phone that has never opened the studio still sees every render on it. Opening one gives the same viewer as the gallery — the same buttons in the same order, because it is the same component and the list of them is the component's, not the page's. Three of them write into the form: Regenerate (its settings, on the tab and operation that made it, and then Run), Load settings (the same without the Run — what the gallery page calls Open in studio, since there it navigates here first) and Load seed only.

The three are told apart by what they touch, which is why they are worded alike. Load settings replaces the form — operation and tab, prompt, engine, size, steps, LoRAs and their weights, and the inputs an edit started from. Load seed only touches nothing but the seed box, keeping the form you have been typing in; it is offered on renders Load settings is not, since a record can have kept the seed and lost the engine. The include seed tickbox beside Load settings is the difference between "this picture, then let me change something" and "these settings, rolled again": unticked, the seed box is left blank and the studio says so. It is remembered, because which of those you are doing you are usually doing all afternoon. Regenerate ignores it and always pins the seed — a regeneration that re-rolled would be a different picture wearing the word. Two more buttons act on the FILE rather than the recipe: Full size opens it in a new tab and Download saves a copy.

Every tab carries a Seed box (blank = random). Every render reports the seed it used — in its gallery entry, and as a Load seed only button in the viewer that fills the box for you — so "that one again, but weaker LoRA" is a real request instead of a re-roll and a hope. This is the tool for diagnosing a LoRA: hold the seed still and change one thing. It covers everything that samples; create_mask and swap_face_fast have no seed because neither is stochastic.

The start from an image and People tabs drive the same tools chat calls — edit, reimagine, inpaint, smart-edit, the three face swaps, mask preview — and have its own LoRA picker for the ones that take them. There the model is fixed by the operation rather than chosen to fit the LoRA, so the list shows only what that operation can bind (and follows the inpaint engine you pick); a selection that stops fitting is dropped rather than silently sent.

Below the form, a live reference explains how checkpoints/checkpoint-merges/fine-tunes (full models = engines) differ from LoRAs (adapters bound to a base architecture), and lists every generation engine, editing pipeline, installed LoRA (with trigger word and chat/API usage), and chat tool — sourced from image-server's /catalog and media-tools' /openapi.json, so it's always current.

3D view (/viz)

The box as a machine: the GPU, RAM, the model store and the CPU, with data moving between them at the rate the counters actually report, and a band of software chips above showing which services are running and what each one is driving.

A chip lights up on evidence, in two tiers. A service that publishes what it is doing gets quoted directly — llama-swap names the model it is holding, image-server and media-tools name their running job, and a ComfyUI graph is reported by the media-tools call that queued it (ComfyUI's own /queue is never polled: that would defeat its idle stop). Everything else — LiteLLM, Magentic-UI, Open WebUI, the tool servers — publishes nothing per request and never will, since most of it is somebody else's code. Those fall back to their systemd unit's cgroup CPU (portal.extraServices.*.unit for your own cards), which is a measurement of that service and nobody else, and they draw their control line to the CPU rather than the card.

The comparison is against each service's own idle floor, not a fixed threshold, and that is measured rather than chosen: the shipped services idle anywhere from 0.03% of a core (most of the Python servers) to 13.8% (magentic-ui, which is simply that busy doing nothing). One threshold would either light the busy ones permanently or never notice the quiet ones working.

What no tier will do is guess. A service with no cgroup — anything socket-activated and currently stopped — reports up/down and nothing else, because this view is read to find out who is holding the card, and an invented "busy" is worse than no answer.

The card (/gpu)

One GPU, several people, and every engine here was written as if it were alone with it. Each resolved contention by evicting whatever was there: a render asks llama-swap to unload before it starts, and a chat model starting asked systemd to stop ComfyUI. Both were correct for one person. The second was destructive even for one person — a coding harness left running would stop a video render its own owner had forgotten about, mid-graph, and the render died with its VRAM.

So every path that wants the card asks the portal first, and is told either the id of a lease or what is in the way. A lease carries an expiry its holder renews while it works, which is the part that matters: a crashed render gives the card back by failing to renew rather than by somebody noticing. It is not a scheduler — there is no queue and no fairness. A blocked request waits out gpu.arbitrate.waitSeconds (20 by default, in case the card comes free) and is then refused with a message naming what is in the way.

The one asymmetry is deliberate, in both directions:

wants the card free a render is running chat is running somebody's hold
a render yes yours: yes · theirs: no yes yours: yes · theirs: no
chat / an agent yes no yes yours: yes · theirs: no
training yes no no yours: yes · theirs: no
takes no card¹ yes yes yes yes

¹ embed, rerank, asr, voice, fast-cpu — everything the catalog marks serve.group = "resident" or serve.device = "cpu", which is the same predicate genai-gpu-claim uses to decide who may skip yielding.

A render takes the card from chat safely — it calls /unload first and waits for the drain — so gating renders on chat would disable the studio for most of the day over a collision that cannot happen. Chat taking it from a render is not safe at all, and that is the direction that was losing work. Set gpu.arbitrate.chatBlocksRender = true to make the card strictly exclusive instead. Anything that does not take the card is exempt entirely — the resident set (~2.6GB of embeddings, rerank and STT), because refusing a memory lookup while somebody renders would break retrieval for no gain, and the CPU models (voice, fast-cpu), which use no VRAM at all and exist precisely so a spoken turn never has to fight for the card. Your own render does not block your next one — the engines have always queued those, and queueing beats refusing.

Refusals look the same everywhere: 503 with a Retry-After, which the OpenAI, Anthropic and Vercel SDKs already back off on, so a coding harness handles a busy card without being taught anything. 429 would be wrong — it means "slow down", not "somebody else has the machine".

Holding the card (gpu.holds.*) is for working in bursts, where the gaps between renders are exactly when somebody else's agent session moves in. It lasts until released, until holds.idleMinutes passes with no work started under it (15 by default — the forgotten-tab guard, and what makes a manual lock safe to offer), or until holds.maxMinutes. An administrator can break one, and the person whose hold it was is told so on their own page rather than left to work out why their next render was refused.

A hold shows in the nav on every page, as a red on the activity pill beside the health pill — because a hold is the state you forget you are in, which is the whole reason it has an idle timer. It rides on the pill rather than replacing it: an idle-but-held box reads held by you in red, while a hold during a render keeps the render's line and adds the marker, since the render is the more urgent fact and both are true. The marker survives the narrow layout that hides the pill's text, and it disappears when the poll fails — an absent marker must not read as "not held" when nothing was learned.

Stop reaches every surface from one button: studio jobs, image-server renders, chat-tool calls, and in-flight chat or agent streams — closing the upstream socket, which is the only cancellation llama-server offers. Your own work from anywhere; anybody's if you administer the box.

Who you see. An administrator sees who is on the card, because arbitrating a shared machine means knowing who to talk to. Nobody but the owner sees the label or the engine — a prompt is somebody's work, not a fact about the schedule — so everyone else gets "someone else · a video render · 4m". That line is drawn in the API, not the page.

What this does not cover, and it matters. ComfyUI's own interface has no accounts, so a render started there is attributed to nobody and (with gpu.cancel.comfyuiUi) stoppable only by an administrator. The engine ports on the LAN — :8893, :8894, :8188 — take work from anyone who can reach them; genai-gpu-claim refuses to evict a busy ComfyUI and asks the arbiter before a model starts, which covers anything that goes straight to :8080, but the whole arrangement is cooperative. This keeps people from walking into each other. It is not a boundary against somebody determined.

It fails open, everywhere. An unreachable portal means every engine behaves exactly as it did before any of this existed. A box that cannot render because a status page is down would be a worse failure than an unarbitrated one.

Coding harnesses reach the LLM through LiteLLM, which builds a fresh request and drops what the client sent — so general_settings.forward_client_headers_to_llm_api is on, and a harness that sets one of gpu.identityHeaders (X-Genai-User, or Open WebUI's X-OpenWebUI-User-Id) gets named. One that sets none is a single shared unattributed bucket: it may use a free card, it may never hold one, and anybody else's hold refuses it. Never the owner — the failure mode is being turned away, not inheriting somebody's session.

Admin (/admin)

Only present when identity.mode = "trusted-header" — a single-user box has one occupant and nothing to divide, so the page is not routed there rather than rendering a table with one row. It is also left out of the nav for anyone who is not an administrator: the API behind it refuses them whoever reaches it, and an entry that answers "you are not an administrator" is furniture rather than navigation.

One row per person the box has seen, and for each: whether they administer it, whether they are restricted to models that refuse what their vendor made them refuse, whether they may hold the GPU, and a column for each group of settings an optional module declares. Every control is tri-state — on, off, or default, and the third is not decoration. Losing it would make "I decided this" indistinguishable from "I happen to agree with the box", which is the thing the page exists to show; clearing a control deletes the setting rather than storing today's default, so a later change to the host config still moves that person.

Administrator resolves in this order, and the row says which one it landed on:

a setting made here wins over everything, so a decision is not silently undone by an SSO role changing
the platform's role header an SSO admin works on arrival, with no rebuild
identity.admins the operator's build-time list

You cannot remove your own access — the API refuses it and the control is disabled rather than left to fail. This page is the only door, so doing it would leave the box with one fewer administrator and no way for that person to undo it, possibly none at all.

Two limits worth knowing. The role header describes the caller and nobody else, so the box learns what the platform thinks of someone only when that person makes a request; until then their row reads not an administrator, and an explicit setting is the only way to decide about somebody in advance. And there is no user table — people appear here once they have opened the portal or made something, so somebody who has only ever used another service is not listed yet.

Being an administrator does not include reading anybody's renders. The gallery is filtered by ownership for everyone, operator included: private by default has to mean private from the operator too, or it is a setting rather than a property. Prune remains an admin operation and still walks the whole store — deleting by a rule without being shown anything is a different power from browsing, and it is the one an operator running out of disk needs.

An unidentified request is nobody, not the owner. If the proxy does not send the identity header, the caller owns nothing, sees nothing, is not an administrator whatever the role header says, and cannot write — the gallery says so rather than looking empty. The one exception is loopback: the box's own services (the seeder, media-tools, the voice server) reach this port with no proxy in front of them to be labelled by, and act on the box's behalf. This is the failure that has to fail closed. When it did not, an authenticated stranger whose header the proxy happened to omit became the owner — the owner's whole gallery, the owner's admin rights, and no second name anywhere on this page to hint that anything had gone wrong.

Unfiltered models (uncensored.*)

Some models decline whole categories of request — a base model's safety training, or a vendor's political filtering. Others are the same weights with that direction removed, or were never trained to refuse at all. A catalog entry says which with uncensored = true, and this stack reads it.

Both directions of that choice are worth having. It is why the abliterations are on this box at all: a refusal on a legitimate medical, security or historical question is a wrong answer, not a safe one. It is also a reasonable thing to keep an account away from on a shared machine.

uncensored.showDefault (on) is whether a person is offered what is marked — the portal's model and engine lists, the chat tools' schema, the Open WebUI picker. Turned off, those drop everything marked and a request that names one outright is refused. A DEFAULT, overridable per person on /admin, and phrased as show so that every control on that page reads "on means they get it".

On, because what turning it off can do depends on the medium and pretending otherwise would be the mistake. For chat it yields a real subset — official Qwen/gpt-oss against huihui/Heretic — so restricting somebody to it means something. For images essentially every local checkpoint is unfiltered, so an honest marking marks nearly all of them and this alone would leave a restricted person with almost nothing to render on.

It filters the SOURCE lists, never a rendered schema: list_loras() and the engine tables feed the tool spec, the parameter descriptions and the validation, so one filter covers what is offered and what is accepted, and they cannot disagree. It is not a permission system — :8893 and :8894 are on the LAN and take every engine from anyone who can reach them.

Labels this stack does not read (flags)

Every catalog entry carries a free flags attrset, and nothing here reads any of it. A flag is a claim about a model that some other module acts on — one a host imports beside this one — and it is free-form rather than a fixed set of booleans because the module that cares may not be loaded.

That is the whole point. An entry keeps its labels either way, plugins.<name>.* keeps its settings either way, and turning such a module off is a line in a host config rather than an edit to every entry that mentioned it. LoRA sidecars carry them through as well (lora-add --flag, lora-train install --flag, and mediaModels.<name>.flags via genai-fetch-media), so a reader working from the store sees the same labels as one working from the config.

See Optional modules for what such a module can attach to.

Optional modules (plugins, pluginModules, pluginAssets)

Some behaviour does not belong to everybody who runs this box, and the honest place for it is a module a host imports beside this one. A model catalog was the first of those and needed nothing but the module system: catalogs ship at mkDefault, so another module's entries merge with them. Behaviour is harder, because the services are single-file Python and a page is a single file of HTML — there is nowhere for a second module's code to go. Four options are that nowhere:

option what it supplies
plugins.<name>.* a settings namespace this stack declares and never reads
pluginModules Python loaded into the portal and media-tools, asked for the extension points it implements
pluginAssets browser code concatenated as /assets/plugin.js, which every page loads with the chrome
pluginEnv / pluginPackages configuration and commands for both halves

Plus lib.flaggedNames, lib.torchEnv and lib.fetchMedia, so a module does not carry its own copy of catalog-name flattening, a second gigabyte of torch, or an instruction to run a fetch by hand.

Absence is not an error. plugins is freeform: a host that sets plugins.foo.something = true and then stops importing foo gets a value nothing reads, not an evaluation failure. That is what makes "try the box without it" a one-line change.

A failed import is fatal. These modules exist to filter, refuse and relabel; one that quietly did not load leaves a box that looks configured and behaves as though it were not. The service fails to start and says which file and why.

This is not a sandbox. A plugin module runs inside the portal's own process with its privileges — it is another way to write part of this stack, kept in another repo, and it is trusted exactly as much as this repo is.

The screen as a status display (kiosk.enable)

Off by default. When it's on, the machine's own monitor becomes the status display: a job starts, and if nobody has touched the keyboard for kiosk.idleSeconds (5 minutes by default) the screen wakes and /viz opens full-screen. The job finishes, and kiosk.offDelaySeconds later the browser closes and the monitor goes back to sleep. Touch the machine at any point and it hands the screen straight back — kiosk closed, no sleep scheduled — and stays out of the way until the session goes idle again.

What counts as a job is kiosk.activities, and the default is the work that takes minutes: a render, a graph on the card, weights loading. LLM inference is deliberately not in it — on a desktop that would wake the monitor for every chat turn — so a box whose LLM work does run for many minutes (an agentic coding harness) wants inference added. Without it the display comes up for loading-model and then leaves as the actual work starts, which is the wrong half of the job to watch.

The off-delay is a minute rather than the few seconds it takes to read the last frame, because "idle" is sampled and real work is intermittent at that resolution: a coding run alternates decode with tool calls the GPU sits out, and a multi-step graph goes quiet between steps. The delay is what bridges those gaps, so it wants to be longer than the longest ordinary pause in the work.

A locked session is what makes this feature work or not. A Wayland session lock is exclusive — nothing draws over it — so a kiosk window opened under one is invisible and all you get is a monitor waking up to show a lock screen. The answer is therefore not to draw over the lock but to keep it from engaging: while a job runs the daemon holds systemd-inhibit --what=idle (kiosk.preventLock, on by default), which puts idle in logind's BlockInhibited — the property an idle daemon reads before firing its lock and screen-off timers. It is released the moment the work stops, and on shutdown, since an inhibitor that outlived its daemon would be a box that quietly stopped locking. It is also the one inhibitor that does not suppress the compositor's own idle notifications, so the daemon can still tell when a human comes back; a raw Wayland idle-inhibitor would blind it.

There is no default for "is the session already locked" (kiosk.lockedCommand) and that is deliberate — see the option.

A locked screen is let into, not drawn over. Nothing can be drawn over a Wayland session lock; that exclusivity is the whole point of the protocol. The way in is therefore not a prettier lock screen but an unlock: with kiosk.unlockCommand set, a session that is locked while nobody is there gets opened for the duration of the job and re-locked (relockCommand) the instant either the work stops or a human touches the machine — the screen is always handed back exactly as it was found, because nobody authenticated to open it. Without an unlock command the daemon leaves a locked screen dark, on the grounds that waking a monitor to show a password prompt is worse than leaving it off.

That requires a locker that can be opened by something other than a typed password. DankMaterialShell's has dms ipc call lock unlock; hyprlock does not (PAM only, no D-Bus unlock), so a host using it gets the dark-screen behaviour above.

services.genai-server.kiosk = {
  lockedCommand = ''test "$(dms ipc call lock isLocked)" = true'';
  unlockCommand = "dms ipc call lock unlock";
  relockCommand = "dms ipc call lock lock";
};

A locked session also short-circuits idleSeconds. The lock means the seat already sat untouched for the session's own idle timeout, and that verdict survives this daemon restarting — an idle notification always counts from zero, so measuring presence only with the daemon's own timer made every rebuild ignore an empty, locked seat for another full period. The shortcut lapses the moment the idle watcher reports input: somebody typing at that lock screen is present, and their session is not this daemon's to open.

(Two other routes were built and measured before this one, and neither shipped: a GTK4 + WebKitGTK layer-shell client dies inside swaylock-plugin's nested compositor for want of xdg_wm_base, and windowtolayer in front of chromium rendered cleanly 1 run in 3 — the rest lost the buffer-size race with the acked configure and left swaylock's default white.)

This is the one piece of the stack that runs as a user service rather than a system one, because waking a monitor and knowing whether somebody is present are questions only the compositor can answer. That is also its one precondition: it needs a graphical session to run in, so a rebooted box that nobody has logged into shows nothing. The console at that point belongs to the display manager's greeter — a separate user, its own compositor, its own idle timer — and it is that timer, not this daemon's offDelaySeconds, blanking the screen. The tell is an empty journal: systemctl --user status genai-kiosk says inactive (dead) because the unit never started. unlockCommand is not the way in either, and deliberately: a greeter is a login prompt rather than a lock, and letting yourself into a lock and putting it back is not the same act as authenticating as somebody. Wanting the view unattended from boot means a console autologin. Presence comes from swayidle against ext-idle-notify-v1 (niri, sway, Hyprland and KDE all implement it); the screen is driven by whichever compositor CLI answers a probe — niri, swaymsg, hyprctl, wlopm, xset. Both are overridable (kiosk.idleCommand, kiosk.wakeCommand, kiosk.sleepCommand) for anything else.

kiosk.activities picks what counts as work. The default is the work that takes minutes — generating-media, gpu-busy, loading-model — and deliberately excludes inference, which is true of every chat turn on the box; a monitor that lights up whenever somebody asks a question is a nuisance rather than a display.

Metrics

Prometheus metrics (metrics.enable, default on) come in two layers:

  • :8897/metrics — aggregate gauges from the dashboard's samplers: GPU utilization/VRAM/temperature/power, CPU, RAM, per-service up/down probes, per-model ready/enabled/size.
  • :8080/upstream/<model>/metrics — llama-server's own per-model metrics (tokens/s, KV usage, queue depth) for whichever models are loaded.

Host-side scrape config:

services.prometheus.scrapeConfigs = [
  { job_name = "genai";       static_configs = [{ targets = [ "localhost:8897" ]; }]; }
  { job_name = "genai-qwen";  metrics_path = "/upstream/qwen/metrics";
    static_configs = [{ targets = [ "localhost:8080" ]; }]; }
];

Retired 2026-07: genai-gpu-watchdog (and its watchdog.* options) existed solely to restart whisper-server out of a whisper.cpp CUDA-context wedge — GPU utilization pegged at ~100% with the memory controller idle, forever. It gated on a whisper process holding the card, so replacing whisper.cpp with the asr llama-swap model removed both the failure mode and the watchdog's ability to fire.

Two ways models reach the web, both backed by the local SearXNG (no API keys):

  1. Open WebUI search toggle — preconfigured via env; flip the web-search toggle in any chat and results are injected into context. Works with every model.

  2. Agentic tool calling — fully declarative: the open-webui-declarative-config service enforces the tool-server registration and, for every chat model, attaches the server:web-search tool and sets native function calling (via the admin API; runs after each activation, preserves other UI customizations). Models decide on their own when to search. Any other agent framework can use the same endpoint:

    curl -X POST http://<host>:8891/search -H 'Content-Type: application/json' \
         -d '{"query": "nixos flakes tutorial", "max_results": 5}'
    

Storage

This is a shared service flake: nothing lives in home folders and no username is hardcoded (see CLAUDE.md for the conventions). All model weights live in one store so no file is ever downloaded twice:

/var/lib/genai-models/
  llm/                LLM GGUFs (llama.cpp HF-cache layout; genai-prefetch)
  diffusion_models/   Z-Image, FLUX.1 GGUFs + fp8, FLUX.2, Wan
  checkpoints/        Pony / epiCRealism / Juggernaut, LTX, ACE-Step
  text_encoders/      clip_l, t5xxl, umt5, Qwen/Mistral encoders
  vae/                Z-Image / FLUX.1 / FLUX.2 / SDXL-fix / Wan VAEs
  loras/              shared LoRA store + {trigger,base} sidecars
  pulid/  insightface/  identity-adapter weights
  hf-cache/           shared HF_HOME (training bases, EVA-CLIP, ...)
/var/lib/comfyui, /var/lib/ai-toolkit, /var/lib/lora-jobs   apps + jobs
/var/lib/genai-media/     everything GENERATED (mediaStore.dir)
  <time>-<hex>.png/.mp4     renders, clips, masks — with a .json record
  sources/                  input images kept so an edit can be repeated

The same rule as the model store, applied to the other direction: one tree, written once. The portal writes studio renders there, media-tools writes every tool artifact there, both serve the same files (:8897/studio/ images/<name> and :8894/files/<name> are two doors onto one file), and the gallery lists it. It is deliberately not a service StateDirectory: those live under /var/lib/private at 0700 root, where the other service cannot follow — which is exactly how the box ended up writing every tool-made picture twice, once in each service's private state, with the gallery able to see only one of them. Same permissions as the model store, and for the same reason: 2775 root:genai, so a member of genai can delete their own media.

stable-diffusion.cpp, ComfyUI (via extra-model-paths), and the trainers all read the same files — one clip_l, one FLUX VAE, one Pony checkpoint for both inference and training, and a LoRA that lands once appears in chat, the API, and ComfyUI's loader nodes simultaneously. The catalog is declarative: services.genai-server.llmModels and .mediaModels (add a model = add an option entry, and the units fetch it — a declared model is a downloaded model. The two bulk sets, comfy and h3, are the exception named in comfyui.optInModelSets: they wait for a host to list them in comfyui.modelSets, which is still a declaration, not a command somebody runs. genai-fetch-media <set> is what the units call and what you reach for when debugging one — never the way a set gets enabled). Dirs are root:genai 2775 — members of the genai group manage models and training jobs without sudo (join it the standard NixOS way: users.users.<name>.extraGroups = [ "genai" ]).

A second copy elsewhere (remoteStore)

Off by default. Points the store at an archive on another machine — the same relative layout, one tree, reachable two ways:

services.genai-server = {
  remoteStore = {
    enable = true;
    url  = "http://10.0.0.1:8898/genai-models";  # restore: tried before the internet
    path = "/mnt/genai-archive";                 # cold tier: something to symlink into
  };
  # NFS carries numeric ids. Pin these, and match the export's anonuid/anongid.
  serviceUser = { uid = 890; gid = 890; };
};

Restore (url). Every download tries the archive first and falls through to the catalog URL on any miss, so an incomplete or offline archive costs one HEAD request and nothing else. LLM blobs are still verified against Hugging Face's x-linked-etag sha256 whichever source they came from — the mirror can make a fetch faster, never wronger.

Be honest about what this buys. genai-prefetch measures ~104MB/s against Hugging Face with 16 connections — so unless the LAN link comfortably beats that, restoring from the archive is not obviously faster than downloading again, and the win is availability, not speed: repos get re-quantized and deleted, Civitai versions go early-access, and a local copy is the only one that still answers the same bytes next year.

Measure rather than assume, in both directions. A full-duplex wired link is the easy case; a good 5/6GHz link can land in the same range as the CDN, and a mediocre one is nowhere near. genai-store-sync prints the effective rate of every pass that moves anything, so the first backfill answers this for your link without anyone timing it by hand.

Cold tier (path + tier = "cold"). Per catalog entry, weights stay on the archive and a symlink stands in for them at the canonical store path. Every reader opens a path and none of them inspects where it lands, so ComfyUI, image-server and llama.cpp needed no changes at all. Cold LLMs are served with --no-mmap automatically — demand-paging a GGUF across a network filesystem turns one sequential read into tens of thousands of random faults, and llama.cpp touches every weight during a forward pass anyway, so there is no laziness to win.

It trades disk for load time, every load. llama-swap evicts and reloads on every model switch, so a cold chat model pays the network read each time it comes back — this suits an occasionally-used specialist, not anything in a rotation.

On a shared medium — WiFi, or a link with other traffic on it — budget for the variance rather than the average. A cold model's load time is however long its weights take to cross the link right now, so the same model can load in two minutes at 3am and six at 8pm with nothing having changed. That is survivable for something reached a few times a week and quietly infuriating for anything else.

Keeping it current. A genai-store-sync timer backfills declared weights that exist locally and not on the archive, and carries out a tier change in whichever direction the catalog now says (freeze up, thaw down). A freeze copies, verifies the copy, and only then deletes the local file. Run it by hand to watch, or --dry-run to see what would move:

genai-store-sync --dry-run     # what would move, and which way
genai-store-sync               # do it (the timer runs this)
genai-prune --archive          # extend orphan deletion to the archive copies

The archive must be writable, not a read-only export: llama-server writes multimodal projector sidecars next to the weights it loads.

Guards, because the failure modes here are quiet ones. A tier = "cold" entry with no remoteStore.path fails eval rather than becoming a model that is declared and never fetched. The fetchers skip a cold entry whose archive is unreachable instead of downloading it onto the disk it was moved off. genai-doctor reports an unmounted archive, a read-only one, and any dangling cold symlink — the failure that looks like success, since ls shows the model and only open() disagrees — and the portal's /api/health fails on all three. genai-prune counts cold bytes on the archive's ledger rather than claiming to reclaim local disk it never had.

The other end is a HomeFree app (homefree-genai's genai-archive): one directory on the router's pool, exported read-write over NFS and served read-only over HTTP on the LAN address.

Running on different hardware (hardware.profile)

Profiles cover how much VRAM. Running on a different GPU vendor, architecture, or a multi-GPU box is a larger unsolved problem — see HARDWARE.md for the design and an honest account of what is not built.

The shipped tuning is not advice, it is measurements of one box (32GB card, 128GB RAM). hardware.profile applies a named set of serve.* overrides:

services.genai-server.hardware = { vramGB = 24; profile = "vram24"; };

vram32 is measured and selecting it changes nothing. vram24 and vram16 are derived from KV-cache arithmetic and have never been run — they are a starting point for someone with that card, not a promise, and genai-doctor warns whenever a derived profile is active. If you run one, measure it and send the numbers back.

Checking the box against its config (genai-doctor)

genai-doctor        # reports; changes nothing; exit 1 if something failed

hardware.vramGB and hardware.ramGB are declared rather than detected, because Nix evaluation is pure. genai-doctor is what notices when a declaration is wrong — and it treats the two directions differently: claiming more VRAM than the card has is a failure (models are kept that cannot fit), claiming less is only a warning.

It also checks store permissions and free space, portal health, and reports orphans and stale quants via genai-prune.

Measuring what the stack claims (genai-eval)

Runs suites against the live services and writes a report:

genai-eval                 # everything
genai-eval --list          # what suites exist
genai-eval memory rag      # named suites
genai-eval --warm-only     # skip anything that would load a model

Cases are multi-step HTTP calls with deterministic checks, so the same harness grades a chat model and a service. The shipped suites are regression tests for claims this stack makes: that memory supersedes corrected facts, that hybrid retrieval catches both paraphrase and exact identifiers, that the code sandbox cannot see the host filesystem.

Set evals.compareModels = [ "coder-pro" "qwen-dense" ] to run a champion/challenger coding suite: each model writes a program, the sandbox runs it, and the output is checked — so the ranking is what the code did, not what another model thought of it. Pair the winner with a selector to promote it without touching any client. Results render at :8897/evals.

Thinking models and token budgets. A reasoning model spends the budget thinking before it answers, so a fixed max_tokens shared across a comparison measures whether each model finished thinking in time, not what it can do. The shipped suite allows 4000 and checks finish_reason plus non-empty content at the step that produces the code, so a truncated response fails where the message explains itself. Without that check the first three-way run scored a thinking model 0/3 on problems it never attempted — it had burned all 1200 tokens reasoning and returned empty content, which then substituted into the next step as an empty program. If you add your own comparison cases, budget for the reasoning.

A comparison is also treated as failed when no model clears the suite. Numbers alone are not evidence the harness worked; "nothing could pass this" means either every model is bad or the suite is broken, and both deserve a look before anyone quotes a ranking.

The speech suite is a round trip: it synthesizes a sentence with the CPU voice and transcribes it back, covering TTS, STT and the ASR-preamble strip in one case with no audio fixture in the repo. asr is resident, so it evicts nothing and is safe to run on a busy box.

speech-under-load is the same round trip, but only runs when a chat model is already resident. That is not redundancy: asr was once configured with a KV cache too large to fit beside a warm chat model, so transcription failed on a working box and passed on an idle one — and the plain speech suite could not see it. A test that only runs in the easy condition is not covering the hard one.

A suite whose precondition is missing skips rather than fails — memory-reconcile needs a resident chat model, since reconciliation uses one that is already warm and never loads its own. Exit status is non-zero on real failures, so it works from a timer. Reports go to /var/lib/genai-eval; add your own with evals.extraSuites.

Pruning the store (genai-prune)

The store is meant to be exactly what llmModels and mediaModels declare. genai-prune reports the difference; --delete acts on it:

genai-prune                          # report only (default)
genai-prune --stale-quants           # + old quants inside declared repos
genai-prune --delete --stale-quants  # actually remove them

Undeclared is not the same as garbage. Locally trained LoRAs share a directory with catalog ones and can't be re-downloaded, so the tool only deletes what identifies as catalog-managed (source: "manifest" in the sidecar) and is no longer in the catalog. A LoRA with no sidecar, or one from lora-train/lora-add, is always kept and merely listed.

Two classes are opt-in because they're riskier:

Flag What it removes Why it's gated
--stale-quants a .gguf in a declared repo that isn't the declared tag — e.g. a Q4_K_M left behind when the tag moved to UD-Q4_K_XL latest tags can't be matched locally, so they're never classified
--media-orphans non-LoRA files absent from mediaModels this is also where tooling-fetched weights live (PuLID's EVA-CLIP, Florence-2); deleting them costs a re-download and breaks the feature until it happens

Note that removing a model from the portal alone doesn't stick: genai-models-prefetch re-downloads the full catalog at every boot. To reclaim space permanently, drop the entry from llmModels and prune.

Adding a LoRA (or any other model)

mediaModels entries are shaped like the Civitai model card you are copying from — URL, title, type, base model, the specific model instance to run it on, usage tips, trigger words. Add the entry in your host config; host entries are added to the shipped catalog, so nothing else needs restating:

services.genai-server.mediaModels."ghibli-style.safetensors" = {
  type     = "lora";                                    # Civitai's "Type"
  title    = "Studio Ghibli Style";
  page     = "https://civitai.com/models/433138";       # for humans
  url      = "https://civitai.com/api/download/models/482825";
  base     = "pony";        # its "Base Model": Pony → pony, SDXL 1.0 /
                            # Illustrious → sdxl, Flux.1 D → flux,
                            # Z-Image → zimage, Krea 2 → krea
  engine   = "epicrealism"; # the checkpoint the page showcases it on
  triggers = [ "ghibli style" ];
  strength = 0.8;           # usage tips from the page…
  clipSkip = 2;
  notes    = "Keep the trigger early in the prompt.";
};

type picks the store category (lora/lycorisloras/, checkpointcheckpoints/, vaevae/, …), so dir is only needed for oddities. Rebuild, then genai-fetch-media (or just restart image-server, which runs it) downloads the file and writes its {"trigger","base",...} sidecar. From then on the LoRA is listed in chat by name, routed to a compatible engine, and its trigger word, weight, clip skip and step count are applied automatically — lora-list shows all of it. Metadata-only edits apply on the next fetch without re-downloading; enable = false drops a shipped entry.

Not everything has a URL. Locally trained LoRAs are deployed by lora-train install, and a file you already have on disk by lora-add <file.safetensors> "<trigger words>" [base]; both write the same sidecar. Removing a catalog entry never deletes files from the store — those two paths write into the same directory.

A checkpoint works the same way (type = "checkpoint"), but the image server only runs the engines it knows: point one of the services.genai-server.imageServer.models.* options at the new file.

Testing the flake

nix flake check                        # evaluates every check
nix build .#checks.x86_64-linux.vm     # the stack's wiring, in a VM
nix build .#checks.x86_64-linux.wyoming # the same with Wyoming enabled

These boot a real NixOS VM and assert what actually broke during development: units start and answer, every hosted page renders with the shared nav, every enabled service has a portal card and a health path (or an explicit non-HTTP protocol), MCP bridges its tools with forget confirm-gated, the /svc allowlist 404s an unknown name, an uncataloged /api/pull names the option to add it to, genai-prune reports rather than deletes, and disabling a service really removes its unit, card and proxy entry.

They deliberately do not test the models. There is no GPU in a VM, so llama-swap and everything needing weights is switched off. That is a smaller claim than "the stack works" — and it is the claim worth automating, because every bug this repo has shipped to a live machine was in the wiring, not the models. For the models themselves, see genai-eval, which runs against a real box.

One check is worth running before trusting the others: nix eval .#checks.x86_64-linux.vm.drvPath. A check that fails to evaluate is not a failing test, it is an absent one — and it reports as neither.

Usage

# flake.nix inputs
genai-server.url = "git+https://git.homefree.host/homefree/genai-server";

# host modules
inputs.genai-server.nixosModules.default

# host configuration
services.genai-server.enable = true;
# Grant yourself write access to the model store / LoRAs / training jobs
# (standard NixOS group membership; the flake never names users).
users.users.youruser.extraGroups = [ "genai" ];
# Optional: Civitai API token file (civitai.com/user/account → API Keys).
# Needed by any mediaModels entry with `civitaiToken = true` — Civitai gates
# most downloads behind an account. Without it those entries are skipped
# (with a warning) and everything else works.
services.genai-server.civitaiTokenFile = "/run/secrets/civitai-api-token";

Requires: NVIDIA drivers configured on the host, unfree/CUDA nixpkgs allowed. Note the CUDA llama.cpp build compiles from source (no cache).

Bringing this up on your hardware

Read this before assuming it will fit. This flake is a very good configuration of one machine — a 32GB RTX 5090 with 128GB of RAM — and being honest about that is more useful than a list of requirements that implies otherwise.

What is actually portable today

The flake evaluates on any hardware, including machines with no NVIDIA GPU. That sounds like a low bar; it was not met until 2026-07-31, and a flake you cannot evaluate is one you cannot try. The VM test runs on a driver-less node specifically so this keeps working.

Declare what you have:

services.genai-server.hardware = {
  vramGB = 24;          # what the card actually has
  ramGB  = 64;
  profile = "vram24";   # per-model serving overrides
};

genai-hwscan will write that block for you — or rather, print it:

$ genai-hwscan
# GPU: NVIDIA GeForce RTX 5090 (31GB, CC 12.0)
# RAM: 123GB   /dev/kvm: true   free on /var/lib/genai-models: 457GB

services.genai-server = {
  enable = true;
  hardware = { vramGB = 31; ramGB = 123; profile = "vram32"; };
};

It emits config and never writes it — a tool that edits a NixOS configuration is one that eventually loses somebody's edits. It also reports what it could not determine, and leaves those values commented out rather than guessing: a card with no driver loaded looks exactly like no card, and guessing optimistically keeps models enabled that cannot fit, which fails at load after a long download. --json for scripting.

Then run genai-doctor, which compares that declaration against reality and complains. The two share detection and run in opposite directions: hwscan proposes a declaration for a machine that has none, the doctor checks one that exists. Over-declaring VRAM is a hard failure — a model that does not fit fails at load, badly, after a long download — while under-declaring is only a warning.

serve.minVramGB drops any generated model whose floor exceeds vramGB, so a smaller card gets a smaller fleet rather than a fleet that OOMs.

What is measured and what is arithmetic

Only vram32 is measured. vram24 and vram16 are derived by arithmetic, labelled as such in profiles.nix, and genai-doctor says so out loud when a derived profile is active. Treat their numbers as a starting point you should verify, not as tested configuration. If you measure better ones, they are worth contributing back — genai-eval and the VM test are what make a contributed measurement checkable rather than a claim.

Turning things off

Everything GPU-hungry has an enable:

services.genai-server = {
  comfyui.enable = false;      # torch stack; the big one
  imageServer.enable = false;
  codeSandbox.enable = false;  # needs /dev/kvm
  magenticUi.enable = false;
};

The code sandbox is the model to copy: with no KVM it answers 503 with a reason, never "run it on the host instead". Unavailable-with-an-explanation beats a service that starts and fails at first request.

What will not work yet

  • Non-NVIDIA GPUs. llama-cpp is built with cudaSupport = true unconditionally. On AMD or Intel that build is useless; there is no ROCm, Vulkan or CPU variant selected by config yet.
  • Multiple GPUs. No tensor split, no pinning a model to a card. llama-swap's peers and llama.cpp's RPC workers are the mechanisms and neither is wired up.
  • Small VRAM. The shipped model fleet assumes ~32GB. Below roughly 16GB you will be picking models by hand.

The full analysis — five axes of hardware variation, the capability table each component needs, and a proposed genai-hwscan that reads a machine and emits a config — is in HARDWARE.md. It is a design document, not an implementation, and it says so. The honest cost there is in the measuring, not the coding, and this repo has exactly one machine.

Build times, and serving a cache

The CUDA llama.cpp build compiles from source and takes 2040 minutes on a first build or a nixpkgs bump. Budget for that before your first nixos-rebuild switch.

A machine that has paid that cost can serve it to the others. This box already builds those derivations in order to run them, so serving costs a process and no extra disk:

services.genai-server.binaryCache = {
  enable = true;
  signKeyPath = "/run/secrets/cache-priv-key.pem";
  openFirewall = true;    # LAN or tailnet
};

Generate the key pair first, keeping the private half out of the Nix store:

nix-store --generate-binary-cache-key myhost-1 \
  /var/lib/secrets/cache-priv-key.pem cache-pub-key.pem

Then on each client:

nix.settings = {
  substituters = [ "http://myhost.lan:5000" ];
  trusted-public-keys = [ "myhost-1:<contents of cache-pub-key.pem>" ];
};

Sign it. With no signKeyPath the store is served unsigned, every client refuses those paths unless it turns signature checking off, and the result looks like a working cache until somebody tries to use it — so the flake warns at build time rather than letting you find out later.

It uses harmonia rather than nix-serve: nix-serve is a Perl CGI whose upstream is dormant, and harmonia is the maintained replacement with the same contract.