Levi DeHaan

AI Cluster — Local AI Appliance

Go runtime and control plane for a local AI appliance: model loading, compatible inference APIs, scheduling, durable chat, agents, projects, and a Svelte UI — with Vault for identity/connectors and Helix for encrypted knowledge.

Status: in-progress · 2026-09-13

Overview

AI Cluster is the runtime and control plane for a local AI appliance. It loads and swaps llama.cpp servers, then layers an embedded Svelte UI, OpenAI/Anthropic-compatible APIs, resource scheduling, durable conversations, tools, agents, project workspaces, media workflows, and platform maintenance. Vault authenticates people and ingests external data. Helix stores encrypted searchable knowledge and uses AI Cluster for embeddings and generation. This is a different system from Almagest, which is a native C++ llama.cpp engine.

Technologies

Go, Svelte 5 / TypeScript / Vite, llama.cpp, MCP (stdio + HTTP), SQLite, Vault (TypeScript / PostgreSQL), Helix (Rust / RocksDB / Tantivy / Qdrant), Docker Compose, OpenAI- and Anthropic-compatible HTTP

Core
Go + Svelte 5
Inference
llama.cpp
Knowledge
Helix
Identity
Vault
Human UI
:8080
Machine API
:8081

AI Cluster — Local AI Appliance

AI Cluster is the runtime and control plane for a local AI appliance. It serves the main browser app, compatible inference APIs, model process management, resource scheduling, durable conversations, tools, agents, project workspaces, media workflows, and governed platform maintenance.

The Go proxy still loads and swaps llama.cpp servers. Around that, the product now includes the main UI, compatible inference APIs, scheduling, durable conversations, tools, agents, projects, and platform maintenance.

This is not Almagest. Almagest is a separate native C++ llama.cpp engine. AI Cluster is a Go server with an embedded Svelte 5 UI.

Companion systems

SystemRole
AI Cluster (Core)Model loading, shared inference capacity, request admission, main UI, durable conversations/runs/projects
VaultUser authentication, connector authorization, ingestion workers, event rules
HelixEncrypted searchable documents, keyword/vector/graph retrieval, answers with citations. Gets inference from AI Cluster.

Local-first by default. Remote peers, web search, connectors, downloads, backups, and optional audit anchoring can still leave the host when configured.

What the UI covers

The main app is served at /ui/ (hash routing). Tabs include Playground, Omni (multimodal + voice), Avatar, Models, Backends, Files, Helix Search, Projects, Agents, Studio (ComfyUI), Management, Downloads, Activity, System, Logs, Settings, Remote, API, and MCP.

Helix Search is its own workspace: browse documents, BM25 keyword search, and Ask AI with citations — not just a chat prompt.

Inference surface

Core keeps a stable API while mapping a model choice onto the process that can run it:

  • Text / assistants: /v1/chat/completions, completions, Responses, agentic APIs
  • Anthropic: /v1/messages
  • Embeddings and rerank
  • Audio (transcription / TTS) and images (OpenAI-style and Automatic1111/Forge)
  • Studio/ComfyUI jobs
  • Runtime control under /api/*

A shared API does not mean every backend supports tools, vision, audio, or structured output. Compatibility follows the selected engine.

Clients can wait while a model loads, or opt into async admission (X-AiRouter-Async: 1) with a 202, request ID, and queue/stream URLs. Completion recovery can retrieve a retained result; it does not resume token-by-token after a process restart.

Scheduling

There is no single queue. Separate layers decide model-swap order, memory admission, durable global jobs, agent steps, Vault sync, and Helix indexing.

Observed production combination: model mode auto, global policy setup-wspt, intake enforce. Auto mode considers priority, prompt/output cost, context fit, startup cost, and observed throughput. Memory admission compares footprints against real host pressure, including unified-memory hosts. This is not transparent model sharding or automatic cross-host failover for every request.

Supported engines in the registry include llama.cpp (and MTP), vLLM variants, Ollama, ComfyUI, x-flux, and generic HTTP services. Registry support is not the same as “installed and qualified on this host.”

Agents and projects

Agent profiles are reusable config (model, tools, sandbox, plan). Runs are durable executions with queued/running/paused/completed/failed/stopped states. Modes: adaptive loops or structured plans (steps, branches, gates, parallel work).

Projects are server-managed workspaces with files, runs, previews, and Kanban. Project Builder is a gated pipeline (plan, implement, test, browser proof, bundle) — a passing child run or a model saying “passed” is not the same as host/browser acceptance.

Honesty bound

The source tree, a built binary, a running service, and a verified user workflow are four different states. This write-up is an architectural map as of 2026-09-13, not a claim that every feature has production acceptance.