Levi DeHaan

Almagest — Native Inference Engine

C++ llama.cpp engine with durable native workflows, layered memory, MCP tools, and a live dashboard. Runs Qwen3.8-27B across two RTX 3060s with MTP speculation.

Status: in-progress · 2026-09-10

Overview

Almagest is a native C++ inference engine built on llama.cpp. The binary owns the model, KV cache, durable workflows, and working memory — not a Python harness wrapping a server. A live dashboard watches the same coordinator the TUI uses: health, loaded model, memory in motion, sequence occupancy, tool calls, and native runs. Application tools come from host-configured MCP servers. Compatibility OpenAI chat and session APIs remain, but native work goes through /v1/workflows.

Technologies

C++, llama.cpp / GGML, CUDA, Qwen3.8-27B, MTP speculation, SQLite FTS5, Model Context Protocol, Native TUI, Embedded web dashboard, OpenAI-compatible HTTP

Decode (MTP, 2× RTX 3060)
27.62 tok/s
Vs previous profile
7.55×
Context / sequence
16,384
Shared KV pool
32,768
Sequence slots
6
Model (Q4_K_S)
14.7 GiB

Almagest — Native Inference Engine

Almagest is a C++ inference engine on llama.cpp. The current line (0.5.4) puts a durable workflow coordinator inside the binary: parent/child contexts, RAM offload, versioned working memory, MCP tool calls, approvals, and recovery. The TUI and the web dashboard observe that same coordinator.

Compatibility /v1/chat/completions and session APIs still exist; native work uses /v1/workflows.

What it does

  • One loaded model, many sequences. Official Qwen3.8-27B, quantized locally from verified BF16 to Q4_K_S, with the matching MTP head embedded. Layers split across two NVIDIA RTX 3060 GPUs. Shared Q8_0 K/V pool, 16,384 tokens per sequence, six sequence slots.
  • Native workflows. Nested children, frozen parent documents, RAM checkpoints when capacity is tight, and canonical replay after edits. Child KV tensors are never spliced into a different parent history.
  • Layered memory. Durable task overview, organized working notes, and immutable source evidence in SQLite (FTS5). Search mixes lexical relevance with source age. Automatic consolidation when history outgrows the active window.
  • MCP tools. Seven internal memory/control tools plus host-configured MCP providers (stdio and HTTP). The coding TUI loads file edit, patch, process, and large-PDF tools from a persistent external provider.
  • Live dashboard. Embedded web UI at the engine port (default 8088): Overview, Health, Models, Memory & KV, Tools & connections, Runtime & runs. Polls a bounded snapshot; no coordinator commands on the observation path.

Measured inference

On this two-GPU box, controlled short generations (192 tokens, temperature 0, tools off) moved from 3.66 to 27.62 output tokens/sec with official Q4_K_S + one-token MTP — about 7.55× the previous mixed CPU/GPU profile. Long prompts still pay prefill. Live dashboard rates during multi-sequence native work are lower, because several contexts share the same decode batch.

Clients

  • Web dashboard compiled into the executable — no Node build, CDN, or extra server.
  • almagest-code TUI: native engine by default, resumable sessions, approvals, /workflow graph.
  • HTTP: health, model, native workflows, plus OpenAI-style chat as a focused adapter.

Local development service: binds localhost by default, no multi-user auth.