Almagest — Native Inference Engine
C++ llama.cpp engine with durable native workflows, layered memory, MCP tools, and a live dashboard. Runs Qwen3.8-27B across two RTX 3060s with MTP speculation.
Status: in-progress · 2026-09-10
Overview
Almagest is a native C++ inference engine built on llama.cpp. The binary owns the model, KV cache, durable workflows, and working memory — not a Python harness wrapping a server. A live dashboard watches the same coordinator the TUI uses: health, loaded model, memory in motion, sequence occupancy, tool calls, and native runs. Application tools come from host-configured MCP servers. Compatibility OpenAI chat and session APIs remain, but native work goes through /v1/workflows.
Technologies
C++, llama.cpp / GGML, CUDA, Qwen3.8-27B, MTP speculation, SQLite FTS5, Model Context Protocol, Native TUI, Embedded web dashboard, OpenAI-compatible HTTP
- Decode (MTP, 2× RTX 3060)
- 27.62 tok/s
- Vs previous profile
- 7.55×
- Context / sequence
- 16,384
- Shared KV pool
- 32,768
- Sequence slots
- 6
- Model (Q4_K_S)
- 14.7 GiB
Almagest — Native Inference Engine
Almagest is a C++ inference engine on llama.cpp. The current line (0.5.4) puts a durable workflow coordinator inside the binary: parent/child contexts, RAM offload, versioned working memory, MCP tool calls, approvals, and recovery. The TUI and the web dashboard observe that same coordinator.
Compatibility /v1/chat/completions and session APIs still exist; native work uses /v1/workflows.
What it does
- One loaded model, many sequences. Official Qwen3.8-27B, quantized locally from verified BF16 to Q4_K_S, with the matching MTP head embedded. Layers split across two NVIDIA RTX 3060 GPUs. Shared Q8_0 K/V pool, 16,384 tokens per sequence, six sequence slots.
- Native workflows. Nested children, frozen parent documents, RAM checkpoints when capacity is tight, and canonical replay after edits. Child KV tensors are never spliced into a different parent history.
- Layered memory. Durable task overview, organized working notes, and immutable source evidence in SQLite (FTS5). Search mixes lexical relevance with source age. Automatic consolidation when history outgrows the active window.
- MCP tools. Seven internal memory/control tools plus host-configured MCP providers (stdio and HTTP). The coding TUI loads file edit, patch, process, and large-PDF tools from a persistent external provider.
- Live dashboard. Embedded web UI at the engine port (default
8088): Overview, Health, Models, Memory & KV, Tools & connections, Runtime & runs. Polls a bounded snapshot; no coordinator commands on the observation path.
Measured inference
On this two-GPU box, controlled short generations (192 tokens, temperature 0, tools off) moved from 3.66 to 27.62 output tokens/sec with official Q4_K_S + one-token MTP — about 7.55× the previous mixed CPU/GPU profile. Long prompts still pay prefill. Live dashboard rates during multi-sequence native work are lower, because several contexts share the same decode batch.
Clients
- Web dashboard compiled into the executable — no Node build, CDN, or extra server.
almagest-codeTUI: native engine by default, resumable sessions, approvals,/workflowgraph.- HTTP: health, model, native workflows, plus OpenAI-style chat as a focused adapter.
Local development service: binds localhost by default, no multi-user auth.