Levi DeHaan

The Mobile Sentient Data Center: Mesh Intelligence After the Building Dissolves

A sequel to The Sentient Data Center: AI Cluster on a Strix Halo with 128 GB unified memory, home stacks, solar-powered mobile nodes, mesh and internet peers, compute contracts, and a video-shaped transport.

By Levi DeHaan ·

The Mobile Sentient Data Center: Mesh Intelligence After the Building Dissolves

In The Sentient Data Center I argued that operations would stop being a human firefight and become a hierarchical immune system: tiny models on every server, supervisors in every rack, a central command for the hard cases, and a last-resort out-of-band path when the network itself was the casualty.

That blueprint assumed a building. Raised floors. Cages. Dedicated GPU servers. Cellular OOBM as the SOS channel.

I still believe the hierarchy. I no longer believe it has to live in a building.

I am running AI Cluster as a local AI appliance on an AMD Strix Halo with 128 GB of unified memory. One box holds the control plane, the models, the conversations, the agents, and the knowledge path. It is not a toy. It is not a rack of H100s either. It is the missing middle: a supervisor-class node that draws household power, sits on a desk or in a closet, and can — with a battery and a couple of solar panels — stay up when the grid does not.

Once that node exists, the Sentient Data Center stops being a place you visit and becomes a fabric you carry. Home stacks. Neighborhood Wi‑Fi meshes. Internet-distributed jobs. Unique AI systems that do not clone each other, but borrow from each other. That is the sequel.

flowchart LR
  subgraph building [Building: original blueprint]
    T1["Tier 1 tiny LLMs on servers"]
    T2["Tier 2 rack supervisors"]
    T3["Tier 3 central GPU cluster"]
    T1 --> T2 --> T3
  end
  subgraph mesh [Mesh: after the building]
    P["Phones, laptops, Jetsons"]
    H["Home Strix Halo / AI Cluster"]
    N["Neighborhood mesh"]
    I["Internet peers"]
    P --> H
    H --> N
    N --> I
  end
  building -.-> mesh

What the original blueprint got right

The original post had three principles that still hold.

Hierarchy and specialization. A 1B log sensor should not do data-center-wide RCA. A 70B+ command model should not sit on every box burning watts for a disk-full alert. Specialization is how you stay fast.

Resilience by design. Try the center first, then the local supervisor, then isolation, then OOBM. The state machine from Section 3 of the original still maps: Normal, Unreachable, Fallback, Total Isolation.

Security by default. Agents as Non-Human Identities. Ephemeral credentials. Sandboxes. Zero Trust. If a node can act, it can also be captured. The blast radius has to be designed, not hoped.

What changes is the substrate. The original sized Tier 2 as an in-rack GPU server (24–48 GB VRAM, 7B–70B models) and Tier 3 as a building-scale cluster. That was honest TCO for 2025 hardware. It also locked the architecture to real estate, cooling plants, and a NOC.

Unified-memory APUs, better kernels, and a control plane that already accounts for shared RAM/GPU pressure change the unit of deployment from “rack” to “appliance.”

The Strix Halo as a household supervisor

AMD’s Strix Halo class (Ryzen AI Max, Radeon 8060S-class iGPU, up to 128 GB LPDDR5X) is interesting for a boring reason: the GPU and the CPU drink from the same pool. There is no 24 GB VRAM cliff with 128 GB of unused DDR sitting next to it. A large quantized model, KV cache, agent working set, Helix indexes, and the OS can coexist if the scheduler is honest about footprints.

AI Cluster already treats that as a first-class problem. Admission is not “is there a free GPU?” It compares model and job footprints against real host pressure, including unified-memory hosts. Residency is a loaded backend, not a marketing diagram. Auto scheduling weighs priority, prompt/output cost, context fit, startup cost, and observed throughput. Clients can wait for a model to load, or opt into async admission (202, request id, queue/stream URLs) the same way the original Sentient DC used HTTP 202 so Tier 1 would not block on Tier 3.

That is the rack supervisor, collapsed into a silent box.

On this node I run:

Role in the original blueprintWhat runs here
Tier 2 supervisorAI Cluster (Go control plane, embedded Svelte UI, llama.cpp and peer backends)
Local knowledge / RAGHelix (encrypted documents, BM25 + vectors, citations)
Identity and ingestVault (people, connectors, event rules)
Specialist engineAlmagest as a different C++ llama.cpp runtime for durable native workflows — not a clone of the cluster
Edge sensorsLaptops, phones, Jetsons, cameras, log shippers

The important sentence: these are not one model copied N times. They are unique systems with different contracts that talk.

flowchart TB
  Human[People] --> UI["AI Cluster UI"]
  Sensors["Phones, Jetsons, cameras"] --> Core["AI Cluster on Strix Halo"]
  UI --> Core
  Core --> Schedule["Admission and residency"]
  Schedule --> Local["llama.cpp and local backends"]
  Schedule --> Peers["Mesh and internet peers"]
  Core <--> Vault["Vault identity and ingest"]
  Core <--> Helix["Helix knowledge"]
  Helix -->|"embeddings and generation"| Core
  Core <--> Alma["Almagest native engine"]

Remap the three tiers onto the world

Keep the cognitive hierarchy. Move the hardware.

Tier 1 — Edge sensors that move. Tiny models on phones, laptops, Jetson-class boards, NVRs, cars. Continuous local analysis: logs, radio, camera, health telemetry, farm sensors, radio nets. They escalate. They do not wait for a building.

Tier 2 — The home (or van) supervisor. One Strix Halo-class appliance, or a small stack of them. AI Cluster loads models, admits work, keeps conversations and agent runs durable, talks to Helix and Vault. This is the rack supervisor without the rack.

Tier 3 — The mesh command, not the campus command. When the home node is stuck — too big a context, a model it does not host, a job that needs 40 peers — it escalates sideways and upward into whatever is reachable: other homes on a Wi‑Fi repeater mesh, a friend’s node across the internet, a paid peer under contract. There is no single Central Command unless a community chooses to run one.

Failover is the original state machine with new links:

stateDiagram-v2
  [*] --> Normal
  Normal --> LocalSupervisor: Home AI Cluster reachable
  LocalSupervisor --> MeshFallback: Neighbor mesh has spare capacity
  MeshFallback --> InternetPeer: Contracted remote peer
  InternetPeer --> Isolation: No path
  Isolation --> StoreAndForward: Queue locally solar-powered
  StoreAndForward --> LocalSupervisor: Link returns
  Isolation --> VideoPath: Encode job as prioritized media stream
  VideoPath --> MeshFallback: WebRTC mesh accepts the stream

The original OOBM was cellular SOS. The mesh version has two OOBM-shaped tools: store-and-forward on battery, and a transport that looks like video so congested networks treat the bits as something they already know how to carry.

Stacking a home data center

A single Strix Halo is a supervisor. A closet can still be a mini-cage.

Stack them the way the original stacked racks:

  • Node A — always-on inference and UI (AI Cluster). Human listener on the house network. Machine listener for peers.
  • Node B — knowledge and ingest (Helix + Vault workers). Different failure domain. Encrypted at rest.
  • Node C — specialist engines (Almagest, image/video backends, STT like the Jetson Orin service). Different kernels, different memory layouts.
  • Storage — snapshots, restic-style encrypted backups. Copying a live SQLite or RocksDB file is not a backup. The AI Cluster ops notes already say this; a home DC does not get to forget it.
  • Power — UPS first, then a portable power station, then panels. The goal is not net-zero theater. The goal is the node does not die when the street does.

You do not need a CRAC unit. You need airflow, honest admission, and the discipline to not load five 70B models because the UI made it look free. Unified memory makes that mistake easy. The scheduler has to be the adult in the room.

The mobile node

Some of these boxes can run 24/7 off a portable battery and a couple of solar panels. That is not a footnote. It changes the architecture.

A building data center has a street address, a generator contract, and a fence. A solar node has a trajectory. It is in a driveway, a field, a roof, a van, a disaster zone, a farm. Connectivity is intermittent. Bandwidth is whatever the mesh or the phone uplink is today. The node is still a supervisor for whatever sensors are nearby.

That is the shift from Sentient Data Center to mobile shared intelligence:

  • The fabric must tolerate partition. Isolation is a normal state, not an incident.
  • Work must be durable. AI Cluster already treats agent runs as queued/running/paused/completed/failed — a solar brownout should pause, not invent a success.
  • Knowledge must sync when the link exists, not assume a fat pipe. Vault ingestion and Helix indexing already separate “accepted” from “searchable.” Mobile nodes need that split even more.
  • Identity must roam. Vault’s owner mapping and Helix grants cannot assume the node is always on the home VLAN.

The original “human on the loop” still applies. You do not want an unattended van autonomously accepting every remote job. You want a policy: which peers, which models, how many watts, what is allowed to leave the box.

Unique AIs, not identical clones

Distributed computing over the internet is easy to picture as one Docker image, a million copies, a load balancer. That is not this.

The interesting mesh is heterogeneous:

  • AI Cluster — Go appliance. Loads and swaps llama.cpp (and other) backends. OpenAI/Anthropic-shaped APIs. Agents, projects, playground, omni, MCP. Admission against real memory.
  • Almagest — native C++ engine. Durable parent/child workflows, layered SQLite memory, MCP tools, live dashboard. Different KV story, different coordinator.
  • Helix — encrypted search and answers with citations. Gets inference from the cluster; it is not a second chatbot.
  • Vault — people and connectors. Events that can fire owner-bound agent runs. Not a model host.
  • Edge specialists — STT, vision, radio, PLC, medical devices, farm controllers.

They should interact, not merge. A phone’s tiny model asks the home cluster. The cluster asks Almagest for a long native workflow. Helix answers with sources. A neighbor’s node has a model you do not. A contracted peer has spare decode. Each system keeps its contract. You compose them the way the original composed tiers: escalate, do not flatten.

Borrowing compute is then a capability handshake, not an SSH trust-me:

sequenceDiagram
  participant A as Home AI Cluster
  participant C as Contract ledger
  participant B as Peer node
  A->>C: Offer job footprint watts deadline model class
  C-->>A: Quote and escrow terms
  A->>B: WebRTC data channel or video-path stream
  B->>B: Admit against local pressure
  B-->>A: Result artifact plus receipt
  A->>C: Settle or slash on timeout

Crypto contracts belong here as receipts and escrow, not as a personality. You want: who offered watts, who consumed them, what model class ran, whether the result hash matched, whether the deadline was missed. You do not want a mesh that only works if everyone trusts the same Discord. A signed capability plus a timeout is enough to start. Tokens can meter it later.

Neighborhood meshes and internet peers

Two overlays, one scheduler.

Shared Wi‑Fi repeater meshes when density is high. Apartments, campuses, villages, disaster camps, farms with line-of-sight. Latency is LAN-ish. You can share KV-warm models, not just HTTP. AI Cluster already has remote peers; a mesh is “peers that happen to be 30 ms away and might vanish when someone unplugs a repeater.”

Internet-distributed jobs when density is low. Higher RTT. No shared KV. You send a packaged job: model id or adapter, context pack, tool policy, deadline, max watts. The peer admits or refuses. This is the original async 202 pattern at planetary scale.

WebRTC is the practical glue for both. It already punches NATs, already multiplexes media and data, already has congestion control, already runs in browsers and in native stacks. A mesh of AI Cluster nodes does not need a new VPN religion to say hello. Data channels for control and artifacts. Media channels when you want the network to treat the payload as video (next section). STUN/TURN as the OOBM of connectivity.

The original hierarchical failover still applies: prefer the home supervisor, then the mesh, then the internet peer, then queue locally.

A transmission path that looks like video

This is the speculative piece, and I want it labeled as such.

Most networks, from coffee-shop Wi‑Fi to carrier cores, have spent twenty years learning to not drop video. They give it buffers, queues, and sometimes explicit QoS. Bulk “file transfer” and “unknown UDP” get the leftovers. If agent context packs, KV pages, and model shards travel as ordinary downloads, they lose to Netflix in the same household.

So: encode the job as a video stream.

Not a metaphor. An elementary stream whose pixels (or a reserved SEI/side-data channel) carry framed, checksummed payloads: context packs, tool results, small adapters, control messages. To the network it is WebRTC or SRT or HLS. To the nodes it is a reliable message bus with the priority profile of a face.

Properties you actually want:

  • Priority — rides the path operators already protect.
  • NAT traversal — WebRTC.
  • Congestion — existing video CC, not a custom UDP you will never get through a hotel firewall.
  • Cover — looks like a call. Useful in hostile or naive networks. Also a dual-use problem: treat it as a security feature and an abuse vector. Zero Trust still applies. Encrypted payloads. Signed frames. No “priority” without authentication.
  • Degradation — if the “video” is throttled, the codec of the payload can drop non-essential shards the way a video codec drops B-frames. KV extras first. Logs later. Control never.

Call it a research protocol, not a product claim. The Sentient DC proposed cellular OOBM before most shops had it wired. This is the mesh-era equivalent: make the immune system’s messages look like a thing the network is afraid to bufferbloat.

flowchart TD
  Job["Context pack and control"] --> Frame["Frame into media-shaped packets"]
  Frame --> RTC["WebRTC / SRT"]
  RTC --> Net["Networks that already prioritize video"]
  Net --> Peer["Peer AI Cluster"]
  Peer --> Admit["Admit against local watts and RAM"]
  Admit --> Result["Receipt plus artifact"]

What thousands of fast, weird machines enable

Once inference is cheaper, kernels are tighter, and hardware is more purpose-built (APUs, NPUs, tiny accelerators in phones and radios), the mesh is not a curiosity. It is a labor market for cognition. A non-exhaustive list of jobs that get better when the node is near the problem and can still phone a friend:

  • Disaster and field medicine — STT, protocol lookup, radio-to-text, offline Helix of manuals, escalate imaging to a neighbor with a bigger model.
  • Farms and infrastructure — local vision on pests and leaks; home supervisor correlates; mesh shares a disease model without uploading every frame to a hyperscaler.
  • Science on the cheap — nights of batched folding, search, or simulation packed into whoever has solar surplus.
  • Local journalism and records — transcribe meetings, index public docs in Helix, keep the corpus in the town.
  • Accessibility — omni/voice on the appliance, not a round trip to a coast.
  • Industrial and home ops — the original Sentient DC workflow (detect, diagnose, remediate) on a closet that happens to run the house, the shop, or the boat.
  • Education — a classroom cluster that does not die when the district filter does.

The original promised MTTR measured in machine time. The mesh promises coverage: intelligence where the building never was.

Security when the cage is a backpack

The Zero Trust chapter of the original post does not get a holiday.

  • Every node is an NHI. Vault-style identity, short-lived credentials, no “it’s on my LAN.”
  • Sandboxes for agent tools stay mandatory. A borrowed job must not inherit the host’s file roots.
  • Helix encryption at rest still matters when the node can be stolen from a van.
  • Video-path transport is encrypted and authenticated or it is a gift to whoever is on the repeater.
  • Contracts settle results, not root shells.
  • Human on the loop for privileged system profiles — the AI Cluster management agent is not a toy on a public mesh.

Mobile does not mean careless. It means the threat model includes theft, mesh poisoning, and “the peer that was honest yesterday.”

A practical path, not a big bang

The original rollout was phased: observe, then recommend, then act. Same here.

  1. One supervisor. AI Cluster on a Strix Halo-class box. Local models, local UI, Helix, Vault. Learn admission on unified memory. This is now.
  2. Stack. Second box for knowledge or a specialist engine. Encrypted backups that you have actually restored.
  3. Battery. Measure watts. Size the station and panels for the scheduler’s worst-case residency, not the idle UI.
  4. One peer. WebRTC or a boring WireGuard link to a friend. Async jobs only. Receipts.
  5. Mesh. Repeater network when there are enough nodes in a building or a block. Prefer local.
  6. Contracts. Meter watts and deadlines. Then, and only then, open the front door wider.
  7. Video-path prototype. Encode a context pack as a stream. Prove it through a hostile NAT. Do not start here.

I am not claiming the mesh is done. I am claiming the unit of the Sentient Data Center has already changed. The building was a convenient chassis. The chassis is now a 128 GB APU, a battery, and whatever you can reach.

The immune system does not need a raised floor. It needs honest schedulers, unique specialists, and a way to talk when the only path left looks like a video call.