Levi DeHaan

Petco SimpleAI (S3 Model Loading)

Heavily customized LocalAI fork with asynchronous S3 model loading, real-time status tracking, two-level queue architecture, health monitoring, CUDA 12.0 optimization, and advanced build system for GPU-accelerated AI model serving.

Status: completed · 2024-04-01

Overview

Petco SimpleAI is a heavily customized version of LocalAI that includes asynchronous S3 model loading, real-time status tracking, advanced two-level queue architecture, comprehensive health monitoring, CUDA 12.0 optimization, and a sophisticated build system for GPU-accelerated AI model serving with automatic model hot-swapping and intelligent load balancing.

Technologies

Go, C++, CUDA, Docker, Kubernetes, LocalAI, S3 Integration, GRPC, Protocol Buffers, Makefile, CMake, NVIDIA CUDA 12.0, GPU Memory Management, Queue Management, Load Balancing, Health Monitoring, Metrics Collection, Real-time Status, YAML Configuration, Model Hot-swapping, Containerization, Microservices Architecture, Build Automation, Environment Management, Concurrency Control, Error Handling, Performance Optimization, Scalability, Infrastructure Management, DevOps Tools

Model Loading
S3 Asynchronous
Queue Architecture
Two-Level
Build Optimization
CUDA 12.0
Monitoring
Real-time Health

Petco - SIMPLEAI Base with S3 Model Loading

Internal AI server implementation for running local LLM models with advanced S3 integration and real-time model loading capabilities.

Overview

This is a heavily customized version of LocalAI that includes:

  • S3 Model Loading: Automatic downloading and hot-loading of models from S3 storage
  • Real-time Status Tracking: Live progress monitoring and model availability status
  • Advanced Queue System: Two-level queue architecture with intelligent load balancing
  • Health Monitoring: Comprehensive system metrics and resource tracking
  • CUDA 12.0 Optimization: Custom build system optimized for NVIDIA GPUs

Key Features

S3 Model Loading System

  • Asynchronous model downloads that don't block server startup
  • YAML-first prioritization for immediate model configuration loading
  • Incremental model availability - models become usable as soon as downloaded
  • Real-time progress tracking via REST API endpoint

Advanced Architecture

  • Global queue manager for request distribution across multiple servers
  • Local queue managers for individual model server resource allocation
  • Health-based load balancing and automatic failover
  • Resource usage optimization and monitoring

Monitoring & Metrics

  • Real-time health metrics collection
  • Resource utilization tracking
  • Error rate monitoring and categorization
  • Performance trend analysis

Prerequisites

Ensure you have the following installed on your build system:

  1. GCC 12 or later
  2. Python 3.10 or later with pip
  3. Go 1.22.6 or later
  4. Docker
  5. CUDA 12.0 toolkit (for CUDA builds)

Required Python Packages

Install the following Python packages:

S3 Model Loading System

This system provides advanced S3 integration for automatic model downloading and loading.

Configuration

Configure S3 access via environment variables or `.env` file:

```bash

S3 Configuration

S3_BUCKET_URL=s3://your-bucket-name S3_PREFIX=models/ # Optional: S3 key prefix S3_REGION=us-west-2 # AWS region S3_ENDPOINT=https://custom-s3-endpoint.com # Optional: Custom S3 endpoint S3_FILE_LIST=model1.yaml,model2.yaml # Optional: Specific files to download ```

Model Storage Structure

Organize your models in S3 with this structure:

``` s3://your-bucket/models/ ├── llama-7b.yaml # Model configuration ├── llama-7b.ggml # Model weights ├── codellama.yaml ├── codellama.ggml └── ... ```

Real-time Status Tracking

Monitor download progress and model availability:

```bash

Check S3 download status

curl http://localhost:8080/s3/status

Response includes:

- Overall download progress

- Individual file status

- Ready models (available for inference)

- Error information if downloads fail

```

Model Availability

Models become available for inference as soon as both their YAML configuration and model files are downloaded. The system:

  1. Prioritizes YAML files for immediate configuration loading
  2. Loads model configurations as YAML files are downloaded
  3. Tracks model readiness in real-time
  4. Makes models available without requiring server restart

API Endpoints

  • `GET /s3/status` - Get current S3 download status and model availability
  • `GET /models/loaded` - List currently loaded models
  • `GET /backend/monitor` - System health and resource metrics

Building with build.sh

The primary method for building this system is using the `build.sh` script. This script handles:

  1. Version Management

    • Automatically increments version numbers
    • Allows manual version override
    • Maintains version tracking in `.last_version` file
  2. Build Configuration

    • Sets up CUDA build environment (CUDA 12.0)
    • Configures build parameters:
      • Uses P2P tags
      • Enables CUBLAS support
      • Sets up external GRPC backends for various AI models
      • Configures Go (1.22.6) and CMake (3.26.4)
  3. Docker Image Creation

    • Builds using ubuntu:22.04 as base image
    • Creates CUDA-enabled container
    • Tags image with version
    • Pushes to ECR repository
  4. Kubernetes Integration

    • Updates deployment manifests with new version
    • Prepares for immediate deployment

Usage

To build the system:

```bash

1. Generate Protocol Buffer files

make protogen

2. Prepare the build environment

make prepare

3. Run the build script

./build.sh ```

The script will:

  1. Prompt for version confirmation
  2. Build Docker image with CUDA support
  3. Tag and push to ECR
  4. Update Kubernetes manifests

Build Process Details

The build process involves several key steps:

  1. Protocol Buffer Generation

    • Required for gRPC communication
    • Must be done before the main build
    • Uses `grpcio-tools` for Python protobuf generation
  2. Environment Preparation

    • Clones and initializes required submodules
    • Sets up Go module dependencies
    • Configures CUDA build environment
  3. Docker Build

    • Uses multi-stage build for optimization
    • Includes CUDA support for GPU acceleration
    • Installs and configures all required backends

Troubleshooting Common Build Issues

  1. Protocol Buffer Generation Failures

    • Ensure `grpcio-tools` is installed: `pip install grpcio-tools`
    • Run `make protogen` before the main build
  2. Go Module Issues

    • If you see "updates to go.mod needed", run `go mod tidy`
    • Ensure all Go dependencies are properly initialized
  3. CUDA Build Problems

    • Verify CUDA toolkit installation
    • Check CUDA version compatibility (currently using CUDA 12.0)
    • Ensure proper GPU drivers are installed

CUDA Build local binary

```bash make BUILD_TYPE=cublas build-minimal ```

Docker CUDA Build

```bash make BUILD_TYPE=cublas docker-cuda12 ```

CPU-only Build

```bash

CPU only image:

docker run -ti --name local-ai -p 8080:8080 localai/localai:latest-cpu

Nvidia GPU:

docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-gpu-nvidia-cuda-12

CPU and GPU image (bigger size):

docker run -ti --name local-ai -p 8080:8080 localai/localai:latest

AIO images (it will pre-download a set of models ready for use, see https://localai.io/basics/container/)

docker run -ti --name local-ai -p 8080:8080 localai/localai:latest-aio-cpu ```

Running the Server

With CUDA Support

```bash make clean ```

  1. Build the project: ```bash make build-minimal ```

Building the Project

Understanding the Build System

The project uses a sophisticated build system that creates different variants (CUDA, AVX2) from a base template:

  1. Base templates are stored in `/backend/cpp/llama/`
  2. Build variants (cuda, avx2) are created from this base
  3. `prepare.sh` handles file distribution and patching
  4. Variant-specific compilation occurs with appropriate CMAKE_ARGS

Key Build Files

  • `/backend/cpp/llama/prepare.sh`: Central script for file distribution
  • `/backend/cpp/llama/llama_client_slot.hpp`: Core client slot definitions
  • `/backend/cpp/llama/server_metrics.hpp`: Server metrics tracking
  • `/backend/cpp/llama/utils.hpp`: Utility functions and logging
  • `/backend/cpp/llama/grpc-server.cpp`: Main server implementation

Code Organization

Core Structs

The system uses several key structs for different purposes:

  1. `slot_params` (in `utils.hpp`):

    • Basic slot configuration
    • Stream settings
    • Cache settings
    • Seed and prediction parameters
  2. `common_sampler_params` (from llama.cpp):

    • Sampling parameters (top_k, top_p)
    • Temperature control
    • Penalty settings
    • Mirostat parameters

These structs are used together in `llama_client_slot` to provide comprehensive request handling and text generation control.

Core Components

The system is built around several key components and structures:

  1. Request Processing

    • `slot_params`: Basic slot configuration for request handling
    • `common_sampler_params`: Advanced text generation control
    • Integrated queue management system
    • Health monitoring and metrics collection
  2. Sampling Parameters The system now supports comprehensive text generation control through `common_sampler_params`:

    • Temperature and dynamic temperature control
    • Top-k and top-p sampling
    • Penalty settings for repetition and frequency
    • Mirostat sampling parameters
    • Grammar-constrained generation
    • Token biasing and penalty systems
  3. Queue Management

    • Two-level queue architecture (global and local)
    • Resource-aware request distribution
    • Health-based load balancing
    • Graceful shutdown and restart capabilities

Build Process Flow

  1. `make clean` removes previous builds
  2. Base directory is copied to create variant (e.g., llama-cuda)
  3. `prepare.sh` copies and patches necessary files
  4. CMake configures with variant-specific flags
  5. Compilation occurs with specified optimizations

Troubleshooting Builds

  • Check that all files exist in `/backend/cpp/llama/`
  • Verify `prepare.sh` has correct copy commands
  • Ensure CMAKE_ARGS match your system
  • Monitor build logs for missing dependencies

Required installed packages

  • `gcc-10`
  • `cuda`
  • `protobuf`
  • `cmake`
  • `make`
  • `go`
  • `upx`

Changes Made

The following changes were made to support building on Arch Linux:

Makefile Changes

Here are all the changes made to the Makefile:

  1. Added shell and path cleanup: ```makefile SHELL := /bin/bash

Function to clean PATH of anaconda/conda

define CLEAN_PATH $(shell PATH="$$(echo "$$PATH" | tr ':' '\n' | grep -v -e anaconda -e conda | tr '\n' ':')"; echo "$$PATH") endef

Clean environment variables that might interfere with the build

unexport CONDA_PREFIX unexport CONDA_DEFAULT_ENV unexport CONDA_PYTHON_EXE unexport CONDA_EXE unexport CONDA_SHLVL

Set clean PATH

export PATH := $(CLEAN_PATH) ```

  1. Added gcc-10 configuration: ```makefile

Force gcc-10 usage

override CC := /usr/bin/gcc-10 override CXX := /usr/bin/g++-10 export PATH := /usr/lib/gcc/x86_64-pc-linux-gnu/10/:/usr/lib/gcc/x86_64-pc-linux-gnu/10/:/usr/lib/gcc/x86_64-pc-linux-gnu/10/:$(PATH)

Add compiler flags to force gcc-10

override CFLAGS += -B/usr/lib/gcc/x86_64-pc-linux-gnu/10/ override CXXFLAGS += -B/usr/lib/gcc/x86_64-pc-linux-gnu/10/ override LDFLAGS += -B/usr/lib/gcc/x86_64-pc-linux-gnu/10/ ```

  1. Updated CUDA configuration for Arch Linux: ```makefile

CUDA setup

export CUDA_HOME := /opt/cuda export CUDA_PATH := /opt/cuda export LD_LIBRARY_PATH := $(CUDA_HOME)/targets/x86_64-linux/lib:$(LD_LIBRARY_PATH) export LIBRARY_PATH := $(CUDA_HOME)/targets/x86_64-linux/lib/stubs:$(LIBRARY_PATH) export PATH := $(CUDA_HOME)/bin:$(PATH)

Default to cublas build type if not specified

BUILD_TYPE ?= cublas

Update CUDA library path for build

CUDA_LIBPATH?=$(CUDA_HOME)/targets/x86_64-linux/lib/ ```

  1. Added parallel build support: ```makefile #set threads -j to 24 export MAKEFLAGS += -j24 ```

go.mod Changes

Added local replacements for dependencies to speed up build process.

go.mod Changes

Added local replacements for dependencies: ```go replace github.com/donomii/go-rwkv.cpp => /storage/Projects/simpleai/sources/go-rwkv.cpp replace github.com/ggerganov/whisper.cpp => /storage/Projects/simpleai/sources/whisper.cpp replace github.com/ggerganov/whisper.cpp/bindings/go => /storage/Projects/simpleai/sources/whisper.cpp/bindings/go replace github.com/go-skynet/go-bert.cpp => /storage/Projects/simpleai/sources/go-bert.cpp replace github.com/M0Rf30/go-tiny-dream => /storage/Projects/simpleai/sources/go-tiny-dream replace github.com/mudler/go-piper => /storage/Projects/simpleai/sources/go-piper replace github.com/mudler/go-stable-diffusion => /storage/Projects/simpleai/sources/go-stable-diffusion replace github.com/go-skynet/go-llama.cpp => /storage/Projects/simpleai/sources/go-llama.cpp ```

Architecture Overview

Two-Level Queue System

The system implements a sophisticated two-level queue architecture:

Global Queue Manager (Go)

  • Receives all initial user requests
  • Handles request distribution across multiple model servers
  • Manages server selection based on load and health
  • Tracks request lifecycle across the entire system
  • Implements retry and failover logic

Local Queue Manager (C++)

  • Receives requests forwarded from the global manager
  • Manages local CPU/GPU queues for actual request processing
  • Handles resource allocation within a single model server
  • Processes requests and returns results back to global manager

Integration Points

  • Global manager uses gRPC to forward requests to appropriate local managers
  • Local managers process requests and send responses back through gRPC
  • Health metrics and status updates flow from local to global managers
  • Global manager maintains overall system state and coordinates between servers

Health Monitoring System

Comprehensive monitoring includes:

  • Real-time Metrics: Queue depth, request processing times, resource utilization
  • Error Tracking: Categorized error monitoring with automatic recovery
  • Resource Monitoring: CPU, GPU, and memory usage tracking
  • Performance Analysis: Trend analysis and bottleneck identification

Configuration Options

Environment variables for system configuration:

```bash

Queue Configuration

QUEUE_MAX_SIZE=1000 # Maximum queue size QUEUE_SERVER_RESTART_TIMEOUT=300s # Server restart timeout QUEUE_DEAD_REQUEST_TIMEOUT=60s # Dead request timeout QUEUE_MAX_RESTART_ATTEMPTS=3 # Maximum restart attempts QUEUE_RETRY_BACKOFF_MIN=1s # Minimum retry backoff QUEUE_RETRY_BACKOFF_MAX=30s # Maximum retry backoff

S3 Configuration

S3_BUCKET_URL=s3://your-bucket-name S3_PREFIX=models/ S3_REGION=us-west-2 S3_ENDPOINT=https://custom-endpoint.com

Model Configuration

MODEL_PATH=/models # Local model storage path THREADS=4 # Number of inference threads CONTEXT_SIZE=2048 # Maximum context size ```


System Workflow

S3 Model Loading Workflow

  1. Server Startup: Server starts and begins S3 downloads in background
  2. YAML Priority: YAML configuration files are downloaded first
  3. Configuration Loading: Each YAML file is loaded as soon as downloaded
  4. Model Readiness: Models become available when both config and weights are ready
  5. Status Tracking: Real-time updates via `/s3/status` endpoint

Request Processing Workflow

  1. Request Arrival: User requests arrive at global queue manager
  2. Server Selection: Global manager selects optimal server based on health/load
  3. Request Forwarding: Request forwarded to selected server's local queue
  4. Processing: Local queue processes request using appropriate model
  5. Response: Results returned through the same path

Monitoring Workflow

  1. Metric Collection: Local managers collect performance metrics
  2. Aggregation: Global manager aggregates metrics across all servers
  3. Analysis: System analyzes trends and identifies bottlenecks

Troubleshooting

S3 Connection Issues

Problem: S3 downloads fail with authentication errors

Solutions:

  1. Verify AWS credentials are properly configured
  2. Check S3 bucket permissions
  3. Ensure correct region and endpoint settings
  4. Verify network connectivity to S3 endpoint

Problem: Models not appearing after S3 download

Solutions:

  1. Check `/s3/status` endpoint for download progress
  2. Verify YAML files are valid and contain correct model configurations
  3. Check server logs for configuration loading errors
  4. Ensure model files are in correct format (GGML, GGUF, etc.)

Build Issues

Problem: CUDA build failures on Arch Linux

Solutions:

  1. Ensure CUDA 12.0 toolkit is properly installed
  2. Check that `gcc-10` and `g++-10` are available
  3. Verify CUDA library paths in Makefile
  4. Run `make clean` before rebuilding

Problem: Protocol buffer generation failures

Solutions:

  1. Install `grpcio-tools`: `pip install grpcio-tools`
  2. Run `make protogen` before main build
  3. Check Python version (3.10+ required)

Runtime Issues

Problem: High memory usage

Solutions:

  1. Adjust `THREADS` environment variable
  2. Monitor resource usage via `/backend/monitor`
  3. Check for memory leaks in model configurations
  4. Consider using smaller model variants

Problem: Poor performance or timeouts

Solutions:

  1. Check queue configuration settings
  2. Monitor server health via health endpoints
  3. Adjust context size and thread settings
  4. Verify GPU utilization if using CUDA builds

Model Issues

Problem: Models not loading correctly

Solutions:

  1. Verify model file integrity and format
  2. Check YAML configuration syntax
  3. Ensure model path permissions are correct
  4. Review server logs for specific error messages

Problem: Model inference failures

Solutions:

  1. Check model compatibility with LocalAI
  2. Verify model file is not corrupted
  3. Ensure sufficient system resources (RAM, VRAM)
  4. Check model-specific configuration requirements

Summary of Changes

  1. Makefile:

    • Added environment cleanup for conda/anaconda
    • Configured gcc-10 as the compiler with proper paths
    • Updated CUDA paths to match Arch Linux installation
    • Added proper library paths for CUDA
    • Added parallel build support with 24 threads
    • Fixed linking issues with CUDA libraries
  2. go.mod:

    • Added local replacements for all dependencies to use local sources
  3. Build Process:

    • Now uses the correct CUDA library paths
    • Properly links against CUDA libraries
    • Uses gcc-10 for compilation
    • Builds in parallel for faster compilation
  4. S3 Integration:

    • Added asynchronous S3 download system
    • Implemented YAML-first prioritization
    • Added real-time status tracking via `/s3/status`
    • Integrated automatic model configuration loading
    • Added ready model tracking for incremental availability
  5. Architecture Enhancements:

    • Implemented two-level queue system (global/local)
    • Added comprehensive health monitoring
    • Enhanced request distribution and load balancing
    • Added automatic failover and retry mechanisms

Message Flow:

a. Global Queue Manager (Go): + Receives all initial user requests + Handles request distribution across servers + Manages server selection based on load and health + Tracks request lifecycle across the entire system + Implements retry and failover logic

b. Local Queue Manager (C++): + Receives requests forwarded from the global manager + Manages local CPU/GPU queues for actual request processing + Handles resource allocation within a single model server + Processes requests and returns results back to global manager

Integration Points:

+ The global manager uses gRPC to forward requests to appropriate local managers
+ Local managers process requests and send responses back through gRPC
+ Health metrics and status updates flow from local to global managers
+ The global manager maintains overall system state and coordinates between servers

to make BERT you need to go to /storage/Projects/simpleai/sources/go-bert.cpp

then run make make libgobert then cd to bert.ccp and run make bert