Petco SimpleAI (S3 Model Loading)
Heavily customized LocalAI fork with asynchronous S3 model loading, real-time status tracking, two-level queue architecture, health monitoring, CUDA 12.0 optimization, and advanced build system for GPU-accelerated AI model serving.
Status: completed · 2024-04-01
Overview
Petco SimpleAI is a heavily customized version of LocalAI that includes asynchronous S3 model loading, real-time status tracking, advanced two-level queue architecture, comprehensive health monitoring, CUDA 12.0 optimization, and a sophisticated build system for GPU-accelerated AI model serving with automatic model hot-swapping and intelligent load balancing.
Technologies
Go, C++, CUDA, Docker, Kubernetes, LocalAI, S3 Integration, GRPC, Protocol Buffers, Makefile, CMake, NVIDIA CUDA 12.0, GPU Memory Management, Queue Management, Load Balancing, Health Monitoring, Metrics Collection, Real-time Status, YAML Configuration, Model Hot-swapping, Containerization, Microservices Architecture, Build Automation, Environment Management, Concurrency Control, Error Handling, Performance Optimization, Scalability, Infrastructure Management, DevOps Tools
- Model Loading
- S3 Asynchronous
- Queue Architecture
- Two-Level
- Build Optimization
- CUDA 12.0
- Monitoring
- Real-time Health
Petco - SIMPLEAI Base with S3 Model Loading
Internal AI server implementation for running local LLM models with advanced S3 integration and real-time model loading capabilities.
Overview
This is a heavily customized version of LocalAI that includes:
- S3 Model Loading: Automatic downloading and hot-loading of models from S3 storage
- Real-time Status Tracking: Live progress monitoring and model availability status
- Advanced Queue System: Two-level queue architecture with intelligent load balancing
- Health Monitoring: Comprehensive system metrics and resource tracking
- CUDA 12.0 Optimization: Custom build system optimized for NVIDIA GPUs
Key Features
S3 Model Loading System
- Asynchronous model downloads that don't block server startup
- YAML-first prioritization for immediate model configuration loading
- Incremental model availability - models become usable as soon as downloaded
- Real-time progress tracking via REST API endpoint
Advanced Architecture
- Global queue manager for request distribution across multiple servers
- Local queue managers for individual model server resource allocation
- Health-based load balancing and automatic failover
- Resource usage optimization and monitoring
Monitoring & Metrics
- Real-time health metrics collection
- Resource utilization tracking
- Error rate monitoring and categorization
- Performance trend analysis
Prerequisites
Ensure you have the following installed on your build system:
- GCC 12 or later
- Python 3.10 or later with pip
- Go 1.22.6 or later
- Docker
- CUDA 12.0 toolkit (for CUDA builds)
Required Python Packages
Install the following Python packages:
S3 Model Loading System
This system provides advanced S3 integration for automatic model downloading and loading.
Configuration
Configure S3 access via environment variables or `.env` file:
```bash
S3 Configuration
S3_BUCKET_URL=s3://your-bucket-name S3_PREFIX=models/ # Optional: S3 key prefix S3_REGION=us-west-2 # AWS region S3_ENDPOINT=https://custom-s3-endpoint.com # Optional: Custom S3 endpoint S3_FILE_LIST=model1.yaml,model2.yaml # Optional: Specific files to download ```
Model Storage Structure
Organize your models in S3 with this structure:
``` s3://your-bucket/models/ ├── llama-7b.yaml # Model configuration ├── llama-7b.ggml # Model weights ├── codellama.yaml ├── codellama.ggml └── ... ```
Real-time Status Tracking
Monitor download progress and model availability:
```bash
Check S3 download status
curl http://localhost:8080/s3/status
Response includes:
- Overall download progress
- Individual file status
- Ready models (available for inference)
- Error information if downloads fail
```
Model Availability
Models become available for inference as soon as both their YAML configuration and model files are downloaded. The system:
- Prioritizes YAML files for immediate configuration loading
- Loads model configurations as YAML files are downloaded
- Tracks model readiness in real-time
- Makes models available without requiring server restart
API Endpoints
- `GET /s3/status` - Get current S3 download status and model availability
- `GET /models/loaded` - List currently loaded models
- `GET /backend/monitor` - System health and resource metrics
Building with build.sh
The primary method for building this system is using the `build.sh` script. This script handles:
-
Version Management
- Automatically increments version numbers
- Allows manual version override
- Maintains version tracking in `.last_version` file
-
Build Configuration
- Sets up CUDA build environment (CUDA 12.0)
- Configures build parameters:
- Uses P2P tags
- Enables CUBLAS support
- Sets up external GRPC backends for various AI models
- Configures Go (1.22.6) and CMake (3.26.4)
-
Docker Image Creation
- Builds using ubuntu:22.04 as base image
- Creates CUDA-enabled container
- Tags image with version
- Pushes to ECR repository
-
Kubernetes Integration
- Updates deployment manifests with new version
- Prepares for immediate deployment
Usage
To build the system:
```bash
1. Generate Protocol Buffer files
make protogen
2. Prepare the build environment
make prepare
3. Run the build script
./build.sh ```
The script will:
- Prompt for version confirmation
- Build Docker image with CUDA support
- Tag and push to ECR
- Update Kubernetes manifests
Build Process Details
The build process involves several key steps:
-
Protocol Buffer Generation
- Required for gRPC communication
- Must be done before the main build
- Uses `grpcio-tools` for Python protobuf generation
-
Environment Preparation
- Clones and initializes required submodules
- Sets up Go module dependencies
- Configures CUDA build environment
-
Docker Build
- Uses multi-stage build for optimization
- Includes CUDA support for GPU acceleration
- Installs and configures all required backends
Troubleshooting Common Build Issues
-
Protocol Buffer Generation Failures
- Ensure `grpcio-tools` is installed: `pip install grpcio-tools`
- Run `make protogen` before the main build
-
Go Module Issues
- If you see "updates to go.mod needed", run `go mod tidy`
- Ensure all Go dependencies are properly initialized
-
CUDA Build Problems
- Verify CUDA toolkit installation
- Check CUDA version compatibility (currently using CUDA 12.0)
- Ensure proper GPU drivers are installed
CUDA Build local binary
```bash make BUILD_TYPE=cublas build-minimal ```
Docker CUDA Build
```bash make BUILD_TYPE=cublas docker-cuda12 ```
CPU-only Build
```bash
CPU only image:
docker run -ti --name local-ai -p 8080:8080 localai/localai:latest-cpu
Nvidia GPU:
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-gpu-nvidia-cuda-12
CPU and GPU image (bigger size):
docker run -ti --name local-ai -p 8080:8080 localai/localai:latest
AIO images (it will pre-download a set of models ready for use, see https://localai.io/basics/container/)
docker run -ti --name local-ai -p 8080:8080 localai/localai:latest-aio-cpu ```
Running the Server
With CUDA Support
```bash make clean ```
- Build the project: ```bash make build-minimal ```
Building the Project
Understanding the Build System
The project uses a sophisticated build system that creates different variants (CUDA, AVX2) from a base template:
- Base templates are stored in `/backend/cpp/llama/`
- Build variants (cuda, avx2) are created from this base
- `prepare.sh` handles file distribution and patching
- Variant-specific compilation occurs with appropriate CMAKE_ARGS
Key Build Files
- `/backend/cpp/llama/prepare.sh`: Central script for file distribution
- `/backend/cpp/llama/llama_client_slot.hpp`: Core client slot definitions
- `/backend/cpp/llama/server_metrics.hpp`: Server metrics tracking
- `/backend/cpp/llama/utils.hpp`: Utility functions and logging
- `/backend/cpp/llama/grpc-server.cpp`: Main server implementation
Code Organization
Core Structs
The system uses several key structs for different purposes:
-
`slot_params` (in `utils.hpp`):
- Basic slot configuration
- Stream settings
- Cache settings
- Seed and prediction parameters
-
`common_sampler_params` (from llama.cpp):
- Sampling parameters (top_k, top_p)
- Temperature control
- Penalty settings
- Mirostat parameters
These structs are used together in `llama_client_slot` to provide comprehensive request handling and text generation control.
Core Components
The system is built around several key components and structures:
-
Request Processing
- `slot_params`: Basic slot configuration for request handling
- `common_sampler_params`: Advanced text generation control
- Integrated queue management system
- Health monitoring and metrics collection
-
Sampling Parameters The system now supports comprehensive text generation control through `common_sampler_params`:
- Temperature and dynamic temperature control
- Top-k and top-p sampling
- Penalty settings for repetition and frequency
- Mirostat sampling parameters
- Grammar-constrained generation
- Token biasing and penalty systems
-
Queue Management
- Two-level queue architecture (global and local)
- Resource-aware request distribution
- Health-based load balancing
- Graceful shutdown and restart capabilities
Build Process Flow
- `make clean` removes previous builds
- Base directory is copied to create variant (e.g., llama-cuda)
- `prepare.sh` copies and patches necessary files
- CMake configures with variant-specific flags
- Compilation occurs with specified optimizations
Troubleshooting Builds
- Check that all files exist in `/backend/cpp/llama/`
- Verify `prepare.sh` has correct copy commands
- Ensure CMAKE_ARGS match your system
- Monitor build logs for missing dependencies
Required installed packages
- `gcc-10`
- `cuda`
- `protobuf`
- `cmake`
- `make`
- `go`
- `upx`
Changes Made
The following changes were made to support building on Arch Linux:
Makefile Changes
Here are all the changes made to the Makefile:
- Added shell and path cleanup: ```makefile SHELL := /bin/bash
Function to clean PATH of anaconda/conda
define CLEAN_PATH $(shell PATH="$$(echo "$$PATH" | tr ':' '\n' | grep -v -e anaconda -e conda | tr '\n' ':')"; echo "$$PATH") endef
Clean environment variables that might interfere with the build
unexport CONDA_PREFIX unexport CONDA_DEFAULT_ENV unexport CONDA_PYTHON_EXE unexport CONDA_EXE unexport CONDA_SHLVL
Set clean PATH
export PATH := $(CLEAN_PATH) ```
- Added gcc-10 configuration: ```makefile
Force gcc-10 usage
override CC := /usr/bin/gcc-10 override CXX := /usr/bin/g++-10 export PATH := /usr/lib/gcc/x86_64-pc-linux-gnu/10/:/usr/lib/gcc/x86_64-pc-linux-gnu/10/:/usr/lib/gcc/x86_64-pc-linux-gnu/10/:$(PATH)
Add compiler flags to force gcc-10
override CFLAGS += -B/usr/lib/gcc/x86_64-pc-linux-gnu/10/ override CXXFLAGS += -B/usr/lib/gcc/x86_64-pc-linux-gnu/10/ override LDFLAGS += -B/usr/lib/gcc/x86_64-pc-linux-gnu/10/ ```
- Updated CUDA configuration for Arch Linux: ```makefile
CUDA setup
export CUDA_HOME := /opt/cuda export CUDA_PATH := /opt/cuda export LD_LIBRARY_PATH := $(CUDA_HOME)/targets/x86_64-linux/lib:$(LD_LIBRARY_PATH) export LIBRARY_PATH := $(CUDA_HOME)/targets/x86_64-linux/lib/stubs:$(LIBRARY_PATH) export PATH := $(CUDA_HOME)/bin:$(PATH)
Default to cublas build type if not specified
BUILD_TYPE ?= cublas
Update CUDA library path for build
CUDA_LIBPATH?=$(CUDA_HOME)/targets/x86_64-linux/lib/ ```
- Added parallel build support: ```makefile #set threads -j to 24 export MAKEFLAGS += -j24 ```
go.mod Changes
Added local replacements for dependencies to speed up build process.
go.mod Changes
Added local replacements for dependencies: ```go replace github.com/donomii/go-rwkv.cpp => /storage/Projects/simpleai/sources/go-rwkv.cpp replace github.com/ggerganov/whisper.cpp => /storage/Projects/simpleai/sources/whisper.cpp replace github.com/ggerganov/whisper.cpp/bindings/go => /storage/Projects/simpleai/sources/whisper.cpp/bindings/go replace github.com/go-skynet/go-bert.cpp => /storage/Projects/simpleai/sources/go-bert.cpp replace github.com/M0Rf30/go-tiny-dream => /storage/Projects/simpleai/sources/go-tiny-dream replace github.com/mudler/go-piper => /storage/Projects/simpleai/sources/go-piper replace github.com/mudler/go-stable-diffusion => /storage/Projects/simpleai/sources/go-stable-diffusion replace github.com/go-skynet/go-llama.cpp => /storage/Projects/simpleai/sources/go-llama.cpp ```
Architecture Overview
Two-Level Queue System
The system implements a sophisticated two-level queue architecture:
Global Queue Manager (Go)
- Receives all initial user requests
- Handles request distribution across multiple model servers
- Manages server selection based on load and health
- Tracks request lifecycle across the entire system
- Implements retry and failover logic
Local Queue Manager (C++)
- Receives requests forwarded from the global manager
- Manages local CPU/GPU queues for actual request processing
- Handles resource allocation within a single model server
- Processes requests and returns results back to global manager
Integration Points
- Global manager uses gRPC to forward requests to appropriate local managers
- Local managers process requests and send responses back through gRPC
- Health metrics and status updates flow from local to global managers
- Global manager maintains overall system state and coordinates between servers
Health Monitoring System
Comprehensive monitoring includes:
- Real-time Metrics: Queue depth, request processing times, resource utilization
- Error Tracking: Categorized error monitoring with automatic recovery
- Resource Monitoring: CPU, GPU, and memory usage tracking
- Performance Analysis: Trend analysis and bottleneck identification
Configuration Options
Environment variables for system configuration:
```bash
Queue Configuration
QUEUE_MAX_SIZE=1000 # Maximum queue size QUEUE_SERVER_RESTART_TIMEOUT=300s # Server restart timeout QUEUE_DEAD_REQUEST_TIMEOUT=60s # Dead request timeout QUEUE_MAX_RESTART_ATTEMPTS=3 # Maximum restart attempts QUEUE_RETRY_BACKOFF_MIN=1s # Minimum retry backoff QUEUE_RETRY_BACKOFF_MAX=30s # Maximum retry backoff
S3 Configuration
S3_BUCKET_URL=s3://your-bucket-name S3_PREFIX=models/ S3_REGION=us-west-2 S3_ENDPOINT=https://custom-endpoint.com
Model Configuration
MODEL_PATH=/models # Local model storage path THREADS=4 # Number of inference threads CONTEXT_SIZE=2048 # Maximum context size ```
System Workflow
S3 Model Loading Workflow
- Server Startup: Server starts and begins S3 downloads in background
- YAML Priority: YAML configuration files are downloaded first
- Configuration Loading: Each YAML file is loaded as soon as downloaded
- Model Readiness: Models become available when both config and weights are ready
- Status Tracking: Real-time updates via `/s3/status` endpoint
Request Processing Workflow
- Request Arrival: User requests arrive at global queue manager
- Server Selection: Global manager selects optimal server based on health/load
- Request Forwarding: Request forwarded to selected server's local queue
- Processing: Local queue processes request using appropriate model
- Response: Results returned through the same path
Monitoring Workflow
- Metric Collection: Local managers collect performance metrics
- Aggregation: Global manager aggregates metrics across all servers
- Analysis: System analyzes trends and identifies bottlenecks
Troubleshooting
S3 Connection Issues
Problem: S3 downloads fail with authentication errors
Solutions:
- Verify AWS credentials are properly configured
- Check S3 bucket permissions
- Ensure correct region and endpoint settings
- Verify network connectivity to S3 endpoint
Problem: Models not appearing after S3 download
Solutions:
- Check `/s3/status` endpoint for download progress
- Verify YAML files are valid and contain correct model configurations
- Check server logs for configuration loading errors
- Ensure model files are in correct format (GGML, GGUF, etc.)
Build Issues
Problem: CUDA build failures on Arch Linux
Solutions:
- Ensure CUDA 12.0 toolkit is properly installed
- Check that `gcc-10` and `g++-10` are available
- Verify CUDA library paths in Makefile
- Run `make clean` before rebuilding
Problem: Protocol buffer generation failures
Solutions:
- Install `grpcio-tools`: `pip install grpcio-tools`
- Run `make protogen` before main build
- Check Python version (3.10+ required)
Runtime Issues
Problem: High memory usage
Solutions:
- Adjust `THREADS` environment variable
- Monitor resource usage via `/backend/monitor`
- Check for memory leaks in model configurations
- Consider using smaller model variants
Problem: Poor performance or timeouts
Solutions:
- Check queue configuration settings
- Monitor server health via health endpoints
- Adjust context size and thread settings
- Verify GPU utilization if using CUDA builds
Model Issues
Problem: Models not loading correctly
Solutions:
- Verify model file integrity and format
- Check YAML configuration syntax
- Ensure model path permissions are correct
- Review server logs for specific error messages
Problem: Model inference failures
Solutions:
- Check model compatibility with LocalAI
- Verify model file is not corrupted
- Ensure sufficient system resources (RAM, VRAM)
- Check model-specific configuration requirements
Summary of Changes
-
Makefile:
- Added environment cleanup for conda/anaconda
- Configured gcc-10 as the compiler with proper paths
- Updated CUDA paths to match Arch Linux installation
- Added proper library paths for CUDA
- Added parallel build support with 24 threads
- Fixed linking issues with CUDA libraries
-
go.mod:
- Added local replacements for all dependencies to use local sources
-
Build Process:
- Now uses the correct CUDA library paths
- Properly links against CUDA libraries
- Uses gcc-10 for compilation
- Builds in parallel for faster compilation
-
S3 Integration:
- Added asynchronous S3 download system
- Implemented YAML-first prioritization
- Added real-time status tracking via `/s3/status`
- Integrated automatic model configuration loading
- Added ready model tracking for incremental availability
-
Architecture Enhancements:
- Implemented two-level queue system (global/local)
- Added comprehensive health monitoring
- Enhanced request distribution and load balancing
- Added automatic failover and retry mechanisms
Message Flow:
a. Global Queue Manager (Go): + Receives all initial user requests + Handles request distribution across servers + Manages server selection based on load and health + Tracks request lifecycle across the entire system + Implements retry and failover logic
b. Local Queue Manager (C++): + Receives requests forwarded from the global manager + Manages local CPU/GPU queues for actual request processing + Handles resource allocation within a single model server + Processes requests and returns results back to global manager
Integration Points:
+ The global manager uses gRPC to forward requests to appropriate local managers
+ Local managers process requests and send responses back through gRPC
+ Health metrics and status updates flow from local to global managers
+ The global manager maintains overall system state and coordinates between servers
to make BERT you need to go to /storage/Projects/simpleai/sources/go-bert.cpp
then run make make libgobert then cd to bert.ccp and run make bert