Levi DeHaan

LlamaLoadBalancer - AI Model Load Balancer

FastAPI-based intelligent load balancer for multiple Llama language models with Redis queues, PostgreSQL persistence, Prometheus metrics, and specialized model routing for finance, programming, log analysis, and creative tasks.

Status: completed · 2024-10-01

Overview

LlamaLoadBalancer is a sophisticated FastAPI-based application that serves as an intelligent load balancer for multiple Llama language models. It provides a unified API interface for accessing specialized AI models optimized for tasks like financial analysis, programming assistance, log analysis, and creative writing, with Redis-based queuing, PostgreSQL persistence, and comprehensive monitoring.

Technologies

Python, FastAPI, Redis, PostgreSQL, Prometheus, Docker, NVIDIA CUDA, Llama.cpp, GPU Memory Management, Model Management, Queue Management, API Design, Microservices Architecture, Request Routing, Performance Monitoring, Data Persistence, Concurrent Processing, Error Handling, Structured Logging, Configuration Management, Containerization, Load Balancing, Intelligent AI Routing, Enterprise Deployment

Model Types Supported
6 Specialized
Concurrent Requests
30+ Max
API Response Time
< 2 seconds
GPU Memory Efficiency
Optimized

LlamaLoadBalancer

A sophisticated AI model load balancer and API server that provides intelligent routing and management of multiple Llama language models for specialized tasks.

I built this while at Petco

Overview

LlamaLoadBalancer is a FastAPI-based application that serves as an intelligent load balancer for multiple Llama language models. It provides a unified API interface for accessing various specialized AI models, each optimized for different types of tasks such as financial analysis, programming assistance, log analysis, and creative writing.

Features

🤖 Multiple Specialized AI Models

  • Finance Model: Specialized for financial data analysis and transaction classification
  • Programmer Model: Expert coding assistant for software development tasks
  • Log Analysis Models: Security-focused log parsing and threat detection
  • Creative Models: Various creative and role-playing AI personas
  • Question Answering: Advanced Q&A capabilities with structured responses

🔄 Intelligent Load Balancing

  • Redis-based Queue Management: Efficient request queuing and processing
  • Model Hot-swapping: Dynamic model loading/unloading based on demand
  • Request Prioritization: Smart request handling and timeout management
  • Concurrent Processing: Multi-threaded request processing

📊 Monitoring & Observability

  • Prometheus Integration: Built-in metrics and monitoring
  • Structured Logging: Comprehensive logging with request tracking
  • Request History: Full audit trail of all requests and responses
  • Performance Metrics: Response time tracking and optimization

🗄️ Data Persistence

  • PostgreSQL Integration: Persistent storage of requests and responses
  • Request Tracking: Complete request lifecycle management
  • Response Caching: Efficient response storage and retrieval

Architecture

Core Components

Model Manager (`app/models/model_manager.py`)

  • Singleton pattern for centralized model management
  • Dynamic model loading with GPU optimization
  • Specialized prompt engineering for each model type
  • Memory management and garbage collection

Queue Manager (`app/queues/queue_manager.py`)

  • Redis-based request queuing system
  • Background processing with threading
  • Request deduplication and status tracking
  • Timeout and retry mechanisms

API Layer (`app/requests/`)

  • v1_completions: Standard completion API
  • v2_completions: Enhanced completion API with additional features
  • v1_intel_chain: Advanced reasoning chain for complex tasks
  • RESTful endpoints with proper error handling

Supported Model Types

Model TypePurposeContext SizeSpecial Features
`FINANCE`Financial analysis30K tokensJSON structured responses
`PROGRAMMER`Code generation30K tokensMulti-language support
`LOGGEN`Log analysis30K tokensSecurity threat detection
`HIGHERFORM`General AI30K tokensFlexible response formats
`QUEENBASED`Game strategy20K tokensTactical decision making
`METALLAMA3C`General purpose8K tokensFast inference

API Usage

Basic Completion Request

```bash curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ -d '{ "model": "PROGRAMMER", "messages": [ {"role": "user", "content": "Write a Python function to calculate fibonacci numbers"} ], "temperature": 0.7, "max_tokens": 1000 }' ```

Streaming Response

```bash curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ -d '{ "model": "FINANCE", "messages": [ {"role": "user", "content": "Analyze this transaction data"} ], "stream": true }' ```

GET Request (Simple)

```bash curl "http://localhost:8000/v1/completions?prompt=Hello&prompt=model=PROGRAMMER" ```

Installation & Setup

Prerequisites

  • GPU: NVIDIA GPU with CUDA 11.7+ support
  • Docker: For containerized deployment
  • Redis: For request queuing
  • PostgreSQL: For data persistence

Environment Configuration

Create a `.env` file:

```env REDIS_HOST=localhost REDIS_PORT=6380 DATABASE_URL=postgresql://user:password@localhost:5433/mydatabase ```

Docker Deployment

  1. Build the container: ```bash docker build -t llama-loadbalancer . ```

  2. Run with GPU support: ```bash docker run -d \ --name llama-loadbalancer \ --gpus all \ -p 8000:8000 \ -e REDIS_HOST=your-redis-host \ -e DATABASE_URL=your-postgres-url \ llama-loadbalancer ```

Local Development

  1. Install dependencies: ```bash pip install -r requirements.txt ```

  2. Set up Redis and PostgreSQL

  3. Run the application: ```bash uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload ```

Configuration

Model Configuration

Models are configured in `app/models/model_manager.py`:

```python model_paths = { "PROGRAMMER": "/path/to/programming-model.gguf", "FINANCE": "/path/to/finance-model.gguf", # ... other models } ```

Queue Configuration

Adjust queue settings in `app/queues/queue_manager.py`:

```python processing_threshold = 30 # Maximum concurrent requests retry_limit = 3 # Retry attempts timeout = 300 # Request timeout in seconds ```

Use Cases

🏢 Enterprise Applications

  • Customer Service: Intelligent chatbots with specialized knowledge
  • Content Generation: Automated content creation for marketing
  • Code Review: Automated code analysis and suggestions
  • Financial Analysis: Transaction classification and risk assessment

🔒 Security & Monitoring

  • Log Analysis: Real-time security threat detection
  • Anomaly Detection: Unusual pattern identification
  • Compliance Monitoring: Automated compliance checking

🎮 Gaming & Entertainment

  • NPC AI: Intelligent non-player characters
  • Story Generation: Dynamic content creation
  • Game Strategy: AI-powered game mechanics

📚 Education & Research

  • Tutoring Systems: Personalized learning assistance
  • Research Assistance: Literature review and analysis
  • Language Learning: Conversational AI tutors

Performance Optimization

GPU Memory Management

  • Models are automatically loaded/unloaded based on demand
  • Memory-efficient inference with optimized parameters
  • Garbage collection after model operations

Request Optimization

  • Redis-based queuing prevents system overload
  • Request batching for improved throughput
  • Intelligent model selection based on task requirements

Monitoring & Scaling

  • Prometheus metrics for performance monitoring
  • Horizontal scaling support with Redis clustering
  • Load balancing across multiple GPU instances

Contributing

  1. Fork the repository
  2. Create a feature branch
  3. Add comprehensive tests
  4. Submit a pull request with detailed description

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

For support and questions:

  • Create an issue in the GitHub repository
  • Check the documentation for common solutions
  • Review the example configurations

Roadmap

  • Support for additional model architectures
  • Enhanced streaming capabilities
  • Kubernetes deployment templates
  • Advanced caching mechanisms
  • Multi-language model support
  • Enhanced security features