LlamaLoadBalancer - AI Model Load Balancer
FastAPI-based intelligent load balancer for multiple Llama language models with Redis queues, PostgreSQL persistence, Prometheus metrics, and specialized model routing for finance, programming, log analysis, and creative tasks.
Status: completed · 2024-10-01
Overview
LlamaLoadBalancer is a sophisticated FastAPI-based application that serves as an intelligent load balancer for multiple Llama language models. It provides a unified API interface for accessing specialized AI models optimized for tasks like financial analysis, programming assistance, log analysis, and creative writing, with Redis-based queuing, PostgreSQL persistence, and comprehensive monitoring.
Technologies
Python, FastAPI, Redis, PostgreSQL, Prometheus, Docker, NVIDIA CUDA, Llama.cpp, GPU Memory Management, Model Management, Queue Management, API Design, Microservices Architecture, Request Routing, Performance Monitoring, Data Persistence, Concurrent Processing, Error Handling, Structured Logging, Configuration Management, Containerization, Load Balancing, Intelligent AI Routing, Enterprise Deployment
- Model Types Supported
- 6 Specialized
- Concurrent Requests
- 30+ Max
- API Response Time
- < 2 seconds
- GPU Memory Efficiency
- Optimized
LlamaLoadBalancer
A sophisticated AI model load balancer and API server that provides intelligent routing and management of multiple Llama language models for specialized tasks.
I built this while at Petco
Overview
LlamaLoadBalancer is a FastAPI-based application that serves as an intelligent load balancer for multiple Llama language models. It provides a unified API interface for accessing various specialized AI models, each optimized for different types of tasks such as financial analysis, programming assistance, log analysis, and creative writing.
Features
🤖 Multiple Specialized AI Models
- Finance Model: Specialized for financial data analysis and transaction classification
- Programmer Model: Expert coding assistant for software development tasks
- Log Analysis Models: Security-focused log parsing and threat detection
- Creative Models: Various creative and role-playing AI personas
- Question Answering: Advanced Q&A capabilities with structured responses
🔄 Intelligent Load Balancing
- Redis-based Queue Management: Efficient request queuing and processing
- Model Hot-swapping: Dynamic model loading/unloading based on demand
- Request Prioritization: Smart request handling and timeout management
- Concurrent Processing: Multi-threaded request processing
📊 Monitoring & Observability
- Prometheus Integration: Built-in metrics and monitoring
- Structured Logging: Comprehensive logging with request tracking
- Request History: Full audit trail of all requests and responses
- Performance Metrics: Response time tracking and optimization
🗄️ Data Persistence
- PostgreSQL Integration: Persistent storage of requests and responses
- Request Tracking: Complete request lifecycle management
- Response Caching: Efficient response storage and retrieval
Architecture
Core Components
Model Manager (`app/models/model_manager.py`)
- Singleton pattern for centralized model management
- Dynamic model loading with GPU optimization
- Specialized prompt engineering for each model type
- Memory management and garbage collection
Queue Manager (`app/queues/queue_manager.py`)
- Redis-based request queuing system
- Background processing with threading
- Request deduplication and status tracking
- Timeout and retry mechanisms
API Layer (`app/requests/`)
- v1_completions: Standard completion API
- v2_completions: Enhanced completion API with additional features
- v1_intel_chain: Advanced reasoning chain for complex tasks
- RESTful endpoints with proper error handling
Supported Model Types
| Model Type | Purpose | Context Size | Special Features |
|---|---|---|---|
| `FINANCE` | Financial analysis | 30K tokens | JSON structured responses |
| `PROGRAMMER` | Code generation | 30K tokens | Multi-language support |
| `LOGGEN` | Log analysis | 30K tokens | Security threat detection |
| `HIGHERFORM` | General AI | 30K tokens | Flexible response formats |
| `QUEENBASED` | Game strategy | 20K tokens | Tactical decision making |
| `METALLAMA3C` | General purpose | 8K tokens | Fast inference |
API Usage
Basic Completion Request
```bash curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ -d '{ "model": "PROGRAMMER", "messages": [ {"role": "user", "content": "Write a Python function to calculate fibonacci numbers"} ], "temperature": 0.7, "max_tokens": 1000 }' ```
Streaming Response
```bash curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ -d '{ "model": "FINANCE", "messages": [ {"role": "user", "content": "Analyze this transaction data"} ], "stream": true }' ```
GET Request (Simple)
```bash curl "http://localhost:8000/v1/completions?prompt=Hello&prompt=model=PROGRAMMER" ```
Installation & Setup
Prerequisites
- GPU: NVIDIA GPU with CUDA 11.7+ support
- Docker: For containerized deployment
- Redis: For request queuing
- PostgreSQL: For data persistence
Environment Configuration
Create a `.env` file:
```env REDIS_HOST=localhost REDIS_PORT=6380 DATABASE_URL=postgresql://user:password@localhost:5433/mydatabase ```
Docker Deployment
-
Build the container: ```bash docker build -t llama-loadbalancer . ```
-
Run with GPU support: ```bash docker run -d \ --name llama-loadbalancer \ --gpus all \ -p 8000:8000 \ -e REDIS_HOST=your-redis-host \ -e DATABASE_URL=your-postgres-url \ llama-loadbalancer ```
Local Development
-
Install dependencies: ```bash pip install -r requirements.txt ```
-
Set up Redis and PostgreSQL
-
Run the application: ```bash uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload ```
Configuration
Model Configuration
Models are configured in `app/models/model_manager.py`:
```python model_paths = { "PROGRAMMER": "/path/to/programming-model.gguf", "FINANCE": "/path/to/finance-model.gguf", # ... other models } ```
Queue Configuration
Adjust queue settings in `app/queues/queue_manager.py`:
```python processing_threshold = 30 # Maximum concurrent requests retry_limit = 3 # Retry attempts timeout = 300 # Request timeout in seconds ```
Use Cases
🏢 Enterprise Applications
- Customer Service: Intelligent chatbots with specialized knowledge
- Content Generation: Automated content creation for marketing
- Code Review: Automated code analysis and suggestions
- Financial Analysis: Transaction classification and risk assessment
🔒 Security & Monitoring
- Log Analysis: Real-time security threat detection
- Anomaly Detection: Unusual pattern identification
- Compliance Monitoring: Automated compliance checking
🎮 Gaming & Entertainment
- NPC AI: Intelligent non-player characters
- Story Generation: Dynamic content creation
- Game Strategy: AI-powered game mechanics
📚 Education & Research
- Tutoring Systems: Personalized learning assistance
- Research Assistance: Literature review and analysis
- Language Learning: Conversational AI tutors
Performance Optimization
GPU Memory Management
- Models are automatically loaded/unloaded based on demand
- Memory-efficient inference with optimized parameters
- Garbage collection after model operations
Request Optimization
- Redis-based queuing prevents system overload
- Request batching for improved throughput
- Intelligent model selection based on task requirements
Monitoring & Scaling
- Prometheus metrics for performance monitoring
- Horizontal scaling support with Redis clustering
- Load balancing across multiple GPU instances
Contributing
- Fork the repository
- Create a feature branch
- Add comprehensive tests
- Submit a pull request with detailed description
License
This project is licensed under the MIT License - see the LICENSE file for details.
Support
For support and questions:
- Create an issue in the GitHub repository
- Check the documentation for common solutions
- Review the example configurations
Roadmap
- Support for additional model architectures
- Enhanced streaming capabilities
- Kubernetes deployment templates
- Advanced caching mechanisms
- Multi-language model support
- Enhanced security features