FluentD LLM Server - Intelligent Log Analysis System
End-to-end log analysis solution combining FluentD log aggregation with Large Language Models (LLMs) for real-time security threat detection. Features custom-trained binary classification models using Microsoft Phi-2 for automated incident response and system anomaly detection.
Status: in-progress · 2024-11-01
Overview
FluentdLLM is an end-to-end log analysis solution that combines FluentD log aggregation with Large Language Models to automatically detect security threats and system anomalies in real-time. The project leverages custom-trained binary classification models to flag suspicious log entries and provides the foundation for automated incident response.
Technologies
Python, PyTorch, Hugging Face Transformers, FluentD, Ruby, Docker, Kubernetes, Microsoft Phi-2, GGUF, CUDA, Distributed Training, Connection Pooling, REST API, OpenAI Compatible API, Batch Processing, Error Handling, Logging, Configuration Management, Model Quantization, LoRA Fine-tuning
- Training Loss
- 0.2883
- Model Size (Quantized)
- 2.7GB
- Batch Processing
- 1-1000 logs
- Response Time
- < 1s
FluentD LLM Server
An intelligent log analysis system that combines FluentD log aggregation with Large Language Models (LLMs) to automatically detect security threats and system anomalies in real-time. The project leverages custom-trained binary classification models to flag suspicious log entries and provides the foundation for automated incident response.
What This Project Does
FluentdLLM is an end-to-end log analysis solution that:
- Collects logs from multiple sources across cloud and on-premises infrastructure using FluentD agents
- Analyzes logs in real-time using a lightweight, custom-trained binary classification model based on Microsoft's Phi-2
- Flags suspicious activities automatically with high accuracy (trained to detect security threats, unauthorized access attempts, and system anomalies)
- Provides a foundation for automated incident response and code-level patch generation
- Scales efficiently in Kubernetes environments with proper resource management
How It Works
Core Architecture
- Log Collection Layer: FluentD agents deployed across infrastructure collect and forward logs
- Intelligent Processing: Custom FluentD plugin sends logs to an LLM server for real-time analysis
- Binary Classification: Fine-tuned Phi-2 model (1.5B parameters) classifies logs as normal (0) or threat (1)
- Action Pipeline: Flagged logs trigger downstream processing for incident response
Technical Implementation
[Log Sources] → [FluentD Agents] → [FluentD Server + Custom Plugin] → [LLM Server] → [Classification] → [Response Actions]
The system uses:
- Custom FluentD Plugin (
out_llm.rb): Interfaces with LLM server using OpenAI-compatible API - Fine-tuned Phi-2 Model: Binary classifier trained on security-relevant log data
- GGUF Format: Optimized model format for efficient inference
- Connection Pooling: Efficient batch processing with retry mechanisms
Current Implementation Status
✅ Completed Components
1. FluentD Integration (Phase 1a - Complete)
- Custom FluentD Plugin: Production-ready
out_llm.rbplugin with OpenAI API compatibility - Connection Management: Connection pooling, retry mechanisms, and batch processing
- Configuration: Flexible configuration with adjustable batch sizes, timeouts, and logging levels
- Testing Environment: Complete Docker-based test setup with log generators
2. Model Training and Deployment (Phase 1b - Complete)
- Base Model: Microsoft Phi-2 fine-tuned for binary log classification
- Training Infrastructure:
- Distributed training across 3x RTX 3060 GPUs (12GB VRAM each)
- 4-bit quantization using BitsAndBytesConfig for memory efficiency
- LoRA (Low-Rank Adaptation) for parameter-efficient fine-tuning
- Training Results:
- Final training loss: 0.2883 (stable convergence)
- Model size: ~2.7GB (quantized from original ~5.4GB)
- Inference speed: Optimized for real-time log processing
- Model Formats: Multiple output formats including GGUF for production deployment
3. Training Data Pipeline
- Data Generation: Automated log generation scripts for testing
- Data Processing: Comprehensive data cleaning and preprocessing utilities
- Training Datasets: Multiple iterations of training data with security-focused examples
- Validation: Model testing utilities for classification accuracy verification
🔄 In Progress Components
4. LLM Server Integration (Phase 1c/1d)
- Model deployment to production LLM server
- Integration testing with FluentD plugin
- Performance optimization for high-throughput log processing
📋 Planned Components (Future Phases)
Phase 2: Advanced Processing Pipeline
- Vector Database: ChromaDB integration for contextual code analysis
- Agent System: Automated error analysis and code patch generation
- GitLab Integration: Repository analysis and automated pull request creation
- Notification System: Multi-channel alerting (Teams, email, etc.)
Phase 3: Production Deployment
- Kubernetes Deployment: Auto-scaling microservices architecture
- Monitoring: Prometheus/Grafana integration for system observability
- Security: mTLS communication and access controls
Ways to Use This Project
1. Real-Time Security Monitoring
Deploy FluentdLLM as a security monitoring solution:
# Start the complete test environment
./test-environment.sh start
# Monitor security events in real-time
./test-environment.sh llm-logs
Use Cases:
- Detect unauthorized access attempts
- Identify suspicious API calls or SQL injection attempts
- Monitor for credential harvesting attempts
- Flag unusual system behavior patterns
2. Development and Training
Use the training pipeline to create custom models:
# Train a new model with your data
cd llm_server/internal/train_model/
python TrainModelBinary.py --json_file your_data.jsonl --model_name microsoft/phi-2 --output_dir ./custom_model
# Test the trained model
python test_model_binary_hf.py
3. Integration with Existing Infrastructure
Integrate FluentdLLM into existing log processing pipelines:
- FluentD Configuration: Use the custom
out_llmplugin in your FluentD setup - API Integration: Connect to the LLM server using OpenAI-compatible REST API
- Batch Processing: Configure batch sizes and processing intervals based on log volume
4. Research and Development
The project provides a foundation for log analysis research:
- Model Experimentation: Try different base models (Phi-2, Llama, etc.)
- Training Data: Use the data generation utilities to create domain-specific datasets
- Performance Analysis: Benchmark different quantization and optimization techniques
Technical Architecture
Core Components
FluentD Plugin (fluentd/plugin/lib/fluent/plugin/out_llm.rb)
- Language: Ruby with Faraday HTTP client and connection pooling
- Features:
- OpenAI-compatible API integration
- Configurable batch processing (default: 100 logs per batch)
- Exponential backoff retry mechanism (up to 3 retries)
- Connection pooling for efficient resource usage
- Comprehensive logging and error handling
Model Training Pipeline (llm_server/internal/train_model/)
- Base Model: Microsoft Phi-2 (2.7B parameters)
- Training Framework: PyTorch with Hugging Face Transformers
- Optimization Techniques:
- 4-bit quantization with BitsAndBytesConfig
- LoRA (Low-Rank Adaptation) for parameter-efficient fine-tuning
- Distributed training across multiple GPUs
- Gradient checkpointing for memory efficiency
Data Processing Utilities
- Log Generation: Realistic log entry generation for testing (
docker/log-generator/) - Data Cleaning: Unicode normalization and control character removal
- Format Conversion: Tools for converting between training data formats
- Model Testing: Comprehensive testing utilities for model validation
Directory Structure
FluentdLLM/
├── fluentd/ # FluentD plugin and configuration
│ ├── plugin/lib/fluent/plugin/out_llm.rb # Custom LLM output plugin
│ └── config/test.conf # FluentD configuration
├── llm_server/internal/ # Model training and processing
│ ├── train_model/ # Training pipeline and utilities
│ │ ├── TrainModelBinary.py # Main training script
│ │ ├── test_model_binary_hf.py # Model testing utility
│ │ └── model_out/ # Trained model outputs
│ └── training_data/ # Training datasets
├── docker/ # Docker configurations
│ ├── fluentd/Dockerfile # FluentD container setup
│ └── log-generator/ # Test log generation
├── docker-compose.yml # Complete test environment
└── test-environment.sh # Test automation script
Performance Characteristics
Model Performance
- Training Loss: 0.2883 (final, stable convergence)
- Model Size: ~2.7GB (4-bit quantized from 5.4GB original)
- Inference Speed: Optimized for real-time processing
- Memory Usage: <6GB VRAM per GPU during training
- Accuracy: Trained on security-focused log classification tasks
System Performance
- Batch Processing: Configurable batch sizes (1-1000 logs)
- Throughput: Designed for high-volume log processing
- Latency: Sub-second response times for batch classification
- Scalability: Kubernetes-ready architecture
Getting Started
Prerequisites
- Hardware: GPU with 6GB+ VRAM recommended (tested on 3x RTX 3060 12GB)
- Software: Docker, Docker Compose, Python 3.8+
- Development Environment: Linux/Unix (tested on Arch Linux/Manjaro)
Quick Start
-
Clone and Setup:
git clone <repository-url> cd FluentdLLM -
Run Test Environment:
# Start the complete test environment chmod +x test-environment.sh ./test-environment.sh start # View logs and monitor system ./test-environment.sh fluentd-logs ./test-environment.sh llm-logs -
Train Custom Model (optional):
cd llm_server/internal/train_model/ # Install dependencies pip install -r requirements.txt # Train with your data python TrainModelBinary.py \ --json_file your_training_data.jsonl \ --model_name microsoft/phi-2 \ --output_dir ./custom_model
Configuration
FluentD Plugin Configuration
Edit fluentd/config/test.conf:
<match system.logs>
@type llm
llm_server_url http://your-llm-server:8080/v1
batch_size 50 # Logs per batch
flush_interval 10s # Processing frequency
max_retry 3 # Retry attempts
temperature 0.1 # Model creativity (lower = more deterministic)
max_tokens 50 # Response length limit
</match>
Model Training Parameters
Key training parameters in TrainModelBinary.py:
- Batch Size: 16 (adjustable based on available VRAM)
- Learning Rate: 1e-5 (with cosine annealing)
- Epochs: 15 (with early stopping based on loss)
- LoRA Rank: 16 (balance between efficiency and performance)
- Quantization: 4-bit NF4 for memory efficiency
Development Roadmap
✅ Completed (Phase 1)
- FluentD Plugin: Production-ready with comprehensive error handling
- Model Training: Fine-tuned Phi-2 model with 0.2883 final loss
- Testing Infrastructure: Complete Docker-based test environment
- Data Pipeline: Training data generation and processing utilities
🔄 Current Focus (Phase 1c/1d)
- LLM Server Integration: Deploy model to production inference server
- Performance Optimization: Optimize batch processing and response times
- Integration Testing: End-to-end validation with real log data
📋 Upcoming (Phase 2-5)
- Vector Database: ChromaDB integration for contextual analysis
- Agent Systems: Automated code analysis and patch generation
- Kubernetes Deployment: Production-ready microservices architecture
- Monitoring: Prometheus/Grafana observability stack
- Security: mTLS, access controls, and audit logging
Technical Specifications
Development Environment
- OS: Arch Linux (Manjaro) with Zsh + Anaconda
- Hardware: 3x RTX 3060 (12GB VRAM each), Ryzen 5 CPU, 128GB RAM
- GPU Acceleration: CUDA with distributed training support
- Model Format: Hugging Face → GGUF conversion pipeline
Integration Points
- FluentD: Custom Ruby plugin with OpenAI API compatibility
- LLM Server: RESTful API compatible with OpenAI chat completions
- Training Pipeline: PyTorch with Hugging Face Transformers
- Deployment: Docker containers with Kubernetes manifests (planned)
This comprehensive log analysis system revolutionizes security monitoring through intelligent real-time threat detection, advanced machine learning integration, and scalable infrastructure designed for enterprise-grade log processing and automated incident response.