Voice Master - Advanced AI Voice Assistant
Sophisticated multi-modal voice assistant with Whisper.cpp, Piper TTS, and advanced computer command processing. Features multi-model AI integration, project management, specialized agents, and high-performance audio processing with CUDA acceleration.
Status: in-progress · 2025-10-05
Overview
Voice Master is a comprehensive voice assistant platform that combines cutting-edge speech recognition, multiple AI models, computer automation, and intelligent project management. The system features both a Python implementation for basic functionality and an advanced C++ implementation with extensive enterprise-level features.
Technologies
C++, Python, CUDA, PortAudio, Whisper.cpp, Piper TTS, LocalAI, OpenAI API, MCP (Model Context Protocol), JSON Processing, Multi-threading, Async Processing, CMake, ONNX Runtime, WebRTC VAD, SoundDevice, NumPy, SciPy, Ring Buffer Audio, Real-time Processing, Plugin Architecture
- Speech Recognition Speed
- 10x real-time
- TTS Generation Speed
- Near real-time
- Memory Efficiency
- < 2GB RAM
- Audio Latency
- < 100ms
Voice Master - Advanced AI Voice Assistant
A sophisticated, multi-modal voice assistant system with advanced AI integration, computer command processing, project management, and extensive customization capabilities. Built for high-performance Linux systems with GPU acceleration.
🎯 Project Overview
Voice Master is a comprehensive voice assistant platform that combines cutting-edge speech recognition, multiple AI models, computer automation, and intelligent project management. The system features both a Python implementation for basic functionality and an advanced C++ implementation with extensive enterprise-level features.
Key Features
Core Voice Assistant Capabilities
- Real-time Speech Recognition: Powered by Whisper.cpp with CUDA GPU acceleration
- Multi-Model AI Integration: Support for multiple AI models including DeepSeek, Claude, and local models
- Streaming Text-to-Speech: High-quality voice synthesis with Piper TTS
- Keyword Detection: Configurable activation phrases with custom prompts
- Audio Feedback System: Star Trek-inspired chime system for user feedback
Advanced Computer Command System
- Inline Command Processing: Extract and execute computer commands from speech
- Mode-Based Tool Filtering: Different tool sets for project, idea, stock, and notation modes
- Send Mode: Conversation buffering for complex multi-turn interactions
- Agent Switching: Dynamic model switching mid-conversation
- Command History Management: Undo, repeat, and conversation manipulation
Project Management Integration
- Project Workspaces: Organized project-specific conversation storage
- Artifact Generation: Automatic README.md and documentation creation
- Task Management: Integration with project planning and todo systems
- Context Preservation: Maintain conversation context across sessions
Specialized Agents
- Idea Agent: Advanced ideation and brainstorming capabilities
- Stock Agent: Financial analysis and market intelligence
- Project Agent: Development task management and code generation
- Research Agent: Web search and content analysis
Technical Architecture
- MCP Integration: Model Context Protocol for tool integration
- Async Processing: Multi-threaded audio processing and transcription
- Memory System: Long-term conversation persistence and retrieval
- Plugin Architecture: Extensible tool and agent system
🚀 Current Implementation Status
Python Implementation (voice_assistant/)
Status: ✅ Production Ready
- Basic voice assistant with keyword detection
- Whisper.cpp integration with GPU acceleration
- Piper TTS for high-quality speech synthesis
- LocalAI server integration
- Real-time audio processing
C++ Implementation (cpp_voice_assistant/)
Status: 🔄 Active Development
- Advanced computer command processor
- Multi-agent system with mode switching
- Project management and artifact generation
- MCP server integration
- Comprehensive audio management system
📋 Development Roadmap
Immediate Priorities (High Priority)
- Main.cpp Refactoring: Break down 1,467-line monolithic file into modular components
- AudioManager Class: Ring buffer management, VAD, and audio I/O
- TranscriptionProcessor: Async Whisper integration with worker threads
- ConversationManager: State tracking and phrase detection
- SystemOrchestrator: Main application coordination class
Phase 1: Tool System Enhancement
- Remove legacy tools (file_operations, audio_control, system_info)
- Enhance web search capabilities with content fetching
- Implement website content extraction and summarization
- Add search result filtering and ranking
Phase 2: Computer Commands System
- Implement inline command parsing (
computer: truncate last) - Send mode for conversation buffering
- Dynamic agent switching mid-conversation
- Advanced command processing pipeline
Phase 3: Project Mode
- Project workspace creation and management
- Project-specific conversation storage
- Automatic artifact generation (README.md, specs)
- Project metadata and tagging system
Phase 4: Idea Agent & Memory
- Specialized conversational AI for ideation
- Long-term conversation persistence
- Conversation search and retrieval
- Idea categorization and relationship mapping
Phase 5: Advanced Features
- Multi-search engine support
- Workflow automation and custom commands
- Voice interface improvements
- Enhanced web integration
🏗️ Architecture Overview
System Components
┌─────────────────────────────────────────────────────────────────┐
│ Voice Master System │
├─────────────────────────────────────────────────────────────────┤
│ ┌─────────────┐ ┌──────────────┐ ┌─────────────────────┐ │
│ │ AudioManager│ │Transcription │ │ ConversationManager │ │
│ │ - Ring buf │ │Processor │ │ - State tracking │ │
│ │ - VAD │ │ - Whisper │ │ - Phrase detection │ │
│ │ - Audio I/O │ │ - Async proc │ │ - Model activation │ │
│ └─────────────┘ └──────────────┘ └─────────────────────┘ │
├─────────────────────────────────────────────────────────────────┤
│ ┌─────────────────────┐ ┌─────────────────┐ ┌─────────────┐ │
│ │ComputerCommandProc │ │ LLMService │ │ProjectManager│ │
│ │ - Command parsing │ │ - Multi-model │ │ - Workspaces │ │
│ │ - Mode filtering │ │ - Agent switch │ │ - Artifacts │ │
│ │ - Audio feedback │ │ - Context mgmt │ │ - Metadata │ │
│ └─────────────────────┘ └─────────────────┘ └─────────────┘ │
├─────────────────────────────────────────────────────────────────┤
│ ┌──────────────┐ ┌─────────────┐ ┌──────────────────────┐ │
│ │MCP Manager │ │Audio Player │ │ SystemOrchestrator │ │
│ │ - Tool intg │ │ - TTS coord │ │ - Main coordinator │ │
│ │ - Plugin sys │ │ - Chime sys │ │ - Component mgmt │ │
│ └──────────────┘ └─────────────┘ └──────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
Technology Stack
Core Technologies
- Speech Recognition: Whisper.cpp with CUDA acceleration
- Text-to-Speech: Piper TTS with ONNX Runtime GPU support
- AI Models: LocalAI server with multiple model support
- Audio Processing: PortAudio with custom ring buffer implementation
- Build System: CMake with GPU-optimized compilation
Python Implementation
- Audio: sounddevice, numpy, scipy
- AI Integration: OpenAI client library
- VAD: webrtcvad for voice activity detection
- Async Processing: threading and queue modules
C++ Implementation
- Audio: PortAudio C API
- JSON: nlohmann/json
- HTTP: Custom OpenAI client implementation
- Async: std::thread and std::atomic
- MCP: Model Context Protocol integration
🛠️ Installation & Setup
Prerequisites
Hardware Requirements
- Linux (Arch Linux/Manjaro recommended)
- NVIDIA GPU with CUDA 12.3+ support
- Microphone and speakers
- Minimum 16GB RAM (32GB+ recommended)
Software Dependencies
- System: portaudio, ffmpeg, cmake, git
- Python: 3.12+ with uv package manager
- CUDA: 12.3+ with nvcc compiler
- LocalAI: Running at
http://localhost:8080/v1
Python Implementation Setup
-
Install System Dependencies
sudo pacman -S portaudio ffmpeg cmake git uv -
Install Python Dependencies
uv pip install sounddevice==0.4.6 openai==1.0.0 numpy==1.26.0 webrtcvad==2.0.10 onnxruntime-gpu -
Build Whisper.cpp with CUDA
git clone https://github.com/ggerganov/whisper.cpp.git cd whisper.cpp make WHISPER_CUDA=1 ./models/download-ggml-model.sh base.en cd bindings/python && uv pip install . -
Setup Piper TTS
git clone https://github.com/rhasspy/piper.git cd piper/src/python uv pip install . # Download voice models from Piper releases
C++ Implementation Setup
-
Install C++ Dependencies
sudo pacman -S cmake nlohmann-json portaudio -
Build the Project
cd cpp_voice_assistant mkdir build && cd build cmake .. -DCMAKE_BUILD_TYPE=Release -DWHISPER_CUDA=ON make -j$(nproc)
⚙️ Configuration
Python Configuration (voice_assistant/config.json)
{
"keywords": [
{
"phrase": "hey assistant",
"prompt_template": "You are a helpful assistant. Respond to: {transcription}",
"model": "deepseek-r1-distill-llama-8b-Large"
}
],
"audio_settings": {
"sample_rate": 44100,
"channels": 1,
"frame_duration": 30,
"input_device": 0
},
"localai_endpoint": "http://localhost:8080/v1",
"models": {
"whisper": {
"model_path": "models/whisper_base.en"
},
"piper": {
"model_path": "models/piper_voice",
"voice": "en-us-kathleen-low"
}
}
}
C++ Configuration
Configuration is handled through the config/ directory with JSON files for different components and agents.
🎮 Usage
Basic Python Usage
cd voice_assistant
python main.py
Advanced C++ Usage
cd cpp_voice_assistant/build
./voice_assistant
Available Commands
Computer Commands
computer: truncate last- Remove last message from conversationcomputer: send mode- Switch to manual send modecomputer: send- Send buffered conversation to LLMcomputer: change agent- Switch between AI modelscomputer: clear conversation- Reset conversation historycomputer: repeat last- Repeat last AI response
Mode Commands
project mode- Switch to project management modeidea mode- Switch to creative ideation modestock mode- Switch to financial analysis modenotation mode- Switch to free-form dictation mode
🔧 Troubleshooting
Audio Issues
- Check microphone with
pactl list sources - Verify audio device selection in configuration
- Ensure PortAudio is properly installed
GPU/CUDA Issues
- Verify CUDA installation with
nvcc --version - Check GPU availability with
nvidia-smi - Ensure compatible driver versions
LocalAI Issues
- Verify LocalAI is running:
curl http://localhost:8080/v1/models - Check model loading status
- Verify endpoint configuration
🤝 Contributing
This project is actively developed with the following contribution guidelines:
- Code Style: Follow existing patterns and conventions
- Testing: Add tests for new features
- Documentation: Update README and inline documentation
- Architecture: Maintain clean separation of concerns
Development Workflow
- Check the TODO.md for current priorities
- Create feature branches from main
- Implement features following the established patterns
- Add comprehensive tests
- Update documentation
- Submit pull requests with detailed descriptions
📚 Documentation
- TODO.md: Comprehensive development roadmap and task tracking
- ProjectPlan.md: Detailed implementation planning and architecture
- Component Documentation: Inline documentation in source files
- Configuration Guides: Detailed setup and configuration instructions
🔐 Security & Privacy
- Local Processing: All AI inference happens locally
- No External Dependencies: No cloud services or external APIs required
- Configurable Endpoints: All connections are configurable
- Audio Privacy: Audio processing happens locally with no external transmission
📈 Performance
Benchmarks
- Transcription Speed: ~10x real-time with GPU acceleration
- TTS Generation: Near real-time streaming synthesis
- Memory Usage: Optimized for systems with 16GB+ RAM
- CPU Usage: Multi-threaded processing for low latency
Optimization Features
- CUDA GPU acceleration for Whisper and TTS
- Ring buffer audio processing for low latency
- Async transcription processing
- Memory-mapped model loading
- Optimized audio resampling
🌟 Future Vision
Voice Master aims to become the most advanced open-source voice assistant platform, featuring:
- Multi-modal Interface: Voice, text, and gesture integration
- Advanced AI Agents: Specialized agents for different domains
- Enterprise Integration: Plugin system for business applications
- Extensible Architecture: Easy addition of new capabilities
- Cross-platform Support: Linux, Windows, and macOS compatibility
📄 License
This project is open source under the MIT License. See individual component licenses for third-party dependencies.
🙏 Acknowledgments
- Whisper.cpp: Georgi Gerganov's incredible speech recognition system
- Piper TTS: High-quality text-to-speech synthesis
- LocalAI: Local AI inference server
- PortAudio: Cross-platform audio I/O
- CUDA: NVIDIA's GPU computing platform
Voice Master - Empowering human-AI interaction through advanced voice technology and intelligent automation.
This comprehensive voice assistant platform revolutionizes human-AI interaction through advanced speech recognition, multi-modal AI integration, and intelligent automation designed for enterprise and personal productivity.