Levi DeHaan

Voice Master - Advanced AI Voice Assistant

Sophisticated multi-modal voice assistant with Whisper.cpp, Piper TTS, and advanced computer command processing. Features multi-model AI integration, project management, specialized agents, and high-performance audio processing with CUDA acceleration.

Status: in-progress · 2025-10-05

Overview

Voice Master is a comprehensive voice assistant platform that combines cutting-edge speech recognition, multiple AI models, computer automation, and intelligent project management. The system features both a Python implementation for basic functionality and an advanced C++ implementation with extensive enterprise-level features.

Technologies

C++, Python, CUDA, PortAudio, Whisper.cpp, Piper TTS, LocalAI, OpenAI API, MCP (Model Context Protocol), JSON Processing, Multi-threading, Async Processing, CMake, ONNX Runtime, WebRTC VAD, SoundDevice, NumPy, SciPy, Ring Buffer Audio, Real-time Processing, Plugin Architecture

Speech Recognition Speed
10x real-time
TTS Generation Speed
Near real-time
Memory Efficiency
< 2GB RAM
Audio Latency
< 100ms

Voice Master - Advanced AI Voice Assistant

A sophisticated, multi-modal voice assistant system with advanced AI integration, computer command processing, project management, and extensive customization capabilities. Built for high-performance Linux systems with GPU acceleration.

🎯 Project Overview

Voice Master is a comprehensive voice assistant platform that combines cutting-edge speech recognition, multiple AI models, computer automation, and intelligent project management. The system features both a Python implementation for basic functionality and an advanced C++ implementation with extensive enterprise-level features.

Key Features

Core Voice Assistant Capabilities

  • Real-time Speech Recognition: Powered by Whisper.cpp with CUDA GPU acceleration
  • Multi-Model AI Integration: Support for multiple AI models including DeepSeek, Claude, and local models
  • Streaming Text-to-Speech: High-quality voice synthesis with Piper TTS
  • Keyword Detection: Configurable activation phrases with custom prompts
  • Audio Feedback System: Star Trek-inspired chime system for user feedback

Advanced Computer Command System

  • Inline Command Processing: Extract and execute computer commands from speech
  • Mode-Based Tool Filtering: Different tool sets for project, idea, stock, and notation modes
  • Send Mode: Conversation buffering for complex multi-turn interactions
  • Agent Switching: Dynamic model switching mid-conversation
  • Command History Management: Undo, repeat, and conversation manipulation

Project Management Integration

  • Project Workspaces: Organized project-specific conversation storage
  • Artifact Generation: Automatic README.md and documentation creation
  • Task Management: Integration with project planning and todo systems
  • Context Preservation: Maintain conversation context across sessions

Specialized Agents

  • Idea Agent: Advanced ideation and brainstorming capabilities
  • Stock Agent: Financial analysis and market intelligence
  • Project Agent: Development task management and code generation
  • Research Agent: Web search and content analysis

Technical Architecture

  • MCP Integration: Model Context Protocol for tool integration
  • Async Processing: Multi-threaded audio processing and transcription
  • Memory System: Long-term conversation persistence and retrieval
  • Plugin Architecture: Extensible tool and agent system

🚀 Current Implementation Status

Python Implementation (voice_assistant/)

Status: ✅ Production Ready

  • Basic voice assistant with keyword detection
  • Whisper.cpp integration with GPU acceleration
  • Piper TTS for high-quality speech synthesis
  • LocalAI server integration
  • Real-time audio processing

C++ Implementation (cpp_voice_assistant/)

Status: 🔄 Active Development

  • Advanced computer command processor
  • Multi-agent system with mode switching
  • Project management and artifact generation
  • MCP server integration
  • Comprehensive audio management system

📋 Development Roadmap

Immediate Priorities (High Priority)

  • Main.cpp Refactoring: Break down 1,467-line monolithic file into modular components
  • AudioManager Class: Ring buffer management, VAD, and audio I/O
  • TranscriptionProcessor: Async Whisper integration with worker threads
  • ConversationManager: State tracking and phrase detection
  • SystemOrchestrator: Main application coordination class

Phase 1: Tool System Enhancement

  • Remove legacy tools (file_operations, audio_control, system_info)
  • Enhance web search capabilities with content fetching
  • Implement website content extraction and summarization
  • Add search result filtering and ranking

Phase 2: Computer Commands System

  • Implement inline command parsing (computer: truncate last)
  • Send mode for conversation buffering
  • Dynamic agent switching mid-conversation
  • Advanced command processing pipeline

Phase 3: Project Mode

  • Project workspace creation and management
  • Project-specific conversation storage
  • Automatic artifact generation (README.md, specs)
  • Project metadata and tagging system

Phase 4: Idea Agent & Memory

  • Specialized conversational AI for ideation
  • Long-term conversation persistence
  • Conversation search and retrieval
  • Idea categorization and relationship mapping

Phase 5: Advanced Features

  • Multi-search engine support
  • Workflow automation and custom commands
  • Voice interface improvements
  • Enhanced web integration

🏗️ Architecture Overview

System Components

┌─────────────────────────────────────────────────────────────────┐
│                    Voice Master System                          │
├─────────────────────────────────────────────────────────────────┤
│  ┌─────────────┐  ┌──────────────┐  ┌─────────────────────┐    │
│  │ AudioManager│  │Transcription │  │ ConversationManager │    │
│  │ - Ring buf  │  │Processor     │  │ - State tracking    │    │
│  │ - VAD       │  │ - Whisper     │  │ - Phrase detection  │    │
│  │ - Audio I/O │  │ - Async proc  │  │ - Model activation  │    │
│  └─────────────┘  └──────────────┘  └─────────────────────┘    │
├─────────────────────────────────────────────────────────────────┤
│  ┌─────────────────────┐  ┌─────────────────┐  ┌─────────────┐  │
│  │ComputerCommandProc  │  │   LLMService    │  │ProjectManager│  │
│  │ - Command parsing   │  │ - Multi-model   │  │ - Workspaces │  │
│  │ - Mode filtering    │  │ - Agent switch  │  │ - Artifacts  │  │
│  │ - Audio feedback    │  │ - Context mgmt  │  │ - Metadata   │  │
│  └─────────────────────┘  └─────────────────┘  └─────────────┘  │
├─────────────────────────────────────────────────────────────────┤
│  ┌──────────────┐  ┌─────────────┐  ┌──────────────────────┐   │
│  │MCP Manager   │  │Audio Player │  │  SystemOrchestrator  │   │
│  │ - Tool intg  │  │ - TTS coord │  │  - Main coordinator   │   │
│  │ - Plugin sys │  │ - Chime sys │  │  - Component mgmt     │   │
│  └──────────────┘  └─────────────┘  └──────────────────────┘   │
└─────────────────────────────────────────────────────────────────┘

Technology Stack

Core Technologies

  • Speech Recognition: Whisper.cpp with CUDA acceleration
  • Text-to-Speech: Piper TTS with ONNX Runtime GPU support
  • AI Models: LocalAI server with multiple model support
  • Audio Processing: PortAudio with custom ring buffer implementation
  • Build System: CMake with GPU-optimized compilation

Python Implementation

  • Audio: sounddevice, numpy, scipy
  • AI Integration: OpenAI client library
  • VAD: webrtcvad for voice activity detection
  • Async Processing: threading and queue modules

C++ Implementation

  • Audio: PortAudio C API
  • JSON: nlohmann/json
  • HTTP: Custom OpenAI client implementation
  • Async: std::thread and std::atomic
  • MCP: Model Context Protocol integration

🛠️ Installation & Setup

Prerequisites

Hardware Requirements

  • Linux (Arch Linux/Manjaro recommended)
  • NVIDIA GPU with CUDA 12.3+ support
  • Microphone and speakers
  • Minimum 16GB RAM (32GB+ recommended)

Software Dependencies

  • System: portaudio, ffmpeg, cmake, git
  • Python: 3.12+ with uv package manager
  • CUDA: 12.3+ with nvcc compiler
  • LocalAI: Running at http://localhost:8080/v1

Python Implementation Setup

  1. Install System Dependencies

    sudo pacman -S portaudio ffmpeg cmake git uv
    
  2. Install Python Dependencies

    uv pip install sounddevice==0.4.6 openai==1.0.0 numpy==1.26.0 webrtcvad==2.0.10 onnxruntime-gpu
    
  3. Build Whisper.cpp with CUDA

    git clone https://github.com/ggerganov/whisper.cpp.git
    cd whisper.cpp
    make WHISPER_CUDA=1
    ./models/download-ggml-model.sh base.en
    cd bindings/python && uv pip install .
    
  4. Setup Piper TTS

    git clone https://github.com/rhasspy/piper.git
    cd piper/src/python
    uv pip install .
    # Download voice models from Piper releases
    

C++ Implementation Setup

  1. Install C++ Dependencies

    sudo pacman -S cmake nlohmann-json portaudio
    
  2. Build the Project

    cd cpp_voice_assistant
    mkdir build && cd build
    cmake .. -DCMAKE_BUILD_TYPE=Release -DWHISPER_CUDA=ON
    make -j$(nproc)
    

⚙️ Configuration

Python Configuration (voice_assistant/config.json)

{
  "keywords": [
    {
      "phrase": "hey assistant",
      "prompt_template": "You are a helpful assistant. Respond to: {transcription}",
      "model": "deepseek-r1-distill-llama-8b-Large"
    }
  ],
  "audio_settings": {
    "sample_rate": 44100,
    "channels": 1,
    "frame_duration": 30,
    "input_device": 0
  },
  "localai_endpoint": "http://localhost:8080/v1",
  "models": {
    "whisper": {
      "model_path": "models/whisper_base.en"
    },
    "piper": {
      "model_path": "models/piper_voice",
      "voice": "en-us-kathleen-low"
    }
  }
}

C++ Configuration

Configuration is handled through the config/ directory with JSON files for different components and agents.

🎮 Usage

Basic Python Usage

cd voice_assistant
python main.py

Advanced C++ Usage

cd cpp_voice_assistant/build
./voice_assistant

Available Commands

Computer Commands

  • computer: truncate last - Remove last message from conversation
  • computer: send mode - Switch to manual send mode
  • computer: send - Send buffered conversation to LLM
  • computer: change agent - Switch between AI models
  • computer: clear conversation - Reset conversation history
  • computer: repeat last - Repeat last AI response

Mode Commands

  • project mode - Switch to project management mode
  • idea mode - Switch to creative ideation mode
  • stock mode - Switch to financial analysis mode
  • notation mode - Switch to free-form dictation mode

🔧 Troubleshooting

Audio Issues

  • Check microphone with pactl list sources
  • Verify audio device selection in configuration
  • Ensure PortAudio is properly installed

GPU/CUDA Issues

  • Verify CUDA installation with nvcc --version
  • Check GPU availability with nvidia-smi
  • Ensure compatible driver versions

LocalAI Issues

  • Verify LocalAI is running: curl http://localhost:8080/v1/models
  • Check model loading status
  • Verify endpoint configuration

🤝 Contributing

This project is actively developed with the following contribution guidelines:

  1. Code Style: Follow existing patterns and conventions
  2. Testing: Add tests for new features
  3. Documentation: Update README and inline documentation
  4. Architecture: Maintain clean separation of concerns

Development Workflow

  1. Check the TODO.md for current priorities
  2. Create feature branches from main
  3. Implement features following the established patterns
  4. Add comprehensive tests
  5. Update documentation
  6. Submit pull requests with detailed descriptions

📚 Documentation

  • TODO.md: Comprehensive development roadmap and task tracking
  • ProjectPlan.md: Detailed implementation planning and architecture
  • Component Documentation: Inline documentation in source files
  • Configuration Guides: Detailed setup and configuration instructions

🔐 Security & Privacy

  • Local Processing: All AI inference happens locally
  • No External Dependencies: No cloud services or external APIs required
  • Configurable Endpoints: All connections are configurable
  • Audio Privacy: Audio processing happens locally with no external transmission

📈 Performance

Benchmarks

  • Transcription Speed: ~10x real-time with GPU acceleration
  • TTS Generation: Near real-time streaming synthesis
  • Memory Usage: Optimized for systems with 16GB+ RAM
  • CPU Usage: Multi-threaded processing for low latency

Optimization Features

  • CUDA GPU acceleration for Whisper and TTS
  • Ring buffer audio processing for low latency
  • Async transcription processing
  • Memory-mapped model loading
  • Optimized audio resampling

🌟 Future Vision

Voice Master aims to become the most advanced open-source voice assistant platform, featuring:

  • Multi-modal Interface: Voice, text, and gesture integration
  • Advanced AI Agents: Specialized agents for different domains
  • Enterprise Integration: Plugin system for business applications
  • Extensible Architecture: Easy addition of new capabilities
  • Cross-platform Support: Linux, Windows, and macOS compatibility

📄 License

This project is open source under the MIT License. See individual component licenses for third-party dependencies.

🙏 Acknowledgments

  • Whisper.cpp: Georgi Gerganov's incredible speech recognition system
  • Piper TTS: High-quality text-to-speech synthesis
  • LocalAI: Local AI inference server
  • PortAudio: Cross-platform audio I/O
  • CUDA: NVIDIA's GPU computing platform

Voice Master - Empowering human-AI interaction through advanced voice technology and intelligent automation.

This comprehensive voice assistant platform revolutionizes human-AI interaction through advanced speech recognition, multi-modal AI integration, and intelligent automation designed for enterprise and personal productivity.