Levi DeHaan

STT Service for NVIDIA Jetson Orin

GPU-accelerated Faster-Whisper STT service with async queue, YouTube ingestion, and systemd ops on Jetson Orin.

By Levi DeHaan ·

STT Service for NVIDIA Jetson Orin

High-performance Speech-to-Text service optimized for NVIDIA Jetson Orin Developer Edition using Faster-Whisper with CUDA acceleration.

Why this matters

Building reliable, fast STT on edge devices is hard. This project delivers a production-ready service that:

  • Runs fully on-device with CUDA acceleration (Jetson Orin)
  • Handles YouTube ingestion, file uploads, and long-form audio
  • Persists jobs and results with an async queue and SQLite
  • Exposes a clean REST API you can call from any client
  • Monitors real-time GPU/CPU usage with a beautiful terminal UI

It’s designed for speed, resilience, and easy operations.

Features

  • GPU-Accelerated transcription (CUDA, float16)
  • YouTube video transcription with audio extraction
  • Asynchronous processing via job queue and IDs
  • Persistent storage in SQLite
  • Multiple audio formats: WAV, MP3, M4A, FLAC, OGG, OPUS, MP4, MKV, WEBM
  • RESTful API for upload, YouTube, status, and results
  • Performance metrics: processing time, confidence, segments
  • Real-time GPU/CPU monitoring with graphs and trend lines
  • Systemd integration and smart startup scripts

System requirements

  • NVIDIA Jetson Orin Developer Edition
  • CUDA 12.x
  • Ubuntu 20.04/22.04
  • Python 3.8+
  • 4GB+ available RAM

Quick start

1) Setup

# Clone and setup
git clone <repository-url> # I will update this document with this url when I get it cleaned up and pushed.
cd TTSsystem

# Full automated setup (recommended)
make setup

# OR manual setup
make install-deps
make install-python
make all

2) Start service

Option A — Systemd (recommended)

sudo cp stt-service.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable stt-service
sudo systemctl start stt-service

# Helper controls
./stt-service-control.sh status
./stt-service-control.sh logs
./stt-service-control.sh youtube <url>

Option B — Manual (development)

./start_service.sh         # Start on port 8080
./start_service.sh 3000    # Custom port
./start_service.sh restart # Restart
./start_service.sh stop    # Stop

3) Monitor GPU/CPU

python3 monitor_gpu.py      # Real-time, visual graphs
python3 monitor_gpu.py 0.5  # Faster refresh

4) Test transcription

# Local file
python3 test_client.py path/to/audio.wav

# YouTube (JSON body — not form data)
curl -X POST -H "Content-Type: application/json" \
  -d '{"url":"https://www.youtube.com/watch?v=VIDEO_ID"}' \
  http://localhost:8080/youtube

API endpoints

  • POST /youtube
curl -X POST -H "Content-Type: application/json" \
  -d '{"url":"https://www.youtube.com/watch?v=VIDEO_ID"}' \
  http://localhost:8080/youtube

Response (queued):

{
  "job_id": "job_1757939346_605ce77a",
  "status": "queued",
  "message": "YouTube audio downloaded and queued for transcription",
  "video_info": {
    "duration": 2823,
    "file_size": 541938218,
    "title": "Video Title Here",
    "uploader": "Channel Name"
  }
}
  • POST /upload (multipart)
curl -F "audio=@/path/to/audio.wav" http://localhost:8080/upload
  • GET /results/{job_id}
curl http://localhost:8080/results/job_1694648400_a1b2c3d4
  • GET /status
curl http://localhost:8080/status
  • GET /health
curl http://localhost:8080/health

Architecture overview

flowchart LR
  %% Define nodes with quoted labels to avoid special-char issues
  A["HTTP Client"]
  B["HTTP Server (C++/Crow)"]
  C["Job Queue (C++ MT)"]
  D["GPU Transcriber (Faster-Whisper)"]
  E["SQLite DB"]
  F["Jetson Orin GPU"]

  %% Connect nodes
  A --> B
  B --> C
  C --> D
  B --> E
  D -- CUDA --> F
  D --> E
  • HTTP server in C++ (Crow), backed by a multi-threaded job queue
  • GPU-accelerated transcription via Faster-Whisper (PyTorch, float16)
  • SQLite stores jobs, status, and results (including segments)
  • CUDA runtime on Jetson Orin (device cuda:0)

GPU acceleration and performance

  • Dedicated CUDA env: .venv_cuda
  • Optimized inference: transcriber_gpu.py
  • Precision: float16 for speed and memory
  • Typical speedup vs CPU: ~4.1x
  • Processing speed: ~1.19 MB/s on large files
  • Memory usage: ~535MB for a 48‑min video (base model)

Tuning and configuration

Edit transcriber_gpu.py to customize:

  • Model size: tiny, base, small, medium, large
  • Compute type: float16, int8_float16, int8
  • Device: cuda (auto-detected)
  • Decoding params: beam size, temperature, VAD, etc.

Tips:

  • For low memory, try small or tiny + int8_float16
  • For quality, use base/small with VAD filtering enabled

Monitoring

  • monitor_gpu.py — Beautiful real-time terminal dashboard

    • GPU utilization, memory, temperature, power
    • CPU per-core usage and frequencies
    • Memory usage charts and sparklines
    • Historical trend graphs
  • tegrastats — Lightweight alternative

tegrastats --interval 1000
watch -n 1 tegrastats

Operations and service management

Systemd (production):

./stt-service-control.sh status
./stt-service-control.sh restart
./stt-service-control.sh logs
./stt-service-control.sh test
./stt-service-control.sh youtube <url>

sudo journalctl -u stt-service -f

Manual script (dev):

./start_service.sh         # Start on port 8080
./start_service.sh 3000    # Custom port
./start_service.sh restart # Restart
./start_service.sh stop    # Stop

What the startup script does:

  • Detects GPU/CUDA and selects the correct Python venv (.venv_cuda)
  • Runs health checks and endpoint tests
  • Handles port conflicts and logging

Troubleshooting

  • Different YouTube videos return identical transcripts

    • Cause: wrong file pick from downloader
    • Fix: youtube_downloader.py updated to properly match downloaded files by title
  • JSON parse errors with -- in output

    • Cause: stderr corruption
    • Fix: redirect stderr (2>/dev/null) and use contextlib.redirect_stdout(sys.stderr) where appropriate
  • Truncated transcription text

    • Cause: legacy truncation at 1000 chars
    • Fix: removed truncation — full JSON always returned
  • Port conflict (8080)

lsof -ti:8080 | xargs -r kill -9
sleep 1
make -C build && ./build/stt_service &
  • Wrong virtual environment
    • Symptom: GPU transcriber fails
    • Fix: ensure .venv_cuda is used everywhere

Component tests

source .venv_cuda/bin/activate

# YouTube downloader
python3 youtube_downloader.py "https://www.youtube.com/watch?v=VIDEO_ID" uploads 2>/dev/null

# GPU transcriber
python3 transcriber_gpu.py "uploads/audio_file.wav" base

# Check files
ls -la uploads/ | tail -5

Verify different videos produce different files:

bash -c 'source .venv_cuda/bin/activate && python3 youtube_downloader.py "URL1" uploads' | jq -r '.audio_file'
bash -c 'source .venv_cuda/bin/activate && python3 youtube_downloader.py "URL2" uploads' | jq -r '.audio_file'

Build and development

make all        # Release build
make debug      # Debug build
make clean      # Clean

# CMake path if needed
rm -rf build && mkdir build
cd build && cmake .. -DCMAKE_POLICY_VERSION_MINIMUM=3.5 && make -j4

# Dependencies and checks
make check

File structure

TTSsystem/
├── start_service.sh
├── stt-service-control.sh
├── stt-service.service
├── monitor_gpu.py
├── transcriber_gpu.py
├── transcriber.py
├── youtube_downloader.py
├── build/stt_service
├── .venv_cuda/
├── .venv/
├── logs/
├── uploads/
├── results/
└── src/
    ├── server/
    ├── transcriber/
    ├── database/
    └── queue/

Real‑world performance (Jetson Orin)

  • ~4.1× faster than CPU-only
  • ~1.19 MB/s processing throughput (large files)
  • ~535MB RAM usage on a 48‑min video
  • ~2.5s model load time (base)
  • Strong accuracy with VAD + beam search

Closing thoughts

This service is purpose‑built for the NVIDIA Jetson Orin — fast, reliable, and easy to operate. If you’re building on-device speech experiences, this gives you a robust foundation with CUDA acceleration, async pipelines, and observability out of the box.