STT Service for NVIDIA Jetson Orin
High-performance Speech-to-Text service optimized for NVIDIA Jetson Orin Developer Edition using Faster-Whisper with CUDA acceleration.
Why this matters
Building reliable, fast STT on edge devices is hard. This project delivers a production-ready service that:
- Runs fully on-device with CUDA acceleration (Jetson Orin)
- Handles YouTube ingestion, file uploads, and long-form audio
- Persists jobs and results with an async queue and SQLite
- Exposes a clean REST API you can call from any client
- Monitors real-time GPU/CPU usage with a beautiful terminal UI
It’s designed for speed, resilience, and easy operations.
Features
- GPU-Accelerated transcription (CUDA, float16)
- YouTube video transcription with audio extraction
- Asynchronous processing via job queue and IDs
- Persistent storage in SQLite
- Multiple audio formats: WAV, MP3, M4A, FLAC, OGG, OPUS, MP4, MKV, WEBM
- RESTful API for upload, YouTube, status, and results
- Performance metrics: processing time, confidence, segments
- Real-time GPU/CPU monitoring with graphs and trend lines
- Systemd integration and smart startup scripts
System requirements
- NVIDIA Jetson Orin Developer Edition
- CUDA 12.x
- Ubuntu 20.04/22.04
- Python 3.8+
- 4GB+ available RAM
Quick start
1) Setup
# Clone and setup
git clone <repository-url> # I will update this document with this url when I get it cleaned up and pushed.
cd TTSsystem
# Full automated setup (recommended)
make setup
# OR manual setup
make install-deps
make install-python
make all
2) Start service
Option A — Systemd (recommended)
sudo cp stt-service.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable stt-service
sudo systemctl start stt-service
# Helper controls
./stt-service-control.sh status
./stt-service-control.sh logs
./stt-service-control.sh youtube <url>
Option B — Manual (development)
./start_service.sh # Start on port 8080
./start_service.sh 3000 # Custom port
./start_service.sh restart # Restart
./start_service.sh stop # Stop
3) Monitor GPU/CPU
python3 monitor_gpu.py # Real-time, visual graphs
python3 monitor_gpu.py 0.5 # Faster refresh
4) Test transcription
# Local file
python3 test_client.py path/to/audio.wav
# YouTube (JSON body — not form data)
curl -X POST -H "Content-Type: application/json" \
-d '{"url":"https://www.youtube.com/watch?v=VIDEO_ID"}' \
http://localhost:8080/youtube
API endpoints
- POST
/youtube
curl -X POST -H "Content-Type: application/json" \
-d '{"url":"https://www.youtube.com/watch?v=VIDEO_ID"}' \
http://localhost:8080/youtube
Response (queued):
{
"job_id": "job_1757939346_605ce77a",
"status": "queued",
"message": "YouTube audio downloaded and queued for transcription",
"video_info": {
"duration": 2823,
"file_size": 541938218,
"title": "Video Title Here",
"uploader": "Channel Name"
}
}
- POST
/upload(multipart)
curl -F "audio=@/path/to/audio.wav" http://localhost:8080/upload
- GET
/results/{job_id}
curl http://localhost:8080/results/job_1694648400_a1b2c3d4
- GET
/status
curl http://localhost:8080/status
- GET
/health
curl http://localhost:8080/health
Architecture overview
flowchart LR
%% Define nodes with quoted labels to avoid special-char issues
A["HTTP Client"]
B["HTTP Server (C++/Crow)"]
C["Job Queue (C++ MT)"]
D["GPU Transcriber (Faster-Whisper)"]
E["SQLite DB"]
F["Jetson Orin GPU"]
%% Connect nodes
A --> B
B --> C
C --> D
B --> E
D -- CUDA --> F
D --> E
- HTTP server in C++ (Crow), backed by a multi-threaded job queue
- GPU-accelerated transcription via Faster-Whisper (PyTorch, float16)
- SQLite stores jobs, status, and results (including segments)
- CUDA runtime on Jetson Orin (device
cuda:0)
GPU acceleration and performance
- Dedicated CUDA env:
.venv_cuda - Optimized inference:
transcriber_gpu.py - Precision:
float16for speed and memory - Typical speedup vs CPU: ~4.1x
- Processing speed: ~1.19 MB/s on large files
- Memory usage: ~535MB for a 48‑min video (base model)
Tuning and configuration
Edit transcriber_gpu.py to customize:
- Model size:
tiny,base,small,medium,large - Compute type:
float16,int8_float16,int8 - Device:
cuda(auto-detected) - Decoding params: beam size, temperature, VAD, etc.
Tips:
- For low memory, try
smallortiny+int8_float16 - For quality, use
base/smallwith VAD filtering enabled
Monitoring
-
monitor_gpu.py— Beautiful real-time terminal dashboard- GPU utilization, memory, temperature, power
- CPU per-core usage and frequencies
- Memory usage charts and sparklines
- Historical trend graphs
-
tegrastats— Lightweight alternative
tegrastats --interval 1000
watch -n 1 tegrastats
Operations and service management
Systemd (production):
./stt-service-control.sh status
./stt-service-control.sh restart
./stt-service-control.sh logs
./stt-service-control.sh test
./stt-service-control.sh youtube <url>
sudo journalctl -u stt-service -f
Manual script (dev):
./start_service.sh # Start on port 8080
./start_service.sh 3000 # Custom port
./start_service.sh restart # Restart
./start_service.sh stop # Stop
What the startup script does:
- Detects GPU/CUDA and selects the correct Python venv (
.venv_cuda) - Runs health checks and endpoint tests
- Handles port conflicts and logging
Troubleshooting
-
Different YouTube videos return identical transcripts
- Cause: wrong file pick from downloader
- Fix:
youtube_downloader.pyupdated to properly match downloaded files by title
-
JSON parse errors with
--in output- Cause: stderr corruption
- Fix: redirect
stderr(2>/dev/null) and usecontextlib.redirect_stdout(sys.stderr)where appropriate
-
Truncated transcription text
- Cause: legacy truncation at 1000 chars
- Fix: removed truncation — full JSON always returned
-
Port conflict (8080)
lsof -ti:8080 | xargs -r kill -9
sleep 1
make -C build && ./build/stt_service &
- Wrong virtual environment
- Symptom: GPU transcriber fails
- Fix: ensure
.venv_cudais used everywhere
Component tests
source .venv_cuda/bin/activate
# YouTube downloader
python3 youtube_downloader.py "https://www.youtube.com/watch?v=VIDEO_ID" uploads 2>/dev/null
# GPU transcriber
python3 transcriber_gpu.py "uploads/audio_file.wav" base
# Check files
ls -la uploads/ | tail -5
Verify different videos produce different files:
bash -c 'source .venv_cuda/bin/activate && python3 youtube_downloader.py "URL1" uploads' | jq -r '.audio_file'
bash -c 'source .venv_cuda/bin/activate && python3 youtube_downloader.py "URL2" uploads' | jq -r '.audio_file'
Build and development
make all # Release build
make debug # Debug build
make clean # Clean
# CMake path if needed
rm -rf build && mkdir build
cd build && cmake .. -DCMAKE_POLICY_VERSION_MINIMUM=3.5 && make -j4
# Dependencies and checks
make check
File structure
TTSsystem/
├── start_service.sh
├── stt-service-control.sh
├── stt-service.service
├── monitor_gpu.py
├── transcriber_gpu.py
├── transcriber.py
├── youtube_downloader.py
├── build/stt_service
├── .venv_cuda/
├── .venv/
├── logs/
├── uploads/
├── results/
└── src/
├── server/
├── transcriber/
├── database/
└── queue/
Real‑world performance (Jetson Orin)
- ~4.1× faster than CPU-only
- ~1.19 MB/s processing throughput (large files)
- ~535MB RAM usage on a 48‑min video
- ~2.5s model load time (base)
- Strong accuracy with VAD + beam search
Closing thoughts
This service is purpose‑built for the NVIDIA Jetson Orin — fast, reliable, and easy to operate. If you’re building on-device speech experiences, this gives you a robust foundation with CUDA acceleration, async pipelines, and observability out of the box.