Resume
Levi DeHaan’s experience in applied AI, forward deployed engineering, agentic processes, automation, DevOps, and cloud infrastructure.
Forward Deployed AI Engineer (Applied AI)
Forward Deployed AI Engineer (Applied AI) at Glean. I enable clients to get real value from AI by building agents, complex integrations, and hybrid ecosystems — currently carrying three times the average client load while exceeding quarterly targets. Previously Automation and AI Systems Architect and DevOps/SRE/CIS manager at Petco, where I cut ticket closure from six months to hours and stood up multi-cloud LLM serving.
Experience
Forward Deployed AI Engineer (Applied AI) — Glean
2025-11-01 – Present
- Enable clients to maximize AI value by building agents, complex integrations, and hybrid AI ecosystems
- Manage three times the average client load while exceeding quarterly targets
- Provide technical guidance across software stacks, automation, and infrastructure
- Debug customer AI infrastructure at the code level, including sandbox-orchestrator crashes
- Design multi-VPC AWS Transit Gateway architectures for enterprise MCP server connectivity
- QuestionMonitor: Swift + Python macOS app with a sub-250ms VAD/whisper/SetFit pipeline that cut hallucinations 90–95%.
- VideoEditorPro: Local-first React/Node/FFmpeg editor with canvas encoding, compositing, and TTS.
- GTMExtensions: Chrome Native Messaging + TypeScript pipeline bridging extensions to macOS desktop apps over local TCP.
- Meeting Prep: Menu bar app with EventKit calendar sync and GRDB/SQLite for durable agent meeting prep.
- Agentic Workflows: Multi-step agent architectures with advanced classification for client deployments.
- Core Infra: Debugged sandbox-orchestrator crashes at the code level and designed multi-VPC AWS Transit Gateway architectures for enterprise MCP connectivity.
Agentic process education and workflow automation: Educated thousands on how to build agentic processes, while designing and building software to automate my own workflows.
Architected QuestionMonitor, a production Swift/Python macOS app with a sub-250ms VAD/whisper/SetFit pipeline: Reduced hallucinations by 90-95%
Built VideoEditorPro, GTMExtensions, and Meeting Prep as internal Glean-powered applications: Shipped local-first video, extension-to-desktop, and calendar agent tooling
Technologies: Swift, SwiftUI, Python, PyTorch, TypeScript, React, Node.js, AWS, MCP
Principal Architect & Consultant — AutomatonAutomation.ai & DeHaan Consulting
2018-12-01 – Present
- Architect high-performance inference and agent platforms
- Build real-time market data and autonomous trading systems
- Deliver log-analysis and visual stream-processing products
- Almagest: Native C++ llama.cpp engine with durable workflows, layered memory, MCP tools, and a live dashboard.
- AI Cluster: Separate Go server on llama.cpp with native MCP tool selection, agent memory, assistant workflows, and omni voice chat.
- BifrostPipe: C++ Kafka pipeline for high-frequency WebSocket market data and low-latency transforms.
- Alpaca Auto Trader: 24/7 two-stage trading system with news/research, Twitter signals, QMD search, and TimesFM forecasts.
- FluentD LLM Server: FluentD + Phi-2 on Kubernetes for real-time threat detection and automated response.
- Visual & Video Pipelines: Drag-and-drop Kafka stream builder and a YouTube financial video analyzer using DeepSeek.
Almagest: native C++ llama.cpp engine with durable workflows, layered memory, and MCP tools: Parent/child contexts and a live dashboard on two RTX 3060s
AI Cluster: Go server on llama.cpp with native MCP tool selection and agent memory: Agents, assistant, omni voice chat, and high-speed context memory
Alpaca Auto Trader: 24/7 two-stage autonomous trading system: Scanner model plus execution with news, Twitter, QMD, and TimesFM
Technologies: Go, C++, llama.cpp, Kafka, Kubernetes, Python, Phi-2
Automation and AI Systems Architect — Petco
2024-10-01 – 2025-03-01
- Designed and deployed LLM serving across AWS, Azure, and an in-house GPU cluster
- Built standardized, repeatable AI workflows for development teams
- Replaced a Google Ads service with an in-house AI workflow in a week
- LLM: Built advanced LLM systems using open source technologies like llama.cpp and transformers for incredibly fast inference on low cost systems.
- AI Everywhere: Deployed on multiple clouds, with Kubernetes backends, utilized K8's api for distributing inference. Can run on CPU and GPU using all resources. Memory and model management interfaces through Simple Ai Router.
- Enhanced Applications: Developed and deployed a cost-effective AI based solution to replace a google ads service, leveraging AI workflows and Simple AI backend service in a week.
- AI Assist: Mastered AI-assisted coding tools like Windsurf, Cursor, and Claude Desktop, establishing usage methodologies for effective and fast development.
- Consistent: Created standardized workflows for AI integration across development teams.
- Scalable Systems: Designed and deployed scalable AI infrastructure, reducing costs while enhancing capabilities.
Multi-cloud LLM serving with llama.cpp, transformers, and Simple AI Router: Fast inference on low-cost CPU and GPU with Kubernetes distribution
Technologies: llama.cpp, Kubernetes, AWS, Azure, Python, Go
Manager - DevOps COE/SRE/Cloud Infrastructure Services + AI Systems Architect — Petco
2022-08-01 – 2024-10-01
- Led DevOps, SRE, and Cloud Infrastructure Services within six months of joining
- Built AI tools including Ticket Whisperer and PetChat
- Deployed Azure OpenAI fine-tuned models for ticket automation
- Implemented custom Backstage integrated with Terraform and internal auth
- Rapid Advancement: Within a span of six months, I was entrusted with the leadership of the DevOps, SRE, and Cloud Infrastructure Services (CIS) teams, reflecting my capability to drive results and manage diverse teams.
- Developer Engagement: Initiated an internal advertising campaign within the first two months, educating hundreds of developers about the capabilities and tools of the DevOps team. This resulted in an increased collaboration across teams.
- AI Everywhere: My team(s) spearheaded the development of several AI based tools: Ticket Whisperer, PetChat, and many more teams use our team for AI to build their own apps.
- Operational Efficiency: Centralized our logging and statistics into a unified dashboard, streamlining the SRE's collaboration with the DevOps team and ensuring data-driven solutions.
- AI Integration: We deployed an Azure OpenAI Fine-Tuned Curie and then Davinci model, automating ticket responses for administrative tasks. This innovation liberated resources, allowing us to focus on more strategic initiatives.
- Cloud Deployment: Successfully implemented a custom Backstage system, integrating it with our Terraform tools to facilitate seamless deployments across various cloud environments. This system is also integrated with our internal user management for efficient authorization.
- AI Chat Framework: Designed and launched an advanced AI chatbot framework, powered by a customized Langchain agent tying into confluence. This framework provides instant answers to complex, context-specific queries, transforming hours of research into minutes.
- Model Fine-Tuning: Established a robust pipeline to fine-tune models on azure openai for our Ticket Whisperer system.
- Pet-Centric Solutions: Pioneered the development of AI agents tailored for pet-related tools, enhancing both customer experience and providing veterinarians with contextual information instantly.
- Team Transformation: Before my leadership, a team I was given to manage had a ticket closure time of up to six months. Within two months, I revolutionized the process, reducing response times to mere hours and automating numerous systems. This transformation led to a significant reduction in daily tickets and allowed us to reallocate resources more efficiently.
- Collaborative Leadership: As my role evolved into a more strategic one, I ensured smooth operations by onboarding talent to manage day-to-day reporting. This allowed me to focus on crafting solutions and collaborating with other teams to kickstart projects.
Cut a team's ticket closure time from up to six months to hours: Automated systems and reallocated engineering capacity
Technologies: Azure OpenAI, Kubernetes, Terraform, Backstage, LangChain, Python
Director - Cloud Infrastructure — Healthpointe Solutions
2021-03-01 – 2022-08-01
- Built base infrastructure with Terraform on AWS/Azure
- Created Terraform Manager, Imaging Manager, and TerraMatch
- Deployed Databricks and patient-data pipelines in Python and Scala
- HITRUST security implementations and executive security advisory
- Stood up the company's base infrastructure with Terraform on AWS/Azure, continuing the Enable Data engagement
- Built Terraform Manager for push-button deploy of hundreds of configurations to any cloud
- Built Imaging Manager with Packer and Ansible, including security, virus, and test gates before promoting images
- Built TerraMatch so only Terraform-deployed systems are allowed to run
- Designed Databricks on private Azure/AWS clouds, plus Python/Scala pipelines for healthcare and insurance patient data
- Dell ECS POC: Kubernetes + Tableau so Dell could offer HPS products to healthcare providers
- HITRUST security implementations; C-suite security/network advisor for international client travel
- Mentored teammates on whatever they needed to be successful
Dell ECS POC with Kubernetes and Tableau for healthcare providers: Enabled Dell to offer HPS products on ECS
Technologies: Terraform, AWS, Azure, Databricks, Packer, Ansible, Kubernetes
Sr. Systems Architect — Enable Data Consulting
2018-10-01 – 2021-03-01
- Built Greenwich.HR labor-market intelligence platform on thousands of ECS jobs
- Converted IRI Worldwide production R workloads to Databricks Spark
- Built HealthPointe multi-cloud baseline with Terraform, Ansible, and Jenkins
- Built Greenwich.HR's labor-market intelligence platform: thousands of ECS jobs with queues, reporting, and customer management
- Converted an IRI Worldwide production R job to Databricks Spark; runtime dropped to 15% of production
- Built HealthPointe's multi-cloud Terraform/Ansible/Jenkins baseline with automated security scanning and detection
- Ran technical interviews for client companies, from Spark to systems management
IRI Worldwide R-to-Spark conversion: Runtime dropped to 15% of production
Technologies: AWS ECS, Databricks, Spark, Terraform, Ansible, Jenkins
Sr. DevOps Architect specializing in Kubernetes — Nielsen
2018-02-01 – 2018-09-01
- Designed Azure AKS/ACS interface to launch big data jobs into Kubernetes via Jenkins
- Introduced CI/CD practices and trained other departments
- Built a Vue.js/Express GUI for cluster and Jenkins job management
- Designed Azure AKS/ACS interface to launch big data jobs into Kubernetes clusters
- Created faster workflow to reduce job run times
- Implemented CI/CD practices for faster code delivery
- Built GUI interface for kubernetes/jenkins system using Vue.js/Express
- Trained other departments on deployment methodologies
Technologies: Kubernetes, Azure AKS, Jenkins, Vue.js, Express
DevOps Manager — Simple Energy
2015-08-01 – 2017-04-01
- Created new DevOps workflow focused on improving automation and system deployments
- Reduced deploy time from 3-4 weeks to less than a day, and platform updates to minutes from hours
- Implemented Docker/Mesos to increase uptime on volatile systems
- Instantiated the use of Jenkins/GitLab to increase developer productivity and accountability
- Reduced time to test for developers and reduced complexity to get to production
- Improved testing strategies and implemented automated recovery for simple systems
- Reduced costs across the board by automating services and deployments
- Worked on various big data systems with Scala and Python for energy usage analytics
Technologies: Docker, Mesos, Jenkins, GitLab, Scala, Python
Sr. Developer / Sr. Systems Administrator / DevOps Manager — Promet Source
2015-01-01 – 2015-07-01
- Helped design and build the DevOps team from the ground up
- Transferred servers from clients to a unified infrastructure for smoother deploys
- Designed and built new server infrastructure using Google Cloud, saving money
- Implemented new tool stacks to increase team efficiency and automation
- Reduced deployment time from days to hours with improved infrastructure
- Managed several development projects in PHP, JavaScript, Java, and Python
- Worked on a new Drupal module integrating Tika and Solr to index millions of PDFs
- Built Jenkins automation cluster to streamline operations
- Implemented Elasticsearch cluster for centralized logging
- Developed Node.js applications for PDF conversion and IMAP mail
Technologies: GCP, PHP, JavaScript, Java, Python, Drupal, Tika, Solr, Jenkins, Elasticsearch, Node.js
Systems Administrator / DevOps Engineer — Time Warner
2014-05-01 – 2014-11-01
- Used Jenkins, Docker, Mesos, and Puppet to implement CI/CD processes
- Built new onboarding processes with Vagrant for faster developer setup
- Introduced Docker as a tool for development and system administration
- Designed big-data workflows using Docker to containerize platforms
- Created clusters within single machines to test larger deployments
- Designed workflows around Mesos/Marathon and built UI in Angular.js
Technologies: Jenkins, Docker, Mesos, Puppet, Vagrant, Marathon, Angular.js
Skills & expertise
Programming
Python, JavaScript, TypeScript, Go, C++, Swift, Java, C#, PHP, SQL, Scala
AI & machine learning
AI Agents, MCP, LLM Fine-tuning, RAG, Machine Learning, AI Serving Infrastructure, llama.cpp, PyTorch, LangChain, Hugging Face, AI Agents & MCP, LLM Serving & Fine-tuning, Forward Deployed Engineering, Kubernetes & Multi-cloud, Real-time Trading Systems, XR / VR / AR
Web, desktop & testing
React, Next.js, SwiftUI, Vue.js, Node.js, Express, AppKit, Chrome Native Messaging, Cordova, Playwright
Cloud platforms
AWS, EKS, ECS, EC2, S3, Transit Gateway, Azure, AKS, Azure OpenAI, Databricks, GCP, GKE, Compute
Infrastructure & delivery
Docker, Kubernetes, Mesos, FluentD, Prometheus, Jenkins, GitLab CI, Terraform, Ansible
Data & databases
PostgreSQL, MySQL, SQLite, Cassandra, Redis, Kafka, Apache Arrow, Spark
Networking & virtualization
Xen, VMware, Cisco, HAProxy, WebSockets, gRPC
Systems & delivery
Linux, Windows, Server Administration, Datacenter Management, Framework Design, Project Management
Extended reality
Unity, HoloLens, A-Frame, WebXR, VR Development, 3D Modeling
Additional tools used in roles
Phi-2, Backstage, Packer, AWS ECS, Azure AKS, GitLab, Drupal, Tika, Solr, Elasticsearch, Puppet, Vagrant, Marathon, Angular.js
Selected projects
QuestionMonitor
Production macOS app bridging a SwiftUI frontend to a Python PyTorch backend over a local WebSocket server. Sub-250ms ML pipeline with Silero VAD (90–95% fewer hallucinations), faster-whisper, and custom SetFit/transformer models.
VideoEditorPro
Local-first video editor using React, Node.js, and FFmpeg pipelines for pixel-perfect canvas encoding, compositing, and text-to-speech generation.
Alpaca Auto Trader (AAT)
Autonomous 24/7 trading system with a two-stage pipeline: a fast 7B scanner identifies candidates and executes trades. Intelligence stack includes news/research streaming, Twitter signals, QMD search across 100K+ documents, TimesFM forecasting, and automated trade autopsies.
Almagest
Native C++ llama.cpp engine with durable workflows, layered SQLite memory, MCP tools, and a live dashboard. Qwen3.8-27B on two RTX 3060s with MTP.
AI Cluster
Separate Go inference cluster on llama.cpp. Native MCP tool selection, high-speed agent context memory, agents, assistant, and omni voice chat.
GTMExtensions
End-to-end training pipeline using Chrome Native Messaging, TypeScript 5.x, and Webpack 5 to securely bridge browser extensions to macOS desktop apps over local TCP sockets.
Meeting Prep
macOS menu bar app with EventKit calendar sync and GRDB/SQLite for durable, structured AI agent meeting preparation.
Ticket Whisperer
Petco IT ticket automation integrated with ServiceNow. Fine-tuned Azure OpenAI models parse, categorize, and route tickets without human intervention.
Simple AI Router (SAR)
Go reverse proxy with autoscaling that routes between GPU backends, monitors load and memory, and scales inference across clouds and datacenters.
FluentD LLM Server
Log analysis with FluentD plus custom-trained Microsoft Phi-2 models on Kubernetes for real-time security threat detection and automated anomaly response.
BifrostPipe
High-performance C++ Kafka pipeline for real-time, high-frequency market data over WebSockets with low-latency transformation.
ORCA VR - Server Management Platform
Designed and developed software platform for managing server systems in virtual reality, providing immersive monitoring and management capabilities.
Hololens Oil Refinery Inspection
Designed Microsoft Hololens application for making oil refinery inspections easier to verify and manage, improving safety and efficiency.
VR Garage Platform
Built a platform enabling users to easily create VR video/image/3D model experiences using a web-based interface. Used JHipster microservices and A-Frame for the frontend, with Kubernetes and fabric8 for the backend.
Education
Colorado Mountain College — Coursework in programming, writing, and mathematics