Levi DeHaan

Resume

Levi DeHaan’s experience in applied AI, forward deployed engineering, agentic processes, automation, DevOps, and cloud infrastructure.

Forward Deployed AI Engineer (Applied AI)

Forward Deployed AI Engineer (Applied AI) at Glean. I enable clients to get real value from AI by building agents, complex integrations, and hybrid ecosystems — currently carrying three times the average client load while exceeding quarterly targets. Previously Automation and AI Systems Architect and DevOps/SRE/CIS manager at Petco, where I cut ticket closure from six months to hours and stood up multi-cloud LLM serving.

Experience

Forward Deployed AI Engineer (Applied AI) — Glean

2025-11-01 – Present

  • Enable clients to maximize AI value by building agents, complex integrations, and hybrid AI ecosystems
  • Manage three times the average client load while exceeding quarterly targets
  • Provide technical guidance across software stacks, automation, and infrastructure
  • Debug customer AI infrastructure at the code level, including sandbox-orchestrator crashes
  • Design multi-VPC AWS Transit Gateway architectures for enterprise MCP server connectivity
  • QuestionMonitor: Swift + Python macOS app with a sub-250ms VAD/whisper/SetFit pipeline that cut hallucinations 90–95%.
  • VideoEditorPro: Local-first React/Node/FFmpeg editor with canvas encoding, compositing, and TTS.
  • GTMExtensions: Chrome Native Messaging + TypeScript pipeline bridging extensions to macOS desktop apps over local TCP.
  • Meeting Prep: Menu bar app with EventKit calendar sync and GRDB/SQLite for durable agent meeting prep.
  • Agentic Workflows: Multi-step agent architectures with advanced classification for client deployments.
  • Core Infra: Debugged sandbox-orchestrator crashes at the code level and designed multi-VPC AWS Transit Gateway architectures for enterprise MCP connectivity.

Agentic process education and workflow automation: Educated thousands on how to build agentic processes, while designing and building software to automate my own workflows.

Architected QuestionMonitor, a production Swift/Python macOS app with a sub-250ms VAD/whisper/SetFit pipeline: Reduced hallucinations by 90-95%

Built VideoEditorPro, GTMExtensions, and Meeting Prep as internal Glean-powered applications: Shipped local-first video, extension-to-desktop, and calendar agent tooling

Technologies: Swift, SwiftUI, Python, PyTorch, TypeScript, React, Node.js, AWS, MCP

Principal Architect & Consultant — AutomatonAutomation.ai & DeHaan Consulting

2018-12-01 – Present

  • Architect high-performance inference and agent platforms
  • Build real-time market data and autonomous trading systems
  • Deliver log-analysis and visual stream-processing products
  • Almagest: Native C++ llama.cpp engine with durable workflows, layered memory, MCP tools, and a live dashboard.
  • AI Cluster: Separate Go server on llama.cpp with native MCP tool selection, agent memory, assistant workflows, and omni voice chat.
  • BifrostPipe: C++ Kafka pipeline for high-frequency WebSocket market data and low-latency transforms.
  • Alpaca Auto Trader: 24/7 two-stage trading system with news/research, Twitter signals, QMD search, and TimesFM forecasts.
  • FluentD LLM Server: FluentD + Phi-2 on Kubernetes for real-time threat detection and automated response.
  • Visual & Video Pipelines: Drag-and-drop Kafka stream builder and a YouTube financial video analyzer using DeepSeek.

Almagest: native C++ llama.cpp engine with durable workflows, layered memory, and MCP tools: Parent/child contexts and a live dashboard on two RTX 3060s

AI Cluster: Go server on llama.cpp with native MCP tool selection and agent memory: Agents, assistant, omni voice chat, and high-speed context memory

Alpaca Auto Trader: 24/7 two-stage autonomous trading system: Scanner model plus execution with news, Twitter, QMD, and TimesFM

Technologies: Go, C++, llama.cpp, Kafka, Kubernetes, Python, Phi-2

Automation and AI Systems Architect — Petco

2024-10-01 – 2025-03-01

  • Designed and deployed LLM serving across AWS, Azure, and an in-house GPU cluster
  • Built standardized, repeatable AI workflows for development teams
  • Replaced a Google Ads service with an in-house AI workflow in a week
  • LLM: Built advanced LLM systems using open source technologies like llama.cpp and transformers for incredibly fast inference on low cost systems.
  • AI Everywhere: Deployed on multiple clouds, with Kubernetes backends, utilized K8's api for distributing inference. Can run on CPU and GPU using all resources. Memory and model management interfaces through Simple Ai Router.
  • Enhanced Applications: Developed and deployed a cost-effective AI based solution to replace a google ads service, leveraging AI workflows and Simple AI backend service in a week.
  • AI Assist: Mastered AI-assisted coding tools like Windsurf, Cursor, and Claude Desktop, establishing usage methodologies for effective and fast development.
  • Consistent: Created standardized workflows for AI integration across development teams.
  • Scalable Systems: Designed and deployed scalable AI infrastructure, reducing costs while enhancing capabilities.

Multi-cloud LLM serving with llama.cpp, transformers, and Simple AI Router: Fast inference on low-cost CPU and GPU with Kubernetes distribution

Technologies: llama.cpp, Kubernetes, AWS, Azure, Python, Go

Manager - DevOps COE/SRE/Cloud Infrastructure Services + AI Systems Architect — Petco

2022-08-01 – 2024-10-01

  • Led DevOps, SRE, and Cloud Infrastructure Services within six months of joining
  • Built AI tools including Ticket Whisperer and PetChat
  • Deployed Azure OpenAI fine-tuned models for ticket automation
  • Implemented custom Backstage integrated with Terraform and internal auth
  • Rapid Advancement: Within a span of six months, I was entrusted with the leadership of the DevOps, SRE, and Cloud Infrastructure Services (CIS) teams, reflecting my capability to drive results and manage diverse teams.
  • Developer Engagement: Initiated an internal advertising campaign within the first two months, educating hundreds of developers about the capabilities and tools of the DevOps team. This resulted in an increased collaboration across teams.
  • AI Everywhere: My team(s) spearheaded the development of several AI based tools: Ticket Whisperer, PetChat, and many more teams use our team for AI to build their own apps.
  • Operational Efficiency: Centralized our logging and statistics into a unified dashboard, streamlining the SRE's collaboration with the DevOps team and ensuring data-driven solutions.
  • AI Integration: We deployed an Azure OpenAI Fine-Tuned Curie and then Davinci model, automating ticket responses for administrative tasks. This innovation liberated resources, allowing us to focus on more strategic initiatives.
  • Cloud Deployment: Successfully implemented a custom Backstage system, integrating it with our Terraform tools to facilitate seamless deployments across various cloud environments. This system is also integrated with our internal user management for efficient authorization.
  • AI Chat Framework: Designed and launched an advanced AI chatbot framework, powered by a customized Langchain agent tying into confluence. This framework provides instant answers to complex, context-specific queries, transforming hours of research into minutes.
  • Model Fine-Tuning: Established a robust pipeline to fine-tune models on azure openai for our Ticket Whisperer system.
  • Pet-Centric Solutions: Pioneered the development of AI agents tailored for pet-related tools, enhancing both customer experience and providing veterinarians with contextual information instantly.
  • Team Transformation: Before my leadership, a team I was given to manage had a ticket closure time of up to six months. Within two months, I revolutionized the process, reducing response times to mere hours and automating numerous systems. This transformation led to a significant reduction in daily tickets and allowed us to reallocate resources more efficiently.
  • Collaborative Leadership: As my role evolved into a more strategic one, I ensured smooth operations by onboarding talent to manage day-to-day reporting. This allowed me to focus on crafting solutions and collaborating with other teams to kickstart projects.

Cut a team's ticket closure time from up to six months to hours: Automated systems and reallocated engineering capacity

Technologies: Azure OpenAI, Kubernetes, Terraform, Backstage, LangChain, Python

Director - Cloud Infrastructure — Healthpointe Solutions

2021-03-01 – 2022-08-01

  • Built base infrastructure with Terraform on AWS/Azure
  • Created Terraform Manager, Imaging Manager, and TerraMatch
  • Deployed Databricks and patient-data pipelines in Python and Scala
  • HITRUST security implementations and executive security advisory
  • Stood up the company's base infrastructure with Terraform on AWS/Azure, continuing the Enable Data engagement
  • Built Terraform Manager for push-button deploy of hundreds of configurations to any cloud
  • Built Imaging Manager with Packer and Ansible, including security, virus, and test gates before promoting images
  • Built TerraMatch so only Terraform-deployed systems are allowed to run
  • Designed Databricks on private Azure/AWS clouds, plus Python/Scala pipelines for healthcare and insurance patient data
  • Dell ECS POC: Kubernetes + Tableau so Dell could offer HPS products to healthcare providers
  • HITRUST security implementations; C-suite security/network advisor for international client travel
  • Mentored teammates on whatever they needed to be successful

Dell ECS POC with Kubernetes and Tableau for healthcare providers: Enabled Dell to offer HPS products on ECS

Technologies: Terraform, AWS, Azure, Databricks, Packer, Ansible, Kubernetes

Sr. Systems Architect — Enable Data Consulting

2018-10-01 – 2021-03-01

  • Built Greenwich.HR labor-market intelligence platform on thousands of ECS jobs
  • Converted IRI Worldwide production R workloads to Databricks Spark
  • Built HealthPointe multi-cloud baseline with Terraform, Ansible, and Jenkins
  • Built Greenwich.HR's labor-market intelligence platform: thousands of ECS jobs with queues, reporting, and customer management
  • Converted an IRI Worldwide production R job to Databricks Spark; runtime dropped to 15% of production
  • Built HealthPointe's multi-cloud Terraform/Ansible/Jenkins baseline with automated security scanning and detection
  • Ran technical interviews for client companies, from Spark to systems management

IRI Worldwide R-to-Spark conversion: Runtime dropped to 15% of production

Technologies: AWS ECS, Databricks, Spark, Terraform, Ansible, Jenkins

Sr. DevOps Architect specializing in Kubernetes — Nielsen

2018-02-01 – 2018-09-01

  • Designed Azure AKS/ACS interface to launch big data jobs into Kubernetes via Jenkins
  • Introduced CI/CD practices and trained other departments
  • Built a Vue.js/Express GUI for cluster and Jenkins job management
  • Designed Azure AKS/ACS interface to launch big data jobs into Kubernetes clusters
  • Created faster workflow to reduce job run times
  • Implemented CI/CD practices for faster code delivery
  • Built GUI interface for kubernetes/jenkins system using Vue.js/Express
  • Trained other departments on deployment methodologies

Technologies: Kubernetes, Azure AKS, Jenkins, Vue.js, Express

DevOps Manager — Simple Energy

2015-08-01 – 2017-04-01

  • Created new DevOps workflow focused on improving automation and system deployments
  • Reduced deploy time from 3-4 weeks to less than a day, and platform updates to minutes from hours
  • Implemented Docker/Mesos to increase uptime on volatile systems
  • Instantiated the use of Jenkins/GitLab to increase developer productivity and accountability
  • Reduced time to test for developers and reduced complexity to get to production
  • Improved testing strategies and implemented automated recovery for simple systems
  • Reduced costs across the board by automating services and deployments
  • Worked on various big data systems with Scala and Python for energy usage analytics

Technologies: Docker, Mesos, Jenkins, GitLab, Scala, Python

Sr. Developer / Sr. Systems Administrator / DevOps Manager — Promet Source

2015-01-01 – 2015-07-01

  • Helped design and build the DevOps team from the ground up
  • Transferred servers from clients to a unified infrastructure for smoother deploys
  • Designed and built new server infrastructure using Google Cloud, saving money
  • Implemented new tool stacks to increase team efficiency and automation
  • Reduced deployment time from days to hours with improved infrastructure
  • Managed several development projects in PHP, JavaScript, Java, and Python
  • Worked on a new Drupal module integrating Tika and Solr to index millions of PDFs
  • Built Jenkins automation cluster to streamline operations
  • Implemented Elasticsearch cluster for centralized logging
  • Developed Node.js applications for PDF conversion and IMAP mail

Technologies: GCP, PHP, JavaScript, Java, Python, Drupal, Tika, Solr, Jenkins, Elasticsearch, Node.js

Systems Administrator / DevOps Engineer — Time Warner

2014-05-01 – 2014-11-01

  • Used Jenkins, Docker, Mesos, and Puppet to implement CI/CD processes
  • Built new onboarding processes with Vagrant for faster developer setup
  • Introduced Docker as a tool for development and system administration
  • Designed big-data workflows using Docker to containerize platforms
  • Created clusters within single machines to test larger deployments
  • Designed workflows around Mesos/Marathon and built UI in Angular.js

Technologies: Jenkins, Docker, Mesos, Puppet, Vagrant, Marathon, Angular.js

Skills & expertise

Programming

Python, JavaScript, TypeScript, Go, C++, Swift, Java, C#, PHP, SQL, Scala

AI & machine learning

AI Agents, MCP, LLM Fine-tuning, RAG, Machine Learning, AI Serving Infrastructure, llama.cpp, PyTorch, LangChain, Hugging Face, AI Agents & MCP, LLM Serving & Fine-tuning, Forward Deployed Engineering, Kubernetes & Multi-cloud, Real-time Trading Systems, XR / VR / AR

Web, desktop & testing

React, Next.js, SwiftUI, Vue.js, Node.js, Express, AppKit, Chrome Native Messaging, Cordova, Playwright

Cloud platforms

AWS, EKS, ECS, EC2, S3, Transit Gateway, Azure, AKS, Azure OpenAI, Databricks, GCP, GKE, Compute

Infrastructure & delivery

Docker, Kubernetes, Mesos, FluentD, Prometheus, Jenkins, GitLab CI, Terraform, Ansible

Data & databases

PostgreSQL, MySQL, SQLite, Cassandra, Redis, Kafka, Apache Arrow, Spark

Networking & virtualization

Xen, VMware, Cisco, HAProxy, WebSockets, gRPC

Systems & delivery

Linux, Windows, Server Administration, Datacenter Management, Framework Design, Project Management

Extended reality

Unity, HoloLens, A-Frame, WebXR, VR Development, 3D Modeling

Additional tools used in roles

Phi-2, Backstage, Packer, AWS ECS, Azure AKS, GitLab, Drupal, Tika, Solr, Elasticsearch, Puppet, Vagrant, Marathon, Angular.js

Selected projects

QuestionMonitor

Production macOS app bridging a SwiftUI frontend to a Python PyTorch backend over a local WebSocket server. Sub-250ms ML pipeline with Silero VAD (90–95% fewer hallucinations), faster-whisper, and custom SetFit/transformer models.

VideoEditorPro

Local-first video editor using React, Node.js, and FFmpeg pipelines for pixel-perfect canvas encoding, compositing, and text-to-speech generation.

Alpaca Auto Trader (AAT)

Autonomous 24/7 trading system with a two-stage pipeline: a fast 7B scanner identifies candidates and executes trades. Intelligence stack includes news/research streaming, Twitter signals, QMD search across 100K+ documents, TimesFM forecasting, and automated trade autopsies.

Almagest

Native C++ llama.cpp engine with durable workflows, layered SQLite memory, MCP tools, and a live dashboard. Qwen3.8-27B on two RTX 3060s with MTP.

AI Cluster

Separate Go inference cluster on llama.cpp. Native MCP tool selection, high-speed agent context memory, agents, assistant, and omni voice chat.

GTMExtensions

End-to-end training pipeline using Chrome Native Messaging, TypeScript 5.x, and Webpack 5 to securely bridge browser extensions to macOS desktop apps over local TCP sockets.

Meeting Prep

macOS menu bar app with EventKit calendar sync and GRDB/SQLite for durable, structured AI agent meeting preparation.

Ticket Whisperer

Petco IT ticket automation integrated with ServiceNow. Fine-tuned Azure OpenAI models parse, categorize, and route tickets without human intervention.

Simple AI Router (SAR)

Go reverse proxy with autoscaling that routes between GPU backends, monitors load and memory, and scales inference across clouds and datacenters.

FluentD LLM Server

Log analysis with FluentD plus custom-trained Microsoft Phi-2 models on Kubernetes for real-time security threat detection and automated anomaly response.

BifrostPipe

High-performance C++ Kafka pipeline for real-time, high-frequency market data over WebSockets with low-latency transformation.

ORCA VR - Server Management Platform

Designed and developed software platform for managing server systems in virtual reality, providing immersive monitoring and management capabilities.

Hololens Oil Refinery Inspection

Designed Microsoft Hololens application for making oil refinery inspections easier to verify and manage, improving safety and efficiency.

VR Garage Platform

Built a platform enabling users to easily create VR video/image/3D model experiences using a web-based interface. Used JHipster microservices and A-Frame for the frontend, with Kubernetes and fabric8 for the backend.

Education

Colorado Mountain College — Coursework in programming, writing, and mathematics

Download the full résumé (PDF)