AI Engineer · Knoxville, TN

I build the
infrastructure
that runs AI.

Two GPU servers, six custom-tuned models, an autonomous agent framework, and a knowledge graph spanning 600K+ nodes — all running in my homelab, deployed into production, and built from scratch. I'm an AI engineer who doesn't just use models: I deploy them, fine-tune them, and build the systems around them that make them actually useful.

35B Production model
4 GPUs across 2 servers
94% Tool call reduction
18M+ Weekly watch minutes
2x
Intel Arc Pro B70 — vLLM XPU server
2x
AMD Radeon 7800 XT — ROCm server
100+
Tokens/sec generation throughput
5K+
Tokens/sec prefill throughput

What I work with

LLM Inference & Deployment
vLLM Intel XPU · AMD ROCm Qwen3.6 35B Q8 Qwen3.5 2B · 0.8B Cactus Needle 26M tool-caller INT4 · NVFP4 quantization HAProxy LLM load balancing Triton inference backend
Custom Models & Training
Cactus Needle 26M — sub-ms tool selector Qwen3.5 0.8B — intent router classifier Qwen3.6 35B Q4 — web extract / compaction Qwen3.6 35B — primary agent reasoning BERT-NER fine-tuning (90%+ acc, 8 entity types) bge-small-en-v1.5 embedding pipeline Embedding quality monitoring
Agent Frameworks
Hermes Agent — autonomous multi-agent MCP Model Context Protocol Intent routing & semantic tool filtering Knowledge router architecture mem0ai persistent memory Self-improving agent loops
Knowledge & Data
GraphRAG · unified RAG · wiki RAG ChromaDB vector store 600K+ node codebase knowledge graph Leiden community detection Crawl4AI web extraction
Infrastructure
Proxmox VE · PCIe GPU passthrough Docker · systemd microservices GitHub Actions CI/CD Cloudflare Zero Trust tunnel NGINX · HAProxy · reverse proxy Python · FastAPI · Node.js
CTV Platforms
Roku BrightScript · SceneGraph Samsung Tizen · LG webOS Fire TV · Android TV · Kotlin HLS · DASH · FFmpeg pipelines New Relic · LaunchDarkly

Where I've shipped

TEGNA Senior Roku & CTV Developer
Jul 2025 – Present
Knoxville, TN · 49-station streaming platform
Scaled to 18M watch minutes & 1.3M sessions weekly; reduced crash ratio 7× (0.02% → 0.005%) — top producer in department
Reduced analytics bill 98.7% while retaining core data; uncovered New Relic anomaly saving $1M/year
Led complete Roku UI/UX overhaul across 49 stations — single-handedly, zero production issues
Built Fire TV + Android TV app in Kotlin/Jetpack Compose in 2 weeks — 80% UI, 50% backend
Root-caused Roku OS 15.2 crash regression across 18M+ weekly plays; shipped Vizbee casting with zero defects
Served as interim CTV team lead; onboarded 4-person team, documented architecture decisions
Social Media Ministries Roku Developer (Freelance)
2023 – Present
Remote
Built full-stack Roku app from concept to Roku Store certification — BrightScript, SceneGraph, deep linking
Designed content pipeline for sermons, devotionals, and on-demand streaming
Manteno Church of the Nazarene OTT Developer & DevOps
Jul 2023 – Jul 2025
Manteno, IL
Built Roku app from zero to Store: certified, deep linking, responsive UI/UX
Provisioned bare-metal + cloud: Dell servers, hypervisor, Ubuntu REST API, reverse proxy, VPN, DNS
Automated video ingest with FFmpeg titling, thumbnailing, JSON feed generation
Netdata + Uptime Kuma monitoring; weekly security scans, automated backups

AI infrastructure I built

Two dedicated GPU servers. Every model, microservice, and tool deployed and maintained in-house. No cloud credits — just hardware, Linux, and an unreasonable amount of vLLM config tuning.

vLLM Intel XPU

AI Server 1 — Intel Arc

2× Arc Pro B70 · PCIe 5.0 M.2→x16 riser
Qwen3.6 35B Q8 — primary agent model, ~100+ t/s gen, 5K-15K+ t/s prefill
Cactus Needle 26M — dedicated sub-ms tool selector, 94% call reduction
Qwen3.5 0.8B — lightweight intent router, reduces tool surface 21→2
vLLM XPU with tensor parallelism 2, reasoning loop mitigation, HAProxy load balancer
ROCm vLLM

AI Server 2 — AMD ROCm

2× Radeon 7800 XT · ROCm stack
Qwen3.6 35B Q4 — primary agent + web extract / compaction (REST API)
mem0ai — persistent memory with ChromaDB, scene extraction, persona synthesis
BERT-NER fine-tuning pipeline — 90%+ on 8 entity types (FILE_PATH, URL, SKILL_NAME, etc.)
Embedding quality monitoring via bge-small-en-v1.5 with structured log output

✦ Cross-stack pipeline

Hermes Agent — 4 systemd microservices (Intent Router 8765, Entity Extractor 8766, Sentiment 8768, MCP Matcher 8770) + 15min health cron
Knowledge Graph — 600K+ nodes, Leiden clustering, GraphRAG community reports, cross-codebase semantic search
8 microservices — Intent Router, Entity Extractor, Supra Title, Sentiment, Web Extract, MCP Matcher, Health Monitor, Knowledge Router
Proxmox VE — multi-VM with PCIe passthrough, GitHub Actions CI/CD runner for 6 CTV platforms

Systems I've built

Hermes Agent System
Autonomous agent framework with intent routing, multi-agent delegation, semantic tool filtering, and knowledge graph integration. 4 systemd services, health monitoring, 94% tool call reduction.
PythonFastAPIvLLMMCPsystemd
Custom Small Model Stack
Six purpose-tuned models: Cactus Needle 26M for tool calling, Qwen3.5 0.8B for routing, Qwen3.6 35B Q4 for web extraction, Qwen3.6 35B Q8 for reasoning. Each optimized for its specific role.
Cactus Needle 26MQwen3.5 0.8BQwen3.6 35BvLLM
Knowledge Router & GraphRAG
600K+ node knowledge graph with fuzzy entity search, Leiden clustering, community reports, and unified RAG across codebase memory, wiki, and docs for semantic code search.
GraphRAGChromaDBBM25Vector Search
Web Extract & Compaction Service
Qwen3.6 35B Q4 microservice on ROCm for web content extraction, compaction, and structuring. REST endpoints at /extract, /compact, /structure with LangChain @tool() integration.
FastAPIQwen3.6 35BROCmLangChain
mem0ai Persistent Memory
Custom memory system with NER retraining (BERT-base, 90%+ on 8 entity types), ChromaDB vector store, scene extraction, persona synthesis, and embedding quality monitoring.
mem0aiChromaDBBERTtransformers
CTV AI Knowledge Base
Cross-platform knowledge system covering Roku, Fire TV, Android TV, Tizen, webOS. Indexed official docs and community patterns for semantic CTV development search.
RokuKotlinMCPCodebase Memory

Credentials

freeCodeCamp
Web Development, JavaScript Algorithms & Data Structures
CompTIA A+
Hardware & Troubleshooting Certification
CompTIA Network+
Networking Fundamentals Certification
CompTIA CIOS
IT Operations & Infrastructure Certification

Let's build something.

I'm looking for an AI Engineer role where I can ship production systems, deploy LLMs, and build models that actually drive products.