Sr./Staff Forward Deployed Engineer
As a Sr./Staff Forward Deployed Engineer focused on AI Infrastructure at Groq, you will work at the frontier of large-scale AI systems, taking complex customer infrastructure programs from requirements to working production environments. You’ll help bring some of the newest accelerator and infrastructure technologies into production, spanning next-generation NVIDIA GPU systems alongside Groq’s purpose-built inference platform.
Based in one of Groq's hiring hubs (Dallas, San Francisco or New York City area); remote work is allowed until the local office opens, then onsite.
What you'll do:
- Own technical execution across complex customer engagements, from discovery and architecture through PoCs, demos, deployment, cluster bring-up, validation, acceptance, production readiness, and operational handoff.
- Translate incomplete or ambiguous customer requirements into practical architectures, implementation plans, test criteria, runbooks, and concrete engineering actions.
- Work hands-on across Linux, bare-metal infrastructure, Kubernetes and Slurm, networking, storage, observability, automation, and Groq platform integrations to bring customer environments online and resolve issues.
- Support large-scale GPU and LPX deployments, including infrastructure bring-up, cluster health and performance validation, workload testing, benchmarking, failure isolation, and production-readiness evidence.
- Understand customer AI workloads well enough to reason about training and inference behavior, concurrency, throughput, latency, data movement, caching, scheduling, and infrastructure bottlenecks.
- Lead technical portions of customer discovery, architecture reviews, demonstrations, and proofs of concept, clearly explaining design choices, tradeoffs, performance results, and risks to both engineering and business stakeholders.
Full details and application on the official Groq careers page.