Head of Inference

Canyon Code
Canyon Code

Software Engineering · Full-time

San Francisco, CA, USA

Posted on Sep 14, 2026

Why This Role, Why Now

Running enterprise AI workflows is one of the four pillars of Canyon Code, alongside Design, Deploy, and Optimize. Run is where we take open-weight models, put them on capacity we have procured, and serve the endpoints our customers point their workflows at.

We are ready for someone to own that pillar full time. This is a founding-team role, and the person in it will own the inferencing motion at Canyon Code and build the team underneath them.

About the Role

We are looking for a Head of Inference to own the whole path, working directly with Canyon Code's CEO and Chief Architect: procuring bare-metal GPU capacity from neoclouds and colocation facilities, bringing open-weight models up on that capacity, and serving the endpoints customers actually run against. This is hands-on. You will be in the system, not only directing it, and you will hire the team under you as the motion grows.

You're a fit if you…

  • You have done this end to end, more than once, within the last couple of years: procured bare-metal GPU capacity from a neocloud or colocation facility, brought open-weight models up on it, and served production endpoints off it.

  • You understand this concretely rather than in the abstract. You know the nuances, where it goes wrong, and how long each step actually takes.

  • You have procured bare-metal GPU capacity yourself: evaluating providers, negotiating terms, taking delivery, and getting it production-ready.

  • Deep hands-on experience with vLLM, SGLang, TensorRT-LLM, or NVIDIA Dynamo.

  • You understand KV-cache management, continuous batching, paged attention, speculative decoding, and quantization, and you know which of them actually pay off for a given workload.

  • You have owned a cost-per-token number and moved it.

  • Strong systems and operations skills: Kubernetes, containers, networking, storage, autoscaling, and observability. On bare metal you own more of the stack than you would on a hyperscaler, and you are comfortable with that.

  • Comfortable across heterogeneous accelerators: NVIDIA plus at least one of AMD, Groq, Cerebras, Intel, or TPU.

  • You can hire, mentor, and build a team under you.

  • Self-driven; thrives in ill-defined, ambiguous, early-stage work.

  • You have done this inside an inference provider such as Baseten, Fireworks, Lightning, Together, or Modal.Plus

  • Existing relationships with neoclouds or colocation providers.Plus

  • You have worked on multi-agentic workloads, where a single request fans out across many model calls.Plus

  • Published or open-source work in inference optimization.Plus

Why Join

Own a pillar

Run is one of the four things Canyon Code is built on. This is not a supporting function; it is the pillar that has to work for the rest of the vision to hold.

Founding team

You own the inferencing motion and build the team under you. The architecture, the providers, the hardware mix, and the hiring are yours.

A proven advantage to compound

CanyonOS is already 3x better at harness scaling and 3x better at LLM serving. You start from a measured edge rather than a hypothesis.

The buy-versus-build calls are yours

Capacity, providers, and serving stack decisions sit with you, with the money in the bank to act on them.

Work alongside the founding team

Daily collaboration with a CEO and Chief Architect who have deep experience in ML systems research and AI infrastructure platforms.

Canyon Code is an equal opportunity employer.

Apply for this job

Drag and drop or click to upload.
No
No
Tell us why you are a good fit, add a cover letter or anything else you want to share.
To withdraw or update your application, email applications@getro.com