
Serving inference at scale is a scheduling problem. Requests arrive with wildly different shapes and latency expectations, GPUs are heterogeneous and expensive, and the difference between a platform that's fast and one that's economical usually comes down to how well work gets placed. That system is what you'll own.
You'll design and build the scheduling and routing layer of our platform: how requests get admitted, prioritized, batched, and placed across a heterogeneous multi-cloud GPU fleet, under real multi-tenant load and real latency commitments. This is core-systems work with a clean slate — you'll be making the foundational architectural decisions, not maintaining someone else's, and the quality of those decisions will show up directly in our margins and our customers' latency numbers.
You'll work close to the metal and close to the math. Some days that means reasoning about queueing behavior and control loops on a whiteboard; other days it means profiling Go or Rust until the tail latency comes down. We're a small team, so you'll own systems end-to-end — design, implementation, rollout, and the production reality afterward.
This is an individual-contributor role. We're looking for someone who has built systems like this before and can operate independently from the first week.
Nice to have: Kubernetes and multi-cloud operations experience; open-source contributions to inference, serving, or scheduling projects; experience operating GPU clusters at scale.
The base salary range for this position is $170,000 – $350,000 per year. The range reflects the position across experience levels; actual base salary will be determined by job-related knowledge, skills, experience, and work location, and may fall anywhere within the stated range.
In addition to base salary, this position is eligible for equity in the company, along with medical, dental, and vision coverage, and other benefits.
We are able to sponsor H-1B and other work visas for qualified candidates.
About NeuroSpark
NeuroSpark builds and operates a high-performance AI inference platform that helps enterprises run large language models faster, cheaper, and at scale. Inference infrastructure is the foundation the entire AI application layer runs on — every AI product ultimately depends on how fast, how reliably, and how affordably models can serve their users. Our vision is to make that layer so efficient that compute is never the reason a good AI product fails.
The pay range for this role is:
170,000 - 350,000 USD per year (Santa Clara Office)
Engineering
Santa Clara, CA
Share on: