Sanas

Member of Technical Staff, ML Inference Engineering

Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross-language communication, and more.

Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language.

Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$120M ARR. Our team combines deep expertise in model innovation and systems engineering with a design-minded product engineering culture to build and ship cutting-edge AI models and experiences — entirely in-house.

Sanas is a 130 person team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies. With Sanas, you're not just adopting a product; you're investing in the future of communication.

If you’re looking to have a significant role in roadmapping and driving technical directions, if you’re looking to deploy challenging and big ideas without much overhead or slowness, if you're looking to leave your mark on an ambitious, generational mission to change how the worlds thinks about speech + AI, then Sanas is a well-suited place for you.

About the Role

Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.

We're looking for a deeply hands-on, experienced engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

What You'll Do

Inference Optimizations

●      Implement custom kernels and low-level optimizations to push the absolute limits of GPU compute.

●      Apply graph optimization, operator fusion, and hardware-specific code generation at the ML compiler level.

●      Profile, analyze, and resolve deep system bottlenecks to radically improve latency, throughput, and memory efficiency.

●      Drive model-level execution improvements, including mixed precision and advanced quantization strategies.

Serving Optimizations

●      Write and optimize custom inference server backends to handle complex business logic, scripting, and state management.

●      Own and scale our overarching model serving infrastructure across multi-GPU, multi-node deployments to meet strict on-premise latency budgets.

●      Build robust, fault-tolerant runtime services, complete with comprehensive performance benchmarking and monitoring.


Requirements

●      8+ years of experience writing high-quality, high-performance code, including at least 5 years focused on machine learning systems.

●      Experience with one of C/C++/Rust and Python.

●      Deep familiarity with modern NVIDIA GPU architectures (e.g., Ada Lovelace, Blackwell), CUDA, and low-level system profiling.

●      Hands-on experience with ML compilers and optimization frameworks (e.g., Apache TVM, TorchInductor/Dynamo, TensorRT, Triton).

●      Experience building or extending model serving infrastructure, specifically writing custom C++ backends for Triton Inference Server (or similar serving engines).

●      Fluency in the AI serving stack, from kernels and quantization up to schedulers, state management, and autoscaling.

●      A record of shipping research or systems that other people build on, whether in a lab or in industry.

Nice-to-have:

●      Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes.

●      A research-leaning or systems background in Speech (STT, TTS, S2S) or LLM inference, with work you can point to.

●      Experience operating large-scale, on-premise AI training or inference clusters, including bare-metal provisioning and Kubernetes management.

●      Familiarity with high-performance networking (InfiniBand or RoCE), distributed storage systems, and hardware health monitoring.

●      Experience maintaining or contributing to open-source ML or systems projects.

Engineering

Palo Alto, CA

Share on:

Terms of servicePrivacyCookiesPowered by Rippling