NeuroSpark

Member of Technical Staff, Kernel Engineer

About the Role

NeuroSpark Inc operates an enterprise AI inference platform providing high-throughput, low-latency access to large language models across a distributed, heterogeneous compute fleet.

As a Member of Technical Staff, Kernel Engineer, you will work at the lowest performance-critical layers of NeuroSpark's inference stack, designing and optimizing GPU and accelerator kernels that directly determine model latency, throughput, memory efficiency, and hardware utilization.

You will work across the boundary between model architecture, GPU execution, inference runtimes, and distributed serving systems. The role involves identifying performance bottlenecks in real production workloads, developing hardware-aware optimizations, and integrating those improvements into the systems that serve models at scale.

This is a hands-on individual-contributor role with significant technical ownership. You will work closely with engineers across inference, distributed systems, and infrastructure to push model performance across current and emerging accelerator platforms.

Responsibilities

  • Design, implement, and optimize high-performance GPU kernels for performance-critical AI operations, including GEMM, attention, normalization, quantization, KV-cache operations, and Mixture-of-Experts (MoE) workloads.
  • Develop and optimize kernels using technologies such as CUDA, Triton, C++, PTX, CUTLASS, and related GPU programming frameworks.
  • Profile production inference workloads to identify bottlenecks across compute, memory bandwidth, memory hierarchy, kernel launch overhead, synchronization, and data movement.
  • Optimize GPU execution through techniques including memory coalescing, shared-memory utilization, tiling, warp-level programming, Tensor Core utilization, operator fusion, latency hiding, and compute/communication overlap.
  • Improve end-to-end LLM inference performance across latency, throughput, memory utilization, concurrency, and hardware efficiency, rather than optimizing kernels in isolation.
  • Develop and optimize kernels for modern model architectures, including Transformer-based LLMs, attention variants, MoE models, and emerging model architectures.
  • Implement and evaluate lower-precision execution and quantization strategies, including FP16, BF16, FP8, FP4, INT8, and other hardware-supported formats.
  • Integrate optimized kernels and operators into inference frameworks and internal runtimes built around technologies such as PyTorch, Triton, vLLM, SGLang, TensorRT-LLM, or equivalent systems.
  • Use profiling and performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, roofline analysis, and internal benchmarking infrastructure to diagnose and resolve performance regressions.
  • Build reliable benchmarks, correctness tests, and performance regression tests to ensure kernel improvements remain numerically correct and production-ready.
  • Optimize workloads across multi-GPU and distributed environments, working with the broader infrastructure team on communication, parallelism, and compute efficiency.
  • Help extend NeuroSpark's inference stack across heterogeneous hardware, including NVIDIA GPUs, AMD GPUs, and other current and emerging AI accelerators.
  • Work closely with hardware vendors, inference engineers, and distributed-systems engineers to evaluate new accelerator architectures and translate hardware capabilities into production performance improvements.
  • Contribute to architectural decisions affecting the Company's inference runtime, model execution layer, and hardware-performance roadmap.

Qualifications

  • Strong experience in GPU programming, kernel development, high-performance computing, ML systems, or performance engineering.
  • Proficiency in C++ and hands-on experience with CUDA, Triton, or another accelerator programming model.
  • Strong understanding of modern GPU architecture, including:
    • GPU memory hierarchy
    • Threads, warps, blocks, and grids
    • Shared memory and register usage
    • Memory bandwidth and access patterns
    • Tensor Cores
    • Synchronization and parallel execution
    • Occupancy and instruction-level parallelism
  • Experience profiling and optimizing GPU workloads for latency, throughput, memory usage, and hardware utilization.
  • Experience with performance-critical machine-learning operations such as attention, GEMM, quantization, KV cache, MoE, or other Transformer operators.
  • Familiarity with PyTorch and modern ML inference or training execution stacks.
  • Strong understanding of numerical correctness, floating-point behavior, and mixed-precision computation.
  • Ability to reason from first principles about performance bottlenecks across hardware and software layers.
  • Strong debugging skills and the ability to take performance work from profiling and hypothesis through implementation, benchmarking, and production deployment.
  • Strong written and verbal communication skills and the ability to work effectively in a highly collaborative engineering environment.

Nice to Have

  • Experience with PTX/SASS, CUTLASS, CuTe, CUB, Thrust, or other low-level GPU libraries and programming abstractions.
  • Experience with inference frameworks such as vLLM, SGLang, TensorRT-LLM, FlashInfer, or similar systems.
  • Experience implementing or optimizing FlashAttention or other fused attention kernels.
  • Experience with ROCm / HIP and AMD GPU architectures.
  • Experience optimizing workloads across multi-GPU or multi-node systems, including NCCL and collective communication.
  • Knowledge of ML compiler and runtime systems such as torch.compile, XLA, MLIR, TVM, or related compiler stacks.
  • Experience with distributed training and inference, tensor parallelism, pipeline parallelism, or expert parallelism.
  • Experience optimizing workloads on multiple accelerator architectures or developing hardware-portable kernels.
  • Contributions to open-source projects in GPU kernels, ML systems, inference engines, compilers, or high-performance computing.
  • Experience bringing new model architectures or accelerator platforms into production.

Compensation & Benefits

The expected base salary range for this position is:

$170,000 – $350,000 USD per year

Actual compensation will depend on experience, technical depth, level, and role scope.

This position also includes:

  • Equity / stock options
  • Medical, dental, and vision coverage
  • Unlimited PTO
  • Opportunities for significant technical ownership and impact
  • H-1B and other work visa sponsorship available

About NeuroSpark

NeuroSpark builds and operates a high-performance AI inference platform that helps enterprises run large language models faster, cheaper, and at scale. Inference infrastructure is the foundation the entire AI application layer runs on — every AI product ultimately depends on how fast, how reliably, and how affordably models can serve their users. Our vision is to make that layer so efficient that compute is never the reason a good AI product fails.


Engineering

Santa Clara, CA

Compartir en:

Términos de servicioPrivacidadCookiesPatrocinado por Rippling