SpreeAI

MLOps Engineer

About the team

AI Platform scales SPREEAI's infrastructure: productionizing ML model checkpoints, running the API services behind the Partner Portal, optimizing inference serving, and giving ML Scientists training-as-a-service so they can iterate without managing infrastructure themselves.

About the role

This role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.

What you'll do

  • Design and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly
  • Build CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks
  • Own experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly
  • Monitor training job health (GPU utilization, loss curves, OOM detection) and drive cost governance (spot instances, preemptible VMs, budget alerts) as training costs scale
  • Partner directly with ML Scientists to translate workflow pain points (reproducibility, experiment comparison, checkpoint recovery) into platform abstractions
  • Evaluate and integrate external model providers (SPREEAI uses Byteplus and Fireworks.AI as MaaS providers) into the training and eval platform

What you'll bring

  • Experience building or operating an ML platform: training orchestration (Ray, Kubeflow, Airflow, or a custom solution), Docker, Kubernetes (Jobs/CronJobs, Helm)
  • Python or Go for pipeline orchestration and infrastructure tooling
  • Real experience with experiment tracking and data versioning tools at production scale
  • Comfort owning platform direction, not just executing tickets, and mentoring engineers as the team scales
  • Comfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment

You'll thrive here if

You treat ML Scientists as your customers and actively design the boundary between platform responsibility and scientist responsibility rather than letting it happen by accident. You can make the case for a platform investment that isn't obviously urgent yet, and you're comfortable saying no to a feature request that would compromise platform integrity.

SPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce, revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences. Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology. We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact. If you are passionate about innovation and shaping the future of fashion, SPREEAI offers a platform to make your mark.

Die Gehaltsspanne für diese Rolle ist:

145,000 - 180,000 USD pro year (Hybrid (San Francisco, California, US))

Engineering

Hybrid (San Francisco, California, US)

Teilen auf:

NutzungsbedingungenDatenschutzCookiesPowered by Rippling