Order.co is the System of Action for the Office of the CFO, transforming the way businesses purchase and pay into an intuitive, B2C-like shopping experience. Order.co leverages embedded AI agents and embedded financial products to reinvent the way businesses connect with their vendors.
End users enjoy a seamless, zero-training buying experience, while finance and procurement leaders gain a single platform to orchestrate how the business “should operate”. The result is an all-in-one solution that serves as a gravitational pull for spend and data, automating and eliminating procurement and finance workflows from requisition to reconciliation along the way.
Order.co is on the cutting edge of B2B Agentic Commerce, poised to be the market leader in creating a more predictive, prescriptive, and personalized experience for users.
Founded in 2016 and headquartered in New York City, Order.co oversees nearly half a billion in annualized spend across hundreds of customers like WeWork, SoulCycle, Lume, and [solidcore]. Order.co has raised $75M in funding from industry-leading investors like MIT, Stage 2 Capital, Rally Ventures, 645 Ventures, and more. Order.co has been proudly named a 50 to Watch by Spend Matters and a Best Place to Work by BuiltIn and Inc. Magazine.
Role summary
We are hiring a Staff Applied AI Scientist to own how AI works inside our product, from the architecture of the system through to the business outcome it produces. This is a senior individual contributor role that combines ownership of the AI and machine learning architecture with hands-on applied science. You will decide what the AI system should be, including how models are served, how retrieval and context assembly work, how prompts and model versions are managed, what the guardrails are, and how the whole thing gets evaluated. Then you will build it, ship it, and own how it behaves in production.
Because this role sits between two more familiar ones, it is worth being precise about the scope. It is not a modeling role that ends at a handoff, since you are accountable for the deployed system and the metric it moves rather than for the model in isolation. It is also not a data platform or pipeline role, because while you will define what the data and infrastructure must provide for your AI to work, and you will partner closely with data engineering and platform teams to get there, you will not own the warehouse.
You will be embedded day to day with a product and engineering squad while reporting into the data team, and you will have the head of data as your closest technical partner.
What you will own
- The AI and machine learning architecture. You will design the machine learning and agentic architecture end to end, covering model hosting and serving, prompt and model versioning, retrieval and embeddings, agent tooling, guardrails, and the evaluation framework that tells you whether any of it is working. You are the person who makes these calls and defends them.
- Applied modeling and evaluation. You will choose between deterministic and large language model or agent-based approaches on the merits, with explicit trade-offs across accuracy, latency, cost, and reliability. You will build evaluation that connects offline and online quality to business outcomes and risk controls, rather than treating model metrics as the goal.
- Production delivery and model operations. You will own the full lifecycle, including experimentation, versioning, continuous integration and deployment for both models and prompts, monitoring, drift detection, rollback, and incident readiness. You will define the rollout strategy and the likely failure modes before launch rather than after it.
- Responsible AI in practice. You will build safety guardrails, hallucination mitigation, bias testing, and appropriate handling of sensitive data into the design itself, so that our data privacy and prohibited use commitments are met by construction rather than reviewed at the end.
- The data requirements your AI depends on. You will specify what "AI-ready" means for each initiative, including the training and retrieval data, labeling, feature availability, and vector or search infrastructure you need, and you will work with data engineering and platform to make it real so that models run on governed infrastructure instead of shadow deployments.
- Business outcomes and prioritization. You will turn ambiguous goals into technical bets with clear hypotheses and success criteria, and you will own a portfolio of AI opportunities by identifying the highest-leverage ones, sequencing them, and building the execution path rather than waiting for assigned work.
- Technical direction and influence. You will advise product and engineering leadership on what is feasible, what it costs, what it risks, and what it is likely to return. You will set architecture and delivery patterns that other people reuse, and you will mentor experienced individual contributors on applied AI execution and production quality.
What we are working on now
- Predictive ordering, which means models that change how customers plan and place orders inside real procurement constraints such as organizational structure, budgets, approved catalogs, vendor contracts, and substitution rules.
- Agentic copilots for workflow management, where AI assistance is embedded directly in core product workflows and the interesting problems are tool design, retrieval quality, guardrails, and knowing when to hand control back to a person.
- The evaluation and operations layer underneath both of those, so that we can tell the difference between a model that looks good offline and a capability that actually moves the business.
What we are looking for
- At least 10 years in applied data science, machine learning, or applied AI, with repeated delivery of production systems that moved a business metric you can name.
- Ownership of AI and machine learning system architecture rather than models alone, including serving, retrieval, evaluation, guardrails, and the operational loop around them.
- Real depth in current large language model and agent technology, meaning you know where it works, how it fails, how you evaluate it, and when a simpler deterministic approach is the better answer.
- Demonstrated practice in machine learning operations, including versioning, continuous integration and deployment for models and prompts, monitoring, drift detection, and rollback.
- Portfolio-level ownership, where you have prioritized across competing AI opportunities and built the path forward rather than only executing a roadmap that someone else set.
- Heavy daily use of AI-native engineering workflows across design, coding, debugging, and review for at least the past 18 months, along with the judgment to know where to verify and what not to trust.
- Experience setting model governance, monitoring, and responsible AI standards for a team.
- Working implementation proficiency across at least two cloud or technical ecosystems, for example AWS and GCP.
- A strong quantitative foundation in experimentation, statistical reasoning, and causal thinking.
- The ability to align product, engineering, and operations stakeholders on sequencing and trade-offs when the situation is ambiguous.
Preferred qualifications
- Experience with retrieval systems, vector search, ranking, recommendation, or production personalization.
- Experience with self-hosted or local AI infrastructure, including self-managed agent environments.
- Experience in e-commerce, B2B procurement, vendor management, financial products, or heavy integration with external systems.
What success looks like in the first 6 to 9 months
- Multiple AI capabilities are live in production with clear hypotheses and measured outcomes.
- The model architecture and evaluation approach you established is adopted by other initiatives instead of being reinvented.
- Model operations hold up under real conditions, which means monitoring, drift detection, rollback, and incident playbooks that people actually trust.
- Product and engineering leadership plan against a prioritized view of our AI portfolio that you own.
- Low cycle time is the default, so that every release produces a signal we can evaluate.
Working model
You will work embedded with a product engineering squad on customer-facing capabilities, partnered with a principal-level scientist, and hands-on from day one. We ship iteratively and we judge work by the outcome it produces rather than by the size of the launch.
Interview process
The process consists of a conversation with the hiring manager, a take-home technical design review on a real predictive ordering problem that takes roughly 90 minutes and that we then discuss with you live, technical rounds covering modeling and evaluation as well as architecture and production operations, a conversation about business impact, a conversation about leadership and growth, and a values conversation with our People Ops team.