Why this role exists
We're a technology company building AI-native learning products used by more than seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both run on the Learn Engine, the intelligence that moves a learner from initial understanding to demonstrated mastery. We work with partners like NVIDIA to bring responsible, learning-first AI to the students who need it most.
We know our platform works — internally. Mastery scores go up, engagement is strong, students come back. What we don't yet have is proof that would satisfy someone with no reason to take our word for it: a district procurement reviewer, an academic peer reviewer, a funder deciding whether to bet on us. That proof has to be built differently than an internal dashboard. It needs outcome measures independent of our own scoring, research designs that hold up to outside scrutiny, and evidence tiered against standards we don't get to define ourselves.
That's this role. You'll design the studies, own the instruments, and build the evidence base that lets StudyFetch and Honen make effectiveness claims we can defend to anyone — not just to ourselves. This is a founding-team role. You'll work directly with the people making the decisions, and the standard of proof you set is the one every future study gets held to.
What we believe
Every learner deserves the chance to succeed. StudyFetch started with one idea: high-quality, personalized learning should be within reach for anyone, at any stage of life. Honen carries that belief into the workforce.
Accessible to everyone. Learning should reach every person, whatever their background, role, or resources.
Meet people where they are. Every course adapts to a person's pace, their level, and the way they learn best.
Learning never stops. From a first job to a new career, people keep growing at every stage of life.
We hire people who share this conviction. The work is demanding and the hours can be long, and what sustains you through it is caring whether a real student finally understands the material.
What you'll own
- The evidence strategy. Nobody has told you which effectiveness claims to go prove — that's your call. You'll decide what's worth the investment to demonstrate, which standard to hold it to (ESSA/WWC-aligned for education, equivalent standards for workforce training), and be the one who says plainly when a claim isn't defensible yet.
- The measurement gap. Our internal mastery scores and post-tests are generated and graded by our own platform, which means they can't independently prove the platform works. You'll design the outcome measures and study designs — quasi-experimental, randomized, whatever the question requires — that get us past that ceiling.
- A benchmark that outside researchers will accept. You'll build the rubric, the codebook, and the interrater reliability process (human and AI raters) behind a tutoring-effectiveness benchmark meant for external release — reviewed by people with every incentive to find the holes in it.
- External credibility. You'll recruit and run the independent advisory panel and methods reviewers, lead peer-review and publication efforts, and represent our research to academic partners, procurement evaluators, and press. When we say "externally validated," you're the reason it's true.
- Research that gets cheaper every time. The first study is always the expensive one. You'll build the playbook, the reusable instruments, and the operating patterns so the next study — and the one after that — don't start from zero.
- Ethics and governance across populations. IRB protocols, human-subjects determinations, data-sharing agreements — across K-12, higher-ed, and enterprise learners, each with different rules and different risks.
- The measurement bar before it becomes a launch blocker. You'll sit close enough to product and go-to-market to say, early, what evidence a claim will require — instead of finding out after launch that nobody can defend it.
What we're looking for
You're a strong fit if either of these is true:
- A PhD in Learning Sciences, Educational Psychology, Psychometrics, Applied Statistics, or a related quantitative field, plus 7+ years applying it to real research programs, or
- Fewer credentials on paper and a track record of building evidence that held up to outside scrutiny — ESSA/WWC-aligned or equivalent, published, reviewed, defended. Show us the work.
Beyond that: - You've built evidence for a real audience, not just an internal one. You can walk us through a study or benchmark you designed, what an external reviewer challenged, and what you'd design differently now.
- You're rigorous about causality. Experimental and quasi-experimental design, sample and power, and knowing the difference between "this is proven" and "this is our best current guess."
- You write and speak clearly. You'll present to engineers, to the founders, and to external academic partners in the same week. You state uncertainty as a number when you can and in plain words when you can't.
- You've navigated IRB and human-subjects research before, and you don't treat it as paperwork — you treat it as part of the design.
- The mission is why you're here. What sustains you through the hard weeks is the learner on the other end, the one a defensible study proves we actually helped.
The stack you'll work in
You don't need every item below, but you should be deep in most and able to ramp quickly on the rest:
- Research design: experimental and quasi-experimental methods, pre-registration, sample and power analysis, causal inference
- Measurement: psychometrics, item response theory, rubric and codebook development, interrater reliability (human and LLM-assisted raters)
- Evidence standards: ESSA/WWC tiers or equivalent, IRB protocols, human-subjects research, data-sharing agreements
- Analysis: Python (Pandas, NumPy, SciPy) or R, expert-level SQL, hypothesis testing at scale
- AI/LLM familiarity: eval frameworks, LLM-as-judge and its failure modes, enough fluency to evaluate whether an AI-assisted rating pipeline is trustworthy
- Reporting: technical reports and evidence packages that survive an external reviewer's first read
- Prior experience with children's data and the governance rules that come with it is a real plus, as is experience publishing or presenting research externally.
What to expect
- This is an in-person role at a fast pace, with periods of intense work around major evidence deadlines and external reviews.
- You'll have significant ownership and autonomy with limited oversight. The role suits researchers who do their best work with room to run — there's no team here yet, and you'll be deciding the standard as you go, not inheriting one.
- You'll be the first person in this seat. Some weeks are rubric design, some weeks are an IRB submission, some weeks are defending a study design to an external panel, and you'll have to decide which one matters most that week.
- It's a strong fit for people who have built evidence that changed a decision — a procurement outcome, a published claim, a funding call. If your experience has been mostly internal reports that stayed internal, this likely isn't the right match.
Compensation & benefits
- $140,000–$170,000 base salary, plus equity
- 100% employer-paid Medical, Dental, and Vision; 75% dependent coverage
- 401(k) with employer matching
- Daily team dinner provided in-office
- A small, mission-driven team changing how the world learns
#LI-SF1