Founding ML Research Engineer
San Francisco, CA · Full-time, in person
Hey 👋, we’re Haz and Wyatt, the founders of Pally.
Pally is a personal assistant that lives in your text messages. You text it like a friend. It replies to the messages you’ve been avoiding, books the dinner, chases the invoice, and handles the hundred small things that pile up in a week.
Read our philosophy of workWe’re a small team based in the Bay Area, and we’re hiring people who will make Pally the #1 consumer agent in the world.
The role
You’ll make the models behind Pally better at being someone’s assistant. The whole product is a text thread, so every reply is a model output, and every model output costs latency and money. Which model answers, what it knows, how it uses a browser, and whether it does the right thing when someone tries to trick it are all research problems. You’ll own them, and you’ll ship the answers to real people every week.
This is a hands-on research and engineering role. You’ll be in the eval harness, the training code, and the agent itself. In practice that means:
- Evaluations. Build the harness that turns “did that reply feel right?” into a number: turn-level judges, human preference data, replay of real conversations, and regression suites that run before every deploy. Evals the team trusts to make ship decisions.
- Fine-tuning. Turn real conversations into training and preference data, and post-train and distill models (SFT, preference tuning, distillation) so small, fast models do what frontier models do well.
- Model routing. Train and evaluate a router so small models take most turns and frontier models take the ones that matter. Own the latency, cost, and quality tradeoff per turn.
- Browser agents. Make Pally reliable on sites with no API: navigation, sign-in, checkout, and knowing when to stop and ask. Measure it, then make it better.
- Red teaming. Find the ways Pally can be tricked, leak something, or take an action nobody asked for, from prompt injection in an email to a malicious web page. Build the defenses and the evals that keep them working.
- Agent experience. Bring ML across the agent: memory and retrieval, tool selection, when to act versus ask, and when to reach out first. Read transcripts every day and turn what you see into evals and model changes.
We’ll measure you on one thing: how much better Pally gets at being someone’s assistant, week over week, and whether the evals prove it.
Profile
This is a senior to staff level role. We’re looking for someone who has already done this work at a frontier AI lab or a team operating at that level.
- 4+ years of ML or software engineering, with experience at an AI lab, AI product company, or similar.
- Hands-on experience post-training or fine-tuning LLMs (SFT, preference tuning, distillation) and shipping the results to production.
- Experience designing evals, such as LLM judges, human preference data, or replay of real traffic, that a team trusted to make ship decisions.
- Experience in at least one of: red teaming or adversarial testing of LLMs; browser or computer-use agents; model routing or cascades; production agents with tool use and long-running memory.
- Strong Python and the ML stack for training and evals, and comfortable working in TypeScript in the product codebase.
- Fluency with AI-assisted development workflows, grounded in the fundamentals to verify, correct, and own everything you ship.
- You want founding-team scope. You’ll set the research bar for Pally and help hire the engineers who join after you.
Practical stuff
$200–300k base + significant equity, 401(k), and health insurance. Full-time, in person in San Francisco, CA.
We work exceptionally hard. If work isn’t your primary focus, this isn’t for you. However, we also give you full autonomy over your schedule. We don’t care when you work hard, only that you do.
If this sounds like you, apply below. Tell us how you see the future of personal AI and why you’re excited to work at Pally.
Thanks for reading,
Haz + Wyatt
Apply