Research Engineer, HCI and Agent Experience
San Francisco / Mountain View, CA or remote · Full-time
Hey 👋, we’re Haz and Wyatt, the founders of Pally.
Pally is a personal assistant that lives in your text messages. You text it like a friend. It replies to the messages you’ve been avoiding, books the dinner, chases the invoice, and handles the hundred small things that pile up in a week.
We’re YC S25, based in the Bay Area, and growing fast. It’s the two of us right now, and we’re hiring the people who’ll make Pally the thing everyone has.
The role
You’ll own how Pally thinks and talks. The whole product is a text thread. There are no buttons, no spinners, no dashboard. Every product decision comes out of the model as a sentence, and every sentence costs latency and money. That makes the interface a research problem and an engineering problem at the same time, and you’ll own both.
This is a hands-on engineering role. You’ll be in the agent code, the prompts, the eval harness, and the model layer, shipping to real people every week. In practice that means:
- Agent loops. Design and build how Pally decides to act, ask, wait, or reach out first: the tool-calling loop, memory and context assembly, and the policy for long conversations that run for months.
- Model routing. Fine-tune and evaluate a router so small, fast models take most turns and frontier models take the ones that matter. Own the latency and cost budget per turn.
- Evaluations. Build the harness that turns “did that reply feel right?” into a number: turn-level judges, replay of real conversations, regression suites that run before every deploy.
- Prompting and context. Own the prompts and the context-window design so the model gets the right information in the right order, and nothing else.
- Agent infrastructure. The workers, queues, and tool integrations the loop runs on. You’ll ship it, not hand it off.
- User experience. How the product reads on a phone, from the first reply to a follow-up weeks later. You’ll read transcripts every day and fix what you see.
We’ll measure you on one thing: how much better Pally gets at being someone’s assistant, week over week, and whether the evals prove it.
Profile
You’ve built LLM agents in production, or fine-tuned models, or built an eval pipeline that people actually trusted. Ideally more than one of those. You’re fluent in TypeScript or Python and comfortable in a real backend: queues, databases, deploys, on-call.
You have an HCI instinct and an engineer’s hands. You care how a reply reads on a phone at 7am as much as the architecture behind it, and you can change both before lunch. You read papers and ship the idea the same week.
You want to be close to users, not behind a research org. You’re fluent in the latest AI tools and use them every day. You have strong opinions on how technology should shape the world, and you use Pally and have opinions about that too.
We don’t care about your CV. We care that you’ve built something real and made it better after it met real people.
Practical stuff
Competitive salary and substantial equity. Full-time, in person in San Francisco or Mountain View, CA, or fully remote for exceptional candidates.
We work exceptionally hard. If work isn’t your primary focus, this isn’t for you. But relationships are everything, and that applies to you too. You will be there for your family and friends when they need you. We don’t care when you work hard, only that you do.
If this sounds like you, apply below. Tell us how you see the future of personal AI and send us your resume.
Thanks for reading,
Haz + Wyatt
Apply