Careers / Open role
Founding Member of Technical Staff
San Francisco · Full-time · In-person · 2 openings
We review applications on a rolling basis and will close this listing once we've found the right people.
Who Is Neolithic?
Neolithic is a nonprofit startup in San Francisco. Our primary goal is to mitigate catastrophic risk from AI. To do this, we're building open-source, agentic tools that automate and scale AI safety research.
We build the infrastructure that AI safety research runs on, and the scope is wide: tooling for AI control experiments, pipelines that automate and improve safety evaluations, datasets for training stronger safety monitors, and better plumbing — like context management and catching when agents ignore or botch tests — for the research agents that safety orgs increasingly rely on. If it removes a bottleneck for the field, it's in scope.
Neolithic was founded in 2026 by Leo McKee-Reid, whose background spans scheming evals, tools for automated interpretability, and engineering orbital rockets. We work closely with researchers at leading AI safety organizations, who use and shape what we build.
Why This Role Matters
AI capabilities research is increasingly done by AI. Safety research is still mostly manual. Closing this gap between capabilities and safety is crucial for mitigating existential risk.
The field is full of concrete problems that good engineering can just solve: evaluation logs are slow to interpret and easy to misread, red-teaming is done largely by hand, safety RL env datasets lack scale, and models are increasingly eval-aware. Almost nobody is building shared tooling for any of this, so a single well-built tool can speed up the entire field.
Join this early and you shape more than the tools. As one of the first two people at Neolithic, you set the technical direction, the culture, and the hiring bar for everything that comes after. Few roles offer this much counterfactual impact.
What You'll Do
Build the tools. Most of the job is scoping a tool, quickly building the first version, iterating on researcher feedback, and then maintaining it. Expect a lot of ownership, and treat everything you ship as production software that researchers can depend on.
Work with the users. Researchers at safety organizations are our design partners. You'll work with them to find where their workflows are slow, brittle, or manual, build tooling that removes those bottlenecks, and refine it with them until it works.
Build the org. Help set the technical direction, the org culture, and the hiring bar. We're an early-stage startup, so expect to wear many hats and take ownership of whatever most needs doing.
Some concrete projects and tools you'll work on in your first few months:
- Rogue-deployment datasets. The one already underway, and likely the first thing you'll touch: an automated pipeline that generates realistic samples of AI agents misusing their access inside a frontier lab's systems — training data for stronger cybersecurity monitors.
- Automated safety evaluations. Agentic workflows that take a threat-model description and produce a runnable, verified evaluation.
- Evaluation awareness mitigation. Tooling for building realistic environments: raising environment realism until models can't tell they're being evaluated.
- Agent infrastructure for safety research. Context-management systems and other plumbing that make safety orgs' research agents markedly more capable.
- Automated reward hacking detection. Tools that detect reward hacking in the RL environments foundation models are trained on, mitigating the emergent misalignment it causes.
Who You Might Be
This is one posting for both of our first two hires, and it's intentionally broad as we're more interested in high agency, mission alignment, and fast technical learners. That being said, we're most interested in strong engineers who have experience with agents and developing production software. These are the three types of profiles we're looking for:
- Agents engineer. You have deep, hands-on experience building agents and the infrastructure they run on. You have a strong intuition for what current models can and can't do, and care about building and maintaining robust software.
- Technical generalist. Your head supports the wearing of many hats. You can ship fast, build product, conduct user interviews, and think clearly about what's worth building.
- Researcher turned builder. You have technical AI safety research experience (evals, control, interp, red-teaming), you're trending toward AI-coding power user, and you're excited to spend the next 4 months mostly engineering, with the research side of the role growing over time.
Whoever you are, we expect most of this to sound like you:
- Strong engineer: you've shipped and maintained production code.
- Hands-on with LLMs and agents (scaffolds, evals, or tooling), or something similar with clear evidence that you pick new things up fast.
- An AI power-user: you use frontier models constantly, have a good sense of what they can and can't do, know how to avoid slop, and enjoy experimenting with new workflows.
- High agency: you're able to identify what's most needed for the mission and take the initiative, often without clear guidance. You're excited to come up with new strategies and have a major influence in figuring out how we can scale the field as fast as possible.
- Genuinely motivated by reducing risk from advanced AI.
We'd be especially interested if you have:
- Built or maintained a widely used open-source tool or library.
- Done safety research at a lab, safety org, or fellowship program (MATS, LASR, etc.).
- Founded something, or been engineer #1–3 somewhere at a startup.
- Built serious agent scaffolds or eval pipelines.
If you don't tick every box but believe you'd make Neolithic better, apply anyway. Strong candidates can come from a variety of unconventional backgrounds.
What We Offer
- Counterfactual impact. Two hires will roughly triple this organization's capacity. Almost nothing you build here would have happened anyway.
- Ownership and room to grow. Titles start general on purpose. As we raise and scale, founding hires are first in line for the senior roles that emerge — we'd rather promote from within than hire over you.
- $180K+ USD salary, plus benefits through our fiscal sponsor, BERI.
- San Francisco-based role. US visa sponsorship available.
- Freedom and autonomy. Flexible hours, project ownership, no management layers, and unlimited PTO.
Apply Today
Applying is light: email leo@neolithic.org with your CV and links to any impressive things you've built — repos, papers, tools. A few sentences on why you want to join Neolithic and what you'd bring to the team is enough.
Interview Process
- 15-minute call with Leo
- 3-hour work test
- 50-minute interview
- In-person work trial (paid)
For any questions about the role, email leo@neolithic.org.