Evals Lead
About Aslan
Our national security apparatus was designed for an adversary that congregated and operate in the physical world. The threats disrupting our way of life today have gone faceless, borderless, and beyond the reach of any human operator. Aslan builds the autonomous agents that reach them, unmask them, and stop them, before a life can be lost, a household can be bankrupted, innocence can be stolen, or our can be nation subverted.
In 12 months we've gone from founding to live operational tasking and pilots across multiple national security and law enforcement partners. We've proven it works. This role makes it permanent.
We’ve raised $20M to date from Khosla Ventures, XYZ, BoxGroup, 2048, Liquid 2, Precursor Ventures, and others.
Role
Aslan builds autonomous agents that run for weeks at a time across many external systems, plan their own next steps, and work toward an objective. Scoring them the usual way doesn't work. Individual outputs look fine. Results come back late or not at all. What we need to know is whether the agent chose well at the moment it chose, and nothing off the shelf measures that.
This is our first quality role. You'll operate the platform daily, build the tooling that catches what you catch by hand, and own the decision to ship. We deploy on-premise on a fixed cadence and can't push fixes after delivery.
Responsibilities
Run the platform yourself every day and log what's wrong in enough detail that an engineer can reproduce it.
Sample agent plans, get them rated on whether the agent made the right call, and turn the ratings into a versioned dataset. Engineering uses it as a regression suite. The training side uses it as labels.
Diagnose whether a bad decision came from the model or from the harness handing it the wrong context, and route each to the right owner.
Write and run checks over each agent's full history, and over properties that only hold across the whole set of running agents.
Build a simulation harness so a week of agent operation runs overnight against a candidate build.
Define what counts as a regression between builds. Measure both whether agents last and whether they get anywhere. A build that only makes them cautious is a failed build.
Who You Are
You've owned evals for a shipped LLM or agent product, and you can say what your suite caught and what it missed
You've built a labeled dataset that got used for training, not just for measurement
You've picked a proxy metric because the real signal wasn't available, and you can defend the choice
You know when LLM-as-judge works and when it doesn't
You'd rather write the check than file the ticket
What You Bring
Strong Python
An evals framework you've run in production: Inspect, Promptfoo, Braintrust, Arize, LangSmith or equivalent
Familiarity with agent harness internals: memory, context assembly, tool selection
Preference data, process supervision, or reward modeling experience
Enough statistics to tell a regression from noise on small samples
Bonus Points
Evaluating agents that run for days or weeks, where per-episode metrics don't work
Systems you can't roll back or reset
Graph stores, provenance, lineage tracking
Eligible for a US security clearance
Posted by Aslan on their own careers page — you apply directly, no recruiter in between. View original / apply →