EngRadardirect-apply

Evals Lead

Aslan

Washington, DC Metro Area Full-time Posted 4d ago

About Aslan

Our national security apparatus was designed for an adversary that congregated and operate in the physical world. The threats disrupting our way of life today have gone faceless, borderless, and beyond the reach of any human operator. Aslan builds the autonomous agents that reach them, unmask them, and stop them, before a life can be lost, a household can be bankrupted, innocence can be stolen, or our can be nation subverted.

In 12 months we've gone from founding to live operational tasking and pilots across multiple national security and law enforcement partners. We've proven it works. This role makes it permanent.

We’ve raised $20M to date from Khosla Ventures, XYZ, BoxGroup, 2048, Liquid 2, Precursor Ventures, and others.

Role

Aslan builds autonomous agents that run for weeks at a time across many external systems, plan their own next steps, and work toward an objective. Scoring them the usual way doesn't work. Individual outputs look fine. Results come back late or not at all. What we need to know is whether the agent chose well at the moment it chose, and nothing off the shelf measures that.

This is our first quality role. You'll operate the platform daily, build the tooling that catches what you catch by hand, and own the decision to ship. We deploy on-premise on a fixed cadence and can't push fixes after delivery.

Responsibilities

  • Run the platform yourself every day and log what's wrong in enough detail that an engineer can reproduce it.

  • Sample agent plans, get them rated on whether the agent made the right call, and turn the ratings into a versioned dataset. Engineering uses it as a regression suite. The training side uses it as labels.

  • Diagnose whether a bad decision came from the model or from the harness handing it the wrong context, and route each to the right owner.

  • Write and run checks over each agent's full history, and over properties that only hold across the whole set of running agents.

  • Build a simulation harness so a week of agent operation runs overnight against a candidate build.

  • Define what counts as a regression between builds. Measure both whether agents last and whether they get anywhere. A build that only makes them cautious is a failed build.

Who You Are

  • You've owned evals for a shipped LLM or agent product, and you can say what your suite caught and what it missed

  • You've built a labeled dataset that got used for training, not just for measurement

  • You've picked a proxy metric because the real signal wasn't available, and you can defend the choice

  • You know when LLM-as-judge works and when it doesn't

  • You'd rather write the check than file the ticket

What You Bring

  • Strong Python

  • An evals framework you've run in production: Inspect, Promptfoo, Braintrust, Arize, LangSmith or equivalent

  • Familiarity with agent harness internals: memory, context assembly, tool selection

  • Preference data, process supervision, or reward modeling experience

  • Enough statistics to tell a regression from noise on small samples

Bonus Points

  • Evaluating agents that run for days or weeks, where per-episode metrics don't work

  • Systems you can't roll back or reset

  • Graph stores, provenance, lineage tracking

  • Eligible for a US security clearance

Posted by Aslan on their own careers page — you apply directly, no recruiter in between. View original / apply →

More at Aslan