Snr Applied AI Engineer
Ifs1 is the employer — EngRadar is a job radar, not a recruiter. We track this posting from their own careers page and send you straight there; we never handle applications or CVs.
About the AI Software Factory
The AI Software Factory is a new software-engineering function inside R&D. We build the systems that let AI agents do real engineering work on IFS software: design, implementation, testing, review.
We are after an order-of-magnitude improvement in delivery speed. That is a hypothesis we intend to prove or disprove in production, not a slogan. Faster only counts if the software still works, stays secure, and can be supported afterwards.
Our output is software. Orchestration, evaluation, codebase analysis, verification, and the interfaces that connect all of that to the way engineers actually work. Product teams are our users and our pilot partners. The team is new, so the first pilots and much of the technical architecture are still open questions. You would be joining to turn the proposition into something that runs, rather than to inherit a finished system.
One of our core deliverables is a reusable framework for parallel software engineering at IFS. It has to define how a feature gets broken into work several agents can do at once, how dependencies and shared state limit that, what context, tools and guardrails each agent receives, where an engineer reviews or decides, and how separately produced changes come back together as one releasable result. It must cover the full lifecycle: plan, design, build, test, review, document, operate. Parallel code generation on its own is not the goal.
Who this is for
You can hold agent orchestration, software architecture, evaluation, drift and trust in your head as one system. You look at a codebase and see dependencies, concurrency limits and evidence, where someone else sees a file tree. You like making messy systems measurable, and you are equally willing to build the machinery that does the measuring. Mostly, you want to define and prove a new engineering capability instead of inheriting a finished one.
Why this role?
Pointing one agent at one ticket is increasingly common. Building a repeatable framework where many agents plan, build, test and review parts of one outcome at the same time and still produce coherent software is not, and you would be helping define how it works.
You get a meaningful slice of that capability. Clear responsibility for working components and the evidence behind one pilot, with senior architecture support behind you, instead of a queue of disconnected tickets. What you build is production engineering infrastructure: evaluation, regression and codebase-analysis systems used to decide whether an agentic workflow is safe enough to expand. Those results inform whether a pilot proceeds, where human controls stay necessary, and which practices get adopted more widely across IFS.
Our interview process uses realistic work samples from the Factory's problem space, such as agent-generated changes and evaluation evidence, so both sides can look at actual work instead of talking around it.
The problem this role exists to solve
One agent on one ticket makes one engineer faster. The order-of-magnitude gain comes from many agents working at once, and that is where verification gets hard.
A codebase is a graph. The graph decides what can run in parallel, what has to be sequenced, and how far any change reaches. Model it badly and changes that each pass on their own will compose into a broken system.
Two questions follow:
Is this agent's output correct? Eval suites that mean something, deterministic checks, human validation where judgement cannot be avoided, and regression tests that keep known failures fixed when a model or prompt changes.
Do many agents' outputs compose? A codebase is a graph of dependencies. Running agents at the same time means understanding blast radius, spotting overlapping work, and verifying the merged result rather than approving a collection of individually green pull requests.
You own the senior technical judgement behind both answers, and the calls about where the Factory can safely turn up autonomy and concurrency.
What you'll do
You own the technical architecture for safe parallel agentic engineering: how codebase dependencies are represented, how work is decomposed and scheduled, how concurrent changes are isolated and put back together, and how the composed result is verified. Alongside that, you design how we evaluate any of it. Which claims can be checked deterministically, which need structured human validation, and what evidence a pilot has to produce before anyone should believe a delivery-speed or correctness number.
You draw the delegate / review / own line across the lifecycle. What agents run unattended, what an engineer reviews, what stays human. Then you move that line when the evidence says you can, and not before.
This is hands-on. You will write the orchestration, evaluation, regression, drift-detection and codebase-analysis components, wire them into real codebases and CI pipelines, and dig through failures in pilot workflows that behave like production. The measurement programme is yours too: versioned, reproducible scorecards, held-out test batteries, and enough statistical care (inter-rater agreement, paired significance testing) to separate a real improvement from noise when the decision turns on it.
The trade-offs land with you. Latency, cost and risk budgets for agent-driven changes. Where a human-in-the-loop gate is non-negotiable. What happens when a model upgrade, prompt drift or a tool’s changed behaviour quietly invalidates results you trusted last month, and how you keep provenance so you can tell.
Then there is the part that turns pilots into a capability. Working with product teams to test the approach against different codebases and delivery workflows, separating what generalises from what was a local accommodation, and feeding the rest back into how IFS engineers build.
You also raise the technical bar around you: guiding engineers working against your architecture, reviewing the hard judgement calls, and writing findings up clearly enough that engineering leaders can act on them. That is technical leadership through architecture, review and evidence. You will not have direct reports, and this is not a step towards having them.
In your first six months: stand up the evaluation and verification architecture for a live pilot, and use it to produce a defensible baseline of delivery speed and correctness. Reusable components, limitations you are honest about, and evidence showing where more autonomy or parallelism is safe and where it is not.
What you'll bring
You have built and operated agentic or tool-using AI systems in production, and owned the architecture through at least one serious redesign. A prototype or a single model call does not count.
Very strong software-engineering fundamentals, and the ability to work hands-on in codebases, languages and toolchains you have never seen before. We care about architecture and engineering judgement over stack match. You are an excellent coder, which is how you will know what excellent code from an agent looks like.
You have reasoned about concurrent or distributed work before: dependency graphs, DAG scheduling, isolation, fan-out/fan-in, partial failure, several workers touching shared state.
Evaluation instincts. You can tell a deterministic check from a subjective judgement, design evidence around the behaviour that matters, and spot a metric too weak to carry the claim being made on it.
You have run systems whose behaviour shifts underneath you, through model, prompt, dependency or data changes, and built the controls that catch silent regression.
Judgement under ambiguity. You make the consequential call, explain your reasoning, and change your answer when the evidence moves.
You write and speak clearly. You can turn a complex technical result into a decision without hiding the uncertainty or overselling what has actually been proven.
Show us you have genuinely built in this space. Production work, open source, research, a substantial independent project, technical writing: any of it counts. Depth beats collecting every item on the list.
Nice to have
None of this is required. All of it would help.
Statistical methods used in evaluation: inter-rater agreement, paired significance testing, stratified sampling, experimental design.
Deep experience with static analysis, call-graph extraction, monorepo build graphs, graph algorithms or other codebase-understanding systems.
A strong grasp of MCP, or another tool-calling or agent-orchestration protocol.
Work in regulated, security-sensitive, safety-critical or otherwise high-consequence engineering environments.
Experience building internal developer platforms or shared engineering systems used by several teams with different stacks and constraints.
If you are strong on production agentic systems, software architecture and rigorous evaluation, apply anyway when some of the adjacent areas are a stretch.
We embrace flexibility and hybrid work opportunities to support diverse needs and lifestyles, while also valuing inclusive workplace experiences. By fostering a sense of community, we drive innovation, strengthen connections, and nurture belonging. Our commitment ensures you can work in a way that suits you best, while also engaging with colleagues to share ideas and build meaningful relationships.
Posted by Ifs1 on their own careers page — you apply directly, no recruiter in between. View original / apply →