Applied AI Engineer
Build the harness behind Lucy: agent loops, tools, context, code execution, and evaluations that make autonomous work reliable.
About Epiminds
We’re building Lucy, an AI system that does marketing work: investigating performance, finding opportunities, and carrying approved changes through to execution.
We’re a small team in Stockholm building the environment that lets agents do this work reliably across tools, data, and ongoing tasks.
The Problem
A capable model still needs the right environment to do useful work. It needs access to tools, relevant context, a place to execute code, feedback on what happened, and clear boundaries for when to continue or stop.
Those choices compound. A tool response can overwhelm the context. A handoff can lose a constraint. A failed command can send an agent into a loop. An earlier assumption can become the basis for an unsupported action.
Our agents already delegate to specialists, run code in sandboxes, and retrieve client knowledge. The challenge is engineering the harness around them so they can complete longer, harder tasks with less intervention.
The Role
You’ll own the agent harness: the execution loop, tool interfaces, context assembly, memory, delegation, and feedback that turn model capability into dependable behavior.
You’ll work directly with the founders and engineers, moving between runtime code, traces, and evaluations. The backend engineer owns the underlying compute and durable execution infrastructure; you own how agents use that environment to get the job done.
What You’ll Do
Build the agent loop. Shape how agents plan, call tools, inspect results, recover from errors, and decide when they’re finished.
Engineer tool interfaces. Design schemas, discovery, outputs, and error feedback that help agents choose and use tools correctly.
Manage context across long tasks. Decide what enters the prompt, what stays in files, what gets retrieved, and what survives a handoff.
Make code execution productive. Give agents useful ways to run Python and shell commands, inspect artifacts, and correct failed attempts.
Build the improvement cycle. Turn failures from real runs into reproducible evaluations. Measure task completion, correctness, latency, and cost as you change the harness.
Preserve control. Make approval boundaries, stopping conditions, and action outcomes clear to the agent.
What We’re Looking For
You’re a strong software engineer who has built and shipped systems using language models. You understand how tool design, context, and feedback affect agent behavior, and you can investigate failures across a full execution trace.
You’re comfortable with TypeScript and Python, designing experiments, and maintaining the runtime code behind them. Experience building agent frameworks, coding agents, execution environments, or evaluation systems is especially relevant.
Our stack includes TypeScript, Node.js, the Vercel AI SDK, multiple model providers, LangSmith, PostgreSQL, and sandboxed Python execution.
Why Join Now
The harness determines how much of a model’s capability reaches the user. You’ll own that layer at Epiminds, with room to rethink how agents work and direct responsibility for proving the result.
What we offer
Competitive salary and meaningful equity.
Healthcare, retirement savings, and wellbeing support.
Office catering.
Hiring process
Recruiter screening.
Technical task.
Code interview.
Trial with us in the office.
We welcome applicants of all backgrounds.
- Department
- Engineering
- Locations
- Stockholm