Backend Engineer
Build the systems autonomous agents depend on: durable execution, safe recovery, isolated compute, and reliable data.
About Epiminds
We’re building Lucy, an AI system that investigates marketing performance and carries approved changes through to execution.
Behind that work is a backend coordinating model calls, external APIs, persistent state, and code running in sandboxes. We’re a small team in Stockholm building the infrastructure this product needs.
The Problem
Agent workloads are difficult to predict. One request can become several parallel investigations. A task can pause for approval and continue later. A worker can disappear after an external API accepts a change but before the result is saved.
Retrying may duplicate the action. Giving up may leave the work unfinished. Keeping everything ready reduces latency but consumes idle compute.
The system has to know what happened, who owns the next step, and what is safe to repeat—while keeping data isolated and costs under control.
The Role
You’ll own the execution infrastructure, data systems, and reliability behind Lucy.
You’ll work across application code, databases, and cloud infrastructure, taking changes from design through deployment and operation. Your decisions will shape how the product handles more work and recovers when things fail.
What You’ll Do
Make execution durable. Build scheduling, queues, leases, cancellation, and recovery for work that outlives an individual request.
Make retries safe. Preserve execution state and prevent duplicate events from repeating external actions.
Run agent compute efficiently. Improve sandbox startup, worker scaling, resource limits, and cleanup.
Build dependable data flows. Handle pagination, partial failures, changing provider contracts, and datasets that should be streamed rather than held in memory.
Make failures explainable. Connect traces, logs, and cost data so the team can find bottlenecks and resolve incidents.
What We’re Looking For
You’ve operated backend systems in production and owned their reliability. You understand database design, distributed state, queues, and the difference between a failed request and a failed operation.
You’re comfortable with containers and cloud infrastructure. You can explain what breaks under load, how recovery works, and when a simpler design is enough.
Our stack includes TypeScript, Node.js, Supabase/PostgreSQL, Redis, Temporal, Cloud Run, GKE, Modal, Docker, and Pulumi.
Why Join Now
Execution is central to the product. You’ll shape how agents run, how their work survives failures, and how infrastructure cost changes as usage grows.
What we offer
Competitive salary and meaningful equity.
Healthcare, retirement savings, and wellbeing support.
Office catering.
Hiring process
Recruiter screening.
Technical task.
Code interview.
Trial with us in the office.
We welcome applicants of all backgrounds.
- Department
- Engineering
- Locations
- Stockholm