AYN

Senior Site Reliability Engineer (SRE)

sparta-commodities · London · hybrid

We Are Sparta

Sparta is the AI leader for commodity trading, home to Leonidas AI, the first operating system for oil and commodity traders.

Backed by $42 million in Series B funding led by One Peak, with participation from returning investors FirstMark and Singular, the same investors behind AI leaders like Legora, Dataiku, and Synthesia. Most companies are building a chatbot on top of data. We've built an agent with an actual reasoning layer, grounded in years of proprietary market data and in-house trading expertise nobody else has.

We're scaling fast, and it's an exciting time to join. From independents to multinationals, the world's leading commodity teams use Sparta and Leonidas AI to move faster, trade smarter, and stay ahead of the market.

At Sparta, you'll be trusted to take ownership, backed by a team who wants you to succeed. AI is how we build here, not just what we sell. Good ideas move fast, and yours could go from staging to production in a couple of hours, whatever department you're in. You'll be challenged, supported, and given room to grow, building the AI that's changing how commodity trading gets done.

Your work has reach. From shaping Leonidas AI, to powering smarter decisions for traders around the world, to changing how commodity trading is done. And because we're growing fast, you'll accelerate your career with a level of ownership and impact you simply won't get at bigger, slower-moving companies.

At Sparta, we’re on a mission to build the next generation of commodity trading platforms - replacing the fragmented tools that traders typically rely on with a single, powerful dashboard. Our product is a data-driven platform that aggregates real-time feeds from across the commodities domain, transforming them into intuitive, actionable visualisations within one unified interface.

Traders act on what they see in Sparta. That makes reliability a product feature, not an afterthought - and it’s why this role exists.

This is a deliberate 50/50 split. Half your time is backend engineering: designing and building the distributed services and pipelines that move real-time market data through our platform. The other half is reliability engineering: making those systems observable, resilient, and cheap to operate. If you’ve ever shipped a service and then wished you owned how it ran in production, this is that job.

We’ve kept the split explicit rather than tidy. Our Platform Engineering team owns the shared platform, tooling and automation; you’ll be a close partner to them, and the reliability of your own services is yours.

We’re looking for people who thrive in an empowered environment - engineers who are comfortable being given problems to solve rather than solutions to implement. You should enjoy working at pace, value autonomy, and prefer to ask for forgiveness rather than permission.

Whereas Sparta is a remote-first company, for this role we’re looking for someone who values a hybrid working style, which in a typical week could involve spending a couple of days in the office - with flexibility built in.

WHAT YOU’LL BE DOING:

Backend engineering

- Design, build and maintain the backend services behind our real-time and analytical data processing.

- Optimise pipelines and services for low latency, high throughput and scale.

- Own features end to end, from shaping the approach to running them in production.

- Contribute to design reviews with a clear view on the trade-offs.

Reliability engineering

- Own the operational health of your services. Define what healthy means, then measure it.

- Improve observability, monitoring and alerting in Datadog, so problems surface before a trader notices.

- Take part in incident response, then close the loop on the root cause.

- Work with our runtime across Lambda, ECS and EKS, and help move more workloads onto Kubernetes.

- Help operate the data and streaming infrastructure you depend on: Kafka, Flink, Redis/Valkey, RDS, Redshift.

- Extend our infrastructure as code in AWS CDK, and our CI/CD pipelines.

- Reduce toil. Automate the manual, delete the unnecessary, make the next incident less likely.

ABOUT YOU:

- 4+ years as a software or reliability engineer, with production systems you’ve built and supported.

- High ownership tendencies.

- Genuine interest in both halves of this role. You want to write the service and own how it runs.

- Strong in at least one of Kotlin, Java, Python or TypeScript.

- A solid working understanding of AWS - compute, networking, storage, IAM - from running things in production.

- Hands-on with infrastructure as code: AWS CDK, Terraform, CloudFormation or similar.

- Practical experience running container workloads on Kubernetes: deploying, debugging, tuning.

- Experience with CI/CD tooling and a clear view of a good delivery lifecycle.

- A habit of instrumenting what you build - metrics, logging, tracing.

- Experience being on the hook for production, incident response included, and calm when things are on fire.

- Experience defining or working to service-level objectives, or a clear sense of how you’d start.

- A strong urge to own and improve things - to spot what isn’t working and fix it.

- Comfortable with agent-based development tools. We use Claude Code; any equivalent is fine.

- A clear communicator and a pragmatic problem-solver.

NICE TO HAVE EXPERIENCE:

- Building or operating a platform on Kubernetes, EKS especially.

- Datadog specifically.

- Data or streaming systems: Kafka, Flink, Redshift, clustered Postgres.

- Working to error budgets, and using them to make real prioritisation calls.

- Security and compliance best practice in cloud environments.

- Complex distributed environments - high throughput, low latency, large datasets.

- Any exposure to commodities, energy or financial markets. Useful, but we’ll teach you the domain.

Why Join Sparta?At Sparta, our culture is built on innovation, collaboration, and high performance. We thrive in a fast-paced environment where initiative is celebrated, challenges are embraced, and impact is rewarded.

• Equity in a high-growth, VC-backed company

• $42M in Series B funding — strong market traction and investment

• Early-stage growth — build your career as we scale

• Market-leading product — we are redefining commodity trading intelligence

• Top-tier tools and technology to support your success

What's the Culture Like at Sparta?At Sparta, our culture isn't just a set of values on a wall — it's how we work. We embrace a dynamic environment where initiative is celebrated, challenges are met head-on, and innovation thrives. You'll collaborate with people who are passionate about our mission and committed to building outstanding products and outcomes for our customers.

The Sparta Code

Understanding is non-negotiable

Understanding leads to believing. Believing leads to winning.

The 'why' is the difference between compliance and conviction. Before we execute, we ask: why does this matter, who does it serve, what changes if we win? People who understand the mission don't need to be managed.

Ask for forgiveness, not for permission

Most decisions are reversible. Most fears are not.

90% of decisions are reversible. Make them, ship, iterate. Perfection is the enemy of efficiency. Bias for action means moving before you have all the answers, but it isn't recklessness. The rule: tell the rest of us what you did, what changed, what you learned.

Fall in love with the problem

The mission wins. Your idea doesn't have to.

Solutions flatter our ego. They feel like progress. But solutions without problems are opinions, and opinions don't move customers, markets, or teammates. The goal is to find what's true, not to win the room. If your idea gets killed because a better one comes along, that's the system working.

It's your fault. Lucky you.

If it's on you, it's in your control.

True ownership sounds harsh until you feel how liberating it is. If every problem is yours, you stop

Apply on the employer’s site