AYN

Data Engineer

civicmarketplace · London, UK · remote

Why this role existsEvery city, county, and school district buys things. Roads, software, cleaning services, IT infrastructure. The total is somewhere north of two trillion dollars a year. And almost all of it moves through procurement processes designed for a different era: slow; paper-heavy; opaque and exhausting for everyone involved.

Civic Marketplace was built to fix that. We're a modern, data-driven platform where government agencies discover, evaluate, and engage suppliers. Where businesses, especially smaller and growing ones, can actually find and win public sector work without needing a dedicated contracts team to navigate the maze.

We're past the point of proving this works. Agencies are live on the platform, real money moves through it, and we're now combining that marketplace infrastructure with AI in ways that could genuinely transform how procurement works. Not just incrementally but structurally.

None of it works without data that can be trusted. Procurement is a data problem before it is a software problem: who the suppliers actually are, what they can genuinely deliver, which contracts an agency is already entitled to buy from, what this thing cost the county next door. That information exists, scattered across thousands of agency portals, state registries and PDF attachments, in no agreed format, and nobody has assembled it properly. Whoever does gets to define how public money is spent for the next decade. That's this role.

The problem you'd ownThe hard part isn't moving data from one place to another. Any engineer here will tell you the pipelines are the easy half.

The hard part is that public procurement has no shared vocabulary. The same supplier turns up as four different legal entities across three registries, and two of the spellings are wrong. Commodity codes are applied inconsistently, or not at all. A cooperative contract one agency can buy from today is invisible to the agency next door because it was published as a PDF on a portal with no API. Every source you touch is incomplete, inconsistent, or both, and most of it is public record, so you can't quietly correct it. You have to model the mess honestly.

And the bar just moved. We recently launched an MCP integration that lets agencies request quotes through Claude, GPT and Copilot. When an agent answers a procurement question, a wrong answer doesn't read like a bug, it reads like advice, and in public spending, bad advice ends up in a council meeting. That puts weight on freshness, lineage and provenance that most product data layers never carry. As we build out our agentic procurement capabilities, the reliability of that data, and how we prepare it for retrieval, becomes our most critical engineering challenge.

So this is an entity resolution and data trust problem dressed as a pipeline problem, and it's yours to solve. You'd work out what is actually wrong with the data, which is rarely what everyone assumes, then build the fix so it holds for every source we add next. Not a script per source held together by whoever wrote it.

About the roleYou would be our first dedicated data engineering hire, reporting to Mikey Mo, our Head of Engineering. Data work today sits with the product engineering team and gets done alongside shipping features. It works, but it belongs to whoever last touched it, and that isn't a foundation for what comes next.

To give you a sense of the gap: awarded quotes that slip through the platform get manually caught by our commercial team and hand-entered into HubSpot instead. Supplier onboarding data follows the same pattern. So as and when the platform and HubSpot become out of sync, it makes it difficult to say with confidence which one is right.

This is a build-and-do role. You'd collaborate closely with our Head of Engineering to shape the architecture, and you'd also write the pipelines, carry the pager for them, and go and read the raw source when a number looks wrong. If you're looking for a role where you set direction in isolation and hand implementation to other people, this isn't it. If you want substantial responsibility and scope as a function lead alongside our Head of Engineering while still enjoying the craft yourself, it very much is.

You would not be starting from nothing. There's a live platform with real transactions running through it, a product engineering team who know where the bodies are buried, and access most people in this field would have to scrape for: Civic Marketplace is a member of the NIGP Business Council, with real partnerships across councils of governments. Most people doing this job spend their first year getting the data access. You'd start with the door already open.

What you'd ownFive things, and the freedom to decide how.

- Trust in the data. The north star. A golden, trusted dataset that powers Civic Marketplace's analytics, so when someone asks how many quotes were awarded last quarter, or how much GMV we've captured, there's one number, and it's right. You'd own the diagnosis and the fix.

- Ingestion. Pipelines that pull supplier, agency, solicitation and contract data out of sources never designed to be read by anything but a person. Reliable, observable, and cheap enough to add the next source without a meeting about whether it's worth it.

- The canonical model. One supplier, one record, across every spelling, trading name, subsidiary and registry identifier. One shared definition of a contract vehicle, a commodity, an agency, that product, sales and customer success all use and none of them argue with. This is the unglamorous part that makes everything downstream possible.

- The system underneath. Testing, lineage and freshness, so when the platform tells an agency something we know where it came from and when. And self-serve data for product, customer success and analytics, so nobody queues behind an engineer for a number. Everything you build should work for the next source without you in the loop. This is where AI earns its place, turning what is currently hand-checked into something repeatable.

- Agentic data infrastructure. As we scale our use of AI-native procurement, you will own the data layer that powers our agentic workflows. This goes beyond standard retrieval. You will design ingestion and storage strategies, including optimizing vector pipelines like Pinecone for RAG, to ensure our agents have the high-fidelity, context-aware data necessary to perform reliably in production.

You'd work closely with product engineering, customer success and our Growth Lead. Application feature work stays with the product engineering team, so you're not competing for that ground. You own the layer they build on top of.

How we use AIWe use Claude daily, across content pipelines, community response and internal knowledge, and increasingly inside the product itself. That's real, not a line in a job ad.

For this role it cuts two ways. There's how you work: source profiling, schema mapping, test generation, the documentation that otherwise never gets written. Most of that is now automatable, and we expect you to automate it. And there's what you build: the data layer that decides whether agentic procurement is trustworthy or embarrassing. The second is the harder and more interesting problem.

What we care about is not whether you can use it. Everyone says they can. It's whether you reach for it to make something repeatable. The difference between someone who writes pipelines faster and someone who builds the tooling that means the next hundred sources don't each need hand-holding is the difference we're hiring for.

If you've built data quality tooling, evaluation harnesses, retrieval pipelines or automated documentation from scratch, tell us about them. If AI has changed how you think about the work and not just how fast you produce it, we especially want to hear that.

Our stack: Snowflake, Postgres, dbt, Fivetran, HubSpot, PostHog, Metabase.

What you bring

You'll need:- Data engin

Apply on the employer’s site