ABOUT THE COMPANY
We are running a search for an AI-native data platform startup that helps technical distributors and manufacturers turn messy product data (spec sheets, supplier catalogs, PDFs, safety data sheets, ERP/PIM exports) into clean, market-ready product content.
The company is a small, founder-led team with deep prior experience building product-data and catalog platforms, has raised a $3M pre-seed round, and is seeing strong usage growth with enterprise customers.
THE OPPORTUNITY
This is one of the company's first engineering hires. The person in this role will own the data platform at the center of the product - the pipelines, warehouses, and sync systems that ingest, normalize, and enrich catalogs running into the millions of SKUs.
This isn't a toy version of the data problem: customers push through roughly a million products a year, each with up to a thousand attributes, and correctness and cost at that scale is the whole game.
This person will work directly with the CTO to move the company from "an LLM call for every problem" to a platform where LLMs write code that actually gets run - deterministic, tested, and measured.
WHAT THIS PERSON WILL DO
- Build and own the core data platform: ingestion, normalization, and enrichment pipelines across Postgres and BigQuery, with full lineage back to source
- Sync customer systems of record: mirror data from customer ERPs, PIMs, and e-commerce platforms into the warehouse and publish it back out, at a scale of millions of products and thousands of attributes
- Push the company past brute-force LLM calls: design and build the deterministic validation and transformation layer that lets LLMs write code the company runs, stores, sandboxes, and tests, rather than calling a model for every decision
- Design durable, long-running workflows: research agents that run for hours across hundreds of LLM calls and recover cleanly when something breaks
- Take the internal API customer-facing: a versioned, multi-tenant API with proper access control, testing, and docs (likely alongside a CLI)
- Treat AI as a measurement problem: build evals and golden datasets, trace agent calls, and prove that changes actually move the needle
- Own their work end-to-end: from idea to shipped feature to iteration, with direct exposure to customers and their real data problems
WHAT WE'RE LOOKING FOR
Must-Have
- Real, demonstrated data-engineering depth: has built and owned pipelines, warehouses, and ETL/ELT systems at meaningful scale, not just consumed one
- Strong backend fundamentals: can take a system from idea to deployed and stable, comfortable with distributed-systems primitives (queues, workers, idempotency, retries)
- Experience building with LLMs and genuine comfort with non-deterministic systems: knows how to measure outputs and catch regressions in a system that doesn't behave the same way twice
- A track record of shipping production code, not just architecting it
- Comfort with (or fast transferable skills into) TypeScript, Postgres, and BigQuery
Nice to Have
- Background building search, recommendation, ranking, or relevance systems
- Experience with durable-execution or workflow-engine frameworks (Temporal, DBOS, Inngest, Restate, Airflow)
- Big-data and warehouse tooling (Snowflake, dbt, Dataform, SQLMesh, large-scale ETL)
- A research background in ML, IR, or NLP paired with a real track record of shipping
- Early-stage startup experience
BENEFITS & PERKS
Health, dental, and vision coverage. Generous, flexible time off. Flexible working hours and location, fully remote within the US, with a light preference for Pacific/Mountain time zones.




