Founding Data & Machine Learning Lead
Keep Company · Remote (United States) · remote
The Role
This is Keep Company’s first dedicated data hire. It is a hands-on build role: initially weighted toward data engineering and analytics infrastructure, with meaningful ownership of model development, evaluation, and strategy as the foundation matures.
You will own how data moves through Keep Company: from source systems and product events to clean, trusted, member-centric datasets, reporting, and early predictive use cases.
You will partner closely with Product and Engineering to define the signals we should capture, build the infrastructure that makes those signals usable, and evaluate where models can improve matching, recommendations, and intervention decisions.
The first phase of the role will be primarily focused on the foundations: pipelines, data modeling, instrumentation, data quality, and reusable reporting. Over time, you will lead the strategy and development of models that help Keep Company identify risks, improve relationship outcomes, and make more useful recommendations.
What You'll Own
Data engineering and foundations
- Own Keep Company’s analytical data foundation, including the warehouse, transformations, pipelines, data quality, and documentation.
- Establish MotherDuck as a reliable analytical source of truth, including dbt models, source definitions, and durable datasets for internal and in-product use.
- Partner with Engineering to improve how product, program, survey, HRIS, import, and coach data is collected, structured, and maintained.
- Build and maintain pipelines across current inputs, including WorkOS/HRIS integrations, CSV imports, surveys, forms, and platform activity.
- Bring historical data into a clean, member-centric structure that connects attendance, registrations, survey responses, coaching inputs, relationship activity, and outcomes.
- Define and maintain a shared taxonomy for constructs, questions, scales, programs, cohorts, and other historical data that is currently inconsistent.
- Partner with Product and Engineering to define product instrumentation: what actions, milestones, decisions, and outcomes should be captured, where they should live, and how they can be used.
Reporting, measurement, and insight
- Build reusable measurement layers and reporting products rather than one-off analyses.
- Develop reporting that helps clients understand relationship participation, engagement, matching quality, and program outcomes.
- Establish an internal operating cadence for data: trusted dashboards and recurring analysis that leadership uses to make decisions.
- Partner with Product, Programs, Client Success, and leadership to define meaningful metrics and distinguish signals from noise.
- Build analysis that tests and improves Keep Company’s core assumptions, including how relationship design and matching affect engagement, retention, and development.
Model development and data strategy
- Own the strategy for moving from descriptive reporting toward predictive and recommendation capabilities.
- Document, evaluate, and improve the matching logic currently used in the product, including its assumptions, objective functions, scoring logic, failure modes, and quality tradeoffs.
- Build and test early models, heuristics, and prototypes that improve matching, identify risk, or support more useful recommendations.
- Define what additional data is required to support future prediction and recommendation use cases—and prioritize the collection and instrumentation needed to obtain it.
- Develop an experimentation approach for model-backed features: establish baselines, define success criteria, monitor quality, and improve systems over time.
- Identify where a heuristic is sufficient and where a more rigorous model is warranted.
- Build toward future use cases such as identifying low-quality matches before launch, detecting disengagement or at-risk programs early, and recommending the appropriate relationship type for a given employee lifecycle moment.
- Ensure predictive capabilities can be introduced thoughtfully, including privacy protections, tenant isolation, explainability, confidence thresholds, and paths for clients with AI restrictions.
Who You Are
Must Have
- 4-6 years of strong hands-on experience in data engineering, product analytics, decision science and/or data science, at a technology company with a track record of influencing product decisions.
- Experience with transformation tooling such as dbt, and with building reliable ingestion and transformation pipelines.
- Deep SQL fluency and experience designing and maintaining warehouse data models.
- Experience working with messy, incomplete, inconsistent, or historically unstandardized data.
- Strong background designing well-powered A/B tests, diagnosing bias and variance issues, and interpreting results under real-world conditions.
- Ability to translate business and product questions into durable datasets, metrics, reporting, and decision-making tools.
- Experience partnering with Product and Engineering to define instrumentation and turn product behavior into usable data.
- Working knowledge of statistics, experimentation, clustering, matching, and predictive modeling—enough to build or evaluate early models and set a sound modeling roadmap.
- Good judgment about data quality, privacy, sensitive attributes, tenant isolation, and consent-based product design.
- Strong communication skills and comfort explaining data, tradeoffs, uncertainty, and model behavior to non-technical stakeholders.
Preferred
- Experience with MotherDuck, DuckDB, Aurora/Postgres-derived data, S3, Parquet, WorkOS, Hex, or similar tools.
- Experience building recommendation, matching, prioritization, or prediction systems.
- Experience evaluating or improving a heuristic-based system before moving to more formal machine-learning approaches.
- Experience working alongside ML engineers on production models, including offline evaluation, online metrics, and guardrails.
- Experience building data products or analytics capabilities in a B2B SaaS environment.
- Experience with employee, engagement, professional-services, people, or relationship data.
- Experience using modern AI tools or open-weight models responsibly in product or data workflows.
- Experience helping build or lead a growing data, analytics, or BI function.
What Success Looks Like
In the first 6–12 months, you will have:
- Created a trusted, documented analytical foundation for Keep Company’s core data.
- Improved the reliability and usability of historical member, program, survey, and coaching data.
- Defined the most important product and outcome signals Keep Company needs to capture.
- Delivered reusable internal and client-facing reporting built on durable datasets.
- Evaluated the current matching approach and established a clear roadmap for model development.
- Tested at least one early predictive or recommendation-oriented use case with appropriate measurement and safeguards.
- Made data a consistent input into product, client, and leadership decisions.