AYN

Senior Data & AI Engineer

CareDx, Inc. · Brisbane, CA

USD 145,000 to 187,000 a year

CareDx is a leading precision medicine diagnostics company advancing care in transplant, specialty oncology, and cell therapy. Through our innovative portfolio of molecular diagnostics, digital health solutions, AI-powered data and analytics, and patient support services, we partner with healthcare providers, patients, and biopharma organizations to help inform clinical decision-making and improve patient outcomes.

At CareDx, every employee has the opportunity to contribute to innovations that help transform patient care and improve lives.

The Senior Data and AI Engineer will serve as a senior hands-on technical owner for solutions that transform structured and unstructured data from real-world electronic medical records (EMRs) into reliable, standardized clinical data. The role will design and deliver data engineering and clinical natural language processing (NLP) pipelines to integrate, identify, and normalize diseases, diagnoses, laboratory tests and results, demographics, signs and symptoms, medications, procedures, and other information supporting clinical, scientific, and business operations.

This role combines data engineering, clinical NLP, medical terminology and ontology expertise, applied AI, OMOP data modeling, and production software engineering. The engineer will harmonize structured EMR data, interpret clinical narratives, and integrate these sources into a transplant-focused OMOP Common Data Model (CDM). Success requires sound technical judgment, strong clinical domain knowledge, close collaboration with clinicians, and ownership from data integration and validation through deployment and ongoing production support.

Responsibilities:

- Own the design, implementation, testing, deployment, and maintenance of pipelines that integrate structured EMR data and process unstructured clinical narratives.

- Develop entity and relationship extraction workflows for diagnoses, symptoms, laboratory results, demographics, medications, allergies, and procedures, linking findings to relevant values, units, doses, routes, frequencies, anatomical sites, and dates.

- Normalize structured source codes and extracted clinical concepts using ICD, SNOMED CT, RxNorm, LOINC, and other appropriate terminologies; manage local codes, synonyms, abbreviations, ambiguous mappings, terminology versions, and hospital-specific language.

- Implement clinical context interpretation, including negation, uncertainty, temporality, experiencer, and treatment status. Distinguish current conditions from historical, family, hypothetical, and ruled-out findings.

- Handle clinical note structure, copied-forward content, conflicting statements, and information spanning sentences or encounters. Reconcile extracted findings with structured EMR records using defined source and time precedence rules while retaining discrepancies, source evidence, and clinical timelines.

- Build and evaluate hybrid approaches using rules, dictionaries, statistical or deep learning models, and LLMs; apply structured output validation and terminology checks to prevent unsupported facts and invalid codes.

- Partner with clinicians to define clinical data schemas, source mappings, annotation guidelines, reference datasets, and acceptance criteria; resolve annotation disagreements and translate clinical requirements into engineering specifications.

- Measure entity, relationship, assertion, and normalization performance; conduct clinical error analysis and validate generalization across hospitals, specialties, and note types using patient-level data separation and checks for data leakage.

- Preserve source text, note and encounter identifiers, extraction provenance, model and terminology versions, and review status. Establish confidence thresholds and human review workflows for ambiguous or high-impact outputs.

- Build ETL workflows, APIs, and validation controls to harmonize structured EMR fields and NLP outputs, including patient and encounter linkage, units, timestamps, and duplicate records; monitor data quality, throughput, reliability, and model performance, and resolve production issues.

- Document architecture and operational procedures, lead code and design reviews, mentor engineers, and communicate technical tradeoffs to clinical, scientific, product, and engineering stakeholders.

- Follow security, privacy, and compliance requirements for sensitive healthcare data, including HIPAA-aligned practices, least-privilege access, appropriate logging, and secure data handling.

- Design and maintain mappings and transformation pipelines for a transplant-focused OMOP CDM, representing transplant procedures, donor and recipient information, immunosuppressive therapies, laboratory results, rejection episodes, and longitudinal outcomes using appropriate standard concepts and documented extensions where needed.

Qualifications

- Education: Master’s or Ph.D. degree in Computer Science, Computational Linguistics, Biomedical Informatics, Data Science, Engineering, or a related field, or equivalent practical experience.

- Experience: 7+ years of relevant experience in NLP, machine learning, or applied AI engineering, including 3+ years developing clinical NLP solutions using real-world EHR or hospital notes, with demonstrated delivery of production solutions.

- Advanced Python and SQL skills, with experience using pandas or comparable data-processing tools, relational data modeling, ETL orchestration, incremental processing, schema validation, and automated data-quality checks.

- Hands-on experience integrating structured EMR data through HL7 FHIR, HL7 v2, APIs, or database extracts; ability to align patient and encounter records, normalize units and timestamps, manage duplicates, and reconcile conflicting clinical sources.

- Hands-on experience with the OMOP Common Data Model and OHDSI standardized vocabularies, including source-to-standard concept mapping, vocabulary relationships, and ETL development to transform structured EMR data and clinical NLP outputs into appropriate OMOP tables.

- Experience validating OMOP data for conformance, completeness, and plausibility, with traceable source mappings and quality controls for longitudinal clinical records.

- Hands-on experience with clinical entity recognition and linking, relationship extraction, assertion detection, and temporal extraction, including clinical sentence and section detection, abbreviation disambiguation, negation, uncertainty, and experiencer identification.

- Experience implementing clinical NLP pipelines using medspaCy, scispaCy, cTAKES, or comparable frameworks, combining rules, dictionaries, and machine learning based on task requirements and measured performance.

- Practical expertise with ICD, SNOMED CT, RxNorm, LOINC, and UMLS, including terminology services, value sets, semantic types, ontology hierarchies, concept retrieval and disambiguation, local-code crosswalks, and terminology version management; experience with unit normalization using UCUM or equivalent methods.

- Experience applying and evaluating transformer models using PyTorch, TensorFlow, or equivalent frameworks and Hugging Face Transformers or comparable tooling; hands-on experience with LLM-based clinical extraction, schema-constrained outputs, and terminology-grounded validation.

- Strong understanding of medical terminology, common diseases, laboratory measurements, medications, and clinical workflows. Ability to apply clinical reasoning to interpret documentation, distinguish evidence from inference, select appropriate concept specificity, and recognize ambiguous findings requiring clinician review.

- Experience creating clinician-reviewed reference datasets and annotation guidelines; measuring precision, recall, and F1 for extraction, normalization, and assertion tasks; preventing patient-level data leakage; and evaluating performance across source systems and note types.

- Experience building production APIs or services, deploying contain

Apply on the employer’s site