Chicory logo
Chicory·

Senior Data Engineer - Chicory

About the Company

Chicory is the leading contextual advertising platform for CPG & grocery, transforming food content into dynamic commerce experiences that take consumers from inspiration to checkout. Our platform reaches 150 million grocery shoppers each month across a network of 7,000+ websites, food blogs, and apps, including The Kitchn, Food Network, MyFitnessPal, and powers 70+ leading retailers through direct integrations and retail media partnerships. By combining engaging media with contextually relevant, brand-safe content, we help CPG brands and retailers connect with consumers in an active mindset during high-intent moments—whether they’re discovering new recipes, planning meals, or building grocery lists. From curiosity to conversion, Chicory enables brands to show up in the moments that matter most, delivering full-funnel outcomes with ease.

Recognized as a 2025 Best Workplace by Inc. Magazine and as a 2023 Fastest-Growing Private Company in America by Inc. 5000, we believe that what sets Chicory apart is our diverse experiences and skillsets coupled with our shared values.

About The Role

Chicory is hiring a Senior Data Engineer to own and expand the data platform behind our advertising, campaign, product, and financial reporting. You’ll work with high-volume, real-time ad events, partner APIs and files, and internal business systems using GCP, BigQuery, Dataflow, Cloud Composer, Python, SQL, and Superset. The role covers the full data stack, from ingestion and orchestration through warehouse design, data marts, and BI, giving you the opportunity to see how data moves from our advertising platform into reports used across the company.

As Chicory’s Senior Data Engineer, you will have the autonomy to define the next version of our data lake and warehouse architecture and lead its implementation. You’ll remain hands-on, building pipelines and data models, establishing clear metric definitions, and improving how teams access and use data. You will work closely with backend and cloud engineers, as well as Campaign Management, Finance, Product, and leadership. Your work will directly support campaign performance, revenue reporting, publisher payments, and product decisions. This role offers broad technical ownership and the opportunity to establish the engineering practices that will support Chicory’s data systems and future data team.

Responsibilities

  • Data Mapping & Architecture: Map how ad requests, bids, impressions, clicks, campaign metadata, spend, revenue, publisher payments, partner files, and product events move from their source into BigQuery and Superset.
  • Warehouse Design: Choose and implement the new BigQuery structure, including raw, staging, core, and reporting datasets; table grain and keys; naming; partitioning and clustering; incremental loads; and ownership.
  • Pipeline Development: Build and maintain batch and streaming ETL and ELT pipelines with Python, SQL, Apache Beam, and Dataflow using data from Pub/Sub, Cloud Storage, APIs, databases, and files.
  • Orchestration: Build and maintain Cloud Composer and Apache Airflow DAGs with clear dependencies, useful retries and alerts, and tasks that can be rerun safely.
  • Transformation Workflow: Choose and implement a version-controlled SQL transformation workflow using Dataform, dbt, or an equivalent tool, then move important calculations out of scheduled queries and Superset dashboards.
  • Data Marts: Build the BigQuery tables and data marts used for campaign delivery, spend and revenue reconciliation, publisher payments, partner reporting, product analytics, and leadership reporting.
  • Data Quality & Reconciliation: Investigate mismatches between source systems, partner reports, BigQuery, and Superset; correct the logic, repair affected data, and document why the numbers changed.
  • Testing & Observability: Add tests and alerts for late or missing data, duplicate records, unexpected schema changes, broken joins, and totals that do not reconcile.
  • BI Support: Maintain the Superset datasets that power business reporting, review expensive or confusing queries, and make sure shared metrics come from the same maintained model.
  • Cost Optimization: Review BigQuery usage and cost, then improve the tables, queries, schedules, storage, or retention rules responsible for avoidable spend.
  • Governance & Security: Manage data access with GCP IAM, limit access to sensitive fields, and document retention or deletion rules where they apply.
  • Engineering Best Practices: Use Git, automated tests, code review, deployment pipelines, and infrastructure as code so changes can be reviewed, repeated, and rolled back.
  • Cross-Functional Collaboration: Work directly with Engineering, Campaign Management, Finance, Product, Sales, and leadership to define calculations, verify new tables, and decide when old reports can be retired.
  • Incident Response: When a pipeline fails or a report looks wrong, lead the investigation and recovery. Leave runbooks, diagrams, model descriptions, and metric definitions that another engineer can follow.

Qualifications

  • Demonstrated success designing, rearchitecting, and operating a production analytical data platform that includes a data lake or durable raw layer, a cloud data warehouse, governed data models, and business-facing data marts.
  • Deep hands-on knowledge of BigQuery, including schema and table design, partitioning and clustering, query plans and performance, workload patterns, permissions, retention, and cost management.
  • Strong Python and advanced SQL skills, including the ability to write maintainable, tested production code and review work across ingestion, transformation, orchestration, and data-serving layers.
  • Hands-on production experience building batch and streaming systems with Apache Beam and Dataflow and with Apache Airflow and Cloud Composer, including retries, idempotency, dependency management, duplicate and late-arriving events, schema evolution, replay, backfills, monitoring, and failure recovery.
  • Strong data-modeling judgment across dimensional, event, and domain models, with experience creating stable facts, dimensions, slowly changing dimensions, semantic definitions, and reusable analytical datasets.
  • Track record of establishing automated data quality, reconciliation, observability, lineage, alerting, service-level objectives, incident response, and operational runbooks for critical data products.
  • Experience enabling BI and self-service analytics through Superset or a comparable platform, including curated datasets, governed metrics, access patterns, dashboard performance, and stakeholder education.
  • Applied understanding of data governance, least-privilege access, privacy and sensitive-data handling, auditability, retention, and the controls required for financial or revenue-impacting data.
  • Experience with software-engineering practices for data systems, including Git, automated testing, CI/CD, code review, infrastructure as code such as Terraform, and clear technical documentation.
  • Demonstrated ability to operate as an autonomous senior technical owner: assess an unfamiliar environment, make architectural decisions, sequence migrations, resolve ambiguity, lead incidents, and communicate tradeoffs with both engineers and non-technical stakeholders.

Timezone overlap

UTC-8–-4

Open to

US · New York · United States

Sign in to track applications and earn points.

More roles at Chicory

Similar remote roles