Compensation
$180,000-$220,000/yrDescription
About the Role
We are looking for a hands-on Data Engineer to help build and scale our modern data platform. In this role, you will work closely with Finance, Engineering, Product, and Analytics teams to develop reliable, high-performance data pipelines and systems.
You’ll contribute to both batch and real-time data processing using technologies like Databricks, AWS, kafka and several 3rd party data, while helping ensure data quality, accessibility, and usability across the organization. You’ll play a key role in enabling data activation, ensuring that high-quality data flows not only into the warehouse but also outward to business tools such as Salesforce etc. Additionally, you will help power next-generation AI-driven applications, including agent-based systems and AI driven tools using OSS tech, by building robust data foundations and pipelines. This is a great opportunity for someone who enjoys solving data challenges end-to-end from ingestion to insights.
Responsibilities
- Build and maintain scalable data pipelines (batch and streaming) using tools such as Databricks, Spark, Kafka, and AWS services
- Build and maintain pipelines from source systems (Salesforce, billing, product events, API logs) into clean analytics layers
- Design, develop, and optimize ETL/ELT workflows using DBT, PySpark, SQL, and tools like Fivetran
- Work closely with finance in developing Finance data solutions, Finance metrics and forecasting models
- Partner with Finance on revenue accounting, COGS, and margin reporting
- Partner closely with marketing and growth teams to enable data use cases such as segmentation, campaign targeting, and lifecycle analytics
- Develop and maintain reverse ETL pipelines to sync data from the warehouse to tools like Salesforce, HubSpot, Braze, and other downstream systems
- Create and manage curated datasets to support analytics, reporting, and go-to-market initiatives
- Build and maintain dashboards and reporting layers to support marketing and business performance tracking
- Support AI/ML and agent-based applications by preparing and serving high-quality datasets for MCP (Model Context Protocol) integrations and AI driven applications
- Monitor pipeline performance, troubleshoot issues, and ensure high data reliability and quality
- Implement data quality checks, validations, and alerting mechanisms across both ingestion and activation layers
- Collaborate with cross-functional teams to define data contracts and ensure consistency across systems
Qualifications
- 6+ years of experience in data engineering or a related field
- Strong hands-on experience with Databricks, AWS (S3, Glue, Athena, EMR, etc.), and Kafka
- Proficiency in Python (PySpark) and SQL for large-scale data processing
- Experience building and maintaining ETL/ELT pipelines (DBT/Airflow or similar experience preferred)
- Experience with data ingestion tools such as Fivetran (or similar)
- Familiarity with reverse ETL / data activation workflows and syncing data to tools like Salesforce, HubSpot, Braze
- Exposure to or experience with AI/ML data pipelines, including RAG architectures, vector databases, or embeddings workflows
- Familiarity with agent-based systems, MCP integrations, or LLM-powered applications is a strong plus
- Experience working with Finance and building finance specific metrics and pipelines is a strong plus
- Understanding of data modeling and working with large-scale datasets (batch and streaming)
- Experience creating dashboards and supporting reporting workflows (BI tools) for both internal and external audiences
- Strong problem-solving skills and ability to debug production data issues
- Strong communication skills and ability to work collaboratively across teams
Stack
- Posted
- Sep 18, 2026
- Last seen
- Sep 18, 2026
- First seen
- Sep 18, 2026

