Data Engineer
Hi there! AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
Why join us
If you’re looking for a place to grow, make an impact, and work with people who care, we’d love to meet you! :)
About the role
We are looking for a Middle Data Engineer to help modernize a 15-year-old data warehouse into a governed Databricks Lakehouse. You will build batch and streaming pipelines with PySpark and Delta Lake, following a medallion architecture across bronze, silver, and gold layers. This role also uses AI tools like Claude and GitHub Copilot to speed up development.
What you will do
● Design, build, and operate batch and streaming data pipelines on Databricks using PySpark, Delta Lake, and Databricks Workflows.
● Model and maintain a medallion (bronze/silver/gold) architecture serving analytics, reporting, and machine learning consumers.
● Migrate legacy ETL and data warehouse workloads onto the Lakehouse with validated data parity and minimal business disruption.
● Use Claude or Github Copilot as a development accelerator, generating code scaffolding, writing and reviewing tests, creating documentation and prototyping solutions.
● Write clean, well-tested Python and SQL; maintain high standards through code review and documentation.
● Optimize Spark jobs and Delta tables for performance and cost, including partitioning, clustering, caching, and cluster sizing.
● Implement data quality, lineage, and governance controls using Unity Catalog and automated validation checks.
● Debug, troubleshoot, and resolve pipeline failures, data defects, and production incidents.
● Collaborate with DevOps, platform, and analytics engineers on observability, security, and compliance best practices.
Must haves
● 3+ years of professional experience in data engineering, featuring direct expertise with Apache Spark and cloud-based data architectures.
● Strong hands-on experience building data pipelines with Databricks, Apache Spark (PySpark), and Delta Lake.
● Advanced SQL and Python, with strong data modeling skills across dimensional and Lakehouse patterns.
● Experience with streaming ingestion using Structured Streaming, Auto Loader, Kafka, or Event Hubs.
● Experience with workflow orchestration (Databricks Workflows, Airflow, or Azure Data Factory).
● Experience with legacy platform migrations, ETL modernization, or managing data hygiene when porting old systems.
● Strong problem-solving, collaboration, and communication skills.
● Familiarity with Unity Catalog, data governance, access control, and PII handling.
● Experience with dbt or an equivalent transformation framework.
● Familiarity with secure coding standards and industry security best practices.
● Experience delivering production data platforms at scale.
● Upper-intermediate English level.
Nice to haves
● Experience with Infrastructure as Code (IaC) using Terraform and CI/CD using Azure DevOps.
● Experience working with relational databases (specifically PostgreSQL) and data persistence concepts.
● Familiarity with logging and monitoring tools (e.g., Dynatrace, CloudWatch, Databricks system tables).
● Experience working in Agile or team-based development environments preferred.
Perks and benefits
● Professional growth: Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps
● Competitive compensation: We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities
● A selection of exciting projects: Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands
● Flextime: Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office — whatever makes you the happiest and most productive.
Meet Our Recruitment Process
Asynchronous stage — An automated, self-paced track that helps us move faster and give you quicker feedback:
● Short online form to confirm basic requirements
● 30–60 minute skills assessment
● 5-minute introduction video
Synchronous stage — Live interviews
● Technical interview with our engineering team (scheduled at your convenience)
● Final interview with your future teammates
If it’s a match—you’ll get an offer!