logo[мetahunt]
> DOU
middle

Data Engineer

SeeTree
format:Remotecompany:Product
pythonairflowbigquerygoogle clouddockerkubernetessqlci/cdgit
pytorchterraformnode.js
experience3+ years
> full description

Job Title: Software Engineer, Data Pipelines (Mid-Level)

Location: Remote, Ukraine

Team: Data & Platform Engineering — working side-by-side with our Data Science team

Overview

SeeTree is a leading company in the Ag-tech industry, providing per tree intelligence platform to growers to track their trees’ health and productivity.

Using rich sources of information such as drones, satellites, IOT sensors, weather information and more, we can scan and analyze hundreds of millions of trees and provide deep information on each and every one of them!

Our mission is to boost growers and industry ROI by digitally transforming agronomy, operation and decision making.

🎯 The Role

We are looking for a mid-level engineer to own the backend data pipelines that power our Data Science team. You will sit between Data Science and our GCP data platform: building the pipelines that produce the data our models learn from, and taking the models and prototypes the DS team builds and making them run reliably, on schedule, at continental scale.

This is a hands-on Python role. Concretely: satellite and drone imagery lands in cloud storage, and a chain of Airflow DAGs turns it into per-tree and per-field facts in BigQuery — ingestion, raster processing, alignment, detection, feature extraction, model inference, publication. You will build, extend and operate that machinery, and make it a platform the data scientists can move fast on instead of a black box they file tickets against.

🛠 What You’ll Do

  • Build and maintain batch data pipelines in Python, orchestrated with Apache Airflow on Cloud Composer 2 — DAGs, custom operators and sensors, retries, backfills, SLAs

  • Productionize data science work: take notebooks, prototypes and trained models from the DS team and turn them into containerized, scheduled, monitored jobs on GKE and Cloud Run

  • Design and evolve BigQuery datasets — feature tables, training sets, partitioning and clustering strategies, and cost / performance tuning of heavy analytical SQL

  • Own pipeline reliability: idempotent writes, checkpointing and resume for long historical backfills, data-quality checks, alerting, and failure modes that are loud rather than silent

  • Work with large geospatial datasets — Sentinel-2 imagery, drone orthomosaics, elevation models — computing per-field and per-tree statistics that feed downstream models

  • Build and maintain the Python services around the pipelines (Cloud Run, Flask / FastAPI) that expose results to our applications, dashboards and to the data scientists themselves

  • Contribute to architecture, specs and code reviews across a multi-repo, service-oriented GCP platform

⚙️ Our Stack

You do not need every line of this on day one — but this is what the work actually looks like:

  • Languages: Python (primary), SQL, some TypeScript / Node.js

  • Orchestration: Apache Airflow on Google Cloud Composer 2

  • Data: BigQuery, Google Cloud Storage, Firestore, Pub/Sub

  • Compute: Kubernetes (GKE) Jobs, Cloud Run, Cloud Functions, Docker

  • Python libraries: pandas, NumPy, GeoPandas, Shapely, rasterio, PyTorch (models we run)

  • Infra & CI: Terraform, Bitbucket Pipelines, Cloud Monitoring

  • AI tooling: Claude Code, Cursor, MCP servers, spec-driven agent workflows

✅ What We’re Looking For

  • 3+ years of professional experience writing production Python (data-intensive backends, ETL / ELT, or platform work)

  • Hands-on experience with a workflow orchestrator — Airflow strongly preferred; Dagster, Prefect or equivalent is fine if you are willing to learn Airflow properly

  • Strong SQL and real experience with a cloud data warehouse — BigQuery ideally, or Snowflake / Redshift / Databricks. You can look at a query and explain why it costs what it costs

  • Comfortable on a major cloud (GCP preferred; AWS / Azure transferable) and with containers — you can write a Dockerfile and debug a failing Kubernetes job

  • Experience collaborating with data scientists or ML engineers — you have taken someone else’s model or analysis and made it run in production

  • Solid engineering fundamentals: testing, code review, Git, CI / CD, and writing code the next person can read

  • Strong communication in English and the self-direction that remote work requires

🤖 AI-Savvy Must have — What We Actually Mean

This is not a buzzword on our job ad. We run an AI-assisted engineering practice: architectural context and specs live in a central hub repository, and coding agents do a substantial share of the implementation under engineer supervision. We are looking for someone who already works this way — or is visibly hungry to.

  • You use agentic coding tools daily (Claude Code, Cursor, Copilot or similar) and have opinions about where they help and where they don’t

  • You understand context and prompt engineering: writing the spec, the architectural README, the prompt that produces a correct pull request rather than plausible-looking code

  • You review AI-generated code critically — you know its failure modes and you do not merge what you cannot explain

  • You automate your own workflow — custom skills, MCP servers, scripts, CI hooks — instead of repeating manual work

  • Bonus: you have built something with LLM APIs, RAG or MCP in production, or integrated a model into a data pipeline

✨ Nice to Have

  • Geospatial experience: GeoPandas, rasterio, GDAL, PostGIS, Google Earth Engine, or work with Sentinel-2 / Landsat / drone imagery

  • ML engineering exposure: PyTorch, feature engineering at scale, model serving or inference orchestration

  • Infrastructure as code with Terraform, and general GCP infrastructure comfort

  • Node.js / TypeScript — parts of our platform are built on it

  • Background in Ag-tech or interest in data-intensive applications

🌱 Why This Role

  • Real scale and real consequences — hundreds of millions of trees, and growers making decisions on what our pipelines produce

  • Direct partnership with Data Science: short feedback loops, visible impact, no ticket-shuffling layer in between

  • A modern GCP platform and an engineering culture that funds and expects AI-assisted development, not one that quietly tolerates it

  • Remote-first, with a small senior team and genuine ownership of your domain

Відгукнутись на вакансію