> DOU
senior
BinariksData Engineer
format:Remotecompany:Outsource
pythonapache sparksqlkubernetesairflowprometheusgrafanaalertmanagerfivetran
pgvector
> full description
About
Binariks is looking for a highly motivated and skilled Senior Data Engineer.
About the Project: It is an Ingestion Platform.
The candidate will join a team building an ingestion platform for the legal / eDiscovery domain — a system that takes in large volumes of both unstructured documents and structured records from many different client systems and turns them into one clean, searchable model (normalizes it into a canonical data model for downstream search and analysis).
What We’re Looking For
- 7+ years of hands-on experience in data engineering
- Experience working with Python. Strong and production-grade, including testing, packaging, and typing in a shared team codebase
- Designing and building data pipelines that stay reliable as data volume and source count grow
- Hands-on production experience with horizontally scaling processing tools such as Apache Spark (PySpark)
- Data model design: schema design, normalization, and evolving a schema without breaking downstream consumers
- Strong SQL and relational database fundamentals
- Deploying and operating containerized applications on Kubernetes, including resource limits and configuration management
- Workflow orchestration with Airflow or another DAG-based scheduler, including retries, backfills, and dependencies
- Observability stack: Prometheus, Grafana, and Alertmanager, applied to pipeline health and data freshness
- Building and managing search indexes, including mapping design and reindexing
- Integrating with Fivetran or comparable managed connector tooling for third-party source ingestion
- Blob or data lake storage as an immutable raw data store, including partitioning and lifecycle rules
Will be a plus
- Vector indexes and embedding storage, particularly pgvector, for semantic search over documents
- eDiscovery domain experience, including how legal review workflows shape ingestion requirements
- Document processing and text extraction tooling at volume, such as Nuix or Aspose
- Relativity, including its data model, import and export formats, and API
- Chain-of-custody
Your Responsibilities
- Designing and building ingestion and ETL pipelines that stay reliable as data volumes and source counts grow
- Evolving the canonical data model: schema design, normalization, and mapping messy source systems into a shared internal model — including the fields and objects that don’t map cleanly
- Making the mapping auditable, so it’s always traceable how a source record became a canonical one
- Implementing incremental ingestion that absorbs upstream schema changes without full reprocessing
- Building and maintaining the search indexes the platform’s global search depends on — mapping design, reindexing strategy
- Turning a developer-operated pipeline into a self-service, scheduled and monitored system with proper alerting on pipeline health and data freshness
- Working with content-addressable storage and hash-based deduplication of ingested content
We provide the following for our employees
- 18 working days of paid vacation
- 10 working days of sick leave annually (5 days paid at 100% and 5 days at 75% rate of your average monthly salary)
- Flexible work schedule
- Additional days off for special occasions, national holidays off
- A competitive and rewarding salary based on performance appraisals/knowledge evaluation
- Possibility to share and gain knowledge on regular tech talks
- Friendly and professional team
- Innovative projects with advanced technologies
- Remote work
- Accounting service