> Djinni
senior
Backend Engineer
pythonllmlanggraphprompt engineeringawsaws bedrockcloudformationterraformany one ofcloudformationragai agentsmicroservicesci/cdapi
multi-agent systems
> full description
Description
We are looking for a hands-on Senior Python Engineer who can take an AI use case from technical design through implementation and integration into a production environment.
The ideal candidate combines strong Python/software engineering skills with practical experience building LLM and GenAI solutions, and is comfortable working in an enterprise environment with complex business and technical requirements.
Requirements
- 5+ years of software engineering experience (including backend/distributed systems), with 2+ years designing and running LLM/agentic systems in production.
- Expert Python engineering for long-running services: concurrency, resource management, performance profiling, API client resilience (retries, timeouts, backoff).
- Deep working knowledge of model capabilities/limits across tiers, model selection and pinning strategy, cross-region inference profiles, quota and throughput planning.
- Expert-level use of Strands Agents/LangGraph in production, including custom tool adapters, memory/session management, and framework-internals debugging.
- Proven ability to select and implement the right pattern per problem, decompose oversized tasks (map-reduce with checkpointing), and escalate autonomy only on eval evidence.
- Experience owning tool contracts as versioned integration surfaces: schema design, least-privilege scoping, read/write separation, contract testing with API-owning teams.
- Advanced context engineering: prompt layout for cache hit rate, stable-prefix design, per-step context allocation, measured (not guessed) representation choices.
- Experience designing versioned output contracts consumed by downstream systems, including evidence/citation fields and compatibility management.
- Experience designing evaluation methodology end-to-end: statistical treatment of nondeterminism, grader mix and judge calibration, eval-gated CI for prompt/model/config changes, growing eval sets from production failures.
- Experience designing defense-in-depth for production agents: prompt-injection resistance, tenant isolation, PII redaction handling, least-privilege tool access, degradation ladders and failure policies.
- Experience operating agents in production: end-to-end run reconstruction, per-tenant cost attribution, drift monitoring, incident response for AI-specific failures, capacity and quota management.
- Strong AWS experience including Amazon Bedrock (and ideally Bedrock AgentCore: Runtime, Gateway, Memory, Identity), infrastructure as code (CloudFormation/Terraform), and event-driven pipeline design (queues, DLQs, idempotency, checkpointing).
- Experience designing derived-data stores for agent pipelines (run state, checkpoints, memoized intermediate results) with clear source-of-truth boundaries.
- Ability to lead technical discussions, mentor engineers, and align integration contracts with application teams and architects.
Job responsibilities
- Own the design and delivery of production agent pipelines end-to-end, from trigger to published output, including decomposition for oversized inputs.
- Own tool contracts as versioned integration surfaces jointly with API-owning teams; drive contract-first development and contract testing.
- Design the release-bundle process (prompt + model + tools + config versioned and deployed as one unit, stamped on every output).
- Define the evaluation methodology and quality gates; calibrate judges against human experts; grow eval sets from production failures.
- Design the failure and degradation policy (retry, resume, degrade, honest absence), pathological-input handling, and prompt-injection defense-in-depth.
- Own production operations: dashboards, alerts, drift monitoring, incident response, quota/capacity planning, cost optimization.
- Lead code and design reviews; enforce testing standards across agent, tools, and infrastructure code.
- Design and operate shadow/champion-challenger rollout mechanisms; manage pinned-model upgrades with mandatory eval reruns.
- Partner with product and domain experts to define output contracts, quality rubrics, and feedback loops; challenge requirements that harm quality or cost.
- Mentor junior and middle engineers; develop the team’s agent-engineering playbook.