logo[мetahunt]
> DOU
senior

AI Engineer

Xenoss
company:Outsource
pythonpytorchhugging facellmfine-tuningqloraragprompt engineering
vector databases
domainFintech
> full description

Who we are

Xenoss is an AI engineering and integration services company, helping medium to large enterprises run AI transformation end-to-end, from situation analysis and goals framing to data discovery and preparation, pipeline building, model development, retraining pipeline design, solution deployment, and support.

We build a broad spectrum of AI solutions such as user behaviour prediction, content generation, NLP, audience segmentation, pathfinding solutions, AI assistants, edge computer vision, fraud detection, and others.

We work with prominent companies such as Microsoft, Toshiba, AstraZeneca, Activision Blizzard, Verve Group, Voodoo Games, and Telefonica, among others.

We’re included in the top 100 software companies on the Inc. 5000 list.

What is the project

We’re hiring a Senior LLM Engineer to join a long-term In-Call Assistant initiative for a world-leading financial services company.

The project focuses on building a real-time conversational AI system that supports front-office employees during live customer conversations. The system identifies customer needs, objections, buying signals, and required process steps, and provides concise, context-aware recommendations.

You will primarily work on the recommendation generation layer: training and specializing instruction-tuned language models to produce grounded, policy-aligned next-best-action recommendations based on the live conversation and prepared customer context.

The broader solution combines low-latency signal detection, context preparation, specialist recommendation generation, RAG over approved product and policy knowledge, and compliance guardrails.

What will you do

You’ll own implementation and continuous improvement of the recommendation generation models, working closely with the AI Solution Architect and the rest of the AI team.

Core work includes:

  • Building and fine-tuning specialist recommendation models for objection handling, discovery, and product guidance
  • Designing training datasets from historical conversations, outcomes, SME input, and generated supervision
  • Applying SFT, preference optimization, and parameter-efficient fine-tuning approaches such as LoRA / QLoRA
  • Evaluating DPO, KTO, and other post-training methods where appropriate
  • Building RAG capabilities over approved product, policy, and knowledge sources
  • Designing grounding, abstention, and fallback behavior
  • Developing evaluation frameworks for recommendation quality, factual accuracy, relevance, and policy alignment
  • Running systematic error analysis and model improvement cycles
  • Optimizing model serving for latency, throughput, and infrastructure constraints
  • Working with AI, data, MLOps, and client teams to move models from experimentation to production

You’re expected to be deeply hands-on in LLM training, fine-tuning, evaluation, and optimization.

Technology landscape

You’ll operate across the modern LLM and applied AI ecosystem, including:

  • Python
  • PyTorch and Hugging Face
  • Instruction-tuned language models
  • SFT and preference optimization
  • LoRA / QLoRA and PEFT
  • DPO, KTO, and related post-training approaches
  • RAG and knowledge-grounded generation
  • Embeddings and retrieval
  • Prompting and context construction
  • Generative model evaluation
  • Guardrails, grounding, and hallucination control
  • Low-latency LLM inference and serving
  • MLOps, monitoring, and feedback loops

We optimize for measurable recommendation quality, factual grounding, low latency, and enterprise constraints.

Scope of ownership and delivery context

Core ownership

  • Implement and improve the Recommendation Generation model
  • Train and maintain specialist adapters for defined conversation scenarios
  • Build training and evaluation pipelines for generative models
  • Define and test fine-tuning and post-training strategies
  • Develop RAG and grounding mechanisms for approved knowledge sources
  • Establish measurable recommendation quality and factual accuracy
  • Improve abstention, fallback, and policy-compliance behavior
  • Optimize generative models for real-time inference constraints
  • Contribute to production monitoring and continuous model improvement

The proposal describes Model 2 as a shared instruction-tuned base model with router-selected LoRA / QLoRA specialist adapters and a separate RAG path for factual grounding.

Team and delivery context

  • Work closely with the AI Solution Architect and the engineer responsible for signal detection
  • Partner with client SMEs on training data, recommendation quality, and policy validation
  • Collaborate with data engineering and MLOps on training, retrieval, and inference pipelines
  • Contribute to technical decisions through experiments and measurable results
  • Work with incomplete or weakly labeled enterprise datasets and help establish reliable supervision

What should you bring

Must have

  • Strong hands-on experience building and fine-tuning LLM-based systems
  • Strong Python skills
  • Practical experience with PyTorch and Hugging Face
  • Experience with instruction-tuned language models
  • Experience with SFT and parameter-efficient fine-tuning approaches such as LoRA / QLoRA
  • Strong understanding of LLM evaluation and generative model failure modes
  • Experience preparing training datasets for LLM fine-tuning
  • Experience with RAG and knowledge-grounded generation
  • Understanding of grounding, hallucination control, and abstention strategies
  • Experience running systematic experiments and error analysis
  • Understanding of model inference performance, latency, and throughput
  • Comfort working with noisy, incomplete, or weakly labeled enterprise data
  • Ability to communicate clearly with technical teams and domain experts

Nice to have

  • Experience with DPO, KTO, RLHF, or other preference optimization techniques
  • Experience with conversational AI or real-time assistant systems
  • Experience training multiple specialist adapters on a shared base model
  • Experience with synthetic data or LLM-assisted supervision
  • Experience with guardrails and policy-constrained generation
  • Experience with retrieval systems and vector databases
  • Experience with low-latency LLM serving
  • Financial services domain exposure
  • Speech / ASR pipeline familiarity
  • MLOps and production LLM monitoring experience

Operating model

  • Engagement structure: Full-time, long-term B2B contract
  • Work location: Remote, EU-based
  • Time-zone overlap: At least 4 working hours overlapping with the New York team
  • Infrastructure: Client environment only, no external training or data processing environments
  • Data residency: All work executed within the client perimeter
  • Delivery mode: Initial offline prototype followed by controlled live pilot and production evolution
Відгукнутись на вакансію