logo[мetahunt]
> DOU
senior

Data Scientist

Xenoss
company:Outsource
machine learningnlpllmpythonfine-tuningllamahugging facestatisticsdata pipelines
reinforcement learningvllm
experience5+ years
domainAI
> full description

Who we are

Xenoss is an AI engineering and integration services company. We help medium to large enterprises run AI transformation end-to-end: situation analysis and goal framing, data discovery and preparation, pipeline building, model development, retraining pipeline design, deployment, and support.

We build a broad range of AI solutions: user behaviour prediction, content generation, NLP, audience segmentation, pathfinding, AI assistants, edge computer vision, fraud detection, and more.

We work with prominent companies such as Microsoft, Toshiba, AstraZeneca, Activision Blizzard, Verve Group, Voodoo Games, and Telefonica.

We are in the top 100 software companies on the Inc. 5000 list.

What is the project

We are looking for a Senior Data Scientist for a long-term In-Call Assistant initiative with a world-leading financial services company.

The project focuses on building a real-time conversational AI system that supports front-office employees during live customer conversations. The system identifies customer needs, objections, buying signals, and required process steps, and provides concise, context-aware recommendations.

You will primarily work on the signal detection and trigger layer: turning live conversation streams into structured signals, confidence scores, and routing decisions under strict latency requirements.

The broader solution combines low-latency signal detection, context preparation, specialist recommendation generation, RAG over approved product and policy knowledge, and compliance guardrails.

The data is large and messy. It includes over one million speech-to-text call transcripts with PII redaction, ASR errors, and no gold labels.

What you will do

Text data processing

  • Build scalable pipelines to clean and normalise noisy ASR call transcripts.
  • Handle PII-redacted text with typed placeholders. Detect and measure redaction errors.
  • Segment conversations into turns. Repair turns broken by interruptions and overlap.
  • Remove near-duplicate and low-quality samples (MinHash/LSH, embedding similarity, heuristic and model-based quality filters).
  • Mine and cluster objections and agent responses with embeddings and topic models.
  • Build train, validation, and test sets with time-based splits and no leakage.

Supervision and labeling

  • Build training targets from outcome signals: next customer reaction, call conversion, and agent performance.
  • Create labels with LLM-assisted annotation. Validate them with sales subject-matter experts.
  • Apply weak supervision and label-noise detection to weakly labeled data.
  • Build preference datasets: chosen/rejected pairs for DPO and good/bad pools for KTO.
  • Balance data across a taxonomy of around 50 objection types.

LLM fine-tuning and post-training

  • Run and improve supervised fine-tuning (SFT) with LoRA/QLoRA and full-parameter methods.
  • Run preference optimization: DPO, KTO, ORPO, SimPO, and iterative on-policy DPO.
  • Run reinforcement learning at turn level with GRPO (and PPO where useful).
  • Train reward models for response quality and conversion likelihood.
  • Design composite rewards (reward model + LLM judge + rule checks). Detect and limit reward hacking with KL control and reward ensembles.
  • Run best-of-N sampling and reranking as a strong baseline.
  • Train on a single node of 8× H100 80GB GPUs with DeepSpeed ZeRO-3 or FSDP.
  • Prototype on smaller models, then scale to the 100B+ MoE model.

Evaluation

  • Design offline evaluation for generated responses when no gold answer exists.
  • Build LLM-as-judge rubrics: relevance, empathy, clarity, factual accuracy, compliance, and how easy the response is to say aloud.
  • Use pairwise comparison with position swap. Calibrate judges against expert labels.
  • Measure agreement between judges and experts (Cohen’s kappa, Krippendorff’s alpha).
  • Estimate business impact from historical data with causal methods. Control for agent skill, channel, and customer segment.
  • Report results with bootstrap confidence intervals, across time windows and data slices.
  • Run ablations: input context, summary vs full transcript, customer profile on vs off.
  • Check factual accuracy and compliance. Flag invented fees, terms, or numbers.

Collaboration

  • Work closely with the AI Solution Architect, the other data scientist, and client ML teams.
  • Present experiments and results clearly to technical and business stakeholders.

You are expected to be deeply hands-on in data, training, and evaluation.

Scope of ownership

  • Own data quality for training and evaluation.
  • Own a measurable, stable evaluation framework that the client trusts.
  • Deliver improved model weights that beat the current SFT model and the best-of-N baseline.
  • Document every experiment: setup, data, metrics, and decisions.

What you should bring

Must have

  • 5+ years of hands-on experience in ML or NLP. At least 2 years with LLMs.
  • Strong Python. Clean, tested, reproducible code.
  • Hands-on fine-tuning of open-weight LLMs of 7B parameters or more (for example Llama, Qwen, Mistral, GPT-OSS).
  • Practical experience with SFT and at least one preference method (DPO, KTO, ORPO, or similar).
  • Multi-GPU training experience with DeepSpeed or FSDP.
  • Hands-on experience with Hugging Face Transformers, TRL, and PEFT.
  • Experience processing large, noisy text datasets (1M+ documents): cleaning, deduplication, filtering.
  • Experience building supervision from weak, noisy, or incomplete labels.
  • Experience evaluating generative models without gold labels (LLM-as-judge, human evaluation, pairwise comparison).
  • Strong statistics: confidence intervals, significance testing, bias and confounding.
  • Systematic error analysis and structured experiment design.
  • Clear communication with technical teams and domain experts.

Nice to have

  • RL for LLMs: GRPO, PPO, RLHF, or RLAIF. Reward modeling.
  • Fine-tuning Mixture-of-Experts models or models of 70B+ parameters.
  • Knowledge of chat templates and response formats (for example the GPT-OSS harmony format).
  • Experience with vLLM or SGLang for large-scale generation.
  • Causal inference or uplift modeling on observational data.
  • Conversational AI, contact-centre, or sales-call data.
  • Speech and ASR pipelines. Working with ASR errors.
  • PII redaction or privacy-preserving data work.
  • Financial services domain and compliance constraints.
  • Google Cloud (Vertex AI, BigQuery).
  • Publications, open-source contributions, or public fine-tuned models.

Operating model

  • Engagement: full-time, long-term B2B contract
  • Location: remote, EU-based
  • Time-zone overlap: at least 4 working hours with the New York team
  • Infrastructure: client environment only. No external training or data processing
  • Data residency: all work stays within the client perimeter
  • Delivery mode: offline model improvement first, then a controlled live pilot and production evolution
Відгукнутись на вакансію