Machine Learning Engineer
Who we are
Xenoss is an AI engineering and integration services company, helping medium to large enterprises run AI transformation end-to-end, from situation analysis and goals framing to data discovery and preparation, pipeline building, model development, retraining pipeline design, solution deployment, and support.
We build a broad spectrum of AI solutions such as user behaviour prediction, content generation, NLP, audience segmentation, pathfinding solutions, AI assistants, edge computer vision, fraud detection, and others.
We work with prominent companies such as Microsoft, Toshiba, AstraZeneca, Activision Blizzard, Verve Group, Voodoo Games, and Telefonica, among others.
We’re included in the top 100 software companies on the Inc. 5000 list.
What is the project
We’re hiring a Senior AI/ML Engineer to join a long-term In-Call Assistant initiative for a world-leading financial services company.
The project focuses on building a real-time conversational AI system that supports front-office employees during live customer conversations. The system identifies customer needs, objections, buying signals, and required process steps, and provides concise, context-aware recommendations.
You will primarily work on the signal detection and trigger layer: turning live conversation streams into structured signals, confidence scores, and routing decisions under strict latency requirements.
The broader solution combines low-latency signal detection, context preparation, specialist recommendation generation, RAG over approved product and policy knowledge, and compliance guardrails.
What will you do
You’ll own implementation and continuous improvement of the signal detection and trigger model, working closely with the AI Solution Architect and the rest of the AI team.
Core work includes:
- Turning conversation transcripts into structured signals, intents, and trigger events
- Building and training multi-label classification models for live conversation analysis
- Preparing training datasets using LLM-assisted annotation and SME-validated labels
- Developing embeddings, classifiers, and alternative modeling approaches for signal detection
- Designing and tuning confidence thresholds, routing logic, and abstention behavior
- Optimizing inference for low latency and high throughput
- Building evaluation datasets, metrics, and error-analysis workflows
- Running experiments and model comparisons across accuracy, latency, and data requirements
- Investigating false positives, false negatives, and model failure patterns
- Working with AI, data, MLOps, and client teams to move models from experimentation to production
You’re expected to be deeply hands-on in model development, training, evaluation, and optimization.
Technology landscape
You’ll operate across the modern applied AI and ML ecosystem, including:
- Python
- PyTorch and Hugging Face
- NLP and conversational AI
- Embeddings and sentence encoders
- Multi-label classification
- Lightweight neural classifiers
- Classical ML and gradient boosting where appropriate
- LLM-assisted data annotation
- Confidence calibration and threshold optimization
- Model evaluation and error analysis
- Low-latency inference and serving
- MLOps, monitoring, and model lifecycle tooling
We optimize for measurable model quality, low latency, production viability, and enterprise constraints.
Scope of ownership and delivery context
Core ownership
- Implement and improve the Signal Extraction & Trigger model
- Build training and evaluation pipelines for signal detection
- Work with the architect and client SMEs on signal taxonomy and labeling
- Establish measurable model quality across precision, recall, false positives, and false negatives
- Tune confidence thresholds and trigger behavior
- Optimize models for real-time inference constraints
- Contribute to production monitoring and continuous model improvement
The proposal specifically describes Model 1 as an always-on component processing the rolling transcript, producing multi-label confidence scores and routing signals only when calibrated thresholds are met.
Team and delivery context
- Work closely with the AI Solution Architect and the engineer responsible for recommendation generation
- Partner with client SMEs on taxonomy, labeling, and validation
- Collaborate with data engineering and MLOps on training and inference pipelines
- Contribute to technical decisions through experiments and measurable results
- Work with incomplete or weakly labeled enterprise datasets and help establish reliable supervision
What should you bring
Must have
- Strong hands-on experience building and training production-oriented ML systems
- Strong Python skills
- Practical experience with PyTorch and/or Hugging Face
- Experience with NLP, text classification, conversational AI, or similar unstructured-text problems
- Experience training and evaluating classification models
- Strong understanding of precision, recall, F1, false-positive / false-negative trade-offs, and confidence thresholds
- Experience building datasets and supervision from noisy or incomplete labels
- Experience with embeddings and representation learning
- Understanding of model inference performance, latency, and throughput
- Ability to run structured experiments and perform systematic error analysis
- Comfort working with messy enterprise data
- Ability to communicate clearly with technical teams and domain experts
Nice to have
- Experience with real-time conversational AI or call-center systems
- Experience with multi-label classification
- Experience with sentence-transformer or similar embedding models
- Experience with model calibration
- Experience with XGBoost, LightGBM, or similar classical ML approaches
- LLM-assisted labeling or synthetic supervision experience
- Speech / ASR pipeline familiarity
- Financial services domain exposure
- MLOps and production model monitoring experience
- Experience working with large-scale, latency-sensitive inference systems
Operating model
- Engagement structure: Full-time, long-term B2B contract
- Work location: Remote, EU-based
- Time-zone overlap: At least 4 working hours overlapping with the New York team
- Infrastructure: Client environment only, no external training or data processing environments
- Data residency: All work executed within the client perimeter
- Delivery mode: Initial offline prototype followed by controlled live pilot and production evolution