logo[мetahunt]
> DOU

AI Engineer

format:Remotecompany:Outsource
pythonfastapipostgresqlpytestlangchainlanggraphpgvectorqdrantlangfuserag
seowordpressjavascript
domainSaaS
locationLviv, Ukraine
> full description

Our client

WebPros is committed to empowering businesses worldwide through cutting-edge solutions in web hosting, billing automation, infrastructure, server management, and online marketing. Since our founding in 2017, we’ve rapidly grown into a global leader, expanding our robust portfolio to include industry-defining brands such as cPanel & WHM, Plesk, WHMCS, SolusVM, XOVI, SocialBee, Sitejet and Comet Backup.

Today, we power 85 million+ websites across 900,000+ servers worldwide, backed by a 650+ strong team of dedicated professionals spanning multiple continents.

For the role:

You would own the engine: the analysis algorithms, the retrieval, the evaluation harness, and the agents that turn a finding into an applied fix. Not prompt-tweaking on top of someone else’s model, but the logic that decides whether what we tell a customer is true.

That last part is really the job. Anyone can get a language model to produce a plausible recommendation. Proving it is correct is the hard part, because the ground truth shifts every time a model is updated, and it is the reason a customer would pay us rather than ask ChatGPT themselves. If you’ve ever been irritated by AI tools that sound confident and can’t show their work, this is the job where you get to fix that.


Key Responsibilities:

  • You own the analysis engine. The attributes and scoring rules that decide whether a model can understand a business are the core of the product, and keeping them at their sharpest is on you. When model behaviour moves, you’re the one who spots it, works out what it means for our rules, and either ships the change or makes the case for it
  • The evaluation framework. Ground-truth sets, metrics, and regression detection that survives the fact that the thing you’re measuring answers differently every time you run it. Most of the rest of this list is blocked without it
  • Retrieval. Getting text out of real websites, deciding how to chunk it, picking an embedding model, and keeping a path from every claim we make back to the page it came from.
  • The agentic architecture. One agent, several, or a plain tool-use pipeline. You’d work out which fits and say why in terms of cost and latency rather than preference. We haven’t settled this yet
  • Model selection and cost per run. Benchmarking frontier models against open-weight ones, then routing so we only pay for the expensive option where it changes the answer.
  • Guardrails and observability. Validating what goes in and what comes out, resisting prompt injection, catching hallucinations, and enough tracing to know where the money and the latency actually go
  • Turning research into production code. Reading what gets published on LLM and GEO behaviour, and making the case with evidence when it says we should change approach.

Your Qualifications

Must-haves

  • Production Python. It’s our stack and you’ll be in it daily, alongside FastAPI, PostgreSQL and pytest.
  • You’ve used the LLM SDKs directly rather than only through a wrapper, and you know the point where an orchestration framework like LangChain, LangGraph or Pydantic AI stops paying for itself
  • You’ve built retrieval that runs in production, including the messy parts: extraction, chunking, choosing an embedding model, recognising when retrieval isn’t the answer, and a vector store such as pgvector or Qdrant
  • You’ve evaluated an LLM system properly. A real dataset, metrics beyond exact match, and a judge you checked against human labels rather than trusted. Langfuse, Promptfoo and Ragas are the sort of tooling we mean You understand model behaviour well enough tobe sceptical of it. You can form a hypothesis about why retrieval or ranking behaved the way it did, test it, and say plainly what you found
  • Prompt, context and harness engineering. Writing the instruction is the easy part. What counts is what goes into the context window, what you leave out, and what you build around the model. DSPy is one example of the tooling here
  • You’ve designed agentic systems and can say why you picked that architecture over the ones you didn’t
  • You’ve kept one running in production. Tracing, token and cost attribution, latency, error rates, and using those numbers to decide what to fix first. Langfuse, LangSmith and Prometheus sit in this space
  • You know how these systems fail adversarially: prompt injection through content you crawled, output that reads well and is false, and what validation actually catches. Guardrails AI, LLM Guard and NeMo Guardrails are examples


    Nice-to-haves
  • Familiarity with SEO, generative engine optimisation (GEO), or search-visibility
    products
  • Classical NLP: classification, named entity recognition, sentiment analysis
  • WordPress, CMS integrations, schema.org, or llms.txt
  • Crawling and extraction at scale, including JavaScript rendering, rate limits and content
    quality
  • AI governance and risk management practice, such as the EU AI Act

Interview Process

• Screening call with Recruiter (soft skills interview) ~ 20 min
• Hiring Manager Interview (45 min)
• Technical & Culture Interview (60 min)
• Offer

_____

Q & A:

— Does the job come with a probation period, and if so, how long does it last?
Yes, there is a 3-month probation period.
— What is the expected work schedule?
Full-time, flexible. You can work remotely and also you can choose hybrid mode where you can combine working on-site (in Lviv office) and remotely.
— How many vacation and sick days are provided?
Annual paid vacation — 20 working days/ 7 unconfirmed sick days/days off a year.

Social package & benefits:

  • Full medical insurance
  • MacBook & accessories
  • English lessons
  • Accountant assistance
  • Minimal bureaucracy, synergy, and formalities, primarily focusing on effective communication
Відгукнутись на вакансію