ML / AI Engineer (agents, RAG, production)
Who we need
An ML / AI Engineer at JimmyNeuron. You’ll work on conversational agents, RAG, answer-quality evaluation, and embedding models into products and client systems (including on-premise and closed networks).
We need someone who can take AI to production: latency, cost, guardrails, hallucination monitoring, human escalation — not just a pretty Colab prototype.
Responsibilities
- Design and ship LLM agents, RAG, and document/dialog processing pipelines
- Connect models to products: APIs, queues, cache, limits, prompt versions
- Build evaluation: quality metrics, regressions, A/B hypotheses on answers
- Cut inference cost and raise reliability (retry, fallback, handoff)
- Work with product backends and integrations (CRM, messengers, telephony)
- Join client pilots: vision, local LLMs, knowledge bases — by project
Who we’re looking for
- Hands-on experience with LLMs, embeddings, vector DBs, orchestration (LangChain / custom pipelines — either is fine)
- Python and/or TypeScript for production code, not notebooks only
- Understanding of model limits and ability to design around them
- Experience shipping AI to production in at least one commercial product / engagement
- Plus: speech-to-text / call NLP, CV, on-prem GPU, daily Cursor use
What we offer
- Work at the product + custom intersection: ChatNeuron, CommBoost, and client pipelines
- Access to real data and metrics (under NDA), not toy datasets
- Comp and level — based on experience and case portfolio
- Format: Kazan or remote, full-time
How to apply
Telegram @aiuazis or info@jimmyneuron.ru
Subject: “Vacancy: ML/AI”
Attach: 1–3 cases (problem, stack, metrics, your role) and repo / article links if any.