This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Machine Learning Engineer - LLM & GenAI based in India.
Join an innovative engineering team building advanced AI solutions that transform how businesses interact with complex operational data. In this role, you will design and develop production-grade Large Language Model (LLM) applications that enable users to access insights through natural language conversations. You will work on cutting-edge Generative AI technologies, retrieval-augmented generation systems, and intelligent data workflows that create real-world impact. This global opportunity offers the chance to shape AI-powered products from architecture to deployment while collaborating with multidisciplinary teams. If you are passionate about machine learning, conversational AI, and building scalable intelligent systems, this role provides an exciting opportunity to contribute to the future of AI-driven analytics.
Accountabilities
- Design, develop, and maintain LLM-powered backend services using Python and FastAPI to support conversational AI experiences.
- Build and optimize retrieval-augmented generation (RAG) pipelines that connect structured and unstructured data sources with intelligent language models.
- Implement frameworks such as LangChain, LlamaIndex, Haystack, or similar technologies to manage context retrieval, query routing, summarization, and response generation.
- Develop prompt engineering strategies, structured output generation workflows, and intent classification pipelines to improve AI performance and reliability.
- Integrate conversational AI capabilities with existing product interfaces and backend services while ensuring secure and efficient data flows.
- Design, test, and validate AI solutions across diverse analytics use cases, including data summaries, diagnostics, predictive insights, and performance analysis.
- Evaluate and improve retrieval accuracy, response quality, latency, and hallucination rates through continuous testing and optimization.
- Implement caching strategies, schema-based memory solutions, and automated evaluation pipelines to enhance system efficiency.
- Maintain technical documentation covering system architecture, APIs, prompt strategies, deployment processes, and model lifecycle improvements.
Requirements
- 4–6 years of hands-on experience in Machine Learning, with proven experience building production-grade LLM or Generative AI applications.
- Strong proficiency in Python and experience developing backend services using FastAPI.
- Practical experience with LLM application frameworks such as LangChain, LlamaIndex, Haystack, or similar technologies.
- Experience designing retrieval pipelines and working with structured and unstructured data sources.