About the role
Design and build production retrieval-augmented generation systems that answer questions from enterprise content accurately and securely.
What you’ll do
- Design RAG pipelines: ingestion, parsing, chunking, embeddings and hybrid search
- Implement permission-aware retrieval and metadata filtering
- Build evaluation sets and measure retrieval and answer quality
- Optimise latency and cost; implement caching and model routing
- Deploy and monitor RAG services in the cloud
- Mentor junior engineers
What you’ll bring
- 3–6 years of software engineering experience, including at least 1–2 years with LLM applications
- Strong Python and API development skills
- Hands-on experience with vector databases and embeddings
- Experience with at least one of Amazon Bedrock, Azure OpenAI or Vertex AI
- Understanding of evaluation methods for LLM and RAG systems
Nice to have
- Experience with Azure AI Search, OpenSearch or pgvector at scale
- LLMOps and tracing tools (Langfuse, LangSmith)
- Document AI (Textract, Document Intelligence)
Tech stack
What we offer
- A permanent, full-time role
- Enterprise projects for clients in the US and worldwide
- Hands-on work across AI, cloud and modern engineering
- A friendly, collaborative team
- Clear paths for professional growth
Work authorization
This role is based in Hyderabad, India. Applicants must be legally eligible to work in India.
How to apply
Email your résumé to career@inspiredinfotech.com with the job ID IIT-IN-201 in the subject line. Shortlisted candidates will be contacted by email.