ASAI job platform logo
  1. Home
  2. /Jobs
  3. /Machine Learning Engineer Jobs in Ahmedabad
  4. /Senior AI Evaluation & Reliability Engineer
AU
Aubergine

Senior AI Evaluation & Reliability Engineer

Ahmedabad · Senior

Older listing - lower visibility likelyVerified listingPosted 39d ago

Applicants who checked fit first are 3.1× more likely to hear back

Your match scoreCalculated · locked
86Overall
64Skills
97Experience

Your score for this role already exists

ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.

No credit card · 1 tap with Google

What we know about this role

Hiring pulse

MEDIUM

Aubergine is reviewing applications at a steady pace. Expect a standard response time as they evaluate the current pool.

Apply window

First 72 hours

Window passed - posted 39d ago

Early applicants get seen before the pile builds.

Not a repost

The first time we've seen this listing - it hasn't been closed and reopened.

Skills required

18 listed
Multi-Agent SystemsCI/CDVector DatabaseProject ProposalsClutchLarge Language Models (LLMs)Retrieval-Augmented Generation (RAG)GitHub+10 more
Multi-Agent SystemsCI/CDVector DatabaseProject ProposalsClutch

You almost certainly match several of these already. Unlock your skill map to see the matches, the gaps, and what to fix first.

Job description

Why Aubergine

Aubergine is a global transformation and innovation partner, shaping next-gen digital products through consulting-informed execution that integrates strategy, design, and development.

Since 2013, we’ve built 400+ B2B and B2C products worldwide, turning powerful ideas into impact-driven experiences. We are one of the top global B2B companies on Clutch, rated highest among more than 80,000 technology service providers.

With more than 150 digital thinkers, we are home to some of the brightest, most passionate people around the world who are committed to delivering excellence.

We’re not just another workplace. We’re officially Great Place To Work® certified, with an exceptional trust index rating, making Aubergine a community where you can thrive and grow.





Role Overview: Build AI Systems We Can Trust

We are looking for a Senior AI Evaluation & Reliability Engineer who is passionate about solving one of the most important challenges in AI:
How do we know an AI system is actually working, improving, and delivering business value?
You will design and build production-grade evaluation systems for LLMs, RAG applications, and multi-agent systems, while helping enterprise clients and our engineering teams adopt AI with confidence.
This role goes beyond building evals. You will be a trusted AI consultant to clients, a technical mentor to engineers, and a key contributor to our journey towards becoming an AI superagency.

What You Will Own

  1. Make AI Performance & ROI Measurable
    1. Build evaluation strategies that connect AI performance to business outcomes and ROI.
    2. Define quality benchmarks, SLOs, risk thresholds, and success criteria.
    3. Track hallucinations, reliability issues, and production risks.
    4. Build executive-friendly AI quality and ROI scorecards.
    5. Help clients determine where AI should be autonomous, supervised, or avoided.
    6. Don't just measure model accuracy. Measure business impact.
  2. Build Production-Grade Evaluation Pipelines
    1. Architect automated evaluation pipelines for LLMs, RAG, and agentic systems.
    2. Integrate evaluations into CI/CD using GitHub Actions, GitLab CI, or equivalent.
    3. Build regression suites for prompts, models, tools, and workflows.
    4. Measure faithfulness, context precision, answer relevance, semantic drift, and task completion.
    5. Establish statistically meaningful benchmarks and continuously monitor AI quality.
    6. Evaluation should become part of engineering, not a final QA step.
  3. Engineer LLM-as-a-Judge Systems
    1. Design reference-based and reference-free LLM evaluation frameworks.
    2. Create structured rubrics and scoring systems.
    3. Build calibration loops using human-labelled datasets.
    4. Identify and mitigate judge biases such as position, verbosity, and self-preference bias.
    5. Measure judge reliability and optimize evaluation quality, latency, and cost.
  4. Own Evaluation Economics
    1. LLM evaluations can become expensive quickly.
    2. You will:
    3. Design tiered evaluation strategies.
    4. Use deterministic and heuristic graders for simple checks.
    5. Reserve powerful LLM judges for complex evaluations.
    6. Optimise batching, concurrency, and parallel execution.
    7. Track evaluation cost and balance quality, latency, and inference spend.
  5. Benchmark RAG & Agentic Systems
    1. Build systematic benchmarks across:
    2. Retrieval quality, faithfulness, and answer relevance
    3. Embedding, chunking and re-ranking strategies
    4. Vector databases and retrieval architectures
    5. Agent task completion and tool-calling accuracy
    6. State retention, multi-step workflows and failure recovery
    7. Latency, reliability and cost
    8. Build synthetic datasets, golden test suites, and adversarial scenarios to continuously expand evaluation coverage.

Be a Trusted AI Consultant

  • You will work directly with North American enterprise clients as a technical AI advisor.
  • You will:
  • Lead AI architecture and evaluation discussions.
  • Translate complex AI metrics into clear business recommendations.
  • Define AI quality, reliability, and risk frameworks.
  • Present evaluation telemetry and ROI scorecards to technical and business stakeholders.
  • Challenge assumptions and recommend the right AI solution, even when that means saying "don't use AI here."
  • You should be equally comfortable discussing LLM evaluation with engineers and ROI with a CTO/COO/CEO/CFO.

Drive AI Consulting & Presales

Your consulting mindset will extend beyond delivery. You will actively participate in presales and new business opportunities, helping us win strategic AI engagements with enterprise clients.
You will:
  • Participate in discovery calls, solution workshops, and technical discussions with prospective clients.
  • Understand client challenges and identify where AI, agents, RAG, or automation can create meaningful business value.
  • Shape AI solution approaches, evaluation strategies, and technical proposals.
  • Contribute to statements of work, estimates, solution architectures, and project proposals.
  • Build compelling technical narratives and demonstrations that showcase our AI capabilities.
  • Present AI solutions and recommendations to technical and executive stakeholders.
  • Help identify opportunities to expand existing engagements through new AI capabilities.
  • Partner closely with Sales, Delivery, Product, and Engineering teams to convert opportunities into successful engagements.
  • Bring a consultative, outcome-focused approach rather than simply responding to client requirements.
You won't just help us deliver AI solutions. You'll help us win the right AI problems to solve.

Coach Engineers. Raise the Bar.

As our organisation evolves towards an AI superagency, you will help shape how we build AI.
You will:
  • Mentor engineers working on AI and LLM systems.
  • Establish AI engineering and evaluation best practices.
  • Conduct technical workshops and knowledge-sharing sessions.
  • Review AI architectures and evaluation strategies.
  • Help engineers move from prompt experimentation to disciplined AI engineering.
  • Build reusable frameworks, playbooks, and internal accelerators.
  • Stay ahead of emerging AI technologies and bring valuable ideas into the organisation.
Your impact shouldn't be limited to the systems you build. It should be reflected in the engineers you make better.

What We Are Looking For

AI Evaluation & Reliability

  • Proven experience building production-grade AI evaluation systems.
  • Strong practical experience with LLM-as-a-Judge.
  • Experience with calibration, human ground truth, and evaluation bias.
  • Strong understanding of RAG and agent evaluation.
  • Knowledge of Faithfulness, Context Precision, Answer Relevance, Task Completion, and semantic similarity.
  • Experience with DeepEval, Ragas, TruLens, LangSmith, DSPy, Promptfoo, or equivalent.

Engineering

  • Advanced Python and/or TypeScript.
  • Strong backend, API, and asynchronous systems experience.
  • Experience with CI/CD and automated testing.
  • Experience with vector databases such as Pinecone, Qdrant, Milvus, or Chroma.
  • Understanding of LLM infrastructure, observability, and inference economics.

Consulting & Leadership

  • Experience working directly with enterprise clients.
  • Excellent communication and presentation skills.
  • Strong consulting and problem-solving mindset.
  • Ability to translate technical complexity into business outcomes.
  • Strong mentoring and coaching abilities.
  • Curiosity, ownership, and enthusiasm for emerging AI technologies.

Why This Role Matters

The next generation of AI companies won't simply be the ones with access to the best models.
They will be the ones that can measure AI, trust AI, improve AI, and turn AI into business value.
If you want to build that future with us, we'd love to hear from you.

Free · no signup

Get tomorrow's jobs before you have to search

Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.

Join WhatsAppJoin Telegram

No spam. Just jobs and resources.

Why people use ASAI

Someone shared one job with you. ASAI keeps finding the rest.

  • Scored, not searched. Every role ranked against your actual profile.

  • Alerts as often as hourly. Reach new roles while the pile is still small.

  • Skill gaps, spelled out. See exactly which requirements you don't meet yet.

  • Verified jobs, only. Say no to ghost jobs. Your time deserves respect.

More Machine Learning Engineer roles in Ahmedabad

See all

Senior Engineer – Digital Solutions and AI Integration

Adani · Ahmedabad

Senior Python AI/ML Engineer

Inheritx · Ahmedabad

Sr Data and AI specialist

Adani · Ahmedabad

Lead - AI

Adani · Ahmedabad

Keep browsing

All Machine Learning Engineer jobs in Ahmedabad

Two ways in

Applicants who checked fit first are 3.1× more likely to hear back

Your match scoreCalculated · locked
86Overall
64Skills
97Experience

Your score for this role already exists

ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.

No credit card · 1 tap with Google

Free · no signup

Get tomorrow's jobs before you have to search

Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.

Join WhatsAppJoin Telegram

No spam. Just jobs and resources.

Why people use ASAI

Someone shared one job with you. ASAI keeps finding the rest.

  • Scored, not searched. Every role ranked against your actual profile.

  • Alerts as often as hourly. Reach new roles while the pile is still small.

  • Skill gaps, spelled out. See exactly which requirements you don't meet yet.

  • Verified jobs, only. Say no to ghost jobs. Your time deserves respect.

ASAI job platform logo

A job platform finally, balanced in your favour.

Jobs by CityJobs by CompanyGuidesHow We VerifyAboutPrivacy PolicyTerms of Service

Built in India 🇮🇳

© 2026 ASAI. All rights reserved.