Ofer Mendelevitch (Vectara; author of Hands-On RAG for Production and, with Jay Alammar, Hands-On Large Language Models) interviewed by Angelina on TwoSetAI — 70 minutes on production RAG and surviving the AI shift.

RAG isn’t dead — it just got a loop

  • RAG = retrieval + augmented generation, and retrieval stays essential even with agents; the “RAG is dead” claims every two months are mostly people marketing something new
  • Classic RAG is one-shot: query → top chunks → prompt → answer. Agents add a loop: the LLM plans, calls tools (often the RAG pipeline itself) repeatedly, and synthesizes
  • “Talk to my PDF” demos are not production: millions of documents in every format change the problem completely

The pipeline, from ingest to answer

  • Ingest: extract text → chunk → embed → vector store (store the text and page markers too, not just vectors — citations need to point at exact pages)
  • Hybrid search (semantic + BM25/TF-IDF) for what semantic search misses: numbers, product codes, exact strings — “if the source says 90% and your output says 85%, vector search won’t catch it”
  • Query side: top-5/10 results → optional reranking → prompt → grounded answer

Tables, images, and the red button problem

  • Tables are first-class citizens: chunked tables lose their column names (common in medical journals) — store whole tables, retrieve them whole
  • Images: store and return as images, don’t flatten to a description
  • Video: transcription alone loses meaning — “in order to avoid catastrophe, never press this button” means nothing without the visual. Today’s fix: VLM descriptions of short clips correlated with the transcript. Dedicated video embedding models exist but aren’t production-ready yet

When to bother with knowledge graphs

  • Multi-hop questions (“what else did the director of Inception direct?”) defeat semantic search
  • But graphs are expensive to build and maintain — worth it only for high-stakes use cases where a significant share of queries actually need the relationships

Eval: the hard part is the data, not the metric

  • Retrieval eval (did you fetch the right chunks?) needs query→gold-chunk datasets that are brutal to build — and documents keep changing
  • Generation eval compares against curated golden responses
  • Reference-free eval (Jimmy Lin’s Waterloo lab + Vectara): LLM-as-judge scorers like UMBRELLA (0–3 chunk relevance) validated against human correlation — no gold labels required

Build, buy, or rent

  • Build with LangChain/LlamaIndex only if it’s your business and you have the team; complexity compounds (multimodal, graphs, maintenance)
  • RAG/agent-as-a-service (e.g. Vectara) outsources the upkeep; vertical tools are fine when they cover your use case — but watch missing features and data-residency constraints

DevRel as a growth engine

  • Two jobs: teach developers how to use the product, and carry feedback back to the company
  • PLG over expensive sales teams: self-serve product, events, hackathons, real blog posts — “make it your own voice, don’t produce AI slop”
  • Measure directionally, don’t over-engineer attribution: five customers means it’s not working, ten thousand means it is, a thousand is unknowable — same problem founders face reading PMF

The skill that survives

  • Engineers and data scientists become directors, not actors: agents write the code; the remaining critical challenge is deciding what to build and steering where agents are weak (architecture, non-obvious trade-offs)
  • To the high schoolers who feared they made “the most incredibly stupidest mistake” by majoring in CS: graduate with the capability of today’s mid/senior engineer — use college to learn how to wield the AI tools
  • Fundamentals still matter; hiring will have to change — “write Fibonacci in five lines of Python is worthless” — expect AI-augmented interviews
  • His real worry is societal, not technical: how governments and finance distribute the gains

“We’re going to end up in engineering and data science being directors as opposed to actors. The coding agents will write the code.”