RAG & AI agents

Giving models fresh knowledge with retrieval, and the ability to act with tools.

โฑ 7 min read

An LLM's knowledge is frozen at training time, and it can't see your company's private documents. Two big ideas fix this: retrieval-augmented generation (look things up first) and agents (use tools in a loop).

your docs, chunked &embedded in advance ๐Ÿ“šโ“ Questionfrom the user๐Ÿ”ข Embedturn into a vector๐Ÿ—„๏ธ Searchvector database๐Ÿ’ฌ Answerwith citations๐Ÿค– LLMwrites the reply๐Ÿ“ Promptquestion + top chunkstop-k chunks
RAG: embed the question, find the most similar document chunks, paste them into the prompt, then generate a grounded answer.

๐Ÿ“š How RAG works

Index: split documents into chunks, embed each chunk, store the vectors in a vector database.

Retrieve: embed the user's question and fetch the top-k most similar chunks.

Generate: put those chunks in the prompt and ask the model to answer using them, ideally with citations.

๐Ÿ”ง Making RAG good

Chunk size matters: too big and you dilute relevance, too small and you lose context.

Hybrid search (keywords + embeddings) and a re-ranker often beat embeddings alone.

Evaluate retrieval separately from generation: if the right chunk never arrives, the model can't use it.

๐Ÿงฉ Quick quiz

What's the main benefit of RAG over fine-tuning for new facts?

๐Ÿค– Agents: models that act

An agent is an LLM in a loop: think โ†’ choose a tool โ†’ observe the result โ†’ think again, until the task is done.

Tools can be web search, a calculator, a code interpreter, an API, or a browser. The ReAct paper popularised this reason-and-act pattern.

๐Ÿ”Œ Tool calling & protocols

Modern APIs let you describe tools with a JSON schema; the model replies with a structured call, your code runs it, and you feed back the result.

Open standards like the Model Context Protocol (MCP) let any tool or data source plug into any compatible AI app.

๐Ÿงฉ Quick quiz

What is the core loop of an AI agent?

โœจ Before you drift off

  • RAG = retrieve relevant chunks, then generate grounded answers.
  • Chunking, hybrid search and re-ranking make or break RAG.
  • Agents = LLMs that reason and use tools in a loop.
  • MCP standardises how tools plug into AI apps.

๐Ÿ“š Go deeper (free & open)