The Dev Log › AI & Machine Learning
RAG for Web Developers: Let Your App Answer Questions About Your Own Data
By Jezer Niel Blanca, Full Stack Developer ·
·
6 min read
Retrieval-Augmented Generation explained in plain terms, with a practical Laravel implementation: chunking, embeddings, pgvector search and grounded answers.
Every business has knowledge trapped in places nobody reads: help center articles, PDFs, internal wikis, old support tickets. A language model can write beautifully about general topics, but it knows nothing about your refund policy or how your product's billing works. Retrieval-Augmented Generation, usually shortened to RAG, fixes that. Instead of hoping the model knows the answer, you find the relevant pieces of your own content and hand them to the model along with the question. In this post I'll explain RAG in plain terms and show how I build it in a Laravel app.
RAG in One Paragraph
When a user asks a question, you search your own content for the passages most related to it, paste those passages into the prompt, and ask the model to answer using only that context. The model does the language part: reading, summarizing, and phrasing the answer. Your app does the knowledge part: storing content and finding the right pieces. That split is why RAG works so well. You can update knowledge by editing a database row instead of retraining anything.
The pipeline has two halves:
- Ingestion: split documents into chunks, turn each chunk into an embedding, and store it.
- Querying: embed the user's question, find the nearest chunks, and send them to the model with the question.
Embeddings Without the Math Headache
An embedding is a list of numbers that represents the meaning of a piece of text. Texts with similar meanings end up with similar lists of numbers, even when they use different words. "How do I get my money back?" and "refund policy" will land close to each other, which is exactly what keyword search struggles with.
You don't need to understand the maths to use them. You send text to an embeddings endpoint and get a vector back. Here is a small service using Laravel's HTTP client against an OpenAI-compatible API:
<?php
namespace App\Services;
use Illuminate\Support\Facades\Http;
class EmbeddingService
{
/**
* @return array<int, float>
*/
public function embed(string $text): array
{
$response = Http::withToken(config('services.llm.key'))
->timeout(20)
->post(config('services.llm.url').'/embeddings', [
'model' => config('services.llm.embedding_model'),
'input' => $text,
])
->throw()
->json();
return $response['data'][0]['embedding'];
}
}
One important rule: use the same embedding model for ingestion and querying. Vectors from different models are not comparable.
Storing Vectors in Your Database
You don't need a specialized vector database to get started. If you're on PostgreSQL, the pgvector extension adds a vector column type and fast similarity search. A migration can enable it and add the column:
public function up(): void
{
DB::statement('CREATE EXTENSION IF NOT EXISTS vector');
Schema::create('document_chunks', function (Blueprint $table) {
$table->id();
$table->foreignId('document_id')->constrained()->cascadeOnDelete();
$table->foreignId('team_id')->constrained()->cascadeOnDelete();
$table->text('content');
$table->timestamps();
});
DB::statement('ALTER TABLE document_chunks ADD COLUMN embedding vector(1536)');
}
The number in vector(1536) must match the dimensions your embedding model returns, so check your provider's documentation. On MySQL or SQLite, for a small knowledge base you can store embeddings as JSON and compute similarity in PHP. It won't scale to millions of rows, but it's a perfectly good way to prototype.
Chunking: The Step That Makes or Breaks Quality
How you split documents matters more than most people expect. Chunks that are too big bury the relevant sentence in noise. Chunks that are too small lose the context that makes them meaningful. My starting rules:
- Split on natural boundaries first: headings, then paragraphs.
- Aim for chunks of a few hundred words, and let neighbouring chunks overlap slightly so ideas aren't cut in half.
- Keep metadata with each chunk: document title, URL, section heading, and who is allowed to see it.
- Re-embed a document whenever it changes. A queued job triggered from a model observer works nicely.
foreach ($chunker->split($document->body) as $content) {
$vector = $embeddings->embed($document->title."\n\n".$content);
$chunk = $document->chunks()->create([
'team_id' => $document->team_id,
'content' => $content,
]);
DB::update(
'UPDATE document_chunks SET embedding = ?::vector WHERE id = ?',
['['.implode(',', $vector).']', $chunk->id],
);
}
Prefixing the chunk with the document title is a small trick that helps a lot: a chunk that just says "You can cancel within 14 days" becomes much easier to match when it carries the title "Subscription Cancellation Policy".
Retrieval and Answering
At question time, embed the question and ask the database for the closest chunks. With pgvector, the <=> operator returns cosine distance, where smaller means more similar:
$vector = '['.implode(',', $embeddings->embed($question)).']';
$chunks = DocumentChunk::query()
->where('team_id', $user->team_id)
->orderByRaw('embedding <=> ?::vector', [$vector])
->limit(5)
->get(['id', 'document_id', 'content']);
Notice the where('team_id', ...) before the similarity search. Permissions must be applied at retrieval time. If a chunk is never retrieved, it can never leak into an answer.
Then build the prompt. The system message sets the rules and the retrieved chunks become the context:
$context = $chunks->map(fn ($chunk, $index) => "[{$index}] {$chunk->content}")->implode("\n\n");
$messages = [
['role' => 'system', 'content' => 'Answer using only the context below. '
.'If the answer is not in the context, say you don\'t know. '
.'Cite sources using their [number].'."\n\nContext:\n".$context],
['role' => 'user', 'content' => $question],
];
A RAG system that confidently answers from outside its context is worse than no RAG system. Make "I don't know" an acceptable answer.
Making It Good, Not Just Working
Getting a first answer is easy. Getting consistently good answers takes a bit more care:
- Show sources. Link each answer back to the documents it used. Users trust answers they can verify, and you can spot bad retrievals quickly.
- Build a small test set. Write twenty or thirty real questions with the answers you expect, and re-run them whenever you change chunking, models or prompts.
- Combine search types. Semantic search is great at meaning but can miss exact terms like product codes. Mixing in a keyword search for those cases often improves results.
- Log misses. Questions that end in "I don't know" are a free list of content your documentation is missing.
- Cache embeddings for repeated questions so you're not paying for the same work twice.
Wrapping up
RAG is the most practical way to make AI useful for your own business. Chunk your content thoughtfully, store embeddings next to your data, filter by permissions before you search, and tell the model to stay inside the context it was given. Everything else is iteration: better chunks, better tests, better prompts. If you'd like an assistant that genuinely understands your product's content, let's build it together with my team.
Tags: AI, RAG, Embeddings, Laravel, PostgreSQL