The Dev Log › AI & Machine Learning
Running Local LLMs With Ollama for Private Dev Workflows
By Jezer Niel Blanca, Full Stack Developer ·
·
6 min read
How I run open models locally with Ollama, call them from Laravel through an OpenAI-compatible client, generate private embeddings, and where local models fall short.
Cloud AI models are powerful, but they aren't always the right tool. Sometimes the data you're working with shouldn't leave your machine. Sometimes you want to prototype an AI feature without watching an API bill or a daily quota. And sometimes you just want to experiment offline. That's where running a local LLM comes in. In this post I'll show how I use Ollama to run open models locally, wire them into a Laravel app through the same code path I'd use for a hosted provider, and where local models fit, and don't fit, in a real development workflow.
Why Run a Model Locally?
Local models trade raw capability for control. The reasons I reach for them:
- Privacy. Prompts and data never leave your machine, which matters when you're working with client code, internal documents or anything sensitive.
- No per-request cost. Once a model is downloaded, experimenting is free apart from electricity and patience.
- No rate limits. Iterate on a prompt as many times as you like without hitting a quota.
- Offline work. Flights, patchy connections and locked-down networks stop being blockers.
- Predictable versions. A downloaded model doesn't change underneath you, which helps when you're comparing prompts.
The trade-off is honest: smaller local models are generally less capable than the largest hosted models, and speed depends heavily on your hardware. For many tasks such as summarising, classifying, drafting and extracting, a good local model is more than enough.
Getting Started With Ollama
Ollama packages open models with a simple CLI and a local HTTP server. After installing it from the official site, pulling and running a model takes two commands:
ollama pull llama3.2
ollama run llama3.2
The first downloads the model; the second opens an interactive chat in your terminal. A few other commands I use constantly:
ollama list
ollama ps
ollama rm llama3.2
list shows downloaded models, ps shows which ones are currently loaded in memory, and rm frees disk space when you're done with one.
Choosing a Model Size
Open models come in different parameter sizes and quantisation levels. As a rule of thumb, smaller models run faster and need less memory, while larger ones produce better answers but may be too slow or too big for a laptop. I start with a small general-purpose model, test it on my actual task, and only move up a size if the results aren't good enough. The Ollama model library lists the available sizes and tags for each model.
Calling Ollama From Laravel
Ollama runs a local server on port 11434. It exposes its own native API and also an OpenAI-compatible endpoint, which is the one I prefer. It means the same Laravel code can talk to a local model in development and a hosted provider in production, with only config changing.
First, keep the connection details in config/services.php:
'llm' => [
'base_url' => env('LLM_BASE_URL', 'http://localhost:11434/v1'),
'key' => env('LLM_API_KEY', 'ollama'),
'model' => env('LLM_MODEL', 'llama3.2'),
],
Ollama doesn't need a real API key, but the OpenAI-compatible format expects one, so any placeholder works. Then a small client class:
<?php
namespace App\Ai;
use Illuminate\Support\Facades\Http;
class LlmClient
{
/**
* @param array<int, array{role: string, content: string}> $messages
*/
public function chat(array $messages, float $temperature = 0.2): string
{
return Http::withToken(config('services.llm.key'))
->timeout(120)
->post(config('services.llm.base_url').'/chat/completions', [
'model' => config('services.llm.model'),
'temperature' => $temperature,
'messages' => $messages,
])
->throw()
->json('choices.0.message.content', '');
}
}
Notice the generous timeout. Local models on modest hardware can take a while, especially on the first request while the model loads into memory.
Switching Between Local and Hosted
Because everything comes from config, switching providers is just a matter of changing environment variables:
LLM_BASE_URL=http://localhost:11434/v1
LLM_API_KEY=ollama
LLM_MODEL=llama3.2
In production, point those at your hosted provider instead. Your application code doesn't change at all.
Build every AI feature against a provider-agnostic client from day one. It keeps you free to use a local model for private work and a hosted one where you need more power.
Local Embeddings for Private Search
Ollama can also generate embeddings, which power semantic search and retrieval. That makes it possible to build a private search over your own documents without sending them anywhere. Pull an embedding model and call the native embed endpoint:
ollama pull nomic-embed-text
use Illuminate\Support\Facades\Http;
$vectors = Http::timeout(60)
->post('http://localhost:11434/api/embed', [
'model' => 'nomic-embed-text',
'input' => ['How do refunds work?', 'Resetting a password'],
])
->throw()
->json('embeddings');
You get back one vector per input, ready to store and compare. One important rule: embeddings from different models aren't compatible, so always embed your documents and your queries with the same model.
Practical Development Workflows
Here's where local models have genuinely earned a place in how I work.
Prototyping AI Features
When I'm designing a new AI feature, the first dozen versions of the prompt are usually rough. Iterating against a local model costs nothing, so I can refine the prompt structure, the output format and the validation logic before spending anything on a hosted model. Once the shape is right, I test with the production model, because results can differ between models.
Working With Sensitive Data
For tasks like summarising internal notes, generating test fixtures from real-looking data, or drafting documentation from private code, a local model keeps everything on my machine. It's also useful for experimenting with anonymisation before any data goes to a third party.
Running Evals Cheaply
If you keep a set of test inputs for your prompts, running them against a local model gives a fast, free first pass. It won't replace testing against the production model, but it catches obvious regressions early.
Limitations to Keep in Mind
Local models aren't a free replacement for hosted ones, and it's worth being clear-eyed:
- Quality varies. Smaller models can struggle with complex reasoning, long documents and strict output formats. Validate outputs carefully.
- Hardware matters. Performance depends on your CPU, GPU and memory. What feels quick on one machine can be slow on another.
- Context windows are finite. Very long inputs may need chunking.
- Production is a different job. Serving a local model to real users means managing servers, scaling and uptime yourself.
- Check the licence. Open models come with their own licences, and some restrict commercial use. Read them before shipping anything.
Wrapping up
Ollama makes running open models locally simple: pull a model, point an OpenAI-compatible client at localhost, and keep the provider in config so you can switch freely. Use local models for private data, cheap prototyping and quick evals, and use hosted models where you need maximum capability. If you'd like to add private or AI-powered features to your product, I'd be glad to build them with you and my team.
Tags: AI, Ollama, Local LLM, Privacy, Laravel