The Dev Log › AI & Machine Learning
Prompt Engineering for Developers: Structured Outputs, Few-Shot and Evals
By Jezer Niel Blanca, Full Stack Developer ·
·
7 min read
Prompts inside your app run thousands of times against unseen input. Here is how I treat them like code: structured JSON, validation, few-shot examples and simple evals.
Most prompt engineering advice online is written for people chatting with a model in a browser. Developers have a different problem. When a prompt lives inside your application, it runs thousands of times against inputs you never saw, and its output has to be parsed by code that expects a specific shape. That changes everything. In this post I'll share how I treat prompts as real code: clear instructions, structured outputs, few-shot examples, validation, and evals that tell me whether a change made things better or worse.
Treat Prompts Like Code, Not Magic Words
The first mindset shift is that a prompt is a function specification. It has inputs, it should produce a predictable output, and it can have bugs. So I apply the same discipline I apply to any other code:
- Keep prompts in files or classes, not scattered as strings across controllers.
- Version them in Git so you can see exactly what changed and when.
- Parameterise them with clear placeholders instead of string concatenation everywhere.
- Test them against a fixed set of inputs before shipping a change.
In Laravel I usually give each AI task its own small class. It builds the messages, calls the model and returns a typed result. The rest of the app never sees raw prompt text.
Anatomy of a Good System Prompt
A reliable system prompt answers four questions for the model:
- Who are you in this task? A narrow role, like "You classify customer support emails."
- What exactly should you do? Concrete steps, not vague goals.
- What must you never do? Boundaries, such as "never invent order numbers."
- What should the output look like? A precise format.
Vague instructions like "be helpful and accurate" add nothing. Specific instructions like "if the email mentions a refund, set intent to refund" add a lot.
Structured Outputs: Stop Parsing Prose
The single biggest improvement you can make is to stop asking for free text when your code needs data. Many providers that support the OpenAI-compatible chat format let you request JSON, and some let you supply a JSON schema the output must follow. Where that is available, use it. Where it isn't, ask for JSON explicitly and validate it anyway.
<?php
namespace App\Ai;
use Illuminate\Support\Facades\Http;
class SupportEmailClassifier
{
/**
* @return array{intent: string, urgency: string, summary: string}
*/
public function classify(string $email): array
{
$response = Http::withToken(config('services.llm.key'))
->timeout(20)
->post(config('services.llm.base_url').'/chat/completions', [
'model' => config('services.llm.model'),
'temperature' => 0,
'response_format' => ['type' => 'json_object'],
'messages' => [
['role' => 'system', 'content' => $this->systemPrompt()],
['role' => 'user', 'content' => $email],
],
])
->throw();
$data = json_decode($response->json('choices.0.message.content'), true) ?? [];
return $this->validated($data);
}
private function systemPrompt(): string
{
return <<<'PROMPT'
You classify customer support emails for a web application.
Respond with a JSON object containing exactly these keys:
"intent": one of "billing", "bug", "feature_request", "account", "other"
"urgency": one of "low", "normal", "high"
"summary": one sentence, maximum 25 words, in plain English
Only use "high" urgency if the customer cannot use the product at all.
PROMPT;
}
}
Setting temperature to 0 makes the output more consistent for classification-style tasks. It isn't a guarantee of identical results, but it removes a lot of unnecessary variation.
Always Validate the Response
Even with JSON mode, a model can return a value outside your allowed list or leave out a key. I run every response through Laravel's validator, exactly like user input:
use Illuminate\Support\Facades\Validator;
/**
* @param array<string, mixed> $data
* @return array{intent: string, urgency: string, summary: string}
*/
private function validated(array $data): array
{
return Validator::make($data, [
'intent' => ['required', 'in:billing,bug,feature_request,account,other'],
'urgency' => ['required', 'in:low,normal,high'],
'summary' => ['required', 'string', 'max:300'],
])->validate();
}
If validation fails, I either retry once or fall back to a safe default like other with normal urgency. The key point is that bad model output never flows silently into the database.
Model output is untrusted input. Validate it with the same rules you would apply to a form submitted by a stranger.
Few-Shot Examples Beat Long Explanations
When instructions alone aren't producing consistent results, show the model what you want. Few-shot prompting means including a handful of example inputs and ideal outputs before the real input. Models are remarkably good at copying a pattern they can see.
$messages = [
['role' => 'system', 'content' => $this->systemPrompt()],
['role' => 'user', 'content' => 'I was charged twice this month, please fix it.'],
['role' => 'assistant', 'content' => '{"intent":"billing","urgency":"normal","summary":"Customer reports a duplicate charge this month."}'],
['role' => 'user', 'content' => 'The dashboard is blank and I cannot log in at all.'],
['role' => 'assistant', 'content' => '{"intent":"bug","urgency":"high","summary":"Customer cannot log in and sees a blank dashboard."}'],
['role' => 'user', 'content' => $email],
];
Some guidelines for examples:
- Cover the tricky cases, not just the obvious ones. Include the boundary between two categories.
- Keep them short. Every example costs tokens on every single request.
- Make them realistic. Examples that look nothing like real input teach the wrong pattern.
Keep Untrusted Text in Its Place
When your prompt includes text written by users, that text can contain instructions of its own, such as "ignore the rules above." This is prompt injection, and you can't fully prevent it with wording alone. What helps:
- Put user content in the
user message, never mixed into the system prompt.
- Wrap it in clear delimiters and tell the model it is data to analyse, not instructions to follow.
- Restrict what the output can do. If the result is only ever one of five categories, an injected instruction has very little room to cause harm.
Limit What the Model Controls
The safest design is one where even a fully manipulated response can't do damage. A classifier that returns an enum is low risk. A prompt whose output is executed as a database query is high risk. Design for the first kind.
Evals: How You Know a Prompt Got Better
Changing a prompt without measuring is guessing. An eval is simply a fixed list of inputs with expected outputs, run against your prompt so you can compare versions. It doesn't need a fancy platform to be useful.
<?php
namespace App\Console\Commands;
use App\Ai\SupportEmailClassifier;
use Illuminate\Console\Command;
class EvalSupportClassifier extends Command
{
protected $signature = 'ai:eval-support';
protected $description = 'Run the support classifier against labelled examples';
public function handle(SupportEmailClassifier $classifier): int
{
$cases = json_decode(file_get_contents(base_path('evals/support.json')), true);
$passed = 0;
foreach ($cases as $case) {
$result = $classifier->classify($case['email']);
if ($result['intent'] === $case['expected_intent']) {
$passed++;
} else {
$this->warn("Expected {$case['expected_intent']}, got {$result['intent']}: {$case['email']}");
}
}
$this->info("{$passed} of ".count($cases).' cases passed.');
return self::SUCCESS;
}
}
I build the example file from real, anonymised inputs over time. Every time the classifier gets something wrong in production, that input goes into the eval set. The next prompt change has to handle it.
What to Measure
For structured tasks, exact-match accuracy on key fields is usually enough. For free-text outputs like summaries, I check simpler things automatically, such as length limits and required keywords, and review a sample by hand.
Wrapping up
Prompt engineering for developers is mostly software engineering: keep prompts in versioned classes, ask for structured JSON, validate every response, show the model examples of tricky cases, keep user text clearly separated, and measure changes with a small eval set. Do that and AI features stop feeling unpredictable. If you want to add reliable AI features to your product, I'd be happy to build them with you and my team.
Tags: AI, Prompt Engineering, LLM, Laravel