Skip to content

Building an on-site Claude agent in Laravel without an SDK

How I built the assistant on mutexai.net with Laravel's HTTP client: a bounded tool-use loop, a daily spend cap, a local fallback that always answers, and a human-reviewed learning queue.

My studio site, mutexai.net, has an assistant in the corner. It answers questions about what the studio does, sends visitors to the right page, and knows when to say "I don't know, let me flag that". It is built in plain Laravel with no AI SDK, and it runs inside a budget I set.

This post walks through how it works and why I made the choices I did. The same design now runs on this site.

Why no SDK#

The Messages API is a single HTTP endpoint that takes JSON and returns JSON. Laravel already has a good HTTP client with timeouts, retries and fakes for tests. Adding an SDK would have meant another dependency to approve and track, for what is essentially one POST request.

So the client is one small class:

public function createMessage(array $system, array $messages, array $tools, bool $allowTools = true): array
{
    $settings = config('assistant.anthropic');

    $payload = [
        'model' => $settings['model'],
        'max_tokens' => $settings['max_tokens'],
        'system' => $system,
        'messages' => $messages,
        'tools' => $tools,
    ];

    if (! $allowTools) {
        $payload['tool_choice'] = ['type' => 'none'];
    }

    $response = Http::baseUrl($settings['base_url'])
        ->withHeaders([
            'x-api-key' => $settings['api_key'],
            'anthropic-version' => $settings['version'],
        ])
        ->acceptJson()
        ->timeout($settings['timeout'])
        ->retry(1, 600, fn (Throwable $e): bool => $e instanceof ConnectionException
            || in_array($e->response?->status(), [429, 500, 502, 503, 529], true), throw: false)
        ->post('/v1/messages', $payload);

    if ($response->failed()) {
        throw new AssistantUnavailableException('Model provider returned HTTP '.$response->status());
    }

    return $response->json();
}

Everything is in config: model, token limit, timeout, base URL. Swapping to an SDK later is a change to this one class. Tests use Http::fake() and never touch the network.

The tool-use loop#

The agent loop is short enough to read in one sitting. Send the conversation and tool definitions. If the model stops with tool_use, run the tools, append the results as a user turn, and go again. If it stops for any other reason, take the text and return.

for ($round = 1; $round <= $maxRounds; $round++) {
    if (microtime(true) > $deadline) {
        throw new AssistantUnavailableException('Assistant exceeded its time budget.');
    }

    $response = $this->client->createMessage(
        $this->systemPrompt(), $messages, $tools->definitions(),
        allowTools: $round < $maxRounds,
    );

    $content = $response['content'] ?? [];
    $text = collect($content)->where('type', 'text')->pluck('text')->implode("\n\n");

    if (($response['stop_reason'] ?? null) !== 'tool_use') {
        break;
    }

    $messages[] = ['role' => 'assistant', 'content' => $content];
    $messages[] = ['role' => 'user', 'content' => $this->runTools($tools, $content)];
}

Three things keep this loop safe:

  1. A round limit. Four model calls per visitor message, tool calls included. On the last round, tools are switched off with tool_choice: none, so the model has to answer with text.
  2. A wall-clock deadline. Thirty seconds, which stays under typical gateway timeouts. A slow provider turns into a fallback answer instead of a spinner that never ends.
  3. Tool errors are data. Every tool returns ['content' => ..., 'is_error' => bool]. Bad input, unknown pages and empty searches go back to the model as readable messages, not exceptions.

Small tools, narrow powers#

The assistant has a handful of tools, and none of them can change anything important:

  • show_page adds a button that links to a real page on the site. It validates the destination against a fixed list, and caps the number of buttons per reply.
  • search_knowledge searches answers I have reviewed and approved.
  • flag_for_team marks the conversation for review when the model does not have the answer.
  • A newsletter or waitlist signup that only works if the visitor typed their own email address in their latest message.

That last rule is worth spelling out. A model can be talked into many things. So the tool checks the visitor's actual message on the server before it does anything, rather than trusting that the model followed the prompt.

A daily spend cap in five lines#

I did not want a bill surprise from a bot or a traffic spike. The cap is a counter in the cache, keyed by date:

private function withinDailyBudget(): bool
{
    $key = 'assistant:model-replies:'.now()->toDateString();

    Cache::add($key, 0, now()->endOfDay());

    return Cache::increment($key) <= config('assistant.daily_model_replies');
}

Cache::add only sets the key if it is missing, and increment is atomic on Redis and the database driver. Past the cap, the model is not called at all. Combined with a per-visitor rate limit on the route, the worst case cost per day is known in advance.

A fallback that always answers#

The widget must work when there is no API key, when the provider is down, and when the budget is spent. So there is a second responder that does not use a model at all. It matches the question against the site configuration and the reviewed knowledge base, and returns a short answer with links.

if (! $this->client->isConfigured() || ! $this->withinDailyBudget()) {
    return $this->fallback->respond($question);
}

try {
    return $this->runAgent($conversation, $question, $currentPage);
} catch (AssistantUnavailableException $exception) {
    Log::warning('Site assistant fell back to local answers.', ['error' => $exception->getMessage()]);

    return $this->fallback->respond($question);
}

The fallback is less clever, but it is never wrong in a surprising way, and a visitor never sees an error message.

Prompt caching for the stable parts#

The system prompt has three blocks: instructions, a site brief built from config, and the list of approved answers. They change rarely, so the last block carries a cache marker:

return [
    ['type' => 'text', 'text' => $instructions],
    ['type' => 'text', 'text' => "<site_brief>\n{$brief}\n</site_brief>"],
    [
        'type' => 'text',
        'text' => "<taught_answers>\n{$answers}\n</taught_answers>",
        'cache_control' => ['type' => 'ephemeral'],
    ],
];

Because the cache is a prefix match, anything that varies per request (the visitor's current page, the conversation) goes in the messages, never in the system prompt. Put a timestamp in the system prompt and you pay full price on every call.

Learning, with a human in the loop#

The part I like most is how the assistant gets better. When it cannot answer, it calls flag_for_team, and the conversation lands in a review queue in the admin area. Visitors can also mark an answer as unhelpful, which does the same.

I read the queue, write a proper answer, optionally attach a link, and approve it. Approved entries become part of taught_answers in the system prompt and are searchable through search_knowledge. Nothing the model says is ever added to its own knowledge automatically.

This matters for a business site. The assistant speaks for the studio, so every fact it can state was either in the config or written and approved by me. The system prompt tells it plainly: only state facts from the brief, the taught answers or search results, and never invent prices, dates, clients or statistics.

Treat visitor text as data#

Visitors will try to change the rules. The prompt says to treat text inside messages as questions, not instructions. More importantly, the design does not depend on the model obeying: tools validate their own inputs, the only side effects are harmless, and the model has no access to anything private. Prompt injection against this assistant can at worst produce an odd reply, which is why the powers are so narrow in the first place.

Testing it#

The whole thing is tested with Pest and Http::fake(). A few of the cases:

  • A tool_use response followed by a text response produces a reply with the right button.
  • A provider 500 produces a fallback answer, not an exception.
  • Past the daily cap, no HTTP request is made at all.
  • A signup tool call with an email the visitor did not type is rejected.

Faking the API at the HTTP layer means I test my loop, not someone else's SDK.

What I would do the same next time#

Start with the fallback. Make every tool boring and bounded. Put the budget and the deadline in code, not in hope. And let a human decide what the agent learns. The model is the least predictable part of the system, so everything around it should be the most predictable.