Your first eval suite, written by the person who wrote the framework.
You already know the agent needs evals. What you do not have is a fortnight to build a dataset, pick judges, and argue with a baseline. I do this for a small number of Laravel teams at a fixed price, in your repo, from your real traffic.
Nobody wakes up wanting evals. They wake up with one of these.
The bot goes live in three weeks
Support, sales, onboarding: something that talks to customers is about to ship, and "it seemed fine when I tried it" is the entire QA plan. You need a suite before launch, and a number you can show whoever signs it off.
The model is being deprecated
A provider has given you a date. You have a system prompt tuned by feel over six months and no way to tell whether the replacement is better, worse, or differently broken. You need a baseline before you switch.
A client asked for proof
You built the agent for someone else, and they want to know it works. Not a demo: evidence, repeatable, that survives the next prompt change. You need a report that a non-developer can read.
Fixed scope, fixed price, no timesheets.
Each one ends with something you own outright. Prices are a floor; the Sprint quote is fixed once I have seen the repo and the traffic.
Eval Review
I read your agent and a sample of real traffic, then tell you what to test first and how.
- A written plan: which behaviours to eval, what the dataset should look like, which judges fit
- A 45-minute call to walk through it
- Credited against a Sprint if you go ahead within 30 days
Eval Sprint
most asked forYour first suite, built in your repo from your real traffic, passing in CI, handed over.
- A dataset built from your real conversations or logs
- One eval suite covering the behaviours that matter
- A recorded baseline and a CI gate that fails on regression
- A handover session and a short written guide to extending it
Retainer
Someone watching quality so your team does not have to.
- A quarterly regression report
- Model migrations when a provider deprecates what you are on
- New suites as features ship
- Async questions answered within two working days
Prices exclude VAT. The Sprint is spread over four weeks because I do this alongside other work, which suits a go-live a month or more out. If your date is closer than that, say so in the form and I will tell you honestly whether it fits.
Async by default. You keep everything.
A short call
Thirty minutes. What the agent does, what would embarrass you if it went wrong, and when it ships. If it is not a fit I will say so on the call.
Scope and a fixed quote
I look at the repo and a sample of real traffic, then send a one-page scope with the behaviours the suite will cover and a fixed price. No estimate creep.
I work in your repo
On a branch, with pull requests you review. The dataset, the evals, the judges and the CI gate grow where your tests already live. You watch it happen; nothing is built somewhere else and delivered in a zip.
Handover
A walkthrough, a short written guide to adding your next eval, and access revoked. The suite is yours under the same MIT licence as the framework.
- Read access to the repo A branch of my own is enough. I never push to main.
- A sample of real traffic Conversations, logs, support tickets. Anonymised is fine. Real beats invented every time.
- A model key with a spend cap Created by you, capped by you, revoked by you when we are done. Evals run against the real model, so they cost real tokens.
- Someone who knows what "good" looks like An hour or two of their time to label the tricky cases. That person is usually not a developer.
One person. The one who wrote it.
I am Aaron. I wrote vizra/evals, its
dashboard, and Vizra Cloud, and I have been building Laravel applications for a long
time before any of them existed. When you hire this, you get me, not a team with my name
on the proposal.
Every engagement feeds the framework
A pattern that shows up in your suite becomes a judge, a helper, or a docs page. You are paying for evals; the ecosystem gets the generalised version.
You own the output
The suite lives in your repo under MIT. Cancel the retainer and nothing stops working.
No lock-in to Cloud
The suite runs against your own database. Vizra Cloud is optional and stays optional.
Honest about fit
If your agent is not on Laravel, or the problem is not one evals solve, I will say so on the first call rather than the last invoice.
Questions people ask first
Do I need Vizra Cloud? +
No. Everything I build runs on the open-source packages against your own database. Cloud is useful when the whole team wants to see the history, and you can add it later or never.
Who owns the code? +
You do. The suite is committed to your repo under the same MIT licence as the framework. There is nothing to license, renew, or export.
Does the product see my keys after we are done? +
It never did. Vizra Cloud only ever receives finished runs. During an engagement, I use a key you created with a cap you set, and you revoke it at handover.
Which agent frameworks do you work with? +
Anything vizra/evals can target: the Laravel AI SDK, Prism, or a hand-rolled client around an HTTP call. If your agent is a PHP class that returns a response, it can be evaluated.
We are not on Laravel. Can you still help? +
Probably not well, and I would rather say so than take the work. The framework is Laravel-native and that is where I am useful. A short call is free if you want a second opinion on approach.
How soon can you start? +
A Review usually within a week or two. A Sprint depends on what I am already committed to; put your go-live date in the form and I will tell you straight away whether it fits.
What if the agent turns out to be worse than we thought? +
Then the Sprint did its job early. A baseline that says 61% is more useful than a launch that finds out in production, and the suite is exactly the tool for fixing it.
Tell me about the agent.
The extra fields are optional, but each one saves a round of email. I read every enquiry myself and reply within two working days, usually sooner.
Rather try it yourself first?
The quickstart goes from install to a passing eval in about five minutes. Come back if you get stuck.
Read the quickstart →Something else?
Bugs, billing, or a question about the product go through the contact page.
Contact →