Everyone is shipping agents. Almost nobody is measuring them.

A normal test is a promise: given this input, expect that output, every time. Agents break that promise on purpose. Run the same prompt twice and you get two different answers, both plausibly correct, and neither one something you can assert equality against.

So the industry mostly stopped testing. Prompts get tweaked, a couple of examples get tried by hand, and it goes out. Then a model version changes, or someone edits a system prompt, and a quiet regression ships. Nobody notices until a customer does.

Vizra exists because "it seemed fine when I tried it" is not a test, and the fix is not heroic — it is measurement.

What we think is true

01

Sample, do not guess

One run of a nondeterministic system tells you almost nothing. Run every row several times and you get a distribution — which is the difference between "it got worse" and "it rolled badly".

02

Scores, then policy

A score is a measurement. Pass or fail is a decision you make about that measurement, at the run level, where you can see the whole picture. Collapsing the two too early throws away the information you needed.

03

Evals belong in the test suite

Not a separate tool with its own dashboard and its own login. If it lives where your tests live, it runs when your tests run, and it fails your build like anything else.

04

Reasons outlive numbers

A score of 0.61 tells you nothing in three months. The judge's reasoning, the actual response and the tool calls tell you everything — so they get stored, not summarised away.

05

Your data is yours

The framework writes to your database and phones nobody. The hosted service is opt-in, never holds your model keys, and never runs your code. That is a design constraint, not a policy we could quietly change.

06

Free has to be genuinely useful

The open-source package is the whole engine, not a demo with the good parts removed. If it only worked as a funnel into a subscription it would not be worth installing.

Where it came from

Vizra started as an agent framework for Laravel. Building agents turned out to be the easy part; knowing whether a change had made them worse turned out to be the hard part, and nothing in the Laravel ecosystem answered it.

With laravel/ai now the official way to build agents in Laravel, another framework was not what was missing. So Vizra became the thing that was: the evals layer, written as Pest tests, for agents you build however you like.

The other half came from watching non-developers — the people who actually know whether an answer is good — be unable to check anything without booking a developer's afternoon. That is what the Run button in the hosted dashboard is for.