Local-first agent evaluation

Build evaluations
you can trust.

Valcore is a focused workbench for authoring agent judges, building datasets, and running agent evaluations—from a visual UI, your terminal, or CI.

$brew install duncankmckinnon/tap/valcore
Get started

Open source · Apache 2.0 · Python 3.11+

validation runrun_029
94.2%agreement
support-agent-v4✓ threshold passed

One tight loop for better agents.

01 Author the agent judge02 Build the dataset03 Validate and ship
The Valcore workflow

From a judgment call
to a release signal.

Valcore keeps the agent judge, the evidence, and every result connected in one local-first workflow.

01
Define the standard

Author an evaluator

Turn the behavior you care about into an agent judge with explicit inputs, a structured score, and versioned harness capabilities.

Evaluator guide
02
Build the evidence

Shape a dataset

Bring in real Logfire traces, upload existing cases, or generate synthetic edge cases—then label the examples that matter.

Dataset guide
03
Measure the change

Validate, compare, ship

Measure agreement with human labels, compare evaluator versions on the same data, and enforce release thresholds in CI.

Experiment guide
Run locally. Keep control.

Start with your first evaluator.

Open the guide