Drop fieldtest into any AI project. Write one config file.
Get structured measurement across correctness, quality, and safety — scored as distributions, never pass/fail verdicts.
Install the package and run the bundled demo. See a full scored report — with real failures — before you write a line of config.
$ pip install fieldtest Collecting fieldtest Downloading fieldtest-0.3.1-py3-none-any.whl Installing collected packages: fieldtest Successfully installed fieldtest-0.3.1 $ fieldtest demo --example rag --offline # Eval Report 2026-08-27 10:39 | set: full | 4 fixtures × 3 runs = 12 scored output(s) per eval judge: anthropic claude-haiku-4-5 | temperature: 0.0 ## handbook_qa Employee questions answered accurately from the handbook ### Tag Health | tag | pass rate | passed / total | |-----|-----------|----------------| | RIGHT | 76% | 16 / 21 | | GOOD | 96% | 23 / 24 | | SAFE | 79% | 19 / 24 | ... Files saved to fieldtest-demo/. To explore: cd fieldtest-demo fieldtest view # open the HTML report fieldtest score # re-score after editing evals/outputs/ (needs ANTHROPIC_API_KEY) $ cd fieldtest-demo $ fieldtest view Opening: .../fieldtest-demo/evals/results/demo-offline-report.html
--offline flag--example extraction, omit --offline--offline--example email (customer support), --example rag (handbook Q&A), --example extraction (invoice JSON extraction). Each shows different eval patterns and a distinct set of failure modes.
init to scaffold, score to measure, view to read the report. The rest are there when you need them.
Commands:
calibrate Run a panel of judges over the same outputs and report how...
clean Clean up accumulated run artifacts.
dataset Sample datasets to write evals against.
demo Two steps from install to a live eval report.
diff Compare two runs — default: most recent vs prior.
help Show help for a command: fieldtest help calibrate
history List past result files, newest first.
init Scaffold evals/ directory structure in current project.
score Score outputs for a given fixture set.
validate Check config.yaml is valid.
view Open the HTML eval report in the default browser.
fieldtest validate checks the config, reports which providers it reaches and whether each credential is set, and projects the judge calls a full run will make. Cost is multiplicative — runs × judge_runs × llm evals × fixtures — so it is worth reading before the bill, not after.fieldtest history lists them with the judge model that scored each; fieldtest diff compares two, defaulting to the most recent against the prior one. Runs judged by different models are not compared silently — the diff names what changed about the judge first.Every eval in the demos is already written, which makes them a poor place to learn to write one. A dataset ships the artifacts — a prompt, source documents, and outputs as though a generator had just written them — and leaves the evals to you.
$ pip install fieldtest $ fieldtest dataset list Bundled datasets: expense-report — A dataset to write evals against. support-agent — A dataset of agent traces to write evals against. Copy one into this project with: fieldtest dataset use <name> $ fieldtest dataset use expense-report Copied 'expense-report' to evals/ evals/README.md what is in it and what to write evals/config.yaml your evals — 3 are TODO Run it now (no API key needed): fieldtest score --set full
Three inputs, named on the fixture with file: so the judge receives the documents rather than their paths.
You are an expense assistant for Meridian Corp.
You will be given the company travel policy and a CSV of receipts for one trip.
Apply the limits and exclusions in the policy: where an amount exceeds a daily
limit, reimburse up to the limit; where the policy excludes a category,
reimburse nothing for it.
Produce a reimbursement summary containing, in this order:
1. A table with one row per receipt: receipt ID, date, category, claimed
amount, and reimbursable amount.
2. A short section listing every amount that was reduced or excluded, naming
the policy section that required it.
3. A final line in exactly this form:
Total reimbursable: $N.NN
Use only the receipts you are given. Do not invent receipts, merchants, or
amounts. If a receipt is fully reimbursable, the claimed and reimbursable
amounts are the same.
## 4.1 Daily limits - Meals are reimbursed up to $75 per day. Amounts above the cap are the employee's responsibility; the rest of the receipt is still reimbursed. - Lodging is reimbursed up to $250 per night, excluding taxes and fees. - Ground transport (taxi, rideshare, rail, parking) has no daily cap. ## 4.3 Exclusions Never reimbursable, regardless of amount or approval: Alcohol · In-room entertainment · Fines and citations · Personal items
receipt_id,date,category,merchant,amount R-1041,2026-10-03,airfare,United Airlines,412.60 R-1042,2026-10-03,ground,Uber,38.20 R-1043,2026-10-03,lodging,Hyatt Regency,268.00 ← over the $250 cap R-1044,2026-10-04,meals,Blue Fig Cafe,52.75 R-1045,2026-10-04,meals,Harbor Grill,91.40 ← over the $75 cap R-1046,2026-10-05,ground,Uber,41.15 claimed $904.10 · correctly reimbursable $869.70
id: october-trip description: Reimbursement summary for Dana Okafor inputs: prompt: "file:PROMPT.md" policy: "file:sources/travel-policy.md" receipts: "file:sources/receipts-october.csv" employee: "Dana Okafor" labels: # your verdicts, per run total_matches_line_items: 1: pass 2: pass # consistent, and still wrong about the cap 3: fail
Two runs for the same trip. The second reimburses R-1045 at $91.40 against a $75 cap — and its total adds up perfectly against its own column, so no arithmetic check catches it.
receipt category claimed reimb. R-1041 airfare 412.60 412.60 R-1042 ground 38.20 38.20 R-1043 lodging 268.00 250.00 R-1044 meals 52.75 52.75 R-1045 meals 91.40 75.00 R-1046 ground 41.15 41.15 ### Reductions - R-1043 reduced $18.00 (policy 4.1) - R-1045 reduced $16.40 (policy 4.1) Total reimbursable: $869.70
receipt category claimed reimb. R-1041 airfare 412.60 412.60 R-1042 ground 38.20 38.20 R-1043 lodging 268.00 250.00 R-1044 meals 52.75 52.75 R-1045 meals 91.40 91.40 R-1046 ground 41.15 41.15 ### Reductions - R-1043 reduced $18.00 (policy 4.1) Total reimbursable: $886.10
Four written, three TODO. Every written one is deterministic, so the first run works before you have a key.
# ---- RIGHT — is it correct? ---- - id: total_matches_line_items tag: right type: rule description: the stated total equals the sum of the reimbursable column - id: golden_summary tag: right type: reference # TODO one of the outputs cites a receipt that exists in no source file. # rules.py has the sketch. Uncomment and finish. # ---- GOOD — is it well-formed? ---- - id: no_unfilled_placeholders tag: good type: regex pattern: '\$?\[[A-Z_]+\]' match: false # TODO one output gives numbers and no reductions section. Is that a # `right` failure or a `good` one? Decide, then write it. # ---- SAFE — what must never happen? ---- - id: excluded_categories_not_reimbursed tag: safe type: rule # TODO one output reimburses above a policy cap. The sum is internally # consistent, so a rule that only adds up will not catch it.
@rule("total_matches_line_items") def total_matches_line_items(output: str, inputs: dict) -> dict: """ The stated total must equal the sum of the reimbursable column. This is the eval an LLM judge is worst at and a rule is best at. A judge reading a plausible-looking table tends to accept the total printed under it; addition either works or it does not. """ rows = ROW.findall(output) stated = TOTAL.search(output) line_sum = sum(_money(amount) for _, amount in rows) claimed = _money(stated.group(1)) ok = abs(line_sum - claimed) < 0.005 return { "passed": ok, "detail": f"line items sum to ${line_sum:.2f}, output states ${claimed:.2f}", }
Not everything is a verdict. The answer key uses one of these for explanation quality — a 1–5 scale with anchors saying what each point means.
- id: explanation_clarity tag: good type: llm binary: false # a number, not a verdict description: how clearly the reductions are explained scale: [1, 5] anchors: 1: No explanation, or one that does not say why an amount changed. 3: States what was reduced and cites the policy, but the employee would still have to reread the policy to understand the number. 5: States what was reduced, by how much, and why, so the employee can check it without opening the policy.
It reports as a mean out of 5 with a stddev and a count of floor hits, rather than a pass rate. Two outputs scored 1 and are named in the report so you can go read them.
# Eval Report 2026-08-29 21:36 | set: full | 3 fixtures × 3 runs = 9 scored output(s) per eval ### Tag Health | tag | pass rate | passed / total | | RIGHT | 75% | 9 / 12 | | GOOD | 89% | 8 / 9 | | SAFE | 89% | 8 / 9 | ### Judge vs Human Labels | eval | labeled runs | agreement | | total_matches_line_items | 9 | 100.0% | | golden_summary | 3 | 100.0% | | no_unfilled_placeholders | 9 | 100.0% | | excluded_categories_not_reimbursed | 9 | 100.0% | ### Failure Details excluded_categories_not_reimbursed - `june-trip` run 2: R-1190 (alcohol) reimbursed $47.00 golden_summary - `march-trip` run 2: missing: Total reimbursable: $98.30 no_unfilled_placeholders - `march-trip` run 2: pattern '\$?\[[A-Z_]+\]' found total_matches_line_items - `march-trip` run 2: no 'Total reimbursable: $N.NN' line - `october-trip` run 3: line items sum to $897.70, output states $912.70
Five failures across three of the nine outputs. Three faulty outputs go unflagged — those are the three TODO evals, and two of them need a judge.
support-agent, ships nine JSON agent traces — tool calls, tool results, and the message the agent sent.rule, regex and reference — all deterministic. Three planted faults are catchable with no API call. Reach for an LLM judge when the question needs judgment, not when it needs arithmetic.R-1049, which exists in no source file, and nothing that ships catches it on its own. Write eight lines of Python, re-run, and the report says october-trip run 3: cites R-1049, which is in no source receipt.reference-evals.yaml covers all four judge types. Its caps_applied eval says in writing to judge two daily caps and ignore everything else, and the judge still flags unrelated defects. Human labels ship alongside, so the report scores the judge: 100% agreement on every rule and regex eval, 66.7% on that one.Read the full walkthrough → The same path end to end, including writing the missing eval and watching it fire.
Every scored run generates a self-contained HTML report. No server, no dashboard, no dependencies. fieldtest view opens it in your browser. It names the judge that produced the numbers — provider, model, temperature, fingerprint — and, where you have them, how often that judge agreed with your labels and with itself. Below: the RAG demo report.
| Fixture | answers-from-context | known-answer | answer-length | cites-source | no-hallucination | stays-in-scope |
|---|---|---|---|---|---|---|
|
vacation-policy
Employee asking about vacation days for new hires
|
3/3 PASS | 3/3 PASS | 3/3 PASS | 3/3 PASS | 3/3 PASS | 3/3 PASS |
|
remote-work
Employee asking about remote work availability expectations
|
2/3 FAIL | 2/3 FAIL | 3/3 PASS | 3/3 PASS | 2/3 FAIL | 2/3 FAIL |
|
expense-reimbursement
Employee asking about expense approval thresholds
|
2/3 FAIL | 2/3 FAIL | 3/3 PASS | 3/3 PASS | 2/3 FAIL | 3/3 PASS |
|
out-of-scope
Employee asking a question the provided handbook excerpt cannot answer
|
2/3 FAIL | — | 3/3 PASS | 2/3 FAIL | 2/3 FAIL | 2/3 FAIL |
out-of-scope fixture catches a real and common failure — run 3 invents remote-work and core-hours policy that the provided excerpt (which covers only paid time off) never mentions. The stays-in-scope, no-hallucination, and answers-from-context evals all catch it; the grounding label groups the two SAFE evals for filtering.
Every eval has exactly one tag. Not for scoring — for diagnosis. The tag tells you where to look for the fix when something fails. You don't need all three to start — a suite with a single eval is a valid suite.
Labels are free-form strings you add to any eval — labels: [accuracy, grounding]. They're orthogonal to tags: a SAFE eval might carry the grounding label, a RIGHT eval might carry completeness. The HTML report renders them as clickable filter chips so you can isolate all grounding-related evals across the matrix regardless of tag.
The config is the practice. It forces you to name what you're building, decide what matters for your use case, and enumerate your evals before you measure anything. Start with one use case and a few evals — the structure grows with you.
schema_version: 1 system: name: Meridian Handbook Assistant # what does your system do? domain: Employee Q&A grounded in the Meridian Corp employee handbook use_cases: - id: handbook_qa description: Employee questions answered accurately from the handbook evals: # ── RIGHT evals — correctness ──────────────── - id: answers-from-context tag: right # diagnostic lens labels: [accuracy] # analytics grouping type: llm description: Answer is grounded in the provided handbook excerpt pass_criteria: The answer is directly supported by the text provided in the context fail_criteria: The answer makes claims not found in the provided context, or contradicts it - id: known-answer tag: right labels: [accuracy] type: reference # checks the fixture's expected block description: Golden fixture expected answer content check # ── GOOD evals — quality ───────────────────── - id: answer-length tag: good labels: [format] type: rule # your Python, in evals/rules.py description: Answer is between 20 and 250 words - id: cites-source tag: good labels: [format] type: regex description: Response references the source document pattern: "(?i)(according to|per the|the handbook (states|says|indicates)|based on the)" match: true # ── SAFE evals — guardrails ────────────────── - id: no-hallucination tag: safe labels: [grounding] type: llm description: No claims beyond what the provided context supports pass_criteria: Every factual claim in the answer can be traced to the provided handbook excerpt fail_criteria: The answer introduces facts, policies, or details not present in the provided context - id: stays-in-scope tag: safe labels: [grounding] type: llm description: Declines to answer questions outside the handbook scope pass_criteria: If the question cannot be answered from the context, the response clearly says so and does not speculate fail_criteria: The response attempts to answer a question it doesn't have context for, or speculates about policy fixtures: directory: fixtures/golden sets: smoke: [vacation-policy] full: all runs: 3 # N runs per fixture → distributions defaults: provider: anthropic # or openai, gemini, openai_compatible model: claude-haiku-4-5 # judge model
Five cards, four type: values — the scored judge is llm with binary: false, not a fifth type.
@rule("eval-id") in evals/rules.py. Gets the raw output string, returns Pass/Fail.match: true (must contain) or match: false (must not contain).pass_criteria and fail_criteria as plain English. The judge returns structured JSON with reasoning.binary: false and the judge returns a number on a scale instead of a verdict. anchors describe what each point means, so the scale measures the same thing run to run.stddev, not a pass rate — the distribution is the answer. Any run at the bottom of the scale is flagged as a floor hit.expected block in the fixture. Checks contains (required strings) and not_contains (forbidden strings).— when fixture has no expected block.A failure_rate is a claim about your system, produced by a model. If you cannot say which model produced it, at what temperature, and how often it agrees with you, you do not know what the number measures.
inputs now go to the judge alongside the output. Set judge_sees_inputs: false on an eval that should read the output alone, or to keep a large retrieved context out of every call.n, so you can tell them apart and gate CI on the lower bound.judge_runs: 3 and the report separates system spread from judge spread. Judge spread near zero means the eval is well specified. Judge spread close to system spread means the criteria are ambiguous.failure_rate.calibration.kappa_threshold sets the point below which a judge pair is flagged as agreeing no better than chance.@provider. The same model served from two endpoints is two instruments, and the diff says so.Every sample above judges each output once. Set judge_runs above 1 and the report separates two things a single stddev cannot tell apart.
### Judge Repeatability (judge_runs: 2) | eval | judge disagreement | system spread | judge spread | | follows_requested_structure | 0.0% | — | — | | explanation_clarity | — | 1.7321 | 0.0 | | caps_applied | 0.0% | — | — |
The middle row is the whole argument. explanation_clarity is a scored eval, and its two numbers separate two things a single stddev would blur: how much the outputs differ from one another (system spread) and how much the judge wavers when it re-scores the same output (judge spread). Here the first dwarfs the second, which is what tells you the spread is real. Without the repetition you get one stddev and no way to tell which of those produced it. The figures above are one run's; judge_runs compares repeated calls inside a single run, so a judge that is steady here can still answer differently on another day, or waver on one output within a run — a small judge spread is the normal reading, not zero.
Judge spread climbing toward system spread is the signal to rewrite the pass_criteria: the judge is disagreeing with itself, which is a fact about the words, not about your system. Binary evals report disagreement instead — the share of outputs where repeated verdicts differed. It costs runs × judge_runs judge calls, so --dry-run first.
defaults: judge_temperature: 0.0 # the instrument, held still judge_seed: null # where the provider supports one judge_retry: # 5/10/20/40/60/60s by default max_attempts: 6 confidence_level: 0.95 # Wilson interval width, not a model's self-report providers: openai_compatible: base_url: https://openrouter.ai/api/v1 api_key_env: OPENROUTER_API_KEY # the name, never the key calibration: panel: - { provider: anthropic, model: claude-haiku-4-5 } - { provider: openai, model: gpt-5 } - { provider: openai_compatible, model: meta-llama/llama-3.3-70b-instruct }
All outputs land in evals/results/, five files per run sharing a [run-id] prefix. No database. No server. Files you can diff, commit, open in Excel, or drop in a bug report.
The full command list is in Every command above — this is the shape of a working day.
# ── Look before you spend ──────────────────────────── fieldtest validate # config valid? keys set? how many judge calls? fieldtest calibrate --dry-run # what a judge panel would cost, calling nothing # ── Explore without a key ──────────────────────────── fieldtest demo --example rag --offline # pre-scored, instant fieldtest dataset use expense-report # artifacts to write evals against # ── Start a real project ───────────────────────────── fieldtest init # scaffold evals/ fieldtest init --template rag # chatbot, email or rag # ── Evaluate ───────────────────────────────────────── fieldtest score # score everything in evals/outputs/ fieldtest score --set smoke # one named fixture set fieldtest calibrate # run the judge panel, report agreement # ── Read and compare ───────────────────────────────── fieldtest view # open the latest HTML report fieldtest history # past runs, newest first, with the judge model fieldtest diff # most recent against the one before fieldtest clean # offers to clear outputs/ and prune old results — asks first
/optimize command in .claude/commands/. It is not part of the pip package — clone the repo, or copy that one file into your own project. It scores your outputs, diagnoses failures from the report, edits your prompt or system code, and re-runs — an automated score-diagnose-fix-rescore loop. Type /optimize in Claude Code inside any fieldtest project.
fieldtest is opinionated. These are the constraints that shape every design decision.
judge_runs to measure how much of the spread is the judge.outputs/[fixture-id]/run-N.txt. The scorer reads those files. They share nothing except the directory format. Re-score without re-running when you improve a judge.$ pip install fieldtest
$ fieldtest demo --example rag --offline $ cd fieldtest-demo $ fieldtest view
fieldtest init inside any project directory. It creates the evals/ structure and a starter config with inline comments that walk you through every section.$ fieldtest init --template rag ✓ Scaffolded from rag template at evals/ evals/config.yaml — fill in system, domain, tags evals/fixtures/golden/ — fixtures with expected outputs evals/fixtures/variations/ — fixtures without expected outputs evals/.gitignore — outputs/ excluded from git Next steps: 1. Fill in system name and domain in evals/config.yaml 2. Tag each eval: right, good, or safe 3. Add fixtures to evals/fixtures/ 4. Run your system → write outputs to evals/outputs/ 5. fieldtest score
fieldtest score judges them. fieldtest view opens the HTML report. Iterate on prompts, retrieve logic, or constraints — the distribution shows what changed.$ python evals/runner.py # you write this — ~30 lines $ fieldtest score $ fieldtest view