fieldtest v0.3.1 release

Eval practice,
not just eval tooling.

Drop fieldtest into any AI project. Write one config file. Get structured measurement across correctness, quality, and safety — scored as distributions, never pass/fail verdicts.

1
tag per eval — RIGHT, GOOD, or SAFE
0
pass/fail verdicts — distributions only
5
files per run — diff, commit, grep
0
infrastructure — no DB, server, or dashboard

Two commands. No setup.

Install the package and run the bundled demo. See a full scored report — with real failures — before you write a line of config.

Terminal
$ pip install fieldtest
Collecting fieldtest
  Downloading fieldtest-0.3.1-py3-none-any.whl
Installing collected packages: fieldtest
Successfully installed fieldtest-0.3.1

$ fieldtest demo --example rag --offline
# Eval Report
2026-08-27 10:39 | set: full | 4 fixtures × 3 runs = 12 scored output(s) per eval
judge: anthropic claude-haiku-4-5 | temperature: 0.0

## handbook_qa
Employee questions answered accurately from the handbook

### Tag Health
| tag | pass rate | passed / total |
|-----|-----------|----------------|
| RIGHT | 76% | 16 / 21 |
| GOOD | 96% | 23 / 24 |
| SAFE | 79% | 19 / 24 |
...
Files saved to fieldtest-demo/. To explore:
  cd fieldtest-demo
  fieldtest view            # open the HTML report
  fieldtest score           # re-score after editing evals/outputs/  (needs ANTHROPIC_API_KEY)

$ cd fieldtest-demo
$ fieldtest view
Opening: .../fieldtest-demo/evals/results/demo-offline-report.html

Pick your entry point

Offline
Pre-scored results
Uses bundled outputs and results. Full report in under a second. Nothing to configure, no credentials needed.
No API key needed
Add --offline flag
Live extraction
Real scoring, no key
Re-scores the bundled outputs for real. The extraction example's rule and regex evals run fully; its two LLM evals are marked as errors and excluded from rates.
No API key needed
--example extraction, omit --offline
Full live
All evals, LLM judges included
Re-scores the bundled demo outputs with every eval type firing, LLM judges included. fieldtest never calls your system — the demo scores the outputs it ships.
ANTHROPIC_API_KEY required
Omit --offline
Three examples available: --example email (customer support), --example rag (handbook Q&A), --example extraction (invoice JSON extraction). Each shows different eval patterns and a distinct set of failure modes.

Eleven commands, and you need three.

init to scaffold, score to measure, view to read the report. The rest are there when you need them.

$ fieldtest --help
Commands:
  calibrate  Run a panel of judges over the same outputs and report how...
  clean      Clean up accumulated run artifacts.
  dataset    Sample datasets to write evals against.
  demo       Two steps from install to a live eval report.
  diff       Compare two runs — default: most recent vs prior.
  help       Show help for a command: fieldtest help calibrate
  history    List past result files, newest first.
  init       Scaffold evals/ directory structure in current project.
  score      Score outputs for a given fixture set.
  validate   Check config.yaml is valid.
  view       Open the HTML eval report in the default browser.
✅
validate before you spend
fieldtest validate checks the config, reports which providers it reaches and whether each credential is set, and projects the judge calls a full run will make. Cost is multiplicative — runs × judge_runs × llm evals × fixtures — so it is worth reading before the bill, not after.
📉
history and diff
Every run writes a timestamped result set. fieldtest history lists them with the judge model that scored each; fieldtest diff compares two, defaulting to the most recent against the prior one. Runs judged by different models are not compared silently — the diff names what changed about the judge first.

The demos are finished. A dataset is not.

Every eval in the demos is already written, which makes them a poor place to learn to write one. A dataset ships the artifacts — a prompt, source documents, and outputs as though a generator had just written them — and leaves the evals to you.

Terminal
$ pip install fieldtest
$ fieldtest dataset list
Bundled datasets:
  expense-report — A dataset to write evals against.
  support-agent — A dataset of agent traces to write evals against.

Copy one into this project with:  fieldtest dataset use <name>

$ fieldtest dataset use expense-report
Copied 'expense-report' to evals/
  evals/README.md   what is in it and what to write
  evals/config.yaml your evals — 3 are TODO

Run it now (no API key needed):  fieldtest score --set full

What the assistant was given

Three inputs, named on the fixture with file: so the judge receives the documents rather than their paths.

evals/PROMPT.md
You are an expense assistant for Meridian Corp.

You will be given the company travel policy and a CSV of receipts for one trip.

Apply the limits and exclusions in the policy: where an amount exceeds a daily
limit, reimburse up to the limit; where the policy excludes a category,
reimburse nothing for it.

Produce a reimbursement summary containing, in this order:

1. A table with one row per receipt: receipt ID, date, category, claimed
   amount, and reimbursable amount.
2. A short section listing every amount that was reduced or excluded, naming
   the policy section that required it.
3. A final line in exactly this form:

   Total reimbursable: $N.NN

Use only the receipts you are given. Do not invent receipts, merchants, or
amounts. If a receipt is fully reimbursable, the claimed and reimbursable
amounts are the same.
evals/sources/travel-policy.md — excerpt
## 4.1 Daily limits

- Meals are reimbursed up to $75 per day. Amounts above the cap are the
  employee's responsibility; the rest of the receipt is still reimbursed.
- Lodging is reimbursed up to $250 per night, excluding taxes and fees.
- Ground transport (taxi, rideshare, rail, parking) has no daily cap.

## 4.3 Exclusions

Never reimbursable, regardless of amount or approval:
  Alcohol · In-room entertainment · Fines and citations · Personal items
evals/sources/receipts-october.csv — annotated
receipt_id,date,category,merchant,amount
R-1041,2026-10-03,airfare,United Airlines,412.60
R-1042,2026-10-03,ground,Uber,38.20
R-1043,2026-10-03,lodging,Hyatt Regency,268.00   ← over the $250 cap
R-1044,2026-10-04,meals,Blue Fig Cafe,52.75
R-1045,2026-10-04,meals,Harbor Grill,91.40    ← over the $75 cap
R-1046,2026-10-05,ground,Uber,41.15

claimed $904.10 · correctly reimbursable $869.70
evals/fixtures/october-trip.yaml — annotated
id: october-trip
description: Reimbursement summary for Dana Okafor

inputs:
  prompt:   "file:PROMPT.md"
  policy:   "file:sources/travel-policy.md"
  receipts: "file:sources/receipts-october.csv"
  employee: "Dana Okafor"

labels:                          # your verdicts, per run
  total_matches_line_items:
    1: pass
    2: pass      # consistent, and still wrong about the cap
    3: fail

What it produced

Two runs for the same trip. The second reimburses R-1045 at $91.40 against a $75 cap — and its total adds up perfectly against its own column, so no arithmetic check catches it.

run-1.txt — clean
receipt  category   claimed  reimb.
R-1041   airfare    412.60   412.60
R-1042   ground      38.20    38.20
R-1043   lodging    268.00   250.00
R-1044   meals       52.75    52.75
R-1045   meals       91.40    75.00
R-1046   ground      41.15    41.15

### Reductions
- R-1043 reduced $18.00 (policy 4.1)
- R-1045 reduced $16.40 (policy 4.1)

Total reimbursable: $869.70
run-2.txt — meal cap missed
receipt  category   claimed  reimb.
R-1041   airfare    412.60   412.60
R-1042   ground      38.20    38.20
R-1043   lodging    268.00   250.00
R-1044   meals       52.75    52.75
R-1045   meals       91.40     91.40
R-1046   ground      41.15    41.15

### Reductions
- R-1043 reduced $18.00 (policy 4.1)


Total reimbursable: $886.10
$886.10 is arithmetically correct. Catching this one means reading the policy, not adding up — which is the difference between a rule and a judge, and the eval the dataset leaves for you to write.

The evals that ship

Four written, three TODO. Every written one is deterministic, so the first run works before you have a key.

evals/config.yaml — abridged
# ---- RIGHT — is it correct? ----
- id: total_matches_line_items
  tag: right
  type: rule
  description: the stated total equals the sum of the reimbursable column

- id: golden_summary
  tag: right
  type: reference

# TODO one of the outputs cites a receipt that exists in no source file.
#      rules.py has the sketch. Uncomment and finish.

# ---- GOOD — is it well-formed? ----
- id: no_unfilled_placeholders
  tag: good
  type: regex
  pattern: '\$?\[[A-Z_]+\]'
  match: false

# TODO one output gives numbers and no reductions section. Is that a
#      `right` failure or a `good` one? Decide, then write it.

# ---- SAFE — what must never happen? ----
- id: excluded_categories_not_reimbursed
  tag: safe
  type: rule

# TODO one output reimburses above a policy cap. The sum is internally
#      consistent, so a rule that only adds up will not catch it.
evals/rules.py
@rule("total_matches_line_items")
def total_matches_line_items(output: str, inputs: dict) -> dict:
    """
    The stated total must equal the sum of the reimbursable column.

    This is the eval an LLM judge is worst at and a rule is best at. A judge
    reading a plausible-looking table tends to accept the total printed under
    it; addition either works or it does not.
    """
    rows = ROW.findall(output)
    stated = TOTAL.search(output)
    line_sum = sum(_money(amount) for _, amount in rows)
    claimed  = _money(stated.group(1))
    ok = abs(line_sum - claimed) < 0.005
    return {
        "passed": ok,
        "detail": f"line items sum to ${line_sum:.2f}, output states ${claimed:.2f}",
    }

A scored eval, for comparison

Not everything is a verdict. The answer key uses one of these for explanation quality — a 1–5 scale with anchors saying what each point means.

evals/reference-evals.yaml — annotated
- id: explanation_clarity
  tag: good
  type: llm
  binary: false                 # a number, not a verdict
  description: how clearly the reductions are explained
  scale: [1, 5]
  anchors:
    1: No explanation, or one that does not say why an amount changed.
    3: States what was reduced and cites the policy, but the employee
       would still have to reread the policy to understand the number.
    5: States what was reduced, by how much, and why, so the employee
       can check it without opening the policy.

It reports as a mean out of 5 with a stddev and a count of floor hits, rather than a pass rate. Two outputs scored 1 and are named in the report so you can go read them.

The run

$ fieldtest score --set full
# Eval Report
2026-08-29 21:36 | set: full | 3 fixtures × 3 runs = 9 scored output(s) per eval

### Tag Health
| tag   | pass rate | passed / total |
| RIGHT | 75%       | 9 / 12         |
| GOOD  | 89%       | 8 / 9          |
| SAFE  | 89%       | 8 / 9          |

### Judge vs Human Labels
| eval                               | labeled runs | agreement |
| total_matches_line_items           | 9            | 100.0%    |
| golden_summary                     | 3            | 100.0%    |
| no_unfilled_placeholders           | 9            | 100.0%    |
| excluded_categories_not_reimbursed | 9            | 100.0%    |

### Failure Details

excluded_categories_not_reimbursed
- `june-trip` run 2: R-1190 (alcohol) reimbursed $47.00

golden_summary
- `march-trip` run 2: missing: Total reimbursable: $98.30

no_unfilled_placeholders
- `march-trip` run 2: pattern '\$?\[[A-Z_]+\]' found

total_matches_line_items
- `march-trip` run 2: no 'Total reimbursable: $N.NN' line
- `october-trip` run 3: line items sum to $897.70, output states $912.70

Five failures across three of the nine outputs. Three faulty outputs go unflagged — those are the three TODO evals, and two of them need a judge.

📄
Artifacts, not answers
Six of the nine outputs carry a deliberate fault; three are clean. The prompt is a fixture input, so did it do what was asked is an eval you can write. A second dataset, support-agent, ships nine JSON agent traces — tool calls, tool results, and the message the agent sent.
🔌
Runs before you have a key
The filled-in evals are rule, regex and reference — all deterministic. Three planted faults are catchable with no API call. Reach for an LLM judge when the question needs judgment, not when it needs arithmetic.
✍️
You write the missing one
One output cites receipt R-1049, which exists in no source file, and nothing that ships catches it on its own. Write eight lines of Python, re-run, and the report says october-trip run 3: cites R-1049, which is in no source receipt.
🗝️
An answer key that is wrong on purpose
reference-evals.yaml covers all four judge types. Its caps_applied eval says in writing to judge two daily caps and ignore everything else, and the judge still flags unrelated defects. Human labels ship alongside, so the report scores the judge: 100% agreement on every rule and regex eval, 66.7% on that one.

Read the full walkthrough → The same path end to end, including writing the missing eval and watching it fire.

Everything in one file.

Every scored run generates a self-contained HTML report. No server, no dashboard, no dependencies. fieldtest view opens it in your browser. It names the judge that produced the numbers — provider, model, temperature, fingerprint — and, where you have them, how often that judge agreed with your labels and with itself. Below: the RAG demo report.

fieldtest
run 2026-08-27T10-39-43-e327
example rag — Meridian Handbook Assistant
4 fixtures · 6 evals · 3 runs each
2026-08-27
RIGHT
76%
correctness
GOOD
96%
quality
SAFE
79%
guardrails
How to read this
Tags tell you where to look for the fix, not just what failed. RIGHT → prompt or training. GOOD → formatting or tone. SAFE → guardrails.
Filter by label: all accuracy format grounding
Fixture answers-from-context known-answer answer-length cites-source no-hallucination stays-in-scope
vacation-policy
Employee asking about vacation days for new hires
3/3 PASS 3/3 PASS 3/3 PASS 3/3 PASS 3/3 PASS 3/3 PASS
remote-work
Employee asking about remote work availability expectations
2/3 FAIL 2/3 FAIL 3/3 PASS 3/3 PASS 2/3 FAIL 2/3 FAIL
expense-reimbursement
Employee asking about expense approval thresholds
2/3 FAIL 2/3 FAIL 3/3 PASS 3/3 PASS 2/3 FAIL 3/3 PASS
out-of-scope
Employee asking a question the provided handbook excerpt cannot answer
2/3 FAIL — 3/3 PASS 2/3 FAIL 2/3 FAIL 2/3 FAIL
↓ out-of-scope / no-hallucination — click any cell to expand
Run 1 PASS
The output correctly acknowledges that the requested information is not in the provided context and does not introduce unsupported claims about remote work or core hours policies. type: llm · grounding label
Run 2 PASS
The output correctly acknowledges that the provided context does not contain information about remote work or core hours, and appropriately directs the user to other resources rather than inventing unsupported claims. type: llm · grounding label
Run 3 FAIL
The output introduces remote work policies, core hours, and business hour requirements that are not mentioned anywhere in the provided handbook excerpt, which only covers paid time off policies. type: llm · grounding label
Generated by fieldtest · self-contained HTML · no server required
What you're looking at: The out-of-scope fixture catches a real and common failure — run 3 invents remote-work and core-hours policy that the provided excerpt (which covers only paid time off) never mentions. The stays-in-scope, no-hallucination, and answers-from-context evals all catch it; the grounding label groups the two SAFE evals for filtering.

Right / Good / Safe

Every eval has exactly one tag. Not for scoring — for diagnosis. The tag tells you where to look for the fix when something fails. You don't need all three to start — a suite with a single eval is a valid suite.

RIGHT
Is the answer correct?
Correctness relative to ground truth. Did the system answer the question? Did it retrieve the right information? Does the output match the expected reference?
  • Known-answer reference checks
  • Required field presence
  • Addresses the user's actual question
  • Extracted value matches source
Failures point to: prompt, retrieval, training data, or model capability
GOOD
Is the answer well-formed?
Quality beyond correctness. Appropriate tone, format, length, and style. The answer is right — but is it delivered well for this context?
  • Greeting present in support email
  • Response within expected length range
  • Tone appropriate to the customer
  • Cites the source section
Failures point to: prompt instructions, formatting rules, or output post-processing
SAFE
Does it stay in bounds?
Guardrails and constraints. Does the system stay within its defined scope? Does it avoid fabricating, overreaching, or making unauthorized commitments?
  • No hallucinated policy details
  • No invented JSON fields
  • No unauthorized pricing commitments
  • Declines unanswerable questions
Failures point to: system prompt constraints, grounding instructions, or architecture (RAG retrieval scope)

Analytics grouping on top of tags

Labels are free-form strings you add to any eval — labels: [accuracy, grounding]. They're orthogonal to tags: a SAFE eval might carry the grounding label, a RIGHT eval might carry completeness. The HTML report renders them as clickable filter chips so you can isolate all grounding-related evals across the matrix regardless of tag.

One file defines everything.

The config is the practice. It forces you to name what you're building, decide what matters for your use case, and enumerate your evals before you measure anything. Start with one use case and a few evals — the structure grows with you.

config.yaml (rag example — the shipped file, comments added)
schema_version: 1

system:
  name: Meridian Handbook Assistant          # what does your system do?
  domain: Employee Q&A grounded in the Meridian Corp employee handbook

use_cases:
  - id: handbook_qa
    description: Employee questions answered accurately from the handbook

    evals:

      # ── RIGHT evals — correctness ────────────────
      - id: answers-from-context
        tag: right                                # diagnostic lens
        labels: [accuracy]                         # analytics grouping
        type: llm
        description: Answer is grounded in the provided handbook excerpt
        pass_criteria: The answer is directly supported by the text provided in the context
        fail_criteria: The answer makes claims not found in the provided context, or contradicts it

      - id: known-answer
        tag: right
        labels: [accuracy]
        type: reference                            # checks the fixture's expected block
        description: Golden fixture expected answer content check

      # ── GOOD evals — quality ─────────────────────
      - id: answer-length
        tag: good
        labels: [format]
        type: rule                                 # your Python, in evals/rules.py
        description: Answer is between 20 and 250 words

      - id: cites-source
        tag: good
        labels: [format]
        type: regex
        description: Response references the source document
        pattern: "(?i)(according to|per the|the handbook (states|says|indicates)|based on the)"
        match: true

      # ── SAFE evals — guardrails ──────────────────
      - id: no-hallucination
        tag: safe
        labels: [grounding]
        type: llm
        description: No claims beyond what the provided context supports
        pass_criteria: Every factual claim in the answer can be traced to the provided handbook excerpt
        fail_criteria: The answer introduces facts, policies, or details not present in the provided context

      - id: stays-in-scope
        tag: safe
        labels: [grounding]
        type: llm
        description: Declines to answer questions outside the handbook scope
        pass_criteria: If the question cannot be answered from the context, the response clearly says so and does not speculate
        fail_criteria: The response attempts to answer a question it doesn't have context for, or speculates about policy

    fixtures:
      directory: fixtures/golden
      sets:
        smoke: [vacation-policy]
        full: all
      runs: 3                                       # N runs per fixture → distributions

defaults:
  provider: anthropic                            # or openai, gemini, openai_compatible
  model: claude-haiku-4-5              # judge model

Four judges. Closed set.

Five cards, four type: values — the scored judge is llm with binary: false, not a fifth type.

rule
Python rule function
Your own deterministic logic. Register with @rule("eval-id") in evals/rules.py. Gets the raw output string, returns Pass/Fail.
No API calls. Fastest. Best for structural checks that need code.
regex
Pattern match
Tests the output against a regex pattern. Set match: true (must contain) or match: false (must not contain).
No API calls. Zero latency. Exact and predictable.
llm
LLM judge
Binary pass/fail via a second model call. You write pass_criteria and fail_criteria as plain English. The judge returns structured JSON with reasoning.
Most flexible. Per-eval model overrides supported — use Haiku for most, Sonnet for subtle judgments.
llm · scored
Scored LLM judge
Set binary: false and the judge returns a number on a scale instead of a verdict. anchors describe what each point means, so the scale measures the same thing run to run.
Reported as mean and stddev, not a pass rate — the distribution is the answer. Any run at the bottom of the scale is flagged as a floor hit.
reference
Reference comparison
Compares output against an expected block in the fixture. Checks contains (required strings) and not_contains (forbidden strings).
No API calls. Ground-truth fixtures only. Skip row shows — when fixture has no expected block.

Your judge is an instrument. Check it.

A failure_rate is a claim about your system, produced by a model. If you cannot say which model produced it, at what temperature, and how often it agrees with you, you do not know what the number measures.

📌
Pinned by default
Judges run at temperature 0.0, not the provider default. Score the same outputs twice, get the same answer twice. If a model refuses the parameter, the run completes and the report header names it.
👁️
The judge sees the question
Judges used to receive the output and nothing else, so a grounding eval asking whether every claim traces to the source was answering without the source. Fixture inputs now go to the judge alongside the output. Set judge_sees_inputs: false on an eval that should read the output alone, or to keep a large retrieved context out of every call.
🔖
Recorded on every run
Provider, model, temperature, seed, per-eval overrides, and a fingerprint over all of it. Swapping the judge and rescoring used to produce a diff indistinguishable from a regression. Runs with different fingerprints are no longer compared automatically.
📈
Rates carry an interval
One failure in five runs is 20%, and so is twenty in a hundred. Binary evals report a Wilson score interval and n, so you can tell them apart and gate CI on the lower bound.
🔁
Repeatability, measured
Set judge_runs: 3 and the report separates system spread from judge spread. Judge spread near zero means the eval is well specified. Judge spread close to system spread means the criteria are ambiguous.
🧑‍⚖️
Agreement with you
Record what you think the correct verdict is, per eval and per run, in the fixture. The report shows how often the judge agreed, counting false passes separately from false fails. Labels do not affect failure_rate.
⚖️
fieldtest calibrate
Run a panel of judges over the same outputs. Pairwise agreement, Cohen's kappa, Fleiss' kappa. Kappa rather than raw agreement, because two judges that both always answer pass agree 100% of the time — and score a kappa of zero, since none of that agreement is beyond chance. Evals rank by disagreement, most contested first, and calibration.kappa_threshold sets the point below which a judge pair is flagged as agreeing no better than chance.
🔌
Any judge you can reach
Anthropic, OpenAI and Gemini by name; vLLM, Ollama, OpenRouter and anything else speaking the OpenAI protocol by base URL, named in config rather than a shell variable so it lands in the fingerprint. For anything else, register an adapter with @provider. The same model served from two endpoints is two instruments, and the diff says so.

Judging the same output twice

Every sample above judges each output once. Set judge_runs above 1 and the report separates two things a single stddev cannot tell apart.

$ fieldtest score --config evals/reference-evals.yaml --set full  (with judge_runs: 2)
### Judge Repeatability (judge_runs: 2)
| eval                        | judge disagreement | system spread | judge spread |
| follows_requested_structure | 0.0%               | —             | —            |
| explanation_clarity         | —                  | 1.7321        | 0.0          |
| caps_applied                | 0.0%               | —             | —            |

The middle row is the whole argument. explanation_clarity is a scored eval, and its two numbers separate two things a single stddev would blur: how much the outputs differ from one another (system spread) and how much the judge wavers when it re-scores the same output (judge spread). Here the first dwarfs the second, which is what tells you the spread is real. Without the repetition you get one stddev and no way to tell which of those produced it. The figures above are one run's; judge_runs compares repeated calls inside a single run, so a judge that is steady here can still answer differently on another day, or waver on one output within a run — a small judge spread is the normal reading, not zero.

Judge spread climbing toward system spread is the signal to rewrite the pass_criteria: the judge is disagreeing with itself, which is a fact about the words, not about your system. Binary evals report disagreement instead — the share of outputs where repeated verdicts differed. It costs runs × judge_runs judge calls, so --dry-run first.

evals/config.yaml
defaults:
  judge_temperature: 0.0          # the instrument, held still
  judge_seed: null                # where the provider supports one
  judge_retry:                     # 5/10/20/40/60/60s by default
    max_attempts: 6
  confidence_level: 0.95         # Wilson interval width, not a model's self-report

providers:
  openai_compatible:
    base_url: https://openrouter.ai/api/v1
    api_key_env: OPENROUTER_API_KEY   # the name, never the key

calibration:
  panel:
    - { provider: anthropic, model: claude-haiku-4-5 }
    - { provider: openai, model: gpt-5 }
    - { provider: openai_compatible, model: meta-llama/llama-3.3-70b-instruct }

Every scored run writes five files.

All outputs land in evals/results/, five files per run sharing a [run-id] prefix. No database. No server. Files you can diff, commit, open in Excel, or drop in a bug report.

📊
*-data.json
Full structured data: every row, every run, full reasoning text, summary stats, delta vs prior run. Machine-readable. CI parses this for gates.
🌐
*-report.html
Self-contained HTML report. Tag health cards, label filter bar, fixture×eval matrix, click-to-expand cell detail with per-run reasoning. Open in any browser.
📝
*-report.md
Markdown report grouped by RIGHT / GOOD / SAFE. Copy-paste into GitHub issues, Notion pages, or Slack. Delta section shows what changed vs last run.
📄
*-data.csv
Flat table: one row per eval×fixture×run. Tag, labels (pipe-separated), type, passed, score, detail, error. Load directly in Excel or Pandas for ad-hoc slicing.
📋
*-report.csv
Spreadsheet-friendly report view: tag health summary and per-eval matrix. Pairs with the markdown report for teams that want CSV over prose.

The full command list is in Every command above — this is the shape of a working day.

day to day
# ── Look before you spend ────────────────────────────
fieldtest validate               # config valid? keys set? how many judge calls?
fieldtest calibrate --dry-run    # what a judge panel would cost, calling nothing

# ── Explore without a key ────────────────────────────
fieldtest demo --example rag --offline  # pre-scored, instant
fieldtest dataset use expense-report    # artifacts to write evals against

# ── Start a real project ─────────────────────────────
fieldtest init                   # scaffold evals/
fieldtest init --template rag    # chatbot, email or rag

# ── Evaluate ─────────────────────────────────────────
fieldtest score                  # score everything in evals/outputs/
fieldtest score --set smoke      # one named fixture set
fieldtest calibrate              # run the judge panel, report agreement

# ── Read and compare ─────────────────────────────────
fieldtest view                   # open the latest HTML report
fieldtest history                # past runs, newest first, with the judge model
fieldtest diff                   # most recent against the one before
fieldtest clean                  # offers to clear outputs/ and prune old results — asks first
Claude Code users: the fieldtest repository includes an /optimize command in .claude/commands/. It is not part of the pip package — clone the repo, or copy that one file into your own project. It scores your outputs, diagnoses failures from the report, edits your prompt or system code, and re-runs — an automated score-diagnose-fix-rescore loop. Type /optimize in Claude Code inside any fieldtest project.

Opinions we hold.

fieldtest is opinionated. These are the constraints that shape every design decision.

Tool measures. Human judges.
fieldtest does not declare pass or fail. It produces distributions — not verdicts. 83% pass rate on RIGHT evals is information. Whether 83% is acceptable is your engineering call, not the tool's.
One eval per failure mode.
Never bundle multiple concerns into one eval. A bundled eval that passes tells you nothing about which failure modes are absent. Narrow scope = interpretable failures.
Structure before measurement.
The config forces you to name your system, think about what matters, and enumerate failure modes before you run anything. Start with one tag and two evals. The structure scales with you — you can't skip the thinking, but you decide how much to think about first.
Files, not infrastructure.
No database, no server, no dashboard. Results are files. You can diff them, commit them, grep them, open them in any browser. Works from a laptop to CI to enterprise.
N runs capture variance.
A single run tells you almost nothing about a probabilistic system. fieldtest runs each fixture N times and shows distributions. 3 runs is the minimum; smoke suites use 1 for speed.
The judge is an instrument too.
Spread across runs has two sources: your system, and the model scoring it. fieldtest pins the judge at temperature 0, records which judge produced every run, and will not automatically compare runs scored by different ones. Set judge_runs to measure how much of the spread is the judge.
Runner decoupled from scoring.
Your runner writes outputs/[fixture-id]/run-N.txt. The scorer reads those files. They share nothing except the directory format. Re-score without re-running when you improve a judge.

From zero to scored report in two minutes.

1
Install
fieldtest is a plain Python package. No containers, no databases, no accounts required.
$ pip install fieldtest
2
Run the demo
See a real scored report before you write any config. Three examples — email, RAG, extraction — each with real failure modes.
$ fieldtest demo --example rag --offline
$ cd fieldtest-demo
$ fieldtest view
3
Scaffold your own project
Run fieldtest init inside any project directory. It creates the evals/ structure and a starter config with inline comments that walk you through every section.
$ fieldtest init --template rag
✓ Scaffolded from rag template at evals/
  evals/config.yaml       — fill in system, domain, tags
  evals/fixtures/golden/  — fixtures with expected outputs
  evals/fixtures/variations/ — fixtures without expected outputs
  evals/.gitignore        — outputs/ excluded from git

Next steps:
  1. Fill in system name and domain in evals/config.yaml
  2. Tag each eval: right, good, or safe
  3. Add fixtures to evals/fixtures/
  4. Run your system → write outputs to evals/outputs/
  5. fieldtest score
4
Run your system. Score. Repeat.
Your runner calls your system and writes outputs. fieldtest score judges them. fieldtest view opens the HTML report. Iterate on prompts, retrieve logic, or constraints — the distribution shows what changed.
$ python evals/runner.py     # you write this — ~30 lines
$ fieldtest score
$ fieldtest view