Agent Surface

Testing

Datasets and Experiments

Vendor-neutral golden-dataset curation - sourcing, labeling, hygiene, versioning, and drift detection so experiments measure something real

Last verified 2026-09-25

Summary

An is only as good as the dataset behind it. Teams build harnesses, scorers, and CI gates first - and never build the dataset. The result is a green pipeline measuring nothing: a suite that passes because its cases are trivial, stale, or leaked from the same prompts the agent already memorized. This page is the dataset-first counterpart to the platform pages: how to source, label, clean, version, and refresh the golden set that every experiment is pinned to. The tool mechanics (running experiments, diffing results, gating merges) live in the platform and CI pages and are linked rather than restated.

  • Decision rule: curated golden vs. production sampling vs. synthetic generation - matched to decision stakes
  • Sourcing: production as a candidate pool, failure mining, reviewed synthetic edge cases
  • Labeling: rubric-anchored labels, defined labelers, adjudication for disagreement
  • Hygiene: dedup, leakage checks against training/few-shot content, PII scrubbing before storage
  • Versioning: the dataset is a versioned artifact; every experiment is pinned to a version
  • Drift: detect when production moves away from the dataset; re-sample on a schedule

Eval quality is capped by dataset quality. A scorer can be perfect and a harness can be fast, but if the cases do not represent the tasks users actually send - or if they leak answers the agent has already seen - the score is a number with no meaning. The first failure mode of agent evaluation is a dataset that was assembled once, by hand, from whatever examples were nearby, and then frozen while production drifted away from it.

A passing eval suite is evidence that the agent handles the cases in the dataset. Whether that means the agent works depends on the dataset being a faithful, current, leak-free sample of real work - a property you build and maintain, not one you get for free.

Treat the dataset as the primary artifact and the harness as plumbing. The rest of this page is the lifecycle of that artifact.

Decision rule: where cases come from

Three sourcing strategies trade off cost, realism, and coverage. Pick per dataset by what the eval decides.

StrategyUse whenCost per caseRisk
Curated goldenHigh-stakes ship/no-ship gates where every label must be trustedHigh (expert labeled)Small sets miss the long tail
Production samplingYou need distributional realism and volume for comparisonsLow to source, medium to labelReflects current traffic only; needs labeling
Synthetic generationA failure mode is rare or not yet in production (new capability, adversarial edge)Low to generate, high to reviewUnreviewed synthetic cases encode the generator's blind spots

Most mature suites blend all three: a small curated core that gates releases, a larger sampled body that measures regressions distributionally, and a synthetic edge-case layer that probes failure modes production has not yet produced. Synthetic cases never enter the set without human review - an unreviewed generator quietly teaches the eval to accept its own mistakes.

Match dataset size to the stakes of the decision the eval gates. A ship gate needs enough cases that a single flaky run does not flip the verdict (see the statistical-significance treatment in CI integration); an exploratory capability probe can be a few dozen hand-authored frontier tasks.

Decision the eval makesWorking sizeSourcing bias
Ship/no-ship on a core user flow100–300 casesReal production failures plus hand-picked edge cases
Model or prompt comparison50–200 casesStratified production sample across scenario types
Capability probe / exploration20–50 casesFrontier tasks authored by domain experts

Sourcing: production traces to candidate pool

Production is the richest source of realistic cases, but raw traces are not a dataset - they are a candidate pool that has to be mined, labeled, and cleaned.

  1. Sample the trace pool. Pull traces across the distribution you care about (by scenario, tool path, user segment), not just recent or convenient ones. Stratify so rare-but-important scenarios are not drowned out by the common case.
  2. Mine failures first. The highest-value cases are tasks that already broke. Cluster low-scoring or thumbs-down production runs, deduplicate them into distinct failure modes, and promote representatives into a regression body. This mirrors the "start with real failures" discipline in the evaluation framework - the dataset side of the same idea.
  3. Generate the edges you are missing. For failure modes production has not yet exercised (a new tool, an adversarial input class), synthesize candidates, then route every one through human review before it counts as a labeled case.

The output of sourcing is a pool of inputs with candidate expected outcomes - not yet trustworthy labels. Labeling turns the pool into a dataset.

Labeling workflows

A case is input → expected outcome. The expected outcome is a judgment, and judgments need a defined process or they will not be reproducible.

  • Anchor labels to a rubric. Label against the same success criteria the scorer uses, not a labeler's private intuition. A rubric-anchored label ("status == sent and no draft left behind") survives re-labeling.
  • Name who labels. Domain experts for high-stakes curated sets; a calibrated model-based labeler for volume, spot-checked by humans. Record the labeler on the case so a suspect label can be traced to its source.
  • Adjudicate disagreement, do not average it. When two labelers disagree, the disagreement is signal: the case is ambiguous or the rubric is underspecified. Route it to an adjudicator who either resolves it with a rubric clarification or drops the case as genuinely undecidable. Averaging conflicting labels manufactures a fake consensus that will confuse every future comparison.

Calibrating a model-based labeler against human judgment is the same problem as calibrating an judge - see LLM-as-judge for the bias and agreement mechanics.

Dataset hygiene

Before a case is stored, it passes three checks. Skipping them is how a dataset silently rots.

  • Deduplicate. Near-duplicate cases inflate the score of whatever the agent happens to do well on the repeated pattern and starve the tail. Cluster by input similarity and keep one representative per cluster.
  • Check for leakage. A case whose input or expected output appears in the model's few-shot prompt, system prompt, or data measures memorization, not capability. Diff dataset content against training and prompt content and quarantine overlaps. This is the leak that makes a suite pass while the agent fails on anything genuinely new.
  • Scrub PII before storage. Production traces carry real user data. Redact or tokenize personal data at ingestion, before it lands in a versioned store where it is hard to fully expunge later. Curation is also a data-governance surface, not only a quality one.

Versioning and provenance

The dataset is a versioned artifact. Every experiment records the exact dataset version it ran against, and that pin is what makes two experiments comparable at all.

Without pinning, a score change is unattributable: did the agent improve, or did someone add three easy cases to the dataset between runs? With pinning, the causes separate cleanly - a run against a fixed version isolates model/prompt changes, and a deliberate dataset bump is its own reviewable event with its own (which traces were added, who labeled them, why).

Practical shape, regardless of platform:

  • The dataset has an immutable version identifier; edits create a new version rather than mutating the old one.
  • Each case carries provenance: its source (which production incident, which generator), its labeler, and its label timestamp.
  • Every experiment stores the dataset version alongside the model, prompt, and code revision, so a historical result can be reproduced exactly.

For the platform-specific mechanics - pulling a pinned dataset version, running an experiment against it, and reading the diff against a baseline - see the Braintrust worked example, and CI integration for pinning inside a pipeline.

Drift detection

A dataset is a snapshot; production is a moving distribution. Over time the two diverge - new user intents, new tool paths, a shifted input mix - and the eval keeps grading yesterday's traffic. Undetected, this is the slow version of the frozen-dataset failure: the suite stays green while its relevance decays.

  • Compare distributions on a schedule. Periodically measure whether the live input distribution (scenario mix, input length, tool-path frequency) still resembles the dataset. A growing gap is the signal to re-sample.
  • Re-sample, then re-label. Refresh the sampled body from recent production on a cadence tied to how fast your product changes - and route new cases through the same labeling and hygiene gates. A re-sample that skips labeling just imports unlabeled noise.
  • Version the refresh. A re-sampled dataset is a new version, and comparisons that cross the boundary are dataset changes, not agent changes. Keep them separate.

Gates and verdicts

The dataset produces the score; the CI integration page owns what the score does to a merge - baseline comparison, the delta threshold that blocks, and the statistical tests that separate a real regression from run-to-run noise. Two dataset-side obligations make those gates trustworthy:

  • Pin the dataset version in the gate. A gate comparing runs on different dataset versions is comparing nothing. The pipeline must fail on an unpinned or mismatched version before it interprets any delta.
  • Keep the gating core stable. The cases that decide ship/no-ship should change deliberately and rarely. Churning the gate dataset every week makes its verdicts incomparable across time.

Platform landscape

Dataset-and-experiment primitives are converging across platforms: Braintrust, LangSmith, Arize Phoenix, and Weights & Biases Weave all implement versioned datasets, experiments pinned to a dataset version, and diff views comparing runs. They differ in hosting model, trace integration, and human-review tooling, but the curation methodology above is the part none of them can do for you - a managed platform stores and versions the dataset; it does not decide which cases belong in it. This guide uses Braintrust as its worked example; the sourcing, labeling, and hygiene practices port to any of them.

Checklist

  • Each dataset's sourcing strategy is chosen from the decision the eval makes
  • The regression body starts from mined production failures, not hand-invented cases
  • Synthetic cases pass human review before entering the set
  • Labels are anchored to the scorer's rubric and record their labeler
  • Labeler disagreements are adjudicated, not averaged
  • Cases are deduplicated, leakage-checked, and PII-scrubbed before storage
  • The dataset is versioned and every experiment pins a version
  • Production/dataset distribution drift is measured on a schedule and triggers re-sampling

See also

  • /docs/testing/evaluation-framework - capability vs. regression evals; building the first eval set
  • /docs/testing/braintrust - worked example of versioned datasets and pinned experiments
  • /docs/testing/ci-integration - baseline comparison, regression gating, statistical significance
  • /docs/testing/llm-as-judge - calibrating model-based labelers against human judgment
  • /docs/testing/metrics - , tool F1, and the scores a dataset feeds

On this page