A test with yesterday’s answer key
Imagine asking an analytics agent: “How many shoppers bought something last week?” If new events arrive or “last week” advances, yesterday’s correct number is no longer the answer.
Traditional evaluations compare the agent to a frozen reference string. That assumes the correct answer is a constant. But for a live warehouse the answer is a function of time, data, and schema. A perfectly correct agent can be marked wrong by an outdated key; a stale agent can be rewarded for repeating it.
Frozen reference
“Last week: 120 shoppers.” Easy to write, but can expire as the week or warehouse changes.
Executable reference
Run a Python function at grading time; obtain the answer from the relevant data system.
For dynamic data, ground truth is better represented as an algorithm than as an answer. That solves freshness only if the algorithm itself is correct and sees an appropriate data snapshot.
Compute, atomize, compare.
The authors implement an evaluator as a harness skill: a bundle of instructions, scripts and connectors that can run inside an agent environment. Each test case is described by a query U, a reference function G, environment E, and compute settings C (such as endpoints and credentials).
Why split into atomic facts?
An agent might say “Retail had 120; healthcare had 80” or display the same numbers in a table. Comparing characters penalizes the format. Instead, extract claims such as retail_count = 120 and healthcare_count = 80, then compare the claims. The reference extractor sees the user query so it retains only relevant facts; the response extractor is not query-conditioned so irrelevant claims are still counted.
Format independence is the intended property of the factoid layer, not a proof that an LLM extractor will always produce identical claims from every format. Extraction and semantic matching can themselves make mistakes.
Paper Figure 1 shows this flow. The diagram above is an original simplified redraw. The paper’s reference and agent are run in parallel against the live system; in a real deployment, a changing data state can still cause a race unless both read a consistent snapshot.
What does “correct” mean?
Let Fᵣ be the set of claims in the response, Fᵍ the query-relevant reference claims and M the number of one-to-one matched claims. An unmatched response claim is an unsupported or wrong claim; an unmatched reference claim is an omission.
Accuracy = M / (|Fᵣ| + |Fᵍ| − M)The paper calls the final expression “accuracy”; mathematically it is intersection-over-union / Jaccard overlap of the matched fact sets. Scores are multiplied by 9 for reporting.
“Can I trust what you said?”
Four claims, three matched → 3/4 = 75%. Adding an invented claim decreases precision.
“Did you cover what mattered?”
Five expected claims, three matched → 3/5 = 60%. Omitting two needed facts decreases recall.
Edge case: the paper does not specify empty-set conventions in Eq. 2; the interactive toy below returns 0 if a denominator is zero. A production evaluator should explicitly define these cases.
One actual failure, unpacked.
The appendix walks through a financial-services request: find datasets and define a model that flags credit-card accounts at early risk of default. The reference expects the Financial 6M bundle, a billing.paymentFailed event in a 30-day outcome window, and features from the preceding 90 days. It also excludes lastPaymentFailedAt as leakage.
Financial 6M datasets · event-based target · 90-day observation / 30-day outcome · leakage exclusion.
Base Financial datasets · utilization heuristic target · some valid profile features · no leakage warning.
Approximately 5 of 14 response facts matched approximately 12 expected facts: precision ≈ 0.36, recall ≈ 0.42, overlap ≈ 0.24. On the paper’s 0–9 scale the reported values are precision 3.2, recall 3.8, accuracy 2.1 (rounded fact counts and judge output explain small differences).
Does it agree with experts?
The authors built a synthetic, production-shaped stage database spanning retail, finance and healthcare, including decoy datasets. The three principal 26-week bundles contain 353,330 event records, 30,000 profiles, and 82 catalog entries in total (Table 1). The judge study used 55 single-turn cases; 53 were comparable for PASS/FAIL analysis. Three experts labeled responses. The skill and judge both used Claude Sonnet 4.6 in Claude Code.
Annotator agreement, using prevalence-adjusted Gwet’s AC2, was 0.931 at factoid level and 0.883 across stacked dimensions (Table 2). Here is the paper’s Table 3, against the three-annotator human reference:
| Evaluator setup | MCC | Precision | Recall | F1 | Tokens / case |
|---|---|---|---|---|---|
| No explicit ground truth | −0.379 | 0.238 | 0.200 | 0.217 | ~70k |
| Natural-language ground truth | 0.331 | 0.561 | 0.920 | 0.697 | 29.2k |
| Ground truth as code | 0.427 | 0.568 | 1.000 | 0.724 | ~24.6k |
MCC is a correlation between predicted PASS/FAIL and human PASS/FAIL; +1 is perfect agreement, 0 no correlation, −1 opposite predictions. Unlike raw agreement, it considers all four confusion-matrix cells: TP, TN, FP, FN. The improvement is 0.096 absolute MCC, or ~29% relative to 0.331. Token use drops from 29.2k to ~24.6k: ~16% fewer. These are results for the paper’s setup, not guarantees on arbitrary agents.
Break an answer key yourself.
This local simulation uses a tiny made-up event count. Move the evaluation day: watch a frozen answer become stale while an executable reference recomputes the count. Then edit the agent’s claims and see precision and recall move.
function reference(day) {
const week = Math.floor(day / 3); // toy: 3-day “week”
return records.filter(x => x.week === week).length;
}Toy time is deliberately compressed: a “week” lasts three slider days so the change is visible. The code above illustrates the pattern; the JS below actually recomputes from the local records array.
Question: “For the current toy week, give the shopper count and the top segment.” Expected facts below are derived from local toy data. Choose what the agent says. Unchecked = not said.
This demonstration uses exact semantic IDs rather than an LLM: “12 shoppers” and “twelve shoppers” would need normalization to count as the same fact in a real system. The paper uses LLM extraction and one-to-one semantic matching.
# Illustrative Python sketch; not the authors' code.
def ground_truth(api, query, snapshot):
rows = api.events(snapshot=snapshot)
return {"shopper_count": count_unique(rows, query.week),
"top_segment": top_segment(rows, query.week)}
response = agent.answer(query, snapshot=snapshot)
expected = ground_truth(api, query, snapshot=snapshot)
ref_facts = extract_facts(expected, relevant_to=query)
answer_facts = extract_facts(response) # do not filter extras
matches = one_to_one_semantic_match(answer_facts, ref_facts)
precision = matches / len(answer_facts)
recall = matches / len(ref_facts)
overlap = matches / (len(answer_facts) + len(ref_facts) - matches)Use a consistent snapshot where possible. A real evaluator must secure reference execution, handle failures/empty sets, review code correctness, and validate the fact extractor.
Fresh does not mean infallible.
Reference bugs
Missing or wrong branches in the ground-truth function produce confident but wrong scores. Expert authoring and review remain necessary.
Same-model judge
The same model family powers the skill and the evaluator. The authors flag possible self-preference; cross-model replication is needed.
Annotation cost
Writing and maintaining executable reference functions takes more expert effort than writing text references.
Scope and transfer
55 cases from one in-house skill; no reference APIs means this approach may not apply. Token savings may depend on harness caching.
When an agent answers questions about a changing data system, recompute the expected answer from that system at test time, then check the answer claim by claim—but keep auditing the reference computation and the judge.
Read the original.
Primary source: Tamhane et al., Skill-based Agentic Evaluation for Real-time Data Science Tasks, arXiv:2609.16487v1 (2026). This guide was prepared from the 11-page PDF, especially §§3–5, Tables 1–3 and Appendix A. The diagrams, made-up shopper example and interactive simulation here are original teaching aids, not reported experiments.