Skip to content

Evaluation and benchmarking datasets

Function-calling evaluation datasets: schemas, calls and expected arguments

Quick answer

A function-calling evaluation dataset is a held-out set of items. Each item pairs a tool catalog (names, descriptions, JSON Schema parameters) and a user request with a gold answer: the call or calls a correct model makes, their arguments, and any calls it must not make. Scoring checks tool selection, argument values and ordering, and sometimes the end state of a backing system. The hard part is getting realistic schemas and requests from real APIs, not writing the grader.

By SourceX Editorial · Updated

Use this page to specify call-level items. For full multi-step task suites graded on environment state, see agent evaluation task suites. For training data in the same domain, see function-calling fine-tuning data formats and the SourceX capability page on training data for tool use and function calling.

What a function-calling eval item must contain

A usable item contains five things: the tool catalog shown to the model, the conversation up to the decision point, the expected calls, the argument-matching rules, and the calls that would be wrong. Missing any one turns a test into a vibe check. The catalog should use the same representation your runtime uses. With Model Context Protocol servers, each tool has a name, a description, an inputSchema written in JSON Schema, and optionally an outputSchema [2]. Arguments can then be validated mechanically against JSON Schema 2020-12 keywords such as type, required and enum [3].

The gold answer is more than one string. Many requests have several acceptable argument forms: a date as 2026-10-09 or a relative "today" resolved by the harness, an optional limit left out or set to its default. Public leaderboards such as the Berkeley Function Calling Leaderboard (BFCL) [7] handle this by parsing calls into an abstract syntax tree and matching each argument against a list of allowed values. Your item format should carry those allowed-value lists explicitly, not leave them in a grader's head.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "item_id": "fc-ops-0142",
  "slice": "single_call/ambiguous_entity",
  "tools": [
    {"name": "search_tickets", "inputSchema": {"type": "object",
      "properties": {"status": {"enum": ["open", "pending", "closed"]},
                     "account_id": {"type": "string"},
                     "created_after": {"type": "string", "format": "date"}},
      "required": ["account_id"]}},
    {"name": "close_ticket", "inputSchema": {"type": "object",
      "properties": {"ticket_id": {"type": "string"}, "reason": {"type": "string"}},
      "required": ["ticket_id", "reason"]}}
  ],
  "messages": [{"role": "user", "content": "Show me Acme's open tickets since Monday."}],
  "context": {"today": "2026-10-09", "account_lookup": {"Acme": "ACC-7781"}},
  "expected_calls": [
    {"name": "search_tickets",
     "arguments": {"account_id": ["ACC-7781"], "status": ["open"],
                   "created_after": ["2026-10-05"]}}
  ],
  "forbidden_calls": ["close_ticket"],
  "accept_clarifying_question": false,
  "scoring": {"order": "strict", "extra_args": "fail", "string_match": "exact"},
  "provenance": {"schema_source": "vendor_api_v3", "request_source": "rewritten_support_ticket"}
}

Single-call, parallel, multi-turn and stateful items

The four item types test different failure modes and should be counted as separate slices. Reporting one blended accuracy hides which one your model fails. Public suites already separate serial and parallel calls, abstention (no tool fits), and stateful multi-step settings, and practitioner experience suggests single-turn accuracy is often the most direct of these to raise.

  • Single call: one request, one correct call. Tests tool selection among near-duplicates and argument extraction.
  • Parallel or multiple: one request needs several independent calls, such as fetching three account balances. Score as an unordered set.
  • No-call or abstain: no tool in the catalog fits, or a required argument is missing and the right move is a clarifying question.
  • Stateful sequence: later calls depend on earlier results, as in tau-bench, where agents use programmatic APIs that change database state under domain policies [1]. These are graded by comparing end state, which pushes them toward full task suites.

A practical rule: if an item needs a live or simulated backend to grade, it belongs in a task suite. If a static gold call list grades it, it belongs in the call-level set.

Argument-level scoring rules that survive review

Argument scoring should be decided per parameter type before items are written, because graders who improvise rules produce inconsistent gold labels. Exact string match over-penalizes harmless variation, while LLM-as-judge grading drifts. The middle path is typed matching with per-item allowed-value lists.

Parameter typeRecommended ruleCommon failure it catches
Enum (status, currency)Exact match against schema enum [3]Invented values such as "in_progress"
Identifier (account_id, ticket_id)Exact match; never fuzzyHallucinated or reformatted IDs
Date and timeNormalize to ISO 8601 with a fixed today in item contextWrong relative-date resolution, time zone drift
Free text (reason, query)Allowed-value list or a narrow rubricOver-long or policy-violating text
NumericExact, or tolerance stated in the itemUnit errors (cents vs dollars)
Optional parametersAbsent or equal to default both passPenalizing correct omissions
Unknown parametersFail when additionalProperties is falseArguments the API will reject

Report at least four numbers per slice: schema-valid rate (does the call parse and validate), tool-selection accuracy, full-argument accuracy, and forbidden-call rate. Structured output evaluation is the first of these in isolation: a call that fails JSON Schema validation never reaches the API.

Forbidden calls and safety constraints

Every item that sits near a destructive or irreversible tool should list the calls that must not happen, and the forbidden-call rate should be tracked as a release gate. Practitioner guidance on building golden datasets from production failures treats invalid JSON output and forbidden tool calls as failures worth turning into regression items [4]. A model that reaches 95% argument accuracy but calls close_ticket when asked to "look at" a ticket is not shippable.

Write negative items deliberately. Examples include a read request beside a write tool with a similar name, a refund request that exceeds a policy limit, and a request that requires confirmation before a delete_* call. In MCP terms, also include items where the tool result returns isError and the correct next step is to report or retry with corrected arguments rather than repeat the same call [2]. Policy-heavy domains such as support are covered further in policy-following evaluation for customer support agents.

Where realistic schemas and requests come from

The most valuable inputs are real API schemas and real user requests, because synthetic catalogs are too clean. Real enterprise APIs have overlapping endpoints (get_invoice, get_invoice_v2), vague descriptions, deeply nested objects, and parameters whose names do not match how users speak. Benchmarks such as CRMArena-Pro go to considerable lengths to make synthetic CRM data behave realistically [5], which shows how much realism matters for grading.

Useful source material includes OpenAPI specifications for internal and SaaS APIs, MCP server manifests, support tickets and sales emails that map to CRM or ERP actions, and engineering runbooks that describe which command fixes which alert. Requests usually need rewriting to strip personal data and to pin a decision point. The gold call is typically written by someone who knows the system, from what the human operator actually did. If you are building from your own logs, see building a golden evaluation dataset from business records.

For teams that need this material from outside, SourceX sources operational datasets from US companies on request, including support and sales histories, engineering records, and finance and legal workflows. Categories are not inventory, and a request does not guarantee a match. You can describe the schemas and requests you need on the SourceX buyers page.

Buyer checklist for a function-calling eval set

A delivered set is acceptable when its items are graded, sliced, traceable and private. Use this checklist in the specification and again at acceptance; the broader method is in writing an evaluation dataset specification.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Tool catalog per item, in the same format as production (MCP inputSchema or OpenAPI-derived JSON Schema) [2][3]
  • Slice counts stated for single, parallel, abstain, clarifying-question and stateful items
  • Distractor tools present: near-duplicate names, deprecated versions, write tools beside read tools
  • Allowed-value lists per argument, with date anchors fixed in item context
  • forbidden_calls populated for every item near a destructive tool [4]
  • Provenance per item: schema source, request source, who wrote the gold call
  • Double-annotated gold calls on a sample with an adjudication log; an audit of 10 widely used test sets estimated an average label error rate of at least 3.3% [6]
  • Personal data removed from requests and argument values, with the method recorded
  • License terms covering evaluation use, result publication and retention; see evaluation-only data license terms
  • Access controls and canary items to detect leakage; see keeping a private eval set private

Public tool-calling benchmarks vs a private call-level set

Public tool-calling benchmarks are good for comparing base models; a private set built on your own APIs is what predicts production behavior. BFCL [7] and tau-bench [1] define useful categories and grading ideas, but their functions are not your functions, and public items can end up in training corpora. Use public results to shortlist models. Then gate releases on a private set whose catalogs mirror your tool registry and whose items you refresh when schemas change. The trade-offs are covered in private evaluation sets vs public benchmarks and contamination-resistant evaluation design. The full map of eval data types is on the LLM evaluation datasets hub.

Version the set with the API. When search_tickets gains a required parameter, items using the old schema must be migrated or retired, or the eval will penalize correct behavior. Keep a schema hash on each item so stale items are flagged automatically.

Sourcing real API schemas and requests for tool-use evaluation

If your eval set needs real operational requests and the systems behind them, SourceX looks for US businesses that hold the data you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. Tell SourceX what your function-calling eval set needs.

Sources

  1. Sierra Research (arXiv:2406.12045), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  2. Model Context Protocol, "Model Context Protocol specification: Tools (2025-11-25)" (2025). https://modelcontextprotocol.io/specification/2025-11-25/server/tools
  3. JSON Schema, "JSON Schema specification". https://json-schema.org/specification
  4. OneUptime, "How to Build a Golden Evaluation Dataset from Real LLM Production Failures" (2026). https://oneuptime.com/blog/post/2026-08-31-build-golden-evaluation-dataset-production-failures/markdown
  5. Salesforce AI Research (arXiv:2505.18878), "CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions" (2025). https://arxiv.org/pdf/2505.18878
  6. Northcutt, Athalye, Mueller (arXiv:2103.14749), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  7. University of California, Berkeley (PMLR 267, 2025), "The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models" (2025). https://gorilla.cs.berkeley.edu/blogs/8_berkeley_function_calling_leaderboard

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data