Skip to content

Agent, workflow and domain-reasoning data

MCP tool-use data: traces, tool schemas and evaluation tasks for Model Context Protocol agents

Quick answer

An MCP tool-use dataset needs four layers: the tool catalog each server advertised (name, description, JSON input schema, annotations), step-level traces of what the model was offered, what it called, with which arguments and what came back, labeled tool-selection decisions across overlapping servers, and evaluation tasks graded on end state. Public sets are mostly synthetic and built on open-source servers, so realistic enterprise catalogs, permission boundaries and failure cases are the gap buyers usually have to fill.

By SourceX Editorial · Updated

This page is protocol-specific. For the general function-calling picture, start with training data for tool use and function calling and the tool-use data glossary entry; for the wider agent data landscape, see the AI agent training data hub.

What the Model Context Protocol fixes, and what it leaves to your data

MCP standardizes the shape of tool descriptions and call messages, which makes traces portable across servers but says nothing about whether the tools or tasks are realistic. A server exposes tools through tools/list, and the client invokes one through tools/call, both as JSON-RPC messages [1]. Each tool carries a unique name, a human-readable description, an inputSchema in JSON Schema, and optionally an outputSchema and behavioral annotations [1][2].

Results come back as content blocks, optionally with structured content that should validate against the declared output schema, and tool-level failures are signaled with isError: true rather than a protocol error [1]. That distinction matters for training: an isError result is something the model sees and must recover from, while a JSON-RPC error usually never reaches the model.

Pin the spec revision in your data contract. The 2025-06-18 and 2025-11-25 revisions differ in field details, and the newer one requires inputSchema to be a valid JSON Schema object [1][2]. A trace corpus that mixes revisions without a protocol_version field will teach the model inconsistent schemas.

The four data layers an MCP agent learns from

Each layer trains a different behavior, so specify them separately rather than buying "MCP traces" as one undifferentiated blob.

  • Tool catalogs. Snapshots of tools/list per server and per session, including descriptions as written by the server author. Messy, overlapping and underspecified descriptions are the realistic case, and they are what tool selection must handle.
  • Call traces. Ordered steps: user turn, tools in context, model reasoning or plan (if retained), tools/call request, result, and the model's next action. These train argument filling and multi-step chaining.
  • Selection labels. For each step, which tool was correct, which near-miss tools were also offered, and why. These train retrieval and disambiguation when dozens or hundreds of tools are loaded.
  • Evaluation tasks. Held-out prompts with fixed server state and a grader. These measure, rather than train, and must never leak into training.

Operational records from real systems can be converted into the second layer; see API call logs and integration run histories as tool-use training data for that conversion and its limits.

What public MCP datasets cover as of October 2026

Public MCP data is dominated by synthetic trajectories over open-source servers, with a small number of human-curated evaluation sets. TOUCAN generated about 1.5M synthetic tool-agent trajectories from 495 real MCP servers, most of which expose fewer than 10 tools [4]. MCP-Flow builds an automated pipeline for collecting servers and synthesizing training data, and it notes that earlier work relied on a few hand-curated servers [5].

On the evaluation side, Scale AI's MCP-Atlas publishes a 500-task subset under CC-BY-4.0 from a benchmark spanning 36 servers and 220 tools, scoring answers against claims in reference responses [3]. HumanMCP targets tool retrieval with human-like queries rather than templated ones [6]. A separate corpus catalogs MCP implementations on GitHub, which is useful for building catalogs but reflects public, open-source servers [7].

The pattern for buyers: scale is synthetic, servers are public, and catalogs are small per server. What is scarce is enterprise-shaped data, meaning internal systems with dozens of similar tools, role-scoped permissions, inconsistent naming, and real user requests that do not map cleanly to one tool.

Data needPublic coverage (as of October 2026)Typical gap
Synthetic multi-turn trajectoriesLarge (for example TOUCAN [4], MCP-Flow [5])Generated by models, so failure modes reflect the generator
Human-like retrieval queriesPresent (HumanMCP [6])Few enterprise tool names or internal jargon
Graded multi-server tasksPresent (MCP-Atlas subset [3])Public servers; 500 public tasks are easy to contaminate
Enterprise tool catalogs with permissionsSparseInternal systems are rarely published
Real failure and recovery tracesSparseLogs are private and contain personal data
Injection and poisoning cases over MCPPartial, mostly generic agent benchmarks [8]Few cases using MCP descriptions as the attack surface

Tool selection across overlapping servers is its own training target

Tool selection fails differently from argument filling, so it needs labels that record the choice set, not just the chosen call. When a client loads a CRM server, a ticketing server and a data warehouse server, the model may see three tools that can all plausibly "look up a customer." A trace that stores only the winning call cannot teach the model why the others were wrong.

Ask for per-step tools_offered lists with server namespace, the gold tool, labeled hard negatives, and a short rationale. Cover the cases that break production agents: two tools with near-identical descriptions, a read-only tool versus a write tool for the same object (MCP annotations such as read-only and destructive hints help here [1]), a tool the user is not permitted to call, and requests that need no tool at all.

Measure selection separately from task success. Report top-1 accuracy over the offered set, false-call rate on no-tool requests, and accuracy as the number of loaded tools grows; human-like retrieval queries such as HumanMCP's are one way to test selection under realistic phrasing [6].

A step-level record schema for MCP traces

A useful MCP trace stores the offered catalog and the outcome at every step, not just the final transcript. The schema below is a starting point for a data specification; extend it with the fields your training recipe needs, and compare it against the broader guidance in writing an agent data specification.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "trace_id": "trc_000184",
  "protocol_version": "2025-11-25",
  "session": {
    "servers": [
      {"server_id": "crm", "tool_count": 41, "catalog_snapshot_id": "cat_crm_v7"},
      {"server_id": "ticketing", "tool_count": 23, "catalog_snapshot_id": "cat_tkt_v3"}
    ],
    "user_role": "support_tier2",
    "permissions_scope": ["crm:read", "ticketing:read", "ticketing:write"]
  },
  "steps": [
    {
      "step": 3,
      "tools_offered": ["crm.get_account", "crm.search_contacts", "ticketing.get_requester"],
      "tool_chosen": "crm.get_account",
      "gold_tool": "crm.get_account",
      "hard_negatives": ["ticketing.get_requester"],
      "selection_rationale": "account-level SLA needed, not requester profile",
      "arguments": {"account_id": "[ACCOUNT_ID_1]"},
      "arguments_valid_against_inputSchema": true,
      "result": {"isError": false, "structuredContent_valid": true},
      "next_action": "call ticketing.update_priority",
      "latency_ms": 412
    }
  ],
  "outcome": {"task_success": true, "grader": "state_check", "end_state_diff_id": "diff_000184"},
  "redaction": {"method": "tokenized placeholders", "sample_checked": true}
}

If the source system already emits OpenTelemetry, map these fields from GenAI tool-execution spans; note that tool arguments and results are opt-in attributes there, so many production logs will not contain them unless capture was switched on [10]. Record that as a coverage limit in the dataset documentation rather than discovering it after delivery.

Failure, error and recovery cases you should require

Successful traces alone produce an agent that assumes tools work. Real MCP sessions include isError results from bad arguments, expired credentials, rate limits, pagination limits, timeouts and schema drift when a server changes its tool list mid-session [1].

Specify a minimum share of steps where the first call fails, and label the recovery: retry with corrected arguments, switch tools, ask the user, or stop. The detail on sourcing these cases is covered in tool-call errors and recoveries. Keep the original error text; normalizing it away removes the signal the model needs.

Evaluation tasks for MCP agents need fixed state and adversarial cases

MCP evaluation tasks should be graded on the end state of the servers, not on transcript similarity. tau-bench's approach of simulated users, programmatic APIs, written domain policies and database-state grading transfers well to MCP servers, and its pass^k metric exposes agents that succeed only sometimes [9]. Claim-based scoring, as used in MCP-Atlas, suits information-gathering tasks where there is no state change [3].

Include security cases. Tool results and tool descriptions are both untrusted text the model reads, so a malicious server or a poisoned record can carry instructions. AgentDojo pairs user tasks with injection tasks delivered through tool outputs and reports both utility and attack success [8]; build MCP-specific variants in which the payload sits in a tool description, a result field, or a resource returned by a second server. The dedicated page on prompt injection evaluation sets for tool-using agents covers construction and grading.

Keep evaluation sets private and versioned. Public task sets such as the 500-task MCP-Atlas subset are easy to train on accidentally; held-out tasks built on licensed enterprise catalogs are harder to contaminate. For grader design across agent types, see agent evaluation task suites.

Buyer checklist for an MCP tool-use dataset

Use this checklist when writing a request or reviewing a supplier's sample. It assumes you will train or evaluate on the data, not host the supplier's servers.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat to ask forRed flag
Spec revisionprotocol_version on every trace [1][2]Mixed revisions, no version field
Catalog snapshotstools/list output per session, original descriptionsDescriptions rewritten for clarity
Selection labelsOffered set, gold tool, hard negativesOnly the executed call stored
Argument validityValidation result against inputSchemaNo schema stored with the trace
Error coverageShare of steps with isError and labeled recoverySuccess-only traces
PermissionsRole and scope per sessionEvery tool callable by every user
OriginHuman, production log, or synthetic, per traceSynthetic presented as real usage
Security casesInjection and poisoned-description tasks [8]None, or only generic jailbreaks
Personal dataRedaction method and sample checkRaw customer IDs in arguments
RightsLicense covering training, evaluation and derived tasksLicense silent on derived benchmarks

Rights questions are specific for agent data: whether you can replay traces against environments, derive new tasks, or publish benchmark results. License terms for agent data walks through those clauses.

Where operational company data fits

Enterprise-shaped MCP data usually starts as something else: support and sales histories, engineering records, finance and legal workflow logs, and the documents those workflows touch. Those records show which systems people actually queried, in what order, and where they got stuck, which is the raw material for realistic catalogs, selection labels and held-out tasks.

SourceX sources operational datasets like these from US companies on request; it does not hold them in stock, and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and every release is approved by the supplying company. SourceX does not train models and does not source scraped web content. You can describe the traces, catalogs or task sources you need on the SourceX buyer page.

Request MCP tool-use source data from SourceX

SourceX finds US companies that hold the operational records you describe, assesses the data and licensing permissions, and manages the license that defines records, uses, term and delivery. Nothing is contracted until a supplier agrees, and terms are agreed per deal. Describe your MCP tool-use data requirements at https://sourcex.si/buyers.

Frequently asked questions

Is synthetic MCP data enough to train tool selection?

It is enough to teach protocol format and basic chaining, and public sets show it can be generated at large scale [4][5]. It is weaker for selection among many near-duplicate internal tools, because generators tend to write clean, distinct descriptions. Mix synthetic volume with a smaller set grounded in real catalogs and real requests.

How is MCP tool-use data different from function-calling data?

Function-calling data usually fixes one tool list per prompt. MCP data adds server boundaries, dynamic tool lists that can change mid-session, structured results and isError semantics, and the need to choose across servers [1].

Should evaluation tasks include servers the model never saw in training?

Yes. Hold out entire servers, not just tasks, so you measure generalization to new catalogs. Report scores separately for seen and unseen servers.

Sources

  1. Model Context Protocol, "Specification 2025-11-25: Server features, Tools" (2025). https://modelcontextprotocol.io/specification/2025-11-25/server/tools
  2. Model Context Protocol, "Specification 2025-06-18: Server features, Tools" (2025). https://modelcontextprotocol.io/specification/2025-06-18/server/tools
  3. Scale AI on Hugging Face, "MCP-Atlas dataset card (README)". https://huggingface.co/datasets/ScaleAI/MCP-Atlas/blob/main/README.md
  4. arXiv (2510.01179), "TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments" (2025). https://arxiv.org/pdf/2510.01179
  5. arXiv (2510.24284), "MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools" (2025). https://arxiv.org/pdf/2510.24284
  6. arXiv (2602.23367), "HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance" (2026). https://arxiv.org/pdf/2602.23367
  7. arXiv (2607.10123), "A Large-Scale Dataset of MCP Implementations on GitHub" (2026). https://arxiv.org/pdf/2607.10123
  8. arXiv (2406.13352), "AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents" (2024). https://arxiv.org/abs/2406.13352v3
  9. arXiv (2406.12045), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  10. OpenTelemetry, "Inside the LLM Call: GenAI Observability with OpenTelemetry" (2026). https://opentelemetry.io/blog/2026/genai-observability/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data