Skip to content

Fine-tuning and post-training data

Function-calling fine-tuning data: schemas, calls, results and no-call cases

Quick answer

A function-calling SFT set is a JSONL file in which each line holds the tool schemas available for that example, the conversation, assistant turns that emit structured calls with arguments, tool-role messages carrying results, and a final answer grounded in those results. Format alone is not enough. The set must also cover cases where no tool applies, required arguments are missing, several calls run in parallel, multi-turn state carries over, and tools return errors, or the tuned model will call tools reflexively and fail badly when something goes wrong.

By SourceX Editorial · Updated

What one function-calling training record contains

A complete record has five parts: tool definitions, the user turn, an assistant call, a tool result and a final assistant answer. Together AI's format puts a tools list on each JSONL line, lets assistant messages carry structured tool_calls in place of text content, and returns outputs as tool-role messages [2]. Predibase expects tool definitions in the Hugging Face convention: type: "function" plus a name, a description and JSON Schema parameters [3]. Google's Vertex AI tuning guide for Gemini uses its own functionDeclarations and functionCall/functionResponse parts, but the anatomy is the same [1].

The general chat mechanics (roles, system prompts, which tokens carry loss) are covered in our chat fine-tuning data format guide. What is specific to tool use is that the assistant's call is the target being learned. In most pipelines the loss applies to the assistant's tool_calls and final answer, while tool results are context only, because you do not want the model learning to hallucinate API responses.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"tools":[{"type":"function","function":{"name":"get_invoice_status","description":"Look up payment status for one invoice by its ID.","parameters":{"type":"object","properties":{"invoice_id":{"type":"string","pattern":"^INV-[0-9]{6}$"},"include_history":{"type":"boolean","default":false}},"required":["invoice_id"]}}},{"type":"function","function":{"name":"create_ticket","description":"Open a support ticket.","parameters":{"type":"object","properties":{"subject":{"type":"string"},"priority":{"type":"string","enum":["low","normal","high"]}},"required":["subject"]}}}],
 "messages":[
  {"role":"system","content":"You are a billing assistant. Use tools only when needed."},
  {"role":"user","content":"Has INV-204871 been paid yet?"},
  {"role":"assistant","content":null,"tool_calls":[{"id":"call_1","type":"function","function":{"name":"get_invoice_status","arguments":"{\"invoice_id\":\"INV-204871\"}"}}]},
  {"role":"tool","tool_call_id":"call_1","content":"{\"status\":\"paid\",\"paid_on\":\"2026-09-14\",\"amount\":1840.00,\"currency\":\"USD\"}"},
  {"role":"assistant","content":"Yes. INV-204871 was paid in full ($1,840.00) on September 14, 2026."}
 ]}

Note the details that break silently when they are wrong: arguments is a JSON-encoded string in OpenAI-style formats, tool_call_id must match the call id, and optional parameters the user never mentioned (include_history) are omitted rather than filled with defaults.

Coverage categories a function-calling SFT set needs

A usable set mixes at least seven behaviors, not just "user asks, model calls one tool." Google's tuning guide explicitly includes examples where the model answers in text rather than calling a function [1], and public benchmarks such as the Berkeley Function Calling Leaderboard (BFCL) score simple, multiple-function, parallel, multi-turn and irrelevance cases separately [4][5]. If your training mix omits a category, expect the tuned model to regress on it.

CategoryWhat the record teachesTarget answerTypical failure if missing
Single callPick one tool, fill required argsOne tool_calls entryWrong argument types, invented fields
Tool selectionChoose among several similar toolsCorrect name from a crowded listCalls the first plausible tool
No-call (irrelevance)Tools are present but none applyPlain text answerReflexive calls on every turn
Missing argumentsRequired value is absentClarifying question, no callHallucinated IDs, dates, amounts
Parallel callsIndependent lookups in one turnSeveral tool_calls in one messageSerial calls, or one merged call
Sequential and multi-turnOutput of call A feeds call B; state persistsChained calls across turnsRe-asks for known values, loses context
Error resultsTool returns 4xx, timeout or validation errorCorrected retry or honest failure messageRepeats the same failing call

The no-call and missing-argument rows matter most for production. A model tuned only on positive calls learns that the presence of a tools field means "emit a call," which shows up as fabricated customer_id values or calls on small talk. As a starting heuristic, many teams hold no-call and clarification examples at a meaningful share of the mix, then tune the ratio against a held-out irrelevance slice rather than a fixed rule.

How to normalize tool schemas before training

Normalize every tool definition to one schema dialect, ideally a strict subset of JSON Schema, before any record enters the set. The JSON Schema specification defines the keywords you will rely on: type, properties, required, enum, pattern, items and default [6]. Provider formats differ in what they accept, so pick the target serving stack first and convert to it, since Predibase, Together AI and Vertex AI each document their own envelope [1][2][3].

Normalization checklist:

  • One naming convention for functions (snake_case verbs such as get_invoice_status) and no duplicate names across tools in the same record.
  • Descriptions written for the model: what the tool does, when not to use it, and units or formats (ISO 8601 dates, cents vs dollars).
  • required lists that match real API behavior; optional fields stay optional.
  • Enums for closed vocabularies (priority, region, currency) instead of free strings.
  • No $ref chains or oneOf unions unless your serving runtime resolves them identically at inference time.
  • Schema order and tool count shuffled across records so the model does not learn position bias.

A subtle failure: training on schemas that differ from production schemas, even by a renamed parameter, teaches the model the wrong signature. Version tool definitions and record the schema version per example so you can regenerate records when an API changes.

Using real argument values from tool-call logs

Real argument values make examples realistic in ways synthetic generation rarely matches: messy order numbers, partial addresses, ambiguous date phrases and users who give two IDs in one sentence. Teams often source these from production traces; the OpenTelemetry GenAI semantic conventions define spans for model calls and tool execution, and tool-call arguments and results can be captured when the instrumentation opts in [9]. Support tickets, CRM activity and internal ops logs are also natural sources of "user intent paired with the system action a human took."

Logs need three transformations before they are training data. First, map each logged action to a normalized tool schema and re-serialize the arguments. Second, remove or replace personal details such as names, emails, phone numbers and account numbers in both arguments and results, with consistent pseudonyms so a multi-turn chain still links. Third, filter out calls that were wrong in production (the log records what happened, not what should have happened) or relabel them as negative or error-recovery examples.

When the data comes from another company's operational systems, licensing becomes part of the format question: you need the right to use the records for training, and the right to use derived tool schemas that may reveal internal API structure. For the wider sourcing picture, see how to source supervised fine-tuning data and the owner page on training data for tool use and function calling. If you need licensed operational records with real intents and actions, you can describe the records to SourceX; describe the data, not the businesses that might hold it.

How to represent parallel calls and error results

Parallel calls belong in a single assistant message with multiple tool_calls entries, each with its own id, followed by one tool message per call in any order. "What's the weather in Denver and Austin?" should yield two get_weather calls in one turn, not two sequential turns. Train both the parallel case and the look-alike sequential case, where call B needs the output of call A, so the model learns dependency rather than a habit of always fanning out.

Error results should be real-looking payloads: {"error":"invoice_not_found","code":404}, rate-limit messages, schema validation errors and timeouts. The target assistant turn either corrects the call (a fixed ID format, a missing field) or tells the user plainly what failed. Recent research reports that smaller models often repeat a failed call instead of reading the error message, and proposes training methods aimed at recovery [7][8]. SFT records with explicit error-then-repair sequences give a model that signal directly.

Illustrative example: invented to show structure; it does not describe an available dataset.

TurnRoleContent (abbreviated)
1user"Check order 55812 and refund it if it shipped late."
2assistantget_order({"order_id":"55812"})
3tool{"error":"invalid_order_id","hint":"expected ORD-#####"}
4assistantget_order({"order_id":"ORD-55812"})
5tool{"shipped":"2026-09-20","promised":"2026-09-16"}
6assistantissue_refund({"order_id":"ORD-55812","reason":"late_shipment"})
7tool{"refund_id":"RF-9001","status":"queued"}
8assistantPlain-text confirmation citing the refund ID

Keeping benchmark items out of your training set

Exclude public function-calling benchmark items from training and decontaminate against them before you report results. BFCL and similar suites test the same categories your SFT set covers [4][5], so training on their prompts or tool definitions inflates scores without improving real behavior. Run n-gram and normalized-schema matching against benchmark prompts and function names, and keep your own held-out slice drawn from your production tools.

Build that held-out slice per category from the coverage table, with exact-match scoring on function name and argument JSON (after key ordering) plus a pass/fail on no-call and clarification turns. The companion page on function-calling evaluation datasets covers expected-argument scoring in depth, and evaluating a fine-tuning dataset before you buy it covers sample review.

Questions to ask a function-calling data supplier

Ask any supplier for a sample that shows every coverage category, the schema dialect and version, and how arguments were produced. Useful questions:

  • What share of records are no-call, clarification, parallel, multi-turn and error-recovery, and how were those labeled?
  • Are argument values drawn from real operational records, synthetic generation, or both, and which fields were pseudonymized?
  • Were tool definitions written for this dataset or derived from real internal APIs, and do you have rights to share them?
  • Which benchmark suites were checked for overlap, and how?
  • Can records be delivered in your target envelope (OpenAI-style tool_calls, Hugging Face chat template, Vertex functionCall) with a validator script?

For broader structure and output-format work, see structured-output fine-tuning data; for long agent runs beyond single SFT records, see enterprise workflow datasets and agent trajectories. The fine-tuning and post-training data hub links the rest of the cluster.

Sourcing function-calling training data from real business operations

SourceX sources operational datasets from US companies, including support and sales histories, engineering records and finance and legal workflows, and manages the licensing process; data is sourced on request, so a request does not guarantee a match. Personal details are removed or replaced before delivery and every dataset is delivered under a license defining records, uses, term and delivery. If you need real intents and actions to build function-calling records, describe the data you need to SourceX.

Sources

  1. Google Cloud (Vertex AI documentation), "Tune function calling". https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/tune-function-calling
  2. Together AI, "Fine-tuning for function calling". https://docs.together.ai/docs/fine-tuning-function-calling
  3. Predibase, "Function Calling". https://docs.predibase.com/user-guide/fine-tuning/function_calling
  4. Patil et al., ICML 2025 (via ML Anthology), "The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models" (2025). https://mlanthology.org/icml/2025/patil2025icml-berkeley
  5. UC Berkeley Gorilla project, "Berkeley Function-Calling Leaderboard". https://gorilla.cs.berkeley.edu/leaderboard.html
  6. JSON Schema (json-schema.org), "JSON Schema specification". https://json-schema.org/specification
  7. arXiv (Zhang et al., 2026), "Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors" (2026). https://arxiv.org/html/2601.15625v1
  8. arXiv, "Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions" (2025). https://arxiv.org/pdf/2509.18847
  9. OpenTelemetry, "Inside the LLM Call: GenAI Observability with OpenTelemetry" (2026). https://opentelemetry.io/blog/2026/genai-observability/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data