Fine-tuning and post-training data
Function-calling fine-tuning data: schemas, calls, results and no-call cases
Quick answer
A function-calling SFT set is a JSONL file in which each line holds the tool schemas available for that example, the conversation, assistant turns that emit structured calls with arguments, tool-role messages carrying results, and a final answer grounded in those results. Format alone is not enough. The set must also cover cases where no tool applies, required arguments are missing, several calls run in parallel, multi-turn state carries over, and tools return errors, or the tuned model will call tools reflexively and fail badly when something goes wrong.
By SourceX Editorial · Updated
What one function-calling training record contains
A complete record has five parts: tool definitions, the user turn, an assistant call, a tool result and a final assistant answer. Together AI's format puts a tools list on each JSONL line, lets assistant messages carry structured tool_calls in place of text content, and returns outputs as tool-role messages [2]. Predibase expects tool definitions in the Hugging Face convention: type: "function" plus a name, a description and JSON Schema parameters [3]. Google's Vertex AI tuning guide for Gemini uses its own functionDeclarations and functionCall/functionResponse parts, but the anatomy is the same [1].
The general chat mechanics (roles, system prompts, which tokens carry loss) are covered in our chat fine-tuning data format guide. What is specific to tool use is that the assistant's call is the target being learned. In most pipelines the loss applies to the assistant's tool_calls and final answer, while tool results are context only, because you do not want the model learning to hallucinate API responses.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"tools":[{"type":"function","function":{"name":"get_invoice_status","description":"Look up payment status for one invoice by its ID.","parameters":{"type":"object","properties":{"invoice_id":{"type":"string","pattern":"^INV-[0-9]{6}$"},"include_history":{"type":"boolean","default":false}},"required":["invoice_id"]}}},{"type":"function","function":{"name":"create_ticket","description":"Open a support ticket.","parameters":{"type":"object","properties":{"subject":{"type":"string"},"priority":{"type":"string","enum":["low","normal","high"]}},"required":["subject"]}}}],
"messages":[
{"role":"system","content":"You are a billing assistant. Use tools only when needed."},
{"role":"user","content":"Has INV-204871 been paid yet?"},
{"role":"assistant","content":null,"tool_calls":[{"id":"call_1","type":"function","function":{"name":"get_invoice_status","arguments":"{\"invoice_id\":\"INV-204871\"}"}}]},
{"role":"tool","tool_call_id":"call_1","content":"{\"status\":\"paid\",\"paid_on\":\"2026-09-14\",\"amount\":1840.00,\"currency\":\"USD\"}"},
{"role":"assistant","content":"Yes. INV-204871 was paid in full ($1,840.00) on September 14, 2026."}
]}
Note the details that break silently when they are wrong: arguments is a JSON-encoded string in OpenAI-style formats, tool_call_id must match the call id, and optional parameters the user never mentioned (include_history) are omitted rather than filled with defaults.
Coverage categories a function-calling SFT set needs
A usable set mixes at least seven behaviors, not just "user asks, model calls one tool." Google's tuning guide explicitly includes examples where the model answers in text rather than calling a function [1], and public benchmarks such as the Berkeley Function Calling Leaderboard (BFCL) score simple, multiple-function, parallel, multi-turn and irrelevance cases separately [4][5]. If your training mix omits a category, expect the tuned model to regress on it.
| Category | What the record teaches | Target answer | Typical failure if missing |
|---|---|---|---|
| Single call | Pick one tool, fill required args | One tool_calls entry | Wrong argument types, invented fields |
| Tool selection | Choose among several similar tools | Correct name from a crowded list | Calls the first plausible tool |
| No-call (irrelevance) | Tools are present but none apply | Plain text answer | Reflexive calls on every turn |
| Missing arguments | Required value is absent | Clarifying question, no call | Hallucinated IDs, dates, amounts |
| Parallel calls | Independent lookups in one turn | Several tool_calls in one message | Serial calls, or one merged call |
| Sequential and multi-turn | Output of call A feeds call B; state persists | Chained calls across turns | Re-asks for known values, loses context |
| Error results | Tool returns 4xx, timeout or validation error | Corrected retry or honest failure message | Repeats the same failing call |
The no-call and missing-argument rows matter most for production. A model tuned only on positive calls learns that the presence of a tools field means "emit a call," which shows up as fabricated customer_id values or calls on small talk. As a starting heuristic, many teams hold no-call and clarification examples at a meaningful share of the mix, then tune the ratio against a held-out irrelevance slice rather than a fixed rule.
How to normalize tool schemas before training
Normalize every tool definition to one schema dialect, ideally a strict subset of JSON Schema, before any record enters the set. The JSON Schema specification defines the keywords you will rely on: type, properties, required, enum, pattern, items and default [6]. Provider formats differ in what they accept, so pick the target serving stack first and convert to it, since Predibase, Together AI and Vertex AI each document their own envelope [1][2][3].
Normalization checklist:
- One naming convention for functions (
snake_caseverbs such asget_invoice_status) and no duplicate names across tools in the same record. - Descriptions written for the model: what the tool does, when not to use it, and units or formats (ISO 8601 dates, cents vs dollars).
requiredlists that match real API behavior; optional fields stay optional.- Enums for closed vocabularies (
priority,region,currency) instead of free strings. - No
$refchains oroneOfunions unless your serving runtime resolves them identically at inference time. - Schema order and tool count shuffled across records so the model does not learn position bias.
A subtle failure: training on schemas that differ from production schemas, even by a renamed parameter, teaches the model the wrong signature. Version tool definitions and record the schema version per example so you can regenerate records when an API changes.
Using real argument values from tool-call logs
Real argument values make examples realistic in ways synthetic generation rarely matches: messy order numbers, partial addresses, ambiguous date phrases and users who give two IDs in one sentence. Teams often source these from production traces; the OpenTelemetry GenAI semantic conventions define spans for model calls and tool execution, and tool-call arguments and results can be captured when the instrumentation opts in [9]. Support tickets, CRM activity and internal ops logs are also natural sources of "user intent paired with the system action a human took."
Logs need three transformations before they are training data. First, map each logged action to a normalized tool schema and re-serialize the arguments. Second, remove or replace personal details such as names, emails, phone numbers and account numbers in both arguments and results, with consistent pseudonyms so a multi-turn chain still links. Third, filter out calls that were wrong in production (the log records what happened, not what should have happened) or relabel them as negative or error-recovery examples.
When the data comes from another company's operational systems, licensing becomes part of the format question: you need the right to use the records for training, and the right to use derived tool schemas that may reveal internal API structure. For the wider sourcing picture, see how to source supervised fine-tuning data and the owner page on training data for tool use and function calling. If you need licensed operational records with real intents and actions, you can describe the records to SourceX; describe the data, not the businesses that might hold it.
How to represent parallel calls and error results
Parallel calls belong in a single assistant message with multiple tool_calls entries, each with its own id, followed by one tool message per call in any order. "What's the weather in Denver and Austin?" should yield two get_weather calls in one turn, not two sequential turns. Train both the parallel case and the look-alike sequential case, where call B needs the output of call A, so the model learns dependency rather than a habit of always fanning out.
Error results should be real-looking payloads: {"error":"invoice_not_found","code":404}, rate-limit messages, schema validation errors and timeouts. The target assistant turn either corrects the call (a fixed ID format, a missing field) or tells the user plainly what failed. Recent research reports that smaller models often repeat a failed call instead of reading the error message, and proposes training methods aimed at recovery [7][8]. SFT records with explicit error-then-repair sequences give a model that signal directly.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Turn | Role | Content (abbreviated) |
|---|---|---|
| 1 | user | "Check order 55812 and refund it if it shipped late." |
| 2 | assistant | get_order({"order_id":"55812"}) |
| 3 | tool | {"error":"invalid_order_id","hint":"expected ORD-#####"} |
| 4 | assistant | get_order({"order_id":"ORD-55812"}) |
| 5 | tool | {"shipped":"2026-09-20","promised":"2026-09-16"} |
| 6 | assistant | issue_refund({"order_id":"ORD-55812","reason":"late_shipment"}) |
| 7 | tool | {"refund_id":"RF-9001","status":"queued"} |
| 8 | assistant | Plain-text confirmation citing the refund ID |
Keeping benchmark items out of your training set
Exclude public function-calling benchmark items from training and decontaminate against them before you report results. BFCL and similar suites test the same categories your SFT set covers [4][5], so training on their prompts or tool definitions inflates scores without improving real behavior. Run n-gram and normalized-schema matching against benchmark prompts and function names, and keep your own held-out slice drawn from your production tools.
Build that held-out slice per category from the coverage table, with exact-match scoring on function name and argument JSON (after key ordering) plus a pass/fail on no-call and clarification turns. The companion page on function-calling evaluation datasets covers expected-argument scoring in depth, and evaluating a fine-tuning dataset before you buy it covers sample review.
Questions to ask a function-calling data supplier
Ask any supplier for a sample that shows every coverage category, the schema dialect and version, and how arguments were produced. Useful questions:
- What share of records are no-call, clarification, parallel, multi-turn and error-recovery, and how were those labeled?
- Are argument values drawn from real operational records, synthetic generation, or both, and which fields were pseudonymized?
- Were tool definitions written for this dataset or derived from real internal APIs, and do you have rights to share them?
- Which benchmark suites were checked for overlap, and how?
- Can records be delivered in your target envelope (OpenAI-style
tool_calls, Hugging Face chat template, VertexfunctionCall) with a validator script?
For broader structure and output-format work, see structured-output fine-tuning data; for long agent runs beyond single SFT records, see enterprise workflow datasets and agent trajectories. The fine-tuning and post-training data hub links the rest of the cluster.
Sourcing function-calling training data from real business operations
SourceX sources operational datasets from US companies, including support and sales histories, engineering records and finance and legal workflows, and manages the licensing process; data is sourced on request, so a request does not guarantee a match. Personal details are removed or replaced before delivery and every dataset is delivered under a license defining records, uses, term and delivery. If you need real intents and actions to build function-calling records, describe the data you need to SourceX.
Sources
- Google Cloud (Vertex AI documentation), "Tune function calling". https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/tune-function-calling
- Together AI, "Fine-tuning for function calling". https://docs.together.ai/docs/fine-tuning-function-calling
- Predibase, "Function Calling". https://docs.predibase.com/user-guide/fine-tuning/function_calling
- Patil et al., ICML 2025 (via ML Anthology), "The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models" (2025). https://mlanthology.org/icml/2025/patil2025icml-berkeley
- UC Berkeley Gorilla project, "Berkeley Function-Calling Leaderboard". https://gorilla.cs.berkeley.edu/leaderboard.html
- JSON Schema (json-schema.org), "JSON Schema specification". https://json-schema.org/specification
- arXiv (Zhang et al., 2026), "Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors" (2026). https://arxiv.org/html/2601.15625v1
- arXiv, "Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions" (2025). https://arxiv.org/pdf/2509.18847
- OpenTelemetry, "Inside the LLM Call: GenAI Observability with OpenTelemetry" (2026). https://opentelemetry.io/blog/2026/genai-observability/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.