Schemas, packaging and delivery
A Delivery Schema for Chat and Conversation Transcripts
Quick answer
A chat transcript delivery should be two linked record types: one conversation record (channel, start and end timestamps, outcome, satisfaction score, linked ticket) and an ordered stream of turn records (sequence number, speaker role, pseudonymous participant ID, UTC timestamp, text, attachments and system events such as transfers and bot-to-human handoffs). Deliver it as JSON Lines validated by JSON Schema, keep the source system's role and event codes, and convert to chat-template formats yourself after delivery.
By SourceX Editorial · Updated
Why a source-preserving schema beats a training-ready format
Ask for transcripts in a schema that mirrors the source system, not in the messages array your trainer expects, because conversion is lossy and only one direction is reversible. A {"role": "user" | "assistant"} list cannot tell you whether the "assistant" was a bot, a macro, a tier-1 human or a supervisor, when a transfer happened, or how the case ended. Those are the exact signals a support or sales agent model needs for routing, escalation and outcome-conditioned training.
Contact-center platforms already expose richer structure than chat templates. Amazon Connect's chat transcript items carry a ParticipantRole of AGENT, CUSTOMER, SYSTEM, CUSTOM_BOT or SUPERVISOR, plus a Type, a ContentType, an AbsoluteTime in ISO 8601 with milliseconds and a sender DisplayName [1]. Ticket systems add an audit trail: Zendesk records Change and Comment events with public, previous_value, via and author_id fields [5]. A delivery that flattens all of that into alternating turns throws away labels you cannot recreate.
Converting to a training format is a buyer step. Our guide to chat fine-tuning data format, roles and loss masking covers that downstream transform; this page covers what should arrive on disk first.
Conversation-level fields: one row per thread
The conversation record answers "what happened in this thread and how did it end," and every field should be nullable rather than invented. Keep it in a separate file (or table) keyed by conversation_id so you can filter and stratify without parsing every turn.
| Field | Type | Why it matters | Source examples |
|---|---|---|---|
conversation_id | string, stable across deliveries | Join key; dedupe across refreshes | Contact ID, chat session ID, ticket ID |
source_system | enum | Role and event codes differ by platform | amazon_connect, zendesk, salesforce, in_house |
channel | enum | Web chat, SMS, in-app, email-to-chat, social DM behave differently | Channel or via field [5] |
language | BCP 47 tag | Split and evaluate by locale | Detected or configured locale |
started_at, ended_at | RFC 3339 / ISO 8601 UTC | Duration, time-based splits | First and last item timestamps [1] |
queue, skill | string | Routing context for agent policies | Queue or group name, pseudonymized if needed |
outcome_code | enum + outcome_source | Resolution label for reward or eval | Ticket status, disposition code, deal stage |
csat_score, csat_scale | number, string | Satisfaction signal; scale must travel with it | Post-chat survey |
linked_ticket_id, linked_crm_id | string, nullable | Connects chat to later resolution | Ticket or opportunity reference |
handoff_count, transfer_count | integer | Quick filter for escalated threads | Derived from events; document the rule |
redaction_method, redaction_version | string | Tells you which tokens are surrogates | Supplier preparation notes |
Record outcome_source explicitly. "Resolved" set by an agent closing a ticket is a different label from "resolved" inferred because the customer did not return within seven days, and models trained on the two learn different things.
Turn-level fields: order, role and actor type
Each turn record must let you rebuild the exact on-screen sequence and know who produced every line. Minimum fields: conversation_id, seq (gapless integer from 0), item_id from the source, ts (UTC, millisecond precision), role (normalized), source_role (verbatim), actor_type, participant_pseudonym, item_type, content_type, text, attachments[] and visibility.
Keep both seq and ts. Timestamps collide at second precision, arrive out of order from multi-region systems, and some sources only provide offsets: Amazon Connect Contact Lens transcripts locate segments by offset from the start of the contact instead of absolute time [3]. A seq assigned by the supplier from the source's own ordering is what you sort by; ts is what you compute latency from.
Separate role from actor_type. role answers "which side" (customer, agent, system); actor_type answers "what produced it" (human, bot, macro, auto-reply, supervisor whisper). Amazon Connect distinguishes CUSTOM_BOT and SUPERVISOR from AGENT at the source [1], and canned macros typically show up only as metadata on an agent comment. The practical detection rules are in separating bot, macro and human turns in conversation data.
Use visibility for internal notes. Zendesk comments carry a public flag [5]; an internal note such as "customer is on legacy plan, escalate to billing" is valuable for reasoning traces but must never be rendered as something the customer saw.
Events, transfers and bot-to-human handoffs
System events belong in the same ordered stream as messages, typed so you can drop or keep them per experiment. Amazon Connect returns item types including PARTICIPANT_JOINED, PARTICIPANT_LEFT, TRANSFER_SUCCEEDED, TRANSFER_FAILED and CHAT_ENDED, with event content types under application/vnd.amazonaws.connect.event.* [2]. Typing indicators, read receipts and delivery acknowledgements also appear as types in that API [2]; ask the supplier to state whether they were kept or dropped.
Handoffs are the hardest field to get right because most bot platforms record intent, not the actual join. Dialogflow CX's LiveAgentHandoff response is a signal that the conversation should go to a live agent, carrying free-form metadata; Dialogflow itself does not perform the transfer [4]. So a delivered handoff should have two events: handoff_requested (from the bot) and participant_joined for the human, with the gap between them as queue wait. If only the request exists, the customer may have abandoned, and that is an outcome, not missing data.
Recommended item_type vocabulary: message, attachment, handoff_requested, participant_joined, participant_left, transfer_succeeded, transfer_failed, internal_note, field_change, conversation_ended. Map each source code into this list and keep the original in source_item_type.
Illustrative record layout in JSON Lines
JSON Lines is the default container: UTF-8, no byte order mark, one JSON object per line, no blank lines [6]. Ship conversations.jsonl and turns.jsonl (or one nested file if threads are short), each with a JSON Schema (draft 2020-12) that uses required and enum so your loader rejects bad rows on arrival [7]. For analytics, a Parquet copy of the same tables is convenient because the file footer records where each column chunk starts, so readers can load only the columns they need [9].
Illustrative example: invented to show structure; it does not describe an available dataset.
{"conversation_id":"c_8f21","source_system":"amazon_connect","channel":"web_chat","language":"en-US","started_at":"2025-11-03T14:02:11.204Z","ended_at":"2025-11-03T14:19:47.880Z","queue":"billing_tier1","outcome_code":"resolved","outcome_source":"ticket_status_at_close","csat_score":4,"csat_scale":"1-5","linked_ticket_id":"t_55102","handoff_count":1,"transfer_count":0,"redaction_method":"surrogate_replacement","redaction_version":"2025-10-r3"}
{"conversation_id":"c_8f21","seq":0,"item_id":"src_001","ts":"2025-11-03T14:02:11.204Z","role":"customer","source_role":"CUSTOMER","actor_type":"human","participant_pseudonym":"cust_a","item_type":"message","source_item_type":"MESSAGE","content_type":"text/plain","text":"I was charged twice for October.","attachments":[],"visibility":"public"}
{"conversation_id":"c_8f21","seq":1,"item_id":"src_002","ts":"2025-11-03T14:02:12.010Z","role":"agent","source_role":"CUSTOM_BOT","actor_type":"bot","participant_pseudonym":"bot_billing","item_type":"message","source_item_type":"MESSAGE","content_type":"text/plain","text":"I can help with billing. Can you confirm the last four digits of the card?","attachments":[],"visibility":"public"}
{"conversation_id":"c_8f21","seq":4,"item_id":"src_005","ts":"2025-11-03T14:03:40.552Z","role":"system","source_role":"SYSTEM","actor_type":"bot","participant_pseudonym":"bot_billing","item_type":"handoff_requested","source_item_type":"EVENT","content_type":"application/json","text":null,"event":{"reason":"duplicate_charge","target_queue":"billing_tier1"},"attachments":[],"visibility":"system"}
{"conversation_id":"c_8f21","seq":5,"item_id":"src_006","ts":"2025-11-03T14:06:02.917Z","role":"system","source_role":"SYSTEM","actor_type":"system","participant_pseudonym":"agent_17","item_type":"participant_joined","source_item_type":"PARTICIPANT_JOINED","content_type":"application/vnd.amazonaws.connect.event.participant.joined","text":null,"attachments":[],"visibility":"system"}
{"conversation_id":"c_8f21","seq":7,"item_id":"src_008","ts":"2025-11-03T14:07:15.300Z","role":"agent","source_role":"AGENT","actor_type":"human","participant_pseudonym":"agent_17","item_type":"internal_note","source_item_type":"linked_ticket_comment","content_type":"text/plain","text":"Second charge is a pending auth, will drop in 3 days.","attachments":[],"visibility":"internal"}
Note the gaps in seq in this excerpt are only because lines were omitted for space; a real delivery should be gapless, and a gap is a validation failure.
Attachments, redaction markers and long content
Attachments should be referenced, not inlined: attachments[] holds attachment_id, mime_type, bytes, sha256 and a relative path into an attachments/ prefix, with the files listed in the delivery manifest. That lets you verify each file with the checks in dataset manifests and checksums and drop images without rewriting the text files.
Redaction must be visible in the data. Ask for typed surrogates ([EMAIL_1], [ACCOUNT_1]) that stay consistent within a conversation, so "your account [ACCOUNT_1]" in turn 3 still matches turn 9. Untyped *** masks break coreference and make the bot's confirmation steps unlearnable. Expect some residual identifiers anyway; no redaction method is perfect, so plan your own scan.
Watch platform length limits. Amazon Connect caps message or event Content at 16,384 characters [1], and other systems truncate or split long pastes. Ask the supplier to flag truncated: true rather than silently clipping text.
Validation checks to run on arrival
Run structural and semantic checks before any conversion, and reject a delivery batch rather than patching it locally. A shared JSON Schema covers types and enums [7]; the rest needs a small script.
Arrival checklist for conversation deliveries
- Every line parses; files are UTF-8 without BOM; no blank lines [6].
- Every
turns.conversation_idexists inconversations, and vice versa. -
seqis gapless from 0 per conversation; no duplicateitem_id. -
tsis non-decreasing withseqwithin tolerance; note and explain exceptions. - Every
source_roleandsource_item_typemaps to the documented normalized value. - Each
handoff_requestedis followed byparticipant_joinedor a documented abandonment outcome. -
internalturns never haverole: customer. -
outcome_codeandcsat_scoredistributions match the supplier's data card. - Surrogate tokens follow the documented pattern; a regex scan finds no raw emails, phone numbers or card numbers in a sample.
- Attachment hashes match the manifest.
Describe the result in machine-readable metadata. Croissant expresses dataset-level metadata, file resources and record structure in schema.org JSON-LD [8], which suits a two-file conversation dataset; see Croissant metadata for licensed datasets. For ongoing feeds, version the schema and follow schema change handling across recurring deliveries so a new item_type does not silently disappear.
Where this schema fits in sourcing chat data
Put this schema into the technical exhibit of your request so suppliers can say early which fields their systems actually hold. Many sources lack CSAT, some lack outcome codes, and older exports may lose bot versus human distinctions; knowing that before pricing saves a re-delivery. The broader set of delivery decisions is collected in the dataset delivery formats and schemas hub, and the field list above slots directly into JSON Schema dataset validation.
SourceX sources operational datasets, including support and sales histories, from US companies on request, and any match depends on a supplier agreeing to release it. If you are scoping licensed chat logs or customer support transcripts, you can describe the conversation data you need to SourceX. Voice channels have their own conventions; see call transcript.
Request licensed chat transcript data
SourceX looks for US businesses that hold the conversation data you describe, reviews ownership and consents, and delivers under a license that defines the records, allowed uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, with the method recorded. Describe your transcript requirements to SourceX.
Sources
- Amazon Web Services, "Item (Amazon Connect Participant Service API Reference)". https://docs.aws.amazon.com/connect-participant/latest/APIReference/API_Item.html
- Amazon Web Services, "ConnectParticipant get_transcript (Boto3/botocore reference)". https://docs.aws.amazon.com/botocore/latest/reference/services/connectparticipant/client/get_transcript.html
- Amazon Web Services, "Transcript (Amazon Connect Contact Lens API Reference)". https://docs.aws.amazon.com/connect/latest/APIReference/API_connect-contact-lens_Transcript.html
- Google Cloud, "ResponseMessage.LiveAgentHandoff (Dialogflow CX Python reference)". https://docs.cloud.google.com/python/docs/reference/dialogflow-cx/latest/google.cloud.dialogflowcx_v3.types.ResponseMessage.LiveAgentHandoff
- Zendesk, "Ticket Audit events reference". https://developer.zendesk.com/documentation/ticketing/reference-guides/ticket-audit-events-reference/
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- json-schema.org, "JSON Schema Specification". https://json-schema.org/specification
- MLCommons Croissant working group, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- The Apache Software Foundation, "File Format" (Apache Parquet). https://parquet.apache.org/docs/file-format/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.