Agent, workflow and domain-reasoning data
Seed data and state snapshots for enterprise agent sandboxes
Quick answer
Seed data for agent sandbox environments is the database state a simulated enterprise application loads before every episode: linked accounts, contacts, orders, invoices and tickets captured at one instant, with foreign keys intact and real distributions, duplicates and free-text notes preserved. Teams that need realistic difficulty can license a de-identified production snapshot or subset, mask identifiers identically across every table, and confirm the license covers environment builds, repeated resets and anything derived from the records.
By SourceX Editorial · Updated
This page covers the data state only. Tasks and graders are covered in agent evaluation task suites, the other records an RL environment needs in RL environments from business workflows and training data for RL environments, and the wider category in the agent training data hub.
Where generated seed data falls short
Generated seed data is schema-valid and easy to share, but it rarely reproduces the joint distributions, long-tail values and inconsistencies that make enterprise tasks hard, so agents can look more capable in the sandbox than in production. Generation is a common default in sandbox tooling: as of October 2026, Fruxon's sandbox documentation describes a seed, run, inspect and revert loop whose Generate tab turns a plain-language description into JSON payloads, 1 to 100 per call [1]. Public benchmarks lean the same way; the original τ-bench retail and airline domains are synthetic databases behind programmatic APIs [2].
Making generated records realistic is a project of its own. CRMArena-Pro, from Salesforce AI Research, generates data on real Salesforce schemas spanning Service Cloud, Sales Cloud and CPQ, runs every generated record through multi-stage validation, and used expert studies to judge how realistic the data and sandbox were [3]. What generators usually miss is what trips agents:
- Cross-table correlation: discount depth that depends on deal size and quarter-end, or reopen rates clustered on one product version.
- Entity duplicates: one customer stored as two accounts, or one company spread across two systems. August 2026 press coverage of Arga Labs, a startup building resettable replicas of tools such as Salesforce and Workday, describes testing whether an agent recognizes that a Salesforce lead and a separate HubSpot outreach refer to the same company [4].
- Stale state: contacts who left, opportunities open past their close date, closed cases with open child tasks.
- Working free text: abbreviations, pasted email threads, internal codes and references to other records.
- Legacy records: rows that violate today's validation rules because they predate them.
| Seed approach | Where it fits | What it misses or risks | Rights and privacy load |
|---|---|---|---|
| Generated from schema and prompts | Tool coverage, early debugging, data that can never leave a company | Real distributions, duplicates, long-tail values | Low; check the generator model's terms |
| Masked production snapshot or subset | Realistic difficulty for training and evaluation; known outcomes for grading | Quasi-identifiers can survive masking; outcomes leak if the cutoff is wrong | Highest: supplier rights, consents, de-identification, license scope |
| Hybrid: real skeleton, synthetic fill | Volume and rare cases on real structure | Seams where generated rows join real ones | License must cover derived synthetic generation |
The same choice for trajectories is compared in licensed, commissioned or synthetic trajectories; blending methods are in combining licensed and synthetic data.
What a usable enterprise state snapshot contains
A usable snapshot is a closed object graph: every record an agent can reach through a lookup, related list or API call is present, with the reference data and users that give it meaning. For a sales-operations sandbox that means accounts, contacts, opportunities with line items, quotes, price books, cases and activities; for accounts payable, the vendor master, purchase orders, goods receipts, invoices, payment terms and holds; for IT service, incidents, requests, configuration items, assignment groups and knowledge articles. Specify four layers:
- Transactional records: the rows tasks read and change.
- Reference and configuration data: picklist values, approval matrices, business hours, SLA definitions and currency tables. A task that mentions a "Tier 2 SLA" is meaningless without them; see configuration data for enterprise app replicas.
- Users, roles and ownership: owner and assignee fields point at users, who must exist in the sandbox (pseudonymized) with roles, queues and permissions, or permission-scoped tasks cannot be tested.
- History: field-history and audit rows, comments and attachment metadata, needed when tasks ask who changed a value; see field-level audit trails.
Size a snapshot by the queries the agent will run, not by table count: a name search needs enough near-matches to be ambiguous, or every lookup succeeds on the first try. As of October 2026, Salesforce's Trailhead material describes seeding sandboxes from production or from a backup and a "Sample for Coverage" option meant to reflect the spread of values in production [5], while Keepit's guidance advises seeding only the objects you need and limiting parent records to stay within storage limits [6]. The tension is the same: small enough to reset quickly, large enough to keep ambiguity real.
Cutting one point in time across every table
A snapshot is time-consistent when every table reflects the same instant, so no invoice points at a purchase order that does not exist yet and no task's answer is already written into the data. When the supplier controls the database, all tables should come from one transaction: PostgreSQL documents a SERIALIZABLE READ ONLY DEFERRABLE mode that waits for a safe snapshot and suits long-running reports or backups, and a SET TRANSACTION SNAPSHOT command that lets parallel export sessions share one snapshot [7]. SaaS bulk or REST API exports can run for hours, so the supplier should record a cutoff, drop rows created after it, and roll back later field changes from history tables where they exist.
The cutoff also prevents outcome leakage, a timing error that silently inflates scores. If a task asks the agent to approve or reject an invoice, the snapshot must be cut before that decision, with the later approval row, status change and audit entries removed and kept separately as the grading key. RelBench, a benchmark for learning on multi-table relational databases, uses temporal splits for the same reason: models cannot use future data to predict earlier events [8].
Dates need one more decision. Agents read ages, due dates and SLA clocks relative to "now", so builders often shift every timestamp by one constant so the cutoff lands on the sandbox's current date and all intervals survive. For health-adjacent records, HIPAA Safe Harbor removes all elements of dates directly related to an individual except the year [9], so keeping day-level dates, even shifted, is a question for an Expert Determination; dates and ZIP codes under Safe Harbor explains what each option keeps.
Keeping joins intact through subsetting and masking
Referential integrity survives subsetting only if the extract starts from a driver set and follows foreign keys in both directions, and it survives masking only if each identifier gets the same surrogate in every table and text field. Pick driver records (for example, accounts stratified by segment and region), pull their children (contacts, opportunities, cases, orders), then pull every parent those children reference (owners, products, price books) even when it falls outside the sample. Count orphans for declared and undeclared keys; an order number typed into a case subject is an undeclared key, and that is where many breaks hide.
As of October 2026, GRAX's sandbox seeding documentation describes three modes for seeded values: preserve them, randomize them on each seed, or anonymize deterministically so the same input yields the same output, a mode it describes as useful when records must match external systems or be seeded into several environments [10]. Deterministic replacement is the right default for agent sandboxes, because an agent must find the same pseudonymous customer in the CRM, the ERP and an email thread. Keyed hashing or a token vault provides that consistency, and the key is what keeps the data re-linkable, so ask who holds it; methods are compared in pseudonymisation techniques for training data.
Consistency has a legal consequence. Under the GDPR, pseudonymised data that can be attributed to a person using separately kept additional information remains personal data (Article 4(5) and Recital 26) [11]; whether it is personal data for a recipient without the key is covered in pseudonymised data from the recipient's perspective.
Free text needs its own pass, mapped to the same surrogates as the structured fields; the Presidio project, an open-source PII detection and anonymization SDK, states that because it uses trained models it cannot guarantee finding all sensitive information [12]. Loading needs care too: seeding through the application fires validation rules and automation, so legacy records get rejected and triggers send emails or create tasks. Load with automation off and validation bypassed where the platform allows, then reconcile row counts against the manifest.
Preserving messiness without preserving identities
The realism a buyer pays for sits in the same places as re-identification risk: rare combinations, free text and exact timestamps. Rocher, Hendrickx and de Montjoye estimated with a generative model that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes [13]. That is a model estimate, not a count of people re-identified, but it shows why a job title, a small employer, a city and a dated event can single out a person after names are gone. NIST SP 800-188 describes transforming such quasi-identifiers, generating synthetic data from identified data, sharing through protected enclaves, and governance such as a Disclosure Review Board and re-identification studies [14].
Ask the supplier to measure what masking changed, because only the supplier can compare against the source. SDMetrics, the Synthetic Data Vault project's evaluation library, compares synthetic data with the real data it was modeled on at column, column-pair, table and multi-table level without needing to know how the data was produced, and its Quality Report scores column shapes and column-pair trends but is not optimized for ordered sequential data [15]; the same comparison works for a masked snapshot against its source. Pair it with source-side duplicate, null and validation-failure rates to confirm masking did not quietly clean the data. Re-identification risk assessment for licensed datasets lists the methods to request.
A seed manifest a buyer can audit
A seed manifest lets a team reproduce, reset and audit a sandbox state; ask for one with every snapshot version.
Illustrative example: invented to show structure; it does not describe an available dataset.
snapshot_id: crm-erp-midmarket-v3
cutoff_utc: "2026-03-31T23:59:59Z"
date_shift_days: 192 # constant; cutoff maps to sandbox load date
sources:
- system: CRM
export: bulk API
extraction_window: "2026-04-01T00:10Z/2026-04-01T05:42Z"
post_cutoff_changes: rolled back from field history
- system: ERP
export: single read-only serializable transaction
subset:
driver: accounts stratified by segment and region (2,400 of 31,800)
closure: children + every referenced parent
orphans: {declared_keys: 0, text_references_unresolved: 37}
withheld_for_grading: [approvals_after_cutoff, case_closures_after_cutoff, invoice_payment_status]
masking:
method: keyed HMAC surrogates, deterministic across systems and versions
key_holder: supplier
free_text: pattern + NER pass, 500-note manual review
generalized: {contact.title: job_family, account.employee_count: band}
messiness_profile: # source vs masked
duplicate_account_rate: [0.041, 0.041]
null_rate.contact.phone: [0.27, 0.27]
failing_current_validation: [0.06, 0.06]
fidelity_report: quality_report_source_vs_masked.json
load: {automation_disabled: true, validation_bypass: true, post_load_sha256: "<hash>"}
license_ref: "Schedule B: environment build, resets, parallel instances, derived tasks"
For multi-system deliveries, see packaging linked records from multiple business systems.
Acceptance checks on a seed sample
Run these checks on a sample load before committing to the full snapshot; each catches a failure that would otherwise look like an agent error.
| Check | How to run it | Failure it catches |
|---|---|---|
| Graph closure | Count orphans per declared key, then resolve record numbers found in text | Lookups that return nothing; tasks that cannot be completed |
| Cutoff | Search for timestamps after the cutoff (before shifting); confirm withheld outcome rows are absent | Answers already in the state |
| Masking consistency | Trace 20 entities across every table and text field | One customer appearing under two pseudonyms |
| Residual identifiers | PII scan plus manual review of notes and attachment names | Real names or emails inside free text |
| Messiness retained | Compare duplicate, null and validation-failure rates with the source profile | A "cleaned" snapshot that makes tasks easy |
| Load fidelity | Load with automation off; reconcile counts with the manifest | Silently rejected legacy rows |
| Reset determinism | Load, run, restore and hash the state several times | Score noise caused by the environment |
| Task coverage | Count matching records for each planned task template | Tasks with one obvious answer |
General schema, missingness and range tests are listed in validation checks for structured dataset deliveries.
Rights questions specific to sandbox seeding
A standard training-data license often does not describe how environments use data: copied into many instances, reset thousands of times, opened by contractors, and mined for tasks, graders and synthetic records. Settle these points in writing:
- Environment build and hosting: where instances run and how many may run in parallel.
- Resets and copies: unlimited reloads of the same state, and copies held by evaluation partners or annotation contractors.
- Derived synthetic data: whether you may train a generator on the snapshot and keep its output after the term; see generating synthetic data from licensed data.
- Derived tasks and benchmarks: internal use only, shared with partners, or published.
- No re-identification: California's definition of deidentified information requires the business to contractually obligate recipients to comply with its conditions, including not re-identifying [16], so expect that clause from US suppliers relying on it.
- Replicated software: schemas, configuration and screens of the supplier's software vendor carry their own terms; see rights when replicating third-party software.
Clause-level guidance is in license terms for agent data.
Datasets requested through SourceX's buyer process go through rights review, which checks that the business owns or may share the records and that required consents are in place, and are delivered under a license that defines which records are included, what they can be used for, how long the license runs and how delivery happens. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, with the method recorded and a sample checked afterward; no de-identification method is perfect, so put the consistency, cutoff and messiness requirements above in the request. For health records, SourceX requires HIPAA de-identification (Safe Harbor or Expert Determination) before anything is considered for a license.
Building agent sandboxes that need real record populations?
Describe the systems and objects you need, the cutoff, the subset size, the masking consistency and the uses to be licensed, such as environment builds, resets and derived tasks; writing an agent data specification shows how to structure the request. SourceX looks for US businesses that hold those records, checks the data and each supplier's licensing permissions, and manages the license and delivery. Datasets are sourced on request, a request does not guarantee a matching dataset, and nothing is contracted until a supplier agrees. Specify your sandbox seed data with SourceX.
Sources
- Fruxon documentation, "Sandbox Workspace". https://docs.fruxon.com/guides/integrations/sandbox-workspace/
- Yao, Shinn, Razavi, Narasimhan (Sierra), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Salesforce AI Research, "CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions" (2025). https://arxiv.org/pdf/2505.18878
- TechCrunch, "Arga is building a better way to train enterprise AI agents" (2026). https://techcrunch.com/2026/08/26/arga-is-building-a-better-way-to-train-enterprise-ai-agents
- Salesforce Trailhead, "Develop and Test Features Using Salesforce Sandboxes". https://trailhead.salesforce.com/content/learn/modules/salesforce-sandboxes-quick-look/develop-and-test-features-using-salesforce-sandboxes
- Keepit, "Sandbox seeding best practices". https://www.keepit.com/help/salesforce-category/sandbox-seeding-best-practices/
- PostgreSQL Global Development Group, "SET TRANSACTION" (PostgreSQL documentation). https://www.postgresql.org/docs/current/sql-set-transaction.html
- Robinson et al., "RelBench: A Benchmark for Deep Learning on Relational Databases" (2024). https://arxiv.org/pdf/2407.20060
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- GRAX documentation, "Sandbox Seeding". https://documentation.grax.com/docs/sandbox-seeding
- European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
- Microsoft Presidio project, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- Rocher, Hendrickx, de Montjoye, "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- DataCebo (Synthetic Data Vault project), "SDMetrics". https://docs.sdv.dev/sdmetrics
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.