Skip to content

Agent, workflow and domain-reasoning data

Configuration and customization data for realistic enterprise app replicas

Quick answer

An enterprise application replica only behaves like a real deployment if it carries a real tenant's configuration: custom objects and fields, picklist values, record types, validation and workflow rules, permission models and page layouts. Seed records fill the tables, but configuration decides which fields exist, which transitions are legal and who may act. For agent training, source configuration metadata from several real tenants of the same product, sanitized and licensed for environment building, and pair it with matching seed data.

By SourceX Editorial · Updated

Why default demo instances mislead agents

A vanilla sandbox teaches an agent the vendor's defaults, and almost no production tenant runs on defaults. Teams building digital twins of Salesforce, Workday and email clients emphasize preserving permissions and webhooks and seeding the replica from a customer's actual stack [1]. Research environments make the same bet: CRMArena models 16 common CRM object types with latent variables so its data behaves like a working org rather than a fixture [3], and tau-bench pairs databases and APIs with written domain policies the agent must follow [5].

The failure modes are concrete. An agent trained on a stock CRM will try to close an opportunity without the mandatory Loss_Reason__c a validation rule requires, pick a stage value that the tenant renamed, or edit a field that its profile can see but not write. In ServiceNow, it will set an incident state that a business rule immediately reverts, or skip a UI policy that makes a field mandatory only when category equals "Hardware." These errors are invisible on demo data and constant in production.

Fully self-hosted environments such as WebArena showed the value of functional, resettable sites across four domains [4]. Enterprise replicas extend that idea, but realism now depends on the configuration layer, not just the application code. For the record layer, see seed data and state snapshots for enterprise agent sandboxes.

Which configuration artifacts to source, by layer

Source configuration in five layers, because each one constrains a different part of the agent's action space. The table maps the layers to the native artifacts in three common platform families. Names reflect each platform's public metadata model; confirm exact type names against the supplier's version.

LayerWhat it controls for the agentSalesforce (Metadata API types)ServiceNow (tables)Jira / Atlassian
SchemaWhich objects and fields exist, data types, required flagsCustomObject, CustomField, RecordTypesys_db_object, sys_dictionaryCustom fields, issue types, field configurations
Value setsLegal values and dependenciesGlobalValueSet, picklist values, field dependenciessys_choiceField options, cascading select options
Rules and automationWhat happens after an action; what gets rejectedValidationRule, Flow, ApprovalProcess, AssignmentRulesBusiness rules (sys_script), UI policies, flowsWorkflows, transitions, conditions, validators, post functions
AccessWho can read, write or transitionProfile, PermissionSet, sharing rules, role hierarchyACLs (sys_security_acl), roles, groupsPermission schemes, project roles, issue security schemes
PresentationWhat the agent sees on screenLayout, FlexiPage (Lightning pages), compact layoutsForms, views, related listsScreens, screen schemes

The schema and value-set layers are the cheapest to replicate and the most commonly skipped. The rules layer matters most for agent outcomes because it defines rejections and side effects. The presentation layer matters for computer-use agents that read rendered screens; pair it with GUI grounding data from enterprise applications.

How tenant configuration is exported in practice

Most enterprise platforms already have a supported export path for configuration, and buyers should ask for that native format rather than screenshots or spreadsheets. In Salesforce, a Metadata API retrieve is driven by a package.xml manifest listing types and members; wildcards do not cover standard objects, custom fields on standard objects must be named explicitly, and managed-package components carry a namespace prefix, per Salesforce.s Metadata API documentation [7]. A manifest built only with * therefore silently misses Case.EngineeringReqNumber__c-style fields on standard objects, which are often the most business-specific part of an org.

In ServiceNow, customizations are captured as Customer Update records in sys_update_xml under an update set, with per-change history in sys_update_version. Completed update sets can be exported to XML and loaded into another instance. Update sets carry application files, not bulk data, so plan a separate path for seed records. An update set also captures only what changed while it was active, so a tenant's full configuration may span many sets or need a baseline comparison against a clean instance.

Lightweight replicas can work from less. The open-source VEI package builds a programmable replica of an operational stack from a text description or from real Slack, Gmail, Jira or Teams data [2]. That is useful for prototyping, but description-driven configuration reflects what someone remembers about a tenant, not what its rules actually enforce.

Diversity comes from many tenants of the same product

One customized tenant produces an agent that overfits to one company's field names and rules. Coverage comes from sourcing configuration from multiple tenants of the same product across industries, sizes and maturity levels, then holding some tenants out for evaluation. Treat tenant as the split key, the same way you would treat customer or site in other data, so evaluation measures generalization to unseen configurations rather than memorized picklists.

Useful diversity axes include the number of custom objects, depth of approval chains, use of record types, how many profiles versus permission sets are in use, managed packages installed, and whether automation runs through legacy workflow rules, process builders or newer flows. A tenant mid-migration, with both old and new automation active, is realistic and worth including. Document these axes in a dataset card or Data Card so downstream users know what the replica set represents [6].

Confidentiality and sanitization of configuration exports

Configuration exports encode a company's business logic and are commercially sensitive even when they contain no customer records. Approval thresholds, discount ceilings, commission fields, territory names, product codes and partner names reveal how a company sells and operates. Formula fields, Apex or script bodies, email templates and report filters can also embed personal data: a hard-coded manager's email in a flow, a named user in an assignment rule, or an account name in a sharing rule.

Sanitize in place without breaking references. Replace user, queue and group identifiers consistently across every file that references them, rename company-identifying labels while keeping API-name structure, and recompute any cross-references so a deployment still validates. Credentials such as named credentials, connected app secrets, remote site settings and integration endpoints should be stripped entirely. The same redaction discipline used for PII in screen recordings and computer-use trajectories applies to the layouts and labels those screens render.

Licensing needs explicit language for this use. Ask whether the license covers building environments, running agents against them, deriving tasks and publishing benchmark results; license terms for agent data covers the clauses to negotiate. Confirm the supplier owns or may license its own configuration, including whether any components belong to a third-party managed package.

A configuration request specification

A useful request names the product, the layers, the export format and the validation test the delivery must pass. The template below shows the fields a buyer would fill in.

Illustrative example: invented to show structure; it does not describe an available dataset.

request: tenant_configuration_for_agent_replica
product: "CRM platform (version or release named by supplier)"
tenants_requested: 8            # same product, different companies
holdout_tenants: 2              # reserved for evaluation
layers:
  schema: [custom_objects, custom_fields, record_types]
  value_sets: [picklists, global_value_sets, field_dependencies]
  rules: [validation_rules, flows, approval_processes, assignment_rules]
  access: [profiles, permission_sets, sharing_rules, role_hierarchy]
  presentation: [page_layouts, lightning_pages]
format: "native metadata export (e.g., package.xml retrieve or XML update set)"
exclude: [credentials, named_credentials, integration_endpoints, customer_records]
sanitization:
  user_ids: consistent_pseudonyms
  company_labels: renamed_keep_api_structure
  method_recorded: true
validation:
  deploys_cleanly_to_blank_instance: required
  rule_count_preserved: required      # same number of validation rules and flows
  reference_integrity_check: required
pairing: "seed records from same tenant, if available, under same license"
metadata_per_tenant: [industry, org_size_band, managed_packages, automation_generation]

The validation block is the part buyers most often omit. A configuration bundle that does not deploy to a blank instance, or that lost rules during sanitization, produces a replica that looks right and behaves wrong. Fold these fields into your broader agent data specification.

How configuration pairs with trajectories and evaluation

Configuration is the environment, and trajectories and evaluation tasks are what run inside it. A replica built from a tenant's configuration is most valuable when paired with workflow records from the same tenant, such as field-level audit trails or ticket histories reconstructed as trajectories, because replayed actions then hit the same rules that shaped them originally. State-based grading in agent evaluation task suites depends on those rules too: a task is only solvable if the replica enforces the same validation the human faced.

For the wider cluster of workflow, action and decision data, start at the AI agent training data hub, or see SourceX's overview of training data for RL environments and enterprise workflow datasets and agent trajectories. Teams ready to describe the configurations they need can send a request to SourceX.

Source tenant configuration for agent replicas

SourceX sources operational datasets from US companies on request, rather than from stock, and manages licensing and ongoing purchases. Each dataset is reviewed for ownership and consents, personal details are removed or replaced with the method recorded, and every release is approved by the supplying company, though a request does not guarantee a match. Describe the product, layers and tenant mix you need at SourceX for buyers.

Frequently asked questions

Is configuration metadata the same as seed data?

No. Configuration defines the shape and rules of the application, such as fields, picklists, validation, permissions and layouts. Seed data is the records stored inside that shape. A realistic replica needs both, and they should come from the same tenant so records satisfy the rules.

Can synthetic configuration replace real tenant exports?

Partly. Generated configuration can cover common patterns, but it reflects the generator's assumptions about how companies customize software. Real exports capture accumulated, inconsistent and legacy logic, which is exactly what makes production hard for agents.

What should a supplier strip before delivering a configuration export?

Credentials, integration endpoints, connected app secrets and any embedded personal data in formulas, scripts, templates and assignment rules. Company-identifying labels should be renamed consistently while preserving API structure so the bundle still deploys.

Sources

  1. TechCrunch, "Arga is building a better way to train enterprise AI agents" (2026). https://techcrunch.com/2026/08/26/arga-is-building-a-better-way-to-train-enterprise-ai-agents
  2. PyPI, "pyvei (VEI) package". https://pypi.org/project/pyvei/
  3. alphaXiv (overview of arXiv:2411.02305), "CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments" (2024). https://alphaxiv.org/overview/2411.02305v2
  4. arXiv (Zhou, Xu et al., Carnegie Mellon University), "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
  5. arXiv (Sierra Research), "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  6. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  7. Salesforce, "Metadata API Developer Guide". https://developer.salesforce.com/docs/atlas.en-us.api_meta.meta/api_meta/meta_intro.htm
  8. ServiceNow, "System update sets". https://docs.servicenow.com/bundle/vancouver-application-development/page/build/system-update-sets/reference/update-set-transfers.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data