Skip to content

Multimodal and embodied data

Design Files Paired with Production Front-End Code for Design-to-Code Models

Quick answer

A design-to-code dataset pairs a visual design artifact (a design-tool frame, a mockup, or a rendered screenshot) with the front-end code that implemented it, ideally the code that actually shipped. Public benchmarks mostly pair web screenshots with the page's own HTML, which teaches reconstruction, not translation of design intent. Commercially useful data comes from companies that kept design files, component libraries, design tokens, pull requests and design-review threads linked, and that can grant rights to the design and the code together.

By SourceX Editorial · Updated

This guide covers which pair types matter, how to treat the gap between design and shipped code, where rights break, what to scrub, and how to evaluate. It sits in the multimodal data cluster and complements the code dataset buyer's map, which covers code on its own.

Pair types: which design-to-code records to specify

Specify the pair type first, because a model trained on screenshot-to-HTML pairs behaves differently from one trained on design-frame-to-component pairs. The four useful types differ in input modality, target code and how much intent they carry.

  • Design frame to component code. A frame or component from a design file (exported as structured JSON or SVG plus a PNG render) mapped to the React, Vue, SwiftUI, Jetpack Compose or Angular component that implements it. This is the highest-value pair because it carries layer names, auto-layout constraints and variant properties.
  • Rendered screenshot to page code. A screenshot of a shipped screen paired with its HTML/CSS or framework source at the same commit. Easy to generate at scale from a running app and a git history, but it loses the designer's intent.
  • Design tokens to theme code. Token files for color, spacing, typography and radius mapped to the Tailwind config, CSS custom properties or platform theme objects generated from them. Ask whether a supplier's tokens follow the W3C Design Tokens Community Group format or a proprietary tool export, and get the transform scripts that turn them into code.
  • Design revision to code diff. A change between two design versions paired with the pull request that implemented it. This is scarce and the most useful for editing-style tasks ("move the CTA below the fold, update the code").

Design-system repositories (Storybook stories, component docs, visual regression baselines) often hold all four types in one place, which makes a single well-maintained design system more valuable than thousands of loose screenshots.

Why public screenshot-to-code benchmarks are not enough

Public benchmarks measure reconstruction of existing web pages, not translation from a designer's file into a team's codebase conventions. Typical screenshot-to-HTML benchmarks such as Design2Code [8] take crawled web pages, strip external dependencies and ask models to reproduce each page's HTML/CSS from a screenshot. That is a sound test, but the targets are single-file static pages rather than component trees that import a design system.

Production front ends differ in ways that matter for training. Code is split across components, styled with utility classes or CSS-in-JS, bound to state and data fetching, and constrained by accessibility attributes and the company's own lint rules. A model trained only on reconstructed pages tends to inline styles, duplicate components and ignore existing tokens, which is exactly what reviewers in a real codebase reject.

The design-to-implementation delta is the signal

Shipped code rarely matches the design exactly, and recording the differences is where licensed data beats synthetic pairs. Engineers adjust spacing to the token scale, drop variants that were never built, add loading and error states, and fix contrast or focus order. Designers then comment on the implementation in review.

Ask suppliers to keep three things linked per pair: the design version the engineer worked from, the code commit that shipped, and any design-review or code-review comments between them. Review comments on code changes are an established training signal; CodeReviewer trained models to generate review comments and refine code from real change-comment pairs [1]. Design-review comments ("padding should be 16, not 12", "use the secondary button variant") add the same kind of supervision for visual fidelity. For reward modeling on these comments, see code preference and reward data from review outcomes.

Label each pair with a delta class so you can filter: exact match, token-normalized, scope-reduced (design features not built), scope-added (states not in the design), or diverged (later redesign). Diverged pairs are noise for generation but useful negatives for evaluation.

Illustrative pair record

A good record keeps the design artifact, the render, the code, and the review trail joinable by stable identifiers. The schema below shows the fields worth asking for; field names are generic, not tied to any design tool's API.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "pair_id": "dc-000412",
  "pair_type": "design_frame_to_component",
  "design": {
    "file_ref": "design/checkout.json",
    "frame_id": "frame-88:1203",
    "version_id": "v2024-03-14T10:22Z",
    "render_png": "renders/frame-88-1203@2x.png",
    "viewport": {"width": 390, "height": 844},
    "variants": ["state=default", "state=error"],
    "tokens_ref": "tokens/tokens.json"
  },
  "code": {
    "repo": "web-app",
    "commit_sha": "<40-char sha>",
    "paths": ["src/components/CheckoutSummary.tsx", "src/components/CheckoutSummary.test.tsx"],
    "framework": "react@18",
    "styling": "tailwind",
    "storybook_story": "Checkout/Summary/Error",
    "rendered_png": "renders/story-checkout-summary-error.png"
  },
  "delta": {
    "class": "token_normalized",
    "notes": "spacing 12 to 16 per token scale; error icon swapped to design-system variant"
  },
  "reviews": [
    {"source": "design_review", "comment": "Use secondary button variant", "resolved": true},
    {"source": "code_review", "comment": "Replace hex color with token", "resolved": true}
  ],
  "scrub": {"pii_method": "synthetic fixtures", "secrets_scan": "passed", "screenshot_text_review": "sampled"},
  "rights": {"design_owner": "client", "code_owner": "client", "authorization_ref": "auth-17"}
}

Publish the dataset-level description in a machine-readable format so your pipeline can load it: Croissant expresses dataset metadata, file resources and record structure as schema.org JSON-LD [3], and a Data Card captures upstream sources, collection decisions and intended use in prose [4]. For transfer formats and packaging, see dataset delivery formats and schemas.

Rights: design and code often have different owners

Design-to-code pairs frequently have two rightsholders, and a license from only one of them does not cover the pair. A product company usually owns both its design files and its code. Agencies and freelance studios are different: client contracts commonly assign deliverables to the client, so the agency may hold the files without the right to license them. Treat client authorization as a requirement for any agency-sourced pair, and see the software development agencies buyer page for that segment.

Check these layers separately:

  1. Design files. Who owns the file and any embedded assets (stock photography, licensed icon sets, commercial fonts). Font and icon licenses often prohibit redistribution even when the layout is owned.
  2. Code. First-party code versus vendored third-party packages under open-source licenses; for a full-history codebase sale, the proprietary codebases page covers what to ask.
  3. Screenshots of live products. These can include customer content, partner logos and user-generated text that the supplier does not own.
  4. Platform terms. If the design tool vendor itself trains on customer files, its own promises matter; the FTC has warned that breaking commitments not to use customer data for undisclosed purposes such as training may violate laws it enforces [7].

When a record combines several rightsholders, the logic in licensing multimodal records with several rightsholders applies. As of October 2026, California's AB 2013 requires developers of generative AI systems offered to Californians to post documentation about training data [6], so keep each pair's source and rights reference in your manifest.

Scrubbing screenshots, fixtures and code

Design-to-code data leaks personal data and secrets through places text-only pipelines miss. Screenshots of real products show customer names, emails, order numbers and avatars; design files often embed copied production data as placeholder content; test fixtures and seed files carry real records; and code holds API keys, internal hostnames and feature-flag names.

Run these passes before anything leaves the supplier:

  • Screenshots and renders. OCR-based detection of names, emails, phone numbers and account numbers, then replacement with synthetic values, re-rendered rather than blurred so layouts stay realistic. Faces in avatars need separate handling; the multimodal de-identification guide covers screens, faces and metadata together.
  • Design files. Text layers, image fills and comments, including resolved comment threads that may name employees or customers.
  • Code and history. Secret scanning across the full commit history, not just HEAD; enterprise repositories and document-sharing platforms routinely contain credentials that need detection and remediation [2]. Rotate anything found, because redacting a key in the delivered copy does not revoke it.
  • Fixtures and mocks. Replace real seed data with generated records that keep the same shapes and edge cases.

Ask for the method and a sampled check result per batch. No scrub is perfect, so plan your own spot checks on screenshot text.

Evaluating generated UI code

Evaluate by rendering the generated code and comparing it to the design both visually and structurally, then checking whether it would pass review. Public screenshot-to-code benchmarks typically combine automatic visual similarity metrics with human judgment; that is a reasonable baseline but not the whole picture for production code.

A practical evaluation harness scores each output on:

DimensionHow to measureWhat it catches
Visual fidelityRender at the design viewport with Playwright; compare with the design render using perceptual diff and block matchingWrong layout, missing elements, spacing drift
StructureCompare DOM or component trees; count reused design-system components versus re-implemented onesDuplicated components, inline styles
Token useFlag hard-coded hex values, pixel sizes and font stacks that should be tokensTheme breakage, dark-mode bugs
AccessibilityRun axe-core; check roles, labels, focus order and contrastUnlabeled controls, keyboard traps
Build and lintType-check, lint and run existing component testsCode that renders but cannot merge
ResponsivenessRender at several breakpoints and variantsFixed widths, broken states

Hold out a private evaluation slice from the same supplier and freeze it before training to avoid contamination; private evaluation sets for multimodal models explains how. If your model also acts in live interfaces, execution-based environments such as OSWorld test agents on real web and desktop applications [5], and computer-use demonstration collection covers the consent side.

Request checklist for suppliers

A specific request gets better matches than "UI data". Describe the data, not the companies you think hold it.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Pair types wanted, in priority order (frame to component, screenshot to page, tokens to theme, revision to diff)
  • Frameworks and styling (for example React with Tailwind, SwiftUI, Compose) and minimum share of design-system reuse
  • Platforms and viewports (web desktop, mobile web, iOS, Android)
  • Linkage required: design version ID, commit SHA, PR and review comments
  • Delta labels and whether diverged pairs are included
  • Design-system assets: tokens format, Storybook, visual regression baselines
  • Rights: owner of design and code per pair, client authorization for agency work, third-party fonts, icons and images excluded or licensed
  • Scrubbing: screenshot text, design-file placeholders, fixtures, full-history secret scan
  • Documentation: Croissant metadata, data card, per-pair rights reference
  • Evaluation holdout size and freeze date

For a general template, start from the multimodal dataset specification template. Related pairing problems in engineering design are covered in text-to-CAD training data.

How SourceX sources design and code pairs

SourceX sources operational datasets from US companies on request, including engineering records and documents, and manages the licensing process. Nothing is held in stock, so a request for design-to-code pairs is a search among US businesses that hold the described data, and it does not guarantee a match. Every release is approved by the supplying company, and each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. You can describe the pairs you need on the buyers page, or browse source code licensing for AI training and the AI data hub.

Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. SourceX does not train models and does not source scraped web content.

Request design-to-code training data

SourceX finds US companies that hold the design and code records you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything transfers. Nothing is contracted until a supplier agrees. Describe your pair types, frameworks and rights requirements at https://sourcex.si/buyers.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Frequently asked questions

Can I build design-to-code data by screenshotting public websites?

You can build reconstruction pairs that way, which is how several public benchmarks work, but you get no design files, no tokens and no review history, and the rights to the pages are not yours. That data teaches copying rendered pages, not implementing a design system.

How many pairs come from one design system?

It depends on component count, variants and history depth, so ask suppliers for counts by pair type and delta class rather than total screenshots. A design system with long revision history yields many revision-to-diff pairs from a modest number of components.

Should diverged pairs be excluded?

Exclude them from generation training but keep them, labeled, for evaluation and for training models that detect design-implementation drift.

Sources

  1. Li et al. (Microsoft Research Asia and collaborators), "CodeReviewer: Pre-Training for Automating Code Review Activities" (2022). https://arxiv.org/pdf/2203.09095
  2. arXiv, "Using AI/ML to Find and Remediate Enterprise Secrets in Code & Document Sharing Platforms" (2024). https://arxiv.org/html/2401.01754v1
  3. Akhtar et al. (MLCommons), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  4. Pushkarna, Zaldivar, Kjartansson (Google Research), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  5. Xie et al. (XLANG Lab, HKU and collaborators), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
  6. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  7. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  8. arXiv (Si et al.), "Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering" (2024). https://arxiv.org/pdf/2403.03163

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data