Skip to content

Engineering and architecture

Is your firm's published work being used to train AI?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Your firm's published work may already be in AI training data if it was posted publicly: website renderings, award submissions, published drawings and press photos are the most exposed. Internal records such as RFIs, review comments and project financials are normally out of crawlers' reach. Set controls on the public layer, and license the private layer only by consent.

Key takeaways

  • Anything posted on the open web, including renderings and portfolio PDFs, may have been collected by crawlers.
  • Internal records in Deltek, project servers or Bluebeam Studio are out of reach of web scraping unless exposed by mistake.
  • Crawler rules, site terms and Content Credentials signal a firm's preferences but cannot guarantee anyone honors them.
  • Consent-based licensing differs from scraping: the firm chooses the records, approves the use and sets terms in a contract.

What parts of an architecture firm's work are publicly reachable?#

The publicly reachable part of an architecture firm's work is whatever sits on an open web page or in a public file: project pages, portfolio PDFs, renderings, award and competition boards, articles in design publications and social media posts. Planning and zoning submissions posted in public meeting packets can also include drawings.

Photography adds a wrinkle. Architectural photographers often hold copyright in their images and license them to the firm, so the photographer may have as much interest in how those images are reused as the firm does.

What parts of an architecture firm's work are publicly reachable?
MaterialWhere it is publicReachable by crawlersWho controls it
Renderings and project photosFirm website, social media, award sitesUsuallyThe firm, the photographer and sometimes the client
Portfolio and qualifications PDFsWebsite downloads, procurement portalsOftenThe firm
Published drawings and plansDesign publications, competition sitesOftenPublisher and firm under their agreement
Planning submissionsPublic meeting packets and permit portalsSometimesThe agency publishes it; copyright usually stays with the author
Internal project recordsERP, file servers, Bluebeam Studio, emailNo, unless exposed by a public link or misconfigured shareThe firm, subject to client contracts

What stays out of reach of scrapers?#

Internal project records stay out of reach of scrapers because they are not published on the open web. RFI logs, Bluebeam markup sessions, QA/QC comments, Deltek project budgets, staffing plans and internal email live behind logins on servers and cloud applications the firm controls.

That private layer is what makes consent-based licensing distinct. A crawler can collect a rendering, but it cannot see the review comments that shaped the design, the RFI that changed a detail or the staffing decision that kept a project on schedule.

The private layer can still leak in other ways, including a public share link or a misconfigured file server. Staff pasting client drawings into consumer AI tools, or software vendors whose terms allow training on customer content, are more realistic risks to internal records than scraping. An AI use policy, a review of vendor terms and periodic checks of sharing settings address these risks.

Can you tell whether your work was used in training?#

Usually a firm cannot confirm whether a specific rendering was used in training, because most model developers do not publish item-level lists of their training data. Research efforts have tried to document dataset sources: the Data Provenance Initiative, for example, audited 44 data collections spanning more than 1,800 fine-tuning text datasets and recorded their sources and licenses. Audits like that cover published research datasets, not every commercial model.

Some public image datasets can be searched by URL or caption, and some developers publish summaries of their data sources. Treat both as partial evidence. If the question matters for a dispute, counsel can advise on what can be learned and how to preserve it.

What controls can a firm put on its published work?#

A firm can put several controls on its published work, each of which signals a preference or adds friction rather than guaranteeing compliance. Used together, they make the firm's position clear and create a record that it asserted its rights.

Rules outside the United States can matter for firms whose work is published internationally. The European Union, for example, recognizes reservations of rights against text and data mining, and the way a reservation is expressed affects whether it counts. Assess any reliance on foreign rules with counsel.

  • Crawler rules: add robots.txt directives for AI training crawlers that publish their names, and recheck them as new crawlers appear.
  • Site terms: state that site content may not be used for AI training without permission.
  • Image choices: publish lower-resolution images and avoid posting full drawing sets or detailed sheets.
  • Content Credentials: attach provenance and a training preference to images where your tools support it.
  • Publisher agreements: ask design publications how they handle AI crawlers and reuse of submitted material.
  • Legal review: ask counsel about registering copyright in key works and about rules in other jurisdictions.

What can Content Credentials do for renderings and photos?#

Content Credentials can attach a signed provenance record to a rendering or photo and, through a separate assertion, state whether the firm allows AI training. The C2PA specification defines a Content Credential as a cryptographically bound structure that records an asset's provenance, including assertions about origin, modifications and use of AI.

The training preference is no longer part of the core C2PA standard. Version 2.0, released in January 2024, removed the training and data mining assertion, and C2PA's AI and ML guidance now points to the Creator Assertions Working Group's assertion, which can mark an asset as allowed, constrained or not allowed for AI use. Version 2.4, from April 2026, added an AI disclosure assertion.

The limits are practical. A credential states a preference; it does not stop a crawler that ignores it, and embedded metadata can be stripped when images are copied or re-uploaded. The C2PA explainer itself notes that Content Credentials show whether provenance data is intact and signed, not whether it is true.

Scraping and consent-based licensing differ on every point that matters to a firm: who decides, what is included, whether client rights are checked and whether the firm is compensated. The comparison explains why a firm can object to unpermitted scraping and still consider licensing its internal records.

How does scraping compare with consent-based licensing?
QuestionUnpermitted scrapingConsent-based licensing
Who decidesThe crawler operatorThe firm, at every step
What is includedWhatever is public, including client-sensitive imagesRecords the firm selects after a rights review
Client rightsNot consideredChecked against agreements before scope is set
Privacy preparationNonePersonal and confidential details removed
Permitted useUndefinedWritten into the license
CompensationNoneSet in the contract
DocumentationNoneProvenance and approvals recorded

Illustrative: an architecture studio audits its public footprint#

Illustrative: a fictional architecture studio known for cultural buildings reviews its website after a client asks whether project images are being used to train AI. The marketing lead finds old competition boards with detailed plans, drawing excerpts in qualification PDFs and photographers' images posted at full resolution.

The studio removes the detailed plans, republishes images at lower resolution, adds AI crawler rules and site terms, and attaches Content Credentials with a training preference to new renderings. It writes to the client explaining each step and its limits.

Separately, the principals discuss whether the studio's internal review records could be licensed on its own terms. They start with a metadata-only fit check, because that layer was never public and remains entirely under the studio's control.

How SourceX approaches published and private work#

SourceX works only with consent-based licensing of internal operating records; it does not collect public portfolios, and its own rights in a deidentified dataset are set out in the signed supplier agreement. The fit check gathers metadata about systems and record families, and nothing is shared during that initial assessment.

When a firm decides to proceed, it chooses the records and approves each stage of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The SourceX Evidence Packet then documents the provenance and licensing rights of what was included, the permitted use, the privacy record and the release authorization, which is precisely the record that scraping never produces.

Frequently asked questions

Can we take legal action if our renderings were used to train a model?

Possibly, but it is an unsettled and fact-specific area. Claims depend on who owns the copyright, where and how the work was copied, and how courts treat training use. Talk to counsel before acting, and preserve evidence of what was published and when.

Will blocking AI crawlers hurt our search visibility?

Not necessarily. Some AI developers publish separate crawler names or control tokens for training, so a firm can block those while still allowing search engines. Others do not separate the two. Review each directive before adding it and check search performance afterward.

Do client agreements limit what we publish about a project?

Often. Some owner agreements require consent before the firm publishes images or drawings, and secure or private projects may forbid publication entirely. Check the agreement before posting, and remember that your photographer's license may limit uses as well.

Should our AI policy cover how staff use client drawings?

Yes. Uploading client drawings or models to public AI tools is a more direct risk to confidential work than web scraping. A policy should name approved tools, prohibit uploading client material to unapproved ones and explain how AI-assisted work is reviewed.

Sources

  • The C2PA Explainer defines a Content Credential, also called a C2PA Manifest, as a cryptographically bound structure that records an asset's provenance, containing assertions about origin, modifications and use of AI; Content Credentials do not judge whether provenance data is true. Source
  • The C2PA 2.0 specification (January 2024) removed the Training and Data Mining assertion from the core standard; C2PA's AI/ML guidance now points to the Creator Assertions Working Group's training-and-data-mining assertion, which can mark an asset as allowed, constrained or not allowed for AI/ML use. Source
  • C2PA specification version 2.4 (April 2026) added an AI Disclosure Assertion (c2pa.ai-disclosure), among other changes. Source
  • The Data Provenance Initiative, a multi-disciplinary volunteer effort, released a first audit covering 44 data collections that span more than 1,800 fine-tuning text-to-text datasets, documenting their sources, licenses, creators and other metadata. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify