Skip to content

Engineering and architecture

How to de-identify RFIs and submittal logs before review

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

To de-identify RFI logs and submittal logs, decide a treatment for every field before touching the text: keep technical fields such as spec section and review action, pseudonymize people and companies consistently, generalize sites and dates, and drop contact details and attachments. Then scan the question and response text, because most identifying details hide there.

Key takeaways

  • Assign every log field one treatment, keep, generalize, pseudonymize or drop, and write the decision down before processing.
  • Consistent pseudonyms preserve who answered whom, which is the part of the record AI developers find useful.
  • Free-text questions, responses and reviewer remarks carry more identifying detail than the structured columns do.
  • Marked-up PDFs, photos and attachment file names need their own review or should stay out of scope.
  • Automated detection finds most names and numbers, but a person who knows the projects should check a sample.

What makes RFI and submittal logs hard to de-identify?#

RFI and submittal logs are hard to de-identify because they mix clean structured columns with free text written by people who assumed only the project team would read it. A typical Procore or Newforma export holds an RFI number, subject, question, proposed solution, official response, responder, ball-in-court, due dates, spec section and drawing references, plus distribution lists and attachments.

Submittal logs add subcontractor and manufacturer fields, submittal type such as shop drawings, product data or samples, the review action, the reviewer and every resubmittal round. Each of those columns identifies a company, a person or a site in a different way, so one blanket rule such as deleting all names breaks the record or leaves gaps.

Field-by-field treatment for RFI and submittal logs#

A field treatment table is the single document that makes de-identification repeatable across projects and reviewers. Agree it with your project managers and counsel once, then apply it to every export rather than deciding project by project.

The defaults below suit an internal review before a licensing conversation. A specific buyer request or a client contract can make any row stricter, never looser.

Field-by-field treatment for RFI and submittal logs
FieldTreatmentReason
RFI and submittal numbersKeep, with project prefix replacedSequence shows how questions build during construction
Project name and numberPseudonymize with one token per projectLets records from the same job stay linked
Site address and owner nameGeneralize to region and building typeLocation and client identity add little technical value
Architect, contractor and consultant staffPseudonymize to role tokens such as Architect PM 1Keeps the conversation readable without naming people
Subcontractor and supplier companiesPseudonymize consistentlyFirm names can reveal the project and commercial terms
Manufacturer and product namesGeneralize to product category unless rights and need are confirmedProduct detail can be useful but may be commercially sensitive
Spec section, sheet and detail referencesKeepThese are the technical anchors a model learns from
Review action and resubmittal roundKeepApproved as noted or revise and resubmit is the outcome signal
DatesShift per project, keeping intervalsTurnaround time stays measurable while the calendar disappears
Emails, phone numbers, signaturesDropNo analytical value and direct identifiers
Cost and schedule impact amountsGeneralize to an impact flagFigures can expose fees and contract values
Attachments and markupsDrop or route to a separate image reviewFile names and title blocks repeat identifiers

Free-text scan checklist#

The free-text scan covers the question, proposed solution, response and reviewer remarks, which is where names, site details and frustration tend to surface. Structured transformations do nothing for a sentence like a superintendent explaining what the owner's facility manager said on site.

Run an automated pass first, then have someone who knows the projects read a sample from every project. Automated tools miss project nicknames, tenant names and building-specific room tags, and only a project manager will recognize them.

  • Greetings, sign-offs and signature blocks that name people or include direct phone lines.
  • References to individuals inside sentences, such as per the call with a named engineer.
  • Client, tenant and project names, including abbreviations and nicknames the team used.
  • Street names, cross streets, parcel references and neighboring property owners.
  • Room tags or occupant names that reveal who will use the building.
  • Security-sensitive content: access control layouts, camera positions, secure rooms and vault details.
  • Dollar figures, fee references and back-charge disputes.
  • Claims language, delay notices or references to counsel, which go to legal review.
  • Remarks about a person's competence or conduct, which should be removed entirely.

Where automated tools help and where they stop#

Automated PII detection is useful for the first pass over thousands of RFI responses, but it does not replace a reviewer. Presidio, for example, is an open-source, MIT-licensed SDK that combines named-entity recognition, regular expressions, rule-based logic and checksums to find personal data in text and images.

Presidio's own documentation warns that automated detection cannot guarantee it finds all sensitive information and that additional protections should be used. That warning applies to every tool. Construction logs are full of proper nouns that look like products, places or people, so build a project glossary of client, tenant and site names and run it alongside any detector.

In what order should the work happen?#

The work order for de-identifying project logs matters because each step depends on the one before it. Exporting before scoping wastes effort, and scanning text before the glossary exists misses the names that matter most.

  • Confirm which projects are in scope after a rights review, and exclude secure, healthcare and client-restricted work.
  • Export full logs with every column, not filtered report views, and record the export date and system.
  • Build a glossary of client, tenant, site and staff names from your project accounting system and team lists.
  • Apply the field treatment table and create one restricted mapping table per project.
  • Run the automated scan and the glossary match over all free-text fields.
  • Have a project manager read a sample from every project and log each correction.
  • Store the final set, the treatment table and the review log together for approval.

Illustrative: an architecture firm prepares a review set#

Illustrative: a fictional architecture firm with offices in two states keeps RFIs in Procore for recent projects and in Newforma for older ones. Its COO wants a sample set of RFI and submittal logs ready for an internal licensing review, without exposing any client.

The firm drafts a treatment table, exports logs from a handful of completed commercial and higher-education projects, and builds a glossary of owner, tenant and site names from its Deltek project list. Healthcare and secure-facility projects are excluded at the start because their room data is too sensitive. A project architect reviews a sample from each project and finds tenant names in room tags that the automated pass missed. The firm adds them to the glossary, reruns the transformations and keeps attachments out of scope for now.

Mistakes that undo the work#

The most damaging mistake is pseudonymizing inconsistently, so the same contractor appears under three tokens and the thread of who asked and who answered is lost. Use one mapping table per project, store it separately with restricted access, and never ship it with the data.

Other common errors include leaving original file names on attachments, shifting dates by different amounts within one project, and treating de-identification as a rights decision. Removing names does not settle whether a client contract allows the records to be licensed; that is a separate review.

How SourceX handles project logs#

SourceX treats de-identification as the Preparation step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Rights review comes first, so effort is spent only on projects that can be licensed, and the supplier approves the transformed sample before anything leaves its control.

The treatment table, glossary approach and reviewer sign-off become the privacy record in the SourceX Evidence Packet, alongside provenance, licensing rights, permitted use and release authorization. Nothing is shared during the initial fit check, which collects only metadata such as systems, years of history and record families.

Frequently asked questions

Should we de-identify logs before the first conversation about licensing?

No. The first conversation needs only metadata: which systems hold RFIs and submittals, how many years are accessible and which project types are involved. De-identification is real work, so it should start after a rights review confirms which projects are in scope and a buyer's requirements are known.

Can we keep consultant firm names if they agree?

Possibly, with written permission, but there is usually little reason to. Consultant identity rarely changes what a model learns from the exchange, while it can point back to a specific project. Role tokens such as structural consultant or MEP consultant keep the useful context.

What about Bluebeam markups and marked-up shop drawings?

Treat them as a separate image review. Title blocks, stamps, signatures, client logos and file names all repeat identifiers, and text in images is easy to miss. Many firms leave markups out of a first package and add them only when a buyer specifically needs visual review history.

How do we show a buyer the logs were de-identified properly?

Keep a written record: the field treatment table, the tools and glossary used, the sampling method, who reviewed each project and what they changed. That record becomes part of the privacy documentation for the package and lets your own counsel verify the work later.

Do shifted dates still let a buyer measure turnaround time?

Yes, if every date in a project moves by the same offset. The interval between an RFI being issued and answered, or a submittal being received and returned, stays intact. Shifting dates differently within one project would destroy that signal.

Sources

  • Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization in text and images, combining named-entity recognition, regular expressions, rule-based logic and checksums with context. Source
  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify