Engineering and architecture
How to de-identify RFIs and submittal logs before review
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
To de-identify RFI logs and submittal logs, decide a treatment for every field before touching the text: keep technical fields such as spec section and review action, pseudonymize people and companies consistently, generalize sites and dates, and drop contact details and attachments. Then scan the question and response text, because most identifying details hide there.
Key takeaways
- Assign every log field one treatment, keep, generalize, pseudonymize or drop, and write the decision down before processing.
- Consistent pseudonyms preserve who answered whom, which is the part of the record AI developers find useful.
- Free-text questions, responses and reviewer remarks carry more identifying detail than the structured columns do.
- Marked-up PDFs, photos and attachment file names need their own review or should stay out of scope.
- Automated detection finds most names and numbers, but a person who knows the projects should check a sample.
What makes RFI and submittal logs hard to de-identify?#
RFI and submittal logs are hard to de-identify because they mix clean structured columns with free text written by people who assumed only the project team would read it. A typical Procore or Newforma export holds an RFI number, subject, question, proposed solution, official response, responder, ball-in-court, due dates, spec section and drawing references, plus distribution lists and attachments.
Submittal logs add subcontractor and manufacturer fields, submittal type such as shop drawings, product data or samples, the review action, the reviewer and every resubmittal round. Each of those columns identifies a company, a person or a site in a different way, so one blanket rule such as deleting all names breaks the record or leaves gaps.
Field-by-field treatment for RFI and submittal logs#
A field treatment table is the single document that makes de-identification repeatable across projects and reviewers. Agree it with your project managers and counsel once, then apply it to every export rather than deciding project by project.
The defaults below suit an internal review before a licensing conversation. A specific buyer request or a client contract can make any row stricter, never looser.
| Field | Treatment | Reason |
|---|---|---|
| RFI and submittal numbers | Keep, with project prefix replaced | Sequence shows how questions build during construction |
| Project name and number | Pseudonymize with one token per project | Lets records from the same job stay linked |
| Site address and owner name | Generalize to region and building type | Location and client identity add little technical value |
| Architect, contractor and consultant staff | Pseudonymize to role tokens such as Architect PM 1 | Keeps the conversation readable without naming people |
| Subcontractor and supplier companies | Pseudonymize consistently | Firm names can reveal the project and commercial terms |
| Manufacturer and product names | Generalize to product category unless rights and need are confirmed | Product detail can be useful but may be commercially sensitive |
| Spec section, sheet and detail references | Keep | These are the technical anchors a model learns from |
| Review action and resubmittal round | Keep | Approved as noted or revise and resubmit is the outcome signal |
| Dates | Shift per project, keeping intervals | Turnaround time stays measurable while the calendar disappears |
| Emails, phone numbers, signatures | Drop | No analytical value and direct identifiers |
| Cost and schedule impact amounts | Generalize to an impact flag | Figures can expose fees and contract values |
| Attachments and markups | Drop or route to a separate image review | File names and title blocks repeat identifiers |
Free-text scan checklist#
The free-text scan covers the question, proposed solution, response and reviewer remarks, which is where names, site details and frustration tend to surface. Structured transformations do nothing for a sentence like a superintendent explaining what the owner's facility manager said on site.
Run an automated pass first, then have someone who knows the projects read a sample from every project. Automated tools miss project nicknames, tenant names and building-specific room tags, and only a project manager will recognize them.
- Greetings, sign-offs and signature blocks that name people or include direct phone lines.
- References to individuals inside sentences, such as per the call with a named engineer.
- Client, tenant and project names, including abbreviations and nicknames the team used.
- Street names, cross streets, parcel references and neighboring property owners.
- Room tags or occupant names that reveal who will use the building.
- Security-sensitive content: access control layouts, camera positions, secure rooms and vault details.
- Dollar figures, fee references and back-charge disputes.
- Claims language, delay notices or references to counsel, which go to legal review.
- Remarks about a person's competence or conduct, which should be removed entirely.
Where automated tools help and where they stop#
Automated PII detection is useful for the first pass over thousands of RFI responses, but it does not replace a reviewer. Presidio, for example, is an open-source, MIT-licensed SDK that combines named-entity recognition, regular expressions, rule-based logic and checksums to find personal data in text and images.
Presidio's own documentation warns that automated detection cannot guarantee it finds all sensitive information and that additional protections should be used. That warning applies to every tool. Construction logs are full of proper nouns that look like products, places or people, so build a project glossary of client, tenant and site names and run it alongside any detector.
In what order should the work happen?#
The work order for de-identifying project logs matters because each step depends on the one before it. Exporting before scoping wastes effort, and scanning text before the glossary exists misses the names that matter most.
- Confirm which projects are in scope after a rights review, and exclude secure, healthcare and client-restricted work.
- Export full logs with every column, not filtered report views, and record the export date and system.
- Build a glossary of client, tenant, site and staff names from your project accounting system and team lists.
- Apply the field treatment table and create one restricted mapping table per project.
- Run the automated scan and the glossary match over all free-text fields.
- Have a project manager read a sample from every project and log each correction.
- Store the final set, the treatment table and the review log together for approval.
Illustrative: an architecture firm prepares a review set#
Illustrative: a fictional architecture firm with offices in two states keeps RFIs in Procore for recent projects and in Newforma for older ones. Its COO wants a sample set of RFI and submittal logs ready for an internal licensing review, without exposing any client.
The firm drafts a treatment table, exports logs from a handful of completed commercial and higher-education projects, and builds a glossary of owner, tenant and site names from its Deltek project list. Healthcare and secure-facility projects are excluded at the start because their room data is too sensitive. A project architect reviews a sample from each project and finds tenant names in room tags that the automated pass missed. The firm adds them to the glossary, reruns the transformations and keeps attachments out of scope for now.
Mistakes that undo the work#
The most damaging mistake is pseudonymizing inconsistently, so the same contractor appears under three tokens and the thread of who asked and who answered is lost. Use one mapping table per project, store it separately with restricted access, and never ship it with the data.
Other common errors include leaving original file names on attachments, shifting dates by different amounts within one project, and treating de-identification as a rights decision. Removing names does not settle whether a client contract allows the records to be licensed; that is a separate review.
How SourceX handles project logs#
SourceX treats de-identification as the Preparation step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Rights review comes first, so effort is spent only on projects that can be licensed, and the supplier approves the transformed sample before anything leaves its control.
The treatment table, glossary approach and reviewer sign-off become the privacy record in the SourceX Evidence Packet, alongside provenance, licensing rights, permitted use and release authorization. Nothing is shared during the initial fit check, which collects only metadata such as systems, years of history and record families.
Frequently asked questions
Should we de-identify logs before the first conversation about licensing?
No. The first conversation needs only metadata: which systems hold RFIs and submittals, how many years are accessible and which project types are involved. De-identification is real work, so it should start after a rights review confirms which projects are in scope and a buyer's requirements are known.
Can we keep consultant firm names if they agree?
Possibly, with written permission, but there is usually little reason to. Consultant identity rarely changes what a model learns from the exchange, while it can point back to a specific project. Role tokens such as structural consultant or MEP consultant keep the useful context.
What about Bluebeam markups and marked-up shop drawings?
Treat them as a separate image review. Title blocks, stamps, signatures, client logos and file names all repeat identifiers, and text in images is easy to miss. Many firms leave markups out of a first package and add them only when a buyer specifically needs visual review history.
How do we show a buyer the logs were de-identified properly?
Keep a written record: the field treatment table, the tools and glossary used, the sampling method, who reviewed each project and what they changed. That record becomes part of the privacy documentation for the package and lets your own counsel verify the work later.
Do shifted dates still let a buyer measure turnaround time?
Yes, if every date in a project moves by the same offset. The interval between an RFI being issued and answered, or a submittal being received and returned, stays intact. Shifting dates differently within one project would destroy that signal.
Sources
- Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization in text and images, combining named-entity recognition, regular expressions, rule-based logic and checksums with context. Source
- Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.