Industry-specific operational data
RFQ and Quote Histories for AI Quoting Agents in Industrial Distribution
Quick answer
RFQ data for AI quoting is the pre-sale record of an industrial distributor or manufacturer: the inbound request as it arrived (email body, PDF, spreadsheet, drawing), the quote lines an inside-sales rep produced from it, every revision, and the outcome, meaning converted to an order, lost with a reason, or expired. Buy all four layers joined by a quote ID. Raw requests without the produced quote cannot train extraction, and quotes without outcomes cannot train pricing.
By SourceX Editorial · Updated
What a usable RFQ and quote history contains
A usable history links each inbound request to the quote document it produced and to the order or loss that followed. Vendors building RFQ agents describe the same pipeline: pull item descriptions, quantities, part numbers and materials from PDFs, BOMs and spreadsheets, map them to ERP SKUs, units of measure and tax codes, and route margin exceptions to a person [1]. Every one of those steps needs a historical example to learn from.
Ask for these layers, with stable keys between them:
- Inbound request. The original MIME message or portal submission, with attachments in native format (PDF, XLSX, CSV, DWG or STEP drawings, scanned BOMs), received timestamp and channel.
- Quote header. Quote number, customer token, branch, rep token, currency, Incoterms or freight terms, validity date and quote status from the ERP (for example, the quote or sales-quotation tables in systems such as Epicor Prophet 21, Infor CloudSuite Distribution or SAP SD).
- Quote lines. Requested text as written, resolved SKU, manufacturer part number, quantity, unit of measure, unit price, cost or margin field, lead time, stock or non-stock flag, and any alternate or substitute offered.
- Revisions. Each version of the quote with a timestamp, so you can see which lines changed and why.
- Outcome. Converted order number and conversion date, partial conversion by line, or a lost-reason code (price, lead time, no bid, customer cancelled) and expiry.
Post-sale orders are a different asset. Purchase-order records are covered on the purchase orders licensing page; this page is about the quote cycle that precedes them.
Why outcome labels decide what you can train
Outcome labels turn a quote archive into supervised data for pricing and prioritization. Quoting tools already suggest prices from historical sales, market trends and customer behavior to trade win probability against margin [2], and vendor playbooks frame slow or inconsistent RFQ response as margin leakage [3]. None of that works without knowing which quotes became orders.
Check three things before you trust the labels. First, the conversion link: many distributors convert a quote to an order in the ERP, which preserves the link, but reps also re-key orders by hand, which breaks it. Second, line-level outcomes: a customer often accepts six of ten lines, so a header-level "won" label overstates price acceptance on the four lost lines. Third, lost-reason coverage: if most lost quotes carry a blank reason or a default value, treat loss reasons as weak labels and rely on conversion only.
For CRM-side context (opportunity stages, rep activity, pipeline history) that can enrich quote outcomes, see sales CRM datasets with pipeline histories and outcomes.
Extraction and line resolution: what the training pairs look like
The core extraction pair is the raw RFQ as input and the rep's final quote lines as target output. This is harder than invoice extraction because requests are written by buyers, not generated by billing systems: free-text lines ("1/2 in 316 SS ball valve, NPT, qty 40"), competitor part numbers, superseded manufacturer numbers and quantities buried in email threads.
Public document benchmarks help as a baseline but do not cover this. DocILE, for example, provides 6.7k annotated business documents with key information extraction and line item recognition tasks [4]; it is useful for testing layout parsing, not for learning how a rep resolves "same as last order" into a SKU. For the invoice-side comparison, see invoice line-item extraction data.
Line resolution is where most production errors appear. Ask suppliers whether the history preserves the requested text next to the resolved SKU, and whether cross-reference tables (customer part number to internal SKU, competitor to equivalent) are available as of the quote date. Without the as-of crosswalk, a model learns mappings that did not exist when the rep quoted.
Agent trajectories and evaluation sets
An evaluation set for a quoting agent should score the whole task from the raw RFQ, not only field extraction. Agent benchmarks such as τ-bench evaluate agents that call programmatic APIs and talk to simulated users while following domain policies [5]; a quoting agent has the same shape, with ERP lookups for price, stock and lead time, and policies on minimum margin and approval thresholds.
Useful metrics to build into the holdout:
- Line-resolution accuracy: share of requested lines mapped to the SKU the rep chose or an accepted alternate.
- Quantity and UoM correctness: catches "40" read as each when the customer meant boxes of 40.
- Quote correctness: price within the policy band and lead time consistent with stock status at quote time.
- Clarification behavior: whether the agent asks the customer about ambiguous lines the rep also had to clarify (visible in the email thread).
- Escalation precision: margin exceptions sent to a human when, and only when, policy requires.
If you also capture rep activity in the quoting screen or spreadsheet, see spreadsheet agent task trajectories. Keep the evaluation split by time and by customer, so a model is never scored on a customer whose earlier quotes it trained on.
Illustrative record schema
A joined record per quote line is the easiest shape to audit and load. Deliver as JSON Lines, where each line is one valid JSON object in UTF-8 without a byte order mark [6], with native attachments in a separate object store referenced by hash.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"rfq_id": "R-000918", "received_at": "2025-03-14T15:02:11Z", "channel": "email",
"customer_token": "C-7f3a", "attachments": [{"sha256": "9b1e...", "type": "application/pdf", "role": "bom"}],
"quote_id": "Q-55120", "quote_version": 2, "line_no": 4,
"requested_text": "1/2in 316SS ball valve NPT qty 40",
"resolved_sku": "BV-316-050-NPT", "mfr_part_token": "M-22c1", "qty": 40, "uom": "EA",
"unit_price_scaled": 1.084, "lead_time_days": 5, "stock_flag": "stock",
"alternate_offered": false, "line_outcome": "won", "order_id_token": "O-81d0",
"lost_reason": null, "outcome_date": "2025-03-19"}
Ship a data dictionary with field definitions, code lists for status and lost reasons, and the date range. A machine-readable card, such as Croissant with its responsible-AI extension, can record preparation steps and labeling choices alongside the files [7].
Sensitivity: pricing, customers and drawings
Quote histories are among the most commercially sensitive records a distributor holds, so expect the supplying company to transform them before release. Common controls include tokenizing customer and rep identities, replacing contact details in email bodies, scaling or indexing prices (as in the unit_price_scaled field above), dropping cost and margin fields or banding them, and aging the data so the newest quotes are a year or more old. Each choice affects what you can learn: indexed prices still support win-probability modeling, while banded margins weaken margin optimization.
Customer drawings and specifications attached to RFQs are usually the customer's intellectual property, not the distributor's. Ask whether drawings are excluded, redacted to title blocks, or covered by rights the supplier can grant. Email threads also carry signatures, phone numbers and third-party names; the email archive licensing page covers that asset class more broadly.
Sourcing checklist for buyers
Use this checklist when you describe the data you need and when you review a supplier's sample.
| Question | Why it matters | Acceptable answer looks like |
|---|---|---|
| Is the raw inbound RFQ preserved with native attachments? | Extraction needs the input as the rep saw it | MIME or portal payloads plus original files, hashed |
| Are quote lines linked to the RFQ and to orders? | Without joins there are no training pairs or labels | Stable quote ID, version and order link per line |
| Are line-level outcomes available? | Header labels mislabel partial wins | Won, lost or expired per line, with dates |
| How are lost reasons captured? | Weak codes produce noisy labels | Code list with coverage rate disclosed |
| Are cross-reference tables available as of quote date? | Prevents leakage of later mappings | Dated crosswalk snapshots |
| How are prices and customers transformed? | Determines which models remain viable | Documented method and a checked sample |
| What happens to customer drawings? | Third-party IP | Excluded, redacted, or covered by stated rights |
| What date range and branches are covered? | Seasonality and regional pricing differ | Range, branch count and gaps stated |
Related procurement data has a different shape: if you are building a buyer-side agent that runs sourcing events, see sourcing-event and RFQ bid data. For other sector datasets, start at industry-specific operational data for AI or the AI data buyer guides hub.
How SourceX handles RFQ and quote history requests
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing and ongoing purchases. Nothing is held in stock, so a request for quote histories starts a search for US businesses that hold the data you describe, and a request does not guarantee a match. You describe the data, not the businesses; every release is approved by the supplying company.
The process runs Find, Assess (the data and its licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Industry context for distributors is on the industrial distribution buyers page, and you can describe your RFQ data requirement to SourceX.
Request RFQ and quote history data
SourceX sources operational datasets, including sales and quote histories, from US companies on request, for AI teams wherever they are based. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Tell SourceX what RFQ data you need.
Frequently asked questions
Can I train a quoting agent on quotes without the original RFQ?
You can train a pricing or win-probability model, but not an extraction or line-resolution model. Those need the request as the customer wrote it, paired with the lines the rep produced.
How much history is useful?
Enough to cover pricing changes, supplier price updates and seasonal demand for the product lines you target. Ask for the date range and branch coverage, and split evaluation data by time.
Do scaled prices still work for pricing models?
A consistent scaling factor or documented index preserves relative differences across customers, quantities and time, which is what win-probability models use. They do not support absolute margin targets.
Sources
- Cassidy, "AI RFQ Data Extraction Agent". https://www.cassidyai.com/solutions/ai-rfq-data-extraction-agent
- SoftwareFinder, "AIQuote". https://softwarefinder.com/sales-tools/aiquote
- HumCommerce, "Your RFQ Process Is Leaking Margin: A Playbook for AI-Driven Quoting". https://humcommerce.com/knowledge-center/your-rfq-process-is-leaking-margin-a-playbook-for-ai-driven-quoting/
- arXiv (Šimsa et al.), "DocILE Benchmark for Document Information Localization and Extraction" (2023). https://arxiv.org/pdf/2302.05658
- arXiv (Sierra Research), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- arXiv (MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.