Document AI data
Chart Understanding Data: Business Charts Paired with Their Source Tables
Quick answer
Chart understanding data pairs a chart image with the exact table it was drawn from, plus questions whose answers are computable from that table. Most large public sets are synthetic or plotted from web tables, so buyers improving document VLMs increasingly look for real business charts, from reports, decks and dashboards, where the source workbook still exists. The strongest labels come from native Office files, whose chart parts carry the plotted values; scanned or flattened charts need human digitization and tolerance-based scoring.
By SourceX Editorial · Updated
What counts as ground truth for a chart
Ground truth for a chart is the data table the renderer actually plotted, not a table someone reconstructed from pixels. When a chart is drawn from a known table, every bar height, series label and axis tick has a verifiable answer, which makes chart-to-table extraction and chart QA scoreable without annotator guesswork. Reconstructed tables, read off the image by a person or a digitizer tool, inherit reading error and should be labeled as a weaker tier.
Practically, a record should carry four aligned layers: the rendered image (PNG at the original export resolution, or the page crop from a PDF), the source table with series and category names, chart metadata (type, axis scales, units, stacking), and question-answer pairs grounded in the table. This is the same alignment idea that recent research sets follow; IBM's ChartNet card, for example, describes samples that tie the image to underlying data and text [1]. For page-level context around the chart, see our guide to document question answering data.
Where public chart sets fall short for business use
Public chart sets are mostly synthetic, which limits how well they represent the charts a document model meets in finance packs, board decks and operations reports. As of October 2026, the ChartNet dataset card describes roughly 1.7 million samples, mostly synthetic with a smaller real-world subset [1]. ChartGalaxy reaches million scale by generating infographic charts from templates mined from real infographics [2]. Older chart QA benchmarks such as ChartQA draw on charts published on the web rather than charts inside corporate documents, so check the domain mix of any set you adopt.
Synthetic generation gives clean labels and wide type coverage, but it tends to miss the failure modes that break production models:
- Corporate styling: brand palettes with near-identical series colors, legends placed inside the plot area, and data labels that overlap.
- Combo and dual-axis charts: a column series on the primary axis with a line on a secondary axis, where the model must bind each value to the right scale.
- Annotations and callouts: "Q3 includes one-time charge" text boxes, target lines and shaded forecast regions.
- Degraded capture: charts printed, scanned, faxed or photographed on a projector, covered in our page on degraded document images.
- Unit and scale traps: values in thousands with "($000s)" in the title, log axes, truncated y-axes and broken axes.
Before treating any public set as a ChartQA alternative, check the domain mix, how charts were rendered, and the license on both images and tables.
Extracting labels from native Office and PDF files
Native Office files are the richest source of chart labels because the chart part stores the plotted series alongside the drawing. In an Office Open XML deck or workbook, each chart lives in a part such as ppt/charts/chart1.xml or xl/charts/chart1.xml; each series (c:ser) holds a category reference (c:cat) and a value reference (c:val), and those references usually carry a cached copy of the values (c:numCache, c:strCache) next to the cell formula (c:f). Decks typically also embed the backing workbook under ppt/embeddings/, so the full table, including rows the chart filtered out, is often recoverable.
Three checks keep these labels honest:
- Cache freshness. The cached values reflect the last time the chart was refreshed; compare cache against the embedded workbook and flag mismatches rather than silently choosing one.
- Hidden and filtered data. Hidden rows, filtered categories and "show data in hidden rows" settings change what is plotted; the label should be the plotted subset, with the full table stored separately.
- Display transforms. Number formats, percent-of-total stacking and axis min/max settings change what the viewer reads; store both raw values and formatted display strings.
PDF exports usually lose the chart part, leaving vector paths or a raster image. Vector paths can sometimes be mapped back to values using axis tick positions, but treat that as reconstructed ground truth. Linked charts that point to an external workbook path are common in finance decks; if the link is broken, only the cache survives. For the workbook side of these files, see spreadsheet and financial model datasets; for full decks used in slide generation, see presentation deck datasets.
A record schema for chart-to-table and chart QA
A usable record ties each chart image to a table, chart metadata, provenance of the label and question-answer pairs with answer types. Keeping provenance at the record level lets you weight native-cache labels differently from human-digitized ones during training and evaluation.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"chart_id": "deck-0142-slide-07-chart-1",
"image": "images/deck-0142-s07-c1.png",
"source_format": "pptx",
"label_provenance": "ooxml_numcache_verified_against_embedding",
"chart_meta": {
"type": "combo_clustered_column_line",
"axes": [
{"id": "y1", "title": "Revenue", "unit": "USD thousands", "scale": "linear"},
{"id": "y2", "title": "Gross margin", "unit": "percent", "scale": "linear"}
],
"annotations": ["FY24 includes acquisition"],
"masking": "values perturbed by agreed multiplier; ratios preserved"
},
"table": {
"columns": ["Quarter", "Revenue (y1)", "Gross margin (y2)"],
"rows": [["Q1", 4120, 41.2], ["Q2", 4385, 42.0], ["Q3", 3990, 39.8], ["Q4", 4710, 43.1]]
},
"qa": [
{"q": "Which quarter had the lowest gross margin?", "a": "Q3", "type": "categorical"},
{"q": "By how much did revenue change from Q3 to Q4?", "a": 720, "unit": "USD thousands", "type": "numeric_computed"}
]
}
Linearize the table the way your target model expects; plot-to-table approaches such as DePlot translate a chart into a linearized table that a language model then reasons over. If you also need cell-level structure from report tables, pair this with table structure recognition data.
Scoring chart extraction and chart QA
Score numeric outputs with tolerance-based metrics, not exact string match, because rounding and display formats make exact match punish correct readings. The common convention in chart QA is relaxed accuracy, which accepts numeric answers within a small relative margin of the gold value while requiring exact match for text. Newer benchmarks tighten this for years and use edit-distance scores such as ANLS for text answers, and plot-to-table work scores table outputs with cell-matching metrics rather than whole-string comparison.
For an evaluation set built from business charts, decide in advance:
| Decision | Options | Why it matters |
|---|---|---|
| Numeric tolerance | 1%, 5%, or absolute per unit | 5% is lenient on margins in percent; a 40.0 vs 42.0 margin reading passes |
| Year and ID handling | Exact match | Prevents "2024" vs "2023" passing under relative tolerance |
| Table matching | Row/column alignment before cell scoring | Swapped series or transposed tables otherwise score zero or falsely high |
| Unit normalization | Score in source units, or require unit in answer | "4.1M" vs "4,120 (USD thousands)" are the same value |
| Masked values | Score against masked table only | Mixing masked and true values corrupts both training and eval |
Keep a held-out split drawn from different suppliers or document families than the training data; charts from one company's template leak style cues across splits.
Confidentiality, masking and what it does to labels
Business charts often show confidential figures, so suppliers may require masking or perturbation, and that changes the value-level labels you train on. Common approaches include scaling all series by an agreed multiplier (preserving ratios, trends and rankings), replacing category names such as customer or product names with stable pseudonyms, and redacting titles. Each choice must be applied to the image and the table together; re-rendering the chart from the masked table is cleaner than editing pixels.
Record the masking method per record, as in the schema above, because it determines which questions stay valid. Ratio and ranking questions survive multiplicative scaling; absolute-value questions only survive if answered against the masked table. Charts can also carry personal data, such as named sales reps on a leaderboard or patient counts by clinic. NIST SP 800-188 discusses the limits of traditional de-identification relative to formal privacy methods [3], and charts built from protected health information need HIPAA de-identification by Safe Harbor or Expert Determination before release [4].
Specifying a request for real business charts
A precise request describes the charts and labels you need, not the companies you want them from. Useful fields: chart types and target mix (including combo, dual-axis, small multiples and scanned charts), source formats (PPTX, XLSX, DOCX, PDF), label provenance tiers you accept, masking you can tolerate, QA types (lookup, comparison, computed, trend), volume ranges, and the intended use (SFT, evaluation or both).
SourceX sources operational datasets from US companies on request, including documents and finance workflows, and manages the licensing process; categories are not inventory, and a request does not guarantee a match. Every release is approved by the supplying company, each dataset is rights-reviewed for ownership and consents, and personal details such as names and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, though no method is perfect. You can describe the charts you need on the SourceX buyer page. For the wider landscape of document tasks, start at the Document AI data hub or the broader multimodal training data guide.
Sourcing business charts paired with their source tables
SourceX looks for US businesses holding the chart-bearing documents and workbooks you describe, then runs Find, Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees. Licensed datasets are delivered under a license defining records, uses, term and delivery, through private, access-controlled workflows. Describe your chart understanding data requirements.
Sources
- IBM Granite on Hugging Face, "ChartNet dataset card (README.md)". https://huggingface.co/datasets/ibm-granite/ChartNet/blob/main/README.md
- arXiv, "ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation" (2025). https://arxiv.org/html/2505.18668v5
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.