Skip to content

Multimodal and embodied data

Checking Image-Text Alignment in Multimodal Training Data

Quick answer

Image-text alignment quality is the share of pairs whose text accurately describes what the image shows. Measure it in two layers: an embedding-similarity score such as CLIP cosine similarity to triage the full set, then a human-audited stratified sample graded on a fully, partly or not-aligned rubric as the acceptance metric. Report both per slice, keep context-only text tagged rather than discarded, and treat machine recaptioning as a tradeoff, not a fix.

By SourceX Editorial · Updated

What "aligned" should mean in your acceptance spec

An image-text pair is aligned when every factual claim in the text is visually verifiable in the image, or is explicitly tagged as context the image cannot show. Write that definition into the spec before you score anything, because a vague definition lets a supplier and your team grade the same sample differently. Distinguish three failure types: wrong pairing (the caption belongs to another image), partial description (true but missing the subject), and hallucinated detail (counts, colors or text that are not present).

Business-sourced pairs add a fourth category that web alt-text rarely has: operational context. A field technician's note saying "replaced gasket, customer reported leak since March" describes the job, not the photo. That text is valuable for grounding and retrieval, so tag it context instead of filtering it out as misaligned. See how this plays out across support tickets with attachments in support tickets with screenshots.

How CLIP score filtering works, and where it fails

CLIP score filtering keeps pairs whose image and text embeddings exceed a cosine-similarity threshold, and it is the common first-pass filter for web-scale image-text data. Public pipelines in the LAION and DataComp lineage set the cutoff either as an absolute cosine value or as a percentile of the pool, scored with a specific encoder such as ViT-B/32 or ViT-L/14. The encoder matters: the same pair gets different scores from different models, so a threshold is only meaningful together with the model that produced it.

Embedding similarity measures compatibility between an image and a short web-style caption. It does not measure whether a detailed, domain-specific description is correct. The common blind spots are:

  • Counting and spatial relations. "Three valves, the left one open" scores close to "two valves" because contrastive embeddings encode object presence better than quantity or position.
  • Text inside the image. Screenshots, labels, gauges and forms carry meaning in rendered text that a general CLIP model reads poorly; a caption quoting an error code may score low even when correct.
  • Domain jargon. Part numbers, clinical shorthand and trade terms sit outside the web vocabulary the encoder learned, so correct expert captions are penalized.
  • Long captions. CLIP's text encoder truncates at a short token limit, so the tail of a detailed description is never scored.
  • Distribution skew. Any threshold removes some subjects, image styles and dialects more than others. A cutoff is a sampling decision with distribution consequences, so compare slice proportions before and after filtering.

Use the score to rank and to find obvious mismatches, not as a pass or fail gate. Calibrate any threshold against human labels from your own domain; a cutoff borrowed from a web-scale paper says nothing about engineering photos or scanned forms.

Complementary automated checks for caption-image mismatch

Pair the embedding score with checks aimed at the specific blind spots above. Each one is cheap enough to run on the full set before human review.

  • OCR agreement. Run OCR on the image and compare any quoted strings in the text using a soft string metric. ANLS, defined for scene-text VQA, was designed for exactly this case, where OCR imperfections make exact-match accuracy too harsh [2].
  • Detector-count agreement. Where captions state counts or named objects, compare them with an open-vocabulary detector's output and flag disagreements rather than auto-rejecting.
  • Duplicate and swap detection. Perceptual hashes (pHash, dHash) on images and near-duplicate hashing on text expose one caption attached to many images, the most common wrong-pairing pattern from template-driven systems.
  • Boilerplate detection. Flag captions that are filenames (IMG_4471.jpg), timestamps, or form defaults such as "see attached".
  • VQA or LLM-judge probing. Ask a vision-language model yes/no questions derived from the caption. Treat this as another noisy signal and validate it against human labels, as covered in using an LLM judge to score training data.

A human-audited rubric as the acceptance metric

The acceptance metric should be the aligned rate on a human-audited, stratified sample, because every automated score inherits its model's blind spots. Draw the sample across CLIP-score deciles and across your business slices, so low-score expert captions and high-score generic ones are both reviewed. Use two raters on an overlap subset and report agreement (Cohen's kappa) so the rate is defensible.

Illustrative example: invented to show structure; it does not describe an available dataset.

GradeDefinitionExample (field service photo)Counts toward
Fully alignedAll visual claims verifiable; no wrong facts"Corroded flange on 2-inch pipe, rust at bolt heads"Aligned rate
Partly alignedCorrect subject, missing or minor wrong detail"Pipe flange" (omits corrosion)Partial rate
Context onlyText describes the job, not the pixels"Customer reported leak since March"Tagged, not penalized
Not alignedWrong image, boilerplate or hallucinated content"See attached" or a caption from another jobMismatch rate

Set thresholds per use before you see the sample. A pre-training pool may tolerate a higher partial rate than a private multimodal evaluation set, where any mismatch corrupts the score. Agree in writing what happens when a delivery misses the bar, for example replacement or re-filtering, during a pilot; the owner guide on running a data pilot with a supplier covers how to structure that.

Recaptioning: higher alignment, new risks

Machine recaptioning usually raises measured alignment, but it replaces source text with model output and can erase the expert vocabulary you were buying. A captioning model describes what it recognizes, so "hydraulic manifold with a weeping O-ring" becomes "metal machine part with a black ring". It also imports the captioner's phrasing habits, which can make your training distribution echo another model.

If you recaption, keep the original text in a separate field, record the captioning model and prompt, and score both versions on the same human sample. A useful pattern is to generate a visual description and keep the original as context, giving the model both grounding and domain language. Note that a CLIP-based score will tend to favor captions written in CLIP-like web language, so do not use it alone to compare original and recaptioned text.

Why filtering comes before purchase sizing

Curated subsets can beat larger noisy sets, so measure alignment before you decide how many pairs to buy. For multimodal instruction tuning, MM-LIMA reports that a small quality-selected subset of instruction data outperformed fine-tuning on the full set [1]. Filtering research on web-scale pre-training pools points the same way: removing poorly aligned pairs often helps more than adding raw volume.

The procurement consequence is direct: price and volume discussions should be based on the expected aligned yield, not raw pair count. Ask for a representative sample, run your pipeline on it, and size the request against the post-filter number; when you describe a need to SourceX buyers, state the aligned yield and rubric you will accept. Broader methods for this are in training data quality metrics and the quality hub.

Reporting alignment per slice in the dataset card

Report alignment metrics per slice, because a single aggregate hides the slices where your model will fail. Slice by source system, image type (photo, screenshot, scan, diagram), caption origin (human, template, recaptioned), domain and caption length. Datasheets for Datasets argues that documenting composition, collection and preprocessing lets consumers judge fitness for use [3]; alignment results belong in that record.

Illustrative example: invented to show structure; it does not describe an available dataset.

alignment_qa:
  definition: "visual claims verifiable in image; context text tagged"
  automated:
    embedding_model: "ViT-L/14"
    score_field: clip_cosine
    ocr_check: anls
  human_audit:
    sample_size: 600
    sampling: "stratified by clip decile x image_type"
    raters: 2
    cohen_kappa: 0.71
  results_by_slice:
    - slice: "image_type=photo, origin=human"
      fully: 0.68
      partly: 0.17
      context_only: 0.09
      not_aligned: 0.06
    - slice: "image_type=screenshot, origin=template"
      fully: 0.41
      partly: 0.22
      context_only: 0.04
      not_aligned: 0.33
  recaptioning: {applied: false}

If an AI system built on the data falls into a high-risk category under the EU AI Act, Article 10 requires its training, validation and testing data sets to meet quality criteria under data governance practices [4]. As of October 2026, the Annex III high-risk dates were reportedly moved to 2 December 2027 by Regulation (EU) 2026/1744, which also amends Article 10 [5]. A per-slice alignment record is one concrete piece of that documentation.

For related planning, see the multimodal data hub, the multimodal data request specification template and the owner guide on evaluating data supplier quality. Buyers comparing open options can also review image and caption datasets for commercial use or start at the AI data hub.

Request image-text data with alignment checks

SourceX sources operational datasets, including documents and new recordings of hands-on work, from US companies on request; a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with diligence materials on source and preparation prepared per dataset. Describe the image-text data and alignment bar you need at SourceX for buyers.

Frequently asked questions

What CLIP score threshold should I use?

There is no portable threshold. Scores shift with the embedding model and with the domain, so a cutoff from a web-scale pipeline does not transfer to engineering photos or scanned forms. Pick a cutoff by plotting your human-audited aligned rate against score deciles on your own data.

Should I discard captions that describe things not in the image?

No, if they are accurate context. Tag them context, exclude them from the mismatch rate, and decide per use whether to train on them. They often carry the domain knowledge that made the data worth licensing.

Can an LLM or VLM judge replace human audit?

Not as the acceptance metric. Model judges share blind spots with the encoders you are checking, so validate any judge against human labels first.

Sources

  1. arXiv, "MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets" (2023). https://arxiv.org/pdf/2308.12067
  2. Biten et al., arXiv, "ICDAR 2019 Competition on Scene Text Visual Question Answering" (2019). https://arxiv.org/pdf/1907.00490
  3. Gebru et al., arXiv, "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
  4. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  5. European Parliament and Council of the European Union (EUR-Lex), "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data