Skip to content

Multimodal and embodied data

Multimodal Meeting Recordings for AI: Video, Audio, Transcripts and Shared Screens

Quick answer

A usable multimodal meeting dataset is a set of synchronized layers, not a pile of MP4 files: per-speaker or mixed audio, video tiles, the screen-share stream, a speaker-attributed transcript, chat, and the artifacts the meeting produced, such as notes, action items and tickets. Buyers should specify each layer, the clock that aligns them, how overlapping speech is scored, and proof that every participant consented to both face and voice use. Shared screens need their own redaction plan.

By SourceX Editorial · Updated

Which layers a meeting assistant actually needs

Specify layers by task, because a summarization model and an active-speaker detector need different streams. Research corpora such as the AMI Meeting Corpus were captured in instrumented rooms with headset and far-field microphones, room cameras and slide capture on a common timeline. Commercial recordings from Zoom, Microsoft Teams or Google Meet rarely look like that: most exports give one mixed audio track, a composited "speaker view" or gallery video, and a VTT or JSON transcript from the platform's own recognizer.

That gap matters for training. A composited gallery view loses per-tile identity, so active-speaker detection needs either separate per-participant video or a tile layout log with timestamps. Mixed audio makes diarization harder than per-participant tracks, which some platforms record as an option and many organizations never enable.

Illustrative example: invented to show structure; it does not describe an available dataset.

LayerTypical source formatWhat to requireCommon failure mode
AudioM4A/WAV, mixed or per-participantSample rate, channel map, per-speaker tracks where availableOnly mixed audio; echo-cancelled artifacts on overlap
VideoMP4 (H.264), gallery or speaker viewPer-tile streams or a tile layout log with participant IDsComposited view with no mapping from tile to speaker
Screen shareSeparate MP4 stream or inset in compositeSeparate stream, share start and stop events, resolutionShare burned into the composite at low resolution
TranscriptVTT, SRT or vendor JSONWord-level timestamps, speaker labels, ASR vs human-corrected flagPlatform ASR presented as ground truth
Diarization referenceRTTMOne turn per line, overlap policy documentedTurns snapped to sentence boundaries, hiding overlap
Chat and reactionsJSON or TXT exportTimestamps on the meeting clock, sender IDsChat exported without timestamps
OutcomesNotes, action items, tickets, CRM updatesLink to meeting ID, author, created timeNotes written days later and never linked back

Alignment and diarization: get the clocks and the overlap rule in writing

The first thing to verify is that every layer shares one clock with a documented offset. Ask for a manifest per meeting with the start time of each stream, its frame rate or sample rate, and any known drift; the time synchronization checklist covers how to spot-check this on delivered samples.

For speaker turns, request RTTM references. Audio-visual diarization challenges such as MISP 2022 score systems against RTTM files that list one speaker turn per line in fixed columns, using diarization error rate (DER) [1]. Reference turns are often derived from close-talk channels rather than the far-field mix [2], which is exactly what per-participant meeting audio can supply. Decide up front whether overlapping speech is scored or excluded and whether a collar is applied, because both choices move DER substantially and a supplier's labels must match your scoring.

Transcript provenance is the other trap. A platform transcript is a model output; if you train on it as ground truth, you inherit its errors on names, jargon and accented speech. Require a field marking which segments were human-corrected and a sample-level word error rate against a corrected subset.

Screen shares, slides and the third parties on screen

Treat the screen-share stream as the highest-risk layer, because it shows content that belongs to people who never joined the call. A sales demo shares a customer's CRM record, an engineering review shares an incident dashboard, a finance meeting shares a spreadsheet of vendor payments. Those frames contain names, emails, account numbers and confidential terms that a transcript-only review never sees.

Buyers should require one of three treatments per meeting: exclude the share stream, redact it with OCR-driven box masking verified on a sample of frames, or keep it only where the shared content is the supplier's own non-personal material such as slide decks. The multimodal de-identification guide and the page on anonymizing video for AI training cover masking methods and residual-risk checks. Slides that are kept are valuable on their own for multimodal RAG evaluation, since questions often hinge on a chart that was never spoken aloud.

Faces and voices in meeting recordings can be biometric data, so consent has to cover every participant, including external guests, for both face and voice and for AI training specifically. Illinois BIPA defines biometric identifiers to include voiceprints and scans of face geometry and requires a written release before collection [3]. Texas requires notice and consent before capturing a voiceprint or face-geometry record for a commercial purpose [6], and Washington's definition centers on data used to identify a specific individual [7].

This is not theoretical for meeting data. As of October 2026, class actions filed in May 2026 in the Northern District of Illinois allege that voiceprints used to train AI fall under BIPA, with no merits ruling yet [4], and AI meeting transcription tools have already drawn BIPA claims [5]. Recording consent is a separate question from biometric consent; the wiretap law overview and the Illinois BIPA guide explain both.

One more check: if recordings pass through a meeting or transcription vendor, confirm the vendor's terms. FTC staff have warned, in a 2024 Office of Technology post, that model-as-a-service companies can face liability for breaking promises not to use customer data for undisclosed purposes such as training [8]. Rights to a recording held in a vendor's cloud usually start with the company that held the meeting, but the vendor's terms can still restrict export or reuse.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Pairing recordings with outcomes for evaluation

The most useful meeting data links what was said to what happened afterward. An action item in the notes, a Jira ticket created ten minutes after the call, or a CRM stage change gives a summarization or assistant model a verifiable target, which is much stronger than a human-written summary alone.

Ask suppliers which downstream systems can be joined by meeting ID or calendar event ID, and how the join was made. Hold back a slice of fully linked meetings as a private multimodal evaluation set so that leaderboard contamination cannot inflate your results. Document the join logic and the annotation methods in a datasheet in the style of Data Cards [9].

A request template for meeting recordings

Write the request as a specification that a supplier's IT and legal teams can answer line by line. The multimodal specification template gives the full structure; for meetings, add these fields.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Meeting types: internal standups, design reviews, customer calls; language and region.
  • Platforms and export path: Zoom, Teams or Meet; cloud recording vs local; available tracks.
  • Layers required: per-speaker audio (yes/no), per-tile video or layout log, screen-share stream, chat, transcript with word timestamps.
  • Annotation: RTTM turns, overlap policy, collar, human-corrected transcript share.
  • Outcomes: notes, action items, tickets or CRM events linked by meeting ID.
  • Privacy: consent records covering all participants and guests for face and voice; screen-share treatment; PII replacement method.
  • Packaging: per-meeting manifest; see video dataset packaging and speech manifest packaging.

Transcript-only needs belong on the meeting transcripts page, and broader video needs on licensed video recordings.

How SourceX handles meeting recording requests

SourceX sources operational datasets from US companies on request; it holds no stock, and a request does not guarantee a match. Every release is approved by the supplying company, each dataset is rights-reviewed for ownership and consents, and personal details such as names, emails and phone numbers are removed or replaced before delivery, with the method recorded and a sample checked. No method is perfect. Buyers can describe the meeting data they need, and work moves through Find, Assess, Agree, Transact and Manage; nothing is contracted until a supplier agrees. For broader context, see the multimodal data hub and the AI data guides.

Request licensed meeting recordings for AI

SourceX looks for US businesses that hold the meeting data you describe, reviews rights and consents, and delivers under a license that defines the records, uses, term and delivery. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe your multimodal meeting dataset requirements.

Frequently asked questions

Is the AMI corpus enough to train a commercial meeting assistant?

It is a strong research benchmark, but its meetings were recorded in instrumented rooms rather than on video-conferencing platforms, and its license terms should be checked at the official source before any commercial use. Commercial assistants usually need platform-native recordings that match production audio and layouts.

Do we need per-speaker audio tracks?

For diarization and active-speaker training, yes where possible, because close-talk channels make cleaner RTTM references [2]. For summarization alone, mixed audio with a good transcript is often sufficient.

Can screen shares be kept if faces are blurred?

Blurring faces does nothing for the shared screen. Screen content needs its own OCR-based redaction or exclusion, verified on sampled frames.

Sources

  1. arXiv, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
  2. arXiv, "Summary of the DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments" (2023). https://arxiv.org/pdf/2311.12564
  3. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  4. Biometric Update, "Tech giants sued under BIPA over voiceprints used to train AI" (2026). https://www.biometricupdate.com/202605/tech-giants-sued-under-bipa-over-voiceprints-used-to-train-ai
  5. Lewis Rice, "AI Transcription Tools Give Rise to BIPA Claims". https://www.lewisrice.com/publications/ai-transcription-tools-give-rise-to-bipa-claims
  6. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  7. Washington State Legislature, "RCW 19.375 Biometric identifiers" (2024). https://lawfilesext.leg.wa.gov/law/RCWArchive/2024/RCW%20%2019%20.375%20%20CHAPTER.htm
  8. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  9. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data